Document content extraction method and system and storage medium

By performing invalid characters and null characters analysis in sectors of damaged DOC documents, combined with character type statistics, the document content is extracted efficiently and accurately, solving the problem of insufficient extraction efficiency and accuracy in the prior art.

CN120354845APending Publication Date: 2025-07-22CHENGDU YIWO TECH DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510498948.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art cannot efficiently and accurately extract document content from damaged DOC documents, especially in the case of damaged file structures, and the accuracy and efficiency of conventional scanning analysis are low.

Method used

By selecting the current sector from the sectors of the damaged DOC document, determining the sector data according to the preset character encoding method, and performing invalid character detection and null character analysis, filtering out alternative characters, combining character type statistics to determine whether the preset conditions meet, and setting the characters that meet the conditions to the document content.

Benefits of technology

Improve the efficiency and accuracy of extracting damaged DOC documents, avoid interference from noisy data, and ensure the correctness of extracted content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354845A_ABST
    Figure CN120354845A_ABST
Patent Text Reader

Abstract

The invention discloses a document content extraction method and system and a storage medium, and belongs to the technical field of data processing. The document content extraction method comprises the following steps: selecting a current sector from a damaged DOC document; determining sector data of the current sector according to a preset character coding mode, and performing invalid character detection on the current sector; if the invalid characters do not exist in the current sector, whether null characters exist in the current sector or not is judged; if the null characters do not exist in the current sector, all the characters in the current sector are set as alternative characters; if null characters exist in the current sector and characters behind any null character are null characters, all characters before the null character at the most front position in the current sector are set as alternative characters; judging whether the character type statistical result meets a preset condition or not; and if yes, setting the alternative characters as the document content. According to the method and the device, the document content of the damaged DOC document can be efficiently and accurately extracted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly relates to a method, a system and a storage medium for extracting document content. Background Art

[0002] Currently, common versions of DOC documents include multiple stream objects, and each stream object contains various data organized in a specific structure for data positioning and storing specific text, pictures, embedded objects and other information. After the above DOC document is damaged, the document content cannot be read through conventional means; in related technologies, the binary data of the damaged DOC document is usually scanned and analyzed to determine the document content, but the accuracy and processing efficiency of the above method are relatively low.

[0003] Therefore, how to efficiently and accurately extract the document content of a damaged DOC document is a technical problem that those skilled in the art need to solve currently. Summary of the Invention

[0004] The purpose of this application is to provide a method, a system and a storage medium for extracting document content, which can efficiently and accurately extract the document content of a damaged DOC document.

[0005] To solve the above technical problem, this application provides a method for extracting document content, including:

[0006] Select a current sector from the sectors corresponding to the damaged DOC document;

[0007] Determine the sector data of the current sector according to a preset character encoding method, and perform invalid character detection on the current sector based on the sector data;

[0008] Perform character analysis on the current sector based on the sector data to obtain alternative characters; wherein, the process of character analysis includes: if there are no invalid characters in the current sector, determine whether there are null characters in the current sector; if there are no null characters in the current sector, set all characters in the current sector as alternative characters; if there are null characters in the current sector and all characters after any null character are null characters, set all characters before the null character with the earliest position in the current sector as alternative characters;

[0009] Judge whether the statistical result of the character types of all alternative characters in the current sector meets a preset condition;

[0010] If so, set all alternative characters in the current sector as the document content of the damaged DOC document.

[0011] Optionally, it further includes:

[0012] If there are invalid characters in the current sector, it is determined that all characters in the current sector are not the document content of the damaged DOC document;

[0013] If there are null characters in the current sector and the characters after any null character are not all null characters, it is determined that all characters in the current sector are not the document content of the damaged DOC document.

[0014] Optionally, determine the sector data of the current sector according to a preset character encoding method, and perform invalid character detection on the current sector based on the sector data, including:

[0015] Determine the first sector data of the current sector in the way of single-byte encoding;

[0016] Use a compressed character quick reference table to perform invalid character detection on each character in the first sector data; wherein, the compressed character quick reference table stores the corresponding relationship between compressed characters and character states, and the character states include valid and invalid, and invalid characters are characters with an invalid character state;

[0017] Correspondingly, perform character analysis on the current sector based on the sector data to obtain alternative characters, including:

[0018] Perform character analysis on the current sector based on the first sector data to obtain the first alternative characters;

[0019] Correspondingly, determine whether the character type statistical results of all alternative characters in the current sector meet the preset conditions, including:

[0020] Determine whether the character type statistical results of all the first alternative characters in the current sector meet the first preset condition;

[0021] Correspondingly, set all alternative characters in the current sector as the document content of the damaged DOC document, including:

[0022] Set all the first alternative characters in the current sector as the document content of the damaged DOC document.

[0023] Optionally, determine whether the character type statistical results of all the first alternative characters in the current sector meet the first preset condition, including:

[0024] Perform character type statistics on all the first alternative characters in the current sector;

[0025] Determine whether the proportion of common characters in the character type statistical results is greater than the first threshold; wherein, the common characters include any one or a combination of several of letters, numbers, and spaces;

[0026] If so, it is determined that the character type statistical results meet the first preset condition;

[0027] If not, it is determined that the character type statistical result does not meet the first preset condition.

[0028] Optionally, after determining whether the character type statistical results of all the first alternative characters in the current sector meet the first preset condition, it further includes:

[0029] If the character type statistical results of all the first alternative characters in the current sector do not meet the first preset condition, the second sector data of the current sector is determined in the manner of double-byte encoded data, and invalid character detection is performed on the current sector based on the second sector data;

[0030] Character analysis is performed on the current sector based on the second sector data to obtain second alternative characters;

[0031] Determine whether the character type statistical results of all the second alternative characters in the current sector meet the second preset condition;

[0032] If so, all the second alternative characters in the current sector are set as the document content of the damaged DOC document;

[0033] If not, it is determined that all the characters in the current sector are not the document content of the damaged DOC document.

[0034] Optionally, determining the sector data of the current sector in accordance with a preset character encoding method, and performing invalid character detection on the current sector based on the sector data includes:

[0035] Determine the second sector data of the current sector in the manner of double-byte encoded data;

[0036] Perform invalid character detection on each character in the second sector data by using a non-compressed character quick reference table; wherein, the non-compressed character quick reference table stores the correspondence between non-compressed characters and character states, the character states include valid and invalid, and the invalid characters are the characters with the character state of invalid;

[0037] Correspondingly, performing character analysis on the current sector based on the sector data to obtain alternative characters includes:

[0038] Perform character analysis on the current sector based on the second sector data to obtain second alternative characters;

[0039] Correspondingly, determining whether the character type statistical results of all the alternative characters in the current sector meet the preset condition includes:

[0040] Determine whether the character type statistical results of all the second alternative characters in the current sector meet the second preset condition;

[0041] Correspondingly, setting all alternative characters in the current sector as the document content of the damaged DOC document includes:

[0042] Setting all second alternative characters in the current sector as the document content of the damaged DOC document.

[0043] Optionally, the non-compressed character quick reference table also stores the correspondence between characters and language types;

[0044] Correspondingly, determining whether the character type statistical result of all second alternative characters in the current sector meets the second preset condition includes:

[0045] Performing character type statistics on all second alternative characters in the current sector based on the non-compressed character quick reference table;

[0046] Determining whether the number of language types in the character type statistical result is less than the second threshold and the preset character proportion is within the preset range;

[0047] If so, determining that the character type statistical result meets the second preset condition;

[0048] If not, determining that the character type statistical result does not meet the second preset condition.

[0049] Optionally, after determining whether the character type statistical result of all second alternative characters in the current sector meets the second preset condition, it further includes:

[0050] If the character type statistical result of all second alternative characters in the current sector does not meet the second preset condition, determining the first sector data of the current sector in the manner of single-byte encoded data, and performing invalid character detection on the current sector based on the first sector data;

[0051] Performing character analysis on the current sector based on the first sector data to obtain first alternative characters;

[0052] Determining whether the character type statistical result of all first alternative characters in the current sector meets the first preset condition;

[0053] If so, setting all first alternative characters in the current sector as the document content of the damaged DOC document;

[0054] If not, determining that all characters in the current sector are not the document content of the damaged DOC document.

[0055] This application also provides a document content extraction system, which includes:

[0056] A sector selection module for selecting the current sector from the sectors corresponding to the damaged DOC document;

[0057] A data determination module, configured to determine sector data of a current sector according to a preset character encoding method, and perform invalid character detection on the current sector based on the sector data;

[0058] A character analysis module, configured to perform character analysis on the current sector based on the sector data to obtain candidate characters; wherein, the process of the character analysis includes: if there is no invalid character in the current sector, determining whether there is a null character in the current sector; if there is no null character in the current sector, setting all characters in the current sector as candidate characters; if there is a null character in the current sector and all characters after any null character are null characters, setting all characters before the null character with the earliest position in the current sector as candidate characters;

[0059] A decision module, configured to determine whether a character type statistical result of all candidate characters in the current sector meets a preset condition; if so, setting all candidate characters in the current sector as the document content of the damaged DOC document.

[0060] The present application further provides a storage medium, on which a computer program is stored, and when the computer program is executed, the steps performed by the above-mentioned document content extraction method are implemented.

[0061] The present application provides a document content extraction method. The present application selects a current sector from sectors corresponding to a damaged DOC document, determines sector data of the current sector according to a preset character encoding method, and then analyzes characters in the current sector based on the sector data. In the process of character analysis, the present application accurately filters out candidate characters by determining whether there are invalid characters and null characters, and the situation of characters after the null characters. After obtaining the candidate characters, the present application also performs character type statistics on the candidate characters, and sets all candidate characters in the current sector as the document content of the damaged DOC document when the character type statistical result meets a preset condition. In the above process, it is not necessary to perform byte-by-byte scanning on the binary data of the DOC document, but to judge the validity of characters by analyzing invalid characters and null characters, which improves the detection efficiency and accuracy. In addition, after determining the candidate characters, the present application judges the current sector based on the character type statistical result, which can avoid the interference of noise data and ensure the correctness of the extracted content. Therefore, the present application can efficiently and accurately extract the document content of the damaged DOC document. The present application also provides a document content extraction system and a storage medium at the same time, which have the above-mentioned beneficial effects and will not be elaborated here. Description of the Drawings

[0062] To more clearly illustrate the embodiments of the present application, the accompanying drawings required for the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.

[0063] Figure 1 It is a flowchart of a method for extracting document content provided by an embodiment of the present application;

[0064] Figure 2 It is a flowchart of a method for extracting document content of a DOC document provided by an embodiment of the present application;

[0065] Figure 3 It is a schematic diagram of a detection process for a compressed main document provided by an embodiment of the present application;

[0066] Figure 4 It is a schematic diagram of a detection process for an uncompressed main document provided by an embodiment of the present application. Detailed implementation manners

[0067] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.

[0068] In this solution, the damaged DOC document is a Word document before the Office2007 version (a version of an office software), also known as an MS-DOC document, which is constructed based on a file format called MS-CFB (Microsoft -Compound File Binary File Format). The DOC document using the MS-CFB format is a structured storage compound file that organizes various objects together by implementing a simplified file system inside the file. In the MS-CFB format, objects are divided into storage objects and stream objects, similar to directories and files in a file system.

[0069] Office documents prior to Office 2007, whether they are Word (a word processing software), Excel (a spreadsheet software), or PowerPoint (a presentation software) documents, are all built on the basis of MS-CFB, only with different storage objects and stream objects. For MS-DOC documents, they contain key stream objects such as WordDocument Stream (document stream), Table Stream (table stream), and Data Stream (data stream).

[0070] In the field of data recovery and file repair, repairing MS-DOC documents is a relatively difficult task. The difficulties mainly come from the complexity of the MS-DOC document format, which includes both the complexity of the underlying file structure MS-CFB, the complexity of the unique stream objects and data structures of MS-DOC, as well as the association and dependency relationships between various structures. Repairing MS-DOC documents involves the repair of various contents, including text, pictures, tables, attributes, and embedded objects. And the document content (i.e., text content) is a very important part of MS-DOC documents, usually containing a large amount of useful information.

[0071] To read the document content from an undamaged MS-DOC document, the following steps are required: Parse the file according to the MS-CFB structure to obtain the file header information, directory entry information, and information about the WordDocumentStream and Table Stream of the MS-CFB file; Read the FIB information from the WordDocument Stream to obtain the offset position information of Clx in the TableStream; Read the Clx (a structure describing the offset) information from the specified offset position in the Table Stream, and parse Clx to obtain the position distribution of the main document in the WordDocument Stream; According to the position distribution information of the main document, read the main document content from the WordDocument Stream (it may be necessary to read at different positions because the main document may be split into many fragments); If only the document content is needed, special characters can be removed from the read main document content.

[0072] FIB, that is, File Information Block, is stored at the beginning of the WordDocument Stream. The FIB structure contains relevant information about the document and specifies the positions of all other data in the file.

[0073] Clx is stored in the Table Stream, and the offset position of Clx in the Table Stream is specified by the fcClx field in the FIB. By parsing the Clx structure, the position of the main document in the WordDocument Stream can be located.

[0074] The Main Document is the main document, and all document contents are included in the main document. In addition, the main document also contains some special characters (such as paragraph marker characters for indicating paragraph breaks, line break characters for indicating line breaks, and anchor characters for indicating pictures, etc.). The main document may be split into many parts in the WordDocument Stream, that is, the data of the main document is not necessarily continuous in the stream. Its position information is described by the Clx structure in the Table Stream. In this embodiment, the document content read from the damaged document is the content of the above-mentioned main document Main Document.

[0075] However, for an MS-DOC document with a damaged file structure, even if the content of the main document in it is not damaged (or only partially damaged), but other related data structures are damaged, then it will be very difficult or even impossible to retrieve the document content according to the general process.

[0076] In view of the technical problems existing in the above related technologies, this embodiment provides a new solution to extract the document content from the damaged DOC document. Please refer to the following Figure 1 , Figure 1 which is a flowchart of a method for extracting document content provided by an embodiment of the present application. The specific steps may include:

[0077] S101: Select the current sector from the sectors corresponding to the damaged DOC document;

[0078] Among them, this embodiment can be applied to an electronic device with data processing capabilities. Before this step, the operation of the damaged DOC document can be determined. The above-mentioned damaged DOC document refers to a DOC document that cannot be normally opened or read due to reasons such as storage medium damage, file transmission error, software failure, or virus infection, resulting in partial loss or damage of the file structure or data.

[0079] The above damaged DOC document is a document based on the MS-CFB file format, and MS-CFB implements a simplified file system. The first 512 bytes of the damaged DOC document are the file header, and the Sector Shift field in it indicates the sector size of the file. Here, the sector does not refer to the sector of the disk, but the sector concept of the MS-CFB file format. The damaged DOC document manages data in units of sectors. In this embodiment, an MS-CFB file can be regarded as a disk partition that implements a file system. A sector size may be 512 or 4096 bytes. The sector starting from the file offset position 0 is used to store the file header of the damaged DOC document. In addition, the remaining content of the damaged DOC document is numbered starting from 0 according to the sector size, which are sector 0, sector 1, sector 2... until the end of the document. Therefore, all the content of the damaged DOC document corresponds to multiple sectors. In this step, the current sector can be selected from all the sectors corresponding to the damaged DOC document, and then the document content extraction operations of S102-S105 are performed on the current sector. After the extraction of the current sector is completed, the next sector can be selected to repeat the operations of S101-S105 until all the sectors corresponding to the damaged DOC document are processed. As a feasible implementation manner, in this embodiment, the sectors corresponding to the damaged DOC document can be sequentially selected as the current sector in ascending order of sector numbers.

[0080] S102: Determine the sector data of the current sector according to a preset character encoding method, and perform invalid character detection on the current sector based on the sector data;

[0081] Among them, there may be multiple bytes in the current sector, such as 512 bytes or 4096 bytes. In this step, the bytes in the current sector can be processed according to a preset character encoding method to obtain the corresponding sector data.

[0082] The preset character encoding method stipulates the mapping relationship between characters and byte sequences, such as 1 character occupying 1 byte, 1 character occupying 2 bytes, etc. Specifically, the above preset character encoding method may include single-byte encoding, double-byte encoding and other methods. In single-byte encoding, 1 character occupies 1 byte; in double-byte encoding, 1 character occupies 2 bytes.

[0083] The sector data obtained by different character encoding methods is also different. After using the preset character encoding method to determine the sector data of the current sector, the invalid character detection and the character analysis operation of S103 can be performed using the sector data corresponding to the preset character encoding method.

[0084] This step can perform an invalid character detection on the current sector based on the sector data to determine the invalid characters in the current sector under a preset character encoding method. This step can preset a table storing invalid characters and determine the invalid characters in the current sector by looking up the table.

[0085] This step can detect whether the current sector contains invalid characters without determining which characters in the current sector belong to the invalid characters. Specifically, after the first detection of an invalid character, the detection of the remaining characters can be stopped and the operation process of this embodiment can be ended.

[0086] S103: Perform character analysis on the current sector based on the sector data to obtain candidate characters;

[0087] This embodiment can perform invalid character detection and character analysis on all the encoding results of the current sector, or can perform invalid character detection and character analysis on partial encoding results of the current sector. Specifically, this embodiment can process all the bytes of the current sector according to a preset character encoding method, that is, the sector data is the encoding result of all the bytes of the current sector. This embodiment can also process partial bytes of the current sector according to a preset character encoding method, and perform the character analysis operation of S103 once for each obtained character. Therefore, the sector data obtained in this step can be a single character encoded from the current sector through a preset character encoding method.

[0088] This step obtains candidate characters by performing character analysis on the sector data. The process of the above character analysis includes the following strategies:

[0089] If there are invalid characters in the current sector, it is determined that all the characters in the current sector are not the document content of the damaged DOC document, that is, there are no candidate characters.

[0090] If there are no invalid characters in the current sector, it is judged whether there are null characters in the current sector.

[0091] If there are no null characters in the current sector, all the characters in the current sector are set as candidate characters.

[0092] If there are null characters in the current sector and all the characters after any null character are null characters, all the characters before the null character with the earliest position in the current sector are set as candidate characters; here, it is equivalent to setting all the characters other than the null characters in the current sector as candidate characters.

[0093] If there are null characters in the current sector and not all the characters after any null character are null characters, it is determined that all the characters in the current sector are not the document content of the damaged DOC document, that is, there are no candidate characters.

[0094] The above null character can be a pre-specified character, such as 0x00 or 0x0000.

[0095] Specifically, the process of detecting the invalid characters and analyzing the characters to determine the alternative characters is as follows in steps A1 - A7:

[0096] Step A1: Detect invalid characters for the current sector based on the sector data.

[0097] Step A2: Determine whether there are invalid characters in the current sector; if so, go to step A6; if not, go to step A3.

[0098] Step A3: Determine whether there are null characters in the current sector; if so, go to A4; if not, go to step A7.

[0099] Step A4: Determine whether all the characters after any null character are also null characters; if so, go to step A5; if not, go to step A6.

[0100] Step A5: Set all the characters before the earliest null character in the current sector as the alternative characters.

[0101] Step A6: Determine that all the characters in the current sector are not the document content of the damaged DOC document, that is, there are no alternative characters.

[0102] Step A7: Set all the characters in the current sector as the alternative characters.

[0103] S104: Determine whether the statistical result of the character types of all the alternative characters in the current sector meets the preset conditions; if so, go to S105; if not, end the process.

[0104] Among them, the document content of the DOC document contains multiple characters, and the statistical result of the character types meets certain characteristics; for example, the proportion of common characters such as letters, numbers, and spaces in the document content is greater than 60%, the number of language types of all the characters in the document content is less than or equal to 3, and the proportion of special characters is less than 20%. The above special characters can be pre-specified characters used to implement specific functions or identify specific information, such as the paragraph marker character for indicating paragraph breaks, the line break character for indicating line breaks, and the anchor character for indicating pictures, etc. The statistical result of the character types can be statistical information about common characters, special characters, the language types to which the characters belong, etc.

[0105] Before this step, character type statistics can be performed on all candidate characters in the current sector to obtain a character type statistics result, and then the character type statistics results can be compared. If the preset conditions are met, all candidate characters in the current sector are set as the document content of the damaged DOC document; if the preset conditions are not met, the operation process of this embodiment can be ended, or a new preset character encoding method can be determined and the operation of S102 can be executed again.

[0106] S105: Set all candidate characters in the current sector as the document content of the damaged DOC document.

[0107] Among them, after setting all candidate characters in the current sector as the document content of the damaged DOC document, it can be determined whether all sectors corresponding to the damaged DOC document have been detected; if so, all document contents of the damaged DOC document are generated according to all the determined document contents; if not, the next sector is selected and the operations of S101 - S105 are repeated. The document content of the damaged DOC document described in this embodiment refers to the content of the main document of the damaged DOC document.

[0108] In this embodiment, the current sector is selected from the sectors corresponding to the damaged DOC document, and the sector data of the current sector is determined according to the preset character encoding method, and then the characters in the current sector are analyzed based on the sector data. During the character analysis process, this embodiment accurately filters out candidate characters by determining whether there are invalid characters and null characters, and the characters following the null characters. After obtaining the candidate characters, this embodiment also performs character type statistics on the candidate characters, and when the character type statistics result meets the preset conditions, all candidate characters in the current sector are set as the document content of the damaged DOC document. In the above process, it is not necessary to perform a byte-by-byte scan of the binary data of the DOC document, but to judge the validity of the characters by analyzing the invalid characters and null characters, which improves the detection efficiency and accuracy. In addition, after determining the candidate characters, this embodiment judges the current sector based on the character type statistics result, which can avoid the interference of noise data and ensure the correctness of the extracted content. Therefore, this embodiment can efficiently and accurately extract the document content of the damaged DOC document.

[0109] As a further introduction to Figure 1 the corresponding embodiment, the content stored in the sector corresponding to the DOC document can be a compressed document or a non-compressed document; the same DOC document can include sectors storing compressed documents and sectors storing non-compressed documents at the same time. The character encoding method corresponding to the compressed document is single-byte encoding, and the character encoding method corresponding to the non-compressed document is double-byte encoding.

[0110] The "Retrieving Text" section of Microsoft's MS-DOC technical documentation describes the method of retrieving document content from MS-DOC. From this method, it can be seen that the document content is divided into compressed and uncompressed types. In the compressed mode, the text consists of ANSI (a character code) characters, with each character occupying one byte. In the uncompressed mode, the text consists of Unicode (Universal Character Set) characters, with each character occupying two bytes. The compression flag is specified in the Clx structure. An MS-DOC document may contain multiple segments of document content, and the compression flag for each segment may be different. In addition to the document content, the main document also contains some special characters, and these special characters also follow the compression flag rules. That is to say, depending on the compression flag, special characters may occupy one or two bytes. Common special characters are: 0x01 (which is 0x0001 in the uncompressed mode and occupies two bytes, the same applies below) represents the placeholder for an embedded picture, 0x07 represents the paragraph marker character, 0x0B represents the line break character, 0x0C represents the section marker character, and so on.

[0111] Based on Microsoft's documentation, the inferences from the above content, and verification with a large number of real MS-DOC documents, the main document in MS-DOC has the following storage characteristics in the file:

[0112] Characteristic 1: Each segment of the main document must be stored starting from a sector-aligned position. That is to say, the position offset of the start of each segment of the main document in the file must be an integer multiple of the sector size.

[0113] Characteristic 2: For document content that can be encoded in a single byte, such as Latin languages, its storage method in each segment of the main document may be either compressed or uncompressed. The encoding method for compressed storage is usually ISO8859-1. The encoding method for uncompressed storage is Unicode.

[0114] Characteristic 3: For document content that is not suitable for single-byte encoding, such as East Asian languages, its storage method in each segment of the main document is uncompressed, and the encoding method is Unicode.

[0115] Characteristic 4: The characters in the main document do not contain 0x00 (in the compressed storage mode) or 0x0000 (in the uncompressed storage mode), unless it has reached the end of a text segment. Then, the remaining data in the sector where the end position of the text segment is located will be filled with 0x00.

[0116] Feature 5: Some characters cannot appear in the main document. In this article, such characters are called invalid characters. For example, 0x02 (the non-compressed form is 0x0002, the same below), 0x05, 0x0A, 0x0E, etc. Since these characters are neither visible characters nor special characters defined by MS-DOC, they should not appear in the main document.

[0117] Feature 6: For a section of the main document content, the characters it contains tend to follow a certain probability distribution. For example, the proportion of language characters (letters or words) and punctuation marks in all characters of the main document should not be too low, while the proportion of special characters in all characters of the main document should not be too high. It is also unlikely that the characters of different languages contained in a section of document content are too messy (such as more than three languages).

[0118] Specifically, this embodiment can determine the document content of the damaged DOC document in the following several ways:

[0119] Method 1: First, determine the sector data of the current sector in the single-byte encoding method, and perform invalid character detection and character analysis to determine the first alternative character; if the character type statistical result of the first alternative character meets the preset conditions, set the first alternative character as the document content of the damaged DOC document. If the character type statistical result of the alternative character does not meet the preset conditions, determine the sector data of the current sector in the double-byte encoding method, and perform invalid character detection and character analysis to determine the second alternative character; if the character type statistical result of the second alternative character meets the preset conditions, set the second alternative character as the document content of the damaged DOC document; if the character type statistical result of the second alternative character does not meet the preset conditions, determine that the current sector does not store the document content of the damaged DOC document.

[0120] Method 2: First, determine the sector data of the current sector in the double-byte encoding method, and perform invalid character detection and character analysis to determine the second alternative character; if the character type statistical result of the second alternative character meets the preset conditions, set the second alternative character as the document content of the damaged DOC document; if the character type statistical result of the second alternative character does not meet the preset conditions, determine the sector data of the current sector in the single-byte encoding method, and perform invalid character detection and character analysis to determine the first alternative character; if the character type statistical result of the first alternative character meets the preset conditions, set the first alternative character as the document content of the damaged DOC document; if the character type statistical result of the first alternative character does not meet the preset conditions, determine that the current sector does not store the document content of the damaged DOC document.

[0121] Method 3: Determine the sector data of the current sector according to the single-byte encoding and double-byte encoding methods respectively, and perform invalid character detection and character analysis to determine alternative characters. Judge whether the character type statistical results of all alternative characters in the current sector meet the preset conditions; if so, set all alternative characters in the current sector as the document content of the damaged DOC document.

[0122] As a further introduction to the above Scheme 1, the extraction of document content can be achieved through the following steps B1 - B8:

[0123] Step B1: Select the current sector from the sectors corresponding to the damaged DOC document.

[0124] Step B2: Determine the first sector data of the current sector according to the single-byte encoding method.

[0125] Step B3: Use the compressed character quick reference table to perform invalid character detection on each character in the first sector data.

[0126] Among them, the compressed character quick reference table stores the corresponding relationship between compressed characters and character statuses, and the character statuses include valid and invalid. An invalid character is a character with an invalid character status.

[0127] Step B4: Perform character analysis on the current sector based on the first sector data to obtain the first alternative characters.

[0128] Specifically, if there are no invalid characters in the current sector, judge whether there are null characters in the current sector (that is, the first sector data); if there are no null characters in the current sector, set all characters in the current sector as the first alternative characters; if there are null characters in the current sector and all characters after any null character are also null characters, set all characters before the earliest null character in the current sector as the first alternative characters.

[0129] Step B5: Judge whether the character type statistical results of all first alternative characters in the current sector meet the first preset condition; if so, set all first alternative characters in the current sector as the document content of the damaged DOC document; if not, go to Step B6.

[0130] Specifically, in this step, the character type statistics of all first alternative characters in the current sector can be performed; judge whether the proportion of common characters in the character type statistical results is greater than the first threshold; among them, the common characters include any one or several combinations of letters, numbers, and spaces; if so, it is determined that the character type statistical results meet the first preset condition; if not, it is determined that the character type statistical results do not meet the first preset condition.

[0131] Step B6: If the character type statistics results of all the first alternative characters in the current sector do not meet the first preset condition, determine the second sector data of the current sector in the manner of double-byte encoded data, and perform invalid character detection on the current sector based on the second sector data.

[0132] Specifically, in this step, the second sector data of the current sector can be determined in the manner of double-byte encoded data; use the non-compressed character quick reference table to perform invalid character detection on each character in the second sector data; wherein, the non-compressed character quick reference table stores the corresponding relationship between non-compressed characters and character statuses, and the character statuses include valid and invalid, and an invalid character is a character with an invalid character status.

[0133] Step B7: Perform character analysis on the current sector based on the second sector data to obtain second alternative characters.

[0134] Specifically, if there are no invalid characters in the current sector, determine whether there are null characters in the current sector (i.e., the second sector data); if there are no null characters in the current sector, set all the characters in the current sector as the second alternative characters; if there are null characters in the current sector and all the characters after any null character are null characters, set all the characters before the earliest null character in the current sector as the second alternative characters.

[0135] Step B8: Determine whether the character type statistics results of all the second alternative characters in the current sector meet the second preset condition; if so, set all the second alternative characters in the current sector as the document content of the damaged DOC document; if not, determine that all the characters in the current sector are not the document content of the damaged DOC document.

[0136] Specifically, the non-compressed character quick reference table also stores the corresponding relationship between characters and language types. In this step, character type statistics can be performed on all the second alternative characters in the current sector based on the non-compressed character quick reference table; determine whether the number of language types in the character type statistics results is less than the second threshold and the preset character proportion is within the preset range; if so, determine that the character type statistics results meet the second preset condition; if not, determine that the character type statistics results do not meet the second preset condition. The above number of language types refers to the total number of language types of all the second alternative characters in the current sector.

[0137] As a further introduction to the above Solution 2, the extraction of document content can be implemented through the following steps C1 - C8:

[0138] Step C1: Select the current sector from the sectors corresponding to the damaged DOC document.

[0139] Step C2: Determine the second sector data of the current sector in the manner of double-byte encoded data.

[0140] Step C3: Use the non-compressed character quick reference table to perform invalid character detection on each character in the second sector data.

[0141] Among them, the non-compressed character quick reference table stores the correspondence between non-compressed characters and character statuses. The character statuses include valid and invalid. An invalid character is a character with an invalid character status.

[0142] Step C4: Perform character analysis on the current sector based on the second sector data to obtain second alternative characters.

[0143] Step C5: Determine whether the character type statistical result of all the second alternative characters in the current sector meets the second preset condition; if so, set all the second alternative characters in the current sector as the document content of the damaged DOC document; if not, proceed to Step C6.

[0144] Specifically, the non-compressed character quick reference table also stores the correspondence between characters and language types. In this step, the character type of all the second alternative characters in the current sector can be statistically analyzed based on the non-compressed character quick reference table; determine whether the number of language types in the character type statistical result is less than the second threshold and the preset character proportion is within the preset range; if so, determine that the character type statistical result meets the second preset condition; if not, determine that the character type statistical result does not meet the second preset condition. The above-mentioned preset characters can include language characters and / or special characters. If the number of language types is less than the second threshold, the proportion of language characters is within the first preset range, and the proportion of special characters is within the second preset range, then the character type statistical result meets the second preset condition, otherwise the character type statistical result does not meet the second preset condition.

[0145] Step C6: If the character type statistical result of all the second alternative characters in the current sector does not meet the second preset condition, determine the first sector data of the current sector in the manner of single-byte encoded data, and perform invalid character detection on the current sector based on the first sector data.

[0146] Step C7: Perform character analysis on the current sector based on the first sector data to obtain first alternative characters.

[0147] Step C8: Determine whether the character type statistical result of all the first alternative characters in the current sector meets the first preset condition; if so, set all the first alternative characters in the current sector as the document content of the damaged DOC document; if not, determine that all the characters in the current sector are not the document content of the damaged DOC document.

[0148] The following describes the process described in the above embodiments through examples in actual applications.

[0149] In the related art, the document content is usually extracted from a damaged DOC document in the following manner: directly scan and analyze based on the binary data of the damaged DOC document itself, and read the document content of the damaged DOC document, thereby completing the extraction of the damaged DOC document.

[0150] When there are documents of multiple language types (i.e., language types), especially languages of small language types (such as German documents, documents in Romance languages, Greek documents, Coptic and other various language documents), the above related technologies cannot accurately identify and extract the corresponding document content. At the same time, since the document content of a DOC document has two storage methods: compressed text and uncompressed text. For a DOC document, it generally contains multiple document content segments, and the storage methods of these segments are defined separately. It is possible that some document content is compressed and some document content is uncompressed. For a DOC document whose entire document content is stored in a compressed manner, the above related technologies cannot analyze and process it. Since the text in a DOC document is continuous in sectors, if there are some invalid characters at the start position or in the middle position, then the remaining content of the sector must not be the content of the main document data. If the above means are used for analysis and extraction, there may be misjudgment, and then the remaining content of the sector will continue to be analyzed and extracted, resulting in a waste of extraction time.

[0151] In view of the technical problems existing in the above related technologies, this embodiment proposes a method for extracting document content from an MS-DOC document. By creating a compressed character quick reference table and an uncompressed character quick reference table to quickly identify and judge valid characters and invalid characters, and reading the MS-DOC file sector into the memory buffer, and detecting whether the read sector data is a compressed main document and an uncompressed main document, it can extract the document content in the MS-DOC document with very high accuracy and cover a wider range of languages. No matter how severely damaged the MS-DOC document is, as long as the main document content in it is not completely damaged, the document content can be extracted from this MS-DOC document by using this embodiment.

[0152] Please refer to Figure 2 , Figure 2 which is a flowchart of a method for extracting document content of a DOC document provided by an embodiment of the present application, and specifically includes the following steps:

[0153] S201: Start detecting the document content.

[0154] S202: Create a compressed character quick reference table.

[0155] The compressed character quick reference table contains 256 characters. The values of 22 characters with indexes 0x02, 0x05, 0x06, etc. can be set to false (i.e., the characters corresponding to these indexes cannot appear in the main document of MS-DOC), and the values of other characters are set to true. The value of a character in the compressed character quick reference table is the character status, where true indicates that the character is valid and false indicates that the character is invalid.

[0156] S203: Create an uncompressed character quick reference table.

[0157] The uncompressed character quick reference table contains 65,536 characters, with each element occupying one byte. The highest bit identifies whether the corresponding character is a valid character in the MS-DOC main document, and the remaining 7 bits are used to identify the language to which the corresponding character belongs (custom part, for example, 0 indicates not a language character, 1 indicates a Latin language character, 2 indicates a Greek and Coptic language character, 3 indicates a Chinese language character, etc.). Languages or characters rarely used in Unicode can be defined as invalid characters. The value of a character in the uncompressed character quick reference table is the character status, where true indicates that the character is valid and false indicates that the character is invalid.

[0158] S204: Read data of one sector (512 bytes) from the MS-DOC document into the memory buffer;

[0159] S205: Determine whether the end of the document has been read; if so, end the detection; if not, go to S206;

[0160] S206: Execute the compressed main document detection process on the data in the read sector buffer.

[0161] S207: Determine whether the sector data is a compressed main document; if so, remove the special characters in the main document and then output or save the document content; if not, go to S208.

[0162] S208: Execute the uncompressed main document detection process on the data in the read sector buffer.

[0163] S209: Determine whether the sector data is an uncompressed main document; if so, remove the special characters in the main document and then output or save the document content; if not, go to S210.

[0164] S210: Discard the data in the current sector buffer and go to S204.

[0165] Please refer to Figure 3 , Figure 3 which is a schematic diagram of a compressed main document detection process provided by an embodiment of the present application, and the process includes:

[0166] Start detecting the current sector buffer; read a character (one byte) from the current sector buffer.

[0167] Determine whether the buffer has been read completely.

[0168] If the buffer has not been read completely, query in the compressed character quick reference table using the value of the character as the index. Determine whether it is a valid character; if it is not a valid character, determine that the current sector is not the compressed main document and end the detection; if it is a valid character, determine whether 0x00 has been read before. If 0x00 has been read before, determine whether the value of the currently read character is 0x00; if it is 0x00, enter the step of reading a character from the current sector buffer; if it is not 0x00, determine that the current sector is not the compressed main document and end the detection. If 0x00 has not been read before, determine whether the value of the character is 0x00; if it is 0x00, mark that 0x00 has been read; if it is not 0x00, classify and count the read characters.

[0169] If the buffer has been read completely, calculate the sector character distribution; determine whether the proportion of common characters exceeds 60%; if so, determine that the current sector is the compressed main document and end the detection; if not, determine that the current sector is not the compressed main document and end the detection.

[0170] Please refer to Figure 4 , Figure 4 which is a schematic diagram of a non-compressed main document detection process provided by an embodiment of the present application, and the process includes:

[0171] Start detecting the current sector buffer; read a Unicode character (two bytes) from the current sector buffer.

[0172] Determine whether the buffer has been read completely.

[0173] If the buffer has not been read completely, query in the non-compressed character quick reference table using the value of the character as the index. Determine whether it is a valid character; if it is not a valid character, determine that the current sector is not the non-compressed main document and end the detection; if it is a valid character, determine whether 0x0000 has been read before. If 0x0000 has been read before, determine whether the value of the currently read character is 0x0000; if it is 0x0000, enter the step of reading a Unicode character from the current sector buffer; if it is not 0x0000, determine that the current sector is not the non-compressed main document and end the detection. If 0x0000 has not been read before, determine whether the value of the character is 0x0000; if it is 0x0000, mark that 0x0000 has been read; if it is not 0x0000, classify and count the read characters.

[0174] If the buffer has been read completely, calculate the sector character distribution.

[0175] Determine whether the number of language types exceeds 3; if the number of language types exceeds 3, it is determined that the current sector is not an uncompressed main document, and the detection ends; if the number of language types does not exceed 3, determine whether the proportion of language characters exceeds 50%; if the proportion of language characters does not exceed 50%, it is determined that the current sector is not an uncompressed main document, and the detection ends; if the proportion of language characters exceeds 50%, determine whether the proportion of special characters is less than 35%; if the proportion of special characters is less than 35%, it is determined that the current sector is an uncompressed main document, and the detection ends; if the proportion of special characters is not less than 35%, it is determined that the current sector is not an uncompressed main document, and the detection ends.

[0176] Specifically, this embodiment includes the following steps D1 - D7:

[0177] Step D1: Create a compressed character quick reference table (mainly used to quickly identify invalid characters). This quick reference table is an array of boolean type, containing a total of 256 elements. Each of these elements corresponds to a character in the ISO8859 - 1 encoding table. The value of the element being true indicates that the corresponding character is a valid character (including special characters defined by MS - DOC), and the value of the element being false indicates that the corresponding character is an invalid character.

[0178] Specifically, among these 256 elements, the values of the elements with indexes 0x02, 0x05, 0x06, 0x0A, 0x0E, 0x0F, 0x10, 0x11, 0x12, 0x16, 0x17, 0x18, 0x19, 0x1A, 0x1B, 0x1C, 0x1D, 0x1F, 0x7F, 0x80, 0x81, 0x82 are set to false (false), and the values of the remaining other elements are all set to true (true).

[0179] Step D2: Create an uncompressed character quick reference table (mainly used to quickly identify invalid characters and mark the language type to which the characters belong). This quick reference table is also an array, containing a total of 65536 elements, and each of these elements corresponds to a character in the Unicode encoding table.

[0180] The type of an array element is an 8-bit integer value (occupying one byte). The highest bit indicates whether the corresponding character is a valid character (special characters defined by MS-DOC), and the remaining 7 bits are used to define the language to which the character belongs (custom part. For example, 0x00 indicates not a language character, 0x01 indicates Latin language characters, 0x02 indicates Greek and Coptic characters, 0x03 indicates Chinese characters, 0x04 indicates Japanese characters, etc.). In addition to defining the invalid characters defined in the "compressed character quick reference table" as invalid characters in the current table, extremely rarely used languages or characters in Unicode (such as some extremely rarely used symbols) can also be defined as invalid characters.

[0181] Step D3: Read the data of one sector (512 bytes) from the MS-DOC file into the memory buffer (read sequentially starting from the beginning of the file). If the end of the file has been reached, end the entire processing flow.

[0182] Step D4: Detect the sector data read and determine whether it is the compressed main document content. The specific methods include D401 - D406:

[0183] D401: Read a character (each character occupies one byte) sequentially from the buffer. If there are no more characters to read in the buffer, proceed to step D406 to continue execution.

[0184] D402: Query in the compressed character quick reference table using the value of the character as the index. If the character is an invalid character, the data of the current sector is not the compressed main document, and end the detection.

[0185] D403: If 0x00 has been read: Check whether the current character is 0x00. If it is 0x00, proceed to D401 to continue execution; if it is not 0x00, the current sector is not the compressed main document, and end the detection.

[0186] If 0x00 has not been read: Detect whether the current character is 0x00. If it is 0x00, mark it (mark that 0x00 has been read) and proceed to step 1) to continue execution; if it is not 0x00, enter D404 for classification and counting.

[0187] D404: Classify and count the characters read: If the character is a letter, increment the letter count; if the character is a number, increment the number count; if it is a space, increment the space count.

[0188] D405: Proceed to D401 to continue execution.

[0189] D406: Consider characters of letters, numbers, and spaces as "common characters". If the proportion of "common characters" exceeds a certain ratio (e.g., 60%), then the current sector is the compressed main document; otherwise, it is not. After the judgment, end the data detection of the current sector.

[0190] Step D5: If the data read is the compressed main document, then special characters can be removed and this part of the content can be saved or output, and then go to step D3 to continue the process.

[0191] Step D6: If the data read is not the compressed main document, then continue the detection to determine whether it is the content of the uncompressed main document. The specific methods include D601 - D606:

[0192] D601: Read one character (each character occupies two bytes) from the buffer in sequence. If there are no more characters to read in the buffer, then go to D606 to continue the execution.

[0193] D602: Query in the uncompressed character quick reference table with the value of the character as the index. If the character is an invalid character, then the data in the current sector is not the uncompressed main document, and end the detection.

[0194] D603: If 0x0000 has been read: Check whether the current character is 0x0000. If it is 0x0000, then go to D601 to continue the execution; if it is not 0x0000, then the current sector is not the uncompressed main document, and end the detection.

[0195] If 0x0000 has not been read: Detect whether the current character is 0x0000. If it is 0x0000, then make a mark (mark that 0x0000 has been read) and go to D601 to continue the execution; if it is not 0x0000, then enter D604 for classification and counting.

[0196] D604: Classify and count the characters read: Increment the corresponding language character statistical count for different language characters; if it is a special character, then increment the special character statistical count by one.

[0197] D605: Go to D601 to continue the execution.

[0198] D606: If the number of language types exceeds a certain number (such as more than three), then the current sector is not the uncompressed main document, and end the data detection of the current sector. If the total number of language characters accounts for more than a certain ratio (e.g., 50%) of the valid characters, and the total number of special characters accounts for less than a certain ratio (e.g., 35%) of the valid characters, then the data in the current sector is the uncompressed main document; otherwise, it is not. After the judgment, end the data detection of the current sector.

[0199] Step D7: If the read data is an uncompressed main document, special characters can be removed, and then this part of the content can be saved or output, and then go to step D3 to continue processing.

[0200] In this embodiment, when extracting the content of an MS-DOC document, whether the document content is stored in the main document in a compressed manner or an uncompressed manner, it can be effectively identified and extracted. This method has the advantages of a wide coverage of language types and accurate recognition.

[0201] Through the compressed document detection process: This embodiment can accurately identify and extract the content of Latin language documents stored in a compressed manner, such as languages like English, French, German, Spanish, and Portuguese; through the uncompressed document detection process: This embodiment can accurately identify and extract the content of multi-language documents stored in an uncompressed manner, such as almost all common languages like Chinese, Japanese, Korean, Greek, Cyrillic, and Latin languages. The content of Latin language documents can be stored in a compressed manner (ISO-8859 encoding, each character occupies one byte) or in an uncompressed manner (Unicode encoding, one character occupies two bytes).

[0202] Based on the characteristics of the MS-DOC document itself, this embodiment adopts a method of identifying document content in units of sectors (512 bytes), combines two created character quick reference tables, and includes corresponding valid and invalid flags, as well as language classification information in the quick reference tables, which can quickly identify, judge, and classify the read content, and has the advantages of high recognition efficiency and high accuracy. Based on the characteristics of the MS-DOC document itself, this embodiment optimizes the algorithm for characters with a value of 0 during text recognition. When a character with a value of 0 appears in a sector, the remaining content of the sector should all be characters with a value of 0, otherwise this sector is not a sector of the main document and the document content does not need to be extracted. This optimization not only improves the recognition efficiency but also reduces the misjudgment rate. For example, for the data of a certain sector, even if it contains some readable text information but violates this rule, it means that the current sector is not a sector of the main document, and this readable text information is not the document content that needs to be extracted.

[0203] The embodiment of the present application also provides a document content extraction system, including:

[0204] A sector selection module for selecting the current sector from the sectors corresponding to the damaged DOC document;

[0205] A data determination module for determining the sector data of the current sector according to a preset character encoding method, and performing invalid character detection on the current sector based on the sector data;

[0206] A character analysis module for performing character analysis on the current sector based on the sector data to obtain candidate characters; wherein, the process of character analysis includes: if there are no invalid characters in the current sector, determining whether there are null characters in the current sector; if there are no null characters in the current sector, setting all characters in the current sector as candidate characters; if there are null characters in the current sector and all characters after any null character are also null characters, setting all characters before the earliest null character in the current sector as candidate characters.

[0207] A decision-making module for determining whether the character type statistical result of all candidate characters in the current sector meets a preset condition; if so, setting all candidate characters in the current sector as the document content of the damaged DOC document.

[0208] Furthermore, the character analysis module is also used to determine that all characters in the current sector are not the document content of the damaged DOC document if there are invalid characters in the current sector; and is also used to determine that all characters in the current sector are not the document content of the damaged DOC document if there are null characters in the current sector and not all characters after any null character are null characters.

[0209] Furthermore, the process by which the data determination module determines the sector data of the current sector according to a preset character encoding method and performs invalid character detection on the current sector based on the sector data includes: determining the first sector data of the current sector in a single-byte encoding manner; performing invalid character detection on each character in the first sector data by using a compressed character quick reference table; wherein, the compressed character quick reference table stores the correspondence between compressed characters and character states, the character states include valid and invalid, and an invalid character is a character with an invalid character state.

[0210] Correspondingly, the process by which the character analysis module performs character analysis on the current sector based on the sector data to obtain candidate characters includes: performing character analysis on the current sector based on the first sector data to obtain first candidate characters.

[0211] Correspondingly, the process by which the decision-making module determines whether the character type statistical result of all candidate characters in the current sector meets a preset condition includes: determining whether the character type statistical result of all first candidate characters in the current sector meets a first preset condition.

[0212] Correspondingly, the process by which the decision-making module sets all candidate characters in the current sector as the document content of the damaged DOC document includes: setting all first candidate characters in the current sector as the document content of the damaged DOC document.

[0213] Further, the process by which the decision-making module determines whether the character type statistical result of all first alternative characters in the current sector meets the first preset condition includes: statistically analyzing the character types of all first alternative characters in the current sector; determining whether the proportion of common characters in the character type statistical result is greater than a first threshold; wherein the common characters include any one or a combination of several of letters, numbers, and spaces; if so, it is determined that the character type statistical result meets the first preset condition; if not, it is determined that the character type statistical result does not meet the first preset condition.

[0214] Further, after the decision-making module determines whether the character type statistical result of all first alternative characters in the current sector meets the first preset condition, the operations it further performs include: if the character type statistical result of all first alternative characters in the current sector does not meet the first preset condition, determining the second sector data of the current sector in the manner of double-byte encoded data, and performing invalid character detection on the current sector based on the second sector data; performing character analysis on the current sector based on the second sector data to obtain second alternative characters; determining whether the character type statistical result of all second alternative characters in the current sector meets the second preset condition; if so, setting all second alternative characters in the current sector as the document content of the damaged DOC document; if not, determining that all characters in the current sector are not the document content of the damaged DOC document.

[0215] Further, the process by which the data determination module determines the sector data of the current sector according to the preset character encoding method and performs invalid character detection on the current sector based on the sector data includes: determining the second sector data of the current sector in the manner of double-byte encoded data; performing invalid character detection on each character in the second sector data using a non-compressed character quick reference table; wherein the non-compressed character quick reference table stores the correspondence between non-compressed characters and character states, and the character states include valid and invalid, and an invalid character is a character with an invalid character state.

[0216] Correspondingly, the process by which the character analysis module performs character analysis on the current sector based on the sector data to obtain alternative characters includes: performing character analysis on the current sector based on the second sector data to obtain second alternative characters.

[0217] Correspondingly, the process by which the decision-making module determines whether the character type statistical result of all alternative characters in the current sector meets the preset condition includes: determining whether the character type statistical result of all second alternative characters in the current sector meets the second preset condition.

[0218] Correspondingly, the process in which the decision module sets all alternative characters in the current sector as the document content of the damaged DOC document includes: setting all second alternative characters in the current sector as the document content of the damaged DOC document.

[0219] Further, the non-compressed character quick reference table also stores the correspondence between characters and language types;

[0220] The process in which the decision module determines whether the character type statistical result of all second alternative characters in the current sector meets the second preset condition includes: performing character type statistics on all second alternative characters in the current sector based on the non-compressed character quick reference table; determining whether the number of language types in the character type statistical result is less than the second threshold and the preset character proportion is within the preset range; if so, determining that the character type statistical result meets the second preset condition; if not, determining that the character type statistical result does not meet the second preset condition.

[0221] Further, after the decision module determines whether the character type statistical result of all second alternative characters in the current sector meets the second preset condition, the operations further performed include: if the character type statistical result of all second alternative characters in the current sector does not meet the second preset condition, determining the first sector data of the current sector in the manner of single-byte encoded data, and performing invalid character detection on the current sector based on the first sector data; performing character analysis on the current sector based on the first sector data to obtain first alternative characters; determining whether the character type statistical result of all first alternative characters in the current sector meets the first preset condition; if so, setting all first alternative characters in the current sector as the document content of the damaged DOC document; if not, determining that all characters in the current sector are not the document content of the damaged DOC document.

[0222] Since the embodiments in the system part correspond to the embodiments in the method part, for the embodiments in the system part, please refer to the description of the embodiments in the method part, which will not be elaborated here.

[0223] The present application also provides a storage medium on which a computer program is stored, and when the computer program is executed, the steps provided in the above embodiments can be implemented. The storage medium may include: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

[0224] The various embodiments in the specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts among the embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description in the method section. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of this application.

[0225] It should also be noted that in this specification, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the existence of another identical element in the process, method, article or device comprising the element.

Claims

1. A method for extracting document content, characterized in that, Including: Select the current sector from the sectors corresponding to the damaged DOC document; Determine the sector data of the current sector according to the preset character encoding method, and perform invalid character detection on the current sector based on the sector data; Perform character analysis on the current sector based on the sector data to obtain alternative characters; wherein, the process of the character analysis includes: if there are no invalid characters in the current sector, determine whether there are null characters in the current sector; if there are no null characters in the current sector, set all the characters in the current sector as alternative characters; if there are null characters in the current sector and all the characters after any null character are null characters, set all the characters before the null character with the earliest position in the current sector as alternative characters; Judge whether the character type statistical result of all alternative characters in the current sector meets the preset conditions; If so, set all alternative characters in the current sector as the document content of the damaged DOC document.

2. The method for extracting document content according to claim 1, wherein It also includes: If there are invalid characters in the current sector, determine that all characters in the current sector are not the document content of the damaged DOC document; If there are null characters in the current sector and all the characters after any null character are not all null characters, determine that all characters in the current sector are not the document content of the damaged DOC document.

3. The method for extracting document content according to claim 1, wherein Determine the sector data of the current sector according to the preset character encoding method, and perform invalid character detection on the current sector based on the sector data, including: Determine the first sector data of the current sector in the way of single-byte encoding; Use the compressed character quick reference table to perform invalid character detection on each character in the first sector data; wherein, the compressed character quick reference table stores the corresponding relationship between compressed characters and character states, and the character states include valid and invalid, and the invalid character is the character with the character state of invalid; Correspondingly, perform character analysis on the current sector based on the sector data to obtain alternative characters, including: Perform character analysis on the current sector based on the first sector data to obtain the first alternative characters; Correspondingly, judge whether the character type statistical result of all alternative characters in the current sector meets the preset conditions, including: Judge whether the character type statistical result of all the first alternative characters in the current sector meets the first preset conditions; Correspondingly, set all alternative characters in the current sector as the document content of the damaged DOC document, including: Set all the first alternative characters in the current sector as the document content of the damaged DOC document.

4. The method for extracting document content according to claim 3, characterized in that, Judge whether the character type statistical result of all the first alternative characters in the current sector meets the first preset conditions, including: Perform character type statistics on all the first alternative characters in the current sector; Judge whether the proportion of common characters in the character type statistical result is greater than the first threshold; wherein, the common characters include any one or several combinations of letters, numbers and spaces; If so, determine that the character type statistical result meets the first preset conditions; If not, determine that the character type statistical result does not meet the first preset conditions.

5. The method for extracting document content according to claim 3, wherein After judging whether the character type statistical result of all the first alternative characters in the current sector meets the first preset conditions, it also includes: If the character type statistics results of all the first alternative characters in the current sector do not meet the first preset condition, determine the second sector data of the current sector in the manner of double-byte encoded data, and perform invalid character detection on the current sector based on the second sector data; Perform character analysis on the current sector based on the second sector data to obtain second alternative characters; Determine whether the character type statistics results of all the second alternative characters in the current sector meet the second preset condition; If so, set all the second alternative characters in the current sector as the document content of the damaged DOC document; If not, determine that all the characters in the current sector are not the document content of the damaged DOC document.

6. The method for extracting document content according to claim 1, characterized in that Determine the sector data of the current sector according to the preset character encoding method, and perform invalid character detection on the current sector based on the sector data, including: Determine the second sector data of the current sector in the manner of double-byte encoded data; Perform invalid character detection on each character in the second sector data by using the non-compressed character quick reference table; wherein, the non-compressed character quick reference table stores the correspondence between non-compressed characters and character states, and the character states include valid and invalid, and the invalid characters are the characters with the character state of invalid; Correspondingly, perform character analysis on the current sector based on the sector data to obtain alternative characters, including: Perform character analysis on the current sector based on the second sector data to obtain second alternative characters; Correspondingly, determine whether the character type statistics results of all the alternative characters in the current sector meet the preset condition, including: Determine whether the character type statistics results of all the second alternative characters in the current sector meet the second preset condition; Correspondingly, set all the alternative characters in the current sector as the document content of the damaged DOC document, including: Set all the second alternative characters in the current sector as the document content of the damaged DOC document.

7. The method for extracting document content according to claim 6, wherein The non-compressed character quick reference table also stores the correspondence between characters and language types; Correspondingly, determine whether the character type statistics results of all the second alternative characters in the current sector meet the second preset condition, including: Perform character type statistics on all the second alternative characters in the current sector based on the non-compressed character quick reference table; Determine whether the number of language types in the character type statistics results is less than the second threshold and the proportion of preset characters is within the preset range; If so, determine that the character type statistics results meet the second preset condition; If not, determine that the character type statistics results do not meet the second preset condition.

8. The method for extracting the document content according to claim 6, wherein After determining whether the character type statistics results of all the second alternative characters in the current sector meet the second preset condition, it also includes: If the character type statistics results of all the second alternative characters in the current sector do not meet the second preset condition, determine the first sector data of the current sector in the manner of single-byte encoded data, and perform invalid character detection on the current sector based on the first sector data; Perform character analysis on the current sector based on the first sector data to obtain first alternative characters; Determine whether the character type statistics results of all the first alternative characters in the current sector meet the first preset condition; If so, set all the first alternative characters in the current sector as the document content of the damaged DOC document; If not, determine that all characters in the current sector are not the document content of the damaged DOC document.

9. A document content extraction system, characterized in that, Including: A sector selection module, configured to select the current sector from the sectors corresponding to the damaged DOC document; A data determination module, configured to determine the sector data of the current sector according to a preset character encoding method, and perform invalid character detection on the current sector based on the sector data; A character analysis module, configured to perform character analysis on the current sector based on the sector data to obtain alternative characters; wherein, the process of the character analysis includes: if there is no invalid character in the current sector, determine whether there is a null character in the current sector; if there is no null character in the current sector, set all characters in the current sector as alternative characters; if there is a null character in the current sector and all characters after any null character are null characters, set all characters before the null character with the earliest position in the current sector as alternative characters; A decision module, configured to determine whether the character type statistical result of all alternative characters in the current sector meets a preset condition; if so, set all alternative characters in the current sector as the document content of the damaged DOC document.

10. A storage medium, characterized in that, The storage medium stores computer-executable instructions, and when the computer-executable instructions are loaded and executed by a processor, the steps of the document content extraction method according to any one of claims 1 to 8 are implemented.