Document content extraction method and device and related equipment
By performing classification and configuration file-driven extraction and filtering of HTML documents, the problem of poor extraction of HTML documents is solved, and effective recognition and extraction of multiple formats is achieved.
Patent Information
- Application Number
- CN202510357371.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, the content extraction effect of HTML documents is poor, and other formats other than text content cannot be effectively identified.
By classifying the to be processed documents, obtaining the target configuration file, extracting and filtering the document content in different formats according to the extraction rules and processing rules of the configuration file, and generating the target extraction text.
It realizes the recognition of multiple formats in HTML documents, and improves the extraction effect of document content.
Smart Images

Figure CN120257974A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of document processing, and particularly to a method, an apparatus and related devices for extracting document content. Background Art
[0002] An HTML document is a file written in Hyper Text Markup Language (HTML) and is used to create and structure web page content. When actually using an HTML document to write product requirements, there are often a lot of irrelevant or invalid interference information. It is necessary to clean the HTML document in advance and extract the required requirement information before generating more effective test cases.
[0003] Currently, when extracting content from an HTML document, only text content can be extracted, and other formats in the HTML document cannot be effectively recognized. Therefore, the extraction effect of the HTML document is poor. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide a method, an apparatus and related devices for extracting document content to solve the problem of poor extraction effect of HTML documents in the prior art. The specific technical solutions are as follows:
[0005] In the first aspect of the present invention, first, a method for extracting document content is provided, which is applied to a document cleaning system. The method includes:
[0006] Classify the document content of the document to be processed to obtain a plurality of data sets, and each data set is a set of document content of a format type;
[0007] Obtain a target configuration file that matches the document to be processed;
[0008] For each data set, extract the content of the document content in the data set according to the target configuration file to obtain the first text content corresponding to each document content, and screen and process the first text content according to the target configuration file to obtain the second text content. The target configuration file includes extraction rules and processing rules corresponding to each format type included in the document to be processed. The first text content is the text obtained by extracting content according to the extraction rules corresponding to the data set in the target configuration file, and the second text content is the text obtained by screening and processing according to the processing rules corresponding to the data set in the target configuration file;
[0009] Combine the second text contents corresponding to the plurality of data sets to obtain a target extraction text.
[0010] In a second aspect, an apparatus for extracting document content is provided. The apparatus includes:
[0011] A classification module, configured to classify the document content of the document to be processed, obtaining a plurality of data sets, where each data set is a set of document content of a format type;
[0012] An acquisition module, configured to acquire a target configuration file that matches the document to be processed;
[0013] An extraction module, configured to extract the content of the document in the data set according to the target configuration file to obtain a first text content corresponding to each document content, and perform screening processing on the first text content according to the target configuration file to obtain a second text content. Wherein, the target configuration file includes extraction rules and processing rules corresponding to each format type included in the document to be processed. The first text content is the text obtained by extracting the content according to the extraction rules corresponding to the data set in the target configuration file, and the second text content is the text obtained by performing screening processing according to the processing rules corresponding to the data set in the target configuration file;
[0014] A combination module, configured to combine the second text contents corresponding to the plurality of data sets to obtain a target extraction text.
[0015] In a third aspect of the embodiments of the present application, an electronic device is further provided, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method for extracting document content as described in any one of the first aspects are implemented.
[0016] In a fourth aspect of the embodiments of the present application, a computer-readable storage medium is further provided. A program or instruction is stored on the readable storage medium. When the program or instruction is executed by a processor, the steps of the method for extracting document content as described in any one of the first aspects are implemented.
[0017] In a fifth aspect of the embodiments of the present application, a computer program product is further provided. The computer program product is stored in a storage medium. The computer program product is executed by at least one processor to implement the steps of the method for extracting document content as described in any one of the first aspects.
[0018] An embodiment of the present invention provides a method, apparatus, and related device for extracting document content. The method includes: classifying the document content of a document to be processed to obtain a plurality of data sets, where each data set is a set of document content of a format type; obtaining a target configuration file that matches the document to be processed; for each data set, extracting the content of the document content in the data set according to the target configuration file to obtain first text content corresponding to each document content, and filtering the first text content according to the target configuration file to obtain second text content, where the target configuration file includes extraction rules and processing rules corresponding to each format type included in the document to be processed, the first text content is text obtained by extracting content according to the extraction rules corresponding to the data set in the target configuration file, and the second text content is text obtained by filtering according to the processing rules corresponding to the data set in the target configuration file; combining the second text content corresponding to the plurality of data sets to obtain target extraction text. In this application, the document content of the document to be processed is classified by format type to obtain a plurality of data sets, and extraction rules and processing rules for processing the plurality of data sets are determined according to the target configuration file, completing text extraction and filtering of the plurality of data sets to obtain target extraction text, thereby realizing the recognition of multiple formats in the document to be processed and improving the extraction effect of the document content. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art.
[0020] Figure 1 It is a flowchart of the method for extracting document content in an embodiment of the present application;
[0021] Figure 2 It is a processing flowchart in an embodiment of the present application;
[0022] Figure 3 It is a structural schematic diagram of the apparatus for extracting document content in an embodiment of the present application;
[0023] Figure 4 It is a structural schematic diagram of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0025] Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts depict the steps as sequential processes, many of the steps can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operations are completed, but it can also have additional steps not included in the figures. The process can correspond to a method, function, procedure, subroutine, subprogram, and so on.
[0026] In addition, terms such as "first", "second", etc. may be used herein to describe various directions, actions, steps, or elements, etc., but these directions, actions, steps, or elements are not limited by these terms. These terms are only used to distinguish the first direction, action, step, or element from another direction, action, step, or element. For example, without departing from the scope of the present application, the first order request can be referred to as the second order request, and similarly, the second order request can be referred to as the first order request. Both the first order request and the second order request are order requests, but they are not the same order request. The terms "first", "second", etc. should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second" may explicitly or implicitly include one or more of such features. In the description of the present application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0027] The embodiments of the present application provide a method for extracting document content, which is applied to a document cleaning system, as Figure 1 shown, the method includes:
[0028] Step 101, classify the document content of the document to be processed to obtain a plurality of data sets, and each data set is a set of document content of a format type.
[0029] In this embodiment, the execution subject of the present application can be the server side. After receiving the document processing request sent by the user through the server side, the document to be processed is extracted according to the document processing request. Specifically, the server side is a document cleaning system. Exemplarily, the document cleaning system of the present application can include a front-end page, a service interface, business logic, and data layers. The document to be processed in this embodiment is an HTML document, where the HTML content is a document composed of different blocks defined by multi-level tags. Among them, different blocks respectively correspond to the declaration, boundary, title, text, picture, table, link, etc. of the document.
[0030] In this embodiment, the format types corresponding to different blocks are different. For example, when classifying the document to be processed, the content of the document to be processed can be divided into formats such as text, pictures, and tables. Specifically, the text content included in the document to be processed can be multiple pictures, multiple paragraphs of text, multiple tables, etc. After classification, the multiple pictures, multiple paragraphs of text, and multiple tables are combined separately to obtain corresponding multiple data sets, that is, they can be a picture set, a text set, and a table set. In this embodiment, the format types corresponding to different data sets are different.
[0031] Step 102: Obtain a target configuration file that matches the document to be processed.
[0032] In this embodiment, the target configuration text includes the extraction rules and processing rules corresponding to each format type included in the document to be processed. Among them, the extraction rule is used to indicate the rule for text extraction from the data set, such as which parts of the text to extract, etc. The processing rule is used to indicate the rule for filtering the data set, such as retaining some content in the data set or discarding some content in the data set.
[0033] Exemplarily, the target configuration file is mainly used to determine the extraction rules and processing rules for different formats in the document to be processed. The target configuration file includes the extraction rules and processing rules for different formats, as well as corresponding custom parameters, processing methods, and corresponding custom parameters. Taking a json string as an example, the specific format is:
[0034]
[0035]
[0036] Among them: typeName represents the detailed type identifier, typeDesc is the description, analysisMethod is the corresponding extraction rule, analysisKeyword is the keyword customized for the extraction rule, separated by English commas or Chinese commas, analysisAttr is the specific attribute value it contains (such as thickness, font size, color, underline, strikethrough, italic, etc.), processMethod is the corresponding processing rule, processParams is the customized parameter of the processing rule, and needProcess indicates whether it needs to be recognized.
[0037] It should be noted that the target configuration text can be input by the user or matched according to the document information of the document to be processed to obtain the target configuration text that matches the document to be processed. Specifically, the extraction rules, processing rules, and the parameters corresponding to the extraction rules and processing rules for the corresponding format types are mapped in the target configuration file, thereby determining which contents in the requirement document are valid (such as text, tables, pictures, etc.) and which are invalid (such as titles, backgrounds, revenues, strikethrough text, etc.). In the case where the target configuration file needs to be changed, the extraction rules and processing rules can be dynamically adjusted by adjusting the input parameters.
[0038] Step 103: For each data set, extract the document content in the data set according to the target configuration file to obtain the first text content corresponding to each document content, and screen and process the first text content according to the target configuration file to obtain the second text content, where the target configuration file includes the extraction rules and processing rules corresponding to each format type included in the document to be processed, the first text content is the text obtained by extracting the content according to the extraction rule corresponding to the data set in the target configuration file, and the second text content is the text obtained by screening and processing according to the processing rule corresponding to the data set in the target configuration file.
[0039] In this embodiment, for each obtained data set, the document content in the data set is extracted according to the extraction rule corresponding to the data set. For example, if the format type corresponding to the data set is the text format, then the text in the text format is extracted. The extraction method can be keyword and content attribute extraction, that is, the specific content attributes of the document content are determined through keywords, so as to complete the extraction of the text included in the entire document content and generate the first text content. It should be noted that in this embodiment, after the content of the data sets corresponding to different format information is extracted, the final obtained content is all text content.
[0040] For each data set, the first text content corresponding to the data set is screened according to the processing rule corresponding to the data set to obtain the second text content. Among them, after obtaining the multiple first text contents corresponding to the multiple data sets, the texts included in the multiple first text contents are screened and processed to determine the second text content. Exemplarily, for example, when the first text content is the text obtained by text format extraction, the text content can be subdivided according to the specific content into: background introduction, deleted content, temporarily ignored content, key content to focus on, specific requirements, etc. Among them, only the key texts such as the key content to focus on and specific requirements can be retained, while the non-critical texts such as background introduction, deleted content, and temporarily ignored content are deleted.
[0041] Step 104: Combine the second text contents corresponding to the multiple data sets to obtain the target extraction text.
[0042] In this embodiment, after processing all the data sets, the multiple retained second text contents are combined to generate the target extraction text. Among them, combination means organizing the multiple second text contents according to the needs of the content in terms of rank and segmentation, and reorganizing them to generate complete text content. Among them, rank means classifying, marking, or evaluating the text to determine its importance, quality, relevance, complexity, etc.
[0043] This application classifies the document content of the document to be processed by format type, obtains multiple data sets, determines the extraction rules and processing rules for processing the multiple data sets respectively according to the target configuration file, completes the text extraction and screening of the multiple data sets, and obtains the target extraction text, thereby realizing the recognition of multiple formats in the document to be processed and improving the extraction effect of the document content.
[0044] In some feasible implementation manners, Step 101,
[0045] The document cleaning system includes a content parsing module, and classifying the document content of the document to be processed to obtain multiple data sets, including:
[0046] The content parsing module traverses multiple tag contents in the document to be processed, obtains multiple document contents, and the multiple document contents correspond one-to-one to the multiple tag contents. The tag content is used to represent the sorting level of the corresponding text content in the document to be processed, and the document to be processed is a Feishu document;
[0047] The content parsing module classifies the multiple document contents based on the format types corresponding to the multiple document contents to obtain multiple data sets.
[0048] In this embodiment, the document cleaning system includes a content parsing module for classifying the document content of the document to be processed. Here, the document to be processed is taken as an example of an HTML document for illustration. An HTML document is composed of different "blocks" defined by multi-level tags. These "blocks" respectively correspond to the declaration, boundary, title, text, picture, table, link, etc. of the document. The entire document content can be traversed through the tag content in the document to be processed. Exemplarily, in this embodiment, the traversal of the document to be processed can be implemented through a crawler algorithm, which is a program or script for automatically accessing the Internet and extracting information and is used to automatically obtain information. When the crawler algorithm obtains content, the parsing of the document content can be achieved by carrying authentication information. The main logic is to traverse all the tags of the HTML document and determine the level of the content (whether it is a title, which level of title, and the serial numbers of each level to which the content belongs), text, specific attributes (thickness, font size, color, underline, strikethrough, italic, etc.), and content serial number through specific tag values.
[0049] The content parsing module traverses the multiple tag contents in the document to be processed to obtain multiple document contents, and determines the format types corresponding to different document contents based on the obtained multiple document contents, thereby correspondingly generating multiple data sets. Specifically, different content types can be roughly classified into: pure text content, picture content, table content, hyperlink. Other contents such as declaration, boundary, title, video, etc. can be ignored. Specific content can be further subdivided: text content can be subdivided according to specific content into: background introduction, deleted content, temporarily ignored content, key content to be concerned about, specific requirements; picture content can be subdivided into: flow chart, functional prototype diagram, etc.; table content is similar to the combination of text and pictures, generally mainly text and text converted from pictures: background introduction, deleted content, temporarily ignored content, key content to be concerned about, requirements; hyperlinks are divided into: links that need to pay attention to the document content, links that do not need to pay attention to the document content.
[0050] In this embodiment, the content parsing module traverses the multiple tag contents in the document to be processed to obtain multiple document contents, thereby ensuring the integrity of the recognition of the entire document to be processed, and classifying through the format types corresponding to the multiple document contents, thereby forming multiple different data sets and improving the efficiency of generating multiple data sets.
[0051] Optionally, the document cleaning system further includes a content organization module, and the method further includes:
[0052] The content organization module generates a serial number corresponding to each of the document contents based on the multiple tag contents, and the serial number is used to represent the position of the document content in the to-be-processed document; wherein, the data set includes at least one document content and the serial number corresponding to each of the document contents.
[0053] The combining the second text contents corresponding to the multiple data sets to obtain the target extraction text includes:
[0054] The content organization module determines the serial number corresponding to each of the second text contents based on the serial number corresponding to each of the document contents.
[0055] The content organization module combines the second text contents corresponding to the multiple data sets in sequence according to the serial number corresponding to each of the second text contents to obtain the target extraction text.
[0056] In this embodiment, the content organization module sorts each document content according to multiple tag contents, and then combines the corresponding second text contents according to the sorting serial numbers to obtain the target extraction text. Specifically, based on the tag contents obtained by summarizing according to the above embodiments, a serial number corresponding to each of the document contents is generated. Specifically, the generation method of the serial number is in the format of "level 1 serial number_level 2 serial number_level 3 serial number_level 4 serial number_level 5 serial number_sort order of content blocks within the level", and it is determined according to the corresponding level serial number and the sort order of content blocks within the level. If there is no content in the corresponding level, it is represented by 0. For example, the text content "I. Requirement Background" is the title of the first level, and its serial number is "1_0_0_0_0", and the text content "To support XXX and solve XXX problems" is under the scope of "I. Requirement Background" and introduces the specific background of the requirement, so its serial number is "1_0_0_0_1".
[0057] After determining the serial number corresponding to each document content, the serial number corresponding to each second text content is thus determined. According to the serial number corresponding to each second text content, the second text contents corresponding to the multiple data sets are combined in sequence to obtain the target extraction text. Among them, the priority of the serial number is determined according to the size value of the serial number itself, and the smaller the value of the serial number, the more forward the arrangement position.
[0058] By using the method of combining the second text contents according to the serial number in this application to obtain the target extraction text, it can ensure that the text content included in the combined target extraction text is in the same position as the relevant content in the to-be-processed text, ensuring the accuracy of the extraction text.
[0059] Optionally, in step 103, the document cleaning system further includes a content recognition module, and obtaining the target configuration file matching the to-be-processed document includes:
[0060] The content recognition module determines the format types corresponding to the multiple data sets;
[0061] The content recognition module matches among multiple candidate configuration files according to the format types corresponding to the multiple data sets to determine the target configuration file. The target configuration file is any one of the multiple candidate configuration files, and the target configuration file includes the extraction rules and processing rules for the format types corresponding to the multiple data sets.
[0062] In this embodiment, the content recognition module is used to obtain the target configuration file matching the to-be-processed document. Specifically, the format types corresponding to the multiple data sets may be composed of formats such as text format, table format, picture format, and hyperlink format. It should be noted that there may be other formats in other embodiments. In this embodiment, the text format, table format, picture format, and hyperlink format are taken as examples for illustration.
[0063] After obtaining the to-be-processed document, the target configuration file is determined by matching among multiple candidate configuration files. Among the multiple candidate configuration files, the number of format types included in each configuration file is different. For example, some candidate configuration files only include one format type, while some candidate configuration files include three format types. During the matching process, the matching is performed according to the format types included in the to-be-processed document. For example, if the to-be-processed document includes two format types, namely text and picture, then the target configuration file matching it also includes the extraction rules and processing rules corresponding to the two format types of text and picture.
[0064] Specifically, the multiple candidate configuration files are stored in a database, and the extraction rules and processing rules included in each candidate configuration file are different. For example, the first candidate configuration file includes the extraction rules and processing rules for text format and table format, while the second candidate configuration file may include the extraction rules and processing rules for picture format and hyperlink format.
[0065] In this embodiment, it is necessary to match the candidate configuration file including all the format types corresponding to the multiple data sets. For example, if the to-be-processed document includes four types: text format, table format, picture format, and hyperlink format, then the candidate configuration file needs to include at least four types to be determined as the target configuration file. Through the above matching method, the target configuration file of the to-be-processed document can be quickly determined, thereby improving the processing efficiency of the to-be-processed document.
[0066] In other embodiments, the target configuration file can also be input by the user while inputting the document to be processed. Specifically, the user can set the extraction rules and processing rules for the document to be processed by themselves, so that the target extracted text after processing can better meet the user's needs.
[0067] Optionally, the document cleaning system further includes a content analysis module. Before the content recognition module matches in multiple candidate configuration files according to the format types corresponding to the multiple data sets and determines the target configuration file, the method further includes:
[0068] The content analysis module determines multiple preset format types, and the multiple preset format types include at least one of the following: text format, table format, picture format, and hyperlink format;
[0069] The content analysis module analyzes each of the multiple preset format types to determine the extraction rules and processing rules matching the multiple preset format types;
[0070] The content analysis module generates multiple candidate configuration files according to the extraction rules and processing rules corresponding to the multiple preset format types. Each candidate configuration file includes at least one preset format type, and at least one preset format type included in any two candidate configuration files is different.
[0071] In this embodiment, the content analysis module is used to analyze and generate multiple candidate configuration files. Among them, the preset format types corresponding to the multiple data sets can be composed of formats such as text format, table format, picture format, and hyperlink format. It should be noted that there can be other formats in other embodiments. In this embodiment, the text format, table format, picture format, and hyperlink format are taken as examples for illustration.
[0072] After obtaining the multiple preset format types, determine the extraction rules and processing rules that match the multiple preset format types. For example, the extraction rules and processing rules for the text format are different from those for the table format. Therefore, after determining the extraction rules and processing rules that match each preset format type respectively, generate multiple candidate configuration files.
[0073] It should be noted that among the multiple candidate configuration files, each candidate configuration file includes at least one preset format type, and at least one preset format type included in any two candidate configuration files is different, so as to ensure that each document to be processed can definitely match the target configuration file in the multiple candidate configuration files.
[0074] Optionally, when the format type is text format, the extraction rule includes a text extraction rule, and the text extraction rule includes: extracting text from the background content, deletable text, deleted text, and non-deletable text in the data set;
[0075] When the format type is table format, the extraction rule includes a table extraction rule, and the table extraction rule includes: extracting text from the tables in the data set;
[0076] When the format type is picture format, the extraction rule includes a picture extraction rule, and the picture extraction rule includes: extracting text from the deletable pictures and non-deletable pictures in the data set;
[0077] When the format type is hyperlink format, the extraction rule includes a hyperlink extraction rule, and the hyperlink extraction rule includes: extracting text from the deletable hyperlinks and non-deletable hyperlinks in the data set.
[0078] Specifically, when the format type is text format, the extraction rule includes a text extraction rule, and the text extraction rule includes extracting text from the background content in the data set, specifically including: specifying keywords A1, A2... or An and content attributes C1, C2... or Cn, this level, and all content within this level range. It also includes extracting deletable text, specifically including: specifying keywords B1, B2... or Bn and content attributes C1, C2... or Cn, this level, and all content within this level range. It also includes extracting deleted text, specifically including: all content with "strikethrough" in the content attribute and content attributes C1, C2... or Cn. It also includes extracting non-deletable text, specifically including: all other content except the text in the background content, deletable text, and deleted text.
[0079] When the format type is table format, the extraction rule includes a table extraction rule, and the table extraction rule includes identifying all text content in the table.
[0080] When the format type is picture format, the extraction rule includes a picture extraction rule, and the picture extraction rule includes extracting text content from non-deletable pictures. It also includes extracting text content from deletable pictures.
[0081] When the format type is hyperlink format, the extraction rule includes a hyperlink extraction rule, and the hyperlink extraction rule includes extracting text content from non-deletable hyperlinks. It also includes extracting text content from deletable hyperlinks.
[0082] It should be noted that a table is essentially a content block composed of text, pictures, and hyperlinks. When processing the table content, according to the table processing logic, it is also necessary to identify the type of the organized content, and then process it according to the processing logic of the corresponding type.
[0083] In this embodiment, by correspondingly processing document contents in different formats, the extracted content can be made more accurate, and the document extraction efficiency is improved.
[0084] Optionally, when the format type is text format, the processing rule includes a text processing rule, and the text processing rule includes: discarding the text extraction content corresponding to the text, the deletable text, and the deleted text in the background content, and retaining the text extraction content corresponding to the non-deletable text;
[0085] When the format type is table format, the processing rule includes a table processing rule, and the table processing rule includes: combining the text extraction content corresponding to the table;
[0086] When the format type is picture format, the processing rule includes a picture processing rule, and the picture processing rule includes: discarding the text extraction content corresponding to the deletable picture, and retaining the text extraction content corresponding to the non-deletable picture;
[0087] When the format type is hyperlink format, the processing rule includes a hyperlink processing rule, and the hyperlink processing rule includes: discarding the text extraction content corresponding to the deletable hyperlink, and retaining the text extraction content corresponding to the deletable hyperlink.
[0088] Specifically, when the format type is text format, the processing rule includes a text processing rule, and the text processing rule includes: discarding the text extraction content corresponding to the text, the deletable text, and the deleted text in the background content, and retaining the text extraction content corresponding to the non-deletable text. Among them, the text extraction content corresponding to the non-deletable text also needs to retain its corresponding sorting serial number for combination.
[0089] When the format type is in tabular format, the processing rules include tabular processing rules, and the tabular processing rules include: scanning the first row to obtain the titles of each column; then scanning the content of each column row by row in the order from top to bottom and from left to right, performing content recognition and content processing on the content (text, pictures, hyperlinks) of each column to obtain the cleaned content; then organizing the content of each row of the table again in the form of "title: content" for each column, concatenating them with semicolons, and outputting the organized content with each row as a paragraph.
[0090] When the format type is in picture format, the processing rules include picture processing rules, and the picture processing rules include: calling the encapsulated OCR image recognition algorithm for non-deletable pictures and returning the recognized text content, and discarding the content extracted from deletable pictures.
[0091] When the format type is in hyperlink format, the processing rules include hyperlink processing rules, and the hyperlink processing rules include: recursively calling the business logic for non-deletable hyperlinks, cleaning the content of the document corresponding to the hyperlink, and returning the plain text content without further processing of the hyperlinks in the content.
[0092] In other implementable embodiments, the server side is a document cleaning system, and the document cleaning system includes four layers: a front-end page layer, a business interface layer, a business logic layer, and a data layer. Among them, the business logic layer is divided into 6 parts: json configuration file management, HTML content crawling and parsing in HTLM documents, json configuration parsing, content type recognition, content processing, and content reorganization.
[0093] Specifically, in this embodiment, the document to be processed is an HTLM document. The front-end page layer is mainly used to manage the extraction and processing configuration of the document to be processed. Users can generate an extraction request for the document to be processed in the front-end page, and can modify the default configuration of the extraction and processing of the document to be processed, add and delete custom configurations of HTLM documents, and the data is stored in a mysql database and maintained using a configuration table.
[0094] The business interface layer is mainly used for configuration management and triggering of cleaning logic. Specifically, it is used to receive the document to be processed by the user, generate a system task according to the document to be processed, and after triggering the system task, the document cleaning system will perform corresponding processing on the document to be processed.
[0095] The json configuration file management is used to manage multiple candidate configuration files. Among them, each of the candidate configuration files includes at least one preset format type, and at least one preset format type included in any two candidate configuration files is different.
[0096] HTML content crawling and parsing mainly involve extracting and parsing the document to be processed. Specifically, the extraction and parsing include the extraction and parsing of the text, specific attributes, block types, and block numbers of content blocks, etc.
[0097] Json configuration parsing mainly involves obtaining the extraction rules, corresponding processing rules, and custom parameters for different detailed types of content, and determining the custom recognition and processing logic for the corresponding type of content.
[0098] Content type recognition mainly involves extracting the detailed content such as the text, specific attributes, block types, and block numbers of the content blocks obtained from the document to be processed according to the custom configuration.
[0099] Content processing mainly involves processing different detailed types of content. For example: discarding or retaining text content, scanning and organizing methods for table content, recognizing or discarding image content, recognizing or discarding the content of links, etc.
[0100] Content reorganization is to reorganize the processed different detailed contents into a new pure text document according to certain rules and then output it.
[0101] The data layer mainly encapsulates the operation methods for multiple candidate configuration files, including operations such as configuration acquisition, addition, update, and deletion.
[0102] As Figure 2 shown, Figure 2 is the process schematic diagram in the embodiment. Specifically, it includes the following steps: The front-end page layer mainly receives the extraction request of the document to be processed input by the user, reads the content of the document to be processed, and determines the target configuration file. After the business interface layer triggers the extraction request of the document to be processed, according to the document address and the user's configuration id, it extracts the target configuration file corresponding to the document to be processed. Specifically, the system identifies the content format in the document to be processed through the crawler algorithm, extracts the content of different content blocks, extracts the text, specific attributes (thickness, font size, color, underline, strikethrough, italic), block types, block numbers, etc., and stores the extracted content list in the memory. According to the extraction rules, processing rules, and corresponding parameters included in the target configuration file determined by the business logic layer, and according to the extraction rules and processing rules, it processes the document to be processed. For example, it scans the extracted content list and uses the extraction rules and corresponding parameters for different types of content once to extract and determine the detailed type of the corresponding content. According to the extracted content list, the corresponding detailed type, the corresponding processing rules and parameters in the configuration file, it processes the extracted content to obtain the processed result. The data layer grades and segments the processed result according to the needs of the content, reorganizes it, and outputs the final cleaning result.
[0103] This application classifies the document content of the document to be processed by format type, obtains multiple data sets, determines extraction rules and processing rules for processing the multiple data sets respectively according to the target configuration file, completes text extraction and screening of the multiple data sets, and obtains the target extraction text, thereby realizing the recognition of multiple formats in the document to be processed and improving the extraction effect of the document content.
[0104] An embodiment of this application provides an apparatus for extracting document content, as Figure 3 shown. The apparatus 300 for extracting document content includes:
[0105] A classification module 310, configured to classify the document content of the document to be processed, obtain multiple data sets, and each data set is a set of document content of a format type;
[0106] An acquisition module 320, configured to acquire a target configuration file that matches the document to be processed;
[0107] An extraction module 330, configured to perform content extraction on the document content in the data set according to the target configuration file to obtain a first text content corresponding to each document content, and perform screening processing on the first text content according to the target configuration file to obtain a second text content, where the target configuration file includes extraction rules and processing rules corresponding to each format type included in the document to be processed, the first text content is the text obtained by performing content extraction according to the extraction rules corresponding to the data set in the target configuration file, and the second text content is the text obtained by performing screening processing according to the processing rules corresponding to the data set in the target configuration file;
[0108] A combination module 340, configured to combine the second text contents corresponding to the multiple data sets to obtain a target extraction text.
[0109] Optionally, the document cleaning system includes a content parsing module. The content parsing module traverses multiple tag contents in the document to be processed to obtain multiple document contents, and the multiple document contents correspond one-to-one to the multiple tag contents. The tag content is used to represent the sorting level of the corresponding text content in the document to be processed, and the document to be processed is a Feishu document;
[0110] The content parsing module classifies the multiple document contents based on the format types corresponding to the multiple document contents to obtain multiple data sets.
[0111] Optionally, the document cleaning system further includes a content organization module, which generates a serial number corresponding to each of the document contents based on the multiple tag contents, and the serial number is used to represent the position of the document content in the to-be-processed document; wherein, the data set includes at least one document content and the serial number corresponding to each of the document contents;
[0112] The combining the second text contents corresponding to the multiple data sets to obtain the target extraction text includes:
[0113] The content organization module determines the serial number corresponding to each of the second text contents based on the serial number corresponding to each of the document contents;
[0114] The content organization module sequentially combines the second text contents corresponding to the multiple data sets according to the serial number corresponding to each of the second text contents to obtain the target extraction text.
[0115] The document cleaning system further includes a content recognition module, which determines the format type corresponding to the multiple data sets;
[0116] The content recognition module matches in multiple candidate configuration files according to the format type corresponding to the multiple data sets to determine a target configuration file, the target configuration file is any one of the multiple candidate configuration files, and the target configuration file includes extraction rules and processing rules for the format type corresponding to the multiple data sets.
[0117] The document cleaning system further includes a content parsing module, which determines multiple preset format types, and the multiple preset format types include at least one of the following: text format, table format, picture format, and hyperlink format;
[0118] The content parsing module analyzes each of the multiple preset format types respectively to determine extraction rules and processing rules matching the multiple preset format types;
[0119] The content parsing module generates multiple candidate configuration files according to the extraction rules and processing rules corresponding to the multiple preset format types, each candidate configuration file includes at least one preset format type, and at least one preset format type included in any two candidate configuration files is different.
[0120] Optionally, in the case where the format type is text format, the extraction rule includes a text extraction rule, and the text extraction rule includes: extracting text from the text, deletable text, deleted text, and non-deletable text in the background content of the data set;
[0121] When the format type is a table format, the extraction rule includes a table extraction rule, and the table extraction rule includes: extracting text from the tables in the data set;
[0122] When the format type is a picture format, the extraction rule includes a picture extraction rule, and the picture extraction rule includes: extracting text from the deletable pictures and non-deletable pictures in the data set;
[0123] When the format type is a hyperlink format, the extraction rule includes a hyperlink extraction rule, and the hyperlink extraction rule includes: extracting text from the deletable hyperlinks and non-deletable hyperlinks in the data set.
[0124] Optionally, when the format type is a text format, the processing rule includes a text processing rule, and the text processing rule includes: discarding the text extraction content corresponding to the text, the deletable text, and the deleted text in the background content, and retaining the text extraction content corresponding to the non-deletable text;
[0125] When the format type is a table format, the processing rule includes a table processing rule, and the table processing rule includes: combining the text extraction content corresponding to the table into text;
[0126] When the format type is a picture format, the processing rule includes a picture processing rule, and the picture processing rule includes: discarding the text extraction content corresponding to the deletable pictures and retaining the text extraction content corresponding to the non-deletable pictures;
[0127] When the format type is a hyperlink format, the processing rule includes a hyperlink processing rule, and the hyperlink processing rule includes: discarding the text extraction content corresponding to the deletable hyperlinks and retaining the text extraction content corresponding to the deletable hyperlinks.
[0128] This application classifies the document content of the document to be processed by the format type, obtains multiple data sets, determines the extraction rules and processing rules for processing the multiple data sets respectively according to the target configuration file, completes the text extraction and screening of the multiple data sets, and obtains the target extraction text, thereby realizing the recognition of multiple formats in the document to be processed and improving the extraction effect of the document content.
[0129] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of this application, as Figure 4As shown, the electronic device 400 includes a memory 410 and a processor 420. The number of processors 420 in the electronic device 400 can be one or more. Figure 4 Here, one processor 420 is taken as an example; the memory 410 and the processor 420 in the server can be connected through a bus or other means. Figure 4 Here, taking the connection through the bus as an example.
[0130] The memory 410, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the member push method in the embodiments of the present application. The processor 420 executes various functional applications and data processing of the server / terminal / server by running the software programs, instructions, and modules stored in the memory 410, that is, implements the above-mentioned method for extracting document content.
[0131] Among them, the processor 420 is used to run the computer program stored in the memory 410 to implement the following steps:
[0132] Classify the document content of the document to be processed to obtain a plurality of data sets, and each data set is a set of document content of a format type;
[0133] Obtain a target configuration file that matches the document to be processed;
[0134] For each data set, extract the content of the document content in the data set according to the target configuration file to obtain the first text content corresponding to each document content, and screen and process the first text content according to the target configuration file to obtain the second text content. Among them, the target configuration file includes the extraction rules and processing rules corresponding to each format type included in the document to be processed. The first text content is the text obtained by extracting the content according to the extraction rules corresponding to the data set in the target configuration file, and the second text content is the text obtained by screening and processing according to the processing rules corresponding to the data set in the target configuration file;
[0135] Combine the second text contents corresponding to the plurality of data sets to obtain the target extraction text.
[0136] Optionally, the document cleaning system includes a content parsing module. The classification of the document content of the document to be processed to obtain a plurality of data sets includes:
[0137] The content parsing module traverses multiple tag contents in the to-be-processed document to obtain multiple document contents. The multiple document contents correspond one-to-one to the multiple tag contents. The tag contents are used to represent the sorting levels of the corresponding text contents in the to-be-processed document, and the to-be-processed document is a Feishu document;
[0138] The content parsing module classifies the multiple document contents based on the format types corresponding to the multiple document contents to obtain multiple data sets.
[0139] Optionally, the document cleaning system further includes a content organization module, and the method further includes:
[0140] The content organization module generates a sorting serial number corresponding to each of the document contents based on the multiple tag contents. The sorting serial number is used to characterize the position of the document content in the to-be-processed document; wherein, the data set includes at least one document content and the sorting serial number corresponding to each of the document contents;
[0141] The combining the second text contents corresponding to the multiple data sets to obtain the target extraction text includes:
[0142] The content organization module determines the sorting serial number corresponding to each of the second text contents based on the sorting serial number corresponding to each of the document contents;
[0143] The content organization module sequentially combines the second text contents corresponding to the multiple data sets according to the sorting serial number corresponding to each of the second text contents to obtain the target extraction text.
[0144] Optionally, the document cleaning system further includes a content recognition module. The obtaining the target configuration file matching the to-be-processed document includes:
[0145] The content recognition module determines the format types corresponding to the multiple data sets;
[0146] The content recognition module matches in multiple candidate configuration files according to the format types corresponding to the multiple data sets to determine the target configuration file. The target configuration file is any one of the multiple candidate configuration files, and the target configuration file includes the extraction rules and processing rules of the format types corresponding to the multiple data sets.
[0147] Optionally, the document cleaning system further includes a content parsing module. Before the content recognition module matches in multiple candidate configuration files according to the format types corresponding to the multiple data sets to determine the target configuration file, the method further includes:
[0148] The content parsing module determines multiple preset format types, and the multiple preset format types include at least one of the following: text format, table format, picture format, and hyperlink format;
[0149] The content parsing module analyzes the multiple preset format types respectively, and determines extraction rules and processing rules matching the multiple preset format types;
[0150] The content parsing module generates multiple candidate configuration files according to the extraction rules and processing rules corresponding to the multiple preset format types. Each candidate configuration file includes at least one preset format type, and at least one preset format type included in any two candidate configuration files is different.
[0151] Optionally, when the format type is text format, the extraction rule includes a text extraction rule, and the text extraction rule includes: performing text extraction on the text, deletable text, deleted text, and non-deletable text in the background content of the data set;
[0152] When the format type is table format, the extraction rule includes a table extraction rule, and the table extraction rule includes: performing text extraction on the tables in the data set;
[0153] When the format type is picture format, the extraction rule includes a picture extraction rule, and the picture extraction rule includes: performing text extraction on the deletable pictures and non-deletable pictures in the data set;
[0154] When the format type is hyperlink format, the extraction rule includes a hyperlink extraction rule, and the hyperlink extraction rule includes: performing text extraction on the deletable hyperlinks and non-deletable hyperlinks in the data set.
[0155] Optionally, when the format type is text format, the processing rule includes a text processing rule, and the text processing rule includes: discarding the text extraction contents corresponding to the text, deletable text, and deleted text in the background content, and retaining the text extraction content corresponding to the non-deletable text;
[0156] When the format type is table format, the processing rule includes a table processing rule, and the table processing rule includes: performing text combination on the text extraction content corresponding to the tables;
[0157] When the format type is picture format, the processing rule includes a picture processing rule, and the picture processing rule includes: discarding the text extraction content corresponding to the deletable pictures, and retaining the text extraction content corresponding to the non-deletable pictures;
[0158] When the format type is the hyperlink format, the processing rules include hyperlink processing rules, and the hyperlink processing rules include: discarding the extracted content of the text corresponding to the deletable hyperlink, and retaining the extracted content of the text corresponding to the deletable hyperlink.
[0159] This application classifies the document content of the document to be processed by the format type, obtains a plurality of data sets, determines the extraction rules and processing rules for processing the plurality of data sets respectively according to the target configuration file, completes the text extraction and screening of the plurality of data sets, and obtains the target extraction text, thereby realizing the recognition of multiple formats in the document to be processed and improving the extraction effect of the document content.
[0160] The computer-readable storage medium of the embodiments of this application can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.
[0161] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.
[0162] The program code contained on the storage medium can be transmitted by any suitable medium, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0163] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or terminal. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by connecting through the Internet using an Internet service provider).
[0164] Another embodiment of this application provides a computer program product. The computer program product is stored in a storage medium and is executed by at least one processor to implement each process of the above-mentioned embodiment of the method for extracting document content, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0165] Note that the above is only a preferred embodiment of this application and the technical principles applied. Those skilled in the art will understand that this application is not limited to the specific embodiments described here. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of this application. Therefore, although this application has been described in more detail through the above embodiments, this application is not limited to the above embodiments. Without departing from the concept of this application, it can also include more other equivalent embodiments, and the scope of this application is determined by the scope of the appended claims.
Claims
1. A method for extracting document content, characterized in that Applied to a document cleaning system, the method includes: Classifying the document content of the document to be processed to obtain a plurality of data sets, where each data set is a set of document content of a format type; Obtaining a target configuration file that matches the document to be processed; For each data set, extracting the content of the document content in the data set according to the target configuration file to obtain the first text content corresponding to each document content, and screening and processing the first text content according to the target configuration file to obtain the second text content. Among them, the target configuration file includes the extraction rules and processing rules corresponding to each format type included in the document to be processed. The first text content is the text obtained by extracting the content according to the extraction rules corresponding to the data set in the target configuration file, and the second text content is the text obtained by screening and processing according to the processing rules corresponding to the data set in the target configuration file; Combining the second text contents corresponding to the plurality of data sets to obtain a target extraction text.
2. The method according to claim 1, wherein The document cleaning system includes a content parsing module. Classifying the document content of the document to be processed to obtain a plurality of data sets includes: The content parsing module traverses a plurality of tag contents in the document to be processed to obtain a plurality of document contents. The plurality of document contents correspond one-to-one to the plurality of tag contents. The tag content is used to represent the sorting level of the corresponding text content in the document to be processed. The document to be processed is a Feishu document; The content parsing module classifies the plurality of document contents based on the format types corresponding to the plurality of document contents to obtain a plurality of data sets.
3. The method according to claim 2, wherein The document cleaning system further includes a content organization module. The method further includes: The content organization module generates a sorting number corresponding to each document content based on the plurality of tag contents. The sorting number is used to characterize the position of the document content in the document to be processed; where the data set includes at least one document content and the sorting number corresponding to each document content; The combining the second text contents corresponding to the plurality of data sets to obtain a target extraction text includes: The content organization module determines the sorting number corresponding to each second text content based on the sorting number corresponding to each document content; The content organization module sequentially combines the second text contents corresponding to the plurality of data sets according to the sorting number corresponding to each second text content to obtain the target extraction text.
4. The method according to claim 1, wherein The document cleaning system further includes a content recognition module. Obtaining a target configuration file that matches the document to be processed includes: The content recognition module determines the format types corresponding to the plurality of data sets; The content recognition module matches among multiple candidate configuration files according to the format types corresponding to the multiple data sets to determine a target configuration file, where the target configuration file is any one of the multiple candidate configuration files, and the target configuration file includes extraction rules and processing rules for the format types corresponding to the multiple data sets.
5. The method according to claim 4, characterized in that, The document cleaning system further includes a content parsing module. Before the content recognition module matches among multiple candidate configuration files according to the format types corresponding to the multiple data sets to determine a target configuration file, the method further includes: The content parsing module determines multiple preset format types, and the multiple preset format types include at least one of the following: text format, table format, picture format, and hyperlink format; The content parsing module analyzes the multiple preset format types respectively to determine extraction rules and processing rules matched by the multiple preset format types; The content parsing module generates multiple candidate configuration files according to the extraction rules and processing rules corresponding to the multiple preset format types. Each candidate configuration file includes at least one preset format type, and at least one preset format type included in any two candidate configuration files is different.
6. The method according to claim 5, characterized in that When the format type is text format, the extraction rule includes a text extraction rule, and the text extraction rule includes: extracting text from the text, deletable text, deleted text, and non-deletable text in the background content of the data set; When the format type is table format, the extraction rule includes a table extraction rule, and the table extraction rule includes: extracting text from the table in the data set; When the format type is picture format, the extraction rule includes a picture extraction rule, and the picture extraction rule includes: extracting text from the deletable pictures and non-deletable pictures in the data set; When the format type is hyperlink format, the extraction rule includes a hyperlink extraction rule, and the hyperlink extraction rule includes: extracting text from the deletable hyperlinks and non-deletable hyperlinks in the data set.
7. The method according to claim 6, wherein When the format type is text format, the processing rule includes a text processing rule, and the text processing rule includes: discarding the text extraction content corresponding to the text, deletable text, and deleted text in the background content, and retaining the text extraction content corresponding to the non-deletable text; When the format type is table format, the processing rule includes a table processing rule, and the table processing rule includes: combining the text extraction content corresponding to the table into text; When the format type is picture format, the processing rule includes a picture processing rule, and the picture processing rule includes: discarding the text extraction content corresponding to the deletable pictures and retaining the text extraction content corresponding to the non-deletable pictures; When the format type is the hyperlink format, the processing rules include hyperlink processing rules, and the hyperlink processing rules include: discarding the extracted content of the text corresponding to the deletable hyperlink, and retaining the extracted content of the text corresponding to the deletable hyperlink.
8. An apparatus for extracting document content, characterized in that, The device includes: a classification module, configured to classify the document content of the document to be processed to obtain a plurality of data sets, and each data set is a set of document content of a format type; an acquisition module, configured to acquire a target configuration file matching the document to be processed; an extraction module, configured to extract the content of the document content in the data set according to the target configuration file to obtain a first text content corresponding to each document content, and perform screening processing on the first text content according to the target configuration file to obtain a second text content, where the target configuration file includes extraction rules and processing rules corresponding to each format type included in the document to be processed, the first text content is the text obtained by extracting content according to the extraction rules corresponding to the data set in the target configuration file, and the second text content is the text obtained by performing screening processing according to the processing rules corresponding to the data set in the target configuration file; a combination module, configured to combine the second text contents corresponding to the plurality of data sets to obtain a target extraction text.
9. An electronic device, characterized in that, It includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method for extracting document content according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that, A program or instruction is stored on the readable storage medium. When the program or instruction is executed by the processor, the steps of the method for extracting document content according to any one of claims 1 to 7 are implemented.
11. A computer program product, characterized in that, The computer program product is stored in a storage medium. The computer program product is executed by at least one processor to implement the steps in the method for extracting document content according to any one of claims 1 to 7.
Citation Information
Cited By
Defective file and image number association method based on multi-format analysis
CN120873217A
A defect file and image number association method based on multi-format analysis
CN120873217B