File extraction method and device
By identifying tables and text areas in the file, storing layout information using location indexes and filling content, the commonality and efficiency of file extraction methods in the prior art are solved, and efficient file extraction across formats is achieved.
Patent Information
- Application Number
- CN202510558864.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-05
AI Technical Summary
In the prior art, file extraction methods need to be customized and developed according to different format structures, resulting in poor versatility and low efficiency.
By determining the position index of the table area and text area in the file to be extracted, the layout information is stored, and the text content is filled into the text part and the table content is filled into the table structure to achieve format-independent file extraction.
It improves the versatility and efficiency of file extraction, and does not require separate development for different formats, and is suitable for automated extraction of multiple file formats.
Smart Images

Figure CN120430288A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a file extraction method and device. Background Art
[0002] In the existing technology, file extraction is typically performed using third-party packages or software. However, these methods require customized development based on the different structure of the file content, such as text and table content. This results in poor versatility and reduced efficiency. Summary of the Invention
[0003] In view of this, the embodiments of the present disclosure at least provide a file extraction method, device, electronic device and storage medium, which can improve the versatility of file extraction, eliminate the need for development for different formats of content, and improve file extraction efficiency.
[0004] In a first aspect, an embodiment of the present disclosure provides a file extraction method, comprising:
[0005] Determine the table information corresponding to the table area in the file to be extracted, and extract the text content of the text area in the file to be extracted; the table information includes the table structure and the table content;
[0006] Based on the position index of the table area and the text area in the file to be extracted, the layout information corresponding to the file to be extracted is stored in the target storage location;
[0007] The text content is filled into the text part of the layout information, and the table content is filled into the table structure corresponding to the table part of the layout information to obtain the file extraction result.
[0008] Optionally, determining table information corresponding to a table area in the file to be extracted includes:
[0009] Traverse the area blocks corresponding to the file to be extracted to determine the table line information in the file to be extracted;
[0010] Construct cell information based on table line information; the cell information includes text data and element type corresponding to each cell;
[0011] Based on the table line information and cell information, the table information corresponding to the table area in the file to be extracted is determined.
[0012] Optionally, before extracting the text content of the file to be extracted, the following steps are further included:
[0013] Based on the predetermined page height of the file to be extracted and the position information of the current processing area, determining whether the current processing area is a text area;
[0014] In response to the current processing area being a text area, the area range of the text area is determined, and the corresponding vertical coordinate of the text area in the page is used as a position index of the text area.
[0015] Optionally, filling the text content into the text portion of the layout information, and filling the table content into the table structure corresponding to the table portion of the layout information, includes:
[0016] Traverse the position index of the text area and determine whether the target position information is included in the position index;
[0017] In response to the target location information not being included in the location index, determining that the target location is a table portion in the layout information, and filling the table content corresponding to the target location into the table structure of the layout information;
[0018] In response to the target position information being included in the position index, the target position is determined to be a text portion in the layout information, and the text content corresponding to the target position is filled into the text portion of the layout information.
[0019] Optionally, the method further comprises:
[0020] Determine the region information and position index corresponding to other regions in the file to be extracted;
[0021] Based on the position indexes corresponding to other regions, the region information is filled into other parts corresponding to the layout information of the file to be extracted to obtain the file extraction result; wherein the other regions include image regions and symbol regions.
[0022] Optionally, after obtaining the file extraction result, the following steps are also included:
[0023] Determine the type of each file object in the file extraction result;
[0024] Output the file extraction results according to the output format pre-selected by the user and the type of each file object.
[0025] In a second aspect, an embodiment of the present disclosure provides a file extraction device, comprising:
[0026] An extraction module is used to determine the table information corresponding to the table area in the file to be extracted, and extract the text content of the text area in the file to be extracted; the table information includes the table structure and the table content;
[0027] A storage module, configured to store layout information corresponding to the file to be extracted to a target storage location based on a position index of a table area and a text area in the file to be extracted;
[0028] The filling module is used to fill the text content into the text part of the layout information, and fill the table content into the table structure corresponding to the table part of the layout information to obtain the file extraction result.
[0029] In a third aspect, an embodiment of the present disclosure further provides an electronic device comprising: a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the computer device is running, the processor and the memory communicate through the bus, and when the machine-readable instructions are executed by the processor, the steps in the above-mentioned first aspect or any optional implementation of the first aspect are performed.
[0030] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned first aspect or any optional implementation of the first aspect are executed.
[0031] In a fifth aspect, an embodiment of the present disclosure further provides a computer program product, including a computer program, which implements the method of any of the above embodiments when executed by a processor.
[0032] Any of the above aspects or any implementation methods of any of the aspects, by identifying the table area and text area in the file to be extracted, extracts the table information and text content respectively, without relying on a specific file format. Secondly, the position index of the table area and the text area is used to store the layout information of the entire file to the target storage location, thereby realizing a format-independent layout representation. In this way, not only the structure of the document is retained, but also its content and structure are separated, so that subsequent filling operations can be flexibly adapted to files from different sources. Then, by filling the extracted text content into the text part of the layout information, and filling the table content into the table structure corresponding to the table part of the layout information, a complete file extraction result is obtained. Since it is based on a general layout index storage and filling strategy, rather than customizing the parsing logic for a specific format, it can be applied to the extraction of multiple formats of content in the file, reducing the development cost of format adaptation. In addition, the same set of extraction processes can be used for documents of different formats, without the need to develop a separate parsing solution for each format, thereby improving the degree of automation and efficiency of file extraction.
[0033] The effects of the above-mentioned text extraction device, electronic device and storage medium can be found in the description of the above-mentioned text extraction method, which will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the embodiments. The drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to illustrate the technical solutions of the present disclosure. It should be understood that the following drawings only illustrate certain embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, other relevant drawings can be obtained based on these drawings without inventive effort.
[0035] Figure 1 A flowchart of a text extraction method provided by an embodiment of the present disclosure is shown;
[0036] Figure 2 A flowchart of a text extraction method provided by an embodiment of the present disclosure is shown;
[0037] Figure 3 A schematic diagram of a process of a file preheating method provided by an embodiment of the present disclosure is shown;
[0038] Figure 4 A schematic flow chart of a filling method provided by an embodiment of the present disclosure is shown;
[0039] Figure 5 A schematic diagram of a text extraction device provided by an embodiment of the present disclosure is shown;
[0040] Figure 6 An exemplary system architecture is shown in which embodiments of the present disclosure may be applied;
[0041] Figure 7 A schematic diagram of the structure of a computer system of a terminal device or server used to implement an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0042] In order to make the purpose, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. The components of the embodiments of the present disclosure generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure provided in the drawings is not intended to limit the scope of the disclosure for which protection is sought, but merely represents selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present disclosure.
[0043] It should be noted that in the technical solution of the present invention, the collection, use, storage, sharing and transfer of user personal information involved are in compliance with the provisions of relevant laws and regulations, and it is necessary to inform the user and obtain the user's consent or authorization. When applicable, the user's personal information is de-identified and / or anonymized and / or encrypted.
[0044] The above problems and solutions are the results obtained by the inventors after practice and careful research. The discovery process of the above problems and the solutions proposed for the above problems should be the contributions made by the inventors to this disclosure during the disclosure process.
[0045] The technical solutions in the present disclosure will be clearly and completely described below in conjunction with the drawings in the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. The components of the present disclosure generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure provided in the drawings is not intended to limit the scope of the present disclosure for protection, but merely represents selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present disclosure.
[0046] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.
[0047] To facilitate understanding of this embodiment, a text extraction method disclosed in an embodiment of the present disclosure is first described in detail. The text extraction method provided in the embodiment of the present disclosure is generally executed by a computer device with certain computing capabilities. The computer device includes, for example, a terminal device, a server, or other processing device. The terminal device may be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, an in-vehicle device, a wearable device, etc. In some possible implementations, the text extraction method may be implemented by a processor invoking computer-readable instructions stored in a memory.
[0048] See also Figure 1 FIG. 1 is a flowchart of a text extraction method provided by an embodiment of the present disclosure, and the method includes S101 to S103, wherein:
[0049] S101: Determine table information corresponding to a table area in a file to be extracted, and extract text content in a text area in the file to be extracted.
[0050] In the embodiment of the present disclosure, Figure 2 As shown, users can specify files to be extracted by uploading them themselves; or, users can also specify files to be extracted by selecting the directory where the corresponding files are located. The number of files to be extracted can be one or more. When there are multiple files to be extracted, the extraction process of each file to be extracted can be executed in parallel. However, considering the limitation of system resources, an upper limit on the number of files to be extracted can be set in advance. For each file to be extracted, the file pages inside it can be processed in parallel or page by page. The embodiment of the present disclosure does not make specific restrictions on the specific number of files to be extracted and the number of pages inside the files to be extracted. In actual applications, corresponding thresholds can be set according to actual conditions and needs to ensure stable operation of the system.
[0051] In the embodiment of the present disclosure, after determining the files to be extracted, the files to be extracted can be preheated. File preheating refers to processing or loading the files before the files are formally extracted to improve the efficiency of subsequent access. Figure 3As shown, file pre-warming in the disclosed embodiments can involve dividing the entire file to be extracted into regions. Specifically, the total number of pages in the file to be extracted can be first obtained. For each page in the file to be extracted, the page width and page height can be obtained. The page area of each page can then be calculated based on the page width and page height. A determination is then made as to whether the calculated page area exceeds a pre-defined threshold. The threshold here is determined based on the threshold of the type used to store the page. For example, if the page is stored using the int type, when the page area exceeds the threshold corresponding to that type, integer overflow may occur, leading to an exception or an erroneous calculation result. If the page area of the file to be extracted exceeds the threshold, a prompt such as "File Not compliant" can be returned, and the file extraction process terminates. If the page area of the file to be extracted is less than or equal to the threshold, the file can be instantiated based on the page height and page width. Specifically, a buffered image can be created based on the width and height of the current page, and a drawing object can be generated simultaneously. To prevent overlapping drawn content, the drawing object can first be used to clear the page area. A rectangular area can be constructed with the lower left corner as the vertex and the page width and page height as the clearing operation. The current page can then be re-rendered to ensure complete and accurate presentation of the page content. Then, the boundary information of the current page can be obtained and encapsulated using Java graphic shapes, ultimately obtaining a set of boundary information for all pages in the entire file to be extracted. This ensures that all pages are redrawn with the upper left corner as the reference point. After obtaining the boundary information set of the file to be extracted, the set can be traversed to extract the boundary object of each page. The height of the page is calculated based on the boundary information, and the height of the current page is used as the key and the height of the next page as the value. A key-value pair model is constructed and encapsulated into a tuple structure. This establishes a mapping relationship between page heights, ensuring that subsequent steps such as text or table extraction and page redrawing can be performed according to the correct layout, avoiding problems such as content misalignment, omissions, or paging errors.
[0052] In the embodiment of the present disclosure, while traversing the boundary information set to obtain the key-value pair model of all pages, the table area and table information in the file to be extracted can be determined. Specifically, determining the table information corresponding to the table area in the file to be extracted includes: traversing the area block corresponding to the file to be extracted to determine the table line information in the file to be extracted; constructing cell information based on the table line information; the cell information includes text data and element types corresponding to each cell; and determining the table information corresponding to the table area in the file to be extracted based on the table line information and cell information.
[0053] In a specific implementation, the execution subject of the embodiment of the present disclosure can continue to subdivide the page area to obtain multiple area blocks corresponding to the current page until it can no longer be divided further. Each area block can then be traversed and checked to see if its height is less than 1 pixel. If the height meets this condition, the area block can be determined as a horizontal line of the table. After determining the horizontal line of the table, the table line object can be defined based on the horizontal coordinate, vertical coordinate, and estimated horizontal coordinate plus width of the area block on the overall page, which may include the starting position, end position, and vertical coordinate value of the table horizontal line. When a table horizontal line is detected, the execution subject of the embodiment of the present disclosure can record its vertical coordinate value in the vertical coordinate field of the table line entity object and add the object to the table line information. Next, the execution subject of the embodiment of the present disclosure can create a cell information model based on the table line set. The cell information model here can be used to define a table cell list and an element type field; wherein, the cell information in the cell list can include text data in the cell and the element type of the cell. When instantiating a table cell object, the execution subject of the embodiment of the present disclosure can add the table line information to the text data field and set the type of the current element in the element type field, so that when the file content is subsequently extracted, it can be identified that the element belongs to the horizontal line part of the table. After successfully creating the cell information, the execution subject of the embodiment of the present disclosure can separate the table line information and the page area range of the current page mentioned above, and store it in the table information list. When all pages of the document to be extracted are traversed, the table structure and table content of all tables stored in the per-page format are obtained.
[0054] In another possible implementation, a line detection method can be used to determine the rows and columns of a table by detecting horizontal and vertical lines in the page, especially the borders and dividing lines of the table. Alternatively, based on common table typesetting rules, it is also possible to infer which elements belong to the same table through features such as row spacing, column spacing, and alignment. For example, if a row of text is left-aligned, it indicates that it is a column of the table, and so on. It should be noted that the above method of determining the table lines in the table area is only used as an example of a feasible method in the embodiment of the present disclosure, and does not constitute an improper limitation on the present invention. In actual applications, it can be set according to actual conditions or needs. The embodiment of the present disclosure does not make specific limitations on this, and its function shall be realized.
[0055] In the disclosed embodiment, in addition to determining the table information corresponding to the table area in the file to be extracted, it is also necessary to determine the text area in the file to be extracted in order to extract the text content in the text area. Specifically, before extracting the text content of the file to be extracted, the method further includes: determining whether the current processing area is a text area based on the predetermined page height of the file to be extracted and the position information of the current processing area; in response to the current processing area being a text area, determining the area range of the text area, and using the vertical coordinate corresponding to the text area in the page as the position index of the text area.
[0056] In a specific implementation, some basic attributes can be defined first, such as a text index list for recording the position of the text table, a text range list for storing text areas, the current vertical coordinate position of the user to track the current coordinates, and the maximum height and maximum width of each page of the file to be extracted; wherein, the maximum height can be calculated by obtaining the total number of pages and locating the boundary of the last page, that is, the vertical height of the last page. Similarly, the maximum width can be calculated by the boundary information of the last page, that is, the horizontal width of the last page. It should be noted that the above-mentioned pre-defined basic attributes are only used as examples of feasible implementation methods in the embodiments of the present disclosure, and do not constitute an improper limitation on the present invention. In actual applications, the required basic attributes can be defined according to actual conditions and needs. The embodiments of the present disclosure do not make specific limitations on this, and the ability to achieve its functions shall prevail.
[0057] In a specific implementation, after the required basic attributes are defined, since the height range of each page of the file to be extracted is stored in the table information list, the table information list can be traversed to obtain the height range of the currently processed page page by page, and the coordinate conversion is performed in combination with the total height of the file to be extracted. In this process, the height range records the starting height and ending height of the area in the file to be extracted. When processing the file to be extracted, a currently traversed vertical coordinate position can usually also be recorded. When the starting height minus the currently traversed vertical coordinate position is greater than zero, it means that there is a gap between the current starting position and the vertical coordinate position, and there is no table or other identified structured content in this area. Therefore, the gap area can be considered as a pure text area. After determining the text area, the corresponding vertical coordinate of the text area in the page can be recorded as the position of the text area, which is used as a position index of the text area for subsequent content extraction.
[0058] In another possible implementation, it is also possible to determine whether the area is a pure text area by calculating the number of text characters within a unit area. If the number of characters in a certain area is much larger than that in the surrounding area, the area can be identified as a text area. Alternatively, it is also possible to determine whether the corresponding area is a text area by calculating the spacing between adjacent text lines. It should be noted that the above method for determining the text area is only used as an example of a feasible implementation method in the embodiment of the present disclosure, and does not constitute an improper limitation on the present invention. In actual applications, a suitable text area determination method can be selected according to actual needs and circumstances. The embodiment of the present disclosure does not make specific limitations on this, and is based on the ability to achieve its functions.
[0059] In addition, the embodiments of the present disclosure do not make any specific restrictions on the method of extracting text content in the text area. In actual applications, optical character recognition, third-party plug-in recognition, or database query methods can be selected according to actual needs and circumstances to extract the text content in the text area. The embodiments of the present disclosure do not make any specific restrictions on this, and the ability to achieve its functions shall prevail.
[0060] S102: Based on the position indexes of the table area and the text area in the file to be extracted, the layout information corresponding to the file to be extracted is stored in a target storage location.
[0061] As described above, after determining the table area and the text area, the layout information of the file to be extracted can be determined based on the position index of the table area and the text area in the page of the file to be extracted and saved to the target storage location, thereby completing the file preheating step. The detailed method and process have been explained above and will not be repeated here.
[0062] S103: Fill the text content into the text part of the layout information, and fill the table content into the table structure corresponding to the table part of the layout information to obtain a file extraction result.
[0063] In an embodiment of the present disclosure, filling text content into the text part of the layout information, and filling table content into the table structure corresponding to the table part of the layout information, includes: traversing the position index of the text area, and determining whether the target position information is included in the position index; in response to the target position information not being included in the position index, determining that the target position is the table part in the layout information, and filling the table content corresponding to the target position into the table structure of the layout information; in response to the target position information being included in the position index, determining that the target position is the text part in the layout information, and filling the text content corresponding to the target position into the text part of the layout information.
[0064] In specific implementation, Figure 4As shown, a two-tuple object can be encapsulated for each page. The key-value model of this object can consist of value_text and value_table, representing text and table content, respectively. After traversing the entire file to be extracted, a two-tuple list can be obtained, storing the content to be filled in the table area and text area on each page. The area range of each page in the file to be extracted can then be determined, and the cell information encapsulated in the table information on each page can be extracted. When processing a cell, the current position can be first determined as the target location information. The target location information can then be checked to see if the text area's location index contains the target location information. If so, the current position is a text area. Based on the text area's location index, the text content at the corresponding location can be filled into the text portion of the layout information stored in the target storage location. The above steps can then be repeated, using the next location as the target location information. If the text area's location index does not contain the target location information, the current position should be filled with table content. At this point, a two-tuple object can be obtained from the table information based on the location index, and value_table can be retrieved and filled into the layout information's table structure to assign it to the current cell. In addition, when filling cells, you can also define the element type based on the height of the text, and determine the text level based on the corresponding height range, such as headline, subheading, main text, etc., thereby ensuring that the text structure conforms to the original hierarchical information of the file to be extracted.
[0065] In an embodiment of the present disclosure, the method also includes: determining the area information and position index corresponding to other areas in the file to be extracted; filling the area information into other parts corresponding to the layout information of the file to be extracted based on the position index corresponding to the other areas to obtain the file extraction result; wherein the other areas include image areas and symbol areas.
[0066] In a specific implementation, the file to be extracted may include not only text and table areas, but also image areas, symbol areas, formula areas, and any other arbitrary areas. For these areas, each page of the file to be extracted can be pre-drawn by drawing, obtaining the location index of the corresponding area on each page, generating and storing layout information, and then extracting the file content and filling it into the corresponding part of the layout information. The specific method and process are similar to the method and process described above for table and text areas, and will not be repeated here.
[0067] In the embodiment of the present disclosure, after obtaining the file extraction result, the method further includes: determining the type of each file object in the file extraction result; and outputting the file extraction result according to the output format pre-selected by the user and the type of each file object. Figure 2As shown, users can specify the output format when specifying the file to be extracted, and implement different types of file output, such as txt, json, html, etc. In addition, the type of each file object can be determined. If it is a table, a table tag can be added; if it is a title, a corresponding title tag can be added. Finally, the file extraction results are returned to the user.
[0068] According to the second aspect of the embodiment of the present disclosure, Figure 5 As shown, a text extraction device 500 is provided, comprising:
[0069] Extraction module 501, used to determine the table information corresponding to the table area in the file to be extracted, and extract the text content of the text area in the file to be extracted; the table information includes the table structure and the table content;
[0070] A storage module 502 is configured to store layout information corresponding to the file to be extracted in a target storage location based on position indexes of the table area and the text area in the file to be extracted;
[0071] The filling module 503 is used to fill the text content into the text part of the layout information, and fill the table content into the table structure corresponding to the table part of the layout information, so as to obtain the file extraction result.
[0072] Optionally, the extraction module 501 is specifically configured to:
[0073] Traverse the area blocks corresponding to the file to be extracted to determine the table line information in the file to be extracted;
[0074] Construct cell information based on table line information; the cell information includes text data and element type corresponding to each cell;
[0075] Based on the table line information and cell information, the table information corresponding to the table area in the file to be extracted is determined.
[0076] Optionally, the extraction module 501 is specifically configured to:
[0077] Based on the predetermined page height of the file to be extracted and the position information of the current processing area, determining whether the current processing area is a text area;
[0078] In response to the current processing area being a text area, the area range of the text area is determined, and the corresponding vertical coordinate of the text area in the page is used as a position index of the text area.
[0079] Optionally, the filling module 503 is specifically configured to:
[0080] Traverse the position index of the text area and determine whether the target position information is included in the position index;
[0081] In response to the target location information not being included in the location index, determining that the target location is a table portion in the layout information, and filling the table content corresponding to the target location into the table structure of the layout information;
[0082] In response to the target position information being included in the position index, the target position is determined to be a text portion in the layout information, and the text content corresponding to the target position is filled into the text portion of the layout information.
[0083] Optionally, the filling module 503 further includes:
[0084] Determine the region information and position index corresponding to other regions in the file to be extracted;
[0085] Based on the position indexes corresponding to other regions, the region information is filled into other parts corresponding to the layout information of the file to be extracted to obtain the file extraction result; wherein the other regions include image regions and symbol regions.
[0086] Optionally, the device further includes an output device 504; the output device 504 is specifically configured to:
[0087] Determine the type of each file object in the file extraction result;
[0088] Output the file extraction results according to the output format pre-selected by the user and the type of each file object.
[0089] According to the third aspect of an embodiment of the present disclosure, an electronic device for text extraction is provided, comprising: one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by one or more processors, the one or more processors implement the method provided by the first aspect of the embodiment of the present invention.
[0090] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable medium is provided, on which a computer program is stored. When the program is executed by a processor, the method provided by the first aspect of the embodiment of the present disclosure is implemented.
[0091] According to a fifth aspect of an embodiment of the present invention, a computer program product is provided, comprising a computer program, which implements the method of any of the above embodiments when executed by a processor.
[0092] Figure 6 An exemplary system architecture 600 is shown to which the text extraction method or text extraction apparatus implemented in the present disclosure can be applied.
[0093] like Figure 6As shown, system architecture 600 may include terminal devices 601, 602, 603, a network 604, and a server 605. Network 604 is used to provide a medium for communication links between terminal devices 601, 602, 603 and server 605. Network 604 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0094] Users can use terminal devices 601, 602, and 603 to interact with server 605 via network 604 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 601, 602, and 603, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0095] The terminal devices 601 , 602 , and 603 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.
[0096] Server 605 may be a server that provides various services, such as a backend management server (for example only) that supports shopping websites browsed by users using terminal devices 601, 602, and 603. The backend management server may process received text extraction requests and provide feedback (for example only) on the processing results to the terminal devices.
[0097] It should be noted that the text extraction method provided in the embodiment of the present invention is generally executed by the server 605, and accordingly, the text extraction device is generally provided in the server 605. The text extraction method provided in the embodiment of the present invention can also be executed by the terminal devices 601, 602, and 603, and accordingly, the text extraction device can be provided in the terminal devices 601, 602, and 603.
[0098] It should be understood that Figure 6 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0099] Reference below Figure 7 , which shows a schematic structural diagram of a computer system 700 of a terminal device suitable for implementing an embodiment of the present invention. Figure 7 The terminal device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0100] like Figure 7As shown, computer system 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage unit 708 into a random access memory (RAM) 703. Various programs and data required for the operation of system 700 are also stored in RAM 703. CPU 701, ROM 702, and RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to bus 704.
[0101] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, mouse, and the like; an output section 707 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and speakers; a storage section 708 including devices such as a hard disk; and a communication section 709 including a network interface card such as a LAN card or a modem. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. Removable media 711, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 710 as needed, so that computer programs read from the media can be installed in the storage section 708 as needed.
[0102] In particular, according to embodiments disclosed herein, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed herein include a computer program product comprising a computer program embodied on a computer-readable medium, the computer program containing program code for executing the methods illustrated in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 709 and / or installed from removable media 711. When executed by central processing unit (CPU) 701, the computer program performs the aforementioned functions defined in the system of the present invention.
[0103] It should be noted that the computer-readable medium described in the present invention may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. Computer-readable storage media may include, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present invention, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such a propagated data signal may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wireline, optical fiber cable, RF, or any suitable combination thereof.
[0104] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0105] The modules involved in the embodiments of the present invention may be implemented in software or hardware. The modules described may also be provided in a processor. For example, a processor may include an extraction module, a storage module, and a filling module. The names of these modules do not, in some cases, limit the modules themselves. For example, the extraction module may also be described as a module that "determines the table information corresponding to the table area in the file to be extracted, and extracts the text content of the text area in the file to be extracted."
[0106] As another aspect, the present invention further provides a computer-readable medium, which may be included in the device described in the above embodiment; or may exist independently and not be assembled into the device. The computer-readable medium carries one or more programs, and when the one or more programs are executed by a device, the device implements the following method: determining table information corresponding to a table area in a file to be extracted, and extracting text content of a text area in the file to be extracted; the table information includes a table structure and table content; based on the position index of the table area and the text area in the file to be extracted, storing layout information corresponding to the file to be extracted in a target storage location; filling the text content into the text portion of the layout information, and filling the table content into the table structure corresponding to the table portion of the layout information, to obtain a file extraction result.
[0107] Finally, it should be noted that the above embodiments are only specific implementation methods of the present disclosure, which are used to illustrate the technical solutions of the present disclosure, rather than to limit them. The scope of protection of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the above embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above embodiments within the technical scope disclosed in the present disclosure, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure should be based on the scope of protection of the claims.
Claims
1. A file extraction method, characterized in that: include: Determine the table information corresponding to the table area in the file to be extracted, and extract the text content of the text area in the file to be extracted; the table information includes the table structure and the table content; Based on the position index of the table area and the text area in the file to be extracted, storing the layout information corresponding to the file to be extracted in a target storage location; The text content is filled into the text part of the layout information, and the table content is filled into the table structure corresponding to the table part of the layout information, so as to obtain a file extraction result.
2. The method according to claim 1, characterized in that Determine the table information corresponding to the table area in the file to be extracted, including: Traversing the area blocks corresponding to the file to be extracted, and determining the table line information in the file to be extracted; Constructing cell information based on the table line information; the cell information includes text data and element type corresponding to each cell; Based on the table line information and the cell information, the table information corresponding to the table area in the file to be extracted is determined.
3. The method according to claim 1, characterized in that Before extracting the text content of the file to be extracted, it also includes: Based on the predetermined page height of the file to be extracted and the position information of the current processing area, determining whether the current processing area is a text area; In response to the current processing area being a text area, the area range of the text area is determined, and the corresponding vertical coordinate of the text area in the page is used as the position index of the text area.
4. The method according to claim 3, characterized in that Filling the text content into the text part of the layout information, and filling the table content into the table structure corresponding to the table part of the layout information, includes: Traversing the position index of the text area to determine whether the target position information is included in the position index; In response to the target position information not being included in the position index, determining that the target position is a table portion in the layout information, and filling the table content corresponding to the target position into the table structure of the layout information; In response to the target position information being included in the position index, the target position is determined to be a text portion in the layout information, and the text content corresponding to the target position is filled into the text portion of the layout information.
5. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: Determine the region information and position index corresponding to other regions in the file to be extracted; The region information is filled into other parts corresponding to the layout information of the file to be extracted based on the position index corresponding to the other regions to obtain the file extraction result; wherein the other regions include image regions and symbol regions.
6. The method according to any one of claims 1 to 5, characterized in that: After getting the file extraction results, it also includes: Determining the type of each file object in the file extraction result; The file extraction result is output according to the output format pre-selected by the user and the type of each file object.
7. A file extraction device, characterized in that: include: An extraction module, configured to determine table information corresponding to a table region in a file to be extracted, and extract text content from the text region in the file to be extracted; the table information includes a table structure and table content; A storage module, configured to store layout information corresponding to the file to be extracted into a target storage location based on a position index of a table area and a text area in the file to be extracted; The filling module is used to fill the text content into the text part of the layout information, and fill the table content into the table structure corresponding to the table part of the layout information to obtain a file extraction result.
8. An electronic device, characterized in that: include: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
9. A computer-readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
10. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 6.