Cross-page cell merging method and device, electronic equipment and storage medium
Through artificial intelligence technology, image and text features are extracted and fusion of cross-page tables, the problem of insufficient content integrity in cross-page cell merging is solved, and higher merging accuracy is achieved.
Patent Information
- Application Number
- CN202510026134.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art fails to fully consider the complete accuracy of cell content when merging cells across pages, resulting in low merging accuracy.
Using an artificial intelligence-based method, image and text features are extracted through preset merge models and feature fusion is performed on the spreadsheet. Combining object detection and semantic recognition, the layout structure of the spreadsheet is identified and merged to ensure the integrity and accuracy of cell content.
Improve the accuracy of cell merging in spreadsheets, ensuring that the merged table is visually and functionally a whole, fully considering the completeness and accuracy of cell content.
Smart Images

Figure CN120071373A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular, to a method and device for merging cross-page cells, an electronic device, and a storage medium. Background Art
[0002] Cross-page cell merging refers to a method of merging table contents belonging to the same cell in a spreadsheet. Cross-page cell merging can maintain the integrity and accuracy of cell contents, thereby ensuring the accuracy of cross-page table and file parsing.
[0003] Related technologies usually identify table lines in cross-page tables and perform cross-page table merging based on the identified table lines. However, this method can only ensure that all parts of the merged table are visually and functionally a whole, without considering whether the contents of the cells therein are complete and accurate, resulting in a low accuracy of cross-page cell merging. Summary of the Invention
[0004] The main objective of the embodiments of the present application is to propose a method and device for merging cross-page cells, an electronic device, and a storage medium, which can improve the accuracy of cell merging in cross-page tables.
[0005] To achieve the above objective, a first aspect of the embodiments of the present application proposes a method for merging cross-page cells, the method including:
[0006] Obtain a first file image and a second file image of a target document file; wherein, the file page corresponding to the second file image is the file page after the file page corresponding to the first file image, and the first file image includes a first table area;
[0007] Perform table detection on the first file image based on a table area detection sub-model of a preset merging model to obtain first table area data and a first area label of the first table area; wherein, the first table area data is used to characterize the position information of the first table area in the first file image, and the first area label is used to indicate the area type of the first table area;
[0008] Perform image merging on the first file image and the second file image based on the first area label and the first table area data to obtain a cross-page table merged image;
[0009] Perform image feature extraction on the cross-page table merged image based on an image feature extraction sub-model of the preset merging model to obtain cross-page image features;
[0010] Perform text feature extraction on the cross-page table merged image based on a text feature extraction sub-model of the preset merging model to obtain cross-page text features;
[0011] The feature fusion sub-model based on the preset merging model performs feature fusion on the cross-page image feature and the cross-page text feature to obtain a cross-page fusion feature;
[0012] The cell detection sub-model based on the preset merging model performs cell detection on the cross-page fusion feature to obtain target cross-page cell data.
[0013] In some embodiments, the performing image merging on the first document image and the second document image based on the first region label and the first table region data to obtain a cross-page table merged image includes:
[0014] Obtaining a first table sub-label and a first cross-page sub-label from the first region label; wherein, the first table sub-label is used to represent whether the first table region is a cross-page table region or a non-cross-page table region, and the first cross-page sub-label is used to represent whether the first table region is the first page region or the subsequent page region of the cross-page table region of the cross-page table region;
[0015] If the first table sub-label indicates that the first table region is the cross-page table region, performing table detection on the second document image based on the table region detection sub-model to obtain second table region data and a second region label of a second table region in the second document image;
[0016] Selecting a third document image from the second document image based on the second region label and the first cross-page sub-label; wherein, the third document image includes a third table region that is cross-page associated with the first table region, and using the second table region data corresponding to the third document image as the third table region data;
[0017] Performing image merging on the first document image and the third document image based on the first table region data and the third table region data to obtain the cross-page table merged image.
[0018] In some embodiments, the performing image merging on the first document image and the third document image based on the first table region data and the third table region data to obtain the cross-page table merged image includes:
[0019] Extracting a first region image of the first table region from the first document image based on the first table region data;
[0020] Extracting a second region image of the third table region from the third document image based on the third table region data;
[0021] Determine the cross-page area sequence data based on the first cross-page sub-tag and the third area tag of the third table area;
[0022] Perform image stitching on the first area image and the second area image based on the cross-page area sequence data to obtain the cross-page table merged image.
[0023] In some embodiments, the performing image stitching on the first area image and the second area image based on the cross-page area sequence data to obtain the cross-page table merged image includes:
[0024] Remove the cross-page table lines in the first area image and update the first area image;
[0025] Remove the cross-page table lines in the second area image and update the second area image;
[0026] Perform image stitching on the updated first area image and the updated second area image based on the cross-page area sequence data to obtain the cross-page table merged image.
[0027] In some embodiments, the performing feature fusion on the cross-page image feature and the cross-page text feature by the feature fusion sub-model based on the preset merging model to obtain the cross-page fusion feature includes:
[0028] Determine a first text line feature and a second text line feature from the cross-page text feature, where the second text line corresponding to the second text line feature is the next text line of the first text line corresponding to the first text line feature, the first text line feature is used to describe the text characters arranged based on the corresponding text character sequence in the first text line, and the second text line feature is used to describe the text characters arranged based on the corresponding text character sequence in the second text line;
[0029] Perform masking processing on the first text line feature and the second text line feature respectively based on the text character sequence and the preset mask block to obtain a first masked text line feature and a second masked text line feature;
[0030] Update the first text line feature and the second text line feature based on the second text line, and perform masking processing on the updated first text line feature and the second text line feature until masking processing is performed on all text line features in the cross-page text feature to obtain the masked text feature;
[0031] Perform feature fusion on the cross-page image feature and the masked text feature by the feature fusion sub-model to obtain the cross-page fusion feature.
[0032] In some embodiments, obtaining the first file image and the second file image of the target document file includes:
[0033] Obtain a target document file, where the target document file includes a target file page;
[0034] Perform image conversion on the target file page to obtain a target file image;
[0035] Determine the first file image and the second file image from the target file image based on the file page order of the target file page.
[0036] In some embodiments, the method further includes:
[0037] Obtain cell structure data and cell content data from the target cross-page cell data;
[0038] Perform table visualization processing based on the cell structure data and the cell content data to obtain a cross-page merged table.
[0039] To achieve the above object, a second aspect of the embodiments of the present application proposes a cross-page cell merging device, and the device includes:
[0040] An obtaining module, configured to obtain a first file image and a second file image of a target document file; wherein, the file page corresponding to the second file image is the file page after the file page corresponding to the first file image, and the first file image includes a first table area;
[0041] A table detection module, configured to perform table detection on the first file image based on a table area detection sub-model of a preset merging model to obtain first table area data and a first area label of the first table area; wherein, the first table area data is used to characterize the position information of the first table area in the first file image, and the first area label is used to indicate the area type of the first table area;
[0042] A merging module, configured to perform image merging on the first file image and the second file image based on the first area label and the first table area data to obtain a cross-page table merged image;
[0043] A first feature extraction module, configured to perform image feature extraction on the cross-page table merged image based on an image feature extraction sub-model of the preset merging model to obtain cross-page image features;
[0044] A second feature extraction module, configured to perform text feature extraction on the cross-page table merged image based on a text feature extraction sub-model of the preset merging model to obtain cross-page text features;
[0045] A feature fusion module, configured to perform feature fusion on the cross-page image features and the cross-page text features based on a feature fusion sub-model of the preset merging model to obtain cross-page fusion features;
[0046] A cell detection module, configured to perform cell detection on the cross-page fusion features based on a cell detection sub-model of the preset merging model to obtain target cross-page cell data.
[0047] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, where the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect above is implemented.
[0048] To achieve the above object, a fourth aspect of the embodiments of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect above is implemented.
[0049] The cross-page cell merging method, device, electronic device, and storage medium proposed in the embodiments of the present application obtain a first file image and a second file image of a target document file, where the file page corresponding to the second file image is the file page after the file page corresponding to the first file image, and the first file image includes a first table area; further, perform table detection on the first file image based on the table area detection sub-model of a preset merging model to obtain first table area data and a first area label of the first table area, where the first table area data is used to represent the position information of the first table area in the first file image, and the first area label is used to indicate the area type of the first table area; further, perform image merging on the first file image and the second file image based on the first area label and the first table area data to obtain a cross-page table merged image; further, perform image feature extraction on the cross-page table merged image based on the image feature extraction sub-model of the preset merging model to obtain cross-page image features; further, perform text feature extraction on the cross-page table merged image based on the text feature extraction sub-model of the preset merging model to obtain cross-page text features; further, perform feature fusion on the cross-page image features and the cross-page text features based on the feature fusion sub-model of the preset merging model to obtain cross-page fusion features; further, perform cell detection on the cross-page fusion features based on the cell detection sub-model of the preset merging model to obtain target cross-page cell data. Compared with the related technology that only performs cross-page table merging according to the recognized table lines, the embodiments of the present application can ensure that all parts of the merged table are a whole visually and functionally by respectively performing feature extraction on the cross-page table merged image for images and texts and performing feature fusion on the extracted features, and fully consider the integrity and accuracy of the content of the cells therein. Therefore, the embodiments of the present application can improve the accuracy of cell merging in cross-page tables. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is a flowchart of the cross-page cell merging method provided by the embodiments of the present application;
[0051] Figure 2 is Figure 1 a flowchart of step S110 in
[0052] Figure 3 is Figure 1 a flowchart of step S130 in
[0053] Figure 4 is Figure 3 a flowchart of step S340 in
[0054] Figure 5 is Figure 4 a flowchart of step S440 in
[0055] Figure 6A It is a schematic diagram of a cross-page screenshot of the target document file provided by an embodiment of the present application;
[0056] Figure 6B It is a schematic diagram of a cross-page table merging image provided by an embodiment of the present application;
[0057] Figure 7 is Figure 1 a flowchart of step S160 in
[0058] Figure 8 It is another flowchart of the cross-page cell merging method provided by an embodiment of the present application;
[0059] Figure 9 It is a schematic diagram of a visual cross-page merged table provided by an embodiment of the present application;
[0060] Figure 10 It is another schematic diagram of a visual cross-page merged table provided by an embodiment of the present application;
[0061] Figure 11 It is a schematic structural diagram of a cross-page cell merging device provided by an embodiment of the present application;
[0062] Figure 12 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0063] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0064] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first" and "second" in the description, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0065] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0066] First, several nouns involved in the present application are analyzed:
[0067] Artificial Intelligence (AI): It is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. Artificial intelligence is a branch of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing, and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. It also refers to the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0068] Optical Character Recognition (OCR) method refers to obtaining a photo image of a paper document (including various certificates) through an electronic device (such as a scanner or digital camera, etc.), then recognizing the text from the image and converting it into electronic document information, and the OCR method is widely used in the automatic recognition of various documents, certificates, and bills.
[0069] In today's information society, tables, as an important form of data organization and presentation, are widely used in various documents and reports. It is very common for tables to span multiple pages in long documents. At present, a small amount of research has solved the merging of multi-page tables, but there is no mature solution for the merging of cells that span multiple pages. Merging cells that span multiple pages refers to a method of merging the table contents belonging to the same cell in a spreadsheet. Merging cells that span multiple pages can maintain the integrity and accuracy of the cell contents, thereby ensuring the accuracy of multi-page table and file parsing.
[0070] Related technologies usually identify the table lines in a multi-page table and perform multi-page table merging based on the identified table lines. However, this method can only ensure that all parts of the merged table are a whole visually and functionally, without considering whether the contents of the cells are complete and accurate, resulting in a relatively low accuracy of merging cells that span multiple pages.
[0071] Based on this, the embodiments of this application provide a method and device for merging cells that span multiple pages, an electronic device, and a storage medium, which can improve the accuracy of merging cells in a multi-page table.
[0072] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0073] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0074] The cross-page cell merging method provided by the embodiments of the present application relates to the field of artificial intelligence technology. The cross-page cell merging method provided by the embodiments of the present application can be applied to a terminal, or to a server side, or can also be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms; the software can be an application that implements the cross-page cell merging method, etc., but is not limited to the above forms.
[0075] The present application can be used in many general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet-type devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network personal computers (PCs), minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0076] It should be noted that in each specific embodiment of the present application, when it comes to relevant processing based on data related to the identity or characteristics of an object, such as object table information, object behavior data, object historical data, etc., the permission or consent of the object will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain sensitive personal information of an object, the object's separate permission or separate consent will be obtained through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the object's separate permission or separate consent, the necessary object-related data for the normal operation of the embodiments of the present application will be obtained.
[0077] Please refer to Figure 1 , Figure 1 which is an optional flowchart of the cross-page cell merging method provided by the embodiments of the present application. Figure 1 The method in Figure 1 may specifically include but is not limited to steps S110 to S170. The following will introduce these seven steps in detail with reference to
[0078] Step S110, obtain a first file image and a second file image of the target document file;
[0079] Step S120, perform table detection on the first file image based on the table area detection sub-model of the preset merging model to obtain the first table area data and the first area label of the first table area;
[0080] Step S130, perform image merging on the first file image and the second file image based on the first area label and the first table area data to obtain a cross-page table merged image;
[0081] Step S140, perform image feature extraction on the cross-page table merged image based on the image feature extraction sub-model of the preset merging model to obtain cross-page image features;
[0082] Step S150, perform text feature extraction on the cross-page table merged image based on the text feature extraction sub-model of the preset merging model to obtain cross-page text features;
[0083] Step S160, perform feature fusion on the cross-page image features and the cross-page text features based on the feature fusion sub-model of the preset merging model to obtain cross-page fusion features;
[0084] Step S170, perform cell detection on the cross-page fusion features based on the cell detection sub-model of the preset merging model to obtain the target cross-page cell data.
[0085] In steps S110 to S170 of some embodiments, first, table detection is performed by means of image detection to identify the layout structure of a multi-page table. Feature extraction of images and texts is respectively performed on the merged images of the multi-page table, and the extracted features are fused. This can ensure that all parts of the merged table are an integral whole visually and functionally, while fully considering the completeness and accuracy of the content in the cells. Compared with the related art method of merging multi-page tables only based on the recognized table lines, the present application combines object detection and semantic recognition, extracts the visual features and text features of the table respectively, and increases the discriminant features of semantic continuity through the feature fusion method for cell text, improving the recognition and restoration effect of complex table structures. Therefore, the embodiments of the present application can improve the accuracy of cell merging in multi-page tables.
[0086] In step S110 of some embodiments, the target document file refers to the document file that needs to perform multi-page cell merging, and the target document file can be a file format that is not easy to perform table merging, such as a Portable Document Format (PDF), EPUB format, image, etc. The target document file includes multiple target file pages, and these target file pages are arranged according to the file page order (such as page number). The first file image and the second file image respectively correspond to different target file pages in the target document file, and the file page corresponding to the second file image is the file page after the file page corresponding to the first file image. In this way, the present application can first determine the first file image according to a target file page in the target document file, and then continue to extract the image of the next page from the target file as the second file image.
[0087] It should be noted that after obtaining the first file image and the second file image, in order to ensure that the extracted images are suitable for subsequent processing, the present application can also preprocess the obtained file images, such as operations of image cropping, rotation, scaling, etc., to ensure the quality and consistency of the images, and update the first file image and the second file image according to the preprocessed file images.
[0088] It should be noted that the first file image is used to represent the image including the table area. The first table area refers to the area where the table detected in the first file image is located. For example, if a target file page is all text, then cell merging cannot be performed, and it cannot be used as the first file image or the second file image.
[0089] It should be noted that each target file page in the target document file can itself be in image format. After taking the currently detected file page as the first file image, the subsequent file page images can be used as the second file image. Moreover, after detecting the table contained in the first file image and merging the cross-page cells, the image corresponding to the next file page adjacent to the first file image can be used as the new first file image.
[0090] Please refer to Figure 2 , Figure 2 which is the specific flowchart of step S110 provided by the embodiments of the present application. In some embodiments of the present application, step S110 specifically includes but is not limited to steps S210 to S230. The following will introduce these three steps in detail in combination with Figure 2 this.
[0091] Step S210, obtain the target document file;
[0092] Step S220, perform image conversion on the target file page to obtain the target file image;
[0093] Step S230, determine the first file image and the second file image from the target file images based on the file page order of the target file pages.
[0094] In steps S210 to S230 of some embodiments, the target file image refers to the image corresponding to each target file page. The file page order is used to indicate the arrangement order among multiple target file pages in the target document file, and can be determined according to the page numbers. When each target file page in the target document file is not in image format, then it is necessary to first perform image conversion on each target file page to convert the target file page into image format for image processing. For example, a document conversion tool or an image scanning device can be used to convert the target file page into image format, and the image format of the target file image can be JPEG, PNG, or TIFF, etc.
[0095] In the above steps S210 to S230, by converting the target file page into image format, the present application can obtain an ordered set of images corresponding to the target document file, which is convenient for subsequent image feature extraction, thereby improving the accuracy of cell merging in the cross-page table.
[0096] In step S120 of some embodiments, the table area detection sub-model refers to a model used to detect whether there is a table in an image. The first table area data is used to characterize the position information of the first table area in the first document image, and the first area label is used to indicate the area type of the first table area. The first document image may include 0, 1, or more than 1 first table areas. Among them, in this application, the table area detection sub-model is used to perform table detection on the first document image, which can detect whether there is a table in the first document image, and output the specific position information of the table area and the area type of the table. The table area detection sub-model of this application can adopt machine learning or deep learning methods, such as Convolutional Neural Networks (CNN), Region-CNN (R-CNN), etc., without specific limitation.
[0097] It should be noted that the first table area data can locate the position and size of the first table area in the first document image. For example, the boundary of the table area is identified by a bounding box, such as the coordinates of the upper left corner and the lower right corner in the first document image, as well as possible rotation angles, widths, heights, etc. In this way, the table area can be accurately identified from the first document image, and its specific position information in the image can be obtained.
[0098] In step S130 of some embodiments, after determining the first table area data and the first area label in the first document image, the first document image and the second document image can be merged based on the first area label and the first table area data to obtain a cross-page table merged image. The cross-page table merged image at this time is used to characterize the image after merging the cross-page table areas in the target document file. In this way, an image containing a complete cross-page table can be obtained.
[0099] Please refer to Figure 3 , Figure 3 which is the specific flowchart of step S130 provided by the embodiments of this application. In some embodiments of this application, step S130 specifically includes but is not limited to steps S310 to S340. The following will introduce these four steps in detail in combination with Figure 3 this.
[0100] Step S310, obtain the first table sub-label and the first cross-page sub-label from the first area label;
[0101] Step S320, if the first table sub-label indicates that the first table area is a cross-page table area, perform table detection on the second document image based on the table area detection sub-model to obtain the second table area data and the second area label of the second table area in the second document image;
[0102] Step S330: Select a third document image from the second document image based on the second region label and the first cross-page sub-label;
[0103] Step S340: Based on the first table region data and the third table region data, perform image merging on the first document image and the third document image to obtain a cross-page table merged image.
[0104] In steps S310 and S320 of some embodiments, the first region label includes a first table sub-label and a first cross-page sub-label. The first table sub-label is used to characterize whether the first table region is a cross-page table region or a non-cross-page table region, and the first cross-page sub-label is used to characterize whether the first table region is the first page region or the subsequent page region of the cross-page table region of the cross-page table. That is, the first table sub-label is used to determine whether the first table region is a cross-page table region, that is, whether there is a cross-page table in the first document image. The first cross-page sub-label is used to determine whether the first table region determined to be a cross-page table region is the first page region or the subsequent page region.
[0105] It should be noted that if the first table sub-label indicates that the first table region is a cross-page table region and it is necessary to check whether there is a cross-page table region in the second document image after the first document image, the same table region detection sub-model can be used to perform table detection on the second document image to obtain the second table region data and the second region label of the second table region in the second document image. The second table region data is used to characterize the position information of the second table region in the second document image in the image, and the second region label is used to indicate the region type of the second table region. In addition, the second document image may include 0, 1, or more than 1 second table regions. After determining the second region label of the second document image, it is possible to further determine whether the second table region in the second document image is a cross-page table region. And the second document image can be used as the new first document image to query whether there is still a cross-page table in the next document page. This is because a table may be a large table spanning multiple document pages. In this way, multiple cross-page table regions can be determined.
[0106] It should be noted that if the first table sub-label indicates that the first table region is a non-cross-page table region, the cross-page cell merging method of the present application does not need to be used. Then, the second document image can be used as the new first document image, and steps S101 to S170 can be restarted. In addition, in practical applications, the first table sub-label and the first cross-page sub-label can be judged simultaneously, that is, when judging whether the currently detected document page is the first or subsequent page of the cross-page table, the cross-page table region is determined. And when there are two target table region frames in the region of the same document page, they can be distinguished according to the position information of the region frames in the document page to avoid detection problems such as easy overlap.
[0107] In step S330 of some embodiments, further, the present application may select a third document image from the second document image based on the second region label and the first cross-page sub-label. The third document image contains a third table region that is cross-page associated with the first table region, and the second table region data corresponding to the third document image is used as the third table region data. Since the second document image at this time may be the image of the next document page or the image of the next-next document page of the current first document image, therefore, the third document image that is cross-page associated with the first table region may be determined first from the document pages after the first document image. Among them, being cross-page associated with the first table region means belonging to the same table as the first table region. For example, if the second document image contains multiple tables, and the first table region among them is a continuation page region and belongs to the same table as the home page region corresponding to the first table region, then the second table region data corresponding to the first table region in the second document image may be used as the third table region data.
[0108] In step S340 of some embodiments, after determining based on the first table region data and the third table region data, the first document image and the third document image may be merged based on these data to obtain a cross-page table merged image. The cross-page table merged image at this time refers to the merged image of the region images belonging to the same table in the first document image and the third document image.
[0109] Please refer to Figure 4 , Figure 4 which is the specific flowchart of step S340 provided by the embodiments of the present application. In some embodiments of the present application, step S340 specifically includes but is not limited to steps S410 to S440. The following will combine Figure 4 to introduce these four steps in detail.
[0110] Step S410, extract the first region image of the first table region from the first document image based on the first table region data;
[0111] Step S420, extract the second region image of the third table region from the third document image based on the third table region data;
[0112] Step S430, determine the cross-page region order data based on the first cross-page sub-label and the third region label of the third table region;
[0113] Step S440, splice the first region image and the second region image based on the cross-page region order data to obtain a cross-page table merged image.
[0114] In steps S410 to S430 of some embodiments, the present application may use the first table area data (including position and size information) as a guide to crop the first table area from the first document image, and use the image where the first table area is located as the first area image. And use the third table area data (including position and size information) as a guide to crop the corresponding second table area from the third document image, and use the image where the corresponding second table area is located as the second area image. Further, the first cross-page sub-tag and the third area tag may be analyzed, and these tags provide information on whether the table area is the first page area or the continuation page area of a cross-page table, and provide the page number information of the page where the table is located. In this way, according to these tag information, the cross-page area sequence data of the first area image and the second area image to be merged in the final merged image can be determined. And this cross-page area sequence data can ensure that the finally merged table is correct.
[0115] It should be noted that the cross-page area sequence data may be sequence data determined based on page numbers, sequence marks, or other logical sequences, and is not limited.
[0116] In step S440 of some embodiments, further, the first area image and the second area image may be sequentially image-stitched based on the cross-page area sequence data to obtain a cross-page table merged image. And the present application may merge the cross-page table by using the table lines and the edge information of the text in the area images to be merged.
[0117] Please refer to Figure 5 , Figure 5 which is a specific flowchart of step S440 provided by the embodiments of the present application. In some embodiments of the present application, step S440 specifically includes but is not limited to steps S510 to S530. The following will introduce these three steps in detail in combination with Figure 5 this.
[0118] Step S510, removing the cross-page table lines in the first area image and updating the first area image;
[0119] Step S520, removing the cross-page table lines in the second area image and updating the second area image;
[0120] Step S530, performing image stitching on the updated first area image and the updated second area image based on the cross-page area sequence data to obtain a cross-page table merged image.
[0121] In steps S510 to S530 of some embodiments, when merging to obtain a merged cross-page table image, if there is a bottom table line of the cross-page table or a top table line of the next page in the region images to be merged, these two lines can be removed to avoid the influence of redundant table lines on the merged content, thereby improving the accuracy of cell merging in the cross-page table. Specifically, image processing techniques can be used to first remove the table lines for connecting the cross-page in the first-page region and the subsequent-page region of the cross-page table, that is, first identify the cross-page table lines in the first region image and the second region image respectively, and remove these table lines to update the first region image and the second region image. Further, according to the cross-page region sequence data, the updated first region image and the second region image are stitched together to obtain a merged cross-page table image.
[0122] It should be noted that when stitching the updated first region image and the second region image in this application, the updated first region image and the second region image can be aligned first to ensure that the images of the two regions are aligned before stitching. Specifically, it can further include adjusting the size, angle, or position of the image, etc. Then, the updated first region image and the second region image are stitched together in a determined order to form a complete merged cross-page table image. Further, the obtained merged cross-page table image can also be processed for the stitching edge to ensure that the stitching edge is natural and has no obvious traces, so as to update the merged cross-page table image.
[0123] Exemplarily, please refer to Figure 6A , Figure 6A which is a schematic diagram of a cross-page screenshot of the target document file provided by the embodiment of the present application. The cross-page screenshot 610 includes partial images of the first document image 620 and partial images of the second document image 630, and the cross-page screenshot 610 is used to indicate a cross-page table formed by the first document image 620 and the second document image 630 through the cross-page. Among them, the first document image 620 includes a first table region 621, and the second document image 630 includes a second table region 631. Based on the table region detection sub-model of the present application, table detection is respectively performed on the first table region 621 of the first document image 620 and the second table region 631 of the second document image 630. Then, the first table region data of the first table region 621 can be determined, and the first table region 621 is determined as the first-page region of the cross-page table. And the second table region data of the second table region 631 is determined, and the second table region 631 is determined as the subsequent-page region of the cross-page table. Further, based on the above steps S410 to S440 and removing the cross-page table lines, a merged cross-page table image as shown in Figure 6B can be obtained.
[0124] In step S140 of some embodiments, the present application may perform image feature extraction on the cross-page table merging image based on the image feature extraction sub-model of the preset merging model to obtain cross-page image features. Image feature extraction is a process of identifying and describing important information in an image, and the cross-page image features obtained at this time may include color features of the image (such as color histograms, color moments, etc., used to describe the color distribution of the image), texture features (such as gray-level co-occurrence matrices, local binary patterns, etc., reflecting the texture characteristics of the image), shape features (describing the shape by extracting information such as the contour and edges of the image), etc.
[0125] It should be noted that the image feature extraction sub-model of the present application may be based on a deep neural network model or a machine learning model, etc., without limitation.
[0126] In step S150 of some embodiments, meanwhile, the present application may perform text feature extraction on the cross-page table merging image based on the text feature extraction sub-model of the preset merging model, and may focus on considering the cross-page connection part of the table, and extract the text features of the cell part of the cross-page connection to maintain the integrity and coherence of the table data, thereby improving the accuracy of cell merging in the cross-page table.
[0127] It should be noted that the text feature extraction sub-model of the present application may be constructed based on algorithms such as the Bag of Words (BoW), Term Frequency-Inverse Document Frequency (TF-IDF, that is, a method considering Term Frequency and Inverse Document Frequency, which can reduce the influence of common words and highlight important vocabulary), Word Embeddings, OCR, pdf parsing, etc., without limitation.
[0128] In step S160 of some embodiments, further, the present application may perform feature fusion on the cross-page image features and cross-page text features based on the feature fusion sub-model of the preset merging model to integrate the features extracted from two different modalities of images and texts to form a more comprehensive feature representation that can represent the cross-page content. Among them, the cross-page image features may contain visual information, such as shape, color, and texture, etc., while the cross-page text features may contain semantic information. The goal of feature fusion is to merge this information so that the model can simultaneously understand the visual content of the image and the semantic content of the text. In this way, by fusing the image features and text features of the cross-page table area, the present application can improve the model's understanding and processing ability of cross-page information.
[0129] It should be noted that feature fusion can be achieved in various ways, including but not limited to feature concatenation, weighted fusion, attention mechanism, etc. Feature concatenation means directly connecting feature vectors of different modalities, while the attention-based fusion method can learn the importance of different features, which is more flexible and effective, and can be flexibly selected according to actual needs without limitation.
[0130] Please refer to Figure 7 , Figure 7 which is the specific flowchart of step S160 provided by the embodiments of the present application. In some embodiments of the present application, step S160 specifically includes but is not limited to steps S710 to S740. The following will introduce these four steps in detail in combination with Figure 7 this.
[0131] Step S710, determine the first text line feature and the second text line feature from the cross-page text features;
[0132] Step S720, respectively perform masking processing on the first text line feature and the second text line feature based on the text character sequence and the preset mask block to obtain the first masked text line feature and the second masked text line feature;
[0133] Step S730, update the first text line feature and the second text line feature based on the second text line, and perform masking processing on the updated first text line feature and the second text line feature until masking processing is performed on all text line features in the cross-page text features to obtain the masked text features;
[0134] Step S740, perform feature fusion on the cross-page image features and the masked text features based on the feature fusion sub-model to obtain the cross-page fusion features.
[0135] In step S710 of some embodiments, when performing feature fusion in the present application, a feature fusion sub-model can be constructed based on the attention mechanism. The cross-page text features can describe the text features of each row in the merged table corresponding to the cross-page table merging image. Specifically, the present application can first determine the first text row feature and the second text row feature from the cross-page text features. The first text row refers to any row of text in the merged table, and the second text row is the next row of text immediately following the first text row, that is, the second text row corresponding to the second text row feature is the next text row of the first text row corresponding to the first text row feature. The first text row feature is used to describe the text characters, positions, sizes, styles, etc. arranged based on the corresponding text character sequence in the first text row, which means that the first text row feature contains the information of all characters in that row, and this information may be the recognition results, positions, font sizes, etc. of the characters. The second text row feature is used to describe the text characters, positions, sizes, styles, etc. arranged based on the corresponding text character sequence in the second text row.
[0136] It should be noted that since the number of text characters in the first text row and the second text row is not necessarily the same, their text character sequences are not necessarily the same. The text character sequence is used to represent the arrangement order of text characters in the corresponding text row, which can be from left to right, from top to bottom, from right to left, etc., and is not limited. The following embodiments are illustrated with examples from left to right.
[0137] In step S720 of some embodiments, further, the present application can use a preset mask block to perform mask processing on the first text row feature and the second text row feature. The preset mask block may be of a fixed size and is used to cover (or "mask") certain parts of the text sequence, forcing the model to focus on the unmasked parts. The first masked text row feature is used to represent the feature after mask processing on the first text row feature based on the text character sequence corresponding to the first text row and the preset mask block, and the second masked text row feature is used to represent the feature after mask processing on the second text row feature based on the text character sequence corresponding to the second text row and the preset mask block. These features contain partially masked information for subsequent processing.
[0138] It should be noted that for the cell merging scenario, the present application can modify the cross-page image features and cross-page text features through the mask assignment in the attention-based feature fusion sub-model to better adapt to this application scenario. For example, due to the general writing norms of cell text from left to right and from top to bottom, a mask module (i.e., the preset mask block) is used to control the visible information of the text in the modal feature fusion, and only the text on the right side of the current text line and the next text line are visible to perform mask processing on the text row features. And the specific preset mask block can be expressed as formula 1 shown below:
[0139]
[0140] In Formula 1, mask represents a preset mask block, i is the line number of the first text line recognized based on OCR, j is the sequence number of each text character in the first text line in the corresponding text character sequence, seq is the number of text characters in the text character sequence corresponding to the first text line, and n is a natural number greater than 0. For example, if the first text line is "123456789 (line break)" and the next line (i.e., the second text line) is "abcdefg", then when performing feature fusion on cross-page image features and cross-page text features, when the sliding window reaches 3 and after mask processing, the information of the visible text in the first text line is only "456789". In this way, only the text on the right side of the current line and the text on the next line of the current text are visible.
[0141] In step S730 of some embodiments, further, the second text line is used as the new first text line, the first text line and the second text line are re-determined, the first text line features and the second text line features are updated according to the newly determined text lines, and step S720 is repeated until the next text line of the second text line is a preset blank line, thereby completing the mask processing of all text line features in the cross-page text features. And, the masked text line features corresponding to each line text after mask processing are feature-merged to obtain the masked text features of the cross-page table merged image.
[0142] In step S740 of some embodiments, further, a feature fusion sub-model can be used to perform attention mechanism-based feature fusion on the cross-page image features and the masked text features. In this way, the present application can enhance the text features of the cross-page text based on the attention mechanism and mask processing, and perform feature fusion based on the masked text features after text feature enhancement and the cross-page image features, which can improve the accuracy of feature recognition. Then, the fused features are detected by the cell detection sub-model for cell detection. In this way, the accuracy of cell merging in the cross-page table can be effectively improved.
[0143] In the above embodiments, the present application can, while maintaining the sequential relationship between text lines, enhance the model's understanding of text features through mask processing, and finally obtain a comprehensive feature representation that can simultaneously understand image and text information through feature fusion. This feature representation can improve the performance and accuracy of the model when processing cross-page documents.
[0144] In step S170 of some embodiments, further, the present application can analyze the cross-page fusion features through the cell detection sub-model of the preset merging model, so as to accurately identify and locate the positions and boundaries of the cells in the cross-page table merging image. The target cross-page cell data is used for the positions and contents of each cell in the cross-page table corresponding to the cross-page table merging image. In addition, the cell detection sub-model can adopt model structures such as neural network models, machine learning models, PaddleOCR, etc., without limitation.
[0145] It should be noted that the target cross-page cell data obtained by the present application can be a structured data, such as in the format of a JSON dictionary (JavaScript Object Notation, a lightweight data interchange format, easy to read and write, and also easy to be parsed and generated by machines), JSON dictionary format, etc., without limitation. In addition, in practical applications, the method adopted by the present application can be embedded in the parsing of the entire pdf or used as a separate table parsing, depending on the usage scenario. Eventually, the original information of the table structure can be obtained, that is, the output target cross-page cell data can be in various structured formats such as html, csv, markdown, etc.
[0146] Please refer to Figure 8 , Figure 8 which is another optional flowchart of the cross-page cell merging method provided by the embodiments of the present application. In some embodiments of the present application, the method may specifically further include but is not limited to steps S810 to S820. The following will combine Figure 8 to introduce these two steps in detail.
[0147] Step S810, obtaining cell structure data and cell content data from the target cross-page cell data;
[0148] Step S820, performing table visualization processing based on the cell structure data and the cell content data to obtain a cross-page merged table.
[0149] In steps S810 and S820 of some embodiments, the cell structure data is used to describe the structural position information of each cell in the cross-page table corresponding to the cross-page table merging image, such as the boundary coordinates, as well as the relationships between cells, which cells are adjacent or span multiple rows / columns, etc. These information can accurately restore the structure of each cell in the cross-page table. The cell content data is used to describe the text or image content information in each cell in the cross-page table corresponding to the cross-page table merging image. Further, the present application can reconstruct the layout of the table according to the cell structure data, including the division of rows and columns and the positioning of cells, and reconstruct the specific content of each cell in the table according to the cell content data, that is, fill the cell content data into the corresponding cells to ensure that the text or image in each cell matches its position in the table, so as to generate a visual cross-page merged table. Finally, the reconstructed table is output in a visual form, which can be an image, a PDF or an Excel file, etc.
[0150] Through the above steps, the present application can extract structured data from cross-page images and text features and convert them into a table format that is easy to read and process, so as to realize the effective management and analysis of cross-page table data.
[0151] It should be noted that the present application adopts cross-page cell merging based on multi-modal (image and text), which is applicable to both the cell merging of cross-page tables and the cell recognition of wireless tables.
[0152] It should be noted that when generating a visual cross-page merged table, as Figure 9 shown, the cross-page merged table generated by the present application may not display the table lines of each cell, and only use underlines to mark some data. As Figure 10 shown, the cross-page merged table generated by the present application can display the table lines of each cell, and can use different colors to mark the cell data to highlight the key points of concern.
[0153] It should be noted that when training the preset merging model in the present application, a training sample set can be obtained first. The training sample set includes multiple sample data, and each sample data includes a sample file image and a corresponding sample label. The specific content of the sample label at this time can refer to the above-mentioned first area label and will not be elaborated here, so that the preset merging model can be trained according to the labeled training sample set. And the loss function used in the training can be a cross-entropy loss function, an attention loss function, etc., which can be flexibly adjusted according to actual needs and will not be specifically limited here. The training end condition can be when all the sample data in the training sample set are trained.
[0154] A method for merging cells across pages provided by an embodiment of the present application starts from the row and column characteristics of table regularization, combined with the text characteristics of spanning rows and columns, to implement an end-to-end table structure recognition and restoration method to achieve the purpose of merging cells across pages. By fusing image features and text features and combining object detection and post-processing techniques, this method can efficiently and accurately identify and restore tables, especially outstanding in applications such as merging cells across pages and restoring wireless tables. Specifically,
[0155] The specific innovation points include: (1) Feature fusion of object detection and semantic recognition: By combining object detection and semantic recognition, the visual features and text features of the table are extracted respectively, and discriminative features with semantic continuity are added through a feature fusion method for cell text to improve the recognition and restoration effect of complex table structures; (2) Layout recognition and merging processing for tables across pages: For tables across pages, this method first identifies the layout structure of the table across pages through an image detection method, merges the table across pages into a single table, that is, merges the corresponding tables in the image of the table across pages, and then processes it through the above-mentioned feature fusion method of object detection and semantic recognition to achieve the merging of cells across pages; (3) A solution for merging cells across pages in complex PDFs is proposed, and this technical method is also applicable to the recognition and restoration of wireless tables. Therefore, this application can ensure that all parts of the merged table are a whole visually and functionally, and fully consider the integrity and accuracy of the content in the cells, which can improve the accuracy of merging cells in tables across pages.
[0156] Please refer to Figure 11 , Figure 11 which is a schematic structural diagram of a device for merging cells across pages provided by an embodiment of the present application. The device specifically includes:
[0157] An acquisition module 1110, configured to acquire a first file image and a second file image of a target document file; wherein, the file page corresponding to the second file image is the file page after the file page corresponding to the first file image, and the first file image includes a first table area;
[0158] A table detection module 1120, configured to perform table detection on the first file image based on a table area detection sub-model of a preset merging model to obtain first table area data and a first area label of the first table area; wherein, the first table area data is used to represent the position information of the first table area in the first file image;
[0159] A merging module 1130, configured to perform image merging on the first file image and the second file image based on the first area label and the first table area data to obtain a merged image of the table across pages;
[0160] The first feature extraction module 1140 is configured to perform image feature extraction on the cross-page table merging image based on the image feature extraction sub-model of the preset merging model to obtain cross-page image features;
[0161] The second feature extraction module 1150 is configured to perform text feature extraction on the cross-page table merging image based on the text feature extraction sub-model of the preset merging model to obtain cross-page text features;
[0162] The feature fusion module 1160 is configured to perform feature fusion on the cross-page image features and the cross-page text features based on the feature fusion sub-model of the preset merging model to obtain cross-page fusion features;
[0163] The cell detection module 1170 is configured to perform cell detection on the cross-page fusion features based on the cell detection sub-model of the preset merging model to obtain target cross-page cell data.
[0164] It should be noted that the cross-page cell merging device in the embodiments of the present application is used to implement the cross-page cell merging method in the above embodiments. The cross-page cell merging device in the embodiments of the present application corresponds to the foregoing cross-page cell merging method. For the specific processing process, please refer to the foregoing cross-page cell merging method, which will not be elaborated here.
[0165] The embodiments of the present application further provide an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the cross-page cell merging method in the embodiments of the present application. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0166] Please refer to Figure 12 , Figure 12 which schematically shows the hardware structure of an electronic device according to another embodiment. The electronic device includes:
[0167] The processor 1210 can be implemented in a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is configured to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0168] The memory 1220 can be implemented in the form of a Read Only Memory (ROM), a static storage device, a dynamic storage device, or a Random Access Memory (RAM), etc. The memory 1220 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1220 and are called by the processor 1210 to execute the cross-page cell merging method of the embodiments of this application;
[0169] The input / output interface 1230 is used to implement information input and output;
[0170] The communication interface 1240 is used to implement communication interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0171] The bus 1250 transmits information between the various components of the device (such as the processor 1210, the memory 1220, the input / output interface 1230, and the communication interface 1240);
[0172] Among them, the processor 1210, the memory 1220, the input / output interface 1230, and the communication interface 1240 are communicatively connected to each other inside the device through the bus 1250.
[0173] The embodiments of this application also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program is used to make a computer execute the cross-page cell merging method in the above embodiments.
[0174] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0175] The embodiments described in the embodiments of this application are for more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are equally applicable to similar technical problems.
[0176] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.
[0177] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0178] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.
[0179] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0180] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0181] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0182] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0183] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0184] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of this application. And the aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks or optical discs that can store programs.
[0185] The above has illustrated the preferred embodiments of the embodiments of this application with reference to the accompanying drawings, and thus does not limit the scope of rights of the embodiments of this application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall be within the scope of rights of the embodiments of this application.
Claims
1. A method for merging cells across pages, characterized in that: The method comprises: Acquire a first file image and a second file image of a target document file; wherein the file page corresponding to the second file image is a file page subsequent to the file page corresponding to the first file image, and the first file image includes a first table area; Performing table detection on the first document image based on the table area detection sub-model of the preset merged model to obtain first table area data and a first area label of the first table area; wherein the first table area data is used to represent the position information of the first table area in the first document image, and the first area label is used to indicate the area type of the first table area; Based on the first area label and the first table area data, the first document image and the second document image are merged to obtain a cross-page table merged image; Based on the image feature extraction sub-model of the preset merging model, image feature extraction is performed on the cross-page table merged image to obtain cross-page image features; Based on the text feature extraction sub-model of the preset merging model, text features are extracted from the cross-page table merged image to obtain cross-page text features; Based on the feature fusion sub-model of the preset merging model, the cross-page image feature and the cross-page text feature are subjected to feature fusion to obtain a cross-page fusion feature; The cell detection sub-model based on the preset merge model performs cell detection on the cross-page fusion feature to obtain target cross-page cell data.
2. The method according to claim 1, characterized in that The step of merging the first document image and the second document image based on the first area label and the first table area data to obtain a cross-page table merged image includes: Obtaining a first table subtag and a first cross-page subtag from the first area tag; wherein the first table subtag is used to indicate that the first table area is a cross-page table area or a non-cross-page table area, and the first cross-page subtag is used to indicate that the first table area is a cross-page table homepage area or a cross-page table continuation page area of the cross-page table area; If the first table sub-tag indicates that the first table area is the cross-page table area, performing table detection on the second document image based on the table area detection sub-model to obtain second table area data and a second area tag of the second table area in the second document image; Selecting a third document image from the second document image based on the second area tag and the first page-spread subtag; wherein the third document image includes a third table area associated with the first table area page-spread, and using the second table area data corresponding to the third document image as the third table area data; Based on the first table area data and the third table area data, the first document image and the third document image are merged to obtain the cross-page table merged image.
3. The method according to claim 2, characterized in that The step of merging the first file image and the third file image based on the first table area data and the third table area data to obtain the cross-page table merged image includes: extracting a first area image of the first table area from the first document image based on the first table area data; extracting a second area image of the third table area from the third document image based on the third table area data; Determine the cross-page region sequence data based on the first cross-page subtag and the third region tag of the third table region; The first region image and the second region image are spliced based on the cross-page region sequence data to obtain the cross-page table merged image.
4. The method according to claim 3, characterized in that The step of performing image stitching on the first region image and the second region image based on the cross-page region sequence data to obtain the cross-page table merged image includes: Remove the cross-page table lines in the first region image and update the first region image; removing the cross-page table lines in the second region image and updating the second region image; The updated first region image and the updated second region image are spliced based on the cross-page region sequence data to obtain the cross-page table merged image.
5. The method according to claim 1, characterized in that The feature fusion sub-model based on the preset merging model performs feature fusion on the cross-page image feature and the cross-page text feature to obtain the cross-page fusion feature, including: Determining a first text line feature and a second text line feature from the cross-page text feature, wherein a second text line corresponding to the second text line feature is a next text line of the first text line corresponding to the first text line feature, the first text line feature is used to describe text characters in the first text line that are arranged based on a corresponding text character sequence, and the second text line feature is used to describe text characters in the second text line that are arranged based on a corresponding text character sequence; Based on the text character sequence and the preset mask block, respectively masking the first text line feature and the second text line feature to obtain a first masked text line feature and a second masked text line feature; updating the first text line feature and the second text line feature based on the second text line, and performing mask processing on the updated first text line feature and the second text line feature, until all text line features in the cross-page text feature are masked to obtain masked text features; The cross-page image feature and the masked text feature are fused based on the feature fusion sub-model to obtain the cross-page fusion feature.
6. The method according to any one of claims 1 to 5, characterized in that: The step of acquiring the first file image and the second file image of the target document file comprises: Acquire a target document file, wherein the target document file includes a target document page; Performing image conversion on the target file page to obtain a target file image; The first document image and the second document image are determined from the target document image based on a document page sequence of the target document page.
7. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: Acquire cell structure data and cell content data from the target cross-page cell data; Table visualization processing is performed based on the cell structure data and the cell content data to obtain a cross-page merged table.
8. A cross-page cell merging device, characterized in that: The device comprises: An acquisition module, configured to acquire a first file image and a second file image of a target document file; wherein the file page corresponding to the second file image is a file page subsequent to the file page corresponding to the first file image, and the first file image includes a first table area; a table detection module, configured to perform table detection on the first document image based on a table region detection sub-model of a preset merged model, and obtain first table region data and a first region label of the first table region; wherein the first table region data is used to represent position information of the first table region in the first document image, and the first region label is used to indicate a region type of the first table region; a merging module, configured to merge the first file image and the second file image based on the first area label and the first table area data to obtain a cross-page table merged image; A first feature extraction module, configured to extract image features from the cross-page table merged image based on the image feature extraction sub-model of the preset merge model to obtain cross-page image features; A second feature extraction module is used to extract text features from the cross-page table merged image based on the text feature extraction sub-model of the preset merge model to obtain cross-page text features; A feature fusion module, used for performing feature fusion on the cross-page image feature and the cross-page text feature based on the feature fusion sub-model of the preset merging model to obtain a cross-page fusion feature; A cell detection module is used to perform cell detection on the cross-page fusion features based on the cell detection sub-model of the preset merging model to obtain target cross-page cell data.
9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Document table structure identification method oriented to water conservancy large model retrieval enhancement
CN120633613A
A document table structure recognition method for enhanced retrieval of large water conservancy models
CN120633613B