Image table restoration method and device, electronic equipment and storage medium
Through the image table restoration method, the table detection model and clustering processing technology are used to solve the problems of inefficient table restoration and insufficient accuracy in the existing technology, and more efficient and accurate table data restoration is achieved.
Patent Information
- Application Number
- CN202510204968.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is inefficient and prone to errors in the process of restoring table images, making it difficult to adapt to the diversity of table formats and contents, affecting accuracy.
A method of image table restoration is proposed. By obtaining the target table image, using the table detection model to perform cell detection, obtaining character data sets, performing clustering processing and data selection, combining edge character data for cell label detection and data integration, and generating target table data.
Improve the accuracy and efficiency of table restoration, can identify and correctly cluster edge character data that is not clustered correctly, reduce manual operation errors, and adapt to tabular data of various types and scales.
Smart Images

Figure CN120220170A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information processing, and in particular, to an image table restoration method and apparatus, an electronic device, and a storage medium. Background Art
[0002] In the information age, the restoration of a table image refers to the process of converting an image containing a table into an editable structured electronic table format. Therefore, accurate restoration of the table image can ensure the integrity and availability of table data.
[0003] Currently, related technologies usually adopt a manual operation method to restore the table content in a table image, but this method is inefficient and error-prone. Further, related technologies have also proposed a method of restoring the table content in a table image through image processing and machine learning technologies, but this method is difficult to adapt to the diversity of table forms and contents, thus affecting the accuracy of table restoration. Therefore, how to provide an image table restoration method that can improve the accuracy of table restoration is an urgent problem to be solved. Summary of the Invention
[0004] The main purpose of the embodiments of the present application is to propose an image table restoration method and apparatus, an electronic device, and a storage medium, which can improve the accuracy of table restoration.
[0005] To achieve the above object, a first aspect of the embodiments of the present application proposes an image table restoration method, and the method includes:
[0006] Obtain a target table image, where the target table image includes a target table area;
[0007] Perform cell detection on the target table area in the target table image based on a table detection model to determine table detection data corresponding to the target table area, where the table detection data includes cell data of each target cell in the target table area;
[0008] Obtain a character data set in the target table area, where the character data set includes first table character data;
[0009] Perform clustering processing on the first table character data based on the cell data to determine character clustering data of the first table character data, where the character clustering data includes a clustering cell label and clustering evaluation data, the clustering cell label is used to indicate the target cell to which the first table character data is clustered, and the clustering evaluation data is used to characterize the matching degree between the first table character data and the target cell to which it is clustered;
[0010] Perform data selection on the first tabular character data based on the cluster cell labels and the cluster evaluation data to determine the second tabular character data and the edge character data. The second tabular character data is used to represent the correctly clustered first tabular character data, and the edge character data is used to represent the first tabular character data that is not correctly clustered;
[0011] Perform cell label detection on the edge character data based on the second tabular character data to determine the edge cell labels of the edge character data;
[0012] Perform data integration based on the second tabular character data, the cluster cell labels corresponding to the second tabular character data, the edge character data, and the edge cell labels to obtain the target tabular data.
[0013] In some embodiments, the cell data includes initial cell labels and cell position data. The initial cell labels are used to indicate the types of tabular characters stored in the corresponding target cells, and the cell position data is used to indicate the position area of the corresponding target cell in the target tabular area; the first tabular character data includes first character position data, and the first character position data is used to indicate the position information of the corresponding first tabular character data in the target tabular area;
[0014] The clustering process of the first tabular character data based on the cell data to determine the character clustering data of the first tabular character data includes:
[0015] Calculate the relative distance between the cell position data and the first character position data to obtain the first relative distance data, and the first relative distance data is used to describe the positional relationship between the first character position data and the position area corresponding to the cell position data;
[0016] Select the initial cell labels based on the first relative distance data to determine the cluster cell labels of the first tabular character data;
[0017] Determine the cluster evaluation data of the first tabular character data based on the first relative distance data and the cluster cell labels;
[0018] Determine the character clustering data of the first tabular character data based on the cluster evaluation data and the cluster cell labels.
[0019] In some embodiments, the determining the cluster evaluation data of the first tabular character data based on the first relative distance data and the cluster cell labels includes:
[0020] Perform character type detection on the first table character data to obtain table character labels;
[0021] Determine label evaluation data for the first table character data based on the table character labels and the clustered cell labels, where the label evaluation data is used to characterize the matching degree between the table character labels and the clustered cell labels;
[0022] Determine distance evaluation data for the first table character data based on the first relative distance data;
[0023] Determine the clustering evaluation data for the first table character data based on the distance evaluation data and the label evaluation data.
[0024] In some embodiments, the edge character data includes edge character position data, the second table character data includes second character position data, and the second character position data is used to indicate the position information of the corresponding second table character data in the target table area;
[0025] The performing cell label detection on the edge character data based on the second table character data to determine the edge cell label of the edge character data includes:
[0026] Perform relative distance calculation on the edge character position data and the second character position data to obtain second relative distance data, where the second relative distance data is used to describe the positional relationship between the second table character data and the edge character position data;
[0027] Perform cell label detection on the edge character data based on the second relative distance data to determine the edge cell label of the edge character data.
[0028] In some embodiments, the performing cell label detection on the edge character data based on the second relative distance data to determine the edge cell label of the edge character data includes:
[0029] Perform mean calculation on the second relative distance data corresponding to the second table character data belonging to the same clustered cell label to obtain the relative distance mean;
[0030] Determine the edge cell label corresponding to the edge character data based on the relative distance mean.
[0031] In some embodiments, the edge character data includes edge character position data, the cell data includes cell center position data, and the performing cell label detection on the edge character data based on the second table character data to determine the edge cell label of the edge character data further includes:
[0032] Calculate the relative distance between the edge character position data and the cell center position data to obtain third relative distance data, where the third relative distance data is used to describe the positional relationship between the cell center position data and the edge character position data;
[0033] Based on the third relative distance data, perform cell label detection on the edge character data to determine the edge cell label of the edge character data.
[0034] In some embodiments, the obtaining of the target table image includes:
[0035] Obtain an original image, where the original image includes the target table area and first image content, and the first image content is used to describe the content in the original image that does not belong to the target table area;
[0036] Based on the target table area, perform image cropping on the original image to obtain the target table image.
[0037] To achieve the above object, a second aspect of the embodiments of the present application proposes an image table restoration device, where the device includes:
[0038] An image acquisition module, configured to acquire a target table image, where the target table image includes a target table area;
[0039] A cell detection module, configured to perform cell detection on the target table area in the target table image based on a table detection model to determine table detection data corresponding to the target table area, where the table detection data includes cell data of each target cell in the target table area;
[0040] A data acquisition module, configured to acquire a character data set in the target table area, where the character data set includes first table character data;
[0041] A clustering processing module, configured to perform clustering processing on the first table character data based on the cell data to determine character clustering data of the first table character data, where the character clustering data includes a clustering cell label and clustering evaluation data;
[0042] A data selection module, configured to perform data selection on the first table character data based on the clustering cell label and the clustering evaluation data to determine second table character data and edge character data;
[0043] A label detection module, configured to perform cell label detection on the edge character data based on the second table character data to determine the edge cell label of the edge character data;
[0044] A data integration module, configured to perform data integration based on the second tabular character data, the clustering cell labels corresponding to the second tabular character data, the edge character data, and the edge cell labels to obtain target tabular data.
[0045] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect above is implemented.
[0046] To achieve the above object, a fourth aspect of the embodiments of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect above is implemented.
[0047] The image table restoration method and device, electronic device, and storage medium provided by the present application obtain a target table image, which includes a target table area. Then, a table detection model is used to detect cells in the target table area to determine table detection data corresponding to the table area, where the table detection data includes cell data of each target cell in the target table area. Further, a character data set in the target table area is obtained, and the character data set includes first tabular character data. Clustering processing is performed on the first tabular character data based on the cell data to obtain character clustering data, which includes clustering cell labels and clustering evaluation data. Further, based on the clustering cell labels and the clustering evaluation data, data selection is performed on the first tabular character data to determine second tabular character data and edge character data. Further, cell label detection is performed on the edge character data based on the second tabular character data to determine edge cell labels of the edge character data. Finally, the second tabular character data, its corresponding clustering cell labels, the edge character data, and the edge cell labels are integrated to obtain complete target tabular data.
[0048] Compared with the related art, the present application can identify edge character data that is not correctly clustered from the first tabular character data through clustering evaluation data, and further detect the edge character data to correctly cluster the edge character data into the corresponding cell categories. Therefore, the embodiments of the present application can avoid the influence of edge character data that is not correctly clustered on the accuracy of table restoration, and effectively improve the accuracy of table restoration. Description of the Drawings
[0049] Figure 1 is a flowchart of the image table restoration method provided by the embodiments of the present application;
[0050] Figure 2 is Figure 1 the flowchart of step S101 in
[0051] Figure 3 is the detection result diagram of cell detection for the target table area in the target table image provided by the embodiment of the present application based on the table detection model;
[0052] Figure 4 is Figure 1 the flowchart of step S104 in
[0053] Figure 5 is Figure 4 the flowchart of step S403 in
[0054] Figure 6 is Figure 1 the flowchart of step S106 in
[0055] Figure 7 is Figure 6 the flowchart of step S602 in
[0056] Figure 8 is Figure 1 another flowchart of step S106 in
[0057] Figure 9 is the structural schematic diagram of the image table restoration device provided by the embodiment of the present application;
[0058] Figure 10 is the hardware structural schematic diagram of the electronic device provided by the embodiment of the present application. Detailed implementation manners
[0059] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0060] It should be noted that although functional module division is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first" and "second" in the description, claims and the above drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0061] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0062] First, several terms involved in this application are parsed as follows:
[0063] Artificial intelligence (AI): It is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. Artificial intelligence is a branch of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing, and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. It is also a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0064] Faster Regions with Convolutional Neural Networks (Faster R-CNN): It is an object detection algorithm based on deep learning, used for quickly identifying and locating objects in images. Faster R-CNN generates candidate regions through the Region Proposal Network (RPN) and extracts features using the Convolutional Neural Networks (CNN), thus achieving efficient object detection and localization.
[0065] You Only Look Once (YOLO): It is a real-time object detection algorithm based on deep learning, used for quickly identifying and locating objects in images. The YOLO algorithm transforms the object detection task into a single regression problem and directly predicts the class and bounding box position of the object from the input image, thus achieving efficient object detection.
[0066] Densely Connected Convolutional Networks (DenseNet) is a deep convolutional neural network architecture used for efficient feature extraction and image classification. DenseNet enhances feature propagation and gradient flow through dense connections (each layer is connected to all previous layers), thus alleviating the vanishing gradient problem and reducing the number of parameters.
[0067] Natural Language Processing (NLP): It is a technology that uses computers to process and understand human language. It is a branch of artificial intelligence and an interdisciplinary field of computer science and linguistics. NLP includes tasks such as syntactic analysis, semantic understanding, and discourse analysis, and is commonly used in areas such as machine translation, speech recognition, text classification, information extraction, and public opinion analysis.
[0068] Information Extraction: It is a technology that extracts specific information (such as entities, relationships, events, etc.) from text and converts it into structured data. Information extraction is to find useful information from a large amount of text, such as names of people, places, dates, etc., and organize it into a format that is convenient to use.
[0069] Image Caption: It is a technology that generates natural language descriptions for images to help programs understand the semantic content in the images. Image captioning can convert image retrieval into text retrieval and is used for image classification and improving retrieval results. This task requires converting the complex information in the image into natural language, not only identifying the objects in the image but also describing these objects and their relationships in natural language.
[0070] In the information age, tabular data often needs to be integrated and shared. By restoring the table image containing tabular data, tabular data from different sources can be integrated into a unified format. However, the forms of tables are diverse (such as wired tables, wireless tables, tables with few lines, etc.), and the table content may include non-text content such as images, formulas, symbols, as well as special content such as abbreviations, citations, and annotations. Therefore, the accurate restoration of table images can ensure the integrity and usability of tabular data.
[0071] Currently, related technologies usually adopt a manual operation method to restore the table content in table images, but this method is inefficient and error-prone. Further, related technologies have also proposed a method of restoring the table content in table images through image processing and machine learning technologies, but this method is difficult to adapt to the diversity of table forms and content, thus affecting the accuracy of table restoration.
[0072] Based on this, the embodiments of the present application provide an image table restoration method and device, an electronic device, and a storage medium, which can improve the accuracy and efficiency of table restoration.
[0073] The image table restoration method provided by the embodiments of the present application relates to the field of information processing. The image table restoration method provided by the embodiments of the present application can be applied to a terminal, or to a server side, or can also be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the image table restoration method, etc., but is not limited to the above forms.
[0074] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet-type devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0075] Please refer to Figure 1 , Figure 1 which is an optional flowchart of the image table restoration method provided by the embodiments of the present application, Figure 1 and the method in
[0076] Step S101: Obtain a target table image, where the target table image includes a target table area;
[0077] Step S102: Based on a table detection model, determine the cell data of each target cell in the target table area;
[0078] Step S103: Obtain the first table character data in the target table area;
[0079] Step S104: Based on the cell data, perform clustering processing on the first table character data to determine the clustering cell labels and clustering evaluation data of the first table character data;
[0080] Step S105: Based on the clustered cell labels and the clustered evaluation data, perform data selection on the first table character data to determine the second table character data and the edge character data;
[0081] Step S106: Based on the second table character data, perform cell label detection on the edge character data to determine the edge cell labels of the edge character data;
[0082] Step S107: Based on the second table character data, the clustered cell labels corresponding to the second table character data, the edge character data, and the edge cell labels, perform data integration to obtain the target table data.
[0083] In step S101 of some embodiments, the target table image is used to represent the image form of the table indicated by the target table area. The target table area is used to indicate the recorded position area of the table content in the target table image. For example, when the table indicated by the target table area is a wired table, the target table area may indicate the rectangular area enclosed by the boundary lines of the corresponding table; when the table indicated by the target table area is a wireless table, the target table area may indicate the area enclosed by the virtual boundary lines of the corresponding table, and the specific shape of the target table area is not limited.
[0084] It should be noted that the method for obtaining the target table image includes: an image containing a table can be extracted from an electronic document, and the electronic document may contain multiple document images; or, an image can be obtained by only photographing the target table area with a camera, and the specific acquisition method is not limited.
[0085] Please refer to Figure 2 In some embodiments, step S101 may include but is not limited to steps S201 to S202:
[0086] Step S201: Obtain the original image;
[0087] Step S202: Based on the target table area, perform image cropping on the original image to obtain the target table image.
[0088] In step S201 of some embodiments, the original image is image data including a target table area and other non-table content. For example, the original image may be a document image including a table, text, chart, background, or other elements, or a captured image including a table and other content. The specific type of the image is not limited. The original image includes a target table area and first image content, where the first image content is the part of the original image that does not belong to the target table area. For example, the original image may be a page image in an article including descriptive text and a table. At this time, the descriptive text is located outside the table, and the descriptive text may be visual elements unrelated to the table, such as text descriptions, titles, annotations, charts, background patterns, etc. around the table. Therefore, the descriptive text at this time is the first image content, and the area corresponding to the table is the target table area.
[0089] It should be noted that the original image can be obtained by extracting an image including table content and non-table content from an electronic document, or by capturing a picture including table content and non-table content with a camera. The specific acquisition method is not limited.
[0090] In step S202 of some embodiments, the present application may first perform table area detection on the original image through a table area detection model to determine the area position information of the target table area. The area position information at this time may be the area vertices of the target table area, which is used to determine the area range of the target table area. Further, the original image may be cropped according to the area position information of the target table area to obtain a target table image. In this way, the obtained target table image is an image that excludes the first image content and only includes the target table area.
[0091] In step S102 of some embodiments, the table detection model is a deep learning-based algorithm model for automatically identifying and locating the target table area in the target table image. Cell detection refers to the process of identifying and locating each cell within the target table area. Table detection data refers to the structured information extracted from the target table area, which is used to describe the layout of the table indicated by the target table area and the cell data. The table detection data includes the cell data of each target cell in the target table area. For example, the table detection data can be the type of each target cell (such as text, number, or image), the location data (such as vertex coordinates or center point coordinates), and the overall layout information of the table (such as the number of rows, the number of columns, etc.). A target cell refers to a single cell that needs to be identified and extracted within the target table area. For example, in the target table area, each cell can be a target cell, and the specific size of the target cell is not limited. Cell data refers to the specific information of each target cell. For example, the cell data can include an initial cell label and cell location data. Among them, the initial cell label refers to the type of table character stored in the corresponding target cell (such as text, number, or image, etc.); the cell location data refers to the position area of the corresponding target cell in the target table area (such as the coordinates of the vertex of the target cell or the coordinates of the center point, etc.).
[0092] It should be noted that through the table detection model of the present application, the table detection data corresponding to the target table area can be automatically obtained, thereby avoiding the cumbersome process of manual operation for processing a large amount of table data, and significantly improving the processing efficiency and acquisition accuracy of the table data. Moreover, the table detection model of the present application can flexibly process the target table image containing table data from different sources, achieve accurate detection of cells, and thus obtain table data from different sources. The table detection model of the present application has good flexibility and scalability, can adapt to various types and scales of table data, including tables with different formats and different complexities, and ensures high efficiency and accurate performance in diverse application scenarios.
[0093] In some embodiments, the present application can first obtain an image dataset containing tables. For example, a publicly available table image dataset or a self-created table image dataset. Then, the obtained table image dataset is preprocessed, and the preprocessed table image dataset is input into a preset layout detection model for pre-training to obtain a table detection model.
[0094] It should be noted that the preset layout detection model can be constructed based on different algorithms. For example, a layout detection model can be constructed with the Faster R-CNN algorithm as the basic architecture and DenseNet as the feature extractor; or, a layout detection model can be constructed with the YOLO algorithm as the basic architecture and DenseNet as the feature extractor. The specific construction method is not limited.
[0095] It should be noted that the preset layout detection model uses DenseNet as the feature extractor, which can extract more feature information and can better utilize this feature information during the training process, thereby improving the performance of the model. In addition, DenseNet also adopts a smaller growth rate (i.e., fewer newly added channels per layer), reducing the number of model parameters and computational complexity.
[0096] It should be noted that for preprocessing the table image dataset, the table images in the table image dataset can be first translated, rotated, enlarged, and reduced to enhance the adaptability of the model to different table layouts. At the same time, the grayscale value of the image is taken to reduce the interference of the position information and color of the text and charts on detection and recognition. Further, the present application can also add Gaussian blur to the table images in the table image dataset. Through this weighted average operation based on the Gaussian distribution, the image is smoothed to remove noise and detail interference, thereby improving the ability of the model to extract the overall structural features of the table. In addition, the present application can further enhance the table features through neighborhood detection, especially enhancing the features of regions with significant brightness changes such as table borders and text edges, thereby improving the accuracy and robustness of the model when detecting tables.
[0097] Please refer to Figure 3 , Figure 3 which is the detection result diagram of cell detection for the target table area in the target table image provided by the embodiment of the present application. Through Figure 3 it can be seen that the table image contains various different target cells, and each target cell contains a corresponding type (numbers and text). From the visual effect, the table detection model of the present application can achieve precise positioning and recognition of the cell data in each cell of the table under complex data formats, ensuring high precision in the recognition process.
[0098] In step S103 of some embodiments, the character data set refers to the set of all character-related data extracted from the target table area. For example, the character data set may be data such as text, borders, or numbers in the target table area. The character data set includes first table character data, which is the specific information of the table characters extracted from the target table area. For example, the position data of the characters, the specific information or value represented by the characters, and the type of the characters (such as text, numbers, etc.).
[0099] It should be noted that the character data set in the target table area can be obtained through a text parsing tool or a deep learning model, and the specific acquisition method is not limited.
[0100] In step S104 of some embodiments, clustering processing refers to an unsupervised machine learning method used to group the table characters in the table character data so that the table characters in the same cluster are similar to each other. The character clustering data refers to the output data obtained through the clustering processing process, including cluster cell labels and cluster evaluation data. The cluster cell labels are used to indicate the target cells to which the first table character data is clustered, and the cluster evaluation data is used to characterize the matching degree between the first table character data and the target cells to which it is clustered.
[0101] Please refer to Figure 4 , in some embodiments, the cell data includes initial cell labels and cell position data, and the first table character data includes first character position data. Step S104 may include but is not limited to steps S401 to S404:
[0102] Step S401, calculate the relative distance between the cell position data and the first character position data to obtain first relative distance data;
[0103] Step S402, select the initial cell labels based on the first relative distance data to determine the cluster cell labels of the first table character data;
[0104] Step S403, determine the cluster evaluation data of the first table character data based on the first relative distance data and the cluster cell labels;
[0105] Step S404, determine the character clustering data of the first table character data based on the cluster evaluation data and the cluster cell labels.
[0106] In step S401 of some embodiments, the first character position data is used to indicate the position information of the corresponding first table character data in the target table area, such as the coordinates of the edge area or the center coordinates of the character. The first relative distance data is used to describe the positional relationship between the first character position data and the position area corresponding to the cell position data.
[0107] It should be noted that by calculating the relative distance between the cell position data and the first character position data, the first relative distance data can be obtained. The coordinate positions of each cell and each character can be determined through the cell position data and the first character position data. Further, the Euclidean distance formula is used to calculate the distance from each character to each cell. Finally, the calculated distances are recorded to form the first relative distance data. For example, when the first character position data is A, and the data in A is "character 1-1(x 1,1 ,y 1,1 )", "character 1-2(x 1,2 ,y 1,2 )",..., "character i-n(x i,n ,y i,n )", where character i-n represents the character in the i-th column and the n-th row, and i and n represent the index numbers of the character, x i,n represents the x-axis coordinate of the character in the i-th column and the n-th row in the target table area, and y i,n represents the y-axis coordinate of the character in the i-th column and the n-th row in the target table area. The cell data position is B, and the data in B is "cell 1-1(x 1,1 ,y 1,1 )", "cell 1-2(x 1,2 ,y 1,2 )",..., "cell j-m(x j,m ,y j,m )". Among them, cell j-m represents the cell in the j-th column and the m-th row, and j and m represent the index values of the cell, x j,m represents the x-axis coordinate of the cell in the j-th column and the m-th row in the target table area, and y j,m represents the y-axis coordinate of the cell in the j-th column and the m-th row in the target table area. The Euclidean distance formula is shown as formula (1) below:
[0108]
[0109] where d 1-1 represents the relative distance between the character in the first column and the first row and the cell in the first column and the first row.
[0110] Through formula (1), the distance from each character to each cell can be calculated, and the calculated distances are recorded to form the first relative distance data.
[0111] In step S402 of some embodiments, the present application can determine the cell closest to each character through the first relative distance, and then determine the clustering cell label of the first table character data according to the initial cell label corresponding to the cell closest in distance.
[0112] In steps S403 and S404 of some embodiments, the clustering evaluation data can be determined by calculating the minimum distance from each character to each cell through the first relative distance. This minimum distance directly reflects the degree of proximity between the character and the cell. The smaller the distance, the higher the matching degree. For example, if the minimum distance between character 1-1 and cell 2-1 is the shortest, then it can be considered that the matching degree between character 1-1 and cell 2-1 is the highest. Such evaluation results are recorded to form the clustering evaluation data. The character clustering data can be directly obtained through the determined clustering evaluation data and the clustering cell labels.
[0113] Please refer to Figure 5 , in some embodiments, step S403 may include but is not limited to steps S501 to S504:
[0114] Step S501, perform character type detection on the first table character data to obtain table character labels;
[0115] Step S502, determine the label evaluation data of the first table character data based on the table character labels and the clustering cell labels;
[0116] Step S503, determine the distance evaluation data of the first table character data based on the first relative distance data;
[0117] Step S504, determine the clustering evaluation data of the first table character data based on the distance evaluation data and the label evaluation data.
[0118] In step S501 of some embodiments, character type detection refers to identifying and classifying the type of character data, such as numbers, text, dates, etc. The table character labels can be obtained by performing character type detection on the first table character data through a text recognition tool or a deep learning method. The specific manner of character type detection is not limited.
[0119] In step S502 of some embodiments, the label evaluation data is used to characterize the matching degree between the table character labels and the clustering cell labels. For example, if the type of the character is a number and the type of the cell is also a number, then the matching degree is considered high. Another example is that if the type of the character is a number and the type of the cell is text, then the matching degree is considered low.
[0120] In step S503 of some embodiments, the distance evaluation data is determined by comparing the relative distance between the character and the cell with a preset distance threshold. The preset distance threshold is a fixed numerical standard for judging the relationship between the character and the cell. For example, if the relative distance between the character and the cell is less than the preset distance threshold, it can be determined that the matching degree between the character and the cell is relatively high. Or, if the distance between the character and the cell is greater than or equal to the preset distance threshold, the matching degree is determined to be relatively low. Such evaluation results are recorded to form the distance evaluation data. By determining the distance evaluation data, the distance evaluation data and the label evaluation data can be mapped to the same data dimension, facilitating subsequent data processing.
[0121] In step S504 of some embodiments, in order to improve the determination accuracy of the clustering evaluation data, the present application can determine the clustering evaluation data of the first table character data based on the distance evaluation data and the label evaluation data. In this way, from the perspectives of the label and the character content, the association degree between the first table character data and other clustering cell labels can be considered simultaneously. Specifically, the process of determining the clustering evaluation data of the first table character data based on the distance evaluation data and the label evaluation data can be, for example, through summation calculation, weighted calculation, table mapping, etc., without limitation.
[0122] In step S105 of some embodiments, the data selection is performed based on the clustering cell label and the clustering evaluation data, with the aim of classifying the first table character data into two categories: the second table character data and the edge character data. The second table character data is used to represent the first table character data that is correctly clustered. For example, if a character is successfully clustered into a cell, and the clustering evaluation data shows that the matching degree between the character and the cell is very high, then this character will be classified into the second table character data. The edge character data is used to represent the first table character data that is not correctly clustered. For example, if the clustering evaluation data of a character shows that its matching degree with any cell is very low, or the character is wrongly clustered into a cell that does not match its type, then this character will be classified into the edge character data.
[0123] In step S106 of some embodiments, the cell label detection is a process of analyzing and classifying the edge character data by using the known second table character data. The edge cell label refers to the label assigned to the edge character data.
[0124] It should be noted that the cell label detection of the edge character data based on the second table character data can be performed by calculating the relative distance between the edge character position data and the character position data in the second table character data, and judging the edge cell label of the edge character data according to the relative distance.
[0125] Please refer to Figure 6 In some embodiments, the edge character data includes edge character position data, the second tabular character data includes second character position data, and step S106 may further include but is not limited to steps S601 to S602:
[0126] Step S601, calculate the relative distance between the edge character position data and the second character position data to obtain second relative distance data;
[0127] Step S602, based on the second relative distance data, perform cell label detection on the edge character data to determine the edge cell label of the edge character data.
[0128] In step S601 of some embodiments, the edge character position data is the position information of the character data that has not been correctly clustered in the target table area, such as the coordinates of the edge area or the center coordinates of the character. The second character position data is used to indicate the position information of the corresponding second tabular character data in the target table area. The relative distance calculation refers to calculating the distance between the edge character position data and the second character position data to determine their relative position relationship. The second relative distance data is used to describe the position relationship between the second tabular character data and the edge character position data.
[0129] It should be noted that the second relative distance data can be obtained by calculating the distance between the edge character position data and the second character position data through the Euclidean distance formula, and the specific calculation principle is the same as that of step S401.
[0130] In step S602 of some embodiments, the cell label detection refers to the process of using the calculated second relative distance data to assign the correct cell label to the edge character data.
[0131] It should be noted that based on the second relative distance data, the cell label detection of the edge character data can determine the second tabular character data closest to the edge character data through the second relative distance. The cluster cell label to which this closest second tabular character data belongs is the edge cell label of the edge character data.
[0132] Please refer to Figure 7 In some embodiments, step S602 may further include but is not limited to steps S701 to S702:
[0133] Step S701, calculate the mean value of the second relative distance data corresponding to the second tabular character data belonging to the same cluster cell label to obtain the relative distance mean value;
[0134] Step S702, based on the relative distance mean value, determine the edge cell label corresponding to the edge character data.
[0135] In steps S701 to S702 of some embodiments, the mean calculation refers to a mathematical operation of calculating the average of a set of data, and its purpose is to calculate a value that can represent the trend in the second table character data with the same cluster cell label. The relative distance mean is the average calculated for the second relative distance data between the second table character data with the same cluster cell label and the edge character position data.
[0136] It should be noted that to calculate the mean of the second relative distance data corresponding to the second table character data with the same cluster cell label, the second table character data can be classified according to the cluster cell label first. Further, the second relative distance data corresponding to the edge character position data is extracted from each class of data. Finally, the second relative distance data in each class is summed and divided by the number of data to obtain the relative distance mean under the cluster cell label. For example, when the cluster cell labels are a, b, c, the second table character data corresponding to a are D1, D2, the second table character data corresponding to b are D3, D4, D5, the second table character data corresponding to c are D6, D7, D8, and the edge character data is E, the corresponding second relative distances d1, d2, d3, d4, d5, d6, d7, d8 can be determined through each cluster cell label. The relative distance mean da corresponding to a is obtained by summing d1 and d2 and dividing by 2, the relative distance mean db corresponding to b is obtained by summing d3, d4, d5 and dividing by 3, and the relative distance mean dc corresponding to c is obtained by summing d6, d7, d8 and dividing by 3.
[0137] It should be noted that to determine the edge cell label corresponding to the edge character data based on the relative distance mean, the relative distance means under different cluster cell labels can be compared, and the cluster cell label corresponding to the smallest relative distance mean is selected as the edge cell label of the edge character data. For example, by comparing the magnitudes of the relative distance means da, db, dc, the cluster cell label corresponding to the smallest relative distance mean can be selected as the edge cell label of the edge character data E.
[0138] Please refer to Figure 8 , in some embodiments, the edge character data includes edge character position data, the cell data includes cell center position data, and step S106 may further include but is not limited to steps S801 to S802:
[0139] Step S801, perform a relative distance calculation on the edge character position data and the cell center position data to obtain third relative distance data;
[0140] Step S802: Detect the cell labels of the edge character data based on the third relative distance data to determine the edge cell labels of the edge character data.
[0141] In steps S801 to S802 of some embodiments, the cell center position data refers to the position information of the geometric center of each cell in the target table area. The third relative distance data is used to describe the positional relationship between the cell center position data and the edge character position data.
[0142] It should be noted that the third relative distance can be obtained by calculating the distance between the edge character position data and the cell center position data through the Euclidean distance formula, and the specific calculation principle is the same as that of step S401.
[0143] It should be noted that detecting the cell labels of the edge character data based on the third relative distance data can be achieved by comparing the third relative distances of each edge character with the centers of different cells. And select the clustering cell label corresponding to the minimum relative distance as the edge cell label of the edge character data.
[0144] In step S107 of some embodiments, data integration refers to combining scattered and different types of data according to certain rules and logic to form a data whole with a complete structure and clear meaning. The target table data is the data result obtained after data integration, which can accurately restore the original table structure and content, and contains the character information of each cell in the table and the corresponding cell labels.
[0145] It should be noted that based on the second table character data, its corresponding clustering cell labels, edge character data, and edge cell labels for data integration, the basic framework of the target table can be constructed through the clustering cell labels. Further, fill the second table character data into the corresponding positions of the framework according to the corresponding labels to build the main body of the table. Then, add the edge character data to the corresponding places in the table according to the edge cell labels to complete the content. Finally, adjust and optimize the table format according to actual needs to make the data display clearer, more beautiful, and easier to use. For example, when the target table data is C, C contains C1, C2,..., C k multiple target cells. The characters in C1 are "character 1-1(x 1,1 ,y 1,1 )", "character 1-2(x 1,2 ,y 1,2 )", "...", "character 1-n(x 1,n ,y 1,n )", the characters in C2 are "character 2-1(x 2,1 ,y 2,1 )", "character 2-2(x 2,2 ,y2,2 )”, “…”, “Character 2-n(x 2,n , y 2,n )”, …, C k The characters within are “Character k-1(x k,1 , y k,1 )”, “Character k-2(x k,2 , y k,2 )”, “…”, “Character k-n(x k,n , y k,n )”. Among them, C k represents the character content contained in the k-th cell, x k,n represents the x-axis coordinate of the n-th character in the k-th cell, y k,n represents the y-axis coordinate of the n-th character in the k-th cell. By performing two-dimensional sorting based on the coordinates of the characters contained in each cell in C, the complete and ordered cell content can be obtained. Then, according to the labels corresponding to each character, the content of each cell is added to the corresponding position in the table. Finally, a border can be added to construct a wired table.
[0146] It should be noted that through data integration, the embodiments of the present application can integrate the parsed table data into a unified format, such as CSV, Excel, or PDF, etc., which is convenient for sharing and analysis across departments or organizations. This integration not only improves the usability of the data but also provides a basis for applications such as data analysis, data mining, and machine learning. In addition, the embodiments of the present application can also convert the table data into a machine-readable format, enabling more efficient in-depth analysis to discover the patterns, trends, and correlations in the data, thereby providing strong support for decision-making.
[0147] Please refer to Figure 9 , the embodiments of the present application also provide an image table restoration method device, which can implement the above image table restoration method. The device includes:
[0148] An image acquisition module 901, configured to acquire a target table image, where the target table image includes a target table area;
[0149] A cell detection module 902, configured to perform cell detection on the target table area in the target table image based on a table detection model to determine table detection data corresponding to the target table area, where the table detection data includes cell data of each target cell in the target table area;
[0150] A data acquisition module 903, configured to acquire a character data set in the target table area, where the character data set includes first table character data;
[0151] The clustering processing module 904 is configured to perform clustering processing on the first tabular character data based on the cell data, and determine the character clustering data of the first tabular character data, where the character clustering data includes clustering cell labels and clustering evaluation data;
[0152] The data selection module 905 is configured to perform data selection on the first tabular character data based on the clustering cell labels and the clustering evaluation data, and determine the second tabular character data and the edge character data;
[0153] The label detection module 906 is configured to perform cell label detection on the edge character data based on the second tabular character data, and determine the edge cell labels of the edge character data;
[0154] The data integration module 907 is configured to perform data integration based on the second tabular character data, the clustering cell labels corresponding to the second tabular character data, the edge character data, and the edge cell labels, so as to obtain the target tabular data.
[0155] The specific implementation manner of this device is basically the same as that of the specific embodiment of the above image table restoration method, and will not be elaborated here.
[0156] An embodiment of this application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above image table restoration method is implemented. This electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0157] Please refer to Figure 10 , Figure 10 which schematically shows the hardware structure of an electronic device according to another embodiment. The electronic device includes:
[0158] The processor 1001 can be implemented by using a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of this application;
[0159] The memory 1002 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1002 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1002 and are called by the processor 1001 to execute the image table restoration method of the embodiments of this application;
[0160] The input / output interface 1003 is used to implement information input and output;
[0161] The communication interface 1004 is used to implement communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0162] The bus 1005 transmits information between the various components of the device (such as the processor 1001, the memory 1002, the input / output interface 1003, and the communication interface 1004);
[0163] Among them, the processor 1001, the memory 1002, the input / output interface 1003, and the communication interface 1004 achieve communication connections with each other inside the device through the bus 1005.
[0164] The embodiments of this application also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned image table restoration method is implemented.
[0165] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0166] The embodiments described in the embodiments of this application are for more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art can know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are equally applicable to similar technical problems.
[0167] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.
[0168] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0169] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.
[0170] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0171] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a, b, and c", where a, b, and c can be single or multiple.
[0172] In several embodiments provided by this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0173] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0174] In addition, in each embodiment of this application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0175] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of this application. And the aforementioned storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0176] The preferred embodiments of the embodiments of this application have been described above with reference to the accompanying drawings, but this does not limit the scope of the rights of the embodiments of this application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall be within the scope of the rights of the embodiments of this application.
Claims
1. A method for restoring an image table, characterized in that: The method comprises: Acquire a target table image, wherein the target table image includes a target table area; Perform cell detection on the target table area in the target table image based on a table detection model, and determine table detection data corresponding to the target table area, wherein the table detection data includes cell data of each target cell in the target table area; Acquire a character data set in the target table area, wherein the character data set includes first table character data; Performing clustering processing on the first table character data based on the cell data to determine character clustering data of the first table character data, the character clustering data including a clustering cell label and clustering evaluation data, the clustering cell label being used to indicate the target cell to which the first table character data is clustered, and the clustering evaluation data being used to characterize the degree of matching between the first table character data and the target cell to which the first table character data is clustered; Performing data selection on the first table character data based on the cluster cell labels and the cluster evaluation data to determine second table character data and edge character data, wherein the second table character data is used to characterize the first table character data that is correctly clustered, and the edge character data is used to characterize the first table character data that is not correctly clustered; Performing cell label detection on the edge character data based on the second table character data to determine edge cell labels of the edge character data; Data integration is performed based on the second table character data, the cluster cell labels corresponding to the second table character data, the edge character data, and the edge cell labels to obtain target table data.
2. The method according to claim 1, characterized in that The cell data includes an initial cell label and cell position data, wherein the initial cell label is used to indicate the type of table character stored in the corresponding target cell, and the cell position data is used to indicate the position area of the corresponding target cell in the target table area; The first table character data includes first character position data, and the first character position data is used to indicate the position information of the corresponding first table character data in the target table area; The clustering process of the first table character data based on the cell data to determine the character clustering data of the first table character data includes: Performing relative distance calculation on the cell position data and the first character position data to obtain first relative distance data, wherein the first relative distance data is used to describe a positional relationship between the first character position data and a position area corresponding to the cell position data; Selecting the initial cell label based on the first relative distance data to determine the clustered cell label of the first table character data; determining the clustering evaluation data of the first table character data based on the first relative distance data and the clustering cell labels; Character clustering data of the first table character data is determined based on the clustering evaluation data and the clustering cell labels.
3. The method according to claim 2, characterized in that The step of determining the clustering evaluation data of the first table character data based on the first relative distance data and the clustering cell labels comprises: Performing character type detection on the first table character data to obtain a table character label; Determining label evaluation data of the first table character data based on the table character label and the cluster cell label, wherein the label evaluation data is used to characterize a matching degree between the table character label and the cluster cell label; determining distance evaluation data of the first table character data based on the first relative distance data; The cluster evaluation data of the first table character data is determined based on the distance evaluation data and the label evaluation data.
4. The method according to claim 1, characterized in that: The edge character data includes edge character position data, and the second table character data includes second character position data, and the second character position data is used to indicate the position information of the corresponding second table character data in the target table area; The performing cell label detection on the edge character data based on the second table character data to determine the edge cell label of the edge character data includes: Performing relative distance calculation on the edge character position data and the second character position data to obtain second relative distance data, wherein the second relative distance data is used to describe the positional relationship between the second table character data and the edge character position data; Cell label detection is performed on the edge character data based on the second relative distance data to determine the edge cell label of the edge character data.
5. The method according to claim 4, characterized in that The performing cell label detection on the edge character data based on the second relative distance data to determine the edge cell label of the edge character data includes: Calculating the mean of the second relative distance data corresponding to the second table character data belonging to the same cluster cell label to obtain the relative distance mean; The edge cell label corresponding to the edge character data is determined based on the relative distance mean.
6. The method according to claim 1, characterized in that The edge character data includes edge character position data, the cell data includes cell center position data, and the cell label detection is performed on the edge character data based on the second table character data to determine the edge cell label of the edge character data, and further includes: Performing relative distance calculation on the edge character position data and the cell center position data to obtain third relative distance data, wherein the third relative distance data is used to describe the positional relationship between the cell center position data and the edge character position data; Cell label detection is performed on the edge character data based on the third relative distance data to determine the edge cell label of the edge character data.
7. The method according to claim 1, characterized in that The step of obtaining a target table image comprises: Acquire an original image, wherein the original image includes the target table area and first image content, wherein the first image content is used to describe content in the original image that does not belong to the target table area; The original image is cropped based on the target table area to obtain the target table image.
8. An image table restoration device, characterized in that: The device comprises: An image acquisition module, used for acquiring a target table image, wherein the target table image includes a target table area; A cell detection module, configured to perform cell detection on the target table area in the target table image based on a table detection model, and determine table detection data corresponding to the target table area, wherein the table detection data includes cell data of each target cell in the target table area; A data acquisition module, used for acquiring a character data set in the target table area, wherein the character data set includes first table character data; A cluster processing module, used for performing cluster processing on the first table character data based on the cell data, and determining character cluster data of the first table character data, wherein the character cluster data includes cluster cell labels and cluster evaluation data; A data selection module, configured to select the first table character data based on the cluster cell labels and the cluster evaluation data, and determine the second table character data and the edge character data; A label detection module, used for performing cell label detection on the edge character data based on the second table character data, and determining edge cell labels of the edge character data; The data integration module is used to perform data integration based on the second table character data, the cluster cell labels corresponding to the second table character data, the edge character data, and the edge cell labels to obtain target table data.
9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.