Method, apparatus, and storage medium for generating table images and corresponding annotation information
By generating table samples and annotating pictures, the problem of high cost of data acquisition and annotation of table pictures is solved, and an efficient data set generation solution is provided, which is suitable for deep learning model training.
Patent Information
- Application Number
- CN202210203324.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-03
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-03-03
AI Technical Summary
In some fields, table image data acquisition is difficult and labeling is expensive, and the complete data set is lacking for deep learning model training.
By collecting scene information, forming a corpus, defining table parameters and recording it in a configuration file, combining corpus rendering to generate table samples and annotation pictures, extracting annotation information, generating annotation files, and forming a data set for training and verification of neural network models.
It realizes the random generation of table pictures in large batches and obtains labeling information under specified corpus and layout, solves the problem of difficult and high labeling cost for table picture data collection, and provides an efficient data set generation method.
Smart Images

Figure CN114581923B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of tabular data, and particularly to a method, apparatus, and storage medium for generating tabular images and corresponding annotation information. Background Art
[0002] Currently, tables are a commonly used information expression form in various types of documents. With the increasing demand for data processing in various industries and the rapid development of artificial intelligence technology, more enterprises tend to use deep learning models to process tabular data stored in image form, replacing traditional manual parsing. Common tabular tasks include: table detection, table structure parsing, table line detection, table text localization and recognition, etc. Generally, the training of deep learning neural network models requires a large amount of training data, including the tabular images themselves and corresponding annotations. However, in some fields, relevant tabular image data has sensitive or confidential attributes or the cost of image acquisition is relatively high, resulting in a serious lack of training data. On the other hand, even if a large amount of data is collected, the manual annotation of tabular images is extremely costly. The lack of complete tabular images and annotations brings great difficulties to the development of deep learning models.
[0003] In other deep learning fields related to document recognition, attempts have been made to use specified corpora to synthesize training samples to address the problem of scarce readable data. However, there are few public methods for generating tabular image datasets, mainly because the annotation cost after generating samples is too high, and the results cannot be directly used for model training and verification. Therefore, there is an urgent need for a method for generating tabular images and data annotations. Summary of the Invention
[0004] To solve at least one of the problems mentioned in the above background art, the present invention provides a method, apparatus, and storage medium for generating tabular annotation information, which can randomly generate a large number of tabular images in a specified corpus and format, and simultaneously obtain all annotation information in the images, solving the problems of high difficulty in collecting tabular image data and high annotation cost in specific scenarios.
[0005] The specific technical solutions provided by the embodiments of the present invention are as follows:
[0006] In a first aspect, a method for generating a tabular image and corresponding annotation information, the method comprising:
[0007] Collect the corpus corresponding to the scenario information according to the scenario information to form a corpus;
[0008] Define table parameters and record them in a configuration file;
[0009] Render and generate tabular sample images and annotation images in combination with the configuration file and the corpus;
[0010] Extract the annotation images to generate annotation information.
[0011] Further, before generating the labeled image, it further includes: adding and modifying the text background color code in the configuration file to extract the labeling information according to the configuration file.
[0012] Further, it further includes: batch generating the table sample images and the labeling information to form a data set, and using the data set to train and / or verify the neural network model.
[0013] Further, defining the table parameters specifically includes:
[0014] Defining the table structure and / or table content and / or table style according to the scenario information.
[0015] Further, extracting the labeled image and generating the labeling information specifically includes:
[0016] Extracting the labeled image to obtain the text coordinates and the frame line coordinates;
[0017] Combining the configuration file, the text coordinates and the frame line coordinates to obtain the labeling information.
[0018] Further, separating the labeled image to obtain a text labeling image and a table labeling image;
[0019] Extracting the contour coordinates based on the text labeling image to generate text coordinates;
[0020] Performing contour detection on the table labeling image to obtain the frame line coordinates.
[0021] Further, binarizing the table labeling image to obtain a table binarized image;
[0022] Eroding and dilating the table binarized image to separate the horizontal lines in the table binarized image to obtain a horizontal line image;
[0023] Eroding and dilating the table binarized image to separate the vertical lines in the table binarized image to obtain a vertical line image;
[0024] Extracting the frame line coordinates according to the horizontal line image and the vertical line image.
[0025] In a second aspect, there is provided an apparatus for a method of generating based on a table image and corresponding labeling information, the apparatus including:
[0026] A corpus configuration module for collecting the corpus corresponding to the scenario information according to the scenario information to form a corpus;
[0027] A file configuration module for defining the table parameters and recording them in a configuration file;
[0028] A rendering module, configured to collect the configuration file and the prediction library, and render and generate a table sample picture and an annotation picture;
[0029] An extraction module, configured to extract the annotation picture and generate annotation information.
[0030] In a third aspect, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:
[0031] Collect the corpus corresponding to the scenario information according to the scenario information to form a corpus;
[0032] Define table parameters and record them in a configuration file;
[0033] Combine the configuration file and the corpus, and render and generate a table sample picture and an annotation picture;
[0034] Extract the annotation picture and generate annotation information.
[0035] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0036] Collect the corpus corresponding to the scenario information according to the scenario information to form a corpus;
[0037] Define table parameters and record them in a configuration file;
[0038] Combine the configuration file and the corpus, and render and generate a table sample picture and an annotation picture;
[0039] Extract the annotation picture and generate annotation information.
[0040] The embodiments of the present invention have the following beneficial effects:
[0041] 1. According to the data characteristics of scenarios such as banks and financial statements, collect the corpus of the corresponding scenario information to form a corpus. At the same time, according to the table data characteristics of the scenario where data needs to be collected, define the table characteristics and record them in a configuration file. Combine the configuration file and the corpus content, and render and generate a table sample picture and an annotation picture, where the annotation picture and the table sample picture are pictures with the same layout. Extract the annotation picture to obtain annotation information;
[0042] 2. Before generating the annotation information, add and modify the text background color code in the configuration file to realize the extraction of annotation information according to the configuration. When obtaining the position of the text box, modify the text background color in the sample picture code so that the text area is enveloped by a solid red rectangle for easy position extraction;
[0043] 3. Store the obtained annotation information as an annotation file, and combine the annotation file with the obtained table sample pictures to batch generate table sample pictures and annotation files, forming a dataset and a training set. Use the dataset and the training set to train and validate the neural network model respectively. Brief Description of the Drawings
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following-described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0045] Figure 1 A schematic diagram for embodying the method in Embodiment 1 of the present application;
[0046] Figure 2 A schematic diagram for embodying the table sample pictures in the present application;
[0047] Figure 3 A schematic diagram for embodying the annotation pictures in the present application;
[0048] Figure 4 A horizontal line diagram for embodying the present application;
[0049] Figure 5 A vertical line diagram for embodying the present application;
[0050] Figure 6 A schematic diagram of the structure of the server for embodying the present application. Detailed Description of the Embodiments
[0051] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some, rather than all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0052] At present, tables are a common way to express information in various documents. With the growing demand for data processing in various industries and the rapid development of artificial intelligence technology, more companies tend to use deep learning models to process table data stored in the form of images, replacing traditional manual analysis. Common table tasks include: table detection, table structure analysis, table line detection, table text positioning and recognition, etc. The training of general deep learning neural network models requires a large amount of training data, including the table image itself and the corresponding annotations. However, in some fields, the relevant table image data has sensitive or confidential attributes or the image collection cost is high, and there is a serious lack of training data. On the other hand, even if a large amount of data is collected, the manual annotation of table images is extremely costly. The lack of complete table images and annotations brings great difficulties to the development of deep learning models. At the same time, using the chrome kernel to render into images, the absolute coordinates of the rendering result images based on html code are uncontrollable, so the annotation position cannot be directly obtained; for wireless tables and three-line tables, the distinction between lines and cells is not displayed after rendering. Based on the above problems, this application proposes a method, device and storage medium for generating table images and corresponding annotation information, which can randomly generate table images in large quantities under specified corpus and layout, and obtain all annotation information in the images at the same time, solving the problems of difficulty in collecting table image data and high annotation cost in specific scenarios.
[0053] Embodiment 1
[0054] A method for generating a table image and corresponding annotation information is provided, the method comprising:
[0055] Step S1:
[0056] According to the scene information, corpus corresponding to the scene information is collected to form a corpus.
[0057] Specifically, in view of the data characteristics of banks, finance, and insurance, and for specific scenarios where table images need to be generated, relevant corpus is collected as a database for the content filled in the table cells. For example, in the bank reconciliation detail table scenario, the text corpus pool contains the desensitized user names, transaction behavior names, etc. as the corresponding corpus, and for field scenarios with specific encoding rules such as dates and amounts, a corpus generator can be involved to automatically and randomly generate corpus in the corresponding format. By combining the above user input corpus with the corpus generated by the corpus generator, a corpus is formed to support the subsequent content that needs to be filled.
[0058] Step S2:
[0059] Define table parameters and record them in the configuration file.
[0060] Specifically, according to specific scenario information, define the table structure and / or table content and / or table style, and record them in the configuration file. Specifically, defining the table structure includes defining the number of rows and columns of the table, and setting the number of rows and columns of the table to randomly vary within a certain range. Defining the table content includes defining the corpus type of the content of each row or column of the table in units of rows or columns of the table, such as text, date, amount, etc. Defining the table style includes defining the layout characteristics of each row or column of the table, for example: whether there are border lines, text alignment methods (such as left alignment, center alignment, right alignment, etc.), font types and font colors, etc.
[0061] Specifically, except for the first row header, the first column of the table content is the date column, the second column is the text column, and the third column is the amount column. Set the width ratio to 1:3:1. The even rows are gray and the odd rows are blue. The table has no vertical lines but only horizontal lines. This solution can be independently defined in the configuration file according to the actual scenario and the required content.
[0062] Step S3:
[0063] Combine the configuration file and the corpus to render and generate a table sample picture and an annotation picture. Specifically, first combine the configuration file and the corpus to render and generate a table sample picture, then add and modify the text background color code in the configuration file, and then combine the configuration file with the added and modified text background color code and the corpus to render and generate an annotation picture. Use the from_string method in the open-source tool imgkit to render and generate the table sample picture and the annotation picture respectively with the above codes.
[0064] Specifically, the codes for generating the table sample picture and the annotation picture are exactly the same in the HTML part, and ensure that their structures and contents are the same. Add additional text background color modification codes in the CSS code part for generating the annotation picture. Among them, the table sample picture is as Figure 2 shown, and the code for generating the table sample picture is as follows:
[0065] According to the structure and content specified in the configuration file, use a python script (it can also be other languages, which is not limited in the present invention) to generate the corresponding HTML string. When generating, according to the syntax rules of HTML, Combine the above tags with the text content and specify the class label class_id for each column. Additionally, generate the corresponding The CSS code for the label controls the layout effect of the table sample image. When it is necessary to change the style of an entire row or column, the ID selector of CSS can be used for batch setting.
[0066] Specifically, the labeled image is as Figure 3 shown, and the labeled image code is as follows:
[0067] To obtain the position of the text box and modify the background color of the text in the table sample image, add the code for modifying the background color of the text in the CSS section. For example: in <style>标签内添加"span{background-color:red;color:red;}”即可使文字区域被实心的红色矩形框包络,以便进行位置提取。如需对样本图片中不可见的表格线提取坐标,可将对应类别设置为"{border:1pxsolid black;}”。
[0068] 步骤S4:
[0069] 提取标注图片,生成标注信息。
[0070] 针对表格样本图片,如使用chrome内核渲染成图片,获取标注有以下困难:其一,基于html代码的渲染结果图片,其绝对坐标是不可控的,因此无法直接得到标注位置。其二,对于无线表、三线表,其渲染后本身就不显示线和单元格的区分,因此,本申请中采用生成标注图片的方法,通过对标注图片进行检测和提取,获取标注信息。
[0071] 具体包括,步骤S4.1:提取标注图片,得到文字坐标和框线坐标。
[0072] 标注图片中的文字区域为黑色矩形框,表格框线为黑色直线。首先,通过颜色RGB通管分离操作将文字标注和表格标注分离开,得到文字标注图和表格标注图。
[0073] 基于文字标注图提取轮廓坐标,生成文字坐标,基于表格标注图,进行轮廓检测,得到框线坐标。具体的,对文字标注图使用OpenCV中的findContours方法获取外轮廓坐标,即可得到所有文字坐标。接着对表格标注图进行二值化,得到表格二值化图,利用OpenCV对图分别做核为(3,1)的开运算,即对表格二值化图像进行3次腐蚀和3次膨胀,分离出表格二值化图像中的横线,得到横线图;同理,将核替换为(3,1)进行处理,即对表格二值化图像进行3次腐蚀和3次膨胀,分离出表格二值化图像中的竖线,得到竖线图。
[0074] 根据如图4和5所示的横线图和竖线图,提取得到框线坐标,具体的,对横线图和纵线图使用findContours方法获取轮廓坐标,可得到横 / 纵线的端点表示,即得到框线坐标。
[0075] 步骤S4.2:结合所述配置文件、所述文字坐标和所述框线坐标,得到标注信息。
[0076] 具体的,通过解析框线坐标可得到表格范围坐标。配置文件中的HTML文件中已随机生成完毕每个字段的语料内容,例如图3所示的,"Date”,"日期”,"27Feb”,"易办事”,通过图3和图4中的文字框的得出定位框,可以按照顺序将二者一一对应。结合配置文件、文字坐标和框线坐标可得"Date”:"x1,y1,x2,y2”,"日期”:"x1,y1,x2,y2”,既包含文字坐标又包含语料内容的标记信息。
[0077] 在获得标注信息的同时,通过所述配置文件、所述文字坐标和所述框线坐标还能获得单元格位置信息。
[0078] 将获得的单元格位置信息和标注信息输入标注工具,得到的标注信息即是以JSON格式数据。本步骤中通过使用imgkit工具渲染表格,并直接精确提取出文字、单元格、线条等坐标位置,确保了获取标注信息的准确性和效率。
[0079] 步骤S5:
[0080] 批量生成所述表格样本图片和所述标注信息,形成数据集,利用所述数据集对神经网络模型进行训练和 / 或验证。
[0081] 具体的,批量生成JSON格式数据的标注信息,结合标注信息和表格样本图片形成数据集,数据集中包括有验证集,利用数据集对神经网络模型进行训练,利用验证集对神经网络模型进行验证。
[0082] 本实施例中的步骤S1~S5的前后顺序不作限制,工作人员可根据实际情况进行调整。
[0083] 实施例二
[0084] 对应上述实施例,本申请提供一种基于表格图像及对应标注信息的生成方法的装置,所述装置包括:
[0085] 语料配置模块,用于根据场景信息,收集所述场景信息对应的语料形成语料库。
[0086] 文件配置模块,用于对表格参数进行定义,并记录在配置文件中。
[0087] 渲染模块,用于集合所述配置文件和所述预料库,渲染生成表格样本图片和标注图片。
[0088] 提取模块,提取所述标注图片,生成标注信息。
[0089] 还包括模型训练模块,批量生成所述表格样本图片和所述标注信息,形成数据集,利用所述数据集对神经网络模型进行训练和 / 或验证。
[0090] 具体的,在生成标注图片之前,还包括:在配置文件中的添加修改文字底色代码,以实现根据所述配置文件提取标注信息。
[0091] 通过银行、金融报表类等场景的数据特点,收集对应场景信息的语料形成语料库,同时根据需要收集数据的场景的表格数据特点,对表格特性进行定义并记录在配置文件中,结合配置文件和语料库内容,渲染生成表格样本图片和标注图片,其中标注图片和表格样本图片为同一布局图片,提取标注图片,得到标注信息,能够在指定语料和版式的情况下大批量随机生成表格图片,同时获取图片中所有标注信息,解决了特定场景下表格图片数据采集难度大、标注成本高的问题。
[0092] 实施例三
[0093] 提供了一种计算机设备,包括存储器、处理器及存储在存储器上并可在处理器上运行的计算机程序,处理器执行计算机程序时实现以下步骤:
[0094] 根据场景信息,收集所述场景信息对应的语料形成语料库;
[0095] 对表格参数进行定义,并记录在配置文件中;
[0096] 结合所述配置文件和所述语料库,渲染生成表格样本图片和标注图片;
[0097] 提取所述标注图片,生成标注信息。
[0098] 在一个实施例中,提供了一种计算机设备,该计算机设备可以是服务器,其内部结构图可以如图6所示。该计算机设备包括通过系统总线连接的处理器、存储器、网络接口和数据库。其中,该计算机设备的处理器用于提供计算和控制能力。该计算机设备的存储器包括非易失性存储介质、内存储器。该非易失性存储介质存储有操作系统、计算机程序和数据库。该内存储器为非易失性存储介质中的操作系统和计算机程序的运行提供环境。该计算机设备的数据库用于存储语料库数据。该计算机设备的网络接口用于与外部的终端通过网络连接通信。该计算机程序被处理器执行时以实现一种表格图像及对应标注信息的生成方法。
[0099] 本领域技术人员可以理解,图6中示出的结构,仅仅是与本申请方案相关的部分结构的框图,并不构成对本申请方案所应用于其上的计算机设备的限定,具体的计算机设备可以包括比图中所示更多或更少的部件,或者组合某些部件,或者具有不同的部件布置。
[0100] 实施例四
[0101] 在一个本实施例中,提供了一种种计算机可读存储介质,其上存储有计算机程序,计算机程序被处理器执行时实现以下步骤:
[0102] 根据场景信息,收集所述场景信息对应的语料形成语料库;
[0103] 对表格参数进行定义,并记录在配置文件中;
[0104] 结合所述配置文件和所述语料库,渲染生成表格样本图片和标注图片;
[0105] 提取所述标注图片,生成标注信息。
[0106] 本领域普通技术人员可以理解实现上述实施例方法中的全部或部分流程,是可以通过计算机程序来指令相关的硬件来完成,所述的计算机程序可存储于一非易失性计算机可读取存储介质中,该计算机程序在执行时,可包括如上述各方法的实施例的流程。其中,本申请所提供的各实施例中所使用的对存储器、存储、数据库或其它介质的任何引用,均可包括非易失性和 / 或易失性存储器。非易失性存储器可包括只读存储器(ROM)、可编程ROM(PROM)、电可编程ROM(EPROM)、电可擦除可编程ROM(EEPROM)或闪存。易失性存储器可包括随机存取存储器(RAM)或者外部高速缓冲存储器。作为说明而非局限,RAM以多种形式可得,诸如静态RAM(SRAM)、动态RAM(DRAM)、同步DRAM(SDRAM)、双数据率SDRAM(DDRSDRAM)、增强型SDRAM(ESDRAM)、同步链路(Synchlink)DRAM(SLDRAM)、存储器总线(Rambus)直接RAM(RDRAM)、直接存储器总线动态RAM(DRDRAM)、以及存储器总线动态RAM(RDRAM)等。
[0107] 尽管已描述了本发明实施例中的优选实施例,但本领域内的技术人员一旦得知了基本创造性概念,则可对这些实施例作出另外的变更和修改。所以,所附权利要求意欲解释为包括优选实施例以及落入本发明实施例中范围的所有变更和修改。
[0108] 显然,本领域的技术人员可以对本发明进行各种改动和变型而不脱离本发明的精神和范围。这样,倘若本发明的这些修改和变型属于本发明权利要求及其等同技术的范围之内,则本发明也意图包含这些改动和变型在内。< / style>
Claims
1. A method for generating a table image and corresponding annotation information, characterized in that, The method includes: Collecting the corpus corresponding to the scenario information according to the scenario information to form a corpus; Defining table parameters and recording them in a configuration file; Combining the configuration file and the corpus to render and generate table sample pictures and annotation pictures, specifically including adding and modifying text background color codes in the configuration file, and rendering and generating annotation pictures by combining the configuration file with added and modified text background color codes and the corpus; Extracting the annotation pictures to generate annotation information.
2. The method according to claim 1, characterized in that, It also includes: Batch generating the table sample pictures and the annotation information to form a data set, and using the data set to train and / or validate a neural network model.
3. The method according to claim 1, characterized in that Defining the table parameters specifically includes: Defining the table structure and / or table content and / or table style according to the scenario information.
4. The method according to claim 3, characterized in that, Extracting the annotation pictures to generate annotation information specifically includes: Extracting the annotation pictures to obtain text coordinates and frame line coordinates; Combining the configuration file, the text coordinates and the frame line coordinates to obtain the annotation information.
5. The method according to claim 4, wherein Separating the annotation pictures to obtain a text annotation picture and a table annotation picture; Extracting contour coordinates based on the text annotation picture to generate text coordinates; Performing contour detection on the table annotation picture to obtain frame line coordinates.
6. The method according to claim 5, wherein Binarizing the table annotation picture to obtain a table binarized picture; Eroding and dilating the table binarized picture to separate the horizontal lines in the table binarized picture to obtain a horizontal line picture; Eroding and dilating the table binarized picture to separate the vertical lines in the table binarized picture to obtain a vertical line picture; Extracting the frame line coordinates according to the horizontal line picture and the vertical line picture.
7. An apparatus for a generation method based on a table image and corresponding annotation information, characterized in that, The device includes: A corpus configuration module for collecting the corpus corresponding to the scenario information according to the scenario information to form a corpus; A file configuration module for defining table parameters and recording them in a configuration file; A rendering module for combining the configuration file and the corpus to render and generate table sample pictures and annotation pictures; The rendering module is also used for adding and modifying text background color codes in the configuration file, and rendering and generating annotation pictures by combining the configuration file with added and modified text background color codes and the corpus; An extraction module for extracting the annotation pictures to generate annotation information.
8. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Systems, methods, and computer readable media for security in profile utilizing systems
CN102971738A
Electric power field project feature identification method based on deep learning
CN113869054A