Method and system for automatically generating table image and annotation data thereof based on semantics
By constructing single-row headers, multi-row headers, and hierarchical table images, and combining them with few-shot learning of a large language model, we automatically generate table images and their annotation data. This solves the problem of insufficient coverage of table datasets in existing technologies, achieves efficient annotation of complex table scenarios, and enhances multimodal training data, improving the model's adaptability and annotation accuracy.
Patent Information
- Application Number
- CN202510711984.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-05
AI Technical Summary
The scale and quality of existing table datasets are insufficient to meet the needs of complex table parsing tasks. There is a lack of extensive coverage of various table styles in real application scenarios. The diversity of the datasets and the fine-grainedness of the annotations are insufficient, which leads to limited performance of the models in practical applications. Existing methods also find it difficult to generate multimodal training data enhancement and unified annotation standards.
By acquiring single-header and multi-header semantic datasets, constructing single-row headers, multi-row headers, and hierarchical table images, and combining preset basic templates with few-shot learning of large language models, it automatically generates annotation variants, realizes automated processing of table images and their annotation data, supports merging cells and cross-row and cross-column annotation, and generates multi-language, multi-format annotation versions.
It improves the adaptability to various table types, improves the accuracy and richness of annotations, reduces the dependence on large-scale annotated data, and enhances the model's ability to quickly adapt to new scenarios.
Smart Images

Figure CN120599084A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of table data technology, and mainly to a method and system for automatically generating semantic-based table images and their annotated data. Background Art
[0002] Table parsing remains a core challenge in complex text layout recognition, text extraction, and data generation. Both traditional optical recognition technology and the multimodal large-scale models that have emerged in recent years are dedicated to improving accuracy and generalization in this area.
[0003] However, the scale and quality of existing datasets are currently unable to meet the needs of complex table parsing tasks; the number and diversity of currently publicly available table datasets are insufficient to support the generalization capabilities of deep learning models for diverse tables; many datasets focus on specific fields or formats and lack extensive coverage of various table styles in real application scenarios, resulting in limited performance of the model in actual applications; in addition, the diversity of the datasets and the insufficient granularity of the annotations limit the adaptability of the model to real application scenarios; at the same time, most existing datasets only provide basic cell segmentation information, lack deeper text semantic annotations, and have no unified table annotation standards, making it difficult for different datasets to be compatible, further limiting the generalization capabilities of the model.
[0004] For example, the Chinese invention patent with publication number "CN114581923A" discloses a "method, device and storage medium for generating table images and corresponding annotation information", which specifically discloses "a method, device and storage medium for generating table images and corresponding annotation information, relating to the field of table data technology. The method includes: according to scene information, collecting corpus corresponding to the scene information to form a corpus; defining table parameters and recording them in a configuration file; combining the configuration file and the corpus to render and generate table sample images and annotation images; extracting the annotation images to generate annotation information", but this method requires manual definition of table parameters and writing them into the configuration file. When the table structure changes, the configuration needs to be manually modified, which may easily lead to the generated table image and annotation information not matching; in addition, this method only generates annotation information by extracting annotation images, lacks the ability to semantically process the annotation content, and is difficult to meet the annotation requirements of complex tables; at the same time, this method cannot generate multiple annotation variants of the same table, and is difficult to meet the needs of multimodal training data enhancement. Summary of the Invention
[0005] In order to solve the above-mentioned problems existing in the prior art, the present application provides a method and system for automatically generating semantic-based table images and their annotated data.
[0006] The technical solution of this application is as follows:
[0007] In one aspect, the present invention provides a method for automatically generating a semantically-based table image and its annotated data, the method comprising:
[0008] Obtain a single-header semantic data set and a multi-header semantic data set, wherein each piece of semantic data in the single-header semantic data set includes a single-row header and multiple rows of specific values, and each piece of semantic data in the multi-header semantic data set includes multiple rows of headers and multiple rows of specific values;
[0009] Construct table images based on single-header semantic datasets and multi-header semantic datasets, including single-row header images, multi-row header images, and hierarchical table images;
[0010] Based on a preset basic template, basic annotation data is generated; the basic annotation data is input into a preset large language model, and annotation variants are generated using a few-shot learning method; different annotation data are generated by combining the table image, basic template and annotation variants, and all annotation data and table images with different annotation data are obtained.
[0011] Preferably, the method further comprises performing data cleaning on the semantic dataset, wherein the data includes processing missing values, outliers and standardizing data formats.
[0012] Preferably, the single-line header image is specifically:
[0013] Randomly selecting first data from the semantic data set, the selected first data comprising a single-line header and multiple lines of specific values; determining a header type in the selected first data, wherein the header types include SHNM, SHRM, and SHCM;
[0014] If the current header type is a merge task, initialize the merge area and generate the coordinates of the cells in the merge area; determine the merge type of the current merge task. If it is a row merge, randomly merge a preset number of rows in the same column and the text contents of the corresponding cells; otherwise, randomly merge a preset number of columns in the same row and the text contents of the corresponding cells, where the text contents of the corresponding cells are connected with "," or " / ";
[0015] Fill the merged text content into the upper left corner cell of the merged area, and leave all cells except the upper left corner cell of the merged area blank; fill the values of cells in the non-merged area with the specific values of the selected first data;
[0016] Furthermore, if the merged area overlaps with the already merged area, the merged area is overwritten or skipped;
[0017] Iterate the merging process until the preset merging target number or the preset maximum merging number is reached, then stop merging and obtain the merged area;
[0018] If the current table header type is not a merge task or the merge has been completed, perform table dimension calculation on the selected first data to obtain the first table dimension;
[0019] Determine whether the area occupied by the text content in the current cell is larger than the area of the current cell; if so, adjust the text content according to preset constraint rules; the constraint rules include line break rules and font scaling, wherein the line break rule specifically calculates the maximum number of characters that can be accommodated in a single line under the current font size, and automatically wraps the text if the single-line text content reaches the maximum number of characters that can be accommodated; the font scaling specifically reduces the font size if the pixel height occupied by the total number of lines after the text content is wrapped exceeds the preset height of the current cell;
[0020] Based on the first table dimension, a blank table image is created using an image rendering tool to generate a single-row header initial image; and image enhancement is performed on the single-row header initial image to generate a single-row header image.
[0021] Preferably, the multi-line header image is specifically:
[0022] Randomly selecting second data from the semantic data set, the selected second data comprising column headers and multiple rows of specific values;
[0023] Calculate the dimension of the selected second data to obtain the second table dimension, and perform header row distribution, which can be expressed as the formula:
[0024] len(avail_headers)=col*row, row∈[2,0.5*len(headers)],
[0025] In the formula, avail_headers represents the valid column headers; len() represents the length calculation function; row represents the row height of the column header cell; col represents the column height of the column header cell; headers represents the column headers;
[0026] Determining a header type in the selected second data, where the header types include MHNM and MHHM;
[0027] If the current header type is a merge task, randomly generate a target number of merged columns that does not exceed the total number of columns; merge the header cells according to the preset vertical merge range, and use "," to connect the text content of the cells before the merge;
[0028] If the current header type is not a merge task or the merge has been completed, the constraint rules are used to adjust the text content in the merged header cell; based on the second table dimension, a blank table image is created using an image rendering tool to generate a multi-line header initial image; the multi-line header initial image is enhanced to generate a multi-line header image.
[0029] Preferably, the hierarchical table image is specifically:
[0030] Randomly selecting third data from the semantic data set, where the selected third data includes table headers of multiple tables and multiple rows of specific values;
[0031] Calculating the dimension of the selected third data to obtain a third table dimension, merging the first-level headers of the multiple table headers, and filling the merged first-level header cells with the merged text content;
[0032] The constraint rules are used to adjust the text content of the merged first-level header cells; based on the third table dimension, an image rendering tool is used to create a blank table image to generate an initial image of the hierarchical table; and image enhancement is performed on the initial image of the hierarchical table to generate a hierarchical table image.
[0033] Preferably, the basic annotation data includes the number of rows, columns, merged areas, text content in cells, semantic relationships of hierarchical headers, and question-answer triples of the table image, wherein the question-answer triple annotations are specifically the correspondence between questions, answers, and source cells.
[0034] Preferably, the table image, the basic template and the annotation variant are combined to generate different annotation data, and all the annotation data and the table image with different annotation data are obtained, specifically:
[0035] Convert the format of the basic annotation data into the annotation format of the preset large language model, extract the question-answer triples annotations of different annotation tasks, and generate a few-shot learning set;
[0036] Set the parameter group {T, max_tokens, top_p} of the preset large language model, where T represents the diversity parameter, max_tokens represents the text length parameter, and top_p represents the balance parameter;
[0037] The preset large language model identifies the type of the current input annotation task, retrieves the question-answer triplet annotations that match the current input annotation task type in the few-shot learning set, and obtains the input and output formats of the current input annotation task type annotations;
[0038] Generate annotation variants with the same annotation format as the currently input annotation task type based on the preset variant generation rules;
[0039] The table image, basic template and annotation variant are randomly combined according to a combination method, wherein the combination method includes {table image, basic template}, that is, the same table image is associated with multiple types of basic templates, {basic template, annotation variant}, that is, each basic template contains multiple annotation variants, generating different annotation data, and obtaining all annotation data and table images with different annotation data.
[0040] On the other hand, the present invention also proposes a semantic-based automatic generation system for table images and their annotated data, the system comprising a data acquisition module, a table image generation module, an annotation module, and a result output module, wherein:
[0041] The data acquisition module is used to acquire a single-header semantic data set and a multi-header semantic data set, wherein each piece of semantic data in the single-header semantic data set includes a single-row header and multiple rows of specific values, and each piece of semantic data in the multi-header semantic data set includes multiple rows of headers and multiple rows of specific values; and transmit the single-header semantic data set and the multi-header semantic data set to the table image generation module;
[0042] The table image generation module is used to construct a table image based on a single-header semantic dataset and a multi-header semantic dataset, including a single-row header image, a multi-row header image, and a hierarchical table image;
[0043] The annotation module is used to generate basic annotation data based on a preset basic template; input the basic annotation data into a preset large language model and generate annotation variants using a few-shot learning method; combine the table image, the basic template and the annotation variants to generate different annotation data, and obtain all the annotation data and the table image with different annotation data;
[0044] The result output module is used to display all the labeled data and table images with different labeled data.
[0045] On the other hand, the present invention also proposes an electronic device, comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein when the processor executes the program, the method for automatically generating semantic-based table images and their annotation data as described in any embodiment of the present invention is implemented.
[0046] On the other hand, the present invention further proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for automatically generating a semantic-based table image and its annotation data as described in any embodiment of the present invention.
[0047] Compared with the prior art, the present invention has the following beneficial effects:
[0048] 1) The present invention provides a method and system for automatically generating semantically based table images and their annotated data. This method supports the construction of single-row headers, multi-row headers, and hierarchical table images, can cover complex table scenarios, and improves adaptability to diverse table types.
[0049] 2) The present invention provides a method and system for automatically generating semantically-based table images and their annotation data, which can automatically handle complex logic such as merging cells and annotating across rows and columns, avoiding manual configuration errors and improving annotation accuracy;
[0050] 3) The present invention provides a method and system for automatically generating semantically based table images and their annotation data. By combining basic templates with few-shot learning of a large language model to generate annotation variants, the system can automatically generate multi-language, multi-format annotated versions of the same table, thereby enhancing the richness and scalability of annotations, reducing dependence on large-scale annotated data, and improving the model's ability to quickly adapt to new scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 is a flow chart of a method according to an embodiment of the present invention;
[0052] Figure 2 This is a single-row header image according to an embodiment of the present invention;
[0053] Figure 3 is a multi-row header image according to an embodiment of the present invention;
[0054] Figure 4 It is a hierarchical table image according to an embodiment of the present invention. DETAILED DESCRIPTION
[0055] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.
[0056] The present invention provides the following technical solution: a method and system for automatically generating semantic-based table images and their annotated data.
[0057] Example 1
[0058] See Figure 1 This embodiment provides a method for automatically generating a table image and its annotated data based on semantics, and the specific steps include:
[0059] S1. Obtain a single-header semantic dataset, Semantics-TableQA, and a multi-header semantic dataset, Semantics-DuSQL. Each piece of semantic data in the single-header semantic dataset includes a single-row header and multiple rows of specific values. Each piece of semantic data in the multi-header semantic dataset includes multiple rows of headers and multiple rows of specific values.
[0060] The method further comprises performing data cleaning on the semantic dataset, including processing missing values, outliers and standardizing the data format;
[0061] S2. constructing table images based on the single-header semantic dataset and the multi-header semantic dataset, including single-row header images, multi-row header images, and hierarchical table images;
[0062] S21, see Figure 2 The single-row header image is specifically randomly selected from the semantic data set, where the selected first data includes a single-row header and multiple rows of specific values; the header type in the selected first data is judged, where the header types include SHNM, SHRM, and SHCM; the SHNM is specifically a non-merged cell; the SHRM is specifically a row-merged cell; and the SHCM is specifically a column-merged cell;
[0063] If the current header type is a merge task, initialize the merge area and generate the coordinates of the cells in the merge area; the coordinates of the upper left corner cell of the merge area are (R1, C1), and the coordinates of the lower right corner cell are (R2, C2), R1≤R2, C1≤C2; determine the merge type of the current merge task. If it is a row merge, randomly merge a preset number of rows in the same column and the text contents of the corresponding cells; otherwise, randomly merge a preset number of columns in the same row and the text contents of the corresponding cells, where the text contents of the corresponding cells are connected with "," or " / ";
[0064] Fill the merged text content into the upper left corner cell of the merged area, and leave all cells except the upper left corner cell of the merged area blank; fill the values of cells in the non-merged area with the specific values of the selected first data;
[0065] Furthermore, if the merged area overlaps with the already merged area, the merged area is overwritten or skipped;
[0066] Iterate the merging process until the preset merging target number or the preset maximum merging number is reached, then stop merging and obtain the merged area;
[0067] If the current table header type is not a merge task or the merge has been completed, perform table dimension calculation on the selected first data to obtain the first table dimension;
[0068] Determine whether the area occupied by the text content in the current cell is larger than the area of the current cell; if so, adjust the text content according to preset constraint rules; the constraint rules include line break rules and font scaling, wherein the line break rule specifically calculates the maximum number of characters that can be accommodated in a single line under the current font size, and automatically wraps the text if the single-line text content reaches the maximum number of characters that can be accommodated; the font scaling specifically reduces the font size if the pixel height occupied by the total number of lines after the text content is wrapped exceeds the preset height of the current cell;
[0069] Based on the first table dimension, using an image rendering tool to create a blank table image to generate a single-row table header initial image; performing image enhancement on the single-row table header initial image to generate a single-row table header image;
[0070] S22, see Figure 3 , the multi-row header image is specifically a second data randomly selected from the semantic data set, and the selected second data includes a column title and a plurality of rows of specific values;
[0071] Calculate the dimension of the selected second data to obtain the second table dimension, and perform header row distribution, which can be expressed as the formula:
[0072] len(avail_headers)=col*row, row∈[2,0.5*len(headers)],
[0073] In the formula, avail_headers represents the valid column headers; len() represents the length calculation function; row represents the row height of the column header cell; col represents the column height of the column header cell; headers represents the column headers;
[0074] Determine the header type in the selected second data, where the header types include MHNM and MHHM; the MHNM is specifically a header-free merge; the MHHM is specifically a header-merged merge;
[0075] If the current header type is a merge task, randomly generate a target number of merged columns that does not exceed the total number of columns; merge the header cells according to the preset vertical merge range, and use "," to connect the text content of the cells before the merge;
[0076] If the current header type is not a merge task or the merge has been completed, the text content in the merged header cell is adjusted using the constraint rule; based on the second table dimension, a blank table image is created using an image rendering tool to generate a multi-line header initial image; the multi-line header initial image is enhanced to generate a multi-line header image;
[0077] S23, see Figure 4 , the hierarchical table image is specifically randomly selecting third data from the semantic data set, and the selected third data includes table headers of multiple tables and multiple rows of specific values;
[0078] Calculating the dimension of the selected third data to obtain a third table dimension, merging the first-level headers of the multiple table headers, and filling the merged first-level header cells with the merged text content;
[0079] The text content of the merged first-level table header cells is adjusted using the constraint rules; based on the third table dimension, a blank table image is created using an image rendering tool to generate an initial image of the hierarchical table; and image enhancement is performed on the initial image of the hierarchical table to generate a hierarchical table image;
[0080] S3. Generate basic annotation data based on a preset basic template; the basic annotation data includes the number of rows and columns of the table image, the number of merged areas, the text content in the cells, the semantic relationship of the hierarchical header, and the question-answer triples, wherein the question-answer triples are specifically annotated with the corresponding relationship between the question, the answer, and the source cell;
[0081] In this embodiment, the basic template includes a non-positioning template and a positioning template. The non-positioning template includes TSD, TR, and THD, and the positioning template includes TCE, TCL, MCD, and RCE.
[0082] TSD is specifically a table structure parser, which is used to obtain the table structure in the image and return the number of rows and columns in the table; the annotation format is {"rows": <number of rows>, "columns": <number of columns>}
[0083] TR stands for table to HTML, which is used to convert the table content in the image into HTML format and present it in JSON format;
[0084] THD is specifically a header hierarchical parsing, which is used to parse the table structure in the image, determine the hierarchical relationship of the header, and return the header field list according to the rules; the annotation format is a single-layer header: ["column name 1", "column name 2", "column name 3"], that is, [C1, C2, ..., C n ,...,C N ], where C n Indicates the nth column name, N indicates the number of column names, and n indicates the index value of the nth column name; multi-layer header: ["first-level column name-second-level column name", "first-level column name-second-level column name", ...], that is, [C1-C2, ..., C1-C n ,...,C1-C N ];
[0085] TCE stands for cell content extraction, which is used to parse a table and return the text content of cells in a specified row and column. The annotation format is {"cell_content":"<cell content>"}.
[0086] TCL is specifically for cell positioning and content query, which is used to analyze the table image and determine the cell content of each request and its position in the table; the annotation format is {'value':'<cell value> ', 'location': (row index, column index)}; where row index represents the row index, column index represents the column index, value represents the value, cell value represents the cell value, and location represents the cell location;
[0087] MCD stands for Merged Cell Detection, which is used to analyze tables in an image and find all merged cells. The annotation format is [{“top-left”: (R1, C1), “bottom-right”: (R2, C2)}], where (R1, C1) represents the coordinates of the top-left corner of the current region, (R2, C2) represents the coordinates of the top-right corner of the current region, top-left represents the top-left corner, and bottom-right represents the top-right corner.
[0088] RCE is specifically row and column cell retrieval, which is used to retrieve cell data in a specified row or column in a table; the annotation format is {“id”: <index>","cell_list":"<cell data of the specified row or column>"}, where index represents the cell index value of the specified row or column, and cell_list represents the cell data of the specified row or column;
[0089] S4. Input the basic annotation data into a preset large language model and generate annotation variants using a few-shot learning method;
[0090] Convert the format of the basic annotation data into the annotation format of the preset large language model, extract the question-answer triples annotations of different annotation tasks, and generate a few-shot learning set;
[0091] Set the parameter group {T, max_tokens, top_p} of the preset large language model, where T represents the diversity parameter, max_tokens represents the text length parameter, and top_p represents the balance parameter;
[0092] The preset large language model identifies the type of the current input annotation task, retrieves the question-answer triplet annotations that match the current input annotation task type in the few-shot learning set, and obtains the input and output formats of the current input annotation task type annotations;
[0093] Generate annotation variants with the same annotation format as the currently input annotation task type based on the preset variant generation rules;
[0094] S5. Combining the table image, the basic template, and the annotation variants, randomly combining the table image, the basic template, and the annotation variants according to a combination method, wherein the combination method includes {table image, basic template}, i.e., associating multiple basic templates with the same table image, and {basic template, annotation variant}, i.e., each basic template contains multiple annotation variants, generating different annotation data, and obtaining all the annotation data and table images with different annotation data;
[0095] In this embodiment, {tabular image, basic template} specifically generates 10 annotation variants in the non-positioning template, that is, one image can construct 70 different annotation data; {basic template, annotation variant} specifically generates 10 annotation variants in the positioning task, and each annotation variant selects 5 basic templates, which can construct 50 different annotation data.
[0096] Example 2
[0097] This embodiment provides a semantic-based automatic generation system for table images and their annotated data. The system includes a data acquisition module, a table image generation module, an annotation module, and a result output module, wherein:
[0098] The data acquisition module is used to acquire a single-header semantic data set and a multi-header semantic data set, wherein each piece of semantic data in the single-header semantic data set includes a single-row header and multiple rows of specific values, and each piece of semantic data in the multi-header semantic data set includes multiple rows of headers and multiple rows of specific values; and transmit the single-header semantic data set and the multi-header semantic data set to the table image generation module;
[0099] The table image generation module is used to construct a table image based on a single-header semantic dataset and a multi-header semantic dataset, including a single-row header image, a multi-row header image, and a hierarchical table image;
[0100] The annotation module is used to generate basic annotation data based on a preset basic template; input the basic annotation data into a preset large language model and generate annotation variants using a few-shot learning method; combine the table image, the basic template and the annotation variants to generate different annotation data, and obtain all the annotation data and the table image with different annotation data;
[0101] The result output module is used to display all the labeled data and table images with different labeled data.
[0102] Example 3
[0103] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method for automatically generating a semantic-based table image and its annotation data as described in any embodiment of the present invention is implemented.
[0104] Example 4
[0105] This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the method for automatically generating a semantic-based table image and its annotation data as described in any embodiment of the present invention is implemented.
[0106] It is worth noting that the system, electronic device and computer-readable storage medium described in the present invention are all based on the same principles as the method described in Example 1, and will not be repeated here.
[0107] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention's description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.< / index>
Claims
1. A method for automatically generating a semantically-based table image and its annotated data, characterized in that: The method comprises: Acquire a single-header semantic data set and a multi-header semantic data set, wherein each piece of semantic data in the single-header semantic data set includes a single-row header and multiple rows of specific values, and each piece of semantic data in the multi-header semantic data set includes multiple rows of headers and multiple rows of specific values; Construct table images based on single-header semantic datasets and multi-header semantic datasets, including single-row header images, multi-row header images, and hierarchical table images; Based on a preset basic template, basic annotation data is generated; the basic annotation data is input into a preset large language model, and annotation variants are generated using a few-shot learning method; different annotation data are generated by combining the table image, basic template and annotation variants, and all annotation data and table images with different annotation data are obtained.
2. The method for automatically generating a semantically-based table image and its annotated data according to claim 1, characterized in that: The method further includes performing data cleaning on the semantic dataset, including processing missing values, outliers and standardizing data formats.
3. The method for automatically generating a semantically-based table image and its annotated data according to claim 1, characterized in that: The single-row header image is specifically: Randomly selecting first data from the semantic data set, the selected first data comprising a single-line header and multiple lines of specific values; determining a header type in the selected first data, wherein the header types include SHNM, SHRM, and SHCM; If the current header type is a merge task, initialize the merge area and generate the coordinates of the cells in the merge area; determine the merge type of the current merge task. If it is a row merge, randomly merge a preset number of rows in the same column and the text content of the corresponding cells; Otherwise, randomly merge a preset number of columns and text contents of corresponding cells in the same row, where the text contents of corresponding cells are connected with "," or " / "; Fill the merged text content into the upper left cell of the merged area, and leave all cells except the upper left cell of the merged area blank; The values of the cells in the non-merged area are filled with the specific values of the selected first data; Furthermore, if the merged area overlaps with the already merged area, the merged area is overwritten or skipped; Iterate the merging process until the preset merging target number or the preset maximum merging number is reached, then stop merging and obtain the merged area; If the current table header type is not a merge task or the merge has been completed, perform table dimension calculation on the selected first data to obtain the first table dimension; Determine whether the area occupied by the text content in the current cell is larger than the area of the current cell; if so, adjust the text content according to preset constraint rules; the constraint rules include line break rules and font scaling, wherein the line break rule specifically calculates the maximum number of characters that can be accommodated in a single line under the current font size, and automatically wraps the text if the single-line text content reaches the maximum number of characters that can be accommodated; the font scaling specifically reduces the font size if the pixel height occupied by the total number of lines after the text content is wrapped exceeds the preset height of the current cell; Based on the first table dimension, use the image rendering tool to create a blank table image and generate a single-row table header initial image; Image enhancement is performed on the single-row header initial image to generate a single-row header image.
4. The method for automatically generating a semantically-based table image and its annotated data according to claim 1, characterized in that: The multi-line header image is specifically: Randomly selecting second data from the semantic data set, the selected second data comprising column headers and multiple rows of specific values; Calculate the dimension of the selected second data to obtain the second table dimension, and perform header row distribution, which can be expressed as the formula: len(avail_headers)=col*row,row∈[2,0.5*len(headers)], In the formula, avail_headers represents the valid column headers; len() represents the length calculation function; row represents the row height of the column header cell; col represents the column height of the column header cell; headers represents the column header; Determining a header type in the selected second data, where the header types include MHNM and MHHM; If the current header type is a merge task, randomly generate a target number of merged columns that does not exceed the total number of columns; merge the header cells according to the preset vertical merge range, and use "," to connect the text content of the cells before the merge; If the current header type is not a merge task or the merge has been completed, the text content in the merged header cell is adjusted using the constraint rules; Based on the second table dimension, use the image rendering tool to create a blank table image and generate the initial image of the multi-row header; Image enhancement is performed on the multi-line header initial image to generate a multi-line header image.
5. The method for automatically generating a semantically-based table image and its annotated data according to claim 1, characterized in that: The hierarchical table image is specifically: Randomly selecting third data from the semantic data set, where the selected third data includes table headers of multiple tables and multiple rows of specific values; Calculating the dimension of the selected third data to obtain a third table dimension, merging the first-level headers of the multiple table headers, and filling the merged first-level header cells with the merged text content; Using the constraint rules, adjust the text content of the merged first-level header cells; Based on the third table dimension, a blank table image is created using an image rendering tool to generate an initial image of the hierarchical table; Image enhancement is performed on the initial image of the hierarchical table to generate a hierarchical table image.
6. The method for automatically generating a semantically-based table image and its annotated data according to claim 1, characterized in that: The basic annotation data includes the number of rows, columns, merged areas, text content in cells, semantic relationships of hierarchical headers, and question-answer triples of the table image, where the question-answer triple annotations are specifically the correspondence between questions, answers, and source cells.
7. The method for automatically generating a semantically based table image and its annotated data according to claim 1, characterized in that: Combining the table image, the basic template and the annotation variant, different annotation data are generated to obtain all the annotation data and the table image with different annotation data, specifically: Convert the format of the basic annotation data into the annotation format of the preset large language model, extract the question-answer triples annotations of different annotation tasks, and generate a few-shot learning set; Set the parameter group {T, max_tokens, top_p} of the preset large language model, where T represents the diversity parameter, max_tokens represents the text length parameter, and top_p represents the balance parameter; The preset large language model identifies the type of the current input annotation task, retrieves the question-answer triplet annotations that match the current input annotation task type in the few-shot learning set, and obtains the input and output formats of the current input annotation task type annotations; Generate annotation variants with the same annotation format as the currently input annotation task type based on the preset variant generation rules; The table image, basic template and annotation variant are randomly combined according to a combination method, wherein the combination method includes {table image, basic template}, that is, the same table image is associated with multiple types of basic templates, {basic template, annotation variant}, that is, each basic template contains multiple annotation variants, generating different annotation data, and obtaining all annotation data and table images with different annotation data.
8. A semantic-based table image and its annotation data automatic generation system, characterized by: The system includes a data acquisition module, a table image generation module, a marking module and a result output module, wherein: The data acquisition module is used to acquire a single-header semantic data set and a multi-header semantic data set, wherein each piece of semantic data in the single-header semantic data set includes a single-row header and multiple rows of specific values, and each piece of semantic data in the multi-header semantic data set includes multiple rows of headers and multiple rows of specific values; and transmit the single-header semantic data set and the multi-header semantic data set to the table image generation module; The table image generation module is used to construct a table image based on a single-header semantic dataset and a multi-header semantic dataset, including a single-row header image, a multi-row header image, and a hierarchical table image; The annotation module is used to generate basic annotation data based on a preset basic template; input the basic annotation data into a preset large language model and generate annotation variants using a few-shot learning method; combine the table image, the basic template and the annotation variants to generate different annotation data, and obtain all the annotation data and the table image with different annotation data; The result output module is used to display all the labeled data and table images with different labeled data.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method for automatically generating a semantic-based table image and its annotation data as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for automatically generating a semantic-based table image and its annotation data as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Form image and corresponding annotation information generation method and device and storage medium
CN114581923A
Cited By
Complex merging cell table image construction method, system, device and medium
CN122134846A
Method, system and device for constructing complex merged cell table image and medium
CN122134846B