Complex table header one-dimensional automatic processing method and system for multi-dimensional table
By constructing a header tree structure and using a pre-trained model to process complex headers of multi-dimensional tables, the problems of low efficiency and error-proneness in existing technologies are solved, efficient and accurate one-dimensional processing is achieved, and the level of automation of data analysis is improved.
Patent Information
- Application Number
- CN202510727948.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-09-16
AI Technical Summary
Existing technologies are inefficient, error-prone, and lack repeatability when processing complex headers in multidimensional tables. In particular, manual processing methods have high error rates and the probability of repeated work when faced with large amounts of data.
An automated method is used to obtain the basic information of the header cells, build a header tree structure, merge text content layer by layer, and use a pre-trained model to determine semantic similarity and remove redundant information to generate a standardized one-dimensional header.
It realizes one-dimensional processing of complex headers of multi-dimensional tables, reduces the complexity of manual operations, reduces the risk of errors, and improves the automation level and consistency of data analysis.
Smart Images

Figure CN120654669A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to table data processing technology, and in particular to a method and system for automatically processing complex table headers into one dimension for multi-dimensional tables. Background Art
[0002] Scientific and technical literature contains a vast amount of tabular data. Extracting this tabular data and performing data cleaning and integration is a common paradigm for comprehensive research based on literature data integration in the big data era. However, during data integration, the complex headers of multidimensional tables are multidimensional; while the integrated data is stored in the database, the headers of the data entity tables must be one-dimensional. How to achieve one-dimensional processing of complex headers of multidimensional tables is the key issue addressed by this invention.
[0003] Although existing manual processing methods can solve some simple cleaning tasks, they often have significant shortcomings when faced with complex multi-level and multi-dimensional tabular data. First, manual processing is inefficient for complex data, especially when the amount of tabular data is large. Manual cleaning is not only time-consuming and labor-intensive, but also prone to errors. For example, when dealing with merged cells, missing values, and the relationship between different dimensions, manual intervention often makes it easy to miss certain key data or produce inconsistent formats. Second, manual methods lack repeatability and consistency. The results of each processing may vary due to differences in operators. Especially when faced with large amounts of data, the error rate and probability of duplication of work in manual processing methods will be greatly increased. Therefore, the one-dimensional automatic processing of complex headers is particularly important in this context. Summary of the Invention
[0004] In view of the problems existing in the prior art, the purpose of the present invention is to provide a method and system for automatically processing complex table headers into one dimension for multidimensional tables with high data processing efficiency and high accuracy.
[0005] In order to achieve the above-mentioned object of the invention, the present invention provides the following technical solutions:
[0006] A method for automatically converting complex headers of multidimensional tables into one dimension includes the following steps:
[0007] (1) Obtain the target document and extract the table data in the target document to form a table data set T;
[0008] (2) For any table t in the table data set T i , obtain the basic information of all header cells and store it in the header cell information set LC;
[0009] (3) Calculate the header level depth L i , if L i >=2, it means that the table ti If it is a complex header, execute step (4); otherwise, execute step (6);
[0010] (4) Take each header cell in the header cell set LC as a node, and i L i Starting from the header cell of the first layer, the table header tree structure is constructed by tracing back gradually upwards, and the text content of all nodes is merged layer by layer starting from the root node until the bottom node is reached. The merged text content obtained when the bottom node is reached is used as the one-dimensional header of the column where the bottom node is located;
[0011] (5) By judging the semantic similarity of the one-dimensional header, removing the redundant information in the one-dimensional header, applying the one-dimensional header to the table, and generating a normalized one-dimensional table;
[0012] (6) Structuring the table to generate a standardized one-dimensional table;
[0013] (7) Loop through steps (2) to (6) until all table data in the target document are processed.
[0014] Furthermore, the basic information of the header cell includes the text content of the header cell, the standard number of cells span merged by the header cell, the starting position col_start and ending position col_end of the row of the header cell in the table, and the next-level cell children of the current header cell, where the next-level cell children are cells in the same column as the current header cell but in the next row.
[0015] Furthermore, step (3) specifically includes:
[0016] (3-1) Traverse table t i For each row, calculate the ratio of non-numeric cells in each row, rate = number of non-numeric cells / total number of columns;
[0017] (3-2) When the rate is greater than the preset threshold, mark the corresponding line as the header row, and use the number of header rows as the header level depth L i ;
[0018] (3-3) If L i >=2, it means that the table t i If it is a complex header, execute step (4); otherwise, execute step (6).
[0019] Furthermore, step (4) specifically includes:
[0020] (4-1) According to the basic information of each header cell in the header cell set LC, obtain the next level cell children of each header cell;
[0021] (4-2) Take each header cell in the header cell set LC as a node, and i The header of the layer starts to trace back gradually, taking each header cell as the parent node of its next level cell children to build the header tree structure;
[0022] (4-3) Based on the basic information of each header cell, check whether the text content of each header cell is empty. If so, inherit the text content from its nearest upper ancestor node;
[0023] (4-4) Starting from the root node of the header tree structure, the header information in the header tree structure is merged layer by layer using a depth-first traversal algorithm until the bottom node is reached. Each time a node is traversed, the text content of the node is appended to the merged text content of all its ancestor nodes. A preset separator is set between the text contents of each node.
[0024] (4-5) The merged text content obtained when reaching the bottom node is used as the one-dimensional header of the column where the bottom node is located.
[0025] Furthermore, step (5) specifically includes:
[0026] (5-1) Use the pre-trained model to determine whether the texts in the one-dimensional header are synonymous. The specific judgment rules of the pre-trained model are as follows:
[0027] A. Chinese and English professional terms are considered synonymous if they are equivalent;
[0028] B. Standard abbreviations and full names are considered synonymous;
[0029] C. The method name and the measurement result are considered to be synonymous;
[0030] D. Two expressions with the same meaning are considered synonymous;
[0031] (5-2) When the pre-trained model determines that the two one-dimensional headers are synonymous, the two one-dimensional headers are considered redundant. Any one of the synonymous texts is removed to generate a normalized one-dimensional header.
[0032] A system for automatically converting complex headers into one dimension for multi-dimensional tables, including:
[0033] The table extraction module is used to obtain the target document and extract the table data in the target document to form a table data set T;
[0034] The table header cell extraction module is used for each table t in the table data set T i , obtain the basic information of all header cells and store it in the header cell information set LC;
[0035] The layer depth calculation module is used to calculate the layer depth L of the header based on LC i , if L i >=2, it means that the table t i If the header is complex, the one-dimensional header calculation module is executed; otherwise, the second normalization module is executed;
[0036] One-dimensional header calculation module, used to take each header cell in the header cell set LC as a node, from table t i L i Starting from the header cell of the first layer, the table header tree structure is constructed by tracing back gradually upwards, and the text content of all nodes is merged layer by layer starting from the root node until the bottom node is reached. The merged text content obtained when the bottom node is reached is used as the one-dimensional header of the column where the bottom node is located;
[0037] A first normalization module is used to remove redundant information in the one-dimensional header by judging the semantic similarity of the one-dimensional header, apply the one-dimensional header to the table, and generate a normalized one-dimensional table;
[0038] The second normalization module is used to structure the table and generate a normalized one-dimensional table.
[0039] Furthermore, the basic information of the header cell includes the text content of the header cell, the standard number of cells span merged by the header cell, the starting position col_start and ending position col_end of the row of the header cell in the table, and the next-level cell children of the current header cell, where the next-level cell children are cells in the same column as the current header cell but in the next row.
[0040] Furthermore, the hierarchical depth calculation module specifically includes:
[0041] Proportional calculation unit, used to traverse the table t i For each row, calculate the ratio of non-numeric cells in each row, rate = number of non-numeric cells / total number of columns;
[0042] The level depth calculation unit is used to mark the corresponding behavior as the header row when the rate is greater than the preset threshold, and use the number of header rows as the level depth L of the header i ;
[0043] Execution judgment unit, used if L i>=2, it means that the table t i If the header is complex, execute the one-dimensional header calculation module; otherwise, execute the second normalization module.
[0044] Furthermore, the one-dimensional header calculation module specifically includes:
[0045] The cell acquisition unit is used to obtain the next level cell children of each header cell according to the basic information of each header cell in the header cell set LC;
[0046] The header tree construction unit is used to take each header cell in the header cell set LC as a node, starting from the L i The header of the layer starts to trace back gradually, taking each header cell as the parent node of its next level cell children to build the header tree structure;
[0047] The assignment unit is used to retrieve the text content of each header cell based on the basic information of each header cell to see if it is empty. If so, the text content is inherited from its nearest upper ancestor node.
[0048] The text content merging unit is used to merge the header information in the header tree structure layer by layer using a depth-first traversal algorithm starting from the root node of the header tree structure until it reaches the bottom node. Each time a node is traversed, the text content of the node is added to the end of the merged text content of all its ancestor nodes. A preset separator is set between the text contents of each node.
[0049] The one-dimensional header generating unit is used to use the merged text content obtained when reaching the bottom-level node as the one-dimensional header of the column where the bottom-level node is located.
[0050] Furthermore, the first normalization module specifically includes:
[0051] The synonym judgment unit is used to use a pre-trained model to judge whether the texts in the one-dimensional header are synonymous. The specific judgment rules of the pre-trained model are:
[0052] A. Chinese and English professional terms are considered synonymous if they are equivalent;
[0053] B. Standard abbreviations and full names are considered synonymous;
[0054] C. The method name and the measurement result are considered to be synonymous;
[0055] D. Two expressions with the same meaning are considered synonymous;
[0056] The normalization unit is used to remove any one of the synonymous texts and generate a normalized one-dimensional header when the pre-trained model determines that the two one-dimensional headers are redundant.
[0057] Compared with existing technologies, the present invention has the following advantages: it realizes a method for automatically processing complex headers of multidimensional tables into one dimension. This method not only reduces the complexity of manual operations and the risk of errors, but also improves the automation level of data analysis, making the analysis process more efficient and consistent. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 This is a flow chart of a method for automatically converting complex headers into one dimension for a multi-dimensional table provided by an embodiment of the present invention;
[0059] Figure 2 is the original table data to be processed in this embodiment 1;
[0060] Figure 3 This is an example of the header cell information parsed in Example 1;
[0061] Figure 4 This is the header tree structure diagram constructed in this embodiment 1;
[0062] Figure 5 This is the result of converting the multidimensional header into one dimension in this embodiment;
[0063] Figure 6 is the original table data to be processed in this embodiment 2;
[0064] Figure 7 This is an example of the header cell information parsed in Example 2.
[0065] Figure 8 This is the header tree structure diagram constructed in this embodiment 2;
[0066] Figure 9 This is the result of converting the multidimensional header into one dimension in the second embodiment. DETAILED DESCRIPTION
[0067] The technical solutions in the embodiments of the present invention will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present invention.
[0068] Example 1
[0069] The embodiment of the present invention provides a method for automatically processing complex headers into one dimension for multi-dimensional tables, such as Figure 1 As shown, the following steps are included:
[0070] (1) Obtain the target document and extract the table data in the target document to form a table data set T{t i |1,2,…,n}.
[0071] The target document in this example is a geothermochronology paper titled "Fishion Trak Chronology of the Miocene Igneous Rocks of the Ki Peninsula" (Hideki Iwano et al., 2009). The purpose of this experiment is to convert the complex headers of a multidimensional table into a single dimension to facilitate importing the data into a database for unified management. The target document can be extracted using the third-party library Camelot. This document contains three tables, so n = 3.
[0072] (2) For any table t in the table data set T i , get the basic information of all header cells and store it in the header cell information set LC={c j |j=1,2,…,m}, where m represents the number of header cells.
[0073] Among them, the basic information of the header cell includes the text content of the header cell, the standard number of cells span merged by the header cell, the starting position col_start and ending position col_end of the row of the header cell in the table, and the next-level cell children of the current header cell. The next-level cell children are cells in the same column but in the next row as the current header cell.
[0074] Take Table 2 of the target document as an example. Figure 2 As shown, the results of the table after division according to the standard table are as follows Figure 3 As shown, for example, the text of the header cell "Before" is "Before", the number of merged cells span is 2, the starting position is the second row and third column, the ending position is the second row and fourth column, and the next level cell children is "U1(ppm)".
[0075] (3) Calculate the header level depth L i , if L i >=2, it means that the table t i If it is a complex header, execute step (4); otherwise, execute step (6).
[0076] The steps of love include:
[0077] (3-1) Traverse table t i For each row, calculate the ratio of non-numeric cells in each row, rate = number of non-numeric cells / total number of columns;
[0078] (3-2) When the rate is greater than the preset threshold (0.6), mark the corresponding line as the header row, and use the number of header rows as the level depth L of the header i ;
[0079] (3-3) If L i >=2, it means that the table t i If it is a complex header, execute step (4); otherwise, execute step (6).
[0080] In this embodiment, for table t2, the number of rows in the table is 8, and the non-digital ratio of each row is calculated. It can be found that the rate of the first three rows is greater than 0.6, marking the header level depth L i =3,L i >=2, it indicates that the table has a complex header.
[0081] (4) Take each header cell in the header cell set LC as a node, and i L i Starting from the header cell of the layer, the header tree structure is constructed step by step upward, and the text content of all nodes is merged layer by layer starting from the root node until the bottom node is reached. The merged text content obtained when the bottom node is reached is used as the one-dimensional header of the column where the bottom node is located.
[0082] This step specifically includes:
[0083] (4-1) According to the basic information of each header cell in the header cell set LC, obtain the next level cell children of each header cell;
[0084] (4-2) Take each header cell in the header cell set LC as a node, and i The header of the layer starts to trace back gradually, and each header cell is used as the parent node of its next level cell children to build a header tree structure; in this embodiment, the constructed tree structure is as follows Figure 4 As shown, "5%HCI solution (20 hours)" and "Etching (20 hours)" are two parent root nodes. Starting from the bottom table header, the tree structure is constructed by tracing back step by step. For example, the previous level of "U1 (ppm)" is "Before", and the previous level is "5%HCI solution (20 hours)". The previous level of the first "±1σ" is "Before", and the previous level is "5%HCI solution (20 hours)". The previous level of "U4 (ppm)" is "After", and the previous level is "Etching (20 hours)".
[0085] (4-3) Based on the basic information of each header cell, check whether the text content of each header cell is empty. If so, inherit the text content from its nearest upper ancestor node;
[0086] (4-4) Starting from the root node of the header tree structure, the depth-first traversal algorithm is used to merge the header information in the header tree structure layer by layer until the bottom node is reached. Each time a node is traversed, the text content of the node is added to the end of the merged text content of all its ancestor nodes. The text content of each node is separated by a preset separator "_". In this embodiment, the headers of a certain dimension are "5% HCI solution (20 hours) _ Before _ U1 (ppm)", "Etching (20 hours) _ Ratio _ ± 1σ", etc. The table after merging is as follows Figure 5 As shown;
[0087] (4-5) The merged text content obtained when reaching the bottom node is used as the one-dimensional header of the column where the bottom node is located.
[0088] (5) By judging the semantic similarity of the one-dimensional header, the redundant information in the one-dimensional header is removed, and the one-dimensional header is applied to the table to generate a normalized one-dimensional table.
[0089] This step specifically includes:
[0090] (5-1) Use a pre-trained model (e.g., a mature model such as DeepSeek) to determine whether the texts in the one-dimensional header are synonymous. The specific judgment rules of the pre-trained model are as follows:
[0091] A. Chinese and English professional terms are considered synonymous if they are equivalent;
[0092] B. Standard abbreviations and full names are considered synonymous;
[0093] C. The method name and the measurement result are considered to be synonymous;
[0094] D. Two expressions with the same meaning are considered synonymous;
[0095] (5-2) When the pre-trained model determines that the two one-dimensional headers are synonymous, it considers the two headers redundant and removes any of the synonymous texts to generate a normalized one-dimensional header. For example, if the input one-dimensional header is "In-situ dating", then DeepSeek will return true and only retain "In-situ dating".
[0096] (6) Structure the table to generate a standardized one-dimensional table and store it as a CSV file.
[0097] (7) Loop through steps (2) to (6) until all table data in the target document are processed.
[0098] Example 2
[0099] An embodiment of the present invention provides an automatic one-dimensional processing system for complex headers of multidimensional tables, including:
[0100] The table extraction module is used to obtain the target document and extract the table data in the target document to form a table data set T;
[0101] The table header cell extraction module is used for each table t in the table data set T i , obtain the basic information of all header cells and store it in the header cell information set LC;
[0102] The layer depth calculation module is used to calculate the layer depth L of the header based on LC i , if L i >=2, it means that the table t i If the header is complex, the one-dimensional header calculation module is executed; otherwise, the second normalization module is executed;
[0103] One-dimensional header calculation module, used to take each header cell in the header cell set LC as a node, from table t i L i Starting from the header cell of the first layer, the table header tree structure is constructed by tracing back gradually upwards, and the text content of all nodes is merged layer by layer starting from the root node until the bottom node is reached. The merged text content obtained when the bottom node is reached is used as the one-dimensional header of the column where the bottom node is located;
[0104] A first normalization module is used to remove redundant information in the one-dimensional header by judging the semantic similarity of the one-dimensional header, apply the one-dimensional header to the table, and generate a normalized one-dimensional table;
[0105] The second normalization module is used to structure the table and generate a normalized one-dimensional table.
[0106] The experiment in this embodiment is derived from a table in a geological thermochronology related document ( Figure 6 shown), the document name is "Awaji Island, Hyogo Prefecture, and the U-Pb and FT chronology of the Sennan rhyolite type during the Baian Period" (Oogakusato, 2020).
[0107] The basic information of the header cell includes the text content of the header cell, the number of standard cells span that the header cell merges, the starting position col_start and ending position col_end of the row of the header cell in the table, and the children of the next level of cells of the current header cell. The children of the next level of cells are cells in the same column but in the next row as the current header cell. Figure 7 ,by" 207 Pb / 206Pb" as an example, the cell content of c3 is " 207 Pb / 206 Pb", the number of merged cells span is 2, the starting position of c3 is the second row and second column, the ending position is the second row and third column, and the next level cell is "ratio".
[0108] The layer depth calculation module specifically includes:
[0109] Proportional calculation unit, used to traverse the table t i For each row, calculate the ratio of non-numeric cells in each row, rate = number of non-numeric cells / total number of columns;
[0110] The level depth calculation unit is used to mark the corresponding behavior as the header row when the rate is greater than the preset threshold, and use the number of header rows as the level depth L of the header i ;
[0111] Execution judgment unit, used if L i >=2, it means that the table t i If the header is complex, execute the one-dimensional header calculation module; otherwise, execute the second normalization module.
[0112] In this embodiment, the number of rows in the table is 35, and the non-numeric ratio of each row is calculated. It can be found that the rate of the first three rows is greater than 0.6, marking the depth of the header level L i =3,L i >=2, it indicates that the table has a complex header and the one-dimensional header calculation module is executed.
[0113] The one-dimensional header calculation module specifically includes:
[0114] The cell acquisition unit is used to obtain the next level cell children of each header cell according to the basic information of each header cell in the header cell set LC;
[0115] The header tree construction unit is used to take each header cell in the header cell set LC as a node, starting from the L i The header of the layer starts to trace back gradually, and each header cell is used as the parent node of its next level cell children to build a header tree structure; in this embodiment, the constructed tree structure is as follows Figure 8 As shown, "Isotopic ratios" and "U-Pbages" are two parent nodes, and the tree structure is constructed by tracing back from the bottom table header step by step. For example, the previous level of "ratio" is " 207 Pb / 206Pb", the next level up is "Isotopic ratios", the next level up is " 206 Pb / 238 Pb", and the next level up is "U-Pb ages";
[0116] The assignment unit is used to retrieve the text content of each header cell based on the basic information of each header cell to see if it is empty. If so, the text content is inherited from its nearest upper ancestor node.
[0117] The text content merging unit is used to merge the header information in the header tree structure layer by layer using a depth-first traversal algorithm starting from the root node of the header tree structure until it reaches the bottom node. Each time a node is traversed, the text content of the node is added to the end of the merged text content of all its ancestor nodes. A preset separator is set between the text contents of each node.
[0118] The one-dimensional header generating unit is used to use the merged text content obtained when reaching the bottom-level node as the one-dimensional header of the column where the bottom-level node is located.
[0119] In this embodiment, a one-dimensional table header is "Isotopic ratios_ 207 Pb / 206 Pb_ratio”, “U-Pbages_age_2σ”, the table of one-dimensional header is as follows Figure 9 shown.
[0120] The first normalization module specifically includes:
[0121] The synonym judgment unit is used to use a pre-trained model to judge whether the texts in the one-dimensional header are synonymous. The specific judgment rules of the pre-trained model are:
[0122] A. Chinese and English professional terms are considered synonymous if they are equivalent;
[0123] B. Standard abbreviations and full names are considered synonymous;
[0124] C. The method name and the measurement result are considered to be synonymous;
[0125] D. Two expressions with the same meaning are considered synonymous;
[0126] The normalization unit is used to remove any one of the synonymous texts and generate a normalized one-dimensional header when the pre-trained model determines that the two one-dimensional headers are redundant.
[0127] The device provided in the embodiment of the present invention can be used to execute the method provided in the first embodiment of the present invention, and has the corresponding functions and beneficial effects of executing the method.
[0128] It is worth noting that in the embodiment of the above-mentioned determination device, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention.
[0129] The embodiments described above are merely illustrative, wherein the modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, i.e., they may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Those skilled in the art will readily appreciate that each embodiment may be implemented using software plus a necessary general-purpose hardware platform, or may be implemented solely through hardware, as long as the functionality or effect can be achieved.
[0130] It should be understood that the above embodiments and descriptions only describe the principles, main features and advantages of the present invention. Without departing from the spirit and scope of the present invention, the present invention may be subject to various changes and improvements, and these changes and improvements all fall within the scope of protection of the present invention.
Claims
1. A method for automatically converting complex headers into one dimension for multidimensional tables, characterized by: The steps include: (1) Obtain the target document and extract the table data in the target document to form a table data set T; (2) For any table t in the table data set T i , obtain the basic information of all header cells and store it in the header cell information set LC; (3) Calculate the header level depth L i , if L i >=2, it means that the table t i If it is a complex header, execute step (4); otherwise, execute step (6); (4) Take each header cell in the header cell set LC as a node, and i L i Starting from the header cell of the first layer, the table header tree structure is constructed by tracing back gradually upwards, and the text contents of all nodes are merged layer by layer starting from the root node until the bottom node is reached. The merged text content obtained when the bottom node is reached is used as the one-dimensional header of the column where the bottom node is located; (5) By judging the semantic similarity of the one-dimensional header, removing the redundant information in the one-dimensional header, applying the one-dimensional header to the table, and generating a normalized one-dimensional table; (6) Structuring the table to generate a standardized one-dimensional table; (7) Loop through steps (2) to (6) until all table data in the target document are processed.
2. The method for automatically converting complex headers into one dimension for multidimensional tables according to claim 1, characterized in that: The basic information of the header cell includes the text content of the header cell, the standard number of cells span merged by the header cell, the starting position col_start and ending position col_end of the row of the header cell in the table, and the next-level cell children of the current header cell. The next-level cell children are cells in the same column but in the next row as the current header cell.
3. The method for automatically converting complex headers into one dimension for multidimensional tables according to claim 1, characterized in that: Step (3) specifically includes: (3-1) Traverse table t i For each row, calculate the ratio of non-numeric cells in each row, rate = number of non-numeric cells / total number of columns; (3-2) When the rate is greater than the preset threshold, mark the corresponding line as the header row, and use the number of header rows as the header level depth L i ; (3-3) If L i >=2, it means that the table t i If it is a complex header, execute step (4); otherwise, execute step (6).
4. The method for automatically converting complex headers into one dimension for multidimensional tables according to claim 1, characterized in that: Step (4) specifically includes: (4-1) According to the basic information of each header cell in the header cell set LC, obtain the next level cell children of each header cell; (4-2) Take each header cell in the header cell set LC as a node, and i The header of the layer starts to trace back gradually, taking each header cell as the parent node of its next level cell children to build the header tree structure; (4-3) Based on the basic information of each header cell, check whether the text content of each header cell is empty. If so, inherit the text content from its nearest upper ancestor node; (4-4) Starting from the root node of the header tree structure, the header information in the header tree structure is merged layer by layer using a depth-first traversal algorithm until the bottom node is reached. Each time a node is traversed, the text content of the node is appended to the merged text content of all its ancestor nodes. A preset separator is set between the text contents of each node. (4-5) The merged text content obtained when reaching the bottom node is used as the one-dimensional header of the column where the bottom node is located.
5. The method for automatically converting complex headers into one dimension for multidimensional tables according to claim 1, characterized in that: Step (5) specifically includes: (5-1) Use the pre-trained model to determine whether the texts in the one-dimensional header are synonymous. The specific judgment rules of the pre-trained model are as follows: A. Chinese and English professional terms are considered synonymous if they are equivalent; B. Standard abbreviations and full names are considered synonymous; C. The method name and the measurement result are considered to be synonymous; D. Two expressions with the same meaning are considered synonymous; (5-2) When the pre-trained model determines that the two one-dimensional headers are synonymous, the two one-dimensional headers are considered redundant. Any one of the synonymous texts is removed to generate a normalized one-dimensional header.
6. An automatic one-dimensional processing system for complex headers of multi-dimensional tables, characterized by: include: The table extraction module is used to obtain the target document and extract the table data in the target document to form a table data set T; The table header cell extraction module is used for each table t in the table data set T i , obtain the basic information of all header cells and store it in the header cell information set LC; Hierarchical depth calculation module, used to calculate the hierarchical depth L of the header i , if L i >=2, it means that the table t i For complex headers, execute the one-dimensional header calculation module; Otherwise, execute the second normalization module; One-dimensional header calculation module, used to take each header cell in the header cell set LC as a node, from table t i L i Starting from the header cell of the first layer, the table header tree structure is constructed by tracing back gradually upwards, and the text contents of all nodes are merged layer by layer starting from the root node until the bottom node is reached. The merged text content obtained when the bottom node is reached is used as the one-dimensional header of the column where the bottom node is located; A first normalization module is configured to remove redundant information from the one-dimensional header by determining semantic similarity of the one-dimensional header, apply the one-dimensional header to the table, and generate a normalized one-dimensional table; The second normalization module is used to structure the table and generate a normalized one-dimensional table.
7. The method for automatically converting complex headers into one dimension for multidimensional tables according to claim 6, characterized in that: The basic information of the header cell includes the text content of the header cell, the standard number of cells span merged by the header cell, the starting position col_start and ending position col_end of the row of the header cell in the table, and the next-level cell children of the current header cell. The next-level cell children are cells in the same column but in the next row as the current header cell.
8. The method for automatically converting complex headers into one dimension for multidimensional tables according to claim 6, characterized in that: The hierarchical depth calculation module specifically includes: Proportional calculation unit, used to traverse the table t i For each row, calculate the ratio of non-numeric cells in each row, rate = number of non-numeric cells / total number of columns; The level depth calculation unit is used to mark the corresponding behavior as the header row when the rate is greater than the preset threshold, and the number of header rows is used as the level depth L of the header. i ; Execution judgment unit, used if L i >=2, it means that the table t i If the header is complex, execute the one-dimensional header calculation module; otherwise, execute the second normalization module.
9. The method for automatically converting complex headers into one dimension for multidimensional tables according to claim 6, characterized in that: The one-dimensional header calculation module specifically includes: The cell acquisition unit is used to obtain the next level cell children of each header cell according to the basic information of each header cell in the header cell set LC; The header tree construction unit is used to take each header cell in the header cell set LC as a node, starting from the L i The header of the layer starts to trace back gradually, taking each header cell as the parent node of its next level cell children to build the header tree structure; The assignment unit is used to retrieve the text content of each header cell based on the basic information of each header cell to see if it is empty. If so, the text content is inherited from its nearest upper ancestor node. The text content merging unit is used to merge the header information in the header tree structure layer by layer using a depth-first traversal algorithm starting from the root node of the header tree structure until it reaches the bottom node. Each time a node is traversed, the text content of the node is added to the end of the merged text content of all its ancestor nodes. A preset separator is set between the text content of each node. The one-dimensional header generating unit is used to use the merged text content obtained when reaching the bottom-level node as the one-dimensional header of the column where the bottom-level node is located.
10. The method for automatically converting complex headers into one dimension for multidimensional tables according to claim 6, characterized in that: The first normalization module specifically includes: The synonym judgment unit is used to use a pre-trained model to judge whether each text in the one-dimensional header is synonymous. The specific judgment rules of the pre-trained model are: A. Chinese and English professional terms are considered synonymous if they are equivalent; B. Standard abbreviations and full names are considered synonymous; C. The method name and the measurement result are considered to be synonymous; D. Two expressions with the same meaning are considered synonymous; The normalization unit is used to remove any one of the synonymous texts and generate a normalized one-dimensional header when the pre-trained model determines that the two one-dimensional headers are redundant.
Citation Information
Cited By
Image processing method and device
CN121415429A