Spreadsheet preprocessing method and device, computer equipment, readable storage medium and program product

By identifying and tracing merged cells in spreadsheets, the accuracy issue of large language models when processing merged cells is resolved, enabling more efficient spreadsheet data processing.

CN120822501APending Publication Date: 2025-10-21CHN ENERGY NEW ENERGY TECHNOLOGY RESEARCH INSTITUTE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510731750.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

When processing spreadsheets containing merged cells, large language models cannot accurately determine the meaning of the merged cells, resulting in inaccurate processing of complex spreadsheet tasks.

Method used

By obtaining the header cells in the spreadsheet, identifying the properties of the blank cells, and determining that they are merged cells, the direction of the associated cells is traced and the cell contents are obtained, the updated spreadsheet is generated and the token sequence set is filtered to improve accuracy.

Benefits of technology

Improves the accuracy of large language models in processing complex spreadsheets, ensuring the integrity and data quality of spreadsheet content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120822501A_ABST
    Figure CN120822501A_ABST
Patent Text Reader

Abstract

The invention relates to a spreadsheet preprocessing method and device, computer equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: acquiring a to-be-processed spreadsheet, and determining header cells contained in the spreadsheet; if the header cells comprise blank cells, cell attributes of the blank cells are obtained; if the cell attributes represent that the blank cells are merged cells, determining associated cells of the blank cells and tracing directions corresponding to the associated cells; and cell content tracing is performed according to the tracing direction, and cell content of blank cells is determined. By adopting the method, the accuracy of processing the complex spreadsheet task by the large language model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to an electronic spreadsheet preprocessing method, apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Art

[0002] As a structured data storage and presentation tool, spreadsheets cover a wide range of formats, from basic row and column layouts to cell merging and splitting. This rich and diverse format information cannot be covered by simple text descriptions. Together, they construct the unique visual presentation and logical structure of spreadsheets, which is crucial for accurately understanding the table content, performing calculation tasks, and generating expected output results.

[0003] Traditionally, large language models are trained primarily on text data, with their core capabilities focused on semantic understanding, grammatical analysis, and generation of natural language text. However, when training on spreadsheets containing merged cells, large language models cannot directly determine the meaning of these merged cells, as these cells may include blank cells. This can lead to inaccuracies in handling complex spreadsheet tasks. Summary of the Invention

[0004] Based on this, it is necessary to provide a spreadsheet preprocessing method, apparatus, computer equipment, computer-readable storage medium and computer program product that can improve the accuracy of large language models in processing complex spreadsheet tasks in response to the above technical problems.

[0005] In a first aspect, the present application provides a spreadsheet preprocessing method, comprising:

[0006] Obtaining a spreadsheet to be processed, and determining header cells contained in the spreadsheet;

[0007] When the table header cells include blank cells, obtaining cell attributes of the blank cells;

[0008] If the cell attribute indicates that the blank cell is a merged cell, determining an associated cell of the blank cell and a tracing direction corresponding to the associated cell;

[0009] The cell content is traced back according to the tracing direction to determine the cell content of the blank cell.

[0010] In one embodiment, the number of the associated cells is multiple; the number of the cell contents is multiple; and the method further includes:

[0011] generating a plurality of updated electronic spreadsheets according to the contents of each cell;

[0012] For each of the updated electronic forms, generating a token sequence set corresponding to the updated electronic form;

[0013] A token sequence set whose token sequence repetition degree meets a repetition threshold condition is selected from each of the token sequence sets as a target token sequence set of the electronic form.

[0014] In one embodiment, obtaining a spreadsheet to be processed and determining header cells contained in the spreadsheet includes:

[0015] Get the spreadsheet to be processed and the header database;

[0016] Based on the header database, a header cell matching the header database is determined from a plurality of cells of the electronic table.

[0017] In one embodiment, obtaining the electronic form to be processed includes:

[0018] Get the initial spreadsheet to be processed;

[0019] Determining the degree of heterogeneity corresponding to each table row and each table column in the initial electronic spreadsheet;

[0020] The table rows or table columns with a heterogeneity greater than a heterogeneity threshold in each of the heterogeneity degrees are regarded as heterogeneous rows or heterogeneous columns;

[0021] determining homogeneous rows and homogeneous columns in each of the table rows and each of the table columns;

[0022] The homogeneous rows whose cell distances from the heterogeneous rows meet the distance condition are deleted, and the homogeneous columns whose cell distances from the heterogeneous columns meet the distance condition are deleted to obtain an electronic spreadsheet to be processed.

[0023] In one embodiment, tracing back the cell content according to the tracing direction to determine the cell content of the blank cell includes:

[0024] Obtain the next cell content of the next cell located in the tracing direction of the blank cell;

[0025] In the case where the next cell content is blank content, return to the step of obtaining the next cell content of the next cell located in the tracing direction of the blank cell, until the next cell content obtained is non-blank content, and use the non-blank content as the cell content of the blank cell.

[0026] In one embodiment, generating, for each updated electronic form, a token sequence set corresponding to the updated electronic form includes:

[0027] determining an order of cell attributes in the updated spreadsheet;

[0028] For each updated electronic form, a token sequence is generated for non-header content in the updated electronic form based on the cell attribute sequence; each token sequence constitutes a token sequence set.

[0029] In a second aspect, the present application further provides an electronic spreadsheet preprocessing device, comprising:

[0030] A header cell determination module is used to obtain a spreadsheet to be processed and determine the header cells contained in the spreadsheet;

[0031] A cell attribute acquisition module, configured to acquire cell attributes of a blank cell when the table header cells include the blank cell;

[0032] a tracing direction determining module, configured to determine an associated cell of the blank cell and a tracing direction corresponding to the associated cell if the cell attribute indicates that the blank cell is a merged cell;

[0033] The cell content tracing module is used to trace the cell content according to the tracing direction and determine the cell content of the blank cell.

[0034] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.

[0035] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above-described method when executed by a processor.

[0036] In a fifth aspect, the present application further provides a computer program product, which includes a computer program that implements the steps of the above method when executed by a processor.

[0037] The above-mentioned spreadsheet preprocessing method, apparatus, computer equipment, computer-readable storage medium and computer program product obtain a spreadsheet to be processed and determine the header cells contained in the spreadsheet. The header cells and value cells in the spreadsheet can be distinguished first. If the header cells contain blank cells, the cell properties of the blank cells need to be obtained. When the cell properties indicate that the blank cells are merged cells, since the merged cells generally merge the contents of the cells with which the cells are associated, the associated cells of the blank cells and the tracing direction corresponding to the associated cells can be determined. The cell contents are traced according to the tracing direction to determine the cell contents of the blank cells and obtain a new spreadsheet containing all the cell contents, thereby improving the accuracy of the large language model in processing complex spreadsheets. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.

[0039] Figure 1 A diagram showing an application environment of a spreadsheet preprocessing method according to an embodiment;

[0040] Figure 2 1 is a flow chart of a method for preprocessing a spreadsheet in one embodiment;

[0041] Figure 3 is a sample diagram of a spreadsheet in one embodiment;

[0042] Figure 4 is a flow chart of a spreadsheet preprocessing method according to another embodiment;

[0043] Figure 5 is a structural block diagram of an electronic form preprocessing device in one embodiment;

[0044] Figure 6 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0045] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0046] The electronic form preprocessing method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. The terminal 102 communicates with the server 104 via a network. The data storage system can store data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. The terminal 102 can be, but is not limited to, various personal computers, laptops, smart phones, tablets, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart car devices, projection devices, etc. Portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. The server 104 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services. Specifically, during the process of preprocessing the spreadsheet, the server 104 obtains the spreadsheet to be processed from the terminal 102 and determines the header cells contained in the spreadsheet; if the header cells contain blank cells, the cell properties of the blank cells are obtained; if the cell properties indicate that the blank cells are merged cells, the associated cells of the blank cells and the tracing direction corresponding to the associated cells are determined; the cell content is traced according to the tracing direction to determine the cell content of the blank cells.

[0047] In an exemplary embodiment, Figure 2 As shown, a spreadsheet preprocessing method is provided, which is applied to Figure 1 The server 104 in the example is used as an example to illustrate the process, including the following steps S202 to S208.

[0048] Step S202 : obtaining the electronic form to be processed and determining the header cells contained in the electronic form.

[0049] A spreadsheet is a computer application that organizes data in rows and columns, typically consisting of multiple cells. Each cell can store various types of data, such as text, numbers, formulas, and dates. It supports operations such as sorting, filtering, calculations, and charting, and is widely used in fields such as data analysis and project management. Common spreadsheet software includes Microsoft Excel and WPS Spreadsheet. In a spreadsheet, header cells are located at the top (horizontally) or leftmost (vertically) of a table, in one or more rows or columns. They summarize and identify the data stored in the cells below (or to the right of) that row or column, allowing users to quickly understand the meaning of each column (or row) in the table. For example, in a table recording student grades, header cells might include "Student ID," "Name," "Chinese Score," "Math Score," and so on.

[0050] Specifically, in the data processing and analysis workflow, obtaining the spreadsheet to be processed is a crucial first step. After the server successfully retrieves the spreadsheet from the terminal, the next critical task is to identify the header cells contained therein. Header cells serve as "navigational markers" in spreadsheets, providing a clear semantic framework for the data in the entire table. In most standard spreadsheet formats, header cells are typically located in the first row (horizontally) or first column (vertically), but the actual location may vary depending on the table designer's intent. To accurately identify header cells, various strategies can be employed. For example, one can identify them by checking whether the cell content meets header characteristics (such as concise content, generalization, and the absence of specific data values); or by leveraging table formatting information (for example, some header cells may have specific font styles, background colors, or bolding effects). Furthermore, for more complex spreadsheets, pre-set rules or templates can be used to locate header cells, ensuring accurate identification of these cells that are crucial for subsequent data processing. This lays a solid foundation for subsequent operations such as classification, calculation, and analysis of the table data based on header information.

[0051] Step S204: if the header cells include blank cells, obtain the cell attributes of the blank cells.

[0052] A blank cell is a cell in a spreadsheet that contains no data (including text, numbers, formulas, etc.). This can occur for a variety of reasons, such as when a table was created without any data, when data was deleted during processing, or when cells were merged. While blank cells don't display specific data in a spreadsheet, they may contain some cell attribute information. Cell attributes are parameters that describe cell characteristics and formatting. For example, cell attributes can include cell style and background color.

[0053] Specifically, if the header cells contain blank cells, the large language model will not know the specific content meaning of the header cell when querying the header cell. Therefore, it is necessary to determine the cell attributes of the blank cell to facilitate subsequent operations on the blank cell.

[0054] Step S206 : If the cell attribute indicates that the blank cell is a merged cell, then the associated cells of the blank cell and the tracing directions corresponding to the associated cells are determined.

[0055] Among them, merging cells is an operation in spreadsheets, which is to merge two or more adjacent cells into a larger cell. The merged cells are visually presented as a whole, usually only retaining the content of one end cell, and the content of the other merged cells will be hidden. However, in fact, the merged cells are still indexed as multiple separate cells when performing cell indexing, that is, only the content of one end cell can be indexed, and the contents of other cells are blank cells. Associated cells refer to cells that are likely to have an associated relationship with the cell content of the cell, and are generally located in the four directions of the cell, above, below, left, and right. When processing merged cells, the tracing direction refers to the directional guidance for determining the positional relationship between associated cells and cells. It clarifies that when searching or processing associated cells, you should start from the blank cell being processed and follow which direction (such as up, down, left, and right) to locate and identify other cells involved in the merge.

[0056] Specifically, if the cell attributes indicate that a blank cell is a merged cell, this indicates that the blank cell does have content but is displayed as a blank cell after being merged. Therefore, it is necessary to determine the associated cells of the blank cell and the tracing directions corresponding to the associated cells so that subsequent tracing can be performed according to the tracing directions to determine the cell content of the blank cell. For example, the number of associated cells can be one or more, and accordingly, the number of tracing directions can also be one or more.

[0057] Optionally, if the cell attribute indicates that the blank cell is a non-merged cell, that is, the cell content of the blank cell is indeed blank, then the row or column containing the value cell where the cell is located can be deleted.

[0058] Step S208: trace back the cell contents in the tracing direction to determine the cell contents of the blank cells.

[0059] Cell content refers to the specific information stored in each cell in a spreadsheet.

[0060] Specifically, since the blank cell is a merged cell, tracing back in the tracing direction will definitely find other cells with cell content that are merged with the blank cell, and the above cell content can be used as the cell content of the blank cell. For example, since there can be multiple tracing directions, there may also be multiple cell contents of the blank cell. The final target cell content can be determined based on subsequent filtering operations.

[0061] The above-mentioned spreadsheet preprocessing method obtains the spreadsheet to be processed and determines the header cells contained in the spreadsheet. The header cells and value cells in the spreadsheet can be distinguished first. If the header cells contain blank cells, the cell properties of the blank cells need to be obtained. When the cell properties indicate that the blank cells are merged cells, since the merged cells generally merge the contents of the cells with which the cells are associated, the associated cells of the blank cells and the tracing direction corresponding to the associated cells can be determined. The cell contents are traced according to the tracing direction to determine the cell contents of the blank cells and obtain a new spreadsheet containing all the contents, thereby improving the accuracy of the large language model in processing complex spreadsheets.

[0062] In an exemplary embodiment, the number of associated cells is multiple; the number of cell contents is multiple; the spreadsheet preprocessing method also includes: generating multiple updated spreadsheets according to the contents of each cell; for each updated spreadsheet, generating a token sequence set corresponding to the updated spreadsheet; and selecting a token sequence set whose token sequence repetition degree meets the repetition threshold condition from each token sequence set as the target token sequence set of the spreadsheet.

[0063] Among them, updating the spreadsheet is to generate a new version of the spreadsheet after a series of operations or processing on the basis of the original spreadsheet. In this embodiment, a series of operations or processing refers to updating blank cells according to the content of each cell. Token sequence set: In the field of data processing and text analysis, a token is a smallest semantic unit that decomposes raw data. For a spreadsheet, information such as cell content, cell format, and cell position can be converted into tokens. A token sequence is a sequence formed by arranging and combining these tokens according to certain rules. A token sequence set is a combination of multiple such token sequences, which comprehensively describes the characteristics and structural information of a spreadsheet. Token sequence repetition is an indicator used to measure the degree of similarity between token sequences in a token sequence set. For example, Figure 3 As shown, the token sequence can be "region, North China, Shanxi, January, sales, 100", "region, North China, Shanxi, February, sales, 200", "region, North China, Beijing, January, sales, 300". It can be understood that the similarity among the token sequences in the correct token sequence set is the highest.

[0064] Specifically, the number of associated cells can be multiple, and correspondingly, the number of cell contents is also multiple. On this basis, in order to improve the accuracy of spreadsheet preprocessing, multiple updated spreadsheets can be generated according to the contents of each cell, and a token sequence set corresponding to each updated spreadsheet is generated for the updated spreadsheet. Since the similarity among the token sequences in the correct token sequence set is the highest, a token sequence set whose token sequence repetition meets the repetition threshold condition can be selected from each token sequence set as the target token sequence set of the spreadsheet.

[0065] In this embodiment, a token sequence set whose token sequence repetition degree meets a repetition threshold condition is selected from each token sequence set as a target token sequence set for the electronic form, which can improve the accuracy of electronic form preprocessing.

[0066] In an exemplary embodiment, obtaining a spreadsheet to be processed and determining header cells contained in the spreadsheet include: obtaining the spreadsheet to be processed and a header database; and determining, based on the header database, header cells that match the header database from multiple cells of the spreadsheet.

[0067] The header database is a collection specifically used to store and manage spreadsheet header information.

[0068] Specifically, after obtaining the spreadsheet to be processed, since the spreadsheet generally contains multiple cells, which include header cells and value cells, the server cannot directly distinguish them. Therefore, a header database containing multiple types of headers can be obtained. Based on the header database, the header cells that match the header database are determined from the multiple cells of the spreadsheet, and the rest are value cells.

[0069] In this embodiment, the header cells and value cells in the electronic form are distinguished by matching the header database, which can improve the accuracy of the distinction and further improve the accuracy of the electronic form preprocessing.

[0070] In an exemplary embodiment, obtaining a spreadsheet to be processed includes: obtaining an initial spreadsheet to be processed; determining the degree of heterogeneity corresponding to each table row and each table column in the initial spreadsheet; treating table rows or table columns with a degree of heterogeneity greater than a heterogeneity threshold as heterogeneous rows or heterogeneous columns; determining homogeneous rows and homogeneous columns in each table row and each table column; deleting homogeneous rows whose cell distances to heterogeneous rows meet a distance condition, and deleting homogeneous columns whose cell distances to heterogeneous columns meet a distance condition, to obtain the spreadsheet to be processed.

[0071] Heterogeneity is a quantitative metric that measures the diversity, variability, and complexity of data in rows or columns within a spreadsheet. For example, high heterogeneity can occur when a row in a student's report card contains a variety of data types, such as text (name), numbers (95 in math), dates (birthday), and Boolean values ​​(whether or not to repeat a grade), or when a column's data distribution is extremely uneven (e.g., an "age" column contains data ranging from 10 to 80 years old). Low heterogeneity can occur when all rows in a product specification table have "weight" columns that are numerical and concentrated in the 50-60 kg range, or when all columns are standardized text (e.g., model codes). The heterogeneity threshold is a pre-set threshold that distinguishes normal data from heterogeneous data. After calculating the heterogeneity of a table row or column, comparing the heterogeneity with the threshold determines whether the row or column is heterogeneous. Homogeneous rows and columns are those with similar data features and high consistency. These rows or columns exhibit similar patterns or distributions across all fields. Cell distance is the distance between rows or columns. For example, the cell distance between the first and second rows of a spreadsheet is 1. A distance condition is a pre-set criterion used to determine whether the distance between homogeneous rows or columns and heterogeneous rows or columns meets a threshold. For example, if the distance condition is greater than 5, homogeneous rows with a cell distance greater than 5 from heterogeneous rows can be deleted, and homogeneous columns with a cell distance that meets the distance condition from heterogeneous columns can be deleted, resulting in the processed spreadsheet.

[0072] In a specific embodiment, the heterogeneity of each row / column can be calculated using the following formula:

[0073] Among them, H(R / C) is the heterogeneity of rows / columns, is the probability of the i-th unique value appearing, and n is the number of unique values.

[0074] Specifically, in this embodiment, a heterogeneity threshold θ can be set to identify rows / columns with a heterogeneity higher than the heterogeneity threshold as heterogeneous rows and columns, and at the same time mark the heterogeneous rows and columns as structural features. Finally, homogeneous rows and columns that are far away from the structural features (the distance from the cells of the heterogeneous rows meets the distance condition) are removed, the key information is retained, and a simplified version of the spreadsheet is generated.

[0075] In this embodiment, heterogeneous rows and columns are first identified, and then homogeneous rows and columns in the spreadsheet are determined. Homogeneous rows whose cell distances to heterogeneous rows meet the distance conditions are deleted, and homogeneous columns whose cell distances to heterogeneous columns meet the distance conditions are deleted, to obtain a simplified version of the spreadsheet, which can effectively improve the data quality and analysis efficiency of the spreadsheet.

[0076] In an exemplary embodiment, cell content is traced in a tracing direction to determine the cell content of a blank cell, including: obtaining the next cell content of the next cell located in the tracing direction of the blank cell; if the next cell content is blank content, returning to the step of obtaining the next cell content of the next cell located in the tracing direction of the blank cell, until the obtained next cell content is non-blank content, and using the non-blank content as the cell content of the blank cell.

[0077] Non-blank content refers to cells that contain valid data (such as text, numbers, formula calculation results, etc.) and are not empty strings or null.

[0078] Specifically, since the number of cells merged by the merged cell may be more than two, that is, there may be a situation where the merged cell includes multiple blank cells. Therefore, you can first obtain the next cell content of the next cell located in the tracing direction of the blank cell. If the next cell content is blank content, go back to the step of obtaining the next cell content of the next cell located in the tracing direction of the blank cell, and continue to obtain the cell content of the next cell in the tracing direction until the next cell content obtained is non-blank content, and use the non-blank content as the cell content of the blank cell.

[0079] In this embodiment, the specific cell content of the blank cell can be finally determined through cyclic tracing, thereby ensuring the integrity of the content of the electronic spreadsheet.

[0080] In an exemplary embodiment, for each updated spreadsheet, a token sequence set corresponding to the updated spreadsheet is generated, including: determining the order of cell attributes in the updated spreadsheet; for each updated spreadsheet, generating a token sequence for non-header content in the updated spreadsheet based on the order of cell attributes; each token sequence constitutes a token sequence set.

[0081] Among them, the cell attribute order refers to the order of arrangement according to cell attributes, which is generally represented by header cell-value cell. Furthermore, for more complex multi-header merged item spreadsheets, the order of large range first and small range last can be followed. For example, Figure 3 In the table shown, the order of cell attributes is: Region, North China, Shanxi, January, Sales, 100.

[0082] Specifically, in order to make the updated spreadsheet easier to be processed by the large language model, the updated spreadsheet can be converted into the form of a token sequence set. That is, the cell attribute order in the updated spreadsheet can be determined first. For each updated spreadsheet, a token sequence is generated for the non-header content in the updated spreadsheet based on the cell attribute order. In this way, a token sequence set including multiple token sequences can be obtained, which is convenient for the large language model to encode the spreadsheet.

[0083] In some specific embodiments, the token sequence can be represented in the form of key-value pairs: .

[0084] Token represents a token, value represents a numeric value, and address represents an address. For example, if the cell attribute order is: Region, North China, Shanxi, January, Sales, 100, and you enter a value of 100, the address will be: Region, North China, Shanxi, January, Sales.

[0085] In a specific embodiment, Figure 4 As shown, a spreadsheet preprocessing method is also provided, comprising:

[0086] Step S401: obtaining an initial electronic form and a table header database to be processed, and determining the degree of heterogeneity corresponding to each table row and each table column in the initial electronic form;

[0087] Step S402: Table rows or columns with a heterogeneity greater than a heterogeneity threshold in each heterogeneity degree are regarded as heterogeneous rows or columns, and homogeneous rows and columns in each table row or column are determined;

[0088] Step S403 , deleting homogeneous rows whose cell distances to heterogeneous rows satisfy a distance condition, and deleting homogeneous columns whose cell distances to heterogeneous columns satisfy a distance condition, to obtain a spreadsheet to be processed;

[0089] Step S404 , based on the header database, determining a header cell that matches the header database from a plurality of cells in the electronic sheet;

[0090] Step S405: if the header cells include blank cells, obtain the cell attributes of the blank cells;

[0091] Step S406 , if the cell attribute indicates that the blank cell is a merged cell, then determining the associated cells of the blank cell and the tracing direction corresponding to the associated cells;

[0092] Step S407, obtaining the next cell content of the next cell in the tracing direction of the blank cell;

[0093] Step S408: If the next cell content is blank, return to the step of obtaining the next cell content of the next cell in the tracing direction of the blank cell, and continue until the next cell content obtained is non-blank content, and use the non-blank content as the cell content of the blank cell.

[0094] wherein the number of associated cells is multiple; the number of cell contents is multiple;

[0095] Step S409, generating multiple updated electronic tables according to the contents of each cell, and determining the order of cell attributes in the updated electronic tables;

[0096] Step S410 , for each updated electronic form, generating a token sequence for non-header content in the updated electronic form based on the cell attribute sequence;

[0097] Among them, each token sequence constitutes a token sequence set;

[0098] Step S411 : selecting a token sequence set whose token sequence repetition degree meets a repetition threshold condition from each token sequence set as a target token sequence set of the electronic form.

[0099] It should be understood that, although the steps in the flowcharts of the above embodiments are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the flowcharts of the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0100] Based on the same inventive concept, embodiments of the present application also provide a spreadsheet preprocessing device for implementing the aforementioned spreadsheet preprocessing method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more embodiments of the spreadsheet preprocessing device provided below can be found in the above-described limitations of the spreadsheet preprocessing method and will not be further elaborated here.

[0101] In an exemplary embodiment, Figure 5 As shown, a spreadsheet preprocessing device 500 is provided, comprising: a header cell determination module 502, a cell attribute acquisition module 504, a tracing direction determination module 506, and a cell content tracing module 508, wherein:

[0102] The header cell determination module 502 is used to obtain the electronic form to be processed and determine the header cells contained in the electronic form;

[0103] The cell attribute acquisition module 504 is used to acquire the cell attributes of the blank cells when the header cells include blank cells;

[0104] The tracing direction determining module 506 is configured to determine the associated cells of the blank cell and the tracing direction corresponding to the associated cells if the cell attribute indicates that the blank cell is a merged cell;

[0105] The cell content tracing module 508 is used to trace the cell content according to the tracing direction and determine the cell content of the blank cell.

[0106] In an exemplary embodiment, the number of associated cells is multiple; the number of cell contents is multiple. In this embodiment, the electronic form pre-processing device 500 further includes a target token sequence set determination module, including:

[0107] An update spreadsheet generating unit, configured to generate a plurality of update spreadsheets according to the contents of each cell;

[0108] a token sequence set generating unit, configured to generate, for each updated electronic form, a token sequence set corresponding to the updated electronic form;

[0109] The target token sequence set determining unit is used to select a token sequence set whose token sequence repetition degree meets a repetition threshold condition from each token sequence set as the target token sequence set of the electronic form.

[0110] In an exemplary embodiment, the header cell determination module 502 is configured to:

[0111] Get the spreadsheet to be processed and the header database;

[0112] Based on the header database, a header cell matching the header database is determined from a plurality of cells in the spreadsheet.

[0113] In an exemplary embodiment, the header cell determination module 502 is further configured to:

[0114] Get the initial spreadsheet to be processed;

[0115] Determine the degree of heterogeneity corresponding to each row and column in the initial spreadsheet;

[0116] The table rows or columns with a heterogeneity greater than the heterogeneity threshold in each heterogeneity degree are regarded as heterogeneous rows or heterogeneous columns;

[0117] Identify homogeneous rows and columns in each table row and each table column;

[0118] The homogeneous rows whose cell distances to the heterogeneous rows satisfy the distance condition are deleted, and the homogeneous columns whose cell distances to the heterogeneous columns satisfy the distance condition are deleted, to obtain a spreadsheet to be processed.

[0119] In an exemplary embodiment, the cell content tracing module 508 is specifically configured to:

[0120] Get the next cell content of the next cell in the tracing direction of the blank cell;

[0121] If the next cell content is blank, return to the step of obtaining the next cell content of the next cell located in the tracing direction of the blank cell until the next cell content obtained is non-blank content, and use the non-blank content as the cell content of the blank cell.

[0122] In an exemplary embodiment, the token sequence set generation unit is specifically configured to:

[0123] Determine the order in which cell properties in a spreadsheet are updated;

[0124] For each updated spreadsheet, a token sequence is generated for non-header contents in the updated spreadsheet based on the cell attribute sequence; each token sequence constitutes a token sequence set.

[0125] Each module in the electronic spreadsheet preprocessing device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0126] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 6 As shown. The computer device includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals via wired or wireless means, and the wireless means can be implemented via Wi-Fi, a mobile cellular network, near-field communication (NFC), or other technologies. When executed by the processor, the computer program implements a spreadsheet preprocessing method. The display unit of the computer device is used to form a visually visible image, and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse.

[0127] Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0128] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.

[0129] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.

[0130] In one embodiment, a computer program product is provided, comprising a computer program, which implements the steps of the above method when executed by a processor.

[0131] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0132] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), quantum computing-based data processing logic devices, artificial intelligence (AI) processors, and the like.

[0133] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0134] The above embodiments merely illustrate several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art may make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A spreadsheet preprocessing method, characterized in that: The method comprises: Obtaining a spreadsheet to be processed, and determining header cells contained in the spreadsheet; If the table header cells contain blank cells, obtain the cell attributes of the blank cells; If the cell attribute indicates that the blank cell is a merged cell, determining an associated cell of the blank cell and a tracing direction corresponding to the associated cell; The cell content is traced back according to the tracing direction to determine the cell content of the blank cell.

2. The method according to claim 1, characterized in that The number of the associated cells is multiple; the number of the cell contents is multiple; and the method further includes: generating a plurality of updated electronic spreadsheets according to the contents of each cell; For each of the updated electronic forms, generating a token sequence set corresponding to the updated electronic form; A token sequence set whose token sequence repetition degree meets a repetition threshold condition is selected from each of the token sequence sets as a target token sequence set of the electronic form.

3. The method according to claim 1, characterized in that The step of obtaining a spreadsheet to be processed and determining header cells contained in the spreadsheet includes: Get the spreadsheet to be processed and the header database; Based on the header database, a header cell matching the header database is determined from a plurality of cells of the electronic table.

4. The method according to claim 1, wherein The step of obtaining the electronic form to be processed includes: Get the initial spreadsheet to be processed; Determining the degree of heterogeneity corresponding to each table row and each table column in the initial electronic spreadsheet; The table rows or table columns with a heterogeneity greater than a heterogeneity threshold in each of the heterogeneity degrees are regarded as heterogeneous rows or heterogeneous columns; determining homogeneous rows and homogeneous columns in each of the table rows and each of the table columns; The homogeneous rows whose cell distances from the heterogeneous rows meet the distance condition are deleted, and the homogeneous columns whose cell distances from the heterogeneous columns meet the distance condition are deleted to obtain an electronic spreadsheet to be processed.

5. The method according to claim 1, wherein Tracing the cell content according to the tracing direction to determine the cell content of the blank cell includes: Obtain the next cell content of the next cell located in the tracing direction of the blank cell; In the case where the next cell content is blank content, return to the step of obtaining the next cell content of the next cell located in the tracing direction of the blank cell, until the next cell content obtained is non-blank content, and use the non-blank content as the cell content of the blank cell.

6. The method according to claim 2, characterized in that The step of generating, for each updated electronic form, a token sequence set corresponding to the updated electronic form includes: determining an order of cell attributes in the updated spreadsheet; For each updated electronic form, a token sequence is generated for non-header content in the updated electronic form based on the cell attribute sequence; each token sequence constitutes a token sequence set.

7. An electronic spreadsheet preprocessing device, characterized in that: The device comprises: A header cell determination module is used to obtain a spreadsheet to be processed and determine the header cells contained in the spreadsheet; A cell attribute acquisition module, configured to acquire cell attributes of a blank cell when the table header cells include the blank cell; a tracing direction determining module, configured to determine an associated cell of the blank cell and a tracing direction corresponding to the associated cell if the cell attribute indicates that the blank cell is a merged cell; The cell content tracing module is used to trace the cell content according to the tracing direction and determine the cell content of the blank cell.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.