Table formula processing method and device, electronic equipment and storage medium
By obtaining and structured processing of adjacent cell data of target tables in the spreadsheet software, and using pre-trained neural network to predict table formulas, the inefficiency and accuracy problems caused by user input of natural language are solved, and efficient and accurate table data processing is achieved.
Patent Information
- Application Number
- CN202311840836.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-01
AI Technical Summary
In existing spreadsheet software, users need to enter natural language to obtain table formulas, resulting in low efficiency in table data processing and high formula uncertainty and error probability.
By obtaining adjacent cells of the target cell in the target table and structuring the table, the pre-trained neural network predicts table formulas based on adjacent cells and structured data, and directly displays the results at the target cell position.
Table formulas can be obtained without the user entering natural language, which improves data processing efficiency, reduces uncertainty and error probability, and enhances the accuracy of table data processing.
Smart Images

Figure CN120235129A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a method, device, electronic device, and storage medium for processing table formulas. Background Art
[0002] Currently, spreadsheet software (Excel) has been gradually used by people due to its powerful data visualization and analysis capabilities. Among them, table formulas are very powerful tools in spreadsheet software, which can be used to solve various problems such as mathematics, statistics, logic, text, date, and time, and can help users better process and analyze data, etc.
[0003] However, as a very powerful tool in spreadsheet software, table formulas have a certain usage threshold for many users, such as function selection, formula syntax understanding, data reference, function nesting, etc. Such a relatively high usage threshold of table formulas brings certain difficulties to people in using table formulas.
[0004] For this reason, large language models are introduced. Users can input natural language requirements, and the natural language requirements can be converted into corresponding table formulas. Since users need to input natural language requirements to obtain table formulas, it will reduce the efficiency of table data processing, and the resulting table formulas have uncertainty and error probabilities, which will reduce the accuracy of table data processing. Summary of the Invention
[0005] In order to solve the above technical problems that since users need to input natural language requirements to obtain table formulas, it will reduce the efficiency of table data processing, and the resulting table formulas have uncertainty and error probabilities, which will reduce the accuracy of table data processing, embodiments of this application provide a method, device, electronic device, and storage medium for processing table formulas. The specific technical solutions are as follows:
[0006] In the first aspect of the embodiments of this application, first, a method for processing table formulas is provided, and the method includes:
[0007] Obtain a target table, and determine adjacent cells corresponding to a target cell in the target table;
[0008] Perform structured processing on the target table to obtain structured data corresponding to the target table;
[0009] Predict a target table formula according to the target cell, the adjacent cells, and the structured data;
[0010] Display the target table formula at a corresponding position of the target cell.
[0011] In an alternative embodiment, the adjacent cells include upper adjacent cells, lower adjacent cells, left adjacent cells, and right adjacent cells;
[0012] Determining the adjacent cells corresponding to the target cell in the target table includes:
[0013] Determining the first preset number of the upper adjacent cells above the target cell in the target table;
[0014] Determining the second preset number of the lower adjacent cells below the target cell in the target table;
[0015] Determining the third preset number of the left adjacent cells to the left of the target cell in the target table;
[0016] Determining the fourth preset number of the right adjacent cells to the right of the target cell in the target table.
[0017] In an alternative embodiment, the structuring the target table to obtain the structured data corresponding to the target table includes:
[0018] Identifying the table header in the target table and splitting the target table to obtain the table header, each row of cells in the target table, and each column of cells in the target table;
[0019] Adding each row of cells in the target table to the row cell set and adding each column of cells in the target table to the column cell set.
[0020] In an alternative embodiment, the predicting the target table formula based on the target cell, the adjacent cells, and the structured data includes:
[0021] Inputting the target cell, the adjacent cells, the table header, the row cell set, and the column cell set into a pre-trained neural network to obtain a target function and target parameters;
[0022] Combining the target function and the target parameters to obtain the target table formula.
[0023] In an alternative embodiment, the pre-trained neural network includes a pre-trained encoder, a pre-trained decoder, a pre-trained attention network, and a pre-trained large language model;
[0024] The inputting the target cell, the adjacent cells, the table header, the row cell set, and the column cell set into a pre-trained neural network to obtain a target function and target parameters includes:
[0025] Using the pre-trained encoder, encode the table header to obtain a table header embedding vector;
[0026] Using the pre-trained encoder, encode all row cells in the row cell set to obtain a row embedding vector;
[0027] Using the pre-trained encoder, encode all column cells in the column cell set to obtain a column embedding vector;
[0028] Input the target cell, the adjacent cell, the row embedding vector, and the column embedding vector into the pre-trained large language model to obtain a first output result;
[0029] Input the table header embedding vector into the pre-trained attention network to obtain a second output result;
[0030] Using the pre-trained decoder, decode the first output result and the second output result to obtain an objective function and objective parameters.
[0031] In an alternative embodiment, the inputting the target cell, the adjacent cell, the row embedding vector, and the column embedding vector into the pre-trained large language model to obtain a first output result includes:
[0032] Determine a first coordinate corresponding to the target cell and a second coordinate corresponding to the adjacent cell;
[0033] Input the first coordinate, the second coordinate, the row embedding vector, and the column embedding vector into the pre-trained large language model to obtain a first output result.
[0034] In an alternative embodiment, the determining a first coordinate corresponding to the target cell and a second coordinate corresponding to the adjacent cell includes:
[0035] Determine a first position of the target cell in the target table, determine the first position as the coordinate origin, and determine the coordinate origin as the first coordinate corresponding to the target cell;
[0036] Determine a second position of the adjacent cell in the target table, and determine the second coordinate corresponding to the adjacent cell according to the position difference between the first position and the second position.
[0037] In an alternative embodiment, the determining adjacent cells corresponding to a target cell in the target table includes:
[0038] In response to a filling operation for a target cell in the target table, determine the first data filled in the target cell;
[0039] In the case where the first data is preset data, determine the adjacent cells corresponding to the target cell in the target table.
[0040] In an optional embodiment, after displaying the target table formula at the corresponding position of the target cell, the method further includes:
[0041] In response to a re-filling operation for the target cell in the target table, determine the second data re-filled in the target cell;
[0042] In the case where the second data represents a function, display the target function at the corresponding position of the target cell.
[0043] In an optional embodiment, after displaying the target function at the corresponding position of the target cell, the method further includes:
[0044] In response to a re-filling operation for the target cell in the target table, determine the third data re-filled in the target cell;
[0045] In the case where the third data represents a parameter, display the target parameter at the corresponding position of the target cell.
[0046] In an optional embodiment, the method further includes:
[0047] In response to a preset operation for a target column cell in the target table, detect whether at least two identical table formulas are filled in the cells of the target column cell;
[0048] In the case where at least two identical table formulas are filled in the cells of the target column cell, trigger a table formula filling prompt;
[0049] The table formula filling prompt is used to prompt the remaining cells of the target column cell to fill the table formula.
[0050] In an optional embodiment, before executing the method, it further includes:
[0051] Obtain a sample table and determine the adjacent sample cells corresponding to the sample cells in the sample table;
[0052] Identify the sample table headers in the sample table and split the sample table to obtain the sample table headers, each row of sample cells in the sample table, and each column of sample cells in the sample table;
[0053] Add each sample cell in each row of the sample table to the set of row sample cells, and add each sample cell in each column of the sample table to the set of column sample cells;
[0054] Input the sample cell, the adjacent sample cell, the sample table header, the set of row sample cells, and the set of column sample cells into a neural network to obtain a function and parameters;
[0055] Combine the function and the parameters to obtain a table formula;
[0056] Train the neural network based on the table formula, and stop training when the network training termination condition is met to obtain the pre-trained neural network.
[0057] In an alternative embodiment, the adjacent sample cells include an upper adjacent sample cell, a lower adjacent sample cell, a left adjacent sample cell, and a right adjacent sample cell;
[0058] Determining the adjacent sample cells corresponding to the sample cells in the sample table includes:
[0059] Determine the first preset number of upper adjacent sample cells above the sample cell in the sample table;
[0060] Determine the second preset number of lower adjacent sample cells below the sample cell in the sample table;
[0061] Determine the third preset number of left adjacent sample cells to the left of the sample cell in the sample table;
[0062] Determine the fourth preset number of right adjacent sample cells to the right of the sample cell in the sample table.
[0063] In an alternative embodiment, the neural network includes an encoder, a decoder, an attention network, and a large language model;
[0064] The inputting the sample cell, the adjacent sample cell, the sample table header, the set of row sample cells, and the set of column sample cells into a neural network to obtain a function and parameters includes:
[0065] Use the encoder to encode the sample table header to obtain a sample table header embedding vector;
[0066] Use the encoder to encode all row sample cells in the set of row sample cells to obtain a sample row embedding vector;
[0067] Using the encoder, encode all column sample cells in the set of column sample cells to obtain sample column embedding vectors;
[0068] Input the sample cell, the adjacent sample cell, the sample row embedding vector, and the sample column embedding vector into the large language model to obtain a first sample output result;
[0069] Input the sample table header embedding vector into the attention network to obtain a second sample output result;
[0070] Using the decoder, decode the first sample output result and the second sample output result to obtain a function and parameters.
[0071] In an alternative embodiment, the inputting the sample cell, the adjacent sample cell, the sample row embedding vector, and the sample column embedding vector into the large language model to obtain a first sample output result includes:
[0072] Determine a first sample coordinate corresponding to the sample cell and a second sample coordinate corresponding to the adjacent sample cell;
[0073] Input the first sample coordinate, the second sample coordinate, the sample row embedding vector, and the sample column embedding vector into the large language model to obtain a first sample output result.
[0074] In an alternative embodiment, the determining a first sample coordinate corresponding to the sample cell and a second sample coordinate corresponding to the adjacent sample cell includes:
[0075] Determine a first sample position of the sample cell in the sample table, determine the first sample position as the sample coordinate origin, and determine the sample coordinate origin as the first sample coordinate corresponding to the sample cell;
[0076] Determine a second sample position of the adjacent sample cell in the sample table, and determine the second sample coordinate corresponding to the adjacent sample cell according to the position difference between the first sample position and the second sample position.
[0077] In a second aspect of the embodiments of the present application, there is also provided a table formula processing device, the device includes:
[0078] A table acquisition module, configured to acquire a target table;
[0079] A cell determination module, configured to determine an adjacent cell corresponding to a target cell in the target table;
[0080] A table structure processing module for structurally processing the target table to obtain structured data corresponding to the target table;
[0081] A table formula prediction module for predicting a target table formula based on the target cell, the adjacent cells, and the structured data;
[0082] A table formula processing module for displaying the target table formula at a corresponding position of the target cell.
[0083] In an optional embodiment, the adjacent cells include an upper adjacent cell, a lower adjacent cell, a left adjacent cell, and a right adjacent cell;
[0084] The cell determination module is specifically configured to:
[0085] Determine a first preset number of the upper adjacent cells above the target cell in the target table;
[0086] Determine a second preset number of the lower adjacent cells below the target cell in the target table;
[0087] Determine a third preset number of the left adjacent cells to the left of the target cell in the target table;
[0088] Determine a fourth preset number of the right adjacent cells to the right of the target cell in the target table.
[0089] In an optional embodiment, the table structure processing module is specifically configured to:
[0090] Identify the table header in the target table and split the target table to obtain the table header, each row of cells in the target table, and each column of cells in the target table;
[0091] Add each row of cells in the target table to a row cell set, and add each column of cells in the target table to a column cell set.
[0092] In an optional embodiment, the table formula prediction module specifically includes:
[0093] An input sub-module for inputting the target cell, the adjacent cells, the table header, the row cell set, and the column cell set into a pre-trained neural network to obtain a target function and target parameters;
[0094] A target function and target parameter combination sub-module for combining the target function and the target parameters to obtain a target table formula.
[0095] In an alternative embodiment, the pre-trained neural network includes a pre-trained encoder, a pre-trained decoder, a pre-trained attention network, and a pre-trained large language model;
[0096] The input sub-module specifically includes:
[0097] A header encoding unit for encoding the header using the pre-trained encoder to obtain a header embedding vector;
[0098] A row cell encoding unit for encoding all row cells in the row cell set using the pre-trained encoder to obtain a row embedding vector;
[0099] A column cell encoding unit for encoding all column cells in the column cell set using the pre-trained encoder to obtain a column embedding vector;
[0100] A first input unit for inputting the target cell, the adjacent cell, the row embedding vector, and the column embedding vector into the pre-trained large language model to obtain a first output result;
[0101] A second input unit for inputting the header embedding vector into the pre-trained attention network to obtain a second output result;
[0102] A decoding unit for decoding the first output result and the second output result using the pre-trained decoder to obtain an objective function and objective parameters.
[0103] In an alternative embodiment, the first input unit specifically includes:
[0104] A coordinate determination sub-unit for determining a first coordinate corresponding to the target cell and a second coordinate corresponding to the adjacent cell;
[0105] An input sub-unit for inputting the first coordinate, the second coordinate, the row embedding vector, and the column embedding vector into the pre-trained large language model to obtain a first output result.
[0106] In an alternative embodiment, the coordinate determination sub-unit is specifically configured to:
[0107] Determine a first position of the target cell in the target table, determine the first position as the coordinate origin, and determine the coordinate origin as the first coordinate corresponding to the target cell;
[0108] Determine the second position of the adjacent cell in the target table, and determine the second coordinate corresponding to the adjacent cell according to the position difference between the first position and the second position.
[0109] In an optional implementation manner, the cell determination module is specifically configured to:
[0110] In response to a filling operation for a target cell in the target table, determine the first data filled in the target cell;
[0111] In the case where the first data is preset data, determine the adjacent cell corresponding to the target cell in the target table.
[0112] In an optional implementation manner, after displaying the target table formula at the corresponding position of the target cell, the apparatus further includes:
[0113] A function processing module, configured to, in response to a re-filling operation for a target cell in the target table, determine the second data re-filled in the target cell;
[0114] In the case where the second data represents a function, display the target function at the corresponding position of the target cell.
[0115] In an optional implementation manner, after displaying the target function at the corresponding position of the target cell, the apparatus further includes:
[0116] A parameter processing module, configured to, in response to a re-filling operation for a target cell in the target table, determine the third data re-filled in the target cell;
[0117] In the case where the third data represents a parameter, display the target parameter at the corresponding position of the target cell.
[0118] In an optional implementation manner, the apparatus further includes:
[0119] A filling prompt trigger module, configured to, in response to a preset operation for a target column cell in the target table, detect whether at least two identical table formulas are filled in the cells of the target column cell;
[0120] In the case where at least two identical table formulas are filled in the cells of the target column cell, trigger a table formula filling prompt;
[0121] The table formula filling prompt is used to prompt the remaining cells of the target column cell to fill the table formula.
[0122] In an optional implementation manner, the apparatus further includes:
[0123] A sample table acquisition module, configured to acquire a sample table;
[0124] A sample cell determination module, configured to determine adjacent sample cells corresponding to sample cells in the sample table;
[0125] A sample header recognition module, configured to recognize a sample header in the sample table;
[0126] A table splitting module, configured to split the sample table to obtain the sample header, each row of sample cells in the sample table, and each column of sample cells in the sample table;
[0127] A row sample cell and column sample cell addition module, configured to add each row of sample cells in the sample table to a row sample cell set, and add each column of sample cells in the sample table to a column sample cell set;
[0128] A sample input module, configured to input the sample cells, the adjacent sample cells, the sample header, the row sample cell set, and the column sample cell set into a neural network to obtain a function and parameters;
[0129] A function and parameter combination module, configured to combine the function and the parameters to obtain a table formula;
[0130] A network training module, configured to train the neural network based on the table formula, and stop training when a network training termination condition is met to obtain the pre-trained neural network.
[0131] In an optional implementation manner, the adjacent sample cells include an upper adjacent sample cell, a lower adjacent sample cell, a left adjacent sample cell, and a right adjacent sample cell;
[0132] The sample cell determination module is specifically configured to:
[0133] Determine a first preset number of the upper adjacent sample cells above the sample cells in the sample table;
[0134] Determine a second preset number of the lower adjacent sample cells below the sample cells in the sample table;
[0135] Determine a third preset number of the left adjacent sample cells to the left of the sample cells in the sample table;
[0136] Determine a fourth preset number of the right adjacent sample cells to the right of the sample cells in the sample table.
[0137] In an alternative embodiment, the neural network includes an encoder, a decoder, an attention network, and a large language model;
[0138] The sample input module specifically includes:
[0139] A sample header encoding sub-module, configured to use the encoder to encode the sample header to obtain a sample header embedding vector;
[0140] A row sample cell encoding sub-module, configured to use the encoder to encode all row sample cells in the row sample cell set to obtain a sample row embedding vector;
[0141] A column sample cell encoding sub-module, configured to use the encoder to encode all column sample cells in the column sample cell set to obtain a sample column embedding vector;
[0142] A sample input sub-module, configured to input the sample cell, the adjacent sample cell, the sample row embedding vector, and the sample column embedding vector into the large language model to obtain a first sample output result;
[0143] A vector input sub-module, configured to input the sample header embedding vector into the attention network to obtain a second sample output result;
[0144] A decoding sub-module, configured to use the decoder to decode the first sample output result and the second sample output result to obtain a function and parameters.
[0145] In an alternative embodiment, the sample input sub-module specifically includes:
[0146] A sample coordinate determination unit, configured to determine a first sample coordinate corresponding to the sample cell and a second sample coordinate corresponding to the adjacent sample cell;
[0147] A sample input unit, configured to input the first sample coordinate, the second sample coordinate, the sample row embedding vector, and the sample column embedding vector into the large language model to obtain a first sample output result.
[0148] In an alternative embodiment, the sample coordinate determination unit is specifically configured to:
[0149] Determine a first sample position of the sample cell in the sample table, determine the first sample position as the sample coordinate origin, and determine the sample coordinate origin as the first sample coordinate corresponding to the sample cell;
[0150] Determine the second sample position of the adjacent sample cell in the sample table, and determine the second sample coordinates corresponding to the adjacent sample cell according to the position difference between the first sample position and the second sample position.
[0151] In a third aspect of the embodiments of the present application, an electronic device is further provided, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0152] The memory is used to store a computer program;
[0153] The processor is used to implement the table formula processing method described in any one of the above first aspects when executing the program stored on the memory.
[0154] In a fourth aspect of the embodiments of the present application, a storage medium is further provided. Instructions are stored in the storage medium, and when it runs on a computer, the computer is made to execute the table formula processing method described in any one of the above first aspects.
[0155] In a fifth aspect of the embodiments of the present application, a computer program product containing instructions is further provided. When it runs on a computer, the computer is made to execute the table formula processing method described in any one of the above.
[0156] The technical solution provided by the embodiments of the present application is to obtain a target table, determine adjacent cells corresponding to a target cell in the target table, perform structured processing on the target table to obtain structured data corresponding to the target table, predict a target table formula according to the target cell, the adjacent cells, and the structured data, and display the target table formula at the corresponding position of the target cell.
[0157] By determining adjacent cells corresponding to the target cell and performing structured processing on the target table, a target table formula can be predicted according to the target cell, the adjacent cells, and the structured data, and the target table formula is displayed at the corresponding position of the target cell. In this way, the target table formula can be obtained without the need for the user to input natural language, which can improve the efficiency of table data processing. Moreover, the target table formula predicted according to the target cell, the adjacent cells, and the structured data can reduce the uncertainty and error probability of the table formula and improve the accuracy of table data processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0158] The accompanying drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0159] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0160] One or more embodiments are exemplarily illustrated by the pictures in the corresponding accompanying drawings. These exemplary illustrations do not limit the embodiments. Elements with the same reference numerals in the drawings represent similar elements. Unless otherwise stated, the drawings in the figures do not constitute a proportional limitation.
[0161] Figure 1 It is a schematic diagram of the implementation process of a table formula processing method shown in the embodiments of the present application;
[0162] Figure 2 It is a schematic diagram of the implementation process of another table formula processing method shown in the embodiments of the present application;
[0163] Figure 3 It is a schematic diagram of filling "=" in a target cell shown in the embodiments of the present application;
[0164] Figure 4 It is a schematic diagram of displaying a target table formula in the form of a dropdown list shown in the embodiments of the present application;
[0165] Figure 5 It is a schematic diagram of displaying a target function in the form of a dropdown list shown in the embodiments of the present application;
[0166] Figure 6 It is a schematic diagram of displaying a target parameter in the form of a dropdown list shown in the embodiments of the present application;
[0167] Figure 7 It is a schematic diagram of displaying a table formula filling prompt in the form of a dropdown list shown in the embodiments of the present application;
[0168] Figure 8 It is a schematic diagram of the implementation process of a network training method shown in the embodiments of the present application;
[0169] Figure 9 It is a schematic diagram of the structure of a table formula processing device shown in the embodiments of the present application;
[0170] Figure 10 It is a schematic diagram of the structure of an electronic device shown in the embodiments of the present application. Detailed implementation manners
[0171] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.
[0172] The following disclosure provides many different embodiments or examples for implementing different structures of this application. To simplify the disclosure of this application, components and settings of specific examples are described below. Of course, they are merely examples and are not intended to limit this application. In addition, this application may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplification and clarity and does not itself indicate the relationship between the various embodiments and / or settings discussed.
[0173] It should be noted that for the table formulas involved in the embodiments of this application, in essence, they refer to the function formulas in the table, such as the SUM function, AVERAGE function, VLOOKUP function, etc., which are different from specific calculation formulas. For example, the calculation formula can be y = ax + b, and the table formula is different from this type of calculation formula.
[0174] As Figure 1 shown, it is a schematic flowchart of the implementation process of a table formula processing method provided by an embodiment of this application. This method is applied to an electronic device and may specifically include the following steps:
[0175] S101, obtain a target table and determine adjacent cells corresponding to a target cell in the target table.
[0176] In the embodiments of this application, for a user, a table processing software can be used to open a document (such as a table document) to obtain the target table in the document. At this time, the user can perform data filling on the target cell in the target table;
[0177] When data filling (such as filling "=") is performed on the target cell in the target table, the recommendation of the table formula will be triggered to improve the efficiency and accuracy of table data processing. For this purpose, it is necessary to determine adjacent cells corresponding to the target cell in the target table. Among them, the adjacent cells can be directly adjacent or indirectly adjacent to the target cell, which means that the adjacent cells can be cells at a certain distance from the target cell.
[0178] It should be noted that for the adjacent cells corresponding to the target cell, it generally refers to the adjacent cells in the four directions of up, down, left, and right of the target cell, and the adjacent cells include directly adjacent cells and indirectly adjacent cells. The embodiments of the present application do not limit this.
[0179] For example, assume that the target cell in the target table is E11. Then the adjacent cells corresponding to the target cell can be E10 and E09. For the adjacent cells E10 and E09, they are the adjacent cells above the target cell E11, and the adjacent cell E10 is directly adjacent to the target cell E11, while the adjacent cell E09 is indirectly adjacent to the target cell E11. The same applies to the adjacent cells in the other directions.
[0180] S102, perform structured processing on the target table to obtain the structured data corresponding to the target table.
[0181] In the embodiments of the present application, for the target table, structured processing can be performed on the target table, and the structured data corresponding to the target table can be obtained. Performing structured processing on the target table here is essentially to convert the target table into data in a regular format.
[0182] It should be noted that for structured data, it generally refers to three parts: the table header, each row of cells in the target table (the table header may not be included), and each column of cells in the target table. The embodiments of the present application do not limit this.
[0183] For each column of cells in the target table, they exist in the form of key-value pairs. For example, for a certain column of cells, its key refers to the column title of the column of cells, and the value refers to the column content of the column of cells.
[0184] S103, predict the target table formula according to the target cell, adjacent cells, and structured data.
[0185] In the embodiments of the present application, for the target cell in the target table, the adjacent cells corresponding to the target cell, and the structured data obtained by performing structured processing on the target table, the target table formula can be predicted according to the target cell, adjacent cells, and structured data.
[0186] In this way, through the target cell in the target table, the adjacent cells corresponding to the target cell, and the structured data of the target table, the target table formula that may be used in the target cell can be inferred intelligently.
[0187] S104, display the target table formula at the corresponding position of the target cell.
[0188] In the embodiments of the present application, for the predicted target table formula, the target table formula can be displayed at the corresponding position of the target cell to recommend the target table formula, so as to prompt the user to select the target table formula to fill into the target cell. Among them, the corresponding position of the target cell can be, for example, below the target cell, or it can be a pop-up window that uniformly displays the table formulas to be recommended. The embodiments of the present application do not limit this.
[0189] It should be noted that for the target table formula, it is predicted through the target cell in the target table, the adjacent cell corresponding to the target cell, and the structured data of the target table, which can reduce the uncertainty and error probability of the table formula and improve the accuracy of table data processing.
[0190] In addition, the target table formula that may be used in the target cell can be inferred intelligently, and the target table formula can be obtained without the user inputting the requirements in natural language, which can improve the efficiency of table data processing.
[0191] Through the description of the technical solution provided by the embodiments of the present application above, obtain the target table, determine the adjacent cell corresponding to the target cell in the target table, perform structured processing on the target table to obtain the structured data corresponding to the target table, predict the target table formula according to the target cell, the adjacent cell, and the structured data, and display the target table formula at the corresponding position of the target cell.
[0192] By determining the adjacent cell corresponding to the target cell and performing structured processing on the target table, thereby predicting the target table formula according to the target cell, the adjacent cell, and the structured data, and displaying the target table formula at the corresponding position of the target cell. In this way, the target table formula can be obtained without the user inputting the requirements in natural language, which can improve the efficiency of table data processing. Moreover, according to the target cell, the adjacent cell, and the structured data, the predicted target table formula can reduce the uncertainty and error probability of the table formula and improve the accuracy of table data processing.
[0193] As Figure 2 shown, it is a schematic flowchart of the implementation process of another table formula processing method provided by the embodiments of the present application. This method is applied to an electronic device and specifically may include the following steps:
[0194] S201, obtain the target table, and in response to the filling operation for the target cell in the target table, determine the first data filled in the target cell.
[0195] In an embodiment of the present application, for a user, a document (such as a table document) can be opened using a table processing software, so as to obtain a target table in the document. For the data filled in the cells of the target table, the user may process it using a table formula.
[0196] Based on this, the user can perform a filling operation on a target cell (usually referring to a blank cell) in the target table, so as to determine the first data filled in the target cell in response to the filling operation on the target cell in the target table.
[0197] S202, when the first data is preset data, determine the adjacent cells corresponding to the target cell in the target table.
[0198] In an embodiment of the present application, for the first data filled in the target cell, when the first data is preset data (for example, the preset data can be the character "=", or can be text such as "equal" and other data related to the use of table formulas), it indicates that the user hopes to process the data already filled in the cells of the target table using a table formula. At this time, the recommendation of the table formula is triggered to improve the efficiency and accuracy of table data processing. For this purpose, it is necessary to determine the adjacent cells corresponding to the target cell in the target table.
[0199] Among them, the adjacent cells include the upper adjacent cell, the lower adjacent cell, the left adjacent cell, and the right adjacent cell. Thus, determine the first preset number of upper adjacent cells above the target cell in the target table, determine the second preset number of lower adjacent cells below the target cell in the target table, determine the third preset number of left adjacent cells to the left of the target cell in the target table, and determine the fourth preset number of right adjacent cells to the right of the target cell in the target table. Among them, the first preset number, the second preset number, the third preset number, and the fourth preset number can be different from each other or can be the same.
[0200] For example, for the first data filled in the target cell G11, when the first data is "=", as Figure 3 shown, it indicates that the user hopes to process the data already filled in the cells of the target table using a table formula. At this time, the recommendation of the table formula is triggered to improve the efficiency and accuracy of table data processing. For this purpose, determine 10 upper adjacent cells above the target cell G11 in the target table, determine 10 lower adjacent cells below the target cell G11 in the target table, determine 10 left adjacent cells to the left of the target cell G11 in the target table, and determine 10 right adjacent cells to the right of the target cell G11 in the target table.
[0201] It should be noted that for the 10 upper adjacent cells, it refers to the upper adjacent cells G01 - G10. G10 is directly adjacent to the target cell G11, and the rest are indirectly adjacent to the target cell G11. For the 10 lower adjacent cells, it refers to the lower adjacent cells G12 - G21. G12 is directly adjacent to the target cell G11, and the rest are indirectly adjacent to the target cell G11. For the 10 left adjacent cells, it refers to the left adjacent cells A11, B11, C11, D11, E11, F11, etc. and 4 empty left adjacent cells. F11 is directly adjacent to the target cell G11, and the rest are indirectly adjacent to the target cell G11. For the 10 right adjacent cells, it refers to H11, I11, …… Q11. H11 is directly adjacent to the target cell G11, and the rest are indirectly adjacent to the target cell G11.
[0202] S203. Identify the table header in the target table, and split the target table to obtain the table header, each row cell in the target table, and each column cell in the target table.
[0203] S204. Add each row cell in the target table to the row cell set, and add each column cell in the target table to the column cell set.
[0204] In the embodiment of the present application, when the first data is preset data, on the one hand, determine the adjacent cells corresponding to the target cell in the target table as described above. On the other hand, it is necessary to perform structured processing on the target table to obtain the structured data corresponding to the target table. Among them, performing structured processing on the target table means splitting the target table.
[0205] Based on this, identifying the table header in the target table, splitting the target table to obtain the table header, each row cell in the target table, and each column cell in the target table, adding each row cell in the target table to the row cell set, and adding each column cell in the target table to the column cell set means splitting the target table into a table header (such as the first row cell shown in Figure 3 ), each row cell (such as the 2nd - 10th row cells shown in Figure 3 ), each column cell (such as the A - G column cells shown in Figure 3 ), etc. into three parts.
[0206] S205. Input the target cell, adjacent cells, table header, row cell set, and column cell set into a pre - trained neural network to obtain the target function and target parameters.
[0207] S206. Combine the target function and target parameters to obtain the target table formula.
[0208] In an embodiment of the present application, a pre-trained neural network, such as a Transformer, is provided. The target cell, adjacent cells, table header, set of row cells, and set of column cells in the target table can be input into the pre-trained neural network to obtain a target function and target parameters. The target function and target parameters are combined to obtain a target table formula.
[0209] Among them, for the pre-trained neural network, it includes a pre-trained encoder, a pre-trained decoder, a pre-trained attention network (such as attention), and a pre-trained large language model (such as LLM). The pre-trained encoder can compile the rows of cells and columns of cells into content that the pre-trained large language model can understand, and the pre-trained encoder can compile the table header into content that the pre-trained attention network can understand.
[0210] Based on this, using the pre-trained encoder, the table header is encoded to obtain a table header embedding vector. Using the pre-trained encoder, all the row cells in the set of row cells are encoded to obtain a row embedding vector. Using the pre-trained encoder, all the column cells in the set of column cells are encoded to obtain a column embedding vector. The target cell, adjacent cells, row embedding vector, and column embedding vector are input into the pre-trained large language model to obtain a first output result. The table header embedding vector is input into the pre-trained attention network to obtain a second output result. Then, using the pre-trained decoder, the first output result and the second output result are decoded to obtain the target function and target parameters. Then, the target function and target parameters are combined to obtain the target table formula. Among them, the target function and target parameters can exist in groups, which means that there can be multiple groups of target functions and target parameters.
[0211] Among them, for the pre-trained large language model (such as GPT-3), the first output result it outputs is a high-dimensional vector, usually containing thousands of elements. This high-dimensional vector is an abstract representation of the input to the pre-trained large language model, capturing the context and semantic information of the language. The first output result can be obtained based on the input to the pre-trained large language model because the pre-trained large language model has been pre-trained on a large amount of data and has learned the associations of vocabulary, grammar, and semantics. In a specific application, the pre-trained large language model maps the input to this high-dimensional vector, enabling the pre-trained large language model to better understand the semantics and relationships of the data.
[0212] For the pre-trained attention network, the second output result is a set of weights that represent the importance of different positions that the pre-trained attention network focuses on in the input sequence. In the table scenario, the pre-trained attention network may be used to focus on information near the table header, rows, columns, or target cells. The reason for obtaining this second output result based on the input of the pre-trained attention network is that the attention mechanism allows the pre-trained attention network to dynamically adjust the focus of attention, so as to better capture the important parts in the input and provide more targeted information for subsequent tasks.
[0213] For the first output result and the second output result, these two output results provide information at different levels. The first output result provides an understanding of the overall semantics and context, an abstract representation of the entire target table. The second output result provides the focus on specific parts, and can process information more specifically locally, such as the target cell and adjacent cells. By combining these two output results, the pre-trained neural network can simultaneously consider the overall semantic information and local focus points, so as to understand the input more comprehensively. In specific applications, these two output results can be multiplied or concatenated to generate a comprehensive representation that captures the relationship between the overall and layout information. This comprehensive representation can be input into the subsequent pre-trained decoder for predicting the target function and target parameters.
[0214] Generally speaking, the first output result provides a high-level understanding of the entire target table, while the second output result guides the pre-trained neural network to focus on task-related local information. After the two output results are combined, the pre-trained neural network can better understand the semantic relationships in the target table and can more accurately predict the target function and target parameters. The combination of these two output results provides multiple levels of consideration for the input, thereby enhancing the expressive ability and generalization ability of the pre-trained neural network.
[0215] It should be noted that for the table header embedding vector, it refers to a highly abstract form of data information, including data types and potential relationships between data. The row embedding vector contains all the information of a single individual, which can better provide a clear understanding of the individual and identify the semantic information therein. The column embedding vector contains column formats, correlation information between columns, etc. For example, the trend of salary change over time can be obtained through the correlation information between columns. In addition, inputting the table header embedding vector alone into the pre-trained attention network can better understand the relationships and dependencies between table headers.
[0216] In addition, there are two major difficulties in predicting table formulas. One is function prediction, and the other is parameter prediction. To solve the problem of parameter prediction, it is necessary to represent the target cell and its corresponding adjacent cells in the target table in the form of coordinates. To this end, the first coordinate corresponding to the target cell and the second coordinate corresponding to the adjacent cell are determined, and the first coordinate, the second coordinate, the row embedding vector, and the column embedding vector are input into the pre-trained large language model to obtain the first output result. Among them, regarding the internal processing of the pre-trained large language model, the processing process of existing large language models can be referred to.
[0217] Among them, assuming that the target cell is E11 and taking this as the coordinate origin, for the cell D10, compared with the target cell E11, it is shifted one cell to the left and one cell up to obtain the cell D10, and its coordinates can be expressed as (R[-1], C[1]). Based on this, the first position of the target cell in the target table is determined, the first position is determined as the coordinate origin, the coordinate origin is determined as the first coordinate corresponding to the target cell, the second position of the adjacent cell in the target table is determined, and according to the position difference between the first position and the second position, the second coordinate corresponding to the adjacent cell is determined.
[0218] S207, display the target table formula at the corresponding position of the target cell.
[0219] In the embodiments of the present application, for the target table formula, which includes the above-mentioned target function and target parameters, it can be displayed at the corresponding position of the target cell. For example, the target table formula can be displayed in the form of a drop-down list below the target cell, as Figure 4 shown.
[0220] In addition, for the target table formula, the target table formula can be executed to obtain the formula result, which can be displayed together with the target table formula at the corresponding position of the target cell for the user to select. The user can click on the option to fill in the recommended target table formula, or directly use the tab shortcut key to fill in the target table formula at one key, as Figure 4 shown.
[0221] For example, after the target table formula is displayed in the form of a drop-down list below the target cell, it is possible that the user does not select the recommended target table formula. At this time, the user can perform a filling operation on the target cell in the target table again. Therefore, in response to the repeated filling operation on the target cell in the target table, the second data filled in the target cell again is determined. When the second data represents a function, it means that the user manually enters the function. At this time, the target function can be displayed at the corresponding position of the target cell to recommend the target function. For example, the target function can be displayed in the form of a drop-down list below the target cell.
[0222] For example, the user can perform a filling operation on the target cell in the target table again. In response to the repeated filling operation on the target cell in the target table, the second data "V" filled again in the target cell is determined. This second data "V" represents a function, which means that the second data filled again by the user is data related to the function. At this time, the target function can be displayed below the target cell in the form of a drop-down list to recommend the target function, such as Figure 5 shown.
[0223] In addition, for example, after the target function is displayed below the target cell in the form of a drop-down list, it is possible that the user does not select the recommended target function. At this time, the user can perform a filling operation on the target cell in the target table again. In response to the repeated filling operation on the target cell in the target table, the third data filled again in the target cell can be determined. When the third data represents a parameter, it means that the user manually enters the parameter. At this time, the target parameter can be displayed at the corresponding position of the target cell to recommend the target parameter. For example, the target parameter can be displayed below the target cell in the form of a drop-down list.
[0224] For example, the user can perform a filling operation on the target cell in the target table again. In response to the repeated filling operation on the target cell in the target table, the third data "H" filled again in the target cell is determined. The third data "H" represents a parameter, which means that the third data filled again by the user is data related to the parameter. At this time, the target parameter can be displayed below the target cell in the form of a drop-down list to recommend the target parameter, such as Figure 6 shown.
[0225] In this way, through the target cell in the target table, the adjacent cell corresponding to the target cell, and the structured data of the target table, the target table formula that may be used in the target cell can be inferred intelligently, which can reduce the uncertainty and error probability of the table formula and improve the accuracy and efficiency of table data processing. In addition, when the user fills in data related to the function, the target function can be recommended, and when the user fills in data related to the parameter, the target parameter can be recommended, which can further improve the accuracy and efficiency of table data processing.
[0226] In addition, in the embodiments of the present application, in addition to the above-mentioned recommendation of the table formula for the target cell, a table formula recommendation function for column cells is also provided, which means that for column cells, if at least two identical table formulas are filled in the cells, the user can be recommended to fill in the table formula in the remaining cells of the column cell.
[0227] Based on this, in response to the target column cell in the target table (for example, as Figure 7The preset operations (such as the click operation on the target column cell or the selection operation on the cell with filled data in the target column cell) of the cells in column G shown in the figure are used to detect whether at least two identical table formulas are filled in the cells of the target column cell. In the case where at least two identical table formulas are filled in the cells of the target column cell, a table formula filling prompt is triggered. The table formula filling prompt is used to prompt the remaining cells of the target column cell to fill in the table formula. Among them, the table formula filling prompt is displayed in the form of a drop-down list, as Figure 7 shown.
[0228] In addition, the embodiment of the present application also provides a network training method, as Figure 8 shown, which is a schematic flowchart of the implementation process of a network training method provided by the embodiment of the present application. This method is applied to an electronic device and specifically may include the following steps:
[0229] S801, Obtain a sample table and determine the adjacent sample cells corresponding to the sample cells in the sample table.
[0230] In the embodiment of the present application, a sample table is obtained and the adjacent sample cells corresponding to the sample cells in the sample table are determined. Among them, for the adjacent sample cells, it includes the upper adjacent sample cells, the lower adjacent sample cells, the left adjacent sample cells, and the right adjacent sample cells.
[0231] Based on this, determine the first preset number of upper adjacent sample cells above the sample cell in the sample table, determine the second preset number of lower adjacent sample cells below the sample cell in the sample table, determine the third preset number of left adjacent sample cells to the left of the sample cell in the sample table, and determine the fourth preset number of right adjacent sample cells to the right of the sample cell in the sample table.
[0232] For example, determine 10 upper adjacent sample cells above the sample cell in the sample table, determine 10 lower adjacent sample cells below the sample cell in the sample table, determine 10 left adjacent sample cells to the left of the sample cell in the sample table, and determine 10 right adjacent sample cells to the right of the sample cell in the sample table.
[0233] S802, Identify the sample table headers in the sample table and split the sample table to obtain the sample table headers, each row of sample cells in the sample table, and each column of sample cells in the sample table.
[0234] S803, Add each row of sample cells in the sample table to the row sample cell set, and add each column of sample cells in the sample table to the column sample cell set.
[0235] In the embodiments of the present application, for a sample table, it is necessary to perform structured processing on the sample table to obtain the sample structured data corresponding to the sample table. Among them, performing structured processing on the sample table refers to splitting the sample table.
[0236] Based on this, identify the sample table header in the sample table, and split the sample table to obtain the sample table header, each row sample cell in the sample table, and each column sample cell in the sample table. Add each row sample cell in the sample table to the row sample cell set, and add each column sample cell in the sample table to the column sample cell set.
[0237] S804, input the sample cell, adjacent sample cell, sample table header, row sample cell set, and column sample cell set into the neural network to obtain a function and parameters.
[0238] S805, combine the function and parameters to obtain a table formula.
[0239] In the embodiments of the present application, a neural network is provided, such as a transformer, which can input the sample cell, the adjacent sample cell corresponding to the sample cell, the sample table header, the row sample cell set, and the column sample cell set in the sample table into the neural network to obtain a function and parameters, and combine the function and parameters to obtain a table formula.
[0240] Among them, for the neural network, it includes an encoder, a decoder, an attention network (such as attention), and a large language model (such as LLM). The encoder can be used to compile the row sample cells and column sample cells into content that the large language model can understand, and compile the sample table header into content that the attention network can understand.
[0241] Based on this, use the encoder to encode the sample table header to obtain a sample table header embedding vector, use the encoder to encode all the row sample cells in the row sample cell set to obtain a sample row embedding vector, use the encoder to encode all the column sample cells in the column sample cell set to obtain a sample column embedding vector, input the sample cell, adjacent sample cell, sample row embedding vector, and sample column embedding vector into the large language model to obtain a first sample output result, input the sample table header embedding vector into the attention network to obtain a second sample output result, and use the decoder to decode the first sample output result and the second sample output result to obtain a function and parameters. Among them, the function and parameters can exist in the form of groups, which means that there can be multiple groups of functions and parameters.
[0242] Among them, for large language models (such as GPT-3), the first sample output result is a high-dimensional vector, usually containing thousands of elements. This high-dimensional vector is an abstract representation of the input to the large language model, capturing the context and semantic information of the language. The reason for obtaining this first sample output result based on the input to the large language model is that the large language model has been pre-trained on a large amount of data and has learned the associations of vocabulary, grammar, and semantics. In specific applications, the large language model maps the input to this high-dimensional vector, enabling the large language model to better understand the semantics and relationships of the data.
[0243] For the second sample output result of the attention network, it is a set of weights, and these weights represent the importance of different positions that the attention network focuses on in the input sequence. In the table scenario, the attention network may be used to focus on information near the sample table header, row samples, column samples, or sample cells. The reason for obtaining this second sample output result based on the input to the attention network is that the attention mechanism allows the attention network to dynamically adjust the focus of attention, thereby better capturing the important parts in the input and providing more targeted information for subsequent tasks.
[0244] Regarding the first sample output result and the second sample output result, these two sample output results provide information at different levels. The first sample output result provides an understanding of the overall semantics and context, an abstract representation of the entire sample table. The second sample output result provides attention to specific parts and can process information more targeted locally, such as sample cells and adjacent sample cells. By combining these two sample output results, the neural network can simultaneously consider the overall semantic information and local focus points, thereby understanding the input more comprehensively. In specific applications, these two sample output results can be multiplied or concatenated to generate a comprehensive representation that captures the relationship between the overall and layout information. This comprehensive representation can be input into the subsequent decoder for predicting functions and parameters.
[0245] Generally speaking, the first sample output result provides a high-level understanding of the entire sample table, while the second sample output result guides the neural network to focus on local information related to the task. After combining the two sample output results, the neural network can better understand the semantic relationships in the sample table and can more accurately predict functions and parameters. The combination of these two sample output results provides multiple levels of consideration for the input, thereby enhancing the expressive ability and generalization ability of the neural network.
[0246] It should be noted that the sample table header embedding vector refers to a highly abstract representation of data information, including data types and potential relationships between data. The sample row embedding vector contains all the information of a single individual, enabling a better understanding of the individual and recognition of the semantic information therein. The sample column embedding vector contains column formats, correlation information between columns, etc. For example, the trend of salary changes over time can be obtained through the correlation information between columns. In addition, inputting the sample table header embedding vector alone into the attention network for training can better understand the relationships and dependencies between table headers.
[0247] In addition, there are two major difficulties in table formula prediction. One is function prediction, and the other is parameter prediction. To solve the problem of parameter prediction, the sample cells and adjacent sample cells corresponding to the sample cells in the sample table need to be represented in the form of coordinates. Therefore, the first sample coordinates corresponding to the sample cells and the second sample coordinates corresponding to the adjacent sample cells are determined, and the first sample coordinates, the second sample coordinates, the sample row embedding vector, and the sample column embedding vector are input into the large language model to obtain the first sample output result. Among them, regarding the internal processing of the large language model, the processing process of existing large language models can be referred to.
[0248] Among them, assuming that the sample cell is E11 and taking this as the coordinate origin, for the sample cell D10, compared with the sample cell E11, it can be obtained by shifting one cell to the left and one cell up. Its coordinates can be expressed as (R[-1], C[1]). Based on this, the first sample position of the sample cell in the sample table is determined, the first sample position is determined as the sample coordinate origin, the sample coordinate origin is determined as the first sample coordinates corresponding to the sample cell, the second sample position of the adjacent sample cell in the sample table is determined, and the second sample coordinates corresponding to the adjacent sample cell are determined according to the position difference between the first sample position and the second sample position.
[0249] S806, training the neural network based on the table formula and stopping the training when the network training termination condition is met to obtain the pre-trained neural network.
[0250] In the embodiments of the present application, for the table formula, which includes the above functions and parameters, the neural network can be trained based on the table formula, and the training is stopped when the network training termination condition is met to obtain the pre-trained neural network.
[0251] Among them, the neural network includes an encoder, a decoder, an attention network (such as attention) and a large language model (such as LLM), and the encoder, decoder, attention network and large language model can be trained based on the above table formula.
[0252] In addition, for the network training termination condition, for example, it can be that the number of training times reaches a preset threshold, or the accuracy rate or recall rate of the table formula reaches a preset threshold. The embodiments of the present application do not limit this.
[0253] After the above training, a pre-trained neural network can be obtained. Thus, through the target cell in the target table, the adjacent cells corresponding to the target cell, and the structured data of the target table, the pre-trained neural network can intelligently infer the target table formula that may be used in the target cell, which can reduce the uncertainty and error probability of the table formula and improve the accuracy and efficiency of table data processing. In addition, in the case where the user fills in data related to a function, the target function can be recommended, and in the case where the user fills in data related to a parameter, the target parameter can be recommended, which can further improve the accuracy and efficiency of table data processing.
[0254] Corresponding to the above method embodiment, the embodiment of the present application also provides a table formula processing device, as Figure 9 shown. The device may include: a table acquisition module 910, a cell determination module 920, a table structure processing module 930, a table formula prediction module 940, and a table formula processing module 950.
[0255] The table acquisition module 910 is configured to acquire a target table;
[0256] The cell determination module 920 is configured to determine the adjacent cells corresponding to the target cell in the target table;
[0257] The table structure processing module 930 is configured to perform structure processing on the target table to obtain the structured data corresponding to the target table;
[0258] The table formula prediction module 940 is configured to predict a target table formula according to the target cell, the adjacent cells, and the structured data;
[0259] The table formula processing module 950 is configured to display the target table formula at the corresponding position of the target cell.
[0260] In an optional embodiment, the adjacent cells include the upper adjacent cell, the lower adjacent cell, the left adjacent cell, and the right adjacent cell;
[0261] The cell determination module is specifically configured to:
[0262] Determine the first preset number of the upper adjacent cells above the target cell in the target table;
[0263] Determine the second preset number of the lower adjacent cells below the target cell in the target table;
[0264] Determine the third preset number of the left adjacent cells to the left of the target cell in the target table;
[0265] Determine the fourth preset number of the right adjacent cells to the right of the target cell in the target table.
[0266] In an alternative embodiment, the table structuring processing module is specifically configured to:
[0267] Identify the table header in the target table and split the target table to obtain the table header, each row of cells in the target table, and each column of cells in the target table;
[0268] Add each row of cells in the target table to the row cell set, and add each column of cells in the target table to the column cell set.
[0269] In an alternative embodiment, the table formula prediction module specifically includes:
[0270] An input sub-module, configured to input the target cell, the adjacent cells, the table header, the row cell set, and the column cell set into a pre-trained neural network to obtain a target function and target parameters;
[0271] A target function and target parameter combination module, configured to combine the target function and the target parameters to obtain a target table formula.
[0272] In an alternative embodiment, the pre-trained neural network includes a pre-trained encoder, a pre-trained decoder, a pre-trained attention network, and a pre-trained large language model;
[0273] The input sub-module specifically includes:
[0274] A table header encoding unit, configured to encode the table header using the pre-trained encoder to obtain a table header embedding vector;
[0275] A row cell encoding unit, configured to encode all row cells in the row cell set using the pre-trained encoder to obtain a row embedding vector;
[0276] A column cell encoding unit, configured to encode all column cells in the column cell set using the pre-trained encoder to obtain a column embedding vector;
[0277] A first input unit, configured to input the target cell, the adjacent cell, the row embedding vector, and the column embedding vector into the pre-trained large language model to obtain a first output result;
[0278] A second input unit, configured to input the header embedding vector into the pre-trained attention network to obtain a second output result;
[0279] A decoding unit, configured to use the pre-trained decoder to decode the first output result and the second output result to obtain an objective function and objective parameters.
[0280] In an optional embodiment, the first input unit specifically includes:
[0281] A coordinate determination subunit, configured to determine a first coordinate corresponding to the target cell and a second coordinate corresponding to the adjacent cell;
[0282] An input subunit, configured to input the first coordinate, the second coordinate, the row embedding vector, and the column embedding vector into the pre-trained large language model to obtain a first output result.
[0283] In an optional embodiment, the coordinate determination subunit is specifically configured to:
[0284] Determine a first position of the target cell in the target table, determine the first position as the coordinate origin, and determine the coordinate origin as the first coordinate corresponding to the target cell;
[0285] Determine a second position of the adjacent cell in the target table, and determine the second coordinate corresponding to the adjacent cell according to the position difference between the first position and the second position.
[0286] In an optional embodiment, the cell determination module is specifically configured to:
[0287] In response to a filling operation for the target cell in the target table, determine a first data filled in the target cell;
[0288] In the case where the first data is preset data, determine an adjacent cell corresponding to the target cell in the target table.
[0289] In an optional embodiment, after displaying the target table formula at the corresponding position of the target cell, the apparatus further includes:
[0290] A function processing module, configured to, in response to a re-filling operation for the target cell in the target table, determine a second data re-filled in the target cell;
[0291] In the case of the second data representation function, the target function is displayed at the corresponding position of the target cell.
[0292] In an optional implementation, after the target function is displayed at the corresponding position of the target cell, the device further includes:
[0293] A parameter processing module, configured to determine third data refilled in the target cell in response to a refill operation for the target cell in the target table;
[0294] In the case where the third data represents a parameter, the target parameter is displayed at the corresponding position of the target cell.
[0295] In an optional implementation, the device further includes:
[0296] A fill prompt trigger module, configured to detect whether at least two identical table formulas are filled in the cells of the target column cell in response to a preset operation for the target column cell in the target table;
[0297] In the case where at least two identical table formulas are filled in the cells of the target column cell, trigger a table formula fill prompt;
[0298] The table formula fill prompt is used to prompt the remaining cells of the target column cell to fill the table formula.
[0299] In an optional implementation, the device further includes:
[0300] A sample table acquisition module, configured to acquire a sample table;
[0301] A sample cell determination module, configured to determine adjacent sample cells corresponding to the sample cells in the sample table;
[0302] A table header recognition module, configured to recognize the sample table header in the sample table;
[0303] A table splitting module, configured to split the sample table to obtain the sample table header, each row of sample cells in the sample table, and each column of sample cells in the sample table;
[0304] A row sample cell and column sample cell addition module, configured to add each row of sample cells in the sample table to a row sample cell set, and add each column of sample cells in the sample table to a column sample cell set;
[0305] A sample input module for inputting the sample cell, the adjacent sample cell, the sample table header, the set of row sample cells, and the set of column sample cells into a neural network to obtain a function and parameters;
[0306] A function and parameter combination module for combining the function and the parameters to obtain a table formula;
[0307] A network training module for training the neural network based on the table formula and stopping the training when the network training termination condition is satisfied to obtain the pre-trained neural network.
[0308] In an optional embodiment, the adjacent sample cells include upper adjacent sample cells, lower adjacent sample cells, left adjacent sample cells, and right adjacent sample cells;
[0309] The sample cell determination module is specifically configured to:
[0310] Determine the first preset number of the upper adjacent sample cells above the sample cell in the sample table;
[0311] Determine the second preset number of the lower adjacent sample cells below the sample cell in the sample table;
[0312] Determine the third preset number of the left adjacent sample cells to the left of the sample cell in the sample table;
[0313] Determine the fourth preset number of the right adjacent sample cells to the right of the sample cell in the sample table.
[0314] In an optional embodiment, the neural network includes an encoder, a decoder, an attention network, and a large language model;
[0315] The sample input module specifically includes:
[0316] A sample table header encoding sub-module for encoding the sample table header using the encoder to obtain a sample table header embedding vector;
[0317] A row sample cell encoding sub-module for encoding all row sample cells in the set of row sample cells using the encoder to obtain a sample row embedding vector;
[0318] A column sample cell encoding sub-module for encoding all column sample cells in the set of column sample cells using the encoder to obtain a sample column embedding vector;
[0319] A sample input sub-module for inputting the sample cell, the adjacent sample cell, the sample row embedding vector, and the sample column embedding vector into the large language model to obtain a first sample output result;
[0320] A vector input sub-module for inputting the sample table header embedding vector into the attention network to obtain a second sample output result;
[0321] A decoding sub-module for using the decoder to decode the first sample output result and the second sample output result to obtain a function and parameters.
[0322] In an optional implementation manner, the sample input sub-module specifically includes:
[0323] A sample coordinate determination unit for determining a first sample coordinate corresponding to the sample cell and a second sample coordinate corresponding to the adjacent sample cell;
[0324] A sample input unit for inputting the first sample coordinate, the second sample coordinate, the sample row embedding vector, and the sample column embedding vector into the large language model to obtain a first sample output result.
[0325] In an optional implementation manner, the sample coordinate determination unit specifically is used for:
[0326] Determining a first sample position of the sample cell in the sample table, determining the first sample position as the sample coordinate origin, and determining the sample coordinate origin as the first sample coordinate corresponding to the sample cell;
[0327] Determining a second sample position of the adjacent sample cell in the sample table, and determining the second sample coordinate corresponding to the adjacent sample cell according to the position difference between the first sample position and the second sample position.
[0328] An embodiment of the present application further provides an electronic device, as Figure 10 shown, including a processor 101, a communication interface 102, a memory 103, and a communication bus 104, wherein the processor 101, the communication interface 102, and the memory 103 complete communication with each other through the communication bus 104,
[0329] The memory 103 is used for storing a computer program;
[0330] The processor 101, when executing the program stored on the memory 103, implements the following steps:
[0331] Obtain a target table and determine the adjacent cells corresponding to the target cell in the target table; perform structured processing on the target table to obtain the structured data corresponding to the target table; predict a target table formula based on the target cell, the adjacent cells, and the structured data; and display the target table formula at the corresponding position of the target cell.
[0332] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used in the figure to represent it, but it does not mean that there is only one bus or one type of bus.
[0333] The communication interface is used for communication between the above electronic device and other devices.
[0334] The memory may include a Random Access Memory (RAM), or may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0335] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0336] In another embodiment provided by this application, a storage medium is also provided. Instructions are stored in this storage medium, and when it runs on a computer, it causes the computer to execute any one of the table formula processing methods in the above embodiments.
[0337] In another embodiment provided by the present application, a computer program product including instructions is further provided. When it runs on a computer, it causes the computer to execute the table formula processing method described in any one of the above embodiments.
[0338] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a storage medium or transmitted from one storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0339] It should be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0340] Each embodiment in this specification is described in a related manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment.
[0341] The above are only the preferred embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application are all included in the protection scope of the present application.
Claims
1. A method for processing table formulas, characterized in that, The method includes: Obtain a target table and determine adjacent cells corresponding to a target cell in the target table; Perform structured processing on the target table to obtain structured data corresponding to the target table; Predict a target table formula based on the target cell, the adjacent cells, and the structured data; Display the target table formula at a corresponding position of the target cell.
2. The method according to claim 1, characterized in that, The adjacent cells include an upper adjacent cell, a lower adjacent cell, a left adjacent cell, and a right adjacent cell; The determining of adjacent cells corresponding to a target cell in the target table includes: Determine a first preset number of the upper adjacent cells above the target cell in the target table; Determine a second preset number of the lower adjacent cells below the target cell in the target table; Determine a third preset number of the left adjacent cells to the left of the target cell in the target table; Determine a fourth preset number of the right adjacent cells to the right of the target cell in the target table.
3. The method according to claim 1, wherein The performing of structured processing on the target table to obtain structured data corresponding to the target table includes: Identify the header of the target table and split the target table to obtain the header, each row cell in the target table, and each column cell in the target table; Add each row cell in the target table to a row cell set and add each column cell in the target table to a column cell set.
4. The method according to claim 3, characterized in that, The predicting of a target table formula based on the target cell, the adjacent cells, and the structured data includes: Input the target cell, the adjacent cells, the header, the row cell set, and the column cell set into a pre-trained neural network to obtain a target function and target parameters; Combine the target function and the target parameters to obtain a target table formula.
5. The method according to claim 4, wherein The pre-trained neural network includes a pre-trained encoder, a pre-trained decoder, a pre-trained attention network, and a pre-trained large language model; The inputting of the target cell, the adjacent cells, the header, the row cell set, and the column cell set into a pre-trained neural network to obtain a target function and target parameters includes: Use the pre-trained encoder to encode the header to obtain a header embedding vector; Use the pre-trained encoder to encode all row cells in the row cell set to obtain a row embedding vector; Use the pre-trained encoder to encode all column cells in the column cell set to obtain a column embedding vector; Input the target cell, the adjacent cells, the row embedding vector, and the column embedding vector into the pre-trained large language model to obtain a first output result; Input the header embedding vector into the pre-trained attention network to obtain a second output result; Use the pre-trained decoder to decode the first output result and the second output result to obtain a target function and target parameters.
6. The method according to claim 5, wherein Inputting the target cell, the adjacent cell, the row embedding vector, and the column embedding vector into the pre-trained large language model to obtain a first output result includes: Determining a first coordinate corresponding to the target cell and a second coordinate corresponding to the adjacent cell; Inputting the first coordinate, the second coordinate, the row embedding vector, and the column embedding vector into the pre-trained large language model to obtain a first output result.
7. The method according to claim 6, characterized in that, The determining the first coordinate corresponding to the target cell and the second coordinate corresponding to the adjacent cell includes: Determining a first position of the target cell in the target table, determining the first position as the coordinate origin, and determining the coordinate origin as the first coordinate corresponding to the target cell; Determining a second position of the adjacent cell in the target table, and determining the second coordinate corresponding to the adjacent cell according to the position difference between the first position and the second position.
8. The method according to claim 1, characterized in that, The determining the adjacent cell corresponding to the target cell in the target table includes: In response to a filling operation for the target cell in the target table, determining a first data filled in the target cell; In the case where the first data is preset data, determining the adjacent cell corresponding to the target cell in the target table.
9. The method according to claim 4, characterized in that After displaying the target table formula at the corresponding position of the target cell, the method further includes: In response to a re-filling operation for the target cell in the target table, determining a second data re-filled in the target cell; In the case where the second data represents a function, displaying the target function at the corresponding position of the target cell.
10. The method according to claim 9, wherein After displaying the target function at the corresponding position of the target cell, the method further includes: In response to a re-filling operation for the target cell in the target table, determining a third data re-filled in the target cell; In the case where the third data represents a parameter, displaying the target parameter at the corresponding position of the target cell.
11. The method according to claim 1, wherein The method further includes: In response to a preset operation for the target column cell in the target table, detecting whether at least two identical table formulas are filled in the cell of the target column cell; In the case where at least two identical table formulas are filled in the cell of the target column cell, triggering a table formula filling prompt; The table formula filling prompt is used to prompt the remaining cells of the target column cell to fill the table formula.
12. A table formula processing device, characterized in that, The device includes: A table acquisition module for acquiring a target table; A cell determination module for determining the adjacent cell corresponding to the target cell in the target table; A table structuring processing module for performing structuring processing on the target table to obtain structured data corresponding to the target table; A table formula prediction module for predicting a target table formula according to the target cell, the adjacent cell, and the structured data; A table formula processing module for displaying the target table formula at the corresponding position of the target cell.
13. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete their mutual communication through the communication bus; The memory is used to store computer programs; The processor is used to implement the method described in any one of claims 1-11 when executing the program stored on the memory.
14. A storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method described in any one of claims 1-11.