Table recombination method and device, electronic equipment and medium
By identifying and transforming the identifiers in the table cells, splitting and reorganizing the table according to the splitting rules, the problem of users being unable to understand multiple data types is solved, and the user experience is improved.
Patent Information
- Application Number
- CN202510779570.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-10-10
AI Technical Summary
Existing technologies cannot effectively split and reorganize table cells containing multiple data types in a database, resulting in a poor user experience.
By identifying identifiers in table cells, transforming and splitting the identifiers according to splitting rules, and generating new cells to reorganize the table.
It improves users' understanding and experience of table content and meets customized analysis needs.
Smart Images

Figure CN120764503A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of natural language processing (NLP, Neuro-Linguistic Programming), and in particular to a table reorganization method and device, an electronic device and a medium. BACKGROUND
[0002] The vigorous development of information technology means the arrival of the big data-oriented storage era. With the explosive growth of data, the storage form and structure of data have become increasingly complex. For example, data is stored in a database in the form of a table. However, a single cell of the table directly exported from the database contains data of multiple data types, which makes it difficult for users to directly understand the content of the cell from the table of the database, and the user experience is not good. SUMMARY
[0003] The present application provides a table reorganization method, device, electronic device and medium, which provides a new table reorganization scheme, can split the data of different data types contained in the cells of the table, and use the split data to form a reorganized table, thereby improving the user experience.
[0004] In a first aspect, the present application provides a table reorganization method, comprising:
[0005] extracting data of at least one cell belonging to a target name from an original table directly exported from a database;
[0006] identifying identifiers between the data of the at least one cell, and for each identifier of each cell, taking the identifier in the cell that matches a splitting rule as a target symbol; wherein the identifier is a punctuation symbol contained between data matching a plurality of splitting rules corresponding to a table type to which the original table belongs;
[0007] for each target symbol in each cell, modifying the target symbol in the cell according to a splitting rule matching the target symbol in the cell, so that the modified cell uses the same punctuation symbol to distinguish data of different data types; and splitting the data in the modified cell;
[0008] generating a plurality of new cells according to the split data, and forming a reorganized table according to the plurality of new cells.
[0009] The method can extract data of a cell of a target name from an original table, split data of different data types contained in the cell of the table according to identifiers between the cell data and splitting rules matched with the identifiers, and form a reorganized table by using the split data, thereby providing a new table splitting scheme and improving user experience.
[0010] In a possible implementation, the target name is determined in the following manner:
[0011] If the data of the cell in the original table contains a preset data structure, a name corresponding to the data of the cell containing the preset data structure is the target name; or
[0012] The names in the original table are displayed to a user, a selection instruction of the user is received, and a name selected in the selection instruction is taken as the target name; or
[0013] Data types contained in the cell in the original table are determined, the data types contained in the cell in the original table are displayed to the user, a selection instruction of the user is received, and a name corresponding to a data type selected in the selection instruction is taken as the target name.
[0014] In a possible implementation, whether the identifier in the cell matches a splitting rule is determined in the following manner:
[0015] From a plurality of splitting rules, a splitting rule in which a punctuation mark in matched data contains the identifier in the cell is found out;
[0016] When the number of the found splitting rules is a plurality, if there is a splitting rule in the plurality of found splitting rules that has a same data type as a data type to which a context of the identifier in the cell belongs, it is determined that the identifier in the cell matches the splitting rule, otherwise, it is determined that the identifier in the cell does not match the splitting rule;
[0017] When the number of the found splitting rules is one, if a data type of data matched with the found splitting rule is the same as a data type to which a context of the identifier in the cell belongs, it is determined that the identifier in the cell matches the splitting rule, otherwise, it is determined that the identifier in the cell does not match the splitting rule.
[0018] In a possible implementation, the target symbol in the cell is modified according to the splitting rule matched with the target symbol in the cell, and the modification includes:
[0019] if the split rule matched with the target symbol in the cell is a first type of rule, the target symbol in the cell is kept, and a punctuation mark in data before the target symbol in the cell is adjusted according to a punctuation mark in the cell;
[0020] if the split rule matched with the target symbol in the cell is a second type of rule, the target symbol in the cell is replaced with a preset split symbol;
[0021] if the split rule matched with the target symbol in the cell is a third type of rule, the target symbol in the cell is kept.
[0022] In a possible implementation, after the target symbol in the cell is replaced with the preset split symbol, before the data in the modified cell is split, the method further includes:
[0023] determining a data type to which a specific section and / or key character contained in a context of the target character in the replaced cell belong;
[0024] if the data types to which the context of the target character in the replaced cell belong are the same, the target character in the replaced cell is restored;
[0025] if the data types to which the context of the target character in the replaced cell belong are different, the target character in the replaced cell is kept.
[0026] In a possible implementation, after the identifiers between the data in at least one cell are identified, before the identifiers matched with the split rule in the cell are taken as the target symbol, the method further includes:
[0027] for each cell, determining similar fields in the data in the cell;
[0028] processing the similar fields in the data in the cell by using different processing schemes, so that there is no similar field in the data in the processed cell; the processing schemes include deleting the similar fields in the data in the cell and replacing the similar fields in the data in the cell with preset symbols;
[0029] before the data in the modified cell is split, the method further includes:
[0030] comparing the data in the modified cell with the data in the cell before modification;
[0031] if there is a field deleted in the data in the modified cell, the deleted field is restored to the data in the modified cell;
[0032] If the data in the modified cell contains the preset symbol, the field corresponding to the preset symbol is restored to the data in the modified cell.
[0033] In a possible implementation, the plurality of new cells is generated according to the plurality of split data, including:
[0034] For each of the split data, the split data is taken as content in the new cell; or
[0035] The reorganized table template corresponding to the data type to which the original table belongs is displayed to the user, a selection instruction of the user is received, the split data is combined according to a table name corresponding to the reorganized table template in the selection instruction, and the combined content is taken as data in the new cell.
[0036] In a second aspect, an embodiment of the present application provides a table reorganization apparatus, including:
[0037] The extraction module is configured to extract data of at least one cell belonging to a target name from an original table directly exported from a database.
[0038] The identification module is configured to identify identifiers between data of the at least one cell, and take, for each identifier of each cell, the identifier in the cell that matches a split rule as a target symbol; wherein the identifier is a punctuation symbol contained between data matching a plurality of split rules corresponding to a table type to which the original table belongs.
[0039] The split module is configured to, for each target symbol in each cell, modify the target symbol in the cell according to a split rule matching the target symbol in the cell, so that the same punctuation symbol is used to distinguish data of different data types in a modified cell; and split data in the modified cell.
[0040] The reorganization module is configured to generate a plurality of new cells according to the plurality of split data, and to generate a reorganized table according to the plurality of new cells.
[0041] In a third aspect, an embodiment of the present application provides an electronic device, including:
[0042] A processor;
[0043] The processor is configured to execute a computer program or instruction in the memory, so that the table reorganization method as described in any one of the first aspects is executed.
[0044] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, when instructions in the storage medium are executed by a processor, the processor is enabled to perform the table reorganization method according to any one of the first aspect.
[0045] In a fifth aspect, an embodiment of the present application provides a computer program product, the computer program product comprising: computer program code, when the computer program code is run on a computer, the computer is enabled to perform the table reorganization method according to any one of the first aspect.
[0046] In addition, the technical effects brought by any one of the implementation manners of the second aspect to the fifth aspect can refer to the technical effects brought by different implementation manners of the first aspect, which will not be repeated here.
[0047] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF DRAWINGS
[0048] Figure 1 A table splitting schematic diagram provided by the related art;
[0049] Figure 2 A table reorganization system schematic diagram provided by an embodiment of the present application;
[0050] Figure 3 A table reorganization method flowchart schematic diagram provided by an embodiment of the present application;
[0051] Figure 4 A table content extraction schematic diagram provided by an embodiment of the present application;
[0052] Figure 5 A target name cell data extraction schematic diagram provided by an embodiment of the present application;
[0053] Figure 6 A data identifier identification schematic diagram provided by an embodiment of the present application;
[0054] Figure 7 A table reorganization schematic diagram provided by an embodiment of the present application;
[0055] Figure 8 A display page schematic diagram provided by an embodiment of the present application;
[0056] Figure 9 A similar field data processing schematic diagram provided by an embodiment of the present application;
[0057] Figure 10 Another similar field data processing schematic diagram provided by an embodiment of the present application;
[0058] Figure 11 A schematic diagram of locating a target symbol in data of a cell is provided for an embodiment of the present application;
[0059] Figure 12 A schematic diagram of replacing a preset split symbol in data of a cell is provided for an embodiment of the present application;
[0060] Figure 13 A schematic diagram of a restored target symbol in data of a cell is provided for an embodiment of the present application;
[0061] Figure 14 A schematic diagram of data of a similar field after a restoration process is provided for an embodiment of the present application;
[0062] Figure 15 A schematic diagram of table reorganization is provided for an embodiment of the present application;
[0063] Figure 16 A schematic diagram of a structure of an electronic device is provided for an embodiment of the present application. DETAILED DESCRIPTION
[0064] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0065] In the description of the embodiments of the present application, "a plurality of" means two or more than two, unless otherwise specified.
[0066] Hereinafter, the terms "first" and "second" are only for descriptive purposes, and cannot be understood as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first" and "second" can explicitly or implicitly include one or more of the features.
[0067] As mentioned above, most of the data tables directly derived from the database cannot meet the analysis requirements. Therefore, a scheme for disassembling and reorganizing the data tables directly derived from the database is needed.
[0068] The existing table disassembling and reorganizing scheme is to combine the cell merging / splitting according to the different cells Figure 1As shown, the table directly extracted from the database is data with four cell names of achievement code, achievement name, belonging list, and MSS R&D project information. The difference between the "R&D achievement list" and the "full R&D achievement list" in the six cells of the name of the belonging list can be split. The virtual frame C1 is four cells in the third row, and the virtual frame C2 is four cells in the sixth row. The reorganized table is obtained by combining the virtual frame C1 and the virtual frame C2. The name of the table is the same as that of the original table. The fourth cell in the second row, the four cells in the fourth row, the four cells in the fifth row, and the four cells in the seventh row except the virtual frame C1 and the virtual frame C2 are combined to form a new table. The name of the table is the same as that of the original table.
[0069] The cells of the table can contain data of various data types. The user wants to directly form a table by combining data of different data types in different cells. However, the existing scheme cannot split the data in the cells.
[0070] Based on this, the embodiment of the present application provides a new table reorganization scheme, which can split the data of different data types contained in the cells of the table and use the split data to form a reorganized table.
[0071] The object implementation, functional characteristics and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.
[0072] For example, the electronic device in the embodiment of the present application can be a server, which can be a standalone physical server, a server cluster composed of multiple physical servers or a distributed system, a cloud server providing basic cloud computing services such as cloud service, cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, content distribution network (Content Delivery Network, CDN), and big data and artificial intelligence platform; wherein the server can be connected with a device having a display function through a wired or wireless manner, and the present application does not limit the connection manner.
[0073] It can be understood that the specific type of the electronic device in the present application is not limited.
[0074] The possible application scenarios of the table reorganization method provided by the embodiment of the present application will be introduced as follows with reference to the accompanying drawings:
[0075] For example, the electronic device in the embodiment of the present application can be a server, which can be a standalone physical server, a server cluster composed of multiple physical servers or a distributed system, a cloud server providing basic cloud computing services such as cloud service, cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, content distribution network (Content Delivery Network, CDN), and big data and artificial intelligence platform; wherein the server can be connected with a device having a display function through a wired or wireless manner, and the present application does not limit the connection manner. Figure 2As shown in FIG. 1, a schematic diagram of an application scenario in an embodiment of the present application is shown. The embodiment of the present application provides a database system, which includes a plurality of clients 1-n, a server 100, and a database 200. The clients 1-n are connected with the server 100 respectively, and the server 100 is connected with the database 200. The function of the client is to query a table, the function of the server 100 is to receive the information of the table queried by the client, to reorganize the data in the table in the database, and to send the reorganized table to the client. The function of the database 200 is to store data in the form of a table.
[0076] For example, in combination with Figure 2 As shown in FIG. 2, a user uses the client 1, the client 1 is connected with the server 100, the user uses the client 1 to query a table from the database, sends the information of the table queried to the server 100, the server 100 obtains the table from the database, reorganizes the data in the table according to the requirement of the user, sends the reorganized table to the client 1, and the client 1 displays the reorganized table to the user.
[0077] Of course, the method provided by the embodiment of the present application is not limited to Figure 2 The application scenario shown in FIG. 2 can also be used in other possible application scenarios, and the embodiment of the present application is not limited in any way.
[0078] After the application scenario of the embodiment of the present application is introduced, the preferred embodiment of the present application will be further described in detail in combination with the drawings. It should be understood that the preferred embodiment described herein is only used for description and explanation of the present application, and is not used for limiting the present application, and the features in the embodiment of the present application and the embodiment can be combined with each other without conflict.
[0079] The table reorganization method provided by the embodiment of the present application will be specifically introduced in combination with the drawings as follows.
[0080] Figure 3 As shown in FIG. 3, a schematic diagram of the working process of a table reorganization method provided by the embodiment of the present application is shown. As shown in FIG. 4, the specific process of the method is as follows. Figure 3
[0081] S310: Extracting the data of at least one cell belonging to the target name from the original table directly exported from the database.
[0082] The name of the original table is the data of the cell in the first row of the original table, or the data of the cell in the first column of the original table. For example, the data of the four cells in the first row of the original table are achievement code, achievement name, belonging list, and MSS R&D project information, and the target name can be any one or more of the achievement code, the achievement name, the belonging list, and the MSS R&D project information. Figure 1 Figure 1 For example, the target name can be MSS R&D project information.
[0083] For example, Figure 1 The table shown in the figure is arranged from top to bottom and from left to right. Figure 1 The data of the cells in the table are extracted, as shown in Figure 4 As shown, single cells are enclosed in single quotes, each row of data is enclosed in square brackets, all extracted data is enclosed in square brackets, and two rows of data are separated by commas. The data without underline is the data of the four cells in the first row, the data with a single solid underline is the data of the four cells in the second row, the data with a double solid underline is the data of the four cells in the third row, the data with a single dashed underline is the data of the four cells in the fourth row, the data with a single thin wavy underline is the data of the four cells in the fifth row, the data with a single thick wavy underline is the data of the four cells in the sixth row, and the data with a double wavy underline is the data of the four cells in the seventh row.
[0084] The target name is MSS R&D project information. The data types of MSS R&D project information are MSS achievement code, main project information, and sub-project information. Use MSS achievement code, main project information, and sub-project information as new names to obtain the data of 6 cells belonging to MSS R&D project information. Figure 5 As shown, the data of the fourth cell in the second row is:
[0085] XXXXXXXXX:xxxx[Master:XXXXXXXXX,xxxxxxx,xxxxxxxxxx;
[0086] Sub: XXXXXXXXX,xxx,xxxx,xxxxx(xxxxx)].
[0087] The data in the fourth cell of the third row is:
[0088] XXXXXXXXX:xxxx[Master:XXXXXXXXX,xxxxxxxxxxxx,xxxxx;
[0089] Sub:XXXXXXXXX,xxxxx,xxxxxxx(xxxxx)].
[0090] The data in the fourth cell of the fourth row is:
[0091] XXXXXXXXX:xxxx[Master:XXXXXXXXX,xxxxxxxxxxxx,xxxx;
[0092] Sub: XXXXXXXXX,xxxx[xxxx]xxx,x(xxxxx)].
[0093] The data in the fourth cell of the fifth row is:
[0094] XXXXXXXXX: xxx [main: XXXXXXXX, xxxxxxxxxxxx, xxxxxxxx;
[0095] sub: XXXXXXXX, xxxxxxxx, xxxxx (xxxxx)].
[0096] The data in the fourth cell of the sixth row is:
[0097] XXXXXXXXX: xxx [main: XXXXXXXX, xxxxxxxxxxxxx, xxxxx;
[0098] sub: XXXXXXXX, xxx, xxxx, xxxxx; xxx, xx].
[0099] The data in the fourth cell of the seventh row is:
[0100] XXXXXXXXX: xxx [main: XXXXXXXX, xxxxxxxx, xxxxxxxx;
[0101] sub: XXXXXXXX, [xxx], xxxxx, xxxxxxxx, (xxxxx)],
[0102] XXXXXXXXX: xxx [main: XXXXXXXX, xxxxx, xxxxxxxx, xxxxx; sub: XXXXXXXX, xxxxxxxxxxxxx, xxx].
[0103] S320: Identify the identifier between the data of at least one cell, and for each identifier of each cell, take the identifier in the cell that matches the splitting rule as the target symbol; wherein the identifier is a punctuation symbol contained between the data that matches the plurality of splitting rules corresponding to the table type to which the original table belongs.
[0104] Wherein, the punctuation symbol can be a semicolon, parentheses, square brackets, braces, quotation marks, a comma, a period. The identifier can be half of the square brackets, i.e. [; the other half of the square brackets and the comma, i.e. ], the parentheses, i.e. (), the other half of the square brackets], the semicolon, i.e. ;.
[0105] For example, as shown in Figure 6 , when the identifier is half of the square brackets, i.e. [; the other half of the square brackets and the comma, i.e. ], the parentheses, i.e. (), the other half of the square brackets], the semicolon, i.e. ;, the positioning box of the identifier in the data of at least one cell is drawn, and the positioning box isFigure 6 as shown in the solid box in the table.
[0106] The table type can be classified according to different fields, and can also be classified according to tables of different units or companies, and the present application does not make specific limitations. For example, the table type includes a research and development project table, a medical procurement table, and the like. The splitting rule for the research and development project table type is to split different achievements, to split the parallel sub-projects and main projects, to split the sub-projects under the main project, to split the achievement codes and the main project, not to split the subsidiary associated information in the sub-projects, and not to split the subsidiary associated information in the main project. The splitting rule for the medical procurement table type is to split different medical devices.
[0107] The punctuation marks contained in the data that meets the splitting different achievements splitting rule are "]," and "[", the punctuation marks contained in the data that meets the splitting the parallel sub-projects and main projects splitting rule are " ] " and " ], ", the punctuation marks contained in the data that meets the splitting the sub-projects under the main project splitting rule are " ; ", the punctuation marks contained in the data that meets the splitting the achievement codes and the main project splitting rule are "[", the punctuation marks contained in the data that meets the not splitting the subsidiary associated information in the sub-projects splitting rule are " ( ) ", and the punctuation marks contained in the data that meets the not splitting the subsidiary associated information in the main project splitting rule are " ( ) ".
[0108] S330: For each target symbol in each cell, according to the splitting rule matched with the target symbol in the cell, the target symbol in the cell is modified so that the modified cell uses the same punctuation mark to distinguish data of different data types; and the data in the modified cell is split.
[0109] For example, as shown in the table, Figure 6 For example, the data in the fourth cell of the second row is taken as an example:
[0110] XXXXXXXXX: xxxx [main: XXXXXXXX, xxxxxxx, xxxxxxxxx;
[0111] Sub: XXXXXXXX, xxx, xxxx, xxxxx (xxxxx) ].
[0112] The "[" before the field "main" is an identifier, the field containing the identifier is "XXXXXXXXX:xxxx[main:XXXXXXXXX,xxxxxxx,xxxxxxxxxx", the field matches the split rule "split the achievement code and the main project", the field "main" before the field is modified, and the field "XXXXXXXXX:xxxx[main:XXXXXXXXX,xxxxxxx,xxxxxxxxxx" is split into "XXXXXXXXX:xxxx" and "main:XXXXXXXXX,xxxxxxx,xxxxxxxxxx".
[0113] The ";" before the field "sub" is an identifier, the field containing the identifier is "main:XXXXXXXXX,xxxxxxx,xxxxxxxxxx;
[0114] The field matches the split rule "split the sub-projects under the main project", the ";" before the field "sub" is modified, and the field "main:XXXXXXXXX,xxxxxxx,xxxxxxxxxx;
[0115] The field "sub:XXXXXXXXX,xxx,xxxx,xxxxx(xxxxx)" is split into:
[0116] main:XXXXXXXXX,xxxxxxx,xxxxxxxxxx and
[0117] sub:XXXXXXXXX,xxx,xxxx,xxxxx(xxxxx).
[0118] The "(" after the field "sub" is an identifier, the field containing the identifier is "sub:XXXXXXXXX,xxx,xxxx,xxxxx(xxxxx)", the field matches the split rule "do not split the subsidiary information in the sub-projects". According to the "(" after the field "sub", the data inside the "(" and outside the "(" is not split.
[0119] S340: generate a plurality of new cells according to the split data, and form a reorganized table according to the plurality of new cells.
[0120] The split data is data of different data types, for example, the split data of the data in the fourth cell of the second row includes: XXXXXXXX:xxxx, which is of a data type of MSS achievement code, main: XXXXXXXX, xxxxxxx, xxxxxxxxx, which is of a data type of main project information, and sub: XXXXXXXX, xxx, xxxx, xxxxx (xxxxx), which is of a data type of sub project information, and the like. The data in the fourth cell of the third to seventh rows are all processed in the above manner, and a reorganized table is formed according to the 20 split data, as shown in Figure 7 .
[0121] According to the above scheme, the embodiment of the present application provides a new table splitting manner. The data in at least one cell of a target name is obtained, the identifier in the cell data is used, the split rule matched with the identifier is used to split the data in the cell, and the split data is used to reorganize the table. In this way, the data in the cell can be split again, the problem that the data table directly exported from the database cannot meet the relatively customized requirements is solved, and the user experience is improved.
[0122] In some embodiments, the target name is determined in the following manner.
[0123] Manner one: if the data in the cell in the original table contains a preset data structure, the name corresponding to the data in the cell containing the preset data structure is the target name.
[0124] Manner two: the names in the original table are displayed to the user, a selection instruction of the user is received, and the name selected in the selection instruction is taken as the target name.
[0125] Manner three: the data types contained in the cell in the original table are determined, the data types contained in the cell in the original table are displayed to the user, a selection instruction of the user is received, and the name corresponding to the data type selected in the selection instruction is taken as the target name.
[0126] In detail, for manner one, the preset data structure includes a parallel structure and the like. The parallel structure is as shown in Figure 1 . The data in the cell corresponding to the name of MSS research and development project information contains three kinds of information, i.e., MSS achievement code, main project information, and sub project information. Then, the name of MSS research and development project information is taken as the target name. For example, in a medical sales table, the data in the cell corresponding to the name of purchase equipment includes different types of equipment, such as A equipment, B equipment, C equipment, D equipment, E equipment, F equipment, and G equipment. Then, the name of purchase equipment is taken as the target name.
[0127] For the second mode, after directly exporting the original table from the database, the names in the original table are displayed to the user, and the user can directly select the name as the target name. For the third mode, the data type contained in the cell in the original table is determined, the data type is displayed to the user, and the name corresponding to the cell containing the data type selected by the user is taken as the target name.
[0128] The data type contained in the cell in the original table can be matched according to the data type of the similar table, and the matched type is taken as the data type in the cell of the table.
[0129] The similar table can be a table with the same structure as the original table, that is, a similar table, for example, the names of the tables are all achievement codes, achievement names, belonging lists, and MSS R&D project information.
[0130] For example, the data type of the similar table is MSS achievement code, main project information, and sub-project information, so the data type of the data in the cell corresponding to the name of MSS R&D project information is determined as the above data type, the user is displayed with MSS achievement code, main project information, and sub-project information, and after the user selects one or more of them, if MSS achievement code is selected, the name of MSS R&D project information is taken as the target name.
[0131] Exemplarily, the embodiment of the present application provides a display page, which is combined with Figure 8 As shown, the display page includes inputting the table name, and the data type of the data needed by the user or the name in the table. The input table name is the entire name of the table, for example, MSS R&D project statistics table, medical purchase record table, etc. The data type of the data needed by the user can be selected, and then the item of the name in the table will be grayed out, that is, it cannot be operated, and the name corresponding to the cell containing the selected data type is taken as the target name, for example, MSS achievement code, main project information, and sub-project information are selected, and the name in the table is Figure 5 The gray color is the data type selected by the user, and the cell of the name of MSS R&D project information contains the data of the selected data type, and the name of MSS R&D project information is taken as the target name. The data type of the data needed by the user can also not be selected, that is, “none” is selected, and the target name is the name of the cell containing the preset data structure. If the user selects the name in the table, for example, all the names in the table are displayed, achievement code, achievement name, belonging list, and MSS R&D project information, when the user selects MSS R&D project information, MSS R&D project information is taken as the target name. After determining the target name and the original table, the original table is split by using the target name, the split cells are reorganized, and the reorganized table is obtained.
[0132] Due to the similar data in the project information, the similar data can cause two different data to be considered as the same data when matching the split rule, thereby causing identification errors. For example, the MSS achievement code and the main project information have similar information. In order to improve the accuracy, after identifying the identifier between the data of at least one cell, and before taking the identifier in the cell matching the split rule as the target symbol, the method further comprises:
[0133] For each cell, determine the similar field in the data of the cell;
[0134] The similar field in the data of the cell is processed by using different processing schemes, so that there is no similar field in the processed data of the cell. The processing scheme includes deleting the similar field in the data of the cell, and replacing the similar field in the data of the cell with a preset symbol.
[0135] For example, the first three characters of the main project information and the MSS achievement code are similar and are determined as similar characters, or are called sticky characters, that is, replacing “main:” with “$”, removing the first three characters of the MSS achievement code, and temporarily storing the replacement pair in the form of a dictionary. Figure 9 As shown in the formula, the first three characters of the MSS achievement code and the main project information in each cell are similar characters. The similar field in the MSS achievement code “XXXXXXXX: xxxx” is deleted, and the deleted MSS achievement code is “XXXXX: xxxx”. The similar field in the main project information “main: XXXXXXXX, …” in each cell is replaced with “$”, and the replaced main project information is “$XXXXXXXX, …”. Figure 10 As shown in the formula, the similar field in the main project information “main: XXXXXXXX, …” in each cell is deleted.
[0136] In some embodiments, whether the identifier in the cell matches the split rule is determined by:
[0137] From a plurality of split rules, find a split rule whose punctuation symbol in the matching data contains the identifier in the cell;
[0138] When the number of found split rules is more than one, if there is a split rule in the found plurality of split rules that has the same data type as the data type to which the context of the identifier in the cell belongs, it is determined that the identifier in the cell matches the split rule, otherwise, it is determined that the identifier in the cell does not match the split rule.
[0139] When the number of splitting rules found is one, if the data type of the data matching the found splitting rule is the same as the data type of the context of the identifier in the cell, it is determined that the identifier in the cell matches the splitting rule; otherwise, it is determined that the identifier in the cell does not match the splitting rule.
[0140] For example, the data from the fourth cell in the second row to the fourth cell in the seventh row are extracted as follows: Figure 6 As shown, the solid box is the identifier in the cell data. Figure 11 To process the data after similar fields, Figure 11The solid frame in the figure is the target symbol in the cell data, in detail, the lower line is the data of the single solid line, wherein the identifier in the cell is: the two "[" in front of the main project information, the ";" in front of the character "sub", the "(" after the character "sub", the "]" after the character "sub" and the "]", ". Among them, the "[" in front of the main project information is the punctuation symbol possessed by the splitting rule of splitting the achievement code and the main project, the upper context of the "[" in front of the main project information is the data of a single cell, and the lower context is the data of a single cell. Therefore, the identifier does not match the splitting rule, and it is not the target symbol. The second front "[" of the main project information is the punctuation symbol possessed by the splitting rule of splitting the achievement code and the main project, the upper context of the second front "[" of the main project information is "XXXXXX: xxxx, the data type of which is achievement code, and the lower context is: XXXXXXXXXXXX, xxxxxxx, xxxxxxxxxxxx", the data type of which is main project information, which matches the splitting rule of splitting the achievement code and the main project. Therefore, the identifier matches the splitting rule, and it is the target symbol. The ";" in front of the character "sub" is the punctuation symbol possessed by the splitting rule of splitting the sub-projects under the main project, the upper context of which is "XXXXXXXXX, xxxxxxx, xxxxxxxxxxxx, the data type of which is main project information; the lower context of which is: sub: XXXXXXXXXXXX, xxx, xxxx, xxxxx (xxxxx)", the data type of which is sub-project information, which matches the splitting rule of splitting the sub-projects under the main project. Therefore, the identifier matches the splitting rule, and it is the target symbol. The "(" after the character "sub" is the punctuation symbol possessed by the splitting rule of not splitting the subsidiary information in the sub-project, the upper and lower contexts of which are "sub: XXXXXXXXXXXX, xxx, xxxx, xxxxx (xxxxx)", which matches the splitting rule of not splitting the subsidiary information in the sub-project. Therefore, the identifier matches the splitting rule, and it is the target symbol. The "]" after the character "sub" matches the punctuation symbol of the splitting rule of splitting the parallel sub-projects and main projects, the upper and lower contexts of which are "sub: XXXXXXXXXXXX, xxx, xxxx, xxxxx (xxxxx)]", which does not match the splitting rule of splitting the parallel sub-projects and main projects. Therefore, the identifier does not match the splitting rule, and it is not the target symbol. The "]", " after the character "sub" matches the punctuation symbol of the splitting rule of splitting different achievements and splitting the parallel sub-projects and main projects, the upper and lower contexts of which are:
[0141] sub: XXXXXXXXXXXX, xxx, xxxx, xxxxx (xxxxx)]', '[XXXXXXXXX: xxxx, which is the splitting rule of splitting different achievements, the identifier of which matches the splitting rule, and it is the target symbol.
[0142] Similarly, the data underlined by double solid line is the data in the fourth cell of the third row, in which the target symbols in the cell are: "[" immediately before the main project information, ";" immediately before the character "sub", "()" immediately after the character "sub", and "]" immediately after the character "sub". The data underlined by single dash line is the data in the fourth cell of the fourth row, in which the target symbols in the cell are: "[" immediately before the main project information, ";" immediately before the character "sub", "[" and "]" and "()" immediately after the character "sub", and "]" immediately after the character "sub". The data underlined by single thin wavy line is the data in the fourth cell of the fifth row, in which the target symbols in the cell are: "[" immediately before the main project information, ";" immediately before the character "sub", "()" immediately after the character "sub", and "]" immediately after the character "sub". The data underlined by single thick wavy line is the data in the fourth cell of the sixth row, in which the target symbols in the cell are: "[" immediately before the main project information, ";" immediately before the character "sub", ";" and "()" immediately after the character "sub", and "]" immediately after the character "sub". The data underlined by double wavy line is the data in the fourth cell of the seventh row, in which the target symbols in the cell are: "[" immediately before the first main project information, ";" immediately before the character "sub", "[" and "]" and "()" and "]" immediately after the first character "sub", and "[" immediately before the second main project information, ";" immediately before the second character "sub", and the last "]" immediately after the second character "sub".
[0143] In some embodiments, step 330 transforms the implementation of the target symbol in the cell according to the splitting rule matched by the target symbol in the cell, as follows:
[0144] If the splitting rule matched by the target symbol in the cell is the first type of rule, the target symbol in the cell is kept, and the punctuation symbols in the data before the target symbol in the cell are adjusted according to the punctuation symbols in the cell.
[0145] If the splitting rule matched by the target symbol in the cell is the second type of rule, the target symbol in the cell is replaced by a preset split symbol.
[0146] If the splitting rule matched by the target symbol in the cell is the third type of rule, the target symbol in the cell is kept.
[0147] For example, splitting different achievements is the first type of rule, splitting parallel sub-projects and main projects, splitting sub-projects under main projects, and splitting achievements coding and main projects are the second type of rule, and not splitting the subsidiary information in the sub-projects and not splitting the subsidiary information in the main projects are the third type of rule.
[0148] In combinationFigure 12 As shown, the solid box is the preset separator symbol to be replaced, and the data with a single solid line below is the data of the fourth cell in the second row, where the target symbol in the cell is: the "[" immediately before the main project information matches the splitting rule of splitting the result code and the main project, which complies with the second type of rule and is replaced with "','", the ";" before the character "子" matches the splitting rule of splitting the sub-projects under the main project, which complies with the second type of rule and is replaced with "','", the "()" retained by the character "子" matches the splitting rule of not splitting the ancillary information in the sub-project, which complies with the third type of rule, including parentheses, the "]," after the character "子", matches the splitting rule of splitting different results, which complies with the first type of rule and retains the target character. At the same time, the "]" immediately after the parentheses in the sub-project information originally matched the "[" before the main project information, but the "[" before the main project information has been replaced with the preset split symbol, so the "]" here is useless and is deleted. Then the data of the fourth cell in the second row is transformed to: ['XXXXXX:xxxx',
[0149] 'XXXXXXXXX,xxxxxxx,xxxxxxxxxx',
[0150] 'Sub: XXXXXXXXX,xxx,xxxx,xxxxx(xxxxx)'].
[0151] Similarly, the data in the fourth cell of the third row is processed in the same way as the data in the fourth cell of the second row. Then the data in the fourth cell of the third row is transformed into: ['XXXXXX:xxxx', 'XXXXXXXXX,xxxxxxxxxxxx,xxxxx',
[0152] 'Sub: XXXXXXXXX,xxxxx,xxxxxxx(xxxxx)'].
[0153] Similarly, for the data in the fourth cell of the fourth row, the target symbol in the cell is: the "[" immediately before the main project information matches the splitting rule of splitting the result code and the main project, which complies with the second type of rule and is replaced by "','"; the ";" before the character "子" matches the splitting rule of splitting the sub-projects under the main project, which complies with the second type of rule and is replaced by "','"; the "()" after the character "子" matches the splitting rule of not splitting the ancillary information in the sub-project, which complies with the third type of rule and retains the parentheses; the "[" after the character "子" matches the splitting rule of splitting the result code And the main project, in line with the second type of rules, replace it with "','", the "]" after the character "子" matches the splitting rule for splitting the parallel sub-projects and main project, in line with the second type of rules, replace it with "','", the "]," after the character "子" matches the splitting rule for splitting different results, in line with the first type of rules, retain the target character, at the same time, the "]" immediately after the parentheses in the sub-project information was originally matched with the "[" in front of the main project information, but the "[" in front of the main project information has been replaced with the preset segmentation symbol, so the "]" here is useless and will be deleted. Then the data in the fourth cell of the fourth row is transformed into: ['XXXXXX:xxxx', 'XXXXXXXXX,xxxxxxxxxxxx,xxxx', '子:XXXXXXXXX,xxxx', 'xxxx', 'xxx,x(xxxxx)'].
[0154] Similarly, the data of the fourth cell in the fifth row is processed in the same way as the data of the fourth cell in the second row. Then the data of the fourth cell in the fifth row is transformed into: ['XXXXXX:xxxx', 'XXXXXXXXX,xxxxxxxxxx,xxxxxxx',
[0155] 'Sub: XXXXXXXXX,xxxxxxx,xxxxx(xxxxx)'].
[0156] Similarly, the data in the fourth cell of the sixth row, the target symbol in the cell is: the split rule matched with the immediately preceding "[" of the main project information is to split into achievement code and main project, which conforms to the second type of rule, and is replaced by "', '"; the split rule matched with the ";" before the character "child" is to split the sub-projects under the main project, which conforms to the second type of rule, and is replaced by "', '"; the split rule matched with the ";" after the character "child" is to split the parallel sub-projects and the main project, which conforms to the second type of rule, and is replaced by "', '"; the split rule matched with "]" after the immediately following small parentheses in the sub-project information is to split different achievements, which conforms to the first type of rule, and the target character is retained. At the same time, the "]" immediately following the small parentheses in the sub-project information is originally matched with the "[" in front of the main project information, but the "[" in front of the main project information has been replaced by the preset split symbol, so the "]" at this position is useless and is deleted. Then the data in the fourth cell of the sixth row is transformed into: [ 'XXXXXX:xxxx',
[0157] 'XXXXXXXXX, xxxxxxxxxxxx, xxxxx', 'child:XXXXXXXXX, xxx, xxxx, xxxxx', 'xxx, xx'].
[0158] Among them, the data of the fourth cell in the seventh row, among which, the target symbols in the cell are: the "[" immediately before the first main project information, the ";" before the first character "sub", the "[" immediately before the first main project information matches the splitting rule of splitting the achievement code and the main project, which complies with the second type of rules and is replaced with "','", the ";" before the first character "sub" matches the splitting rule of splitting the sub-projects under the main project, which complies with the second type of rules and is replaced with "','", the "()" after the first character "sub" matches the splitting rule of not splitting the ancillary information in the sub-project, which complies with the third type of rules and retains the parentheses, the "[" after the first character "sub" matches the splitting rule of splitting the parallel sub-projects, which complies with the second type of rules and is replaced with "','", the "]" after the first character "sub" matches the splitting rule of splitting the parallel sub-projects, It meets the second type of rules and is replaced by "','". The "]," after the first character "子" matches the splitting rule of splitting different main projects. It meets the second type of rules and is replaced by "','". The "[" immediately before the second main project information matches the splitting rule of splitting the achievement code and the main project. It meets the second type of rules and is replaced by "','". The ";" before the second character "子" matches the splitting rule of splitting the sub-projects under the main project. It meets the second type of rules and is replaced by "','". The "]" at the end of the second character "子" matches the splitting rule of splitting different achievements. It meets the first type of rules and retains the target character. At the same time, the "]" immediately adjacent to the sub-project information originally matched the "[" before the second main project information, but the "[" before the second main project information has been replaced by the preset segmentation symbol. Therefore, the "]" here is useless and is deleted. Then the data of the fourth cell in the seventh row is transformed as follows:
[0159] ['XXXXXX:xxxx','XXXXXXXXX,xxxxxxxx,xxxxxxx',
[0160] 'Sub:XXXXXXXXXX,','xxx',',xxxxx,xxxxxxx,(xxxxx)','XXXXXX:xxxx','XXXXXXXXX,xxxxx,xxxxxxx,xxxxx',
[0161] 'Sub:XXXXXXXXX,xxxxxxxxxxxx,xxx']].
[0162] The above processing is only exemplary, and there may be other processing results. For example, the ";" after the field "子" in the fourth cell of the sixth row can be identified as a separate field, so it can be replaced with "',';','".
[0163] In some embodiments, after replacing the target symbol in the cell with the preset split symbol, before splitting the data in the modified cell, the method further comprises:
[0164] determining the data type to which the context of the target character in the replaced cell belongs;
[0165] restoring the target character in the replaced cell if the data type to which the context of the target character in the replaced cell belongs is the same;
[0166] retaining the target character in the replaced cell if the data type to which the context of the target character in the replaced cell belongs is different.
[0167] For example, the specific segment of the MSS achievement code can be a code unique to the MSS achievement code, and the key character of the MSS achievement code can be a colon (:) used to split the MSS achievement code into two fields. The specific segment of the main project information can be a unique main project content, and the key character of the main project information can be a comma (,) used to split the main project information into multiple fields. The specific segment of the sub-project information can be a unique sub-project content, and the key character of the main project information can be a field containing "sub", a colon (:) and a comma (,).
[0168] For example, in combination with Figure 13 As shown in the data in the fourth cell of the second row, the target symbol in the cell is a left square bracket ([) immediately before the main project information. The context of the left square bracket is only split into two fields using a colon (:), and the data type is the MSS achievement code. The context has a unique main project content or is only split into multiple fields using a comma (,). The data type is the main project information. The target symbol [ is replaced without any problem, and thus does not need to be restored.
[0169] For example, in combination with Figure 13As shown, the real box is the reduced target symbol, and the underlined data is the data of the fourth cell in the fourth row, wherein the [ after the field "child" is replaced by ', ', the context of the [ after the field "child" is "child:XXXXXXXXX,xxxx", which includes "child" and contains ":", contains ", ", is child information, the context of "xxxx" is not any specific segment and / or key character of any data type, so it does not belong to a separate data type, and the [ is restored to ', '.
[0170] Similarly, in combination with Figure 13 As shown, the data in the fourth cell of the seventh row is processed, and the [ after the first field "child" is replaced by ', ', and the context belongs to the same data type, which is restored. The ] after the first field "child" is replaced by ', ', and the context belongs to the same data type, which is restored.
[0171] Exemplarily, in combination with Figure 13 As shown, the data in the fourth cell of the sixth row, wherein the ; after the field "child" is replaced by ', ', and the context of the ; after the field "child" is:
[0172] "child:XXXXXXXXX,xxx,xxxx,xxxxx", which includes "child" and contains ":", contains ", ", is child information, and the context of "xxx,xx" does not conform to any specific segment and / or key character of any data type, so the context belongs to the same data type, and the ','is restored to ;.
[0173] By analogy, the context in other replacements belongs to different data types, and no restoration is performed.
[0174] In some embodiments, before the data in the modified cell is split, the method further comprises:
[0175] Comparing the data in the modified cell with the data in the cell before modification;
[0176] If a field is deleted in the data in the modified cell, the deleted field is restored to the data in the modified cell;
[0177] If the data in the transformed cell contains a preset symbol, the field corresponding to the preset symbol is restored to the data in the transformed cell.
[0178] For example, combined Figure 14 As shown, the data in the transformed cells are restored, the MSS result code "XXXXX:xxxx" is restored to "XXXXXXXX:xxxx", and the main project information "$XXXXXXXX,..." is restored to "Main:XXXXXXXX,...".
[0179] In some embodiments, generating multiple new cells according to the split data includes:
[0180] For each split data, use the split data as the content of a new cell; or
[0181] Display the reorganized table template corresponding to the data type of the original table to the user, receive the user's selection instruction, combine the split data according to the table name corresponding to the reorganized table template in the selection instruction, and use the combined content as the data in the new cell.
[0182] In detail, Figure 1 As shown, the split fields are the MSS achievement code field, the main project information field, and the sub-project information field. These three items are used as the contents of three new cells, 6 projects. Since the fourth cell in the seventh row contains two project achievements, there are two MSS achievement codes, two main project information, and two sub-project information. Therefore, 20 new cells are generated, forming the following: Figure 7 shown.
[0183] For example, when the original table is a medical purchase table, for different hospital purchases of medical equipment, the name of each hospital's multiple medical equipment is a purchase item, and the original table is composed of data of the names of the hospital name, the total number of purchases, and the purchase item. Among them, the data of the purchase item of a hospital is equipment A, equipment B, equipment C, equipment D, equipment E, and equipment F. When the user selects the table name corresponding to the reorganized table template in the instruction as department aa, department bb, and department cc, the equipment A, equipment B, equipment C, equipment D, equipment E, and equipment F can be segmented by using the scheme of the application. The equipment used by department aa is equipment A and equipment C; the equipment used by department bb is equipment B, equipment D, and equipment E; and the equipment used by department cc is equipment F. Then, equipment A and equipment C are taken as the content of a single cell, equipment B, equipment D, and equipment E are taken as the content of a single cell, and equipment F is taken as the content of a single cell. In succession, the different hospitals are classified in the above manner, and a reorganized table can be generated with the names of the hospital name, the total number of purchases, department aa, department bb, and department cc.
[0184] The embodiment of the application can effectively realize the rapid disassembly processing of various tables, and store the disassembled content as one or more sub-tables according to the requirements. The preliminary positioning is realized based on the identifier, the problem of character symbol adhesion replacement is solved by using specific deletion reconstruction, and the problem of repetitive identification is solved by using semantic error correction reorganization.
[0185] As shown in Figure 15 The application further provides a table reorganization device, which comprises:
[0186] The extraction module 1510 is configured to extract data of at least one single cell belonging to a target name from an original table directly exported from a database;
[0187] The identification module 1520 is configured to identify identifiers between the data of the at least one single cell, and take, for each identifier of each single cell, the identifier in the single cell that matches a splitting rule as a target symbol; wherein the identifier is a punctuation symbol contained between data matching a plurality of splitting rules corresponding to a table type to which the original table belongs.
[0188] The splitting module 1530 is configured to, for each target symbol in each single cell, modify the target symbol in the single cell according to a splitting rule matching the target symbol in the single cell, so that the modified single cell uses the same punctuation symbol to distinguish data of different data types; and split the data in the modified single cell.
[0189] The reorganization module 1540 is configured to generate a plurality of new single cells according to the split data, and generate a reorganized table according to the plurality of new single cells
[0190] Optionally, the extraction module 1510 is further configured to:
[0191] if the data of the cell in the original table contains a preset data structure, the name corresponding to the data of the cell containing the preset data structure is the target name; or
[0192] display the names in the original table to the user, receive a selection instruction of the user, and take the name selected in the selection instruction as the target name; or
[0193] determine the data type contained in the cell in the original table, display the data type contained in the cell in the original table to the user, receive a selection instruction of the user, and take the name corresponding to the data type selected in the selection instruction as the target name.
[0194] Optionally, the identification module 1520 is further configured to:
[0195] from a plurality of splitting rules, find out a splitting rule in which the punctuation mark in the matched data contains the identifier in the cell;
[0196] if there is a splitting rule in which the data type of the matched data is the same as the data type to which the context of the identifier in the cell belongs, among the plurality of splitting rules found out, it is determined that the identifier in the cell matches the splitting rule, otherwise, it is determined that the identifier in the cell does not match the splitting rule;
[0197] if the data type of the data matching the splitting rule found out is the same as the data type to which the context of the identifier in the cell belongs, it is determined that the identifier in the cell matches the splitting rule, otherwise, it is determined that the identifier in the cell does not match the splitting rule.
[0198] Optionally, the splitting module 1530 is specifically configured to:
[0199] if the splitting rule matching the target symbol in the cell is the first type of rule, the target symbol in the cell is retained, and the punctuation mark in the data before the target symbol in the cell is adjusted according to the punctuation mark in the cell;
[0200] if the splitting rule matching the target symbol in the cell is the second type of rule, the target symbol in the cell is replaced by a preset split symbol;
[0201] if the splitting rule matching the target symbol in the cell is the third type of rule, the target symbol in the cell is retained.
[0202] Optionally, the splitting module 1530 is further configured to:
[0203] Determine the specific segment contained in the context of the target character in the replaced cell and / or the data type to which the key character belongs;
[0204] If the data type of the context of the target character in the replaced cell is the same, then restore the target character in the replaced cell;
[0205] If the data type of the context of the target character in the replaced cell is different, the target character in the replaced cell is retained.
[0206] Optionally, the identification module 1520 is specifically configured to:
[0207] For each cell, determining similar fields in the data of the cell;
[0208] Using different processing schemes to process similar fields in the data of the cells so that there are no similar fields in the processed data of the cells; wherein the processing schemes include deleting similar fields in the data of the cells and replacing similar fields in the data of the cells with preset symbols;
[0209] The splitting module 1530 is further configured to:
[0210] Compare the data of the transformed cell with the data of the cell before the transformation;
[0211] If a field is deleted in the data of the transformed cell, the deleted field will be restored to the data of the transformed cell;
[0212] If the data in the transformed cell contains a preset symbol, the field corresponding to the preset symbol is restored to the data in the transformed cell.
[0213] Optionally, the reassembly module 1540 is configured to:
[0214] For each split data, use the split data as the content of a new cell; or
[0215] Display the reorganized table template corresponding to the data type of the original table to the user, receive the user's selection instruction, combine the split data according to the table name corresponding to the reorganized table template in the selection instruction, and use the combined content as the data in the new cell.
[0216] In addition, combined Figures 1-15 The table reorganization method and apparatus described in the embodiments of the present invention may be implemented by an electronic device.
[0217] An electronic device, comprising: a processor;
[0218] a memory for storing the processor-executable instructions;
[0219] The processor is configured to execute the instructions to implement the table reorganization method of any of the above.
[0220] Based on the above, an exemplary electronic device structure is proposed Figure 16 .
[0221] The electronic device can include a processor 1610 and a memory 1620 storing computer program instructions.
[0222] In particular, the processor 1610 can include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement one or more embodiments of the present application.
[0223] The memory 1620 can include a mass storage for data or instructions. By way of example and not limitation, the memory 1620 can include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a Universal Serial Bus (USB) drive or a combination of two or more of these. The memory 1620 can be removable and / or non-removable (or fixed) as appropriate. The memory 1620 can be internal or external to the data processing apparatus as appropriate. In a particular embodiment, the memory 1620 is a non-volatile solid-state memory. In a particular embodiment, the memory 1620 includes read-only memory (ROM). The ROM can be mask programmed ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), electrically alterable ROM (EAROM), or flash memory, or a combination of two or more of these, as appropriate.
[0224] The processor 1610 implements the method of any of the above embodiments for performing tasks by reading and executing the computer program instructions stored in the memory 1620.
[0225] In one example, the electronic device can further include a communication interface 1630 and a bus 1640. As shown, the processor 1610, the memory 1620, and the communication interface 1630 are connected through the bus 1640 and complete communication among each other. Figure 16
[0226] The communication interface 1630 is mainly configured to implement the communication between the modules, devices, units and / or equipment in the embodiments of the present application.
[0227] The bus 1640 includes hardware, software, or both, that couples components of the electronic device to each other. As an example and not by way of limitation, the bus can include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand (IB) interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or another suitable bus or a combination of two or more of these. Where appropriate, the bus 1640 can include one or more buses. Although the present application is described and shown with respect to a particular bus, the present application contemplates any suitable bus or interconnect.
[0228] The electronic device can perform the table reorganization method in the embodiments of the present application based on the received task, thereby achieving the table reorganization method and device described in combination Figures 1-15 with the electronic device in the above embodiments.
[0229] In addition, in combination with the electronic device in the above embodiments, the embodiments of the present application can provide a storage medium, when the instructions in the storage medium are executed by the processor of the electronic device, the electronic device can perform the table reorganization method as described in any one of the above.
[0230] The present application is described with reference to flowcharts and / or block diagrams of the method, device (system), and computer program product according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device that implements the functions specified in the flow Figure 1 The device that implements the functions specified in one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0231] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0232] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0233] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0234] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
Claims
1. A table reorganization method, characterized in that: include: Extract data of at least one cell belonging to the target name from the original table directly exported from the database; Identifying an identifier between data of at least one cell, and for each identifier of each cell, using the identifier in the cell that matches a splitting rule as a target symbol; wherein the identifier is a punctuation symbol contained in the data between data that matches multiple splitting rules corresponding to the table type to which the original table belongs; For each target symbol in each cell, according to a splitting rule in the cell that matches the target symbol, transform the target symbol in the cell so that the transformed cell uses the same punctuation mark to distinguish data of different data types; and split the data in the transformed cell; A plurality of new cells are generated according to the split data, and a reorganized table is formed based on the plurality of new cells.
2. The table reorganization method according to claim 1, characterized in that: The target name is determined by: If the data of the cell in the original table includes a preset data structure, the name corresponding to the data of the cell including the preset data structure is used as the target name; or Displaying the names in the original table to the user, receiving a selection instruction from the user, and using the name selected in the selection instruction as the target name; or Determine the data type contained in the cells in the original table, display the data type contained in the cells in the original table to the user, receive a selection instruction from the user, and use the name corresponding to the data type selected in the selection instruction as the target name.
3. The table reorganization method according to claim 1, characterized in that: Determine whether the identifier in the cell matches a split rule by: Finding a splitting rule in which the punctuation mark in the matched data contains the identifier in the cell from a plurality of splitting rules; When there are multiple splitting rules found, if there is a splitting rule among the multiple splitting rules found that sets the data type of the matched data to be the same as the data type to which the context of the identifier in the cell belongs, then it is determined that the identifier in the cell matches the splitting rule; otherwise, it is determined that the identifier in the cell does not match the splitting rule; When the number of splitting rules found is one, if the data type of the data matching the found splitting rule is the same as the data type to which the context of the identifier in the cell belongs, it is determined that the identifier in the cell matches the splitting rule; otherwise, it is determined that the identifier in the cell does not match the splitting rule.
4. The table reorganization method according to claim 1, characterized in that: Transforming the target symbol in the cell according to a splitting rule matching the target symbol in the cell includes: If the splitting rule matching the target symbol in the cell is a first type rule, retaining the target symbol in the cell, and adjusting the punctuation marks in the data preceding the target symbol in the cell according to the punctuation marks in the cell; If the splitting rule matching the target symbol in the cell is a second type of rule, replacing the target symbol in the cell with a preset splitting symbol; If the splitting rule matching the target symbol in the cell is a third type rule, the target symbol in the cell is retained.
5. The table reorganization method according to claim 4, characterized in that: After replacing the target symbol in the cell with a preset segmentation symbol and before splitting the data in the transformed cell, the method further includes: Determine the specific segment contained in the context of the target character in the replaced cell and / or the data type to which the key character belongs; If the data type of the context of the target character in the replaced cell is the same, then restore the target character in the replaced cell; If the data type of the context of the target character in the replaced cell is different, the target character in the replaced cell is retained.
6. The table reorganization method according to any one of claims 1 to 5, characterized in that: After identifying an identifier between data of at least one cell and before using the identifier in the cell that matches the splitting rule as a target symbol, the method further includes: For each cell, determining similar fields in the data of the cell; Using different processing schemes to process similar fields in the data of the cells so that there are no similar fields in the processed data of the cells; wherein the processing schemes include deleting similar fields in the data of the cells and replacing similar fields in the data of the cells with preset symbols; Before splitting the data in the transformed cells, the method further includes: Compare the data of the transformed cell with the data of the cell before the transformation; If a field is deleted in the data of the transformed cell, the deleted field will be restored to the data of the transformed cell; If the data in the transformed cell contains a preset symbol, the field corresponding to the preset symbol is restored to the data in the transformed cell.
7. The table reorganization method according to any one of claims 1 to 5, characterized in that: Generate multiple new cells based on the split data, including: For each split data, use the split data as the content of a new cell; or Display the reorganized table template corresponding to the data type of the original table to the user, receive the user's selection instruction, combine the split data according to the table name corresponding to the reorganized table template in the selection instruction, and use the combined content as the data in the new cell.
8. A table reorganization device, characterized in that: include: An extraction module, configured to extract data of at least one cell belonging to a target name from an original table directly exported from a database; an identification module, configured to identify identifiers between data of at least one cell, and for each identifier of each cell, use the identifier in the cell that matches a splitting rule as a target symbol; wherein the identifier is a punctuation symbol contained in the data between the cells that matches multiple splitting rules corresponding to the table type to which the original table belongs; a splitting module for transforming, for each target symbol in each cell, the target symbol in the cell according to a splitting rule that matches the target symbol in the cell, so that the transformed cell uses the same punctuation mark to distinguish data of different data types; and splitting the data in the transformed cell; The reorganization module is used to generate multiple new cells according to the multiple split data, and form a reorganized table based on the multiple new cells.
9. An electronic device, characterized in that: include: Memory, used to store computer programs or instructions; A processor is configured to execute the computer program or instructions in the memory so that the table reorganization method according to any one of claims 1 to 7 is performed.
10. A computer-readable storage medium, characterized in that When the instructions in the storage medium are executed by a processor, the processor is enabled to perform the table reorganization method according to any one of claims 1 to 7.