Data model splitting method and device, data query method and device, electronic equipment and storage medium

By performing row-level splitting of the data table and distinguishing final and non-final data based on the identification field, the problem of data redundancy and incremental stock data in the prior art is solved, and faster query response time and better data storage strategy are achieved.

CN119938790APending Publication Date: 2025-05-06中国邮政储蓄银行股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510025869.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The splitting method of data tables in the prior art has problems such as data redundancy and failure to distinguish incremental and stock data, resulting in an extended query response time.

Method used

By determining the identification field according to the preset data model, the current data table that stores the historical data is divided into a second data table that stores the final data and a third data table that stores the non-final data, and the current latest full data is stored in the first data table, divided and allocated to the second and third data tables.

Benefits of technology

It effectively reduces redundant data, distinguishes incremental and stock data, improves query response speed, and optimizes data storage strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938790A_ABST
    Figure CN119938790A_ABST
Patent Text Reader

Abstract

The invention discloses a data model splitting method and device, a data query method and device, electronic equipment and a storage medium, the method comprises the following steps: determining an identification field according to a preset data model, the identification field being used for representing final state data and non-final state data; according to the identification field, a current data table for storing historical data is divided into a second data table for storing final-state data and a third data table for storing non-final-state data, and the fields in the second data table and the third data table are the same; the current latest total data is stored in a first data table, the data in the first data table is divided and distributed to the second data table and the third data table, and the fields in the first data table are the same as the fields in the second data table and the fields in the third data table. By means of the brand-new method for splitting the data table, the data query efficiency can be effectively improved, and the data requirement of the response service can be quickly read. The invention further provides a matched data query method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data model optimization processing, and in particular to a data model splitting method, a data query method, a device and an electronic device, and a storage medium. Background Art

[0002] As the financial industry attaches great importance to digital transformation, the use scenarios of data are constantly expanding, and the frequency of data development and query analysis is increasing. Whether it is a data warehouse system or a relational database system, business data will inevitably be processed and organized to provide support for query, analysis and deep integration.

[0003] In the related technology, there are many problems with the way of splitting data tables, such as data redundancy and failure to distinguish between incremental and existing data. Summary of the invention

[0004] The embodiments of the present application provide a data model splitting method, a data query method, a device, an electronic device, and a storage medium to optimize data model splitting and speed up query response time.

[0005] The present application embodiment adopts the following technical solutions:

[0006] In a first aspect, an embodiment of the present application provides a data model splitting method, wherein the splitting method includes:

[0007] Determine an identification field according to a preset data model, where the identification field is used to represent final state data and non-final state data;

[0008] According to the identification field, the current data table storing the historical data is divided into a second data table storing final data and a third data table storing non-final data, wherein the fields in the second data table are the same as those in the third data table;

[0009] The latest full data is stored in the first data table, and the data in the first data table is divided and allocated to the second data table and the third data table. The fields in the first data table are the same as those in the second data table and the third data table.

[0010] In some embodiments, the dividing of the current data table storing historical data into a second data table storing final data and a third data table storing non-final data according to the identification field includes:

[0011] The historical data table is split at the row level, and the data in the historical data table is divided into two types: final state data and non-final state data. The second data table and the third data table are used for separate table storage respectively, and different storage strategies are adopted for the data after the separate tables. The final state data is stored in an incremental flow manner, and the non-final state data is stored in a snapshot manner.

[0012] In some embodiments, storing the latest full data in the first data table, dividing the data in the first data table, and allocating the data to the second data table and the third data table includes:

[0013] All data including final and non-final data up to the present is stored in the first data table. The first data table has the same field content and number of fields as the second data table and the third data table, but different data volumes and data storage formats.

[0014] In some embodiments, storing the latest full data in the first data table, dividing the data in the first data table, and allocating the data to the second data table and the third data table includes:

[0015] Initialize the historical data in the current data table, and obtain all historical snapshot data by processing the full daily data;

[0016] If the original data table lacks a final state data identification field, then add the identification field and update the historical data; if the original data table has a data retention period, then obtain the historical snapshot data within the period;

[0017] Select the full amount of data on the latest snapshot date, filter out the final state data up to the current time through the identification field of the final state data, and store the final state data in the second data table;

[0018] Filter out the final state data of the historical snapshots and delete them, and store the remaining non-final state data of the historical snapshots in the third data table;

[0019] The first data table retains the latest full data and records the latest status as a full table; the second data table retains the unchanged final state data and records the final state as a flow table; the third data table retains all historical non-final state data as a snapshot table.

[0020] In some embodiments, the determining of the identification field according to the preset data model, wherein the identification field is used to characterize the final state data and the non-final state data, includes:

[0021] Declare the master data of the data table and identify the master data information of the data table through the data model;

[0022] Determine whether the data table where the master data information is located includes an identification field for representing the final state;

[0023] The data table where the master data information is located includes an identification field for representing the final state to form a data table corresponding to the preset data model after conversion.

[0024] In a second aspect, an embodiment of the present application further provides a data query method, wherein the split data table obtained by the data model splitting method described in the first aspect is applied, and the query method includes:

[0025] A query based on the current full amount of data, used to query in the first data table;

[0026] A query based on the full final state data of a certain historical day, used for querying in the second data table;

[0027] A query based on the full amount of non-final data of a certain historical day, used for querying in the third data table;

[0028] A query based on the full amount of data on a certain historical day, used to perform data query in combination with the second data table and the third data table;

[0029] A query based on the incremental final state data of a certain day in history, used to query the second data table;

[0030] Based on the incremental data of a certain day in history, it is used to perform data query in combination with the second data table and the third data table.

[0031] In some embodiments, the method further comprises: using the front-end interface SQL query statement mapping to actually execute the SQL query statement,

[0032] The front-end interface SQL query statement includes fields representing query methods for four types of data: full final data of a certain historical day, full non-final data of a certain historical day, incremental final data of a certain historical day, and incremental data of a certain historical day;

[0033] The front-end interface query SQL statement also includes: a data query method for indicating the acquisition of the full amount of data up to the current day or a certain historical day, and a data query date for distinguishing the current day or a certain historical day.

[0034] In a third aspect, an embodiment of the present application further provides a data table splitting device, wherein the splitting device comprises:

[0035] An identification module, used to determine an identification field according to a preset data table model, wherein the identification field is used to represent final state data and non-final state data;

[0036] an initialization module, configured to divide the current data table storing historical data into a second data table storing final data and a third data table storing non-final data according to the identification field, wherein the fields in the second data table are the same as those in the third data table;

[0037] The data allocation module is used to store the latest full data in the first data table, divide the data in the first data table, and allocate it to the second data table and the third data table. The fields in the first data table are the same as those in the second data table and the third data table.

[0038] In a fourth aspect, an embodiment of the present application further provides an electronic device, comprising: a processor; and a memory arranged to store computer executable instructions, wherein the executable instructions, when executed, cause the processor to perform the above method.

[0039] In a fifth aspect, an embodiment of the present application further provides a computer-readable storage medium, which stores one or more programs. When the one or more programs are executed by an electronic device including multiple application programs, the electronic device executes the above method.

[0040] At least one of the above technical solutions adopted in the embodiment of the present application can achieve the following beneficial effects: according to the preset data model, determine the identification field. The identification field is used to characterize the final state data and the non-final state data. Then, according to the identification field, the current data table storing the historical data is divided into a second data table storing the final state data and a third data table storing the non-final state data. The fields in the second data table and the third data table are the same. Finally, the latest full data is stored in the first data table, the data in the first data table is divided, and allocated to the second data table and the third data table, and the fields in the first data table are the same as those in the second data table and the third data table. Through the above method, the data table is row-level split in combination with the business characteristics of the data connotation, the data in the data table is divided into two types of data, the final state and the non-final state, and the data is stored in a separate table, and different storage strategies are adopted for the data after the table is divided. The final state data is stored in an incremental flow mode, and the non-final state data is stored in a snapshot mode. For the latest data, that is, the full amount of data including the final state and the non-final state data up to the current time, a separate full amount storage is performed. Finally, a data model of "multiple data tables" is formed with the same field content and number of fields, but different data volume and data storage format. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0042] Figure 1 This is a flow chart of a data model splitting method in an embodiment of the present application;

[0043] Figure 2 This is a schematic diagram of query mapping results of the data model splitting method in an embodiment of the present application;

[0044] Figure 3 A flowchart of a query method for a data model splitting method in an embodiment of the present application;

[0045] Figure 4 This is a structural diagram of a data model splitting device in an embodiment of the present application;

[0046] Figure 5 It is a schematic diagram of the structure of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION

[0047] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0048] In the field of data application, it is common to integrate data from different database tables into one data table. This wide-format data table has a huge structure. As a relatively important data supply form for data analysis systems, it usually stores historical data of a longer date, and rarely uses a storage method with high maintenance costs such as zippers due to complex processing. Therefore, traditional data tables use snapshot storage methods, which inevitably occupy more storage space. At the same time, when querying and analyzing them, excessive scanning of redundant data will consume more computing resources, resulting in low query and analysis efficiency, so it cannot quickly respond to business needs.

[0049] In the related art, the method of splitting a data table is to split the data table on columns, taking the splitting and combination of dimension columns as the entry point, splitting the data table into multiple small tables with different numbers of columns, that is, mainly reducing the overall data volume by increasing or decreasing the dimension columns, so as to achieve the effect of optimizing storage space and computing performance.

[0050] However, the existing data table splitting method is not applicable to all situations. It does not fundamentally solve the problem of redundant storage of existing data, nor can it completely avoid the situation where small tables are interrelated and consume performance. For example, when there are few dimensional data items or the granularity of the dimension itself is relatively fine, this splitting method cannot effectively reduce the amount of data and the actual use effect is not good.

[0051] In addition, although the data table splitting method in the related art reduces the amount of data overall, it does not distinguish between incremental and stock data. The historical stock data that no longer changes is still stored according to the cycle, which fails to fundamentally solve the redundant storage problem of this part of the data. In addition, the splitting of the data table cannot eliminate the problem of mutual correlation between small tables, and the correlation between tables consumes computing resources. If the splitting is unreasonable, it will run counter to the original design of the data table. In addition, the data table splitting method in the related art cannot query the data of all fields from one table, which changes the query method of the user data table to a certain extent, causing inconvenience to the user.

[0052] In view of the above-mentioned deficiencies, the method in the embodiment of the present application does not split the columns of the data table, but combines the intrinsic business characteristics of the data to split the data table horizontally at the row level and use different storage strategies for storage, which completely avoids the association problem of the data table after it is split according to columns and then combined, and also fundamentally solves the redundant storage problem of the final state data of the historical inventory.

[0053] In addition, the present application also provides a new method for querying data in a data table to meet the needs of querying, analyzing or reprocessing in different business scenarios and speed up the calculation. At the same time, SQL statement mapping is used to optimize writing, realize the seamless splitting of data tables, and optimize the user experience.

[0054] The technical solutions provided by various embodiments of the present application are described in detail below in conjunction with the accompanying drawings.

[0055] The present application embodiment provides a data model splitting method, such as Figure 1 As shown, a schematic diagram of the data model splitting method in an embodiment of the present application is provided, and the method at least includes the following steps S110 to S130:

[0056] Based on the fact that business data has a beginning and an end, and combined with various scenarios of data query in practice, the data table data is re-divided and organized, providing a relatively complete new method for splitting data tables.

[0057] Step S110: determining an identification field according to a preset data model, wherein the identification field is used to represent final data and non-final data.

[0058] The data model specifies the structure of the data, the relationship between the data, and the organization of the data. Generally, the specific form of the data model can be presented through one or more data tables. The data model design process will declare the master data of the data table. The master data information of a data table can be identified through the model design document or development script. As shown in Table 4-1, the basic process of processing the target data table is shown. The fifth column indicates whether the source table is the master table of the processing target data table. The last four columns, association table a, association table b, association method, and association condition, indicate the process of processing the fields of the target data table. From this, it can be determined that the master data range is table_a.

[0059] Table 4-1 Partial information output by data table model design

[0060] Wide table name Wide table fields Source table Source Field Is it the main table? Association table a Association table b Association Association conditions t_table_a field_1 table_a src_field_a Y table_a —— —— —— t_table_a field_2 table_b src_field_b N table_a table_b left join ta.id=tb.id ... ... ... ... ... ... ... ... ... t_table_a field_m table_h src_field_m N table_a table_h left join ta.id=th.id

[0061] After determining that the master data is in the "table_a table", the business system survey and data exploration can determine whether this table is an identification field of the final state. This field must be placed in the data table together with other target fields. According to the model design rules, the different database tables of the business system are extracted, cleaned, converted, and associated, and finally a data table with a large number of fields is formed. As shown in Table 4-2, this data table contains M fields and N records.

[0062] Table 4-2 Data table information

[0063] field_1 field_2 field_3 ... field_m a1 b1 c1 ... m1 a2 b2 c2 ... m2 ... ... ... ... ... an bn cn ... mn

[0064] Specifically, taking the data in the field of banking credit as an example, Table 4-3 shows the detailed data information of the bank's loan order, in which the "Is it settled" field is used to distinguish whether the order data is in the final state, that is, to determine whether the order data will change in the future. Of course, other fields can also be used as identification fields for final state data, such as the "Order Status" field, and one or more fields can be used as identification fields for final state data. Among them, loan order number xxxx002 is a settled order, and the data of this order will no longer be updated or changed.

[0065] Table 4-3 Bank loan order data table data

[0066] Data date Loan order number Loan Date Loan Amount Loan balance Repayment amount Repayment amount on the day ... Is it settled? STATS_DT LOAN_ODR_ID DISTR_DT LOAN_AMT LOAN_BAL RPY_AMT DAY_RPY_AMT IS_PAYOFF 2024-5-29 xxxx001 2024-2-23 10000.00 7000.00 3000.00 2000.00 ... no 2024-5-29 xxxx002 2024-1-18 8000.00 0.00 8000.00 5000.00 ... yes 2024-5-29 xxxx005 2024-5-29 12000.00 12000.00 0.00 0.00 ... no ... ... ... ... ... ... ... ...

[0067] Step S120: dividing the current data table storing historical data into a second data table storing final data and a third data table storing non-final data according to the identification field, wherein the fields in the second data table are the same as those in the third data table.

[0068] Through the above method, the data table is not stored in snapshots according to the traditional method, nor is the data table split into multiple small tables with different numbers of fields. The data table is designed as a data model for the data table with consistent fields but different data volumes and data storage strategies.

[0069] First, initialize the historical data of the current data table M (virtual table), process the full data of each destination, and obtain the data of all historical snapshots. If the original data table lacks the final data identification field, you can add such fields and update the historical values ​​of this field. If the data table has a data retention period, you can obtain the historical snapshot data within the period.

[0070] Then, the full amount of data on the latest snapshot date is selected, and the final state data up to the current time is filtered out through the identification field of the final state data, and the data is stored in data table B (second data table).

[0071] Finally, the final state data of the historical snapshot is filtered out and deleted, so that only the non-final state data of the historical snapshot remains, and it is stored in data table C (the third data table), and the fields of data table C are consistent with the fields of data table B.

[0072] In the above process, the traditional data table M (virtual table) is initially split into three data tables with the same fields, namely data table A, data table B and data table C. Data table M is not a real table, but just a name used to establish its connection with the physical data table.

[0073] Taking a data table in the banking credit field as an example, Table 4-4 corresponds to the traditional data table M, and Table 4-5, Table 4-6 and Table 4-7 correspond to Data Table A, Data Table B and Data Table C respectively.

[0074] Table 4-4 Data in Data Table M

[0075] Data date Loan order number Loan Date Loan Amount Loan balance Repayment amount Repayment amount on the day ... Is it settled? STATS_DT LOAN_ODR_ID DISTR_DT LOAN_AMT LOAN_BAL RPY_AMT DAY_RPY_AMT IS_PAYOFF 2024-5-29 xxxx001 2024-2-23 10000.00 7000.00 3000.00 2000.00 ... no 2024-5-29 xxxx002 2024-1-18 8000.00 0.00 8000.00 5000.00 ... yes 2024-5-29 xxxx003 2024-3-12 15000.00 0.00 15000.00 0.00 ... yes 2024-5-29 xxxx004 2024-3-16 2000.00 0.00 2000.00 0.00 ... yes 2024-5-29 xxxx005 2024-5-29 12000.00 12000.00 0.00 0.00 ... no ... ... ... ... ... ... ... ... 2024-5-28 xxxx001 2024-2-23 10000.00 9000.00 1000.00 0.00 ... no 2024-5-28 xxxx002 2024-1-18 8000.00 5000.00 3000.00 3000.00 ... no 2024-5-28 xxxx003 2024-3-12 15000.00 0.00 15000.00 4000.00 ... yes 2024-5-28 xxxx004 2024-3-12 2000.00 0.00 2000.00 0.00 ... yes ... ... ... ... ... ... ... ... 2022-7-18 ... ... ... ... ... ... ...

[0076] Table 4-5 Data in Data Table A

[0077] Data date Loan order number Loan Date Loan Amount Loan balance Repayment amount Repayment amount on the day ... Is it settled? STATS_DT LOAN_ODR_ID DISTR_DT LOAN_AMT LOAN_BAL RPY_AMT DAY_RPY_AMT IS_PAYOFF 2024-5-29 xxxx001 2024-2-23 10000.00 7000.00 3000.00 2000.00 ... no 2024-5-29 xxxx002 2024-1-18 8000.00 0.00 8000.00 5000.00 ... yes 2024-5-29 xxxx003 2024-3-12 15000.00 0.00 15000.00 0.00 ... yes 2024-5-29 xxxx004 2024-3-16 2000.00 0.00 2000.00 0.00 ... yes 2024-5-29 xxxx005 2024-5-29 12000.00 12000.00 0.00 0.00 ... no ... ... ... ... ... ... ... ... 2022-7-18 ... ... ... ... ... ... ...

[0078] Table 4-6 Data in Data Table B

[0079] Data date Loan order number Loan Date Loan Amount Loan balance Repayment amount Repayment amount on the day ... Is it settled? STATS_DT LOAN_ODR_ID DISTR_DT LOAN_AMT LOAN_BAL RPY_AMT DAY_RPY_AMT IS_PAYOFF 2024-5-29 xxxx002 2024-1-18 8000.00 0.00 8000.00 5000.00 ... yes 2024-5-28 xxxx003 2024-3-12 15000.00 0.00 15000.00 4000.00 ... yes 2024-5-28 xxxx004 2024-3-16 2000.00 0.00 2000.00 0.00 ... yes ... ... ... ... ... ... ... ... 2022-7-18 ... ... ... ... ... ... ...

[0080] Table 4-7 Data in Data Table C

[0081] Data date Loan order number Loan Date Loan Amount Loan balance Repayment amount Repayment amount on the day ... Is it settled? STATS_DT LOAN_ODR_ID DISTR_DT LOAN_AMT LOAN_BAL RPY_AMT DAY_RPY_AMT IS_PAYOFF 2024-5-29 xxxx001 2024-2-23 10000.00 7000.00 3000.00 2000.00 ... no 2024-5-29 xxxx005 2024-5-29 12000.00 12000.00 0.00 0.00 ... no 2024-5-28 xxxx001 2024-2-23 10000.00 9000.00 1000.00 0.00 ... no 2024-5-28 xxxx002 2024-1-18 8000.00 5000.00 3000.00 3000.00 ... no ... ... ... ... ... ... ... ... 2022-7-18 ... ... ... ... ... ... ...

[0082] Among them, the STATS_DT field is the data date, indicating the data batch date. The current latest data date is May 29, 2024, and the smallest historical data date is July 18, 2022; the LOAN_ODR_ID field is the loan order number, indicating the unique number of the loan; the DISTR_DT field is the loan date, indicating the actual date of the business; the LOAN_AMT field is the loan amount, indicating the initial amount of the loan; the RPY_AMT field is the repayment amount, indicating the actual amount of the loan repaid by the customer as of now; the LOAN_BAL field is the loan balance, which is the difference between the LOAN_AMT field and the RPY_AMT field, indicating the remaining amount of the loan as of now. If the balance is 0, it means that the loan has been settled; the DAY_RPY_AMT field indicates the repayment amount on the data date, and the IS_PAYOFF field indicates whether this loan has been settled.

[0083] Step S130, storing the latest full data in the first data table, dividing the data in the first data table, and allocating them to the second data table and the third data table, the fields in the first data table are the same as those in the second data table and the third data table.

[0084] After initializing the historical data in the data table, the latest data of the day needs to be processed again so that it can be connected with the historical data.

[0085] According to the original data processing logic, the full amount of data up to the present is generated and stored in data table A (first data table). The fields of data table A (first data table) are consistent with those of data table B (second data table) and data table C (third data table). It should be noted that data table A only retains the latest full amount of data. The final state and non-final state data of data table A (first data table) are filtered out through the identification field of the final state data, and the final state data of data table A is placed in data table B (second data table), and the non-final state data of data table A (first data table) is placed in data table C (third data table). The operation is repeated every day. It should be noted that the daily cycle is for example only and is not used to limit the scope of protection in the embodiments of the present application.

[0086] After initializing the historical data of the data table and distributing the daily data, three data tables with the same fields are formed, but the data status and storage method are different. Data table A retains the latest full data, and data table B (the second data table) retains the unchanged final state data. These two data tables store the latest state and final state records respectively, among which data table A (the first data table) is the full table, and data table B (the second data table) is the flow table. Data table C (the third data table) retains all historical non-final state data, and relatively more data is retained as a snapshot table.

[0087] The data model after splitting by the above method has a significant reduction in the proportion of redundant data, which will correspondingly increase the speed of data screening. Compared with the existing method of splitting the number of columns in the data table, the splitting method in the embodiment of the present application maintains the number of fields in the original data table, which can prevent the situation of consuming computing performance after splitting and then associating, and effectively ensure the efficiency of data query.

[0088] In the embodiment of the present application, the original data table is designed as a "multiple data tables" format, and the data volume and storage strategy of each data table are different, which facilitates technical performance optimization, and technical developers can optimize according to the characteristics of the database.

[0089] In the embodiment of the present application, the data in the data table is classified and reorganized, and the user can choose the most efficient query method according to the characteristics of the actual business scenario. Therefore, the present invention can effectively improve the efficiency of data query and quickly respond to the data needs of the business.

[0090] Through the above method, based on the characteristics of business data that has a beginning and an end, and combined with various scenarios of data query in practice, the data table data is re-divided and organized, providing a new and more complete method for splitting data tables.

[0091] Through the above method, the data table is not stored in snapshots according to the traditional method, nor is the data table split into multiple small tables with different numbers of fields. The data table is designed as a data model for the data table with consistent fields but different data volumes and data storage strategies.

[0092] Different from the related art, when there are fewer dimension data items or the dimensional granularity itself is relatively fine, the splitting method in the related art cannot effectively reduce the amount of data, and the actual use effect is not good. Through the above method, combined with the inherent business characteristics of the data, the data table is split horizontally at the row level and stored using different storage strategies, which completely avoids the association problem of the data table after being split by column and then combined.

[0093] Different from the related technologies, no distinction is made between incremental and stock data, and the historical stock data that no longer changes is still stored according to the cycle, which fails to fundamentally solve the redundant storage problem of this part of the data. Through the above method, based on the characteristics of business data that has a beginning and an end, and combined with various scenarios of data query in practice, the data table data is re-divided and organized, providing a more complete new method for splitting data tables, thereby fundamentally solving the redundant storage problem of the final state data of the historical stock

[0094] Different from the related art, it is impossible to query the data of all fields from one table, which changes the query method of the user data table to a certain extent, causing inconvenience to the user. Through the above method, the method in the embodiment of the present application designs the original data table into the form of "multiple data tables", and the data volume and storage strategy of each data table are different, which facilitates technical performance optimization. Technical developers can optimize according to the characteristics of the database, such as setting primary keys, establishing partitions, etc.

[0095] In one embodiment of the present application, the current data table storing historical data is divided into a second data table storing final state data and a third data table storing non-final state data according to the identification field, including: row-level splitting of the historical data table, dividing the historical data table data into final state data and non-final state data, and using the second data table and the third data table for separate table storage, and adopting different storage strategies for the data after the separate tables, the final state data is stored in an incremental flow manner, and the non-final state data is stored in a snapshot manner.

[0096] The data table is split at the row level based on the business characteristics of the data content. The data in the data table is divided into final and non-final data and stored in separate tables. Different storage strategies are adopted for the data after table segmentation. The final data is stored in an incremental flow manner, and the non-final data is stored in a snapshot manner.

[0097] In one embodiment of the present application, storing the latest full data in the first data table, dividing the data in the first data table, and allocating it to the second data table and the third data table includes: storing the full data including final and non-final data up to the current time in the first data table, the field content and number of fields in the first data table, the second data table and the third data table are the same, but the data volume and data storage format are different.

[0098] For the latest data, that is, the full amount of data up to the current state, including final and non-final data, a separate full storage is performed. Finally, a data model of "multiple data tables" is formed with the same field content and number of fields, but different data volumes and data storage formats.

[0099] In one embodiment of the present application, the storing of the latest full data in the first data table, dividing the data in the first data table, and allocating them to the second data table and the third data table includes: initializing the historical data in the current data table, and obtaining the data of all historical snapshots by processing the daily full data; if the original data table lacks the final data identification field, adding the identification field and updating the historical data, if the original data table has a data retention period, obtaining the historical snapshot data within the period; selecting the full data of the latest snapshot date, filtering out the final data up to the current state through the identification field of the final data and storing it in the second data table; filtering out the final data of the historical snapshots and deleting them, and storing the remaining non-final data of the historical snapshots in the third data table; the first data table retains the latest full data and records the latest status as a full table; the second data table retains the unchanged final data and records the final status as a flow table; the third data table retains all historical non-final data as a snapshot table.

[0100] After initializing the historical data in the data table, the latest data of the day needs to be processed again to make it connected with the historical data. According to the original data processing logic, the full amount of data up to now is generated and stored in data table A (first data). The fields of data table A (first data) are consistent with those of data table B (second data table) and data table C (third data table).

[0101] After initializing the historical data of the data table and distributing the daily data, three data tables with the same fields are formed, but the data status and storage method are different. Data table A (first data) retains the latest full data, and data table B (second data table) retains the unchanged final state data. These two data tables store the latest state and final state records respectively, among which data table A (first data) is a full table, and data table B (second data table) is a flow table. Data table C (third data table) retains all historical non-final state data, retains relatively more data, and is a snapshot table.

[0102] Preferably, the method in the embodiment of the present application re-divides and organizes the data of the existing data table, splitting it into data table A (first data), data table B (second data table) and data table C (third data table). The data volume and storage strategy of these three data tables are different, which provides convenience for technical performance optimization.

[0103] The data date of data table A (first data) is the current date, and only the latest full data is stored. Therefore, there is no need to impose a separate restriction on the date. You can use the method of setting a primary key to improve optimization. Generally, a commonly used and unique business field is selected as the primary key. For example, in Table 4-5, the loan order number can be set as the primary key.

[0104] Data table B (second data table) is a running table, which stores the final state data within the historical data date, and the data is included in data table A. Its main function is to solve the problem of redundant storage of historical final state data. Its main function is to be used in combination with data table C to achieve the purpose of efficiently obtaining the data of a certain day in history. Therefore, the optimization method of partitioning by date should be adopted for data table B.

[0105] Data table C (the third data table) is a snapshot table that contains non-final data of historical dates. The data volume is large. In addition to the need to extract historical date data, users may also need to view specific data records. Therefore, data table C can be optimized by partitioning and setting primary keys.

[0106] In one embodiment of the present application, the identification field is determined according to a preset data model, and the identification field is used to represent final state data and non-final state data, including: declaring the master data of the data table, identifying the master data information of the data table through the data model; judging whether the data table where the master data information is located includes an identification field for representing the final state; and forming the data table corresponding to the preset data model with the identification field for representing the final state in the data table where the master data information is located.

[0107] The data table model design process will declare the master data of the data table. The master data information of a data table can be identified through model design documents or development scripts. After determining that the master data is in a specific table, the identification field of whether this table is the final state can be obtained through business system research and data exploration. This field must be placed in the data table together with other target fields. According to the model design rules, different database tables of the business system are extracted, cleaned, converted and associated, and finally a data table with a large number of fields is formed.

[0108] The data table retrieval method in the embodiment of the present application is different from that of the traditional data table, and is illustrated below by an SQL query statement.

[0109] like Figure 3 As shown, an embodiment of the present application provides a data query method, wherein the query method is applied to the split data table obtained by the data model splitting method, and includes:

[0110] Step S301, based on the query of the current full data, used to query in the first data table; based on the query of the full final data of a certain historical day, used to query in the second data table; based on the query of the full non-final data of a certain historical day, used to query in the third data table; based on the query of the full data of a certain historical day, used to query data in combination with the second data table and the third data table; based on the query of the incremental final data of a certain historical day, used to query in the second data table; based on the incremental data of a certain historical day, used to query data in combination with the second data table and the third data table.

[0111] Based on the split data tables, a supporting data query method is provided in the embodiment of the present application, and the user can select the optimal query method according to the actual business data requirements. The data query method in the embodiment of the application is illustrated below with Table 4-5, Table 4-6 and Table 4-7 corresponding to Data Table A, Data Table B and Data Table C, respectively.

[0112] It should be noted that the "May 28, 2024" mentioned below is only an example in the query SQL and does not have any specific meaning or indication of actual meaning. Those skilled in the art can use any date for query access.

[0113] Query the current full data:

[0114] The query method of the current latest data is basically the same as the query method of the data table stored in the traditional snapshot. The difference is that the corresponding query SQL in the embodiment of the present application does not need to add a date restriction condition, because the data table A only contains the full amount of data of the current date. The SQL statement for querying the current data in the embodiment of the present application is as follows:

[0115] SELECT * FROM table A [WHERE STATS_DT = "current date"];

[0116] The content in square brackets is optional. Adding or not adding it will not affect the data query results, but not adding it will make the query more efficient.

[0117] For example, if you want to obtain the total balance of all loan orders on the current date, the basic data can be obtained using the following query SQL:

[0118] SELECT LOAN_BAL FROM Table A;

[0119] Query of the final state data of the whole amount on a certain day in history:

[0120] The query of the final state data of the whole amount on a certain day in history can be performed through data table B, as shown below:

[0121] SELECT * FROM data table B WHERE STATS_DT <= "date of a certain day in history";

[0122] For example, if you want to obtain the total loan amount of loan orders settled as of May 28, 2024, the basic data can be obtained using the following query SQL:

[0123] SELECT LOAN_AMT FROM tableB WHERE STATS_DT <= "2024-05-28";

[0124] Query of the full amount of non-final data on a certain day in history:

[0125] The query of the full amount of non-final data on a certain day in history can be performed through data table C, as shown below:

[0126] SELECT * FROM data table C WHERE STATS_DT = "date of a certain day in history";

[0127] For example, if you want to obtain the total loan balance of all loan orders as of May 28, 2024, the basic data can be obtained using the following query SQL. Loan orders with a loan balance of 0 are settled orders and do not need to be included in the basic data for calculation. Therefore, the data in data table B can be excluded for calculation.

[0128] SELECT LOAN_BAL FROM Table C WHERE STATS_DT = "2024-05-28";

[0129] Query of full data for a certain day in history:

[0130] To query the full amount of data for a certain day in history, you need to query data in combination with data table B and data table C. For example:

[0131] SELECT * FROM data table B WHERE STATS_DT <= "date of a certain day in history"

[0132] UNION ALL

[0133] SELECT * FROM data table C WHERE STATS_DT = "date of a certain day in history";

[0134] The upper part of UNION ALL obtains the full final data up to a certain day from data table B, and the lower part obtains the non-final data up to a certain day from data table C.

[0135] For example, if you want to obtain the total loan amount of all loan orders as of May 28, 2024, the basic data can be obtained using the following query SQL:

[0136] SELECT LOAN_AMT FROM Table B WHERE STATS_DT <= "2024-05-28" UNION ALL

[0137] SELECT LOAN_AMT FROM tableC WHERE STATS_DT = "2024-05-28";

[0138] Query of incremental final state data for a certain day in history:

[0139] The query of the incremental final state data of a certain day in history can be queried through data table B, as shown below:

[0140] SELECT * FROM data table B WHERE STATS_DT = "date of a certain day in history";

[0141] For example, if you want to obtain the total repayment amount of the loan order settled on May 28, 2024, the basic data can be obtained using the following query SQL:

[0142] SELECT RPY_AMT FROM tableB WHERE STATS_DT = "2024-05-28";

[0143] Incremental data for a certain day in history:

[0144] To query the incremental data of a certain day in history, you need to query data in combination with data table B and data table C. For example:

[0145] SELECT * FROM data table B WHERE STATS_DT = "date of a certain day in history" UNION ALL

[0146] SELECT * FROM data table C WHERE STATS_DT = "date of a certain day in history";

[0147] The upper part of UNION ALL indicates that the incremental final data of a certain day is obtained from data table B, and the lower part indicates that the non-final data up to a certain day is obtained from data table C. The data in data table C is the incremental data that is currently added or changed, or it may be the non-incremental data that has not changed. Therefore, the data selected by the above query SQL includes the incremental data currently required. If the redundant non-incremental data has an impact, it can be eliminated according to the specific business scenario.

[0148] For example, if you want to obtain the repayment amount of a loan order on May 28, 2024, the basic data can be obtained using the following query SQL:

[0149] SELECT DAY_RPY_AMT FROM tableB WHERE STATS_DT = "2024-05-28" UNION ALL

[0150] SELECT DAY_RPY_AMT FROM tableC WHERE STATS_DT = "2024-05-28";

[0151] The repayment amount on May 28, 2024, will only occur in incremental data and non-incremental non-final data, so there is no need to adjust the above SQL.

[0152] In the embodiment of the present application, a supporting data query method is provided, and on this basis, writing optimization is performed to map slightly complex original SQL statements into simpler conventional SQL statements, which are more in line with user habits and enhance user experience.

[0153] Through the above method, a new data query method is provided for the split data table in the embodiment of the present application, including query methods such as full data and incremental data, final data and non-final data, etc., which provides users with a variety of options for extracting or filtering data. In addition, users can not only adopt the general full data acquisition method, but also use the best method to query the corresponding basic data according to the actual business scenario.

[0154] In one embodiment of the present application, the method also includes: using a front-end interface SQL query statement to map the actual execution of the SQL query statement, the front-end interface SQL query statement includes fields representing query methods for four types of data, namely, full final data of a certain historical day, full non-final data of a certain historical day, incremental final data of a certain historical day, and incremental data of a certain historical day. The front-end interface query SQL statement also includes: a data query method for representing the full data up to the current day or up to a certain historical day, and a data query date for distinguishing the current day or a certain historical day.

[0155] The six basic methods of data query provided above are not very convenient to use in practice. For example, the SQL statement written using the UNION ALL syntax is relatively long, the "<=" and "=" operators need to be used in different situations, and the names of data tables A, B, and C also need to be distinguished. Here, the query methods of the above data tables are improved, that is, the optimization of the data table data query methods is provided to be closer to the user's usage habits and improve the convenience of use.

[0156] Considering that this method of establishing SQL views cannot cover all data query methods, the embodiment of the present application adopts the method of mapping the front-end interface SQL statements to the actual execution of SQL statements. The specific mapping method of SQL statements is as follows Figure 2 As shown in , data table M points to data table A, data table B, data table C and their combination when DATA_TYPE and STATS_DT have different values. It should be noted that data table M is not a real table, but a name used to establish its connection with the entity data table. The field names of data table M are consistent with those of data table A, data table B and data table C. DATA_TYPE has four types: F, U, ZF and ZU, which represent the query methods of the four types of data: full final data of a certain historical day, full non-final data of a certain historical day, incremental final data of a certain historical day and incremental data of a certain historical day. When DATA_TYPE does not appear in the SQL statement, it means to use the two data query methods of full data as of the current day or as of a certain historical day, and use STATS_DT to distinguish the data date of the current day or a certain historical day.

[0157] Obviously, after mapping, the SQL statement for extracting the basic data of the data table is more concise and convenient, and high efficiency of data extraction and analysis is achieved while maintaining the user's SQL writing habits.

[0158] like Figure 2 As shown, the SQL statements of sequence number 1 and sequence number 4 are basically the same, only the date value after STATS_DT is different. The basic data screened by this statement can basically meet the needs of users. If users do not want to consider the optimal query method too much, they do not need to consider the four cases of DATA_TYPE, and can directly use this statement to meet their data needs. Therefore, the method in the embodiment of the present application does not force users to use the optimal query method, but provides users with more choices while being compatible with conventional data acquisition methods.

[0159] The method in the embodiment of the present application not only provides a method for splitting a data table, but also combines the data volume, storage strategy and user usage of the data table to provide a basic idea for technical optimization of the data table.

[0160] The method in the embodiment of the present application maps the slightly complicated SQL statements actually executed into relatively simple SQL statements written by front-end users in accordance with the rules and specifications of the SQL language, without changing the user's usage habits, thereby achieving seamless splitting of data tables.

[0161] The method in the embodiment of the present application does not force the user to consider too much about using the optimal query method, but provides the user with more choices while being compatible with conventional data acquisition methods. This is different from other data table splitting methods.

[0162] The present application embodiment also provides a data table splitting device 400, such as Figure 4 As shown, a schematic diagram of the structure of a data table splitting device in an embodiment of the present application is provided, wherein the data table splitting device 400 at least includes: an identification module 410, an initialization module 420, and a data allocation module 430, wherein:

[0163] In one embodiment of the present application, the identification module 410 is specifically used to: determine an identification field according to a preset data model, and the identification field is used to represent final data and non-final data.

[0164] The data model specifies the structure of the data, the relationship between the data, and the organization of the data. Generally, the specific form of the data model can be presented through one or more data tables. The data model design process will declare the master data of the data table. The master data information of a data table can be identified through the model design document or development script. As shown in Table 4-1, the basic process of processing the target data table is shown. The fifth column indicates whether the source table is the master table of the processing target data table. The last four columns, association table a, association table b, association method, and association condition, indicate the process of processing the fields of the target data table. From this, it can be determined that the master data range is table_a.

[0165] Table 4-1 Partial information output by data table model design

[0166] Wide table name Wide table fields Source table Source Field Is it the main table? Association table a Association table b Association Association conditions t_table_a field_1 table_a src_field_a Y table_a —— —— —— t_table_a field_2 table_b src_field_b N table_a table_b left join ta.id=tb.id ... ... ... ... ... ... ... ... ... t_table_a field_m table_h src_field_m N table_a table_h left join ta.id=th.id

[0167] After determining that the master data is in table_a, business system research and data exploration can determine whether this table is an identification field of the final state. This field must be placed in the data table together with other target fields. According to the model design rules, different database tables of the business system are extracted, cleaned, converted, and associated, and finally a data table with a large number of fields is formed. As shown in Table 4-2, this data table contains M fields and N records.

[0168] Table 4-2 Data table information

[0169] field_1 field_2 field_3 ... field_m a1 b1 C1 ... m1 a2 b2 c2 ... m2 ... ... ... ... ... an bn cn ... mn

[0170] Specifically, taking the data in the field of banking credit as an example, Table 4-3 shows the detailed data information of the bank's loan order, in which the "Is it settled" field is used to distinguish whether the order data is in the final state, that is, to determine whether the order data will change in the future. Of course, other fields can also be used as identification fields for final state data, such as the "Order Status" field, and one or more fields can be used as identification fields for final state data. Among them, loan order number xxxx002 is a settled order, and the data of this order will no longer be updated or changed.

[0171] Table 4-3 Bank loan order data table data

[0172] Data date Loan order number Loan Date Loan Amount Loan balance Repayment amount Repayment amount on the day ... Is it settled? STATS_DT LOAN_ODR_ID DISTR_DT LOAN_AMT LOAN_BAL RPY_AMT DAY_RPY_AMT IS_PAYOFF 2024-5-29 xxxx001 2024-2-23 10000.00 7000.00 3000.00 2000.00 ... no 2024-5-29 xxxx002 2024-1-18 8000.00 0.00 8000.00 5000.00 ... yes 2024-5-29 xxxx005 2024-5-29 12000.00 12000.00 0.00 0.00 ... no ... ... ... ... ... ... ... ...

[0173] In one embodiment of the present application, the initialization module 420 is specifically used to: divide the current data table storing historical data into a second data table storing final data and a third data table storing non-final data according to the identification field, and the fields in the second data table are the same as those in the third data table.

[0174] First, initialize the historical data of the current data table M (virtual table), process the full data of each destination, and obtain the data of all historical snapshots. If the original data table lacks the final data identification field, you can add such fields and update the historical values ​​of this field. If the data table has a data retention period, you can obtain the historical snapshot data within the period.

[0175] Then, the full amount of data on the latest snapshot date is selected, and the final state data up to the current time is filtered out through the identification field of the final state data, and the data is stored in data table B (second data table).

[0176] Finally, the final state data of the historical snapshot is filtered out and deleted, so that only the non-final state data of the historical snapshot remains, and it is stored in data table C (the third data table), and the fields of data table C are consistent with the fields of data table B.

[0177] In the above process, the traditional data table M (virtual table) is initially split into three data tables with the same fields, namely data table A, data table B and data table C. Data table M is not a real table, but just a name used to establish its connection with the physical data table.

[0178] Taking a data table in the banking credit field as an example, Table 4-4 corresponds to the traditional data table M, and Table 4-5, Table 4-6 and Table 4-7 correspond to Data Table A, Data Table B and Data Table C respectively.

[0179] Table 4-4 Data in Data Table M

[0180] Data date Loan order number Loan Date Loan Amount Loan balance Repayment amount Repayment amount on the day ... Is it settled? STATS_DT LOAN_ODR_ID DISTR_DT LOAN_AMT LOAN_BAL RPY_AMT DAY_RPY_AMT IS_PAYOFF 2024-5-29 xxxx001 2024-2-23 10000.00 7000.00 3000.00 2000.00 ... no 2024-5-29 xxxx002 2024-1-18 8000.00 0.00 8000.00 5000.00 ... yes 2024-5-29 xxxx003 2024-3-12 15000.00 0.00 15000.00 0.00 ... yes 2024-5-29 xxxx004 2024-3-16 2000.00 0.00 2000.00 0.00 ... yes 2024-5-29 xxxx005 2024-5-29 12000.00 12000.00 0.00 0.00 ... no ... ... ... ... ... ... ... ... 2024-5-28 xxxx001 2024-2-23 10000.00 9000.00 1000.00 0.00 ... no 2024-5-28 xxxx002 2024-1-18 8000.00 5000.00 3000.00 3000.00 ... no 2024-5-28 xxxx003 2024-3-12 15000.00 0.00 15000.00 4000.00 ... yes 2024-5-28 xxxx004 2024-3-12 2000.00 0.00 2000.00 0.00 ... ... ... ... ... ... ... ... ... 2022-7-18 ... ... ... ... ... ... ...

[0181] Table 4-5 Data in Data Table A

[0182] Data date Loan order number Loan Date Loan Amount Loan balance Repayment amount Repayment amount on the day ... Is it settled? STATS_DT LOAN_ODR_ID DISTR_DT LOAN_AMT LOAN_BAL RPY_AMT DAY_RPY_AMT IS_PAYOFF 2024-5-29 xxxx001 2024-2-23 10000.00 7000.00 3000.00 2000.00 ... no 2024-5-29 xxxx002 2024-1-18 8000.00 0.00 8000.00 5000.00 ... yes 2024-5-29 xxxx003 2024-3-12 15000.00 0.00 15000.00 0.00 ... yes 2024-5-29 xxxx004 2024-3-16 2000.00 0.00 2000.00 0.00 ... yes 2024-5-29 xxxx005 2024-5-29 12000.00 12000.00 0.00 0.00 ... no ... ... ... ... ... ... ... ... 2022-7-18 ... ... ... ... ... ... ...

[0183] Table 4-6 Data in Data Table B

[0184] Data date Loan order number Loan Date Loan Amount Loan balance Repayment amount Repayment amount on the day ... Is it settled? STATS_DT LOAN_ODR_ID DISTR_DT LOAN_AMT LOAN_BAL RPY_AMT DAY_RPY_AMT IS_PAYOFF 2024-5-29 xxxx002 2024-1-18 8000.00 0.00 8000.00 5000.00 ... yes 2024-5-28 xxxx003 2024-3-12 15000.00 0.00 15000.00 4000.00 ... yes 2024-5-28 xxxx004 2024-3-16 2000.00 0.00 2000.00 0.00 ... yes ... ... ... ... ... ... ... ... 2022-7-18 ... ... ... ... ... ... ...

[0185] Table 4-7 Data in Data Table C

[0186] Data date Loan order number Loan Date Loan Amount Loan balance Repayment amount Repayment amount on the day ... Is it settled? STATS_DT LOAN_ODR_ID DISTR_DT LOAN_AMT LOAN_BAL RPY_AMT DAY_RPY_AMT IS_PAYOFF 2024-5-29 xxxx001 2024-2-23 10000.00 7000.00 3000.00 2000.00 ... no 2024-5-29 xxxx005 2024-5-29 12000.00 12000.00 0.00 0.00 ... no 2024-5-28 xxxx001 2024-2-23 10000.00 9000.00 1000.00 0.00 ... no 2024-5-28 xxxx002 2024-1-18 8000.00 5000.00 3000.00 3000.00 ... no ... ... ... ... ... ... ... ... 2022-7-18 ... ... ... ... ... ... ...

[0187] Among them, the STATS_DT field is the data date, indicating the data batch date. The current latest data date is May 29, 2024, and the smallest historical data date is July 18, 2022; the LOAN_ODR_ID field is the loan order number, indicating the unique number of the loan; the DISTR_DT field is the loan date, indicating the actual date of the business; the LOAN_AMT field is the loan amount, indicating the initial amount of the loan; the RPY_AMT field is the repayment amount, indicating the actual amount of the loan repaid by the customer as of now; the LOAN_BAL field is the loan balance, which is the difference between the LOAN_AMT field and the RPY_AMT field, indicating the remaining amount of the loan as of now. If the balance is 0, it means that the loan has been settled; the DAY_RPY_AMT field indicates the repayment amount on the data date, and the IS_PAYOFF field indicates whether this loan has been settled.

[0188] In one embodiment of the present application, the data allocation module 430 is specifically used to: store the current latest full data in the first data table, divide the data in the first data table, and allocate it to the second data table and the third data table. The fields in the first data table are the same as those in the second data table and the third data table.

[0189] After initializing the historical data in the data table, the latest data of the day needs to be processed again so that it can be connected with the historical data.

[0190] According to the original data processing logic, the full amount of data up to the present is generated and stored in data table A (first data table). The fields of data table A (first data table) are consistent with those of data table B (second data table) and data table C (third data table). It should be noted that data table A only retains the latest full amount of data. The final state and non-final state data of data table A (first data table) are filtered out through the identification field of the final state data, and the final state data of data table A is placed in data table B (second data table), and the non-final state data of data table A (first data table) is placed in data table C (third data table). The operation is repeated every day. It should be noted that the daily cycle is for example only and is not used to limit the scope of protection in the embodiments of the present application.

[0191] After initializing the historical data of the data table and distributing the daily data, three data tables with the same fields are formed, but the data status and storage method are different. Data table A retains the latest full data, and data table B (the second data table) retains the unchanged final state data. These two data tables store the latest state and final state records respectively, among which data table A (the first data table) is the full table, and data table B (the second data table) is the flow table. Data table C (the third data table) retains all historical non-final state data, and relatively more data is retained as a snapshot table.

[0192] It can be understood that the above-mentioned data table splitting device can implement each step of the data table splitting method provided in the above-mentioned embodiment, and the relevant explanations about the data table splitting method are applicable to the data table splitting device, which will not be repeated here.

[0193] Figure 5 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present application. Figure 5 At the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. The memory may include a memory, such as a high-speed random access memory (RAM), and may also include a non-volatile memory (non-volatile memory), such as at least one disk storage. Of course, the electronic device may also include hardware required for other services.

[0194] The processor, network interface and memory can be interconnected through an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0195] The memory is used to store the program. Specifically, the program may include a program code, and the program code includes a computer operation instruction. The memory may include a memory and a non-volatile memory, and provides instructions and data to the processor.

[0196] The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, forming a data table splitting device at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations:

[0197] Determine an identification field according to a preset data model, where the identification field is used to represent final state data and non-final state data;

[0198] According to the identification field, the current data table storing the historical data is divided into a second data table storing final data and a third data table storing non-final data, wherein the fields in the second data table are the same as those in the third data table;

[0199] The latest full data is stored in the first data table, and the data in the first data table is divided and allocated to the second data table and the third data table. The fields in the first data table are the same as those in the second data table and the third data table.

[0200] The above application Figure 1 The method performed by the data table splitting device disclosed in the illustrated embodiment can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor or an instruction in the form of software. The above processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as a hardware decoding processor for execution, or a combination of hardware and software modules in the decoding processor for execution. The software module can be located in a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.

[0201] The electronic device may also perform Figure 1 The method executed by the data table splitting device in the embodiment of the present invention is realized by placing the data table splitting device in the Figure 1 The functions of the illustrated embodiment will not be described in detail in the embodiments of the present application.

[0202] The present application also provides a computer-readable storage medium, which stores one or more programs, wherein the one or more programs include instructions, which, when executed by an electronic device including multiple application programs, enable the electronic device to execute Figure 1 The method performed by the data table splitting device in the illustrated embodiment is specifically used to perform:

[0203] Determine an identification field according to a preset data model, where the identification field is used to represent final state data and non-final state data;

[0204] According to the identification field, the current data table storing the historical data is divided into a second data table storing final data and a third data table storing non-final data, wherein the fields in the second data table are the same as those in the third data table;

[0205] The latest full data is stored in the first data table, and the data in the first data table is divided and allocated to the second data table and the third data table. The fields in the first data table are the same as those in the second data table and the third data table.

[0206] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0207] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0208] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0209] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0210] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0211] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0212] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0213] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0214] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0215] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A data model splitting method, wherein: The splitting method comprises: Determine an identification field according to a preset data model, where the identification field is used to represent final state data and non-final state data; According to the identification field, the current data table storing the historical data is divided into a second data table storing final data and a third data table storing non-final data, wherein the fields in the second data table are the same as those in the third data table; The latest full data is stored in the first data table, and the data in the first data table is divided and allocated to the second data table and the third data table. The fields in the first data table are the same as those in the second data table and the third data table.

2. The method of claim 1, wherein: The method of dividing the current data table storing historical data into a second data table storing final data and a third data table storing non-final data according to the identification field includes: The historical data table is split at the row level, and the data in the historical data table is divided into two types: final state data and non-final state data. The second data table and the third data table are used for separate table storage respectively, and different storage strategies are adopted for the data after the separate tables. The final state data is stored in an incremental flow manner, and the non-final state data is stored in a snapshot manner.

3. The method of claim 2, wherein: The storing the latest full data in the first data table, dividing the data in the first data table, and distributing the data to the second data table and the third data table includes: All data including final and non-final data up to the present is stored in the first data table. The first data table has the same field content and number of fields as the second data table and the third data table, but different data volumes and data storage formats.

4. The method of claim 1, wherein: The storing the latest full data in the first data table, dividing the data in the first data table, and distributing the data to the second data table and the third data table includes: Initialize the historical data in the current data table, and obtain all historical snapshot data by processing the full daily data; If the original data table lacks the final data identification field, then add the identification field and update the historical data; if the original data table has a data retention period, then obtain the historical snapshot data within the period; Select the full amount of data on the latest snapshot date, filter out the final state data up to the current time through the identification field of the final state data, and store the final state data in the second data table; Filter out the final state data of the historical snapshots and delete them, and store the remaining non-final state data of the historical snapshots in the third data table; The first data table retains the latest full data and records the latest status as a full table; the second data table retains the unchanged final state data and records the final state as a flow table; the third data table retains all historical non-final state data as a snapshot table.

5. The method of claim 1, wherein: The step of determining an identification field according to a preset data model, wherein the identification field is used to characterize final state data and non-final state data, includes: Declare the master data of the data table and identify the master data information of the data table through the data model; Determine whether the data table where the master data information is located includes an identification field for representing the final state; The data table where the master data information is located includes an identification field for representing the final state to form a data table corresponding to the preset data model after conversion.

6. A data query method, wherein: The query method applied to the split data table obtained by the data model splitting method according to any one of claims 1 to 5 comprises: A query based on the current full amount of data, used to query in the first data table; A query based on the full final state data of a certain historical day, used for querying in the second data table; A query based on the full amount of non-final data of a certain historical day, used for querying in the third data table; A query based on the full amount of data on a certain historical day, used to perform data query in combination with the second data table and the third data table; A query based on the incremental final state data of a certain day in history, used to query the second data table; Based on the incremental data of a certain day in history, it is used to perform data query in combination with the second data table and the third data table.

7. The method according to claim 6, further comprising: Use the front-end interface SQL query statement to map the actual execution of SQL query statements. The front-end interface SQL query statement includes fields representing query methods for four types of data: full final data of a certain historical day, full non-final data of a certain historical day, incremental final data of a certain historical day, and incremental data of a certain historical day; The front-end interface query SQL statement also includes: a data query method for indicating the acquisition of the full amount of data up to the current day or a certain historical day, and a data query date for distinguishing the current day or a certain historical day.

8. A data table splitting device, wherein: The splitting device comprises: An identification module, used to determine an identification field according to a preset data table model, wherein the identification field is used to represent final state data and non-final state data; an initialization module, configured to divide the current data table storing historical data into a second data table storing final data and a third data table storing non-final data according to the identification field, wherein the fields in the second data table are the same as those in the third data table; The data allocation module is used to store the latest full data in the first data table, divide the data in the first data table, and allocate it to the second data table and the third data table. The fields in the first data table are the same as those in the second data table and the third data table.

9. An electronic device, comprising: processor; as well as A memory arranged to store computer executable instructions, which when executed cause the processor to perform the method of any one of claims 1 to 7.

10. A computer-readable storage medium storing one or more programs, which, when executed by an electronic device including a plurality of application programs, causes the electronic device to execute any one of the methods of claims 1 to 7.