Data processing method and device, electronic equipment and storage medium

By obtaining the maintenance date set of historical data tables, analyzing and generating structured query statements, filtering target data from historical data tables and supplementing them into the current data table partition, the problem of manual repair of the data consistency and accuracy of the current data table and historical data tables in the data warehouse is solved, and the burden of automated processing and operation and maintenance is reduced.

CN120353869AActive Publication Date: 2025-07-22BANK OF NINGBO
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510846055.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

The data consistency and accuracy of the current data table and historical data table in the data warehouse need to be manually repaired, resulting in an increase in the operation and maintenance burden, and the existing technology cannot realize automated processing.

Method used

By obtaining the maintenance date set of historical data tables, analyzing the table structure information of the current data table, dividing the field types, and generating structured query statements, traversing the maintenance date set to filter the target data from the historical data table, and filling in the corresponding data partition of the current data table according to the creation date.

Benefits of technology

It realizes automated data consistency repair of data warehouses, reduces operation and maintenance burden, and improves data processing efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353869A_ABST
    Figure CN120353869A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device, electronic equipment and a storage medium, and is applied to the technical field of computers, the method comprises the steps that a maintenance date set of a historical data table is acquired, and the maintenance date set is composed of non-repeated creation dates of maintained data in the historical data table; analyzing table structure information of the current data table to obtain fields and field information in real time, and dividing the fields into public fields, unique fields and partition fields according to the field information and predefined field classification rules; generating an SQL (Structured Query Language) based on the public field, the unique field, the partition field and a pre-configured transcoding field mapping table; traversing each creation date in the maintenance date set, and executing SQL (Structured Query Language) to screen target data from the historical data table; and supplementing the screened target data into the data partition corresponding to the current data table according to the creation date. According to the scheme, the problem that the data consistency and accuracy of the current data table and the historical data table need to be manually repaired after the historical data table is maintained in the source system of the data warehouse is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a data processing method, apparatus, electronic device, and storage medium. Background Art

[0002] In various business systems, as the core data source of a data warehouse, source data has a high complexity of business scenarios and a large volume of data. Thus, the source system frequently performs manual data maintenance operations. In such business scenarios, streaming data usually has a high frequency of incremental data warehousing. Considering the data storage pressure, the source system moves historical data into an archive table for storage, and only retains active data in the current table for storage, physically separating the current database and the historical database at the source system architecture level.

[0003] For a data warehouse, limited by data extraction performance, the current data table (hereinafter referred to as the current table) and the historical data table (hereinafter referred to as the historical table) of the source system are usually incrementally warehoused according to different time dimensions: the current table is partitioned and warehoused according to the data creation date, and the historical table is partitioned and warehoused according to the archive date. When the source system manually maintains archived historical data records (such as data value correction), since the records to be maintained have been migrated to the historical database, the source system only needs to operate on the historical data table. However, on the data warehouse side, the current table is partitioned and stored according to the creation date and retains all historical data. For the current table data of the historical version that has been warehoused, additional manual maintenance is required. Therefore, when the source system maintains the historical data table, in addition to maintaining the historical data table, the data warehouse layer also needs to manually repair the data consistency problem of the current data table.

[0004] Since the data records to be maintained have been archived, when directly extracting partition data of a specific creation date from the current database of the source system, the archived historical records will be missing, resulting in the current data table having to be repaired through manual intervention and unable to be automatically processed through a scheduling system. This operation mode will significantly increase the operation and maintenance burden of the data warehouse, and at the same time, frequent manual maintenance also poses a severe challenge to the consistency guarantee of cross-data tables. Summary of the Invention

[0005] The present disclosure provides a data processing method, apparatus, electronic device, and storage medium to at least solve the above technical problems existing in the prior art.

[0006] According to a first aspect of the present disclosure, there is provided a data processing method, the method including: Obtaining a set of maintenance dates of a historical data table, the set of maintenance dates being composed of non-repeating creation dates of the data to be maintained in the historical data table; Parse the table structure information of the current data table to obtain fields and field information in real time. According to the field information and predefined field classification rules, divide the fields into common fields, unique fields, and partition fields; Generate a structured query statement based on the common fields, unique fields, partition fields, and pre-configured transcoding field mapping table; Traverse each creation date in the maintenance date set, and execute the structured query statement to filter target data from the historical data table; Supplement the filtered target data into the corresponding data partitions of the current data table according to the creation date.

[0007] In an implementable manner, the obtaining the maintenance date set of the historical data table includes: Extract the update time of the historical data table; Execute a date query operation to filter non-duplicate creation dates whose update time is equal to the specified maintenance date and whose creation date is not equal to the specified maintenance date.

[0008] In an implementable manner, the field information includes field name, field type, and field comment; The dividing the fields into common fields, unique fields, and partition fields according to the field information and predefined field classification rules includes: Classify fields whose field names belong to the predefined partition identifier set into partition fields. The predefined partition identifier set includes a data processing task identifier field and a date partition identifier field; Classify fields whose field names contain a predefined source identifier into unique fields; Classify the remaining fields into common fields.

[0009] In an implementable manner, the generating a structured query statement based on the common fields, unique fields, partition fields, and pre-configured transcoding field mapping table includes: Generate a first expression for the common fields: replace the null values in the numeric fields with a preset value and retain the original field type, and directly reference the original field name for non-numeric fields; Generate a second expression for the unique fields: assign a predefined system code to the system source identifier field, assign a fixed table name to the data table identifier field, and assign the current system time to the timestamp field; Generate a third expression for the partition fields: assign the first date format of the current creation date to the date partition field; assign a predefined task name constant value to the task identifier partition field; Identify the transcoding fields specified in the pre-configured transcoding field mapping table and generate the fourth expression: construct the standard code value table association logic with the source system identifier and the encoding type identifier as the association conditions; generate the code value matching result expression, and output a string composed of a delimiter, the original field value, and an exception identifier when the matching fails; directly reference the original field name for non-transcoding fields; Integrate the first expression, the second expression, the third expression, and the fourth expression to obtain the structured query statement.

[0010] In an implementable manner, traversing each creation date in the maintenance date set and executing the structured query statement to filter target data from the historical data table includes: Convert the creation date into a first date format and a second date format; Inject the first date format into the structured query statement and execute the query operation to obtain target data, including: Select target data from the current data table where the partition date is equal to the second date format; Select target data from the historical data table that meets the following conditions: the creation date is equal to the first date format, and the partition date is not less than the first date format.

[0011] In an implementable manner, supplementing the filtered target data into the corresponding data partition of the current data table according to the creation date includes: Inject the first date format and the predefined task identifier constant into the third expression; Based on the third expression, batch-write the target data into the data partition of the corresponding creation date of the current data table.

[0012] According to the second aspect of the present disclosure, a data processing device is provided, and the device includes: An acquisition module for acquiring the maintenance date set of the historical data table, where the maintenance date set is composed of non-repeated creation dates of the data maintained in the historical data table; An analysis module for analyzing the table structure information of the current data table to obtain fields and field information in real time, and dividing the fields into common fields, unique fields, and partition fields according to the field information and the predefined field classification rules; A generation module for generating a structured query statement based on the common fields, unique fields, partition fields, and the pre-configured transcoding field mapping table; An execution module for traversing each creation date in the maintenance date set and executing the structured query statement to filter target data from the historical data table; An insertion module for supplementing the filtered target data into the corresponding data partitions of the current data table according to the creation date.

[0013] In an implementable embodiment, the acquisition module is specifically configured to: extract the update time of the historical data table; perform a date query operation to filter non-duplicate creation dates whose update time is equal to the specified maintenance date and whose creation date is not equal to the specified maintenance date.

[0014] According to a third aspect of the present disclosure, there is provided an electronic device, including: At least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in the present disclosure.

[0015] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method described in the present disclosure.

[0016] The data processing method, device, electronic device and storage medium of the present disclosure, the method includes: obtaining a set of maintenance dates of a historical data table, the set of maintenance dates being composed of non-duplicate creation dates of the data to be maintained in the historical data table; parsing the table structure information of the current data table to obtain fields and field information, and dividing the fields into common fields, unique fields and partition fields according to the field information and a predefined field classification rule; generating a structured query statement based on the common fields, unique fields, partition fields and a pre-configured transcoding field mapping table; traversing each creation date in the set of maintenance dates, and executing the structured query statement to filter target data from the historical data table; inserting the filtered target data into the corresponding data partitions of the current data table according to the creation date. This solution effectively solves the problem that after the source system maintains the historical data table in the data warehouse, the data consistency and accuracy between the current data table and the historical data table need to be manually repaired. Through an automated process, the situation of missing archived historical records due to directly extracting data from the current library of the source system is avoided, realizing automated processing, significantly reducing the operation and maintenance burden of the data warehouse, ensuring cross-data table data consistency, and improving data processing efficiency and accuracy.

[0017] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become readily understood. In the drawings, several embodiments of the present disclosure are shown by way of illustration and not limitation, wherein: In the drawings, the same or corresponding reference numerals denote the same or corresponding parts.

[0019] Figure 1 The implementation process of the data processing method according to the embodiment of the present disclosure is schematically shown Figure 1 ; Figure 2 The implementation process of the data processing method according to the embodiment of the present disclosure is schematically shown Figure 2 ; Figure 3 The implementation process of the data processing method according to the embodiment of the present disclosure is schematically shown Figure 3 ; Figure 4 The implementation process of the data processing method according to the embodiment of the present disclosure is schematically shown Figure 4 ; Figure 5 The implementation process of the data processing method according to the embodiment of the present disclosure is schematically shown Figure 5 ; Figure 6 The implementation process of the data processing method according to the embodiment of the present disclosure is schematically shown Figure 6 ; Figure 7 The structural schematic diagram of the data processing device according to the embodiment of the present disclosure is shown; Figure 8 The composition structural schematic diagram of an electronic device according to the embodiment of the present disclosure is shown. Detailed implementation manners

[0020] To make the objectives, features, and advantages of the present disclosure more obvious and understandable, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present disclosure.

[0021] The present disclosure provides a data processing method, as Figure 1 shown, the method includes: Step 101: Obtain a set of maintenance dates of the historical data table, where the set of maintenance dates is composed of non-repeated creation dates of the data to be maintained in the historical data table.

[0022] In this example, when the source system maintains data in the historical data table (His table), the non-duplicate creation dates being maintained are automatically identified by executing an SQL query (SELECT DISTINCT create_time FROM His table WHERE update_time ='maintenance date' AND create_time <> 'maintenance date'). These dates form the maintenance date set (DATA2), which is used to locate the partitions in the current table where data is missing due to data archiving (for example, when the data created on January 1, 2025 is maintained in May 2025, the partition for January 1, 2025 needs to be repaired). In this way, by replacing manual screening, it is ensured that all historical partitions that need to be repaired are covered.

[0023] Step 102: Analyze the table structure information of the current data table to obtain fields and field information in real time. According to the field information and predefined field classification rules, divide the fields into common fields, unique fields, and partition fields.

[0024] In this example, the common fields are business fields shared by the current data layer (sa) and the historical storage layer (shm) (such as amount, term). The unique fields are technical fields that only exist in the shm layer (such as system source identifier). The partition fields include the task identifier (ETL_JOB) and the date partition identifier (DT).

[0025] Analyzing the table structure information of the current data table involves dynamically reading the metadata of the table to obtain fields and field information, and classifying the fields according to predefined rules. Specifically, the latest structure of the production table, that is, the current data table, is obtained in real time by dynamically executing DESCRIBE, and the fields are classified according to the predefined rules: 1) The partition fields are stored in DATA6; 2) The unique fields are stored in DATA4; 3) The common fields are stored in DATA3. By adopting an automated field classification mechanism, it is possible to adapt to changes in the table structure.

[0026] Step 103: Generate a structured query statement based on the common fields, unique fields, partition fields, and pre-configured transcoding field mapping table.

[0027] In this example, the process of generating a structured query statement includes the construction of four types of expressions, corresponding to the common fields, unique fields, partition fields, and transcoding fields generated by the pre-configured transcoding field mapping table. Finally, all the expressions are integrated to form a complete structured query statement SQL.

[0028] Step 104: Traverse each creation date in the maintenance date set and execute the structured query statement to filter the target data from the historical data table.

[0029] In this example, when traversing the set of maintenance dates, data filtering is performed based on each creation date, and then a structured query statement is executed to filter the target data from the historical data table. In this way, the complete complement of data is ensured through the time dimension.

[0030] Step 105: Supplement the filtered target data into the corresponding data partition of the current data table according to the creation date.

[0031] In this example, the insertion operation of the target data uses partition overwrite to write into the data partition corresponding to the creation date of the current data table. By accurately positioning the repair location and overwriting and replacing the original partition data, the integrity of the data operation is guaranteed and the risk of partial writing is avoided.

[0032] The present disclosure provides a data processing method, including: obtaining a set of maintenance dates of a historical data table, where the set of maintenance dates consists of non-duplicate creation dates of the data maintained in the historical data table; parsing the table structure information of the current data table to obtain fields and field information, and dividing the fields into common fields, unique fields, and partition fields according to the field information and a predefined field classification rule; generating a structured query statement based on the common fields, unique fields, partition fields, and a pre-configured transcoding field mapping table; traversing each creation date in the set of maintenance dates, and executing the structured query statement to filter the target data from the historical data table; inserting the filtered target data into the data partition corresponding to the creation date of the current data table. This solution effectively solves the problem that after the source system maintains the historical data table in the data warehouse, the data consistency and accuracy between the current data table and the historical data table need to be manually repaired. Through an automated process, the situation of missing archived historical records due to directly extracting data from the current library of the source system is avoided, realizing automated processing, significantly reducing the operation and maintenance burden of the data warehouse, while ensuring cross-data table data consistency and improving data processing efficiency and accuracy.

[0033] In one example, the obtaining of the set of maintenance dates of the historical data table, as Figure 2 shown, includes: Step 201: Extract the update time of the historical data table.

[0034] In this example, the update time refers to the business date when the source system performs a maintenance operation on the historical data table, which is used to identify the time point when the data is modified. The update time of the historical data table (S1G_ASSET_REPAY_SERIAL_his) is dynamically extracted through JMP scheduling parameters. The specific implementation is: passing a business date parameter (in the format YYYY-MM-DD, such as 2025-05-10) in the JMP front-end scheduling task, and this parameter directly corresponds to the actual operation date when the source system performs data maintenance.

[0035] Step 202: Perform a date query operation to filter out non-duplicate creation dates where the update time is equal to the specified maintenance date and the creation date is not equal to the specified maintenance date.

[0036] In this example, the non-duplicate creation dates refer to the list of dates after deduplication using SELECT DISTINCT create_time, to avoid repeatedly repairing data for the same date. The specific date filtering query operation involves a dual filtering mechanism: filtering data records where the update time is equal to the specified maintenance date; secondly, excluding newly created data records on the current day by the condition that the creation date is not equal to the specified maintenance date; finally, extracting the set of non-duplicate creation dates (i.e., DATA2) from the records that meet the conditions.

[0037] In one example, the field information includes field name, field type, and field comment; According to the field information and predefined field classification rules, the fields are divided into common fields, unique fields, and partition fields, as Figure 3 shown, including: Step 301: Classify fields whose field names belong to the predefined partition identifier set into partition fields. The predefined partition identifier set includes a data processing task identifier field and a date partition identifier field.

[0038] In this example, the predefined partition identifier set refers to the pre-set key metadata field identifiers, including: Data processing task identifier field (ETL_JOB): The unique identifier for recording the data repair task. Date partition identifier field (DT): Stores the partition date of the data (in the format of year-month-day). By matching the field name with the predefined identifier (such as the field name being DT or ETL_JOB), it is classified as a partition field and stored in DATA6.

[0039] Step 302: Classify fields whose field names contain the predefined source identifier into unique fields.

[0040] In this example, the predefined source identifier refers to the proprietary technical identifier in the historical storage layer (shm), for example: System source field (such as SYS_SRC): Marks the system code of the data source.

[0041] By identifying fields whose field names contain specific technical identifiers (such as prefixes / suffixes like SRC, SYS, etc.) and are not partition fields, they are classified as unique fields and stored in DATA4.

[0042] Step 303: Classify the remaining fields into common fields.

[0043] In this example, the common fields refer to the business fields shared by the current data layer (sa) and the historical storage layer (shm), such as basic business attributes like loan amount (LOAN_AMT), repayment term (TERM), etc. That is, the remaining fields not classified in steps 301 and 302 are defaulted to the common field set and stored in DATA3.

[0044] In one example, generating a structured query statement based on the common fields, unique fields, partition fields, and a pre-configured transcoding field mapping table, as Figure 4 shown, includes: Step 401: Generate a first expression for the common fields: replace the null values in the numeric fields with a preset value and retain the original field type, and directly reference the original field name for non-numeric fields.

[0045] In this example, a first expression is generated for the common fields. Among them: for the numeric fields in the common fields, a null value conversion logic needs to be generated to ensure data integrity; for non-numeric fields, the original field name can be directly referenced to avoid invalid conversion. For example, for a numeric field (such as the amount loan_amount), a null value conversion expression COALESCE(loan_amount, 0) AS loan_amount is generated to retain the original field type; for a non-numeric field (such as the loan note number loan_id), the original field name is directly referenced.

[0046] Step 402: Generate a second expression for the unique fields: assign a predefined system code to the system source identifier field, assign a fixed table name to the data table identifier field, and assign the current system time to the timestamp field.

[0047] In this example, static assignment is performed on the technical identification fields in the unique field set to generate a second expression, where: 1) The system source identifier (such as SYS_SRC) is assigned the predefined system code 'S1G' (marking the source of the online lending system); 2) The data table identifier is assigned a fixed table name constant; 3) The timestamp field (such as ETL_UPD_TIME) is assigned the current system time CURRENT_TIMESTAMP.

[0048] Step 403: Generate a third expression for the partition fields: assign the first date format of the current creation date to the date partition field; assign a predefined task name constant value to the task identifier partition field.

[0049] In this example, a fixed assignment is performed on the partition fields: 1) The date partition field (DT) is dynamically assigned the first date format according to the creation date, that is, the ten-digit format (such as 2025-01-05); 2) The task identification field (ETL_JOB) is assigned the predefined constant 'S1G_ASSET_REPAY_SERIAL'. By injecting the dynamic date parameter into the partition key, accurate partitioning positioning is achieved.

[0050] Step 404: Identify the transcoding fields specified in the pre-configured transcoding field mapping table and generate the fourth expression: Construct the association logic of the standard code value table, with the source system identifier and the encoding type identifier as the association conditions; Generate the code value matching result expression, and output a string composed of a delimiter, the original field value, and an exception identifier when the match fails; Directly reference the original field name for non-transcoding fields.

[0051] In this example, based on the pre-configured mapping table cast_value=['currency_type','T99_CCY_CDE'], the fields to be transcoded (such as currency type, document type) are located; The source system identifier (SRCTAB_CD='S1G') and the encoding type identifier (CDE_TYPE='1') are used as the association conditions; When the match is successful, the standard code value is output, and when the match fails, a string composed of the delimiter #, the original value, and the exception identifier 999 is generated (such as CONCAT('#', original value, '999').

[0052] Step 405: Integrate the first expression, the second expression, the third expression, and the fourth expression to obtain the structured query statement.

[0053] In this example, the previous expressions are integrated to generate the final SQL: The first expression (common field processing), the second expression (unique field assignment), the third expression (partition field assignment), and the fourth expression (transcoding field association) are combined in the execution order to form a complete structured query statement. This integrated statement includes: 1) The field selection list; 2) The association logic between the historical data table and the code value table; 3) The partition filtering condition (WHERE create_time='$date10$').

[0054] In one example, traverse each creation date in the maintenance date set, and execute the structured query statement to filter the target data from the historical data table, such as Figure 5 shown, including: Step 501: Convert the creation date to the first date format and the second date format.

[0055] In this example, each creation date in the date set is converted into two formats: the first date format (DATA10): a ten-digit partitioned format YYYY-MM-DD (such as 2025-01-05), which is used for historical data table timestamp condition filtering and data warehouse partition key assignment; the second date format (DATA8): an eight-digit date format YYYYMMDD (such as 20250105), which is used for current incoming warehouse data table timestamp condition filtering.

[0056] Step 502: Inject the first date format into the structured query statement and execute a query operation to obtain target data.

[0057] Inject the first date format (DATA10) into the structured query statement and execute the query: Dynamic injection: Replace the partition key value in the SQL with $date10$ and then execute the operation to run the pre-built complete SQL (including logic such as field processing and code value association), and filter missing data from the historical data table and the current data table. The specific execution steps are as follows: Step 5021: Select target data from the current data table where the partition date is equal to the second date format.

[0058] In this example, filter data from the current data table where the partition date is equal to the second date format (eight-digit format), and combine the transcoding field association logic (such as currency transcoding) to obtain valid data. For example, limit the partition range by WHERE DT='$date8$' (such as DT='20250105').

[0059] Step 5022: Select target data from the historical data table that meets the following conditions: the creation date is equal to the first date format, and the partition date is not less than the first date format.

[0060] In this example, filter target data from the historical data table that meets both conditions: the creation date is equal to the first date format (ten-digit format) (such as CREATE_TIME = '2025-01-05'); the partition date is not less than the first date format (such as DT>= '2025-01-05'). Through the double timestamp verification of CREATE_TIME and DT, ensure the integrity of the archived data.

[0061] In one example, the filtered target data is filled into the corresponding data partition of the current data table according to the creation date, such as Figure 6 shown, including: Step 601: Inject the first date format and the predefined task identifier constant into the third expression.

[0062] In this example, a first date format (a ten-digit standardized format, such as 2025-01-05) and a predefined task identifier constant (such as "loan_repair_task") are injected into the partition assignment expression (the third expression): the date partition field (DT) is assigned a ten-digit date value (DT = '2025-01-05'); the task identification field (ETL_JOB) is assigned a constant task name.

[0063] Step 602: Based on the third expression, batch-write the target data to the data partition corresponding to the creation date of the current data table.

[0064] In this example, based on the injected third expression, the target data is batch-written to the target partition of the current data table through a partition overwrite operation (INSERT OVERWRITE TABLE... PARTITION (DT='2025-06-05')). By directly associating with the data partition of the creation date (that is, the value of the partition key DT corresponds to the repair date), dirty data is automatically filtered.

[0065] The present disclosure also provides a data processing device, as Figure 7 shown, the device includes: An acquisition module 701, configured to acquire a set of maintenance dates of a historical data table, where the set of maintenance dates consists of non-repeated creation dates of the data maintained in the historical data table; An analysis module 702, configured to analyze the table structure information of the current data table to obtain fields and field information in real time, and divide the fields into common fields, unique fields, and partition fields according to the field information and a predefined field classification rule; A generation module 703, configured to generate a structured query statement based on the common fields, unique fields, partition fields, and a pre-configured transcoding field mapping table; An execution module 704, configured to traverse each creation date in the set of maintenance dates and execute the structured query statement to filter target data from the historical data table; An insertion module 705, configured to supplement the filtered target data into the corresponding data partition of the current data table according to the creation date.

[0066] In one example, the acquisition module 701 is specifically configured to: Extract the update time of the historical data table; Execute a date query operation to filter non-repeated creation dates whose update time is equal to a specified maintenance date and whose creation date is not equal to the specified maintenance date.

[0067] In one example, the field information includes a field name, a field type, and a field comment; specifically, the parsing module 702 is configured to: Classify fields whose field names belong to a predefined partition identifier set into partition fields, where the predefined partition identifier set includes a data processing task identifier field and a date partition identifier field; Classify fields whose field names contain a predefined source identifier into unique fields; Classify the remaining fields into common fields.

[0068] In one example, the generating module 703 is specifically configured to: Generate a first expression for the common fields: replace null values in numeric fields with a preset value and retain the original field type, and directly reference the original field name for non-numeric fields; Generate a second expression for the unique fields: assign a predefined system code to the system source identifier field, assign a fixed table name to the data table identifier field, and assign the current system time to the timestamp field; Generate a third expression for the partition fields: assign the first date format of the current creation date to the date partition field; assign a predefined task name constant value to the task identifier partition field; Identify the transcoding fields specified in the preconfigured transcoding field mapping table and generate a fourth expression: construct a standard code value table association logic with the source system identifier and the coding type identifier as the association conditions; generate a code value matching result expression, and output a string composed of a delimiter, the original field value, and an exception identifier when the matching fails; directly reference the original field name for non-transcoding fields; Integrate the first expression, the second expression, the third expression, and the fourth expression to obtain the structured query statement.

[0069] In one example, the execution module 704 is specifically configured to: Convert the creation date into a first date format and a second date format; Inject the first date format into the structured query statement and perform a query operation to obtain target data, including: Select target data from the current data table where the partition date is equal to the second date format; Select target data from the historical data table that meet the following conditions: the creation date is equal to the first date format, and the partition date is not less than the first date format.

[0070] In one example, the insertion module 705 is specifically configured to: Inject the first date format and a predefined task identifier constant into the third expression; Based on the third expression, batch write the target data into the data partition corresponding to the creation date of the current data table.

[0071] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.

[0072] Figure 8 FIG. shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0073] As Figure 8 shown, the device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0074] A plurality of components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0075] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the data processing method. For example, in some embodiments, the data processing method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the data processing method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the data processing method by any other suitable means (e.g., by means of firmware).

[0076] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0077] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0078] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0079] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0080] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0081] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, can also be a server of a distributed system, or a server incorporating a blockchain.

[0082] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of this disclosure can be achieved, and no limitation is made herein.

[0083] In addition, the terms "first" and "second" are used only for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of this disclosure, "a plurality of" means two or more, unless otherwise specifically defined.

[0084] As described above, the above are only specific embodiments of this disclosure, but the protection scope of this disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by this disclosure can easily think of changes or substitutions, which should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure shall be subject to the protection scope of the claims.

Claims

1. A data processing method, characterized in that, The method includes: Obtaining a set of maintenance dates of a historical data table, where the set of maintenance dates consists of non-duplicate creation dates of the data maintained in the historical data table; Parsing the table structure information of the current data table to obtain fields and field information in real time, and dividing the fields into common fields, unique fields, and partition fields according to the field information and predefined field classification rules; Generating a structured query statement based on the common fields, unique fields, partition fields, and a pre-configured transcoding field mapping table; Traversing each creation date in the set of maintenance dates, and executing the structured query statement to filter target data from the historical data table; Supplementing the filtered target data into the corresponding data partition of the current data table according to the creation date.

2. The method according to claim 1, characterized in that, The obtaining of the set of maintenance dates of the historical data table includes: Extracting the update time of the historical data table; Performing a date query operation to filter non-duplicate creation dates where the update time is equal to the specified maintenance date and the creation date is not equal to the specified maintenance date.

3. The method according to claim 1, wherein The field information includes field name, field type, and field comment; The dividing of the fields into common fields, unique fields, and partition fields according to the field information and predefined field classification rules includes: Classifying fields whose field names belong to a predefined partition identifier set into partition fields, where the predefined partition identifier set includes a data processing task identifier field and a date partition identifier field; Classifying fields whose field names contain a predefined source identifier into unique fields; Classifying the remaining fields into common fields.

4. The method according to claim 1, wherein The generating of the structured query statement based on the common fields, unique fields, partition fields, and a pre-configured transcoding field mapping table includes: Generating a first expression for the common fields: replacing null values in numeric fields with a preset value and retaining the original field type, and directly referencing the original field name for non-numeric fields; Generating a second expression for the unique fields: assigning a predefined system code to the system source identifier field, assigning a fixed table name to the data table identifier field, and assigning the current system time to the timestamp field; Generating a third expression for the partition fields: assigning the first date format of the current creation date to the date partition field; assigning a predefined task name constant value to the task identifier partition field; Identifying the transcoding fields specified in the pre-configured transcoding field mapping table and generating a fourth expression: constructing a standard code value table association logic with the source system identifier and the coding type identifier as the association conditions; generating a code value matching result expression, and outputting a string composed of a delimiter, the original field value, and an exception identifier when the matching fails; directly referencing the original field name for non-transcoding fields; Integrating the first expression, the second expression, the third expression, and the fourth expression to obtain the structured query statement.

5. The method according to claim 4, wherein The traversing of each creation date in the set of maintenance dates and executing the structured query statement to filter target data from the historical data table includes: Converting the creation date into the first date format and the second date format; Injecting the first date format into the structured query statement and performing a query operation to obtain the target data, including: Select target data from the current data table where the partition date is equal to the second date format; Select target data from the historical data table that meets the following conditions: the creation date is equal to the first date format, and the partition date is not less than the first date format.

6. The method according to claim 5, wherein The step of supplementing the filtered target data into the corresponding data partition of the current data table according to the creation date includes: Inject the first date format and the predefined task identifier constant into the third expression; Based on the third expression, batch-write the target data into the data partition of the current data table corresponding to the creation date.

7. A data processing device, characterized in that, The apparatus includes: An acquisition module, configured to acquire a set of maintenance dates of the historical data table, where the set of maintenance dates consists of non-duplicate creation dates of the data maintained in the historical data table; An analysis module, configured to analyze the table structure information of the current data table to obtain fields and field information in real time, and divide the fields into common fields, unique fields, and partition fields according to the field information and the predefined field classification rules; A generation module, configured to generate a structured query statement based on the common fields, unique fields, partition fields, and the pre-configured transcoding field mapping table; An execution module, configured to traverse each creation date in the set of maintenance dates, and execute the structured query statement to filter target data from the historical data table; An insertion module, configured to supplement the filtered target data into the corresponding data partition of the current data table according to the creation date.

8. The device according to claim 7, wherein The acquisition module is specifically configured to: extract the update time of the historical data table; perform a date query operation to filter non-duplicate creation dates where the update time is equal to the specified maintenance date and the creation date is not equal to the specified maintenance date.

9. An electronic device, characterized in that, Comprising: At least one processor; And a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-6.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Partition table historical data storage method and device and computer readable medium

    CN115438049A

  • Zipper table data processing method and device

    CN117149775A

  • Publishing to a data warehouse

    US20200026711A1