Data processing method and device, medium and product

By parsing database disk logs to obtain row-level incremental data, and using logical primary key binding and transaction context registration tables to generate standardized business events, the problem of data distortion in traditional data acquisition methods under microservice architecture is solved, achieving efficient and accurate data processing that meets the needs of the fintech field.

CN122045159APending Publication Date: 2026-05-15INDUSTRIAL AND COMMERCIAL BANK OF CHINA +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INDUSTRIAL AND COMMERCIAL BANK OF CHINA
Filing Date
2026-02-11
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Traditional business data acquisition methods are difficult to keep synchronized under microservices and cloud-native architectures, resulting in data distortion. Furthermore, intrusive solutions suffer from performance fluctuations in high-concurrency scenarios, failing to meet the high requirements of fields such as finance and e-commerce.

Method used

By parsing the logs on the database disk to obtain row-level incremental data, and using logical primary key binding and transaction context registration tables to aggregate the data, standardized business events are generated and target business data is calculated, thus achieving non-intrusive data processing.

Benefits of technology

It achieves transaction-level aggregation and full-link traceability of incremental data at the database row level, ensuring data consistency and accuracy, improving the efficiency of business data determination, reducing system resource consumption, and adapting to the high requirements of the financial technology field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122045159A_ABST
    Figure CN122045159A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device, a medium and a product, relates to the technical field of computers, and can be applied to the field of financial science and technology. The method comprises the following steps: acquiring and analyzing a log on a database disk to obtain at least one row-level incremental data; the row-level incremental data comprises a transaction number; determining a logic main key of row-level incremental data from the data source table set; based on the logic main key, storing the row-level incremental data into a transaction context registration form corresponding to the transaction number; generating a standardized business event based on the transaction context registration form; calculating the standardized business event to obtain target business data corresponding to the business event code; the business event code uniqueness identifies the standardized business event. Through the technical scheme, the determination efficiency of the business data can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology and can be used in the field of financial technology, particularly to a data processing method, apparatus, medium, and product. Background Technology

[0002] There are two main traditional methods for acquiring business data: the first is code-based tracking, which involves inserting statistical logic into the business code. This method is costly to develop, test, and maintain, and is prone to omissions. The second method involves creating triggers or intermediate tables in the database, which can cause lock contention and performance fluctuations in online business processes, making it an intrusive solution. With the popularization of microservices and cloud-native architectures, business systems iterate frequently, making it difficult for tracking solutions to stay synchronized, leading to data distortion. At the same time, high-concurrency scenarios such as finance and e-commerce place higher demands on "zero intrusion, low latency, and traceability." Summary of the Invention

[0003] This invention provides a data processing method, apparatus, medium, and product to solve the problem of low efficiency in determining business data.

[0004] According to one aspect of the present invention, a data processing method is provided, comprising: Obtain and parse the logs on the database disk to get at least one row-level incremental data; the row-level incremental data includes the transaction number. Determine the logical primary key of the row-level incremental data from the data source table set; Based on the logical primary key, the row-level incremental data is stored in the transaction context registration table corresponding to the transaction number; Based on the transaction context registration table, standardized business events are generated; The standardized business events are calculated to obtain the target business data corresponding to the business event code; the business event code uniquely identifies the standardized business event.

[0005] According to another aspect of the present invention, a data processing apparatus is provided, comprising: The data acquisition module is used to acquire and parse the logs on the database disk to obtain at least one row-level incremental data; the row-level incremental data includes the transaction number; The logical primary key determination module is used to determine the logical primary key of the row-level incremental data from the data source table set; The data storage module is used to store the row-level incremental data into the transaction context registration table corresponding to the transaction number based on the logical primary key; The event generation module is used to generate standardized business events based on the transaction context registration table; The data calculation module is used to calculate the standardized business events to obtain the target business data corresponding to the business event code; the business event code uniquely identifies the standardized business event.

[0006] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the data processing method according to any embodiment of the present invention.

[0007] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the data processing method described in any embodiment of the present invention.

[0008] According to another aspect of the present invention, a computer program product is provided, comprising a computer program / instructions that, when executed by a processor, implement the data processing method as described in any embodiment of the present invention.

[0009] This invention adopts a non-intrusive approach, obtaining row-level incremental data by parsing logs on the database disk. Then, through a full-process structured processing including logical primary key binding, transaction context registration table aggregation, standardized business event generation, and target business data calculation, it achieves transaction-level aggregation and full-link traceability of database row-level incremental data. This ensures data consistency and accuracy, improves the efficiency of business data determination, reduces system resource consumption, enhances system maintainability and scalability, and is adapted to the high requirements of data compliance and security in the financial technology field.

[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a flowchart of a data processing method provided in an embodiment of the present invention; Figure 2 This is a flowchart of another data processing method provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a data processing device provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of an electronic device that implements the data processing method of the present invention. Detailed Implementation

[0013] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0014] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0015] Furthermore, it should be noted that the logs and other information collected in the technical solution of this invention are information and data authorized by the user or fully authorized by all parties. The collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data all comply with the relevant laws, regulations and standards of the relevant countries and regions, necessary confidentiality measures have been taken, and they do not violate public order and good morals. Corresponding operation entry points are provided for users to choose to authorize or refuse.

[0016] Figure 1 This is a flowchart of a data processing method provided in an embodiment of the present invention. This embodiment is applicable to processing database logs to obtain business data. The method can be executed by a data processing device, which can be implemented in hardware and / or software. This device can be configured in an electronic device with corresponding data processing capabilities, such as a server. Figure 1 As shown, the method includes: S110. Obtain and parse the logs on the database disk to get at least one row-level incremental data; the row-level incremental data includes the transaction number.

[0017] In this context, the database disk is the physical storage medium for persistently storing files such as logs in the database system. Row-level incremental data refers to row-level data containing only the changed data portions after an insert, update, or delete operation occurs on a record in a database table. A transaction number is a globally unique identifier assigned by the database. The transaction number is used to associate all row-level incremental data within the same transaction.

[0018] Specifically, the log files on the database disk are read directly using a block sequential read method, bypassing the database engine. The obtained logs are parsed into at least one row-level incremental data. Each row-level incremental data is associated with a specific transaction. All row-level incremental data within the same transaction will be associated with the same transaction number. The scattered row-level incremental data can be attributed to the corresponding database transaction through the transaction number.

[0019] S120. Determine the logical primary key of the row-level incremental data from the data source table set.

[0020] The data source table set includes at least one data source table. The data source table is a database table that stores raw business data and serves as the source for row-level incremental data tracing. The data source table has a pre-defined table structure (including fields, indexes, constraints, etc.). When one or more records in the table are inserted, updated, or deleted, corresponding row-level incremental data is generated. The logical primary key uniquely identifies a single record in the data source table and can also be mapped to the identification information of the row-level incremental data.

[0021] Specifically, the corresponding data source table for the row-level incremental data is determined from the data source table set, and the logical primary key of the row-level incremental data is determined based on the data source table. This achieves an accurate match between the row-level incremental data and the original records in the data source table, facilitating subsequent operations such as storing the row-level incremental data.

[0022] S130. Based on the logical primary key, store the row-level incremental data into the transaction context registration table corresponding to the transaction number.

[0023] The transaction context registry is a data table used to store all row-level incremental data under the same transaction number.

[0024] Specifically, based on the logical primary key, row-level incremental data is stored in the transaction context registration table corresponding to the transaction number, thereby realizing the aggregation of scattered row-level incremental data by transaction dimension, ensuring the atomicity of transactions, and enabling centralized management of all row-level incremental data under the same transaction.

[0025] S140. Generate standardized business events based on the transaction context registration table.

[0026] Standardized business events are generated based on data retained in the transaction context register after a transaction is committed. For example, if the final data retained in the transaction context register after a transaction is committed is "Transaction No.: T1, Transaction Status: Committed, Logical Primary Key: A1, Row-level incremental data includes: Deposit Type = 1-year term, Deposit Amount = 5000 yuan, Contract Status = Automatic Renewal, Business Event Code = F1", then the standardized business event is "Business Event Code: F1, Deposit Type: 1-year term, Deposit Amount: 5000 yuan, Contract Status: Automatic Renewal", where F1 indicates that the standardized business event is a fixed deposit and automatic renewal contract event. Thus, standardized business events only retain the event code and business data required by the business side, eliminating technical layers such as logical primary keys and transaction status fields from the transaction context register, adapting to the data management needs at the business level.

[0027] S150. Calculate the standardized business events to obtain the target business data corresponding to the business event code; the business event code uniquely identifies the standardized business event.

[0028] Specifically, the target business data is the business result data obtained by calculating the business fields in standardized business events according to preset business calculation rules. This achieves lightweight and efficient reuse of the target business data, facilitating rapid access to the target business data, adapting to the needs of multiple business scenarios in the fintech field, and improving business response efficiency.

[0029] Optionally, based on the logical primary key, the row-level incremental data is stored in the transaction context registration table corresponding to the transaction number, including: checking whether the logical primary key exists in the transaction context registration table corresponding to the transaction number; if it does not exist, the row-level incremental data is stored in the transaction context registration table to generate a first snapshot; if it exists, the second snapshot corresponding to the logical primary key in the transaction context registration table is updated based on the row-level incremental data.

[0030] The first snapshot refers to the initial data image generated for the business record corresponding to the logical primary key after storing row-level incremental data in the transaction context registry table when it is detected that the logical primary key does not exist in the transaction context registry table corresponding to the transaction number. The second snapshot refers to the data image obtained by updating the business record corresponding to the logical primary key after storing row-level incremental data in the transaction context registry table when it is detected that the logical primary key exists in the transaction context registry table corresponding to the transaction number. For example, if the transaction context registry table corresponding to transaction number T1 does not contain logical primary key A1, then the row-level incremental data of logical primary key A1 (deposit type = 1 year, deposit amount = 5000 yuan, contract status = not signed, business event code = F1) is processed, and the first snapshot generated is "Transaction number: T1, transaction status: not committed, logical primary key: A1, row-level incremental data includes: deposit type = 1 year, deposit amount = 5000 yuan, contract status = not signed, business event code = F1"; subsequently, under the same transaction T1, the row-level incremental data of logical primary key A1 is processed again (the contract status is changed to automatic renewal), and the second snapshot generated is "Transaction number: T1, transaction status: not committed, logical primary key: A1, row-level incremental data includes: deposit type = 1 year, deposit amount = 5000 yuan, contract status = automatic renewal, business event code = F1". By detecting whether the logical primary key exists in the transaction context registration table of the corresponding transaction number, the system avoids the duplicate storage of row-level incremental data of the same business record within the same transaction, ensuring the uniqueness and cleanliness of the data in the transaction context registration table, and thus avoiding data redundancy and duplicate snapshots.

[0031] Optionally, the standardized business events are calculated to obtain the target business data corresponding to the business event code, including: determining the first extraction field and aggregation dimension corresponding to the business event code; the aggregation dimension includes time dimension and spatial dimension; extracting the corresponding first field value from the standardized business events based on the first extraction field; performing aggregation calculation on the first field value based on the aggregation dimension to obtain the initial business data; validating the initial business data to obtain the target business data corresponding to the business event code; the validation includes row-level consistency validation, compliance validation, and backtracking validation.

[0032] The first extraction field is the field associated with the business event code and needs to be extracted from standardized business events. The aggregation dimension is the partitioning dimension for aggregating and calculating the field values ​​extracted from standardized business events. The time dimension includes time ranges divided by hour, day, month, quarter, or year. The spatial dimension includes spatial ranges divided by institution (e.g., head office or branch) or business processing channel (e.g., mobile banking or counter). The first field value is the data value of the first extraction field. Initial business data refers to the calculated result data obtained after aggregating and calculating the first field values ​​extracted from standardized business events based on the aggregation dimension. Row-level consistency verification refers to verifying the consistency of field values ​​between the initial business data obtained from the aggregation calculation and the row-level incremental data in the log (e.g., an alarm is triggered if the error rate of the deposit amount is greater than 0.1%). Compliance verification verifies whether the initial business data meets regulatory requirements (e.g., the amount of tax reduction must be retained to two decimal places, and the precision of log fields must match). Retrospective verification refers to the cross-time dimension historical data consistency verification of initial business data (such as the total amount of automatic renewal in January calculated on a monthly basis; during retrospective verification, it is necessary to check whether this value is consistent with the cumulative amount of each day in January).

[0033] Specifically, the process involves determining the first extraction field and aggregation dimension corresponding to the business event code; extracting the corresponding first field value from standardized business events based on the first extraction field; performing aggregation calculations (such as summation and averaging) on ​​the first field value based on the aggregation dimension to obtain initial business data; validating the initial business data, and obtaining the target business data corresponding to the business event code after successful validation. For example, if the first extraction field is the deposit amount, the aggregation dimensions are January 1, 2026 (time dimension) and counter (spatial dimension), and the business event code is F1 (corresponding to fixed deposit and automatic renewal contract events), then the deposit amounts (5000 yuan, 8000 yuan, and 10000 yuan) of all F1 events processed at the counter on January 1, 2026 are summed to calculate the initial business data, including a deposit amount of 23000 yuan; validating the initial business data, and obtaining the target business data corresponding to the business event code after successful validation, the target business data includes a deposit amount of 23000 yuan. Thus, each business event code corresponds to the first extraction field and aggregation dimension, extracting only the business data required for calculation, avoiding irrelevant data from interfering with the calculation, making the calculation process lightweight and efficient; it can adapt to the common summary requirements of "statistics by time and division by business scope" in financial business, making the calculation logic fit the actual usage needs of different business scenarios.

[0034] This invention adopts a non-intrusive approach, obtaining row-level incremental data by parsing logs on the database disk. Through a full-process structured processing including logical primary key binding, transaction context registration table aggregation, standardized business event generation, and target business data calculation, it achieves transaction-level aggregation and full-link traceability of database row-level incremental data. This ensures data consistency and accuracy, improves the efficiency of business data determination, reduces system resource consumption, enhances system maintainability and scalability, and is adapted to the high requirements of data compliance and security in the financial technology field.

[0035] Figure 2 This is a flowchart of another data processing method provided by an embodiment of the present invention. Based on the above embodiments, this embodiment optimizes the process of "determining the logical primary key of row-level incremental data from the data source table set," providing an optional implementation scheme. For example... Figure 2 As shown, the method includes: S210. Obtain and parse the logs on the database disk to get at least one row-level incremental data; the row-level incremental data includes the transaction number.

[0036] S220. Determine the logical primary key of the row-level incremental data from the data source table set.

[0037] S230. Based on the logical primary key, store the row-level incremental data into the transaction context registration table corresponding to the transaction number.

[0038] S240. Generate standardized business events based on the transaction context registration table.

[0039] S250. Calculate the standardized business events to obtain the target business data corresponding to the business event code; the business event code uniquely identifies the standardized business event.

[0040] Optionally, the logical primary key of the row-level incremental data is determined from the data source table set, including: based on the table name of the row-level incremental data, searching for the primary key of the data source table corresponding to the row-level incremental data; if the primary key is found, using the primary key as the unique identifier key of the row-level incremental data; if the primary key is not found, searching for a unique key from the data source table; if a unique key is found, using the unique key as the unique identifier key of the row-level incremental data; if a unique key is not found, using the database physical row number of the row-level incremental data as the unique identifier key of the row-level incremental data; and performing a hash calculation on the unique identifier key to obtain the logical primary key of the row-level incremental data.

[0041] In this context, the primary key is a field or combination of fields used to uniquely identify each row of data in the data source table. The unique identifier key is used to uniquely identify row-level incremental data. A unique key is a field or combination of fields used to uniquely identify each row of data in the data source table when the primary key is missing. The database physical row number is the physical location number of the row-level incremental data on the database disk. For example, if the row-level incremental data corresponds to a contract table (data source table): if the contract table has a primary key `signId`, then the unique identifier key is `signId`, which is hashed to generate the logical primary key; if the contract table has no primary key but has a unique key `userId`, then the unique identifier key is `userId`, which is hashed to generate the logical primary key; if the contract table has neither a primary key nor a unique key, then the unique identifier key is the database physical row number, which is hashed to generate the logical primary key. This approach ensures the uniqueness of unique identifier keys through a hierarchical search strategy based on "primary key, unique key, and physical row number." Furthermore, hash calculations are used to transform different types of unique identifier keys (fields, field combinations, or physical row numbers) into logical primary keys with a unified format. This solves the problems of inconsistent identifier rules across different data source tables and the absence of primary or unique keys in some tables, thus guaranteeing the global uniqueness and standardized format of logical primary keys.

[0042] Optionally, row-level incremental data may also include database name, table name, operation type, and column data. Accordingly, before determining the logical primary key of the row-level incremental data from the data source table set, the following steps are also included: determining the target triggering condition corresponding to the row-level incremental data based on the database name, table name, and operation type; determining whether the column data meets the target triggering condition, and if so, determining the logical primary key of the row-level incremental data from the data source table set.

[0043] Here, the database name is the name of the database. The table name is the name of the data source table. Operation types include INSERT, DELETE, and UPDATE. Column data is a collection of key-value pairs containing the field names and values ​​of specific columns in the corresponding data source table, contained within the row-level incremental data. Column data may include account number, deposit amount, deposit period, contract status, transaction time, counter number, etc. The target trigger condition is the trigger condition that matches the row-level incremental data.

[0044] Specifically, trigger conditions are pre-defined for different combinations of "database, table, and operation type". The logical primary key will only be determined if the column data of the row-level incremental data meets this trigger condition. For example, the target trigger condition for "database name = funds database, table name = contract table, operation type = update" is "deposit amount greater than or equal to 1000 yuan and contract status is automatic renewal". Only if the column data corresponding to the row-level incremental data meets this trigger condition will the logical primary key determination process begin. This filters out row-level incremental data that has no business significance, reducing the unnecessary overhead of subsequent logical primary key calculations.

[0045] Optionally, based on the database name, table name, and operation type, the target triggering condition corresponding to the row-level incremental data is determined, including: based on the database name and table name, candidate semantic dictionaries are selected from at least one business semantic dictionary; based on the operation type, the target triggering condition is determined from at least one triggering condition in the candidate semantic dictionary.

[0046] The business semantic dictionary refers to a pre-built set of triggering conditions associated with database names and table names. The candidate semantic dictionary is the business semantic dictionary corresponding to row-level incremental data.

[0047] Specifically, based on the database and table names of the row-level incremental data, business semantic dictionaries that match the row-level incremental data are selected as candidate semantic dictionaries from all business semantic dictionaries. Based on the operation type, the target triggering condition is determined from at least one triggering condition in the candidate semantic dictionaries. For example, the triggering conditions of the candidate semantic dictionary may include: when the operation type is INSERT: account status = "normal" and deposit type = "1-year term" and deposit amount ≥ 1000 yuan; when the operation type is UPDATE: the contract status changes from "not contracted" to "automatic renewal" and the associated transaction number is consistent. This forms a three-layer precise matching logic of "global dictionary, database / table-specific dictionary, and operation type-specific condition", which not only ensures dual adaptation of triggering conditions with databases / tables and operation types, but also narrows the scope of condition search and improves the efficiency of determining the target triggering condition.

[0048] Optionally, determining whether the column data meets the target triggering conditions includes: determining at least one column field to be extracted and the corresponding field threshold from the target triggering conditions; obtaining a second field value of at least one column field to be extracted from the column data; comparing the second field value with the field threshold to obtain a comparison result; if at least one second field value meets the corresponding field threshold, then the comparison result is determined to be that the column data meets the target triggering conditions; if there is a second field value that does not meet the corresponding field threshold, then the comparison result is determined to be that the column data does not meet the target triggering conditions.

[0049] The column fields to be extracted are those extracted from the target triggering conditions and whose values ​​need to be extracted from the column data. The second field value refers to the field value extracted from the column data that corresponds to the column field to be extracted. For example, if the target triggering conditions are: account status = "normal" and deposit type = "1 year" and deposit amount ≥ 1000 yuan, then the column fields to be extracted include: account status, deposit type, and deposit amount; if the column data is: account status is normal, deposit type is 1 year, deposit amount is 5000 yuan, contract status is automatic renewal, and interest calculation method is daily interest, then the second field value is normal (account status), 1 year (deposit type), and 5000 yuan (deposit amount); the comparison result is 5000 is greater than 1000, the account status is normal, and the deposit type is 1 year, that is, it is determined that the column data meets the target triggering conditions. By identifying the column fields to be extracted from the target triggering conditions, the verification range of the column data can be accurately locked; then, the corresponding second field value is extracted and compared with the threshold, realizing accurate verification of the column data and the target triggering conditions; comparison is only performed on the fields related to the triggering conditions, avoiding the resource consumption of full column data verification, while ensuring that only row-level incremental data that conforms to business rules enters the subsequent process, thus guaranteeing the accuracy and effectiveness of data processing.

[0050] Optionally, it also includes: according to the period, obtaining the hit rate of the business semantic dictionary based on the number of hits and the total number of changes; when the hit rate is higher than the first threshold or lower than the second threshold, updating the triggering conditions in the business semantic dictionary.

[0051] The hit count refers to the cumulative number of row-level incremental data entries within the period that satisfy the target triggering conditions in the business semantic dictionary. The total number of changes refers to the total number of row-level incremental data entries corresponding to the business semantic dictionary within the period. The hit rate is the ratio of the hit count of the business semantic dictionary to the total number of changes within the period. The period is a pre-defined statistical period. The first threshold and the second threshold refer to two pre-defined critical values ​​used to determine whether the hit rate of the business semantic dictionary is abnormal. The first threshold is the upper limit critical value of the hit rate, and the second threshold is the lower limit critical value of the hit rate.

[0052] Specifically, based on a cycle, the hit rate of the business semantic dictionary is determined by the ratio of the number of hits in the business semantic dictionary to the total number of changes. When the hit rate is higher than the first threshold, it indicates that the triggering conditions are too lenient, allowing almost all data to pass through, including invalid data, and the restrictions on the triggering conditions need to be strengthened. When the hit rate is lower than the second threshold, it indicates that the triggering conditions are too strict, filtering out a lot of valid data, and the triggering conditions need to be relaxed. Thus, the rationality of the conditions is determined through threshold matching calculation, and the triggering conditions are adjusted to ensure that the business semantic dictionary always adapts to the actual business data change characteristics.

[0053] This invention employs a hierarchical search strategy based on "primary key, unique key, and physical row number" to ensure the uniqueness of unique identifier keys. Then, through hash calculation, different types of unique identifier keys (fields, field combinations, or physical row numbers) are transformed into logical primary keys with a unified format. This solves the problems of inconsistent identifier rules across different data source tables and the absence of primary or unique keys in some tables, ensuring the global uniqueness and standardized format of logical primary keys. Only when the column data corresponding to row-level incremental data meets this triggering condition will the logical primary key determination stage begin, filtering out row-level incremental data without business significance and reducing the unnecessary overhead of subsequent logical primary key calculations. This forms a "global dictionary, database..." The three-layer precise matching logic of "table-specific dictionary and operation type-specific conditions" ensures dual compatibility between trigger conditions and database tables and operation types, while narrowing the scope of condition search and improving the efficiency of determining target trigger conditions. By determining the column fields to be extracted from the target trigger conditions, the verification range of column data can be accurately locked. Then, the corresponding second field value is extracted and compared with the threshold, realizing the precise verification of column data with target trigger conditions. Only fields related to trigger conditions are compared, avoiding the resource consumption of full column data verification, while ensuring that only row-level incremental data that conforms to business rules enters the subsequent process, guaranteeing the accuracy and effectiveness of data processing.

[0054] Figure 3 This is a schematic diagram of a data processing device provided in an embodiment of the present invention. This embodiment is applicable to processing database logs to obtain business data. The device can be implemented in hardware and / or software and can be configured in an electronic device with corresponding data processing capabilities, such as a server. Figure 3 As shown, the device includes: Data acquisition module 310 is used to acquire and parse logs on the database disk to obtain at least one row-level incremental data; the row-level incremental data includes the transaction number; The logical primary key determination module 320 is used to determine the logical primary key of row-level incremental data from the data source table set; Data storage module 330 is used to store row-level incremental data into the transaction context registration table corresponding to the transaction number based on the logical primary key; Event generation module 340 is used to generate standardized business events based on the transaction context registration table; The data calculation module 350 is used to calculate standardized business events to obtain the target business data corresponding to the business event code; the business event code uniquely identifies the standardized business event.

[0055] This invention adopts a non-intrusive approach, obtaining row-level incremental data by parsing logs on the database disk. Through a full-process structured processing including logical primary key binding, transaction context registration table aggregation, standardized business event generation, and target business data calculation, it achieves transaction-level aggregation and full-link traceability of database row-level incremental data. This ensures data consistency and accuracy, improves the efficiency of business data determination, reduces system resource consumption, enhances system maintainability and scalability, and is adapted to the high requirements of data compliance and security in the financial technology field.

[0056] Optionally, the logical primary key determination module 320 includes: The primary key lookup unit is used to find the primary key of the data source table corresponding to the row-level incremental data from the data source table set based on the table name of the row-level incremental data. The first identifier key determination unit is used to use the primary key as the unique identifier key for row-level incremental data if a primary key is found. The unique key lookup unit is used to search for a unique key from the data source table if the primary key is not found. The second identifier key determination unit is used to use the unique key as the unique identifier key for row-level incremental data if a unique key is found. The third identifier key determination unit is used to use the database physical row number of the row-level incremental data as the unique identifier key of the row-level incremental data if no unique key is found. The logical primary key determination unit is used to perform hash calculations on the unique identifier key to obtain the logical primary key of the row-level incremental data.

[0057] Optionally, the data storage module 330 includes: The logical primary key detection unit is used to detect whether the logical primary key exists in the transaction context registration table corresponding to the transaction number. The first snapshot generation unit is used to store row-level incremental data into the transaction context registry table and generate the first snapshot if it does not exist. The second snapshot update unit is used to update the second snapshot corresponding to the logical primary key in the transaction context registry table based on row-level incremental data, if it exists.

[0058] Optional, the data computing module 350 includes: The field determination unit is used to determine the first extraction field and aggregation dimension corresponding to the business event code; the aggregation dimension includes time dimension and spatial dimension. The first field value extraction unit is used to extract the corresponding first field value from the standardized business event based on the first extraction field. The initial business data determination unit is used to perform aggregation calculations on the value of the first field based on the aggregation dimension to obtain the initial business data; The target business data determination unit is used to verify the initial business data and obtain the target business data corresponding to the business event code; the verification includes row-level consistency verification, compliance verification and backtracking verification.

[0059] Optionally, row-level incremental data may also include database name, table name, operation type, and column data; Accordingly, the device also includes: a condition detection module; The condition detection module includes: The condition determination sub-unit is used to determine the target triggering conditions corresponding to row-level incremental data based on the database name, table name, and operation type. The condition detection subunit is used to determine whether the column data meets the target triggering condition. If it does, the logical primary key of the row-level incremental data is determined from the data source table set.

[0060] Optionally, the condition determination subunit is specifically used for: filtering candidate semantic dictionaries from at least one business semantic dictionary based on library name and table name; and determining the target triggering condition from at least one triggering condition in the candidate semantic dictionary based on operation type.

[0061] Optionally, the condition detection subunit is specifically used for: determining at least one column field to be extracted and the corresponding field threshold from the target triggering conditions; obtaining a second field value of at least one column field to be extracted from the column data; comparing the second field value with the field threshold to obtain a comparison result; if at least one second field value meets the corresponding field threshold, then the comparison result is determined to be that the column data meets the target triggering conditions; if there is a second field value that does not meet the corresponding field threshold, then the comparison result is determined to be that the column data does not meet the target triggering conditions.

[0062] The data processing apparatus provided in the embodiments of the present invention can execute the data processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0063] According to embodiments of the present invention, the present invention also provides an electronic device, a readable storage medium, and a computer program product.

[0064] Figure 4A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0065] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0066] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0067] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as data processing methods.

[0068] In some embodiments, the data processing method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the data processing method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the data processing method by any other suitable means (e.g., by means of firmware).

[0069] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0070] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0071] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0072] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0073] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0074] A computing system can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product within the cloud computing service system to address the shortcomings of traditional physical hosts and virtual private servers, such as high management difficulty and weak business scalability.

[0075] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0076] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A data processing method, characterized in that, The method includes: Obtain and parse the logs on the database disk to get at least one row-level incremental data; the row-level incremental data includes the transaction number. Determine the logical primary key of the row-level incremental data from the data source table set; Based on the logical primary key, the row-level incremental data is stored in the transaction context registration table corresponding to the transaction number; Based on the transaction context registration table, standardized business events are generated; The standardized business events are calculated to obtain the target business data corresponding to the business event code; the business event code uniquely identifies the standardized business event.

2. The method according to claim 1, characterized in that, Determining the logical primary key of the row-level incremental data from the data source table set includes: Based on the table name of the row-level incremental data, find the primary key of the data source table corresponding to the row-level incremental data from the data source table set; If the primary key is found, then the primary key will be used as the unique identifier key for the row-level incremental data; If the primary key is not found, then a unique key is searched from the data source table; If the unique key is found, then the unique key is used as the unique identifier key for the row-level incremental data; If the unique key is not found, the database physical row number of the row-level incremental data will be used as the unique identifier key of the row-level incremental data. The unique identifier key is hashed to obtain the logical primary key of the row-level incremental data.

3. The method according to claim 1, characterized in that, The step of storing the row-level incremental data into the transaction context registration table corresponding to the transaction number based on the logical primary key includes: Check whether the logical primary key exists in the transaction context registration table corresponding to the transaction number; If it does not exist, the row-level incremental data is stored in the transaction context registration table to generate the first snapshot; If it exists, then based on the row-level incremental data, the second snapshot corresponding to the logical primary key in the transaction context registration table is updated.

4. The method according to claim 1, characterized in that, The calculation of the standardized business events to obtain the target business data corresponding to the business event code includes: Determine the first extraction field and aggregation dimension corresponding to the business event code; the aggregation dimension includes a time dimension and a spatial dimension. Based on the first extracted field, extract the corresponding first field value from the standardized business event; Based on the aggregation dimension, the value of the first field is aggregated and calculated to obtain the initial business data; The initial business data is validated to obtain the target business data corresponding to the business event code; the validation includes row-level consistency validation, compliance validation and backtracking validation.

5. The method according to claim 1, characterized in that, The row-level incremental data also includes the database name, table name, operation type, and column data; Accordingly, before determining the logical primary key of the row-level incremental data from the data source table set, the method further includes: Based on the database name, the table name, and the operation type, determine the target triggering condition corresponding to the row-level incremental data; Determine whether the column data meets the target triggering condition. If it does, determine the logical primary key of the row-level incremental data from the data source table set.

6. The method according to claim 5, characterized in that, The step of determining the target triggering condition corresponding to the row-level incremental data based on the database name, the table name, and the operation type includes: Based on the library name and table name, candidate semantic dictionaries are obtained by filtering from at least one business semantic dictionary; Based on the operation type, a target triggering condition is determined from at least one triggering condition in the candidate semantic dictionary.

7. The method according to claim 5, characterized in that, The step of determining whether the column data meets the target triggering condition includes: Determine at least one column field to be extracted and a field threshold corresponding to the column field to be extracted from the target triggering conditions; Obtain the second field value of the at least one column field to be extracted from the column data; The value of the second field is compared with the field threshold to obtain the comparison result; If at least one second field value satisfies the corresponding field threshold, then the comparison result is determined to be that the column data satisfies the target triggering condition; If a second field value does not meet the corresponding field threshold, then the comparison result is determined to be that the column data does not meet the target triggering condition.

8. A data processing apparatus, characterized in that, The device includes: The data acquisition module is used to acquire and parse the logs on the database disk to obtain at least one row-level incremental data; the row-level incremental data includes the transaction number; The logical primary key determination module is used to determine the logical primary key of the row-level incremental data from the data source table set; The data storage module is used to store the row-level incremental data into the transaction context registration table corresponding to the transaction number based on the logical primary key; The event generation module is used to generate standardized business events based on the transaction context registration table; The data calculation module is used to calculate the standardized business events to obtain the target business data corresponding to the business event code; the business event code uniquely identifies the standardized business event.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the data processing method according to any one of claims 1-7.

10. A computer program product comprising a computer program that, when executed by a processor, implements the data processing method according to any one of claims 1-7.