Historical transaction data backtracking and additional recording method and device

By generating multimodal feature data fingerprints and performing distributed comparison and de-heavy, the data missed crawling problem caused by list changes is solved, automatic backtracking and efficient supplementary recording of historical transaction data is realized, and the accuracy and efficiency of data extraction are improved.

CN120371888APending Publication Date: 2025-07-25CHINA CONSTRUCTION BANK +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510317489.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, data extraction is incomplete due to untimely list management or the need to adjust the scope of business data crawling, manual identification and re-entry work is large and easy to miss, affecting the quality of regulatory submissions.

Method used

By generating multimodal feature fingerprints of historical business data and new business data, distributed computing and multimodal feature comparison and dehydration are used, and they are automatically traced back and loaded into the data reporting system.

Benefits of technology

It realizes automatic backtracking and efficient re-entry of historical transaction data, improves the accuracy and efficiency of data extraction, and reduces the risks of manual intervention and missed crawling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371888A_ABST
    Figure CN120371888A_ABST
Patent Text Reader

Abstract

The invention discloses a historical transaction data backtracking and additional recording method and device which can be used in the technical field of information security, and the method comprises the steps: obtaining at least one to-be-backtracked list which comprises a part which changes compared with a historical list and a predefined part which needs to adjust the business data processing range, the to-be-backtracked list comprises a backtracking starting date; forming a backtracking task for each list to be backtracked, and executing all backtracking tasks; generating a data fingerprint of the historical business data according to the field value of the multi-modal feature of the backtracked historical business data; comparing the data fingerprints of the historical business data with the data fingerprints of the stock business data for duplicate removal; obtaining newly-added business data, and generating a data fingerprint of the newly-added business data according to the field value of the multi-modal feature of the newly-added business data; and loading the historical business data and the newly added business data after comparison and deduplication to a data submission system. According to the invention, effective backtracking and efficient additional recording can be carried out on historical transaction data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of transaction data processing, and in particular to a method and device for retrospectively supplementing historical transaction data. Background Art

[0002] This section aims to provide background or context for the embodiments of the present invention described in the claims. The descriptions herein are not admitted to be prior art merely because they are included in this section.

[0003] For a system that maintains list information and uses the list information to process relevant business data from other data sources for subsequent data management or data reporting, the data extraction method is mainly to extract relevant business data at the batch running point from the upstream at the end of each day through the list information at the time of day cut.

[0004] Taking the related party transaction business scenario as an example, first, business personnel maintain a list of related parties in the system. At the end of each day, based on this maintained list of related parties, all related party transactions that occur within the list and fall within the scope defined by supervision are captured from the upstream business system and loaded into the related party transaction system database for subsequent related party transaction management and supervision report processing.

[0005] However, during the daily business operations, data extraction may be incomplete due to untimely maintenance of the list or the need to adjust the scope of business data capture for a certain past period.

[0006] For such a business data processing and management system based on list management, the sources of business data are relatively numerous, and high requirements are placed on the timeliness and accuracy of subsequent data processing results. If manual identification and supplementary entry are performed, the workload is large and it is easy to miss. For example, related party transactions involve many products, a wide range of transaction types, complex rules, and multiple regulatory requirements, which will lead to difficulties in manual identification, a large workload for supplementary entry, and may cause delays and incompleteness in regulatory reporting, affecting the quality of regulatory reporting. Summary of the Invention

[0007] Embodiments of the present invention provide a method for retrospectively supplementing historical transaction data to effectively retrospectively and efficiently supplement historical transaction data. The method includes:

[0008] Obtain at least one list to be retrospectively processed, where the list to be retrospectively processed includes the part that has changed compared with the historical list and the part that is predefined to need to adjust the scope of business data processing, and the list to be retrospectively processed includes the retrospective start date;

[0009] Generate retrospective tasks for each list to be retrospectively processed, and execute all the retrospective tasks. Each retrospective task is used to retrospectively process the historical business data between the retrospective start date and the current date;

[0010] Generate a data fingerprint of historical business data based on the field values of the multimodal features of the traced historical business data;

[0011] Compare and deduplicate the data fingerprint of the historical business data with the data fingerprint of the existing business data already loaded into the data submission system;

[0012] Obtain new business data, and generate a data fingerprint of the new business data based on the field values of the multimodal features of the new business data, where the new business data is the business data newly added between the last update date and the current date;

[0013] Load the historical business data after comparison and deduplication and the new business data into the data submission system.

[0014] The invention embodiment also provides a historical transaction data tracing and supplementing device for effectively tracing and efficiently supplementing historical transaction data. The device includes:

[0015] A to-be-traced list obtaining module for obtaining at least one to-be-traced list, where the to-be-traced list includes the part that has changed compared with the historical list and the part that predefines the need to adjust the processing scope of business data, and the to-be-traced list contains the tracing start date;

[0016] A tracing task execution module for generating a tracing task for each to-be-traced list and executing all the tracing tasks, where each tracing task is used to trace the historical business data between the tracing start date and the current date;

[0017] A comparison and deduplication module for generating a data fingerprint of the historical business data based on the field values of the multimodal features of the traced historical business data; comparing and deduplicating the data fingerprint of the historical business data with the data fingerprint of the existing business data already loaded into the data submission system;

[0018] A submission module for obtaining new business data, generating a data fingerprint of the new business data based on the field values of the multimodal features of the new business data, where the new business data is the business data newly added between the last update date and the current date; loading the historical business data after comparison and deduplication and the new business data into the data submission system.

[0019] The invention embodiment also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above-mentioned historical transaction data tracing and supplementing method is implemented.

[0020] The invention embodiment also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above-mentioned historical transaction data tracing and supplementing method is implemented.

[0021] An embodiment of the present invention further provides a computer program product, which includes a computer program that, when executed by a processor, implements the above-mentioned method for backtracking and supplementing historical transaction data.

[0022] In an embodiment of the present invention, at least one list to be backtracked is obtained. The list to be backtracked includes a part that has changed compared with the historical list and a part that predefines the need to adjust the scope of business data processing. The list to be backtracked contains the backtracking start date. Backtracking tasks are formed for each list to be backtracked, and all backtracking tasks are executed. Each backtracking task is used to backtrack the historical business data between the backtracking start date and the current date. According to the field values of the multi-modal features of the backtracked historical business data, a data fingerprint of the historical business data is generated. The data fingerprint of the historical business data is compared with the data fingerprint of the stock business data that has been loaded into the data reporting system to remove duplicates. New business data is obtained, and according to the field values of the multi-modal features of the new business data, a data fingerprint of the new business data is generated. The new business data is the business data newly added between the last update date and the current date. The historical business data after comparison and duplicate removal and the new business data are loaded into the data reporting system. Through the above steps, the problem of missed capture of business data caused by changes in the closed list in the data extraction work / system based on list management can be solved, and automatic backtracking and capture of historical business data during the period from the backtracking start date to the time of list information change triggered by the list change can be realized. When comparing and removing duplicates, the field values of multi-modal features are used for comparison, thereby effectively improving the efficiency of data backtracking and supplementing. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. In the drawings:

[0024] Figure 1 is a flowchart of the method for backtracking and supplementing historical transaction data in an embodiment of the present invention;

[0025] Figure 2 is a schematic diagram of the automatic backtracking and supplementation of historical business data in an embodiment of the present invention;

[0026] Figure 3 is a flowchart of generating a data fingerprint of business data in an embodiment of the present invention;

[0027] Figure 4 is a flowchart of data fingerprint comparison and duplicate removal in an embodiment of the present invention;

[0028] Figure 5 This is a schematic diagram of the historical transaction data backfilling device in the embodiments of the present invention;

[0029] Figure 6 This is a schematic diagram of the computer device in the embodiments of the present invention. Detailed implementation manners

[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer and more understandable, the following further describes the embodiments of the present invention in detail with reference to the accompanying drawings. Herein, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention.

[0031] In the technical solutions of this application, the acquisition, storage, use, processing, etc. of data all comply with the relevant regulations of national laws and regulations.

[0032] It should be noted that in the embodiments of this application, some industry-existing solutions such as certain software, components, models, etc. may be mentioned. They should be regarded as exemplary. The purpose is only to illustrate the feasibility in the implementation of the technical solutions of this application, but it does not mean that the applicant has already or necessarily used this solution.

[0033] The incomplete data extraction caused by the untimely maintenance of the list or the need to adjust the business data scraping scope in a certain past period is mainly reflected in:

[0034] 1) In the case where the time when a related party fails to be recognized and entered into the system in a timely manner is later than the time when it actually becomes a related party, in the above scenario, the related transactions during the period from the actual effective date of the related party to the time of entry into the system will be missed.

[0035] 2) A related party may be supervised by multiple regulatory authorities. If the regulatory authority to which the related party belongs increases, the related transactions captured historically may not meet the definition scope of the management of related transactions by the newly added regulatory authority to which the related party belongs, and there will also be a risk of incomplete stock related transaction data.

[0036] Regarding the missed extraction of historical business data due to various changes mentioned in the first point, it could only be manually identified by business personnel and then manually entered before.

[0037] Still taking the related transaction scenario as an example, previously, business personnel needed to collect the transaction data that occurred during the period to be backfilled for the related parties with delayed entry from different upstream business systems and enter it into the related transaction system. There are disadvantages as follows:

[0038] 1) It is necessary to manually identify and judge whether each transaction belongs to a related transaction;

[0039] 2) There are many types of products and transactions involved in related-party transactions, and manual labor is required to capture transactions from multiple different upstream business systems.

[0040] 3) The transaction data stored in the upstream system is different from the regulatory reporting requirements. It is also necessary to map the upstream transaction data one by one according to the regulatory reporting requirements, and finally supplement and enter the collected transaction data into the related-party transaction system.

[0041] The above methods are time-consuming and laborious, with high labor costs and large workloads, and it is easy to miss capturing transactions. Therefore, the embodiments of the present invention propose a historical transaction data backtracking and supplementary recording scheme.

[0042] Figure 1 It is a flowchart of the historical transaction data backtracking and supplementary recording method in the embodiments of the present invention, including:

[0043] Step 101, obtain at least one list to be backtracked. The list to be backtracked includes the part that has changed compared with the historical list and the part that needs to adjust the business data processing scope predefined. The list to be backtracked contains the backtracking start date.

[0044] Step 102, form backtracking tasks for each list to be backtracked, and execute all backtracking tasks. Each backtracking task is used to backtrack the historical business data between the backtracking start date and the current date.

[0045] Step 103, generate a data fingerprint of the historical business data according to the field values of the multi-modal features of the backtracked historical business data.

[0046] Step 104, compare and deduplicate the data fingerprint of the historical business data with the data fingerprint of the stock business data that has been loaded into the data reporting system.

[0047] Step 105, obtain the new business data, and generate a data fingerprint of the new business data according to the field values of the multi-modal features of the new business data. The new business data is the business data newly added between the last update date and the current date.

[0048] Step 106, load the historical business data after comparison and deduplication and the new business data into the data reporting system.

[0049] In the embodiments of the present invention, it can solve the problem of missed capture of business data caused by changes in the related list in the data extraction work / system based on list management, realize automatic backtracking and capture of historical business data from the backtracking start date to the time of list information change triggered by the list change, and use the field values of multi-modal features for comparison during comparison and deduplication, thereby effectively improving the efficiency of data backtracking and supplementary recording.

[0050] The following describes each step in detail.

[0051] In step 101, at least one list to be traced back is obtained. The list to be traced back includes the part that has changed compared with the historical list and the part that is predefined to require adjustment of the business data processing scope. The list to be traced back contains the start date of tracing back.

[0052] According to the real-time business fluctuations and resource load conditions, automatically adjust the business data processing scope in the list to be traced back. For example, when it is detected that the data volume in a certain business area suddenly surges, automatically narrow the tracing back scope of this area to give priority to ensuring the tracing back of key business data; while when the resources are idle, moderately expand the tracing back scope to obtain more comprehensive historical data. This dynamic adjustment mechanism can flexibly optimize the tracing back efficiency in different scenarios and meet the changing needs of the business.

[0053] In step 102, for each list to be traced back, a tracing back task is formed, and all the tracing back tasks are executed. Each tracing back task is used to trace back the historical business data between the start date of tracing back and the current date.

[0054] Figure 2 This is the schematic diagram of the automatic tracing back and supplementary recording of historical business data in the embodiment of the present invention. Taking the daily tracing back as an example, this schematic diagram can form multiple lists to be traced back, and trace back the business data of the upstream business system. The final business database can be the data reporting system.

[0055] Specifically, all the tracing back tasks can be executed in parallel. A multi-core processor or a distributed computing framework can be used to realize the parallel execution of the tracing back tasks; the tracing back tasks are divided according to certain rules (such as business type, time range, etc.) and allocated to multiple computing nodes for simultaneous tracing back operations. Each computing node independently completes a part of the data tracing back task, and finally summarizes the results. For example, when processing large-scale tracing back tasks, through the distributed parallel computing framework, the matching time is shortened from several hours to dozens of minutes, greatly improving the efficiency of data deduplication.

[0056] In addition, an intelligent scheduling algorithm can be introduced. According to the scale of the list to be traced back, data complexity, and system resource status, dynamically allocate the execution priority and execution time window of the tracing back tasks to improve the tracing back efficiency and reduce the waste of system resources. For example, for the list to be traced back with a small scale and simple data, it is preferentially arranged to quickly complete the tracing back task during the system idle period, while for the tracing back task of large-scale complex data, the computing resources are reasonably allocated and executed in stages.

[0057] In one embodiment, the method further includes:

[0058] After all the tracing back tasks are executed, partition the traced back historical business data according to the date, and distribute the historical business data of each partition to different distributed units of the distributed computing platform;

[0059] Generate a data fingerprint of historical business data based on the field values of the multimodal features of the traced historical business data, including:

[0060] For each distributed unit, generate a data fingerprint of historical business data according to the field values of the multimodal features of the traced historical business data in that distributed unit.

[0061] The core idea of the distributed computing platform is to use a partition key to divide data into different blocks, perform calculations in parallel, and then summarize the results. Due to the uniqueness of the occurrence time of a business data, the transaction of different dates is distributed to different distributed units according to the occurrence date of the transaction business data as the partition key, so as to generate fingerprint data efficiently in parallel.

[0062] In step 103, generate a data fingerprint of historical business data according to the field values of the multimodal features of the traced historical business data;

[0063] Figure 3 This is a flowchart for generating a data fingerprint of business data in an embodiment of the present invention. In one embodiment, it further includes:

[0064] Step 301, determine the fields of the multimodal features of the business data according to the type of the business data;

[0065] Step 302, splice the field values of the multimodal features of the business data;

[0066] Step 303, use a hash fusion algorithm to generate a data fingerprint of the business data for the spliced field values of the multimodal features.

[0067] Multiple modal features, such as the timestamp sequence of transaction data, the distribution feature of transaction amounts, and the geographical information associated with transactions, are fused to generate a multimodal data fingerprint. By using a deep learning model to extract and fuse these different modal data, the generated data fingerprint is more unique and representative, greatly improving the accuracy of comparison and deduplication, and is especially suitable for accurately distinguishing similar business data in complex business scenarios.

[0068] The hash fusion algorithm is an algorithm that fuses multiple hash algorithms. Each hash algorithm has its own advantages and disadvantages. By combining different algorithms, the security and uniqueness of the fingerprint can be enhanced. For example, first use the MD5 algorithm to perform a preliminary calculation on the spliced field values of the multimodal features to obtain an intermediate value, and then use the SHA-256 algorithm to perform a secondary operation on this intermediate value to finally generate a data fingerprint. In this way, even in the face of complex data tampering or collision situations, the uniqueness of the data fingerprint can be effectively guaranteed, and the misjudgment probability can be reduced.

[0069] The principle of one of the hash algorithms is given below.

[0070] For the field M of multimodal features of any length after splicing, it is divided into 512-bit information blocks M = {m1, m2,......, m n}, and each information block is further divided into a set of 36-bit information blocks X = {x1, x2,......, x 16}. Each information block calculates a hash value through a logical operation function.

[0071] Define the core calculation function:

[0072]

[0073] The operation logic of the logical operation function h(H, M) (H = {H a , H b , H c , H d}, X = {x1, x2,......, x 16}) is to execute the following function 16 rounds to obtain a new: H = {H a , H b , H c , H d}:

[0074] H a = H b + CLS(F(H b , H c , H d ) + Y k + H a + x i )

[0075] H a = H b + CLS(G(H b , H c , H d ) + Y k + H a + x i )

[0076] H a = H b + CLS(J(H b , H c , H d ) + Y k + H a + x i )

[0077] H a = H b + CLS(I(H b , H c, H d ) + Y k + H a + x i )

[0078] H a = H b + CLS(I(H b , H c , H d ) + Y k + H a + x i )

[0079] where Y k = M q * 16 + k, and CLS is a 32 - bit left - circular shift by s bits.

[0080] Given an initial function value H0 = {H a0 , H b0 , H c0 , H d0}, through the calculation of the hash function recursion H i = h(H i-1 , X i ), the message X = {x1, x2,..., x m} is hashed into the hash value:

[0081] H n = H an || H bn || H cn || H dn

[0082] Business data with consistent field values of multimodal features through the above algorithm has the same unique identifier ID and is stored together.

[0083] In addition, through an intelligent algorithm, according to the business type, transaction characteristics, and data change rules, the fields of multimodal features used to generate data fingerprints can be dynamically determined. In different business scenarios, there are differences in the importance and relevance of the fields of multimodal features. For example, in financial lending business, fields such as loan amount, borrower ID number, and loan term may be fields of multimodal features; while in commodity trading business, commodity number, transaction quantity, and transaction time may be more critical. This algorithm can automatically learn and analyze historical business data, and dynamically adjust the selection of the fields of multimodal features according to the stability, uniqueness, and key degree of the data for business identification, making the data fingerprint more targeted and accurate.

[0084] In one embodiment, the splicing of the field values of multimodal features of business data includes:

[0085] Perform correlation analysis on the fields of multimodal features of business data to obtain the degree of correlation between any two fields;

[0086] Determine the weight of each field in the any two fields according to the comparison result between the degree of correlation between any two fields and multiple degree thresholds;

[0087] Concatenate the field values of the multimodal features of business data according to the weight of each field and the field value.

[0088] Specifically, when concatenating multimodal feature field values, data correlation analysis is introduced. Instead of simply concatenating field values in order, the field values are weighted and concatenated according to the internal correlation logic of business data. For example, in e-commerce transaction data, there are different degrees of correlations between transaction amount and fields such as product category and customer purchase frequency. By analyzing these correlation relationships, higher weights are assigned to the field values with close correlations for concatenation, so that the concatenation result can better reflect the core features of the data and improve the accuracy and distinctiveness of the subsequent generated data fingerprint.

[0089] For example, the degree thresholds include a first threshold, a second threshold, and a third threshold, and the comparison rule can be determined according to the actual situation. For example, if the degree of correlation between the transaction amount and the product category is greater than the first threshold, then the weights of these two fields are set to the maximum weight value. If the degree of correlation between the transaction amount and the customer purchase frequency is between the second threshold and the third threshold, and the weight of the transaction amount has been determined, then the weight of the customer purchase frequency is determined to be the medium weight value.

[0090] In step 104, compare and remove duplicates between the data fingerprint of historical business data and the data fingerprint of the stock business data that has been loaded into the data reporting system;

[0091] The stock business data covers business transaction records from the effective date of the related party to various past time points.

[0092] Figure 4 This is the flowchart of data fingerprint comparison and duplicate removal in the embodiments of the present invention. In one embodiment, comparing and removing duplicates between the data fingerprint of historical business data and the data fingerprint of the stock business data that has been loaded into the data reporting system includes:

[0093] Step 401, query the data fingerprint of the stock business data, and compare the fingerprint prefix of the data fingerprint of the historical business data with the data fingerprint of the stock business data;

[0094] Step 402, when the fingerprint prefixes are inconsistent, determine that the historical business data is different from the stock business data;

[0095] Step 403: When the fingerprint prefixes are the same, compare the complete fingerprints of the data fingerprints of historical business data and the data fingerprints of existing business data;

[0096] Step 404: When the complete fingerprints are different, determine that the historical business data is different from the existing business data; otherwise, determine that the historical business data is the same as the existing business data;

[0097] Step 405: Delete the identical historical business data.

[0098] Through the above hierarchical matching strategy, the fingerprint prefix can be quickly matched first. Fingerprints have certain regularity, and the prefix of the data fingerprint can be the first 8 digits. For those with successful fingerprint prefix matching, the complete fingerprints are accurately compared. This hierarchical matching method greatly reduces the number of complete fingerprint comparisons and improves the matching efficiency. For example, in a large amount of retrospective data, more than 90% of the data that is unlikely to be repeated can be quickly excluded through prefix matching, and only the data fingerprints of the remaining 10% are compared for complete fingerprints, thereby reducing the overall amount of data fingerprint matching and greatly improving the retrospective supplementary recording efficiency.

[0099] In one embodiment, the method further includes:

[0100] After generating the data fingerprint of the business data, encrypt the generated data fingerprint;

[0101] Store the encrypted data fingerprint and the unique identifier of the business data in the distributed hash table of the distributed unit;

[0102] Query the data fingerprint of the existing business data, including:

[0103] According to the unique identifier of the existing business data, query the data fingerprint of the existing business data in the distributed hash table.

[0104] The above embodiment is to prevent the data fingerprint from being maliciously tampered with and encrypt the generated data fingerprint. A high-strength encryption algorithm such as AES (Advanced Encryption Standard) can be used to encrypt the data fingerprint and store it in a secure distributed hash table (DHT). In addition, it can also be stored in a database or a distributed storage system. At the same time, when storing, in addition to associating with the unique identifier of the business data, other relevant meta-information such as business type and retrospective time can also be associated according to requirements, which is convenient for subsequent query and management. For example, store the encrypted fingerprint together with the business batch and transaction timestamp to which the corresponding business data belongs in the distributed hash table (DHT), and quickly locate and retrieve the data fingerprint through the business unique identifier.

[0105] In one embodiment, the method further includes:

[0106] Within the first preset period, the unique identifiers and data fingerprints of service data with a recall frequency greater than the first frequency threshold are stored in the cache;

[0107] For the data fingerprints in the cache, if the recall frequency of a data fingerprint is less than the second frequency threshold within the second preset period, the data fingerprint and the corresponding unique identifier are deleted from the cache;

[0108] Query the data fingerprints of the existing service data, including:

[0109] According to the unique identifier of the existing service data, query the data fingerprint of the existing service data in the cache;

[0110] If the query result is empty, query the data fingerprint of the existing service data in the distributed hash table.

[0111] In the above embodiment, it is realized to cache the data fingerprints that are frequently recalled recently and the unique identifiers of the corresponding service data. When new recalled service data arrives, first search and match in the cache. Since the reading speed of the cache is much higher than that of the distributed hash table (which can be stored in a database), the matching speed can be significantly improved. The cache can adopt an elimination strategy based on time and access frequency. For example, fingerprint data that has not been accessed for a certain period of time is removed from the cache; for frequently accessed fingerprint data, its retention time in the cache is extended. At the same time, the cache is updated regularly to ensure that the data in the cache is consistent with the latest data in the database.

[0112] In addition, parallel matching calculation of data fingerprints can be realized. The recalled data is divided according to certain rules (such as service type, time range, etc.) and distributed to multiple computing nodes for simultaneous matching operations. Each computing node independently completes the matching of a part of the data, and finally summarizes the results. For example, when dealing with large-scale matching, through the distributed parallel computing framework, the matching time is shortened from several hours to dozens of minutes, greatly improving the deduplication efficiency.

[0113] In addition, in extremely rare cases, data fingerprint collisions (that is, different service data generate the same fingerprint) may occur. At this time, start the data conflict resolution process. By deeply comparing all fields of the service data, combined with business logic and historical data records, determine the uniqueness and correctness of the data. For example, by querying relevant information such as transaction flow and business approval records, judge whether two pieces of data with the same fingerprint are truly duplicate data or misjudgments caused by fingerprint collisions. If it is a misjudgment, generate unique fingerprints for the two pieces of data again and perform corresponding processing.

[0114] In step 105, new business data is obtained, and a data fingerprint of the new business data is generated according to the field values of the multi-modal features of the new business data, where the new business data is the business data newly added between the last update date and the current date;

[0115] See Figure 2 , the new business data is the newly added part that can be judged without backtracking. A data fingerprint is also generated for this part of the business data, stored and reported, and subsequent stock business data is formed.

[0116] In step 106, the historical business data after comparison and deduplication and the new business data are loaded into the data reporting system.

[0117] When specifically reporting, encrypted transmission can be used to ensure the security of business data.

[0118] An embodiment of the present invention also proposes a historical transaction data backtracking and supplementary recording device, the principle of which is similar to the historical transaction data backtracking and supplementary recording method, and will not be elaborated here.

[0119] Figure 5 FIG. is a schematic diagram of the historical transaction data backtracking and supplementary recording device in an embodiment of the present invention, including:

[0120] A backtracking list acquisition module 501, configured to acquire at least one backtracking list, where the backtracking list includes a part that has changed compared with the historical list and a part that predefines the need to adjust the business data processing scope, and the backtracking list contains the backtracking start date;

[0121] A backtracking task execution module 502, configured to form a backtracking task for each backtracking list and execute all backtracking tasks, and each backtracking task is used to backtrack the historical business data between the backtracking start date and the current date;

[0122] A comparison and deduplication module 503, configured to generate a data fingerprint of the historical business data according to the field values of the multi-modal features of the backtracked historical business data; compare and deduplicate the data fingerprint of the historical business data with the data fingerprint of the stock business data that has been loaded into the data reporting system;

[0123] A reporting module 504, configured to obtain new business data, generate a data fingerprint of the new business data according to the field values of the multi-modal features of the new business data, where the new business data is the business data newly added between the last update date and the current date; load the historical business data after comparison and deduplication and the new business data into the data reporting system.

[0124] In one embodiment, the device further includes a data fingerprint generation module, configured to:

[0125] After all the backtracking tasks are executed, partition the historical business data of the backtracking according to the date, and distribute the historical business data of each partition to different distributed units of the distributed computing platform;

[0126] Generate a data fingerprint of the historical business data according to the field values of the multimodal features of the historical business data of the backtracking, including:

[0127] For each distributed unit, generate a data fingerprint of the historical business data according to the field values of the multimodal features of the historical business data of the backtracking in the distributed unit.

[0128] In one embodiment, the data fingerprint generation module is used for:

[0129] Determine the fields of the multimodal features of the business data according to the type of the business data;

[0130] Concatenate the field values of the multimodal features of the business data;

[0131] Adopt a hash fusion algorithm to generate a data fingerprint of the business data for the concatenated field values of the multimodal features.

[0132] In one embodiment, the data fingerprint generation module is used for:

[0133] Conduct an association relationship analysis on the fields of the multimodal features of the business data to obtain the degree of association relationship between any two fields;

[0134] According to the comparison result of the degree of association relationship between any two fields and multiple degree thresholds, determine the weight of each field in the any two fields;

[0135] Concatenate the field values of the multimodal features of the business data according to the weight and field value of each field.

[0136] In one embodiment, the comparison and deduplication module is used for:

[0137] Query the data fingerprint of the existing business data, and compare the fingerprint prefix of the data fingerprint of the historical business data with the data fingerprint of the existing business data;

[0138] When the fingerprint prefixes are inconsistent, determine that the historical business data is different from the existing business data;

[0139] When the fingerprint prefixes are consistent, compare the data fingerprint of the historical business data with the complete fingerprint of the data fingerprint of the existing business data;

[0140] When the complete fingerprints are inconsistent, determine that the historical business data is different from the existing business data, otherwise, determine that the historical business data is the same as the existing business data;

[0141] Delete the same historical business data.

[0142] In one embodiment, the device further includes an encryption module for:

[0143] After generating the data fingerprint of the business data, encrypt the generated data fingerprint;

[0144] Store the encrypted data fingerprint and the unique identifier of the business data in a distributed hash table;

[0145] The comparison and deduplication module is used for:

[0146] According to the unique identifier of the existing business data, query the data fingerprint of the existing business data in the distributed hash table.

[0147] In one embodiment, the device further includes a cache module for:

[0148] Store the unique identifier and data fingerprint of the business data whose backtracking frequency is greater than the first frequency threshold within the first preset period in the cache;

[0149] For the data fingerprint in the cache, if the backtracking frequency of the data fingerprint is less than the second frequency threshold within the second preset period, delete the data fingerprint and the corresponding unique identifier from the cache;

[0150] The comparison and deduplication module is used for:

[0151] According to the unique identifier of the existing business data, query the data fingerprint of the existing business data in the cache;

[0152] If the query result is empty, query the data fingerprint of the existing business data in the distributed hash table.

[0153] In summary, the method and device proposed in the embodiments of the present invention have the following beneficial effects:

[0154] Obtain at least one list to be retraced, where the list to be retraced includes the part that has changed compared to the historical list and the part that is predefined to adjust the scope of business data processing, and the list to be retraced contains the retrace start date; create retrace tasks for each list to be retraced, and execute all retrace tasks, where each retrace task is used to retrace the historical business data between the retrace start date and the current date; generate a data fingerprint of the historical business data based on the field values of the multimodal features of the retraced historical business data; compare and deduplicate the data fingerprint of the historical business data with the data fingerprint of the existing business data that has been loaded into the data reporting system; obtain new business data, and generate a data fingerprint of the new business data based on the field values of the multimodal features of the new business data, where the new business data is the business data newly added between the last update date and the current date; load the historical business data after comparison and deduplication and the new business data into the data reporting system. Through the above steps, the problem of missed capture of business data caused by changes in the closed list in the data extraction work / system based on list management can be solved, and automatic retrace capture of historical business data during the period from the retrace start date to the time of list information change triggered by the list change can be realized. When comparing and deduplicating, the field values of multimodal features are used for comparison, thereby effectively improving the efficiency of data retracement and backfilling.

[0155] An embodiment of the present invention also provides a computer device, Figure 6 which is a schematic diagram of the computer device in the embodiment of the present invention. The computer device 600 includes a memory 610, a processor 620, and a computer program 630 stored on the memory 610 and executable on the processor 620. When the processor 620 executes the computer program 630, the above-mentioned historical transaction data retracement and backfilling method is implemented.

[0156] An embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned historical transaction data retracement and backfilling method is implemented.

[0157] An embodiment of the present invention also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the above-mentioned historical transaction data retracement and backfilling method is implemented.

[0158] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0159] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general purpose computers, special purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or more blocks.

[0160] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or more blocks.

[0161] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or more blocks.

[0162] The specific embodiments described above further elaborate on the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for retrospectively supplementing historical transaction data, characterized in that, Including: Obtain at least one list to be traced back, where the list to be traced back includes the part that has changed compared with the historical list and the part that predefines the need to adjust the scope of business data processing, and the list to be traced back contains the start date of tracing back; Create a tracing-back task for each list to be traced back, and execute all the tracing-back tasks. Each tracing-back task is used to trace back the historical business data between the start date of tracing back and the current date; Generate a data fingerprint of the historical business data according to the field values of the multimodal features of the traced-back historical business data; Compare and remove duplicates between the data fingerprint of the historical business data and the data fingerprint of the stock business data that has been loaded into the data reporting system; Obtain new business data, and generate a data fingerprint of the new business data according to the field values of the multimodal features of the new business data. The new business data is the business data newly added between the last update date and the current date; Load the historical business data after comparison and duplicate removal and the new business data into the data reporting system.

2. The method according to claim 1, characterized in that, It also includes: After executing all the tracing-back tasks, partition the traced-back historical business data by date, and distribute the historical business data of each partition to different distributed units of the distributed computing platform; Generate a data fingerprint of the historical business data according to the field values of the multimodal features of the traced-back historical business data, including: For each distributed unit, generate a data fingerprint of the historical business data according to the field values of the multimodal features of the traced-back historical business data in this distributed unit.

3. The method according to claim 1, wherein It also includes: Determine the fields of the multimodal features of the business data according to the type of the business data; Concatenate the field values of the multimodal features of the business data; Use a hash fusion algorithm to generate a data fingerprint of the business data for the concatenated field values of the multimodal features.

4. The method according to claim 3, wherein Concatenating the field values of the multimodal features of the business data includes: Analyze the correlation relationship of the fields of the multimodal features of the business data to obtain the degree of correlation between any two fields; Determine the weight of each field in the any two fields according to the comparison result of the degree of correlation between any two fields and multiple degree thresholds; Concatenate the field values of the multimodal features of the business data according to the weight and field value of each field.

5. The method according to claim 1, characterized in that, Comparing and removing duplicates between the data fingerprint of the historical business data and the data fingerprint of the stock business data that has been loaded into the data reporting system includes: Query the data fingerprint of the stock business data, and compare the fingerprint prefix of the data fingerprint of the historical business data with the data fingerprint of the stock business data; When the fingerprint prefixes are inconsistent, determine that the historical business data is different from the stock business data; When the fingerprint prefixes are consistent, compare the complete fingerprint of the data fingerprint of the historical business data with the data fingerprint of the stock business data; When the complete fingerprints are inconsistent, determine that the historical business data is different from the stock business data, otherwise, determine that the historical business data is the same as the stock business data; Delete the same historical business data.

6. The method according to claim 5, characterized in that, It also includes: After generating the data fingerprint of the business data, encrypt the generated data fingerprint; Store the encrypted data fingerprint and the unique identifier of the business data in a distributed hash table; Query the data fingerprint of the inventory business data, including: Query the data fingerprint of the inventory business data in the distributed hash table according to the unique identifier of the inventory business data.

7. The method according to claim 6, wherein It also includes: The unique identifier and data fingerprint of the business data with a recall frequency greater than the first frequency threshold within the first preset period are stored in the cache; For the data fingerprint in the cache, if the recall frequency of the data fingerprint is less than the second frequency threshold within the second preset period, the data fingerprint and the corresponding unique identifier are deleted from the cache; Query the data fingerprint of the inventory business data, including: Query the data fingerprint of the inventory business data in the cache according to the unique identifier of the inventory business data; If the query result is empty, query the data fingerprint of the inventory business data in the distributed hash table.

8. A historical transaction data backtracking and supplementary recording device, characterized in that, It includes: A pending recall list acquisition module for acquiring at least one pending recall list, where the pending recall list includes the part that has changed compared with the historical list and the part that predefines the need to adjust the business data processing scope, and the pending recall list contains the recall start date; A recall task execution module for generating recall tasks for each pending recall list and executing all recall tasks. Each recall task is used to recall the historical business data between the recall start date and the current date; A comparison and deduplication module for generating the data fingerprint of the historical business data according to the field values of the multimodal features of the recalled historical business data; comparing and deduplicating the data fingerprint of the historical business data with the data fingerprint of the inventory business data that has been loaded into the data reporting system; A reporting module for obtaining new business data, generating the data fingerprint of the new business data according to the field values of the multimodal features of the new business data, where the new business data is the business data newly added between the last update date and the current date; loading the historical business data after comparison and deduplication and the new business data into the data reporting system.

9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the method according to any one of claims 1 to 7.

11. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by the processor, it implements the method according to any one of claims 1 to 7.