Data processing method and device, equipment and storage medium
By performing migration processing and consistency comparison on multi-source heterogeneous data, the problem of data migration and comparison in the financial system is solved, thereby improving system security and processing efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MINSHENG BANKING CORP
- Filing Date
- 2022-08-16
- Publication Date
- 2026-05-26
AI Technical Summary
Existing technologies have failed to effectively address the problem of migrating and comparing multi-source heterogeneous data in financial systems, making it impossible to verify the security of financial systems.
By acquiring data to be processed from heterogeneous data sources, performing migration processing according to preset rules, generating migration data, and comparing the data to be processed with the migration data to determine data consistency, the security of the financial system is verified.
It enables effective migration and comparison of multi-source heterogeneous data, improving the security and data processing efficiency of the financial system.
Smart Images

Figure CN115344556B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the financial field and the data processing field, and in particular to a data processing method, apparatus, device and storage medium. Background Technology
[0002] With the rapid development of the mobile internet, more users have gained a better understanding of financial services through various forms and channels. Therefore, how to leverage financial data to provide financial services and serve financial customers has become a crucial aspect of financial data utilization. The development of financial systems can effectively utilize financial data to provide financial services and serve financial customers.
[0003] In related technologies, financial systems consist of multiple subsystems that can be interconnected and exchange data. During data transmission and aggregation, data transformation processing can verify the correctness, validity, and consistency of the data, thereby verifying the security of the financial system. However, no solutions have been reported in related technologies for migrating and comparing multi-source heterogeneous data. Therefore, when multi-source heterogeneous data is generated in a financial system, the security of the financial system cannot be verified. Summary of the Invention
[0004] This application provides a data processing method, apparatus, device, and storage medium for verifying the security of a financial system by migrating and comparing multi-source heterogeneous data.
[0005] In a first aspect, this application provides a data processing method, comprising: acquiring data to be processed from a heterogeneous data source; performing migration processing on the target field of the data to be processed according to a preset rule to obtain migrated data; and comparing the data to be processed with the migrated data to obtain a comparison result, wherein the comparison result is used to determine whether the data to be processed is consistent with the migrated data.
[0006] In one possible implementation, the target field of the data to be processed is migrated according to preset rules to obtain migrated data, including: performing rule configuration processing on the target field of the data to be processed according to preset rules to obtain rule data; and performing aggregation processing on the rule data according to preset tasks to obtain migrated data.
[0007] In one possible implementation, the data to be processed and the migrated data are compared to obtain a comparison result, including: searching for first data in the data to be processed and searching for second data in the migrated data according to a preset index; if the first data exists in the data to be processed and the second data exists in the migrated data, then the first rule carried by the first data and the second rule carried by the second data are compared to determine whether the first rule and the second rule are the same; if the first rule and the second rule are the same, then the first data and the second data are recorded as consistent in the comparison result.
[0008] In one possible implementation, the method further includes: if the first rule and the second rule are different, traversing each rule in the first rule and the second rule; recording the first data and the second data that have the same rule in the comparison result, as well as the number of the first data and the number of the second data.
[0009] In one possible implementation, the method further includes: if the first data is not present in the data to be processed, and / or the second data is not present in the migration data, then relevant information is recorded in the comparison result. The relevant information includes that the first data is not present in the data to be processed and the second data is present in the migration data, or the first data is present in the data to be processed and the second data is not present in the migration data; or the first data is not present in the data to be processed and the second data is not present in the migration data.
[0010] In one possible implementation, searching for first data in the data to be processed and searching for second data in the migration data according to a preset index includes: generating a migration data table containing the migration data; classifying preset rules to obtain multiple rules; traversing the migration data table using multiple rules to obtain the data size value of the migration data table; if the data size value is less than a preset value, performing a serial inter-table search operation on the migration data table and the data to be processed containing the data to be processed according to the preset index; if the data size value is greater than or equal to the preset value, performing a serial inter-table search operation on the migration data table and the data to be processed according to the preset index, and performing a parallel intra-table search operation.
[0011] In one possible implementation, after obtaining the data to be processed from the heterogeneous data source, the method further includes: performing verification processing on the data to be processed according to the verification rules.
[0012] Secondly, this application provides a data processing apparatus, comprising: an acquisition module for acquiring data to be processed from a heterogeneous data source; a migration module for performing migration processing on the target field of the data to be processed according to preset rules to obtain migrated data; and a comparison module for comparing the data to be processed with the migrated data to obtain a comparison result, wherein the comparison result is used to determine whether the data to be processed is consistent with the migrated data.
[0013] Thirdly, this application provides an electronic device, including: a processor and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the data processing method of the first aspect.
[0014] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the data processing method as described in the first aspect.
[0015] Fifthly, embodiments of this application provide a computer program product, including a computer program, which, when executed by a processor, implements the data processing method of the first aspect.
[0016] The data processing method, apparatus, equipment, and storage medium provided in this application, after obtaining the data to be processed from the heterogeneous data source, perform migration processing on the data to be processed according to preset rules. After obtaining the migrated data, the migrated data and the data to be processed are compared for consistency, thereby determining the consistency comparison result between the data to be processed and the migrated data. In this way, the security of the financial system can be determined based on the consistency comparison result. Attached Figure Description
[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0018] Figure 1 A schematic diagram of the structure of a data processing system provided in an embodiment of this application;
[0019] Figure 2 A flowchart illustrating the data processing method provided in the embodiments of this application;
[0020] Figure 3 This is a schematic diagram illustrating the rule types and migration logic provided in the embodiments of this application;
[0021] Figure 4 This is a schematic diagram illustrating the task types and aggregation logic provided in the embodiments of this application;
[0022] Figure 5 This is a schematic diagram illustrating data verification as provided in an embodiment of this application.
[0023] Figure 6 This is a schematic diagram of the structure of a data processing device provided in an embodiment of this application;
[0024] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0025] The accompanying drawings have illustrated specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concepts of this application to those skilled in the art through reference to specific embodiments. Other drawings can be obtained from these drawings by those skilled in the art without any inventive effort. Detailed Implementation
[0026] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0027] First, let me explain the terms used in this application:
[0028] Multi-source heterogeneous data: Data generated from multiple heterogeneous data sources. A heterogeneous data source refers to multiple data sources with different data structures, access methods, and formats.
[0029] The related technologies provided in the background section have at least the following technical problems:
[0030] With the rapid development and increasing maturity of the financial market, the quantity, scale, profitability, and innovation speed of banking services and products have all greatly improved. Furthermore, with the rapid development of the mobile internet, more users have gained a better understanding of financial services through various forms and channels. Therefore, how to leverage financial data to provide financial services and serve financial customers has become the most important aspect of banks' use of financial data.
[0031] Typically, a financial system can consist of multiple subsystems that are interconnected and exchange data. The data transmission and aggregation process involves multiple modules, multiple database types, data migration, and data transformation. Data transformation and processing can verify the correctness, validity, and consistency of the data, thereby verifying the security of the financial system. However, relevant technologies have not yet reported solutions for database sharding normalization for multi-source heterogeneous data sources, and effective methods for migrating and comparing multi-source heterogeneous data have not been provided.
[0032] To address the aforementioned issues, this application proposes a data processing method. After obtaining the data to be processed from a heterogeneous data source, the data to be processed is migrated according to preset rules. After obtaining the migrated data, the migrated data and the data to be processed are compared for consistency, thereby determining the consistency comparison result between the data to be processed and the migrated data. In this way, the security of the financial system can be determined based on the consistency comparison result.
[0033] In one embodiment, the data processing method can be applied in an application scenario. Figure 1 This is a schematic diagram of the structure of a data processing system provided in an embodiment of this application, such as... Figure 1As shown, the data processing system includes a data source subsystem, a rule set subsystem, a detection subsystem, a task subsystem, and a report subsystem.
[0034] In the above scenario, the data source subsystem can be used to specify relevant fields of the data in the data source, such as database address and / or file address; the detection subsystem can be used to detect the validity of data rules in the rule set subsystem and data in the data source subsystem; the rule set subsystem can be used to specify rules for relevant fields of the data specified by the data source subsystem; the task subsystem can be used to batch execute the migration logic composed of the data source and rules; and the reporting subsystem can be used to compare the consistency of the data before and after migration and provide comparison conclusions.
[0035] In the above scenario, the data source subsystem can specify relevant fields such as database and / or file address, port, account, password, database name, database type, and file type; the detection subsystem can detect the validity of data rules in the rule set subsystem and data in the data source subsystem, including performing directory verification, name verification, character set verification, delimiter verification, and data volume verification; the rule set subsystem can specify the rule set and migration method for relevant fields specified by the data source subsystem, representing the mapping relationship of corresponding fields in the old and new data tables on a field-by-field basis, which can include old and new table names, indexes, field names, filter conditions, rules, etc., where rules can include: translation, mapping, truncation, padding, operation, conversion, etc.; the task subsystem can batch execute the migration logic composed of data sources and rules, by selecting the data source in the data source subsystem and the rule set in the rule set subsystem, and setting an error threshold, the migration logic of rules that meet the error threshold will stop immediately; the reporting subsystem can provide conclusions on the consistency comparison results of data before and after migration, including: statistical data by table dimension, statistical data by rule (field) dimension, error details, etc.
[0036] Specifically, in the above scenario, the data source subsystem can connect to multiple heterogeneous data sources, such as text and databases, to receive data to be processed from these sources. The data source subsystem can also connect to the rule set subsystem to classify the data received by the data source subsystem according to rules. The rule set subsystem classifies the data to be processed according to its source, setting rules for each target field (or data item) of each heterogeneous data source. The rules for the target fields under each heterogeneous data source can be the same or completely different. The task subsystem can perform task aggregation operations on each set of rules for each piece of data to be processed. The task subsystem can include task executors, which can batch execute the migration logic composed of data sources and rules.
[0037] In the above scenario, based on the reporting operations, the reporting subsystem can output standardized report results for the data tables corresponding to each task. These standardized report results can include consistency comparison results for the data corresponding to each task. These consistency comparison results can be used to reflect the migration result of the data for the corresponding task, that is, whether the migration was successful or failed. Therefore, the security of the financial system can be determined based on the migration result reflected in the consistency comparison results; that is, successful migration indicates a high level of security for the financial system, while migration failure indicates a low level of security.
[0038] In light of the above scenarios, the technical solutions of this application and how they solve the aforementioned technical problems will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0039] This application provides a data processing method. Figure 2 A flowchart of the data processing method provided in the embodiments of this application is shown below. Figure 2 As shown, the method includes the following steps:
[0040] S201: Obtain the data to be processed from the heterogeneous data source.
[0041] In this step, there can be multiple heterogeneous data sources, which can have different structures and may include different data presentation types. For example, the data to be processed from a file data source may be presented as text, while the data to be processed from a database data source may be presented as database data. Optionally, different types of data to be processed can have different migration and comparison methods. At least one heterogeneous data source is included, and heterogeneous data sources can also be referred to as financial data sources.
[0042] Optionally, heterogeneous data sources have the function of storing data. Therefore, it is generally accepted that, in the embodiments of this application, heterogeneous data sources may include text data sources and database data sources. Text data sources may include txt type, csv type, etc., and database data sources may include mysql type, gp type, hive type, etc.
[0043] Optionally, after generating financial data, multiple heterogeneous data sources can communicate with the testing system through a distributed network. Since the structures of these multiple heterogeneous data sources are different, in this embodiment, they can be referred to as distributed multi-source heterogeneous data sources. Figure 1 The data source subsystem can be used to receive unprocessed data (or financial data) generated by distributed, multi-source, heterogeneous data sources.
[0044] Optionally, the data to be processed can be data stored in a file and / or data stored in a database, wherein the file type and / or database type can be specified.
[0045] Optionally, data to be processed can be periodically retrieved from heterogeneous data sources according to a preset period.
[0046] S202: According to preset rules, perform migration processing on the target field of the data to be processed to obtain migrated data.
[0047] In this step, the target field in the data to be processed can be specified in advance. The target field can be one or more fields in the data to be processed, or it can be every field in the data to be processed. The preset rules can include translation, mapping, truncation, padding, operation, transformation, etc. You can select one of these rules to perform migration processing on the target field of the data to be processed, thereby obtaining the migrated data.
[0048] S203: Compare the data to be processed with the migrated data to obtain the comparison results.
[0049] In this step, the comparison results are used to determine whether the data to be processed is consistent with the migration data.
[0050] Specifically, after obtaining the migration data, the data to be processed can be compared with the migration data to obtain the consistency comparison result. This comparison result can include the conclusion of whether the data to be processed has been successfully migrated. If the data to be processed is consistent with the migration data, it can be said that the data to be processed has been successfully migrated. If the data to be processed is inconsistent with the migration data, it can be said that the data to be processed has not been successfully migrated.
[0051] Specifically, if the data to be processed is consistent with the data being migrated, it can be determined that the financial system is highly secure; conversely, if the data to be processed is inconsistent with the data being migrated, it can be determined that the financial system is less secure.
[0052] The data processing method provided in this embodiment, after obtaining the data to be processed from the heterogeneous data source, performs migration processing on the data to be processed according to preset rules. After obtaining the migrated data, the migrated data and the data to be processed are compared for consistency, thereby determining the consistency comparison result between the data to be processed and the migrated data. In this way, the security of the financial system can be determined based on the consistency comparison result.
[0053] In one embodiment, migration processing is performed on the target field of the data to be processed according to preset rules to obtain migrated data, including: performing rule configuration processing on the target field of the data to be processed according to preset rules to obtain rule data; and performing aggregation processing on the rule data according to preset tasks to obtain migrated data.
[0054] In this scheme, migration processing can include rule configuration processing and aggregation processing. The target field in the data to be processed can also be the target data item of the data to be processed. A rule can be selected from rules such as translation, mapping, truncation, padding, operation, and transformation to perform rule configuration processing on the target field in the data to be processed, resulting in rule data. For example, configuring translation on the target field yields rule data with the corresponding translated field; or configuring transformation on the target field yields rule data with the corresponding transformed field. Figure 3 As shown. Optionally, each target field can correspond to one rule, and one rule can correspond to multiple target fields.
[0055] Figure 3 This is a schematic diagram of rule types and migration logic provided in the embodiments of this application. Figure 3 In this context, the rule set can include translation, mapping, truncation, padding, operation, and transformation. The data to be processed can include data item 1 from data source 1, data item 2 from data source 1, data item 1 from data source 2, data item 1 from data source 3, data item 1 from data source 5, ..., data item X from data source N. After selecting a rule for each data item in the data to be processed in the rule set and configuring the rule, the corresponding rule data can be obtained. This rule data can include: target data item 1 from data source 1 after padding configuration, target data item 2 from data source 1 after mapping configuration, target data item 3 from data source 2 after transformation configuration, target data item 2 from data source 3 after translation configuration, target data item 5 from data source 1 after translation configuration, ..., target data item X from data source N after configuring the corresponding rule (such as transformation).
[0056] In the above scheme, after obtaining the rule data, in order to improve data processing efficiency, multiple tasks can be set up to group the rule data according to the tasks. After obtaining multiple groups, the rule data of multiple groups can be aggregated to obtain migration data. The preset tasks can be one or more of the set tasks.
[0057] Specifically, each task may include data items from one or more heterogeneous data sources, and one or more rules, such as Figure 4As shown. Optionally, the rule data for each group can correspond to one or more tasks, and each task can correspond to the rule data for one or more groups.
[0058] Figure 4 This is a schematic diagram of task types and aggregation logic provided in the embodiments of this application. Figure 4 In this process, a task set can include multiple tasks, such as Task 1, Task 2, Task 3, Task 4, etc. The rule data is grouped according to tasks, resulting in multiple groups that can include: data item 1 and rule, data item 2 and rule from data source 1; data item 1 and rule, data item 2 and rule from data source 2; data item 1 and rule, data item 2 and rule from data source 3; ..., data item 1 and rule 1, data item 2 and rule 2 from data source N. After aggregating the rule data from multiple groups according to tasks, migration data can be generated. This migration data can be included in the target table.
[0059] Optionally, in Figure 4 In the process of generating migration data, an error threshold can be set to determine errors in the generated migration data. This error threshold can represent the error measurement of each set of rule data. For each set of rule data, if the error threshold is met, the migration processing of that set of rule data stops, and the current error type is investigated. If the error threshold is not met, the migration processing of the next set of rule data continues until each set of rule data is correct, and then the migration processing stops.
[0060] Optionally, by setting up multiple tasks to perform batch migration processing on the data to be processed after completing the rule configuration, the efficiency of migrating the data to be processed can be improved.
[0061] In one embodiment, the data to be processed and the migrated data are compared to obtain a comparison result, including: searching for first data in the data to be processed and searching for second data in the migrated data according to a preset index; if the first data exists in the data to be processed and the second data exists in the migrated data, then the first rule carried by the first data and the second rule carried by the second data are compared to determine whether the first rule and the second rule are the same; if the first rule and the second rule are the same, then the first data and the second data are recorded as consistent in the comparison result.
[0062] In this scheme, after the data to be processed undergoes migration processing, the resulting migrated data corresponds to the data to be processed. The data type or number can remain unchanged. Therefore, the preset index can be the number, type, etc. of the data to be processed and the migrated data. Taking the preset index as the number as an example, the first data is searched in the data to be processed according to the number, and the second data is searched in the migrated data. Then the numbers of the first data and the second data are the same.
[0063] In the above scheme, if the first data and the second data can be found, the rules carried by the first data and the second data can be compared. If they are the same, the comparison result is recorded as the first data and the second data being the same. Then, the search for the next first data and the second data continues in the data to be processed and the migration data, and the rule comparison continues until all data has been compared, or it is determined that the rules carried by the first data and the second data are different. Optionally, the first data and the second data can both include multiple data. For example, the first data is data item 1 and data item 2 in data source 1 of the data to be processed, and the second data is data item 1 and data item 2 in target source 1 of the migration data. The first rule and the second rule can both include multiple rules.
[0064] Optionally, by comparing the data to be processed with the migrated data for consistency, the information recorded in the comparison results can be used to determine whether the data to be processed has been successfully migrated, thereby verifying the security of the financial system.
[0065] In one embodiment, the method further includes: if the first rule and the second rule are different, traversing each rule in the first rule and the second rule; recording the first data and the second data with the same rule, as well as the number of the first data and the number of the second data in the comparison result.
[0066] In this scheme, if it is determined that a certain field (data item) in the first data and the second data carries different rules, then each rule carried in the first data and the second data is traversed and compared to determine and record the first data and the second data with the same rules and the first data and the second data with different rules, as well as record the number of the first data and the second data with the same rules and the number of the first data and the second data with different rules.
[0067] Optionally, by comparing the data to be processed with the migrated data for consistency, the information recorded in the comparison results can be used to determine whether the data to be processed has been successfully migrated, thereby verifying the security of the financial system.
[0068] In one embodiment, the method further includes: if the first data is not present in the data to be processed, and / or the second data is not present in the migration data, then recording relevant information in the comparison result. The relevant information includes that the first data is not present in the data to be processed and the second data is present in the migration data, or the first data is present in the data to be processed and the second data is not present in the migration data; or the first data is not present in the data to be processed and the second data is not present in the migration data.
[0069] In this scheme, when searching for first data in the data to be processed and second data in the migrated data according to a preset index, if the first data is not found in the data to be processed and / or the second data is not found in the migrated data, the search results can be recorded in the comparison results. This allows for the determination of the migration status of the data to be processed based on the comparison results. Therefore, the information recorded in the comparison results can be used to determine whether the migration of the data to be processed was successful, thereby verifying the security of the financial system.
[0070] In one embodiment, searching for first data in the data to be processed and searching for second data in the migration data according to a preset index includes: generating a migration data table containing the migration data; classifying preset rules to obtain multiple rules; traversing the migration data table using the multiple rules to obtain the data size value of the migration data table; if the data size value is less than a preset value, performing a serial inter-table search operation on the migration data table and the data to be processed containing the data to be processed according to the preset index; if the data size value is greater than or equal to the preset value, performing a serial inter-table search operation on the migration data table and the data to be processed according to the preset index, and performing a parallel intra-table search operation.
[0071] In this scheme, a pending data table containing the data to be processed and a migration data table containing the migration data can be generated. When searching for the first data in the pending data and the second data in the migration data, since the pending data and migration data are stored in target files in the target directory or in target tables in the target database, the rule configuration can be loaded and parsed first. Then, for the pending data table, multiple fields can be categorized according to preset rules, and then compared with the relevant fields in the migration data table. The size value of the migration data table is then obtained and compared with a preset value. If the size value of the migration data table is less than the preset value, a serial inter-table lookup operation can be performed between the migration data table and the pending data table. If the size value of the migration data table is greater than or equal to the preset value, a serial inter-table lookup operation and a parallel intra-table lookup operation can be performed, thereby improving the efficiency of data retrieval and further improving the efficiency of consistency comparison between the pending data and the migration data.
[0072] Optionally, if the size of the migration data table is greater than or equal to a preset value, the migration data table and the data table to be processed can be divided into content blocks. Then, the number of blocks is calculated: query the total size and total number of rows of the entire table; the number of blocks = total size / preset value + 1. Next, determine the number of data rows in each block: the number of data rows in each block = total number of rows / number of blocks. Determine the range of each block: sort the data in each table according to the preset index and find the boundary points of each block. Finally, read the data from each table according to the number of blocks. Alternatively, a consistency comparison can be performed between the data to be processed and the migration data based on the number of blocks. When performing content block processing, content block division can be based on the backend processing capacity of the data source system. Optionally, the comparison results can also include the block division method of the data to be processed, facilitating the analysis of comparison results for blocks with the same block division method.
[0073] Optionally, the record report can be updated after each table comparison is completed.
[0074] Optionally, the comparison results may also include a security risk value for the financial system, which can be used to indicate the risk level of the data to be processed in the corresponding block to the back-end processing capacity of the financial system.
[0075] In one embodiment, after obtaining the data to be processed from the heterogeneous data source, the method further includes: performing verification processing on the data to be processed according to the verification rules.
[0076] In this scheme, heterogeneous data sources can also be called financial data sources, and the data to be processed can also be called financial data. It's important to note that, typically, each financial data source generates financial data periodically; this is the most common way financial data is generated—daily or periodically. Financial data generated from the same financial data source can be called standardized financial data. This type of financial data can be directly used for standardized migration and consistency comparison within the financial system without any processing.
[0077] However, non-standardized financial data, i.e., partial or erroneous financial data generated from a specific financial data source at a particular time, refers to a situation where, within a predetermined time period, the process of generating financial data did not occur at the same financial data source; only the generation operation was performed, but with very little or no financial data generated, or with erroneous data. Typically, this type of financial data is incomplete, and may even be invalid, making it unsuitable for standardized migration and consistency comparison within the financial system. In some cases, it may even be judged as data anomaly by the financial system. Furthermore, the inventors have found that this situation occurs frequently in practice and is consistent with reality, thus warranting inclusion in the scope of standardized migration and consistency comparison. Therefore, validation processing can be performed on the data to be processed to identify invalid data.
[0078] Therefore, valid data can be filtered out from the data to be processed according to the verification rules. Subsequently, the valid data can be migrated and compared with the migrated data. This can reduce the amount of data processing, improve the efficiency of data migration and comparison, and improve the accuracy of the comparison results.
[0079] Specifically, the verification process for the data to be processed can be performed according to... Figure 5 The diagram shown is executed. Figure 5 This is a schematic diagram illustrating data verification provided in an embodiment of this application, such as... Figure 5 As shown, after obtaining the data to be processed, the directory of the data to be processed can be validated first. If the directory of the data to be processed is not empty, the name of the data to be processed is validated. If the name of the data to be processed is valid, the character set of the data to be processed is validated. If the character set of the data to be processed matches the preset character set, the delimiter of the data to be processed is validated. If the delimiter of the data to be processed is valid, the data volume of the data to be processed can be validated. If the data volume of the data to be processed meets the preset data volume, the validation result is that the data to be processed is valid data.
[0080] Optionally, when performing validation on the data to be processed, data with empty directories, invalid names, mismatched character sets, and invalid delimiters can be filtered out to obtain valid data.
[0081] Optionally, the data to be processed can be validated for each field.
[0082] Overall, the data processing method provided in this embodiment can perform complete migration and consistency comparison of financial data generated by distributed heterogeneous data sources under different circumstances, thereby improving the security and data processing efficiency of the financial system.
[0083] This application also provides a data processing apparatus. Figure 6 This is a schematic diagram of the structure of a data processing device provided in an embodiment of this application, as shown below. Figure 6 As shown, the data processing device 600 includes:
[0084] The acquisition module 601 is used to acquire data to be processed from heterogeneous data sources;
[0085] Migration module 602 is used to perform migration processing on the target field of the data to be processed according to preset rules to obtain migrated data;
[0086] The comparison module 603 is used to compare the data to be processed with the migration data to obtain the comparison result, which is used to determine whether the data to be processed is consistent with the migration data.
[0087] Optionally, when the migration module 602 performs migration processing on the target field of the data to be processed according to preset rules to obtain migration data, it can be specifically used for: performing rule configuration processing on the target field of the data to be processed according to preset rules to obtain rule data; and performing aggregation processing on the rule data according to preset tasks to obtain migration data.
[0088] Optionally, when the comparison module 603 compares the data to be processed with the migrated data to obtain the comparison result, it can specifically be used to: search for first data in the data to be processed according to a preset index, and search for second data in the migrated data; if the first data exists in the data to be processed and the second data exists in the migrated data, then compare the first rule carried by the first data and the second rule carried by the second data to determine whether the first rule and the second rule are the same; if the first rule and the second rule are the same, then record in the comparison result that the first data and the second data are consistent.
[0089] Optionally, the data processing device 600 further includes a first processing module (not shown), which may be specifically used to: if the first rule and the second rule are different, traverse each rule in the first rule and the second rule; and record the first data and the second data with the same rule, as well as the number of the first data and the number of the second data in the comparison result.
[0090] Optionally, the data processing device 600 further includes a second processing module (not shown), which may be specifically used to: if the first data is not present in the data to be processed, and / or the second data is not present in the migrated data, record relevant information in the comparison result. The relevant information includes that the first data is not present in the data to be processed and the second data is present in the migrated data, or the first data is present in the data to be processed and the second data is not present in the migrated data; or the first data is not present in the data to be processed and the second data is not present in the migrated data.
[0091] Optionally, when the comparison module 603 searches for the first data in the data to be processed and the second data in the migration data according to the preset index, it can specifically be used to: generate a migration data table containing the migration data; classify the preset rules to obtain multiple rules; traverse the migration data table using multiple rules to obtain the data size value of the migration data table; if the data size value is less than the preset value, perform a serial inter-table search operation on the migration data table and the data to be processed containing the data to be processed according to the preset index; if the data size value is greater than or equal to the preset value, perform a serial inter-table search operation on the migration data table and the data to be processed according to the preset index, and perform a parallel intra-table search operation.
[0092] Optionally, the data processing device 600 further includes a verification module (not shown), which can be specifically used to: after acquiring the data to be processed from the heterogeneous data source, perform verification processing on the data to be processed according to the verification rules.
[0093] The data processing device provided in this embodiment is used to execute the technical solution of the data processing method in the aforementioned method embodiment. Its implementation principle and technical effect are similar, and will not be described again here.
[0094] This application also provides an electronic device. Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application, such as... Figure 7 As shown, the electronic device 700 includes:
[0095] The processor 711, the memory 712 which is communicatively connected to the processor 711, and the interaction interface 713;
[0096] The memory 712 is used to store computer-executable instructions that can be executed by the processor 711;
[0097] The processor 711 is configured to execute computer instructions stored in the execution memory 712 to implement the above-mentioned data processing method.
[0098] In the aforementioned electronic device 700, the memory 712, processor 711, and interaction interface 713 are electrically connected directly or indirectly to achieve data transmission or interaction. For example, these components can be electrically connected to each other via one or more communication buses or signal lines, such as a bus connection. The memory 712 stores computer-executable instructions for implementing data processing methods, including at least one software functional module that can be stored in the memory in the form of software or firmware. The processor 711 executes various functional applications and data processing by running the software programs and modules stored in the memory 712.
[0099] The memory can be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), and Electrically Erasable Programmable Read-Only Memory (EEPROM). The memory stores programs, which are then executed by the processor upon receiving execution instructions. Furthermore, the software programs and modules within the memory may include an operating system, which can include various software components and / or drivers for managing system tasks (e.g., memory management, storage device control, power management), and can communicate with various hardware or software components to provide an operating environment for other software components.
[0100] A processor can be an integrated circuit chip with signal processing capabilities. The aforementioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor.
[0101] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the technical solution of the data processing method provided in the foregoing method embodiments.
[0102] This application also provides a computer program product, including a computer program, which, when executed by a processor, is used to implement the technical solution of the data processing method provided in the foregoing method embodiments.
[0103] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the claims.
[0104] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A data processing method, characterized in that, include: Retrieve data to be processed from heterogeneous data sources; According to preset rules, the target fields of the data to be processed are migrated to obtain migrated data; The preset rules include translation, mapping, truncation, padding, operation, and transformation; The data to be processed is compared with the migration data to obtain a comparison result, which is used to determine whether the data to be processed is consistent with the migration data. The process of migrating the target field of the data to be processed according to preset rules to obtain migrated data includes: According to the preset rules, the target fields of the data to be processed are configured to obtain rule data. The rule data is aggregated according to a preset task to obtain the migration data; the preset task is one or more of a set set tasks. The step of comparing the data to be processed with the migrated data to obtain a comparison result includes: The first data is searched in the data to be processed according to a preset index, and the second data is searched in the migration data. If the first data exists in the data to be processed and the second data exists in the migration data, then the first rule carried by the first data and the second rule carried by the second data are compared to determine whether the first rule and the second rule are the same. If the first rule and the second rule are the same, then the comparison results will record that the first data and the second data are consistent.
2. The data processing method according to claim 1, characterized in that, Also includes: If the first rule and the second rule are different, then iterate through each of the first rule and the second rule; The comparison results record the first and second data with the same rules, as well as the number of the first and second data.
3. The data processing method according to claim 1, characterized in that, Also includes: If the first data is not present in the data to be processed, and / or the second data is not present in the migration data, then relevant information is recorded in the comparison result. The relevant information includes the absence of the first data in the data to be processed and the presence of the second data in the migration data, or the presence of the first data in the data to be processed and the absence of the second data in the migration data. Alternatively, the first data may not exist in the data to be processed, and the second data may not exist in the migration data.
4. The data processing method according to any one of claims 1 to 3, characterized in that, The step of searching for first data in the data to be processed according to a preset index and searching for second data in the migrated data includes: Generate a migration data table containing the migration data; The preset rules are categorized to obtain multiple rules; The migration data table is traversed using the aforementioned multiple rules to obtain the data size value of the migration data table; If the data size is less than a preset value, then perform a serial inter-table lookup operation on the migration data table and the data table to be processed containing the data to be processed according to the preset index. If the data size is greater than or equal to the preset value, then a serial inter-table lookup operation is performed on the migration data table and the data table to be processed according to the preset index, and a parallel intra-table lookup operation is performed.
5. The data processing method according to claim 1, characterized in that, After obtaining the data to be processed from the heterogeneous data source, the process further includes: The data to be processed is verified according to the verification rules.
6. A data processing apparatus, characterized in that, include: The acquisition module is used to acquire data to be processed from heterogeneous data sources; The migration module is used to perform migration processing on the target field of the data to be processed according to preset rules to obtain migrated data; The preset rules include translation, mapping, truncation, padding, operation, and transformation; The comparison module is used to compare the data to be processed with the migration data to obtain a comparison result, which is used to determine whether the data to be processed is consistent with the migration data. The migration module is specifically used for: According to the preset rules, the target fields of the data to be processed are configured to obtain rule data. The rule data is aggregated according to a preset task to obtain the migration data; the preset task is one or more of a set set tasks. The comparison module is specifically used for: The first data is searched in the data to be processed according to a preset index, and the second data is searched in the migration data. If the first data exists in the data to be processed and the second data exists in the migration data, then the first rule carried by the first data and the second rule carried by the second data are compared to determine whether the first rule and the second rule are the same. If the first rule and the second rule are the same, then the comparison results will record that the first data and the second data are consistent.
7. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the data processing method as described in any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the data processing method as described in any one of claims 1 to 5.