Transaction data processing method, device, equipment, medium and program product
By splitting large transaction files into sub-files of the same size and importing them into the database in parallel, and adjusting processor utilization, the problem of low import efficiency is solved, and faster data processing time and higher stability are achieved.
Patent Information
- Application Number
- CN202510899051.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-10-03
AI Technical Summary
When importing large amounts of data into a database, there are problems with low import efficiency and long processing time.
The target transaction file is split into multiple sub-files of the same size, and the sub-files are sorted according to the transaction time. After grouping, they are imported into the target database in parallel. By adjusting the processor utilization within the preset range, data is imported in batches, and the transaction commit operation is performed after each batch is completed, and error data is recorded.
It improves the efficiency of data import, reduces processing time, avoids processor resource waste and system crashes, and enhances the stability and reliability of data processing.
Smart Images

Figure CN120743532A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of big data, and specifically to a transaction data processing method, device, equipment, medium and program product. Background Art
[0002] In the related art, when importing large amounts of data into a database, there are problems of low import efficiency and long processing time. Summary of the Invention
[0003] In view of the above problems, the present application provides a transaction data processing method, apparatus, device, medium and program product.
[0004] According to a first aspect of the present application, a transaction data processing method is provided, comprising: splitting a target transaction file into multiple sub-files of the same file size; sorting the multiple sub-files according to the transaction time of the transaction data contained in each sub-file; dividing the multiple sub-files into multiple groups of sub-files based on the arrangement order of the multiple sub-files; each group of sub-files includes at least two sub-files; calling a processor to import the sub-files contained in each group of sub-files into a target database in parallel; wherein, when the processor imports the sub-files contained in each group of sub-files in parallel, the utilization of the processor is within a preset utilization range.
[0005] According to an embodiment of the present application, a calling processor sequentially imports the subfiles contained in each group of subfiles into a target database in parallel, including: taking each subfile in each group of subfiles as a current subfile, and calling the processor to perform the following operations on the current subfile: dividing the transaction data contained in the current subfile into multiple groups of batch data; using a target interface to sequentially import the multiple groups of batch data into the target database through a data stream; wherein, after each group of batch data is imported, the target database is called to perform a transaction commit operation.
[0006] According to an embodiment of the present application, the method further includes: in response to identifying that the import of transaction data in a group of batch data fails, recording the group of batch data in an error register.
[0007] According to an embodiment of the present application, before splitting the target transaction file into multiple sub-files of the same file size, the method also includes: obtaining verification information and feature information of the target transaction file respectively; the verification information includes metadata characterizing the integrity of the target transaction file; and verifying the target transaction file based on the verification information and feature information to obtain a verification result.
[0008] According to an embodiment of the present application, the target transaction file is verified based on the verification information and the characteristic information to obtain a verification result, including: in response to a failure in verification of the target transaction file, re-verifying the target transaction file after waiting for a preset time; wherein the length of the preset time is positively correlated with the number of times the verification of the target transaction file fails.
[0009] According to an embodiment of the present application, after splitting the target transaction file into multiple sub-files of the same file size, the method further includes: deleting the complete target transaction file before splitting.
[0010] According to an embodiment of the present application, the method also includes: obtaining a transaction date interval of the target transaction file; the transaction date interval represents the interval range of the transaction date corresponding to the transaction data of the target transaction file; dividing the target database into multiple partitions according to the transaction date interval; wherein each sub-file is imported into the corresponding partition in the target database.
[0011] A second aspect of the present application provides a transaction data processing device, comprising: a splitting module for splitting a target transaction file into multiple sub-files of the same file size; a sorting module for sorting the multiple sub-files according to the transaction time of the transaction data contained in each sub-file; a division module for dividing the multiple sub-files into multiple groups of sub-files based on the arrangement order of the multiple sub-files; each group of sub-files includes at least two sub-files; an import module for calling a processor to import the sub-files contained in each group of sub-files into a target database in parallel; wherein, when the processor imports the sub-files contained in each group of sub-files in parallel, the utilization of the processor is within a preset utilization range.
[0012] The third aspect of the present application provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.
[0013] The fourth aspect of the present application further provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above method when the computer program or instructions are executed by a processor.
[0014] The fifth aspect of the present application further provides a computer program product, comprising a computer program or instructions, which implement the steps of the above method when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The above contents and other objects, features and advantages of the present application will become more apparent through the following description of the embodiments of the present application with reference to the accompanying drawings, in which:
[0016] Figure 1Schematically illustrates an application scenario diagram of the transaction data processing method, apparatus, device, medium, and program product according to an embodiment of the present application;
[0017] Figure 2 The following schematically shows a flow chart of a transaction data processing method according to an embodiment of the present application;
[0018] Figure 3 The following schematically shows a principle diagram of a transaction data processing method according to an embodiment of the present application;
[0019] Figure 4 A block diagram schematically illustrates a structure of a transaction data processing device according to an embodiment of the present application; and
[0020] Figure 5 A block diagram of an electronic device suitable for implementing a transaction data processing method according to an embodiment of the present application is schematically shown. DETAILED DESCRIPTION
[0021] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present application. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present application. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present application.
[0022] The terms used herein are only for describing specific embodiments and are not intended to limit the present application. The terms "comprise," "include," etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0023] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0024] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0025] In the technical solution of this application, the user information involved (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0026] In some examples, when using Gaussian Database to load merchant net income transaction details list files daily, the file size is extremely large due to the high amount of file data generated daily. When using traditional distributed batch technology to import the file into Gaussian Database, there are problems with low import efficiency and long processing time.
[0027] In view of this, an embodiment of the present application provides a transaction data processing method.
[0028] Figure 1 The application scenario diagram of the transaction data processing method, device, equipment, medium and program product according to the embodiment of the present application is schematically shown.
[0029] like Figure 1 As shown, the application scenario 100 according to this embodiment may include the field of big data. A network 104 is used as a medium for providing a communication link between a first terminal device 101, a second terminal device 102, a third terminal device 103, and a server 105. The network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0030] A user may use a first terminal device 101, a second terminal device 102, or a third terminal device 103 to interact with a server 105 via a network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, or the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).
[0031] The first terminal device 101 , the second terminal device 102 , and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0032] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal devices.
[0033] It should be noted that the transaction data processing method provided in the embodiment of the present application can generally be executed by the server 105. Accordingly, the transaction data processing device provided in the embodiment of the present application can generally be set in the server 105. The transaction data processing method provided in the embodiment of the present application can also be executed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the transaction data processing device provided in the embodiment of the present application can also be set in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.
[0034] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0035] The following will be based on Figure 1 The scene described by Figures 2 to 5 The transaction data processing method according to the embodiment of the present application is described in detail.
[0036] Figure 2 The flowchart of the transaction data processing method according to an embodiment of the present application is schematically shown.
[0037] like Figure 2 As shown, the transaction data processing method of this embodiment includes operations S210 to S240, and the transaction data processing method can be executed by a server.
[0038] In operation S210 , the target transaction file is split into a plurality of sub-files of the same file size.
[0039] In an embodiment of the present application, the target transaction file includes a large amount of transaction data, and the target transaction file needs to be imported into the target database. For example, the target transaction file can be a merchant's net income transaction detail list file.
[0040] Due to the large amount of data in the target transaction file, directly importing the target transaction file into the target database can result in a lengthy import process and a high risk of failure. Therefore, splitting the target transaction file into multiple sub-files of equal size and importing each sub-file separately can avoid these issues.
[0041] For example, for a target transaction file with a file size of 10 GB, the target transaction file can be split into 10 sub-files with a file size of 1 GB each.
[0042] In an embodiment of the present application, a target transaction file may be split into multiple sub-files of the same file size through a SHELL script.
[0043] In operation S220 , the plurality of sub-files are sorted according to the transaction time of the transaction data contained in each sub-file.
[0044] In an embodiment of the present application, multiple subfiles are sorted based on the transaction time of the transaction data contained in each subfile. Each subfile contains multiple transaction data items, each of which includes the transaction time of the transaction. Since the transaction data in the target transaction file is arranged sequentially according to the time of transaction occurrence, the transaction data in each subfile is also arranged sequentially according to time. In other words, the transaction data contained in each subfile is transaction data generated within a time period. The multiple subfiles can be sorted in chronological order based on the earliest transaction time contained in each subfile.
[0045] For example, the target transaction file includes the transaction data of the target merchant within one month. The target transaction file is split into 10 sub-files of the same file size. The 10 sub-files are sorted in chronological order based on the earliest transaction data contained in each sub-file.
[0046] Multiple sub-files are sorted according to the transaction time of the transaction data contained in each sub-file, so that when each sub-file is imported into the target database, it is still imported in the order of the transaction time, which can improve the convenience of subsequent query of transaction data in the target database.
[0047] In operation S230 , the plurality of subfiles are divided into a plurality of groups of subfiles based on the arrangement order of the plurality of subfiles; each group of subfiles includes at least two subfiles.
[0048] When calling a processor to import a sub-file into a target database, due to the small amount of transaction data contained in the sub-file, the processor utilization rate, memory usage rate and other data may be low, resulting in insufficient utilization of the processor.
[0049] In an embodiment of the present application, the subfiles are divided into multiple groups of subfiles based on the order in which the subfiles are arranged, each group of subfiles including at least two subfiles. By grouping the subfiles and importing each group of subfiles into the target database as a unit, processor utilization can be improved.
[0050] For example, for 10 sub-files, the 10 sub-files may be divided into 5 groups, each group including 2 sub-files.
[0051] In operation S240, the processor is called to sequentially import the subfiles contained in each group of subfiles into the target database in parallel. When the processor imports the subfiles contained in each group of subfiles in parallel, the utilization rate of the processor is within a preset utilization rate range.
[0052] In an embodiment of the present application, the calling processor sequentially imports the subfiles contained in each group of subfiles into the target database in parallel. That is, the calling processor sequentially imports the subfiles contained in each group of subfiles into the target database simultaneously, and after the import of the subfiles of the group is completed, the import operation is performed on the next group of subfiles.
[0053] For example, when a group of subfiles includes two subfiles, the processor is called to import the two subfiles into the target database at the same time. After the two subfiles included in the group of subfiles are imported, the processor is called to import the next group of subfiles.
[0054] It should be noted that the number of subfiles included in each subfile group is predetermined based on the amount of transaction data contained in each subfile and the performance of the processor. By adjusting the number of subfiles in each subfile group, the processor utilization remains within a preset utilization range when the processor concurrently imports the subfiles contained in each subfile group. Processor utilization within the preset utilization range indicates that the processor is being efficiently utilized.
[0055] For example, when the processor utilization is low, the number of subfiles included in each group of subfiles may be increased. When the processor utilization is high, the number of subfiles included in each group of subfiles may be reduced.
[0056] In an embodiment of the present application, the preset utilization range may be 70%-85%. By dynamically adjusting the number of sub-files within a group, the processor utilization can be maintained within the preset utilization range, thereby avoiding resource idleness and overload crashes.
[0057] In some examples, when the processor utilization is low (e.g., less than 40%), resources may be wasted. When the processor is under high load (e.g., more than 90%), there is a risk of system crash.
[0058] For example, after adopting the dynamic grouping control of this embodiment, the import time of a 10GB file is shortened from 58 minutes in the traditional method to 22 minutes, and the fluctuation range of processor utilization is stabilized from 40%-95% to 75%-82%.
[0059] In an embodiment of the present application, the method also includes: real-time monitoring of processor utilization; if the utilization is lower than the lower limit of a preset utilization range, increasing the number of sub-files included in the next group of sub-files; if the utilization is higher than the upper limit of the preset utilization range, reducing the number of sub-files included in the next group of sub-files.
[0060] It should be noted that, in this embodiment, the processor may be called to import the sub-file into the target database table, where the target database table refers to a logical data storage unit in the target database.
[0061] By splitting a target file into multiple subfiles, the embodiments of the present application can avoid excessive processor utilization when the processor imports a large target transaction file into a database. After grouping the subfiles, the subfiles contained in each subfile group are imported into the target database in parallel, thereby increasing processor utilization. By adjusting the number of subfiles contained in each subfile group, the processor utilization can be adjusted so that the processor utilization always remains within a preset utilization range. This improves the efficiency of data file importing and effectively reduces data processing time while ensuring the safety of the processor and system operation.
[0062] Figure 3 The schematic diagram shows a principle diagram of a transaction data processing method according to an embodiment of the present application.
[0063] like Figure 3 As shown, the target transaction file is split into 8 subfiles, and the 8 subfiles are sorted to obtain subfiles 1 to 8 sorted according to the transaction time of the transaction data contained in each subfile. The subfiles are grouped to obtain the first group of subfiles (subfile 1, subfile 2), the second group of subfiles (subfile 3, subfile 4), the third group of subfiles (subfile 5, subfile 6), and the fourth group of subfiles (subfile 7, subfile 8). The calling processor sequentially imports the subfiles contained in each group of subfiles into the target database in parallel, for example, Figure 3 As shown, the calling processor imports subfile 1 and subfile 2 included in the first group of subfiles into the target database in parallel.
[0064] In an embodiment of the present application, the target transaction file may be imported into the target database through a batch data import technology (LOADDATA). The target database may include a Gaussian database table (GAUSS table).
[0065] In some embodiments, the calling processor sequentially imports the subfiles contained in each group of subfiles into the target database in parallel, including: taking each subfile in each group of subfiles as the current subfile, and calling the processor to perform the following operations on the current subfile: dividing the transaction data contained in the current subfile into multiple groups of batch data; using the target interface to sequentially import the multiple groups of batch data into the target database through the data stream; wherein, after each group of batch data is imported, calling the target database to perform a transaction commit operation.
[0066] In an embodiment of the present application, the target interface can be an efficient data import interface provided by the target database. For example, the target interface can be an efficient data import interface (CopyManager) provided by Gaussian Database. The target interface uses a streaming interface of copy commands to directly import data in batches. The data does not require function calculation or conversion at the application layer, thus avoiding the communication overhead of row-by-row insertion.
[0067] In an embodiment of the present application, for each sub-file, the transaction data contained in each group of sub-files is divided into multiple groups of batch data, and each group of batch data is imported in sequence.
[0068] For example, the transaction data of each sub-file is split into multiple batches of data, each batch of data includes 5,000 transaction data, and each batch of data is imported into the target database in sequence.
[0069] In an embodiment of the present application, after each batch of data is imported, the target database is called to perform a transaction commit operation (commit), which can avoid the hidden dangers of large transactions.
[0070] For example, each batch of data includes 5,000 transaction data. Every time 5,000 transaction data are successfully imported, a transaction commit operation is performed to permanently save the data in the database and release related resources (such as locks and log space).
[0071] Through the embodiments of the present application, after each batch of data is imported, the target database is called to execute a transaction commit operation, thereby preventing the impact of large transactions on database performance. If the import of a batch of data fails, only the operation for that batch of data is lost, and previously submitted transactions are not affected. Furthermore, batch operations (a batch of data includes multiple transaction data) can reduce the number of transaction commits, avoiding the performance overhead caused by frequent commits and preventing single transactions from being too large.
[0072] In some embodiments, the method further includes: in response to identifying that import of transaction data in a set of batch data fails, recording the set of batch data in an error register.
[0073] In an embodiment of the present application, when it is identified that a transaction data import failure exists in a certain group of batch data, the group of batch data is recorded in an error register.
[0074] For example, each batch of data includes 5,000 transaction data. If one of the 5,000 data in a batch of data has a format error, all 5,000 data in the batch of data will fail, and the transaction will be rolled back, and the error information will be recorded in the error register.
[0075] Through the embodiments of the present application, in the event that erroneous data appears in transaction data, the solution of this embodiment can promptly record the batch data containing the erroneous data, facilitating subsequent troubleshooting of the erroneous data.
[0076] In some embodiments, before splitting the target transaction file into multiple sub-files of the same file size, the method also includes: obtaining verification information and feature information of the target transaction file respectively; the verification information includes metadata representing the integrity of the target transaction file; and verifying the target transaction file based on the verification information and feature information to obtain a verification result.
[0077] In an embodiment of the present application, verification information includes metadata characterizing the integrity of the target transaction file, such as file size and hash value. Feature information characterizes features of the target transaction file, such as file size and hash value. The target transaction file is verified based on the verification information and feature information. If the features of the target transaction file represented by the verification information match those represented by the feature information, the verification result is considered passed.
[0078] For example, read the CHK file (verification information) of the target transaction file, obtain the size of the BIN file (data file of the target transaction file) recorded in the CHK file, first obtain the size of the BIN file (feature information) on the file server, and then compare it with the file size recorded in the CHK file. If the sizes are consistent, the verification is considered to have passed.
[0079] Through the embodiments of the present application, the target transaction file is verified before being split, and subsequent operations are performed after the verification is passed, which can improve the reliability and stability of the transaction data processing method of this embodiment.
[0080] In some embodiments, the target transaction file is verified based on the verification information and the characteristic information to obtain a verification result, including: in response to a failure in verification of the target transaction file, re-verifying the target transaction file after waiting for a preset time; wherein the length of the preset time is positively correlated with the number of times the target transaction file verification fails.
[0081] In an embodiment of the present application, the length of the preset time is positively correlated with the number of verification failures of the target transaction file. An exponential backoff algorithm is introduced into the verification process of the target transaction file. The greater the number of verification failures of the target transaction file, the longer the preset time. If the verification fails each time, the system waits for the preset time and then tries again until the verification finally passes.
[0082] For example, after the verification of the target transaction file fails continuously, the waiting time is gradually extended (such as 1 minute, 2 minutes, 4 minutes...), which can avoid unnecessary repeated verification.
[0083] Through the embodiments of the present application, repeated verification can be avoided, thereby improving the efficiency of the transaction data processing method of this embodiment.
[0084] In some embodiments, after splitting the target transaction file into multiple sub-files of the same file size, the method further includes: deleting the complete target transaction file before splitting.
[0085] In an embodiment of the present application, since the storage space of the file server is limited and the target transaction file to be split is large, the original file (target transaction file) is deleted after the split is successful, and only multiple sub-files of the same file size obtained after the split are retained, thereby releasing the resources of the file server.
[0086] Through the embodiments of the present application, the resources of the file server can be released in a timely manner, thereby improving the reliability and stability of the transaction data processing method of the embodiment.
[0087] In some embodiments, the method further includes: obtaining a transaction date interval of the target transaction file; the transaction date interval represents the interval range of the transaction date corresponding to the transaction data of the target transaction file; dividing the target database into multiple partitions according to the transaction date interval; wherein each sub-file is imported into the corresponding partition in the target database.
[0088] In the embodiment of the present application, by partitioning the target database, it is possible to avoid the situation where all data is stored in one partition of the target database, resulting in low efficiency when querying the target database by index.
[0089] For example, if the transaction date range of the target transaction file is from March 1st to March 10th, the target database can be divided into 10 partitions based on the transaction date range. The first partition corresponds to the transaction data on March 1st, the second partition corresponds to the transaction data on March 2nd, and so on. When each sub-file is imported into the target database, it is imported into the partition corresponding to each sub-file.
[0090] It should be noted that since multiple sub-files have been sorted according to the transaction time of the transaction data contained in each sub-file, when importing the sub-files into the target database, it is only necessary to import the sub-files into each partition of the target database in sequence according to the sorting of the sub-files to achieve matching between the sub-files and the partitions.
[0091] Through the embodiments of the present application, the efficiency of importing data into the target database and the efficiency of querying data can be improved, thereby improving the reliability and stability of the transaction data processing method of the present embodiment.
[0092] In some embodiments, to improve the performance of the transaction data processing method of this embodiment, the performance of the transaction data processing method can be considered from the perspective of four monitoring indicators (including CPU, memory, JVM stack, and database). These four monitoring indicators specifically include: CPU utilization less than 70%, no memory leaks and free memory greater than 30%, no memory overflow or leaks involving the JVM stack, data server usage within preset thresholds, and performance reports meeting expectations. The requirements of these four monitoring indicators can be met by adjusting the number of subfiles included in each subfile group.
[0093] Based on the above transaction data processing method, this application also provides a transaction data processing device. Figure 4 The device is described in detail.
[0094] Figure 4 The structural block diagram of the transaction data processing device according to an embodiment of the present application is schematically shown.
[0095] like Figure 4 As shown, the transaction data processing device 400 of this embodiment includes a splitting module 410 , a sorting module 420 , a partitioning module 430 and an importing module 440 .
[0096] The splitting module 410 is used to split the target transaction file into multiple sub-files of the same file size. In one embodiment, the splitting module 410 can be used to perform the operation S210 described above, which will not be repeated here.
[0097] The sorting module 420 is configured to sort the sub-files according to the transaction time of the transaction data contained in each sub-file. In one embodiment, the sorting module 420 may be configured to perform the operation S220 described above, which will not be described in detail here.
[0098] The division module 430 is configured to divide the plurality of subfiles into a plurality of subfile groups based on the arrangement order of the plurality of subfiles; each subfile group includes at least two subfiles. In one embodiment, the division module 430 may be configured to perform the operation S230 described above, which will not be described in detail here.
[0099] Import module 440 is configured to call a processor to sequentially and concurrently import the subfiles contained in each subfile group into a target database; wherein, when the processor concurrently imports the subfiles contained in each subfile group, the processor utilization is within a preset utilization range. In one embodiment, import module 440 may be configured to perform operation S240 described above, and will not be further described herein.
[0100] According to embodiments of the present application, any multiple modules among the splitting module 410, sorting module 420, partitioning module 430, and importing module 440 can be combined into a single module, or any one of these modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in a single module. According to embodiments of the present application, at least one of the splitting module 410, sorting module 420, partitioning module 430, and importing module 440 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or can be implemented in any one of the three implementation methods of software, hardware, and firmware, or any appropriate combination of these. Alternatively, at least one of the splitting module 410, sorting module 420, partitioning module 430, and importing module 440 can be at least partially implemented as a computer program module that, when executed, performs the corresponding functionality.
[0101] Figure 5 A block diagram of an electronic device suitable for implementing a transaction data processing method according to an embodiment of the present application is schematically shown.
[0102] like Figure 5 As shown, an electronic device 500 according to an embodiment of the present application includes a processor 501, which can perform various appropriate actions and processes based on a program stored in a read-only memory (ROM) 502 or a program loaded from a storage unit 508 into a random access memory (RAM) 503. The processor 501 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 501 may also include onboard memory for caching purposes. The processor 501 may include a single processing unit or multiple processing units for performing different actions of the method flow according to the embodiment of the present application.
[0103] Various programs and data required for the operation of the electronic device 500 are stored in the RAM 503. The processor 501, ROM 502, and RAM 503 are connected to each other via a bus 504. The processor 501 performs various operations of the method flow according to the embodiment of the present application by executing the programs in the ROM 502 and / or RAM 503. It should be noted that the programs may also be stored in one or more memories other than the ROM 502 and the RAM 503. The processor 501 may also perform various operations of the method flow according to the embodiment of the present application by executing the programs stored in the one or more memories.
[0104] According to an embodiment of the present application, electronic device 500 may further include an input / output (I / O) interface 505, which is also connected to bus 504. Electronic device 500 may also include one or more of the following components connected to I / O interface 505: an input section 506 including a keyboard, mouse, etc.; an output section 507 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 508 including a hard disk; and a communication section 509 including a network interface card such as a LAN card or modem. Communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to I / O interface 505 as needed. Removable media 511, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 510 as needed, so that computer programs read from the removable media can be installed into storage section 508 as needed.
[0105] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of this application is implemented.
[0106] According to an embodiment of the present application, a computer-readable storage medium may be a non-volatile computer-readable storage medium, such as, but not limited to, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present application, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present application, a computer-readable storage medium may include the ROM 502 and / or RAM 503 described above and / or one or more memories other than ROM 502 and RAM 503.
[0107] The embodiments of the present application also include a computer program product, which includes a computer program containing program code for executing the method shown in the flowchart. When the computer program product is executed in a computer system, the program code is used to enable the computer system to implement the transaction data processing method provided in the embodiments of the present application.
[0108] The computer program executes the above functions defined in the system / device of the embodiment of the present application when the computer program is executed by the processor 501. According to the embodiment of the present application, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0109] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 509, and / or installed from a removable medium 511. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0110] In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 509, and / or installed from the removable medium 511. When the computer program is executed by the processor 501, the above-mentioned functions defined in the system of the embodiment of the present application are performed. According to the embodiment of the present application, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.
[0111] According to an embodiment of the present application, the program code for executing the computer program provided by the embodiment of the present application can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).
[0112] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of the boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0113] Those skilled in the art will appreciate that the features described in the various embodiments of this application may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in this application. In particular, the features described in the various embodiments of this application may be combined and / or coupled in various ways without departing from the spirit and teachings of this application. All such combinations and / or couplings fall within the scope of this application.
Claims
1. A transaction data processing method, characterized in that: The method comprises: Split the target transaction file into multiple sub-files of the same file size; sorting the multiple sub-files according to the transaction time of the transaction data contained in each sub-file; Sequentially dividing the plurality of subfiles into a plurality of groups of subfiles based on the arrangement order of the plurality of subfiles; each group of subfiles includes at least two subfiles; The calling processor sequentially imports the sub-files contained in each group of sub-files into the target database in parallel; When the processor imports the sub-files included in each group of sub-files in parallel, the utilization rate of the processor is within a preset utilization rate range.
2. The method according to claim 1, characterized in that The calling processor sequentially and concurrently imports the sub-files contained in each group of sub-files into the target database, including: Each subfile in each group of subfiles is used as a current subfile, and the processor is called to perform the following operations on the current subfile: Dividing the transaction data contained in the current sub-file into multiple batches of data; Importing the plurality of batch data sets into the target database in sequence through a data stream using a target interface; After each batch of data is imported, the target database is called to perform a transaction commit operation.
3. The method according to claim 2, characterized in that Also includes: In response to identifying a failure to import transaction data in a set of batch data, the set of batch data is recorded in an error log.
4. The method according to claim 1, wherein Before splitting the target transaction file into a plurality of sub-files of the same file size, the method further includes: respectively obtaining verification information and characteristic information of the target transaction file; wherein the verification information includes metadata representing the integrity of the target transaction file; The target transaction file is verified according to the verification information and the characteristic information to obtain a verification result.
5. The method according to claim 4, characterized in that The verifying the target transaction file according to the verification information and the characteristic information to obtain a verification result includes: In response to a verification failure of the target transaction file, re-verifying the target transaction file after waiting for a preset time; The length of the preset time is positively correlated with the number of times the target transaction file verification fails.
6. The method according to claim 1, characterized in that After splitting the target transaction file into a plurality of sub-files of the same file size, the method further includes: Delete the complete target transaction file before splitting.
7. The method according to claim 1, characterized in that The method further comprises: Acquire a transaction date interval of the target transaction file; the transaction date interval represents a range of transaction dates corresponding to the transaction data of the target transaction file; Dividing the target database into multiple partitions according to the transaction date interval; Each sub-file is imported into a corresponding partition in the target database.
8. A transaction data processing device, characterized in that: The device comprises: A splitting module is used to split the target transaction file into multiple sub-files of the same file size; A sorting module, used to sort the multiple sub-files according to the transaction time of the transaction data contained in each sub-file; a division module, configured to divide the plurality of subfiles into a plurality of groups of subfiles based on the arrangement order of the plurality of subfiles; each group of subfiles includes at least two subfiles; An import module, configured to call a processor to sequentially and in parallel import the sub-files contained in each group of sub-files into a target database; When the processor imports the sub-files included in each group of sub-files in parallel, the utilization rate of the processor is within a preset utilization rate range.
9. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
11. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.