Data processing method and device, computer device and storage medium
By dividing the data processing into time periods and merging incremental and historical data files, the problem of low efficiency in traditional massive data storage is solved, and efficient data extraction is achieved.
Patent Information
- Application Number
- CN202110230018.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-02
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2041-03-02
AI Technical Summary
Traditional methods of storing massive amounts of data are inefficient when data content is required, and cannot process business data that is generated or updated in real time.
By dividing the data into two different time periods, incremental data files are generated and merged with historical full data files to form a second historical full data file, ensuring that the data file for any time period contains all business data up to the end of that period.
It improves data processing efficiency, reduces cumbersome data query processes, and enables direct extraction of the required data content.
Smart Images

Figure CN114996210B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular to a data processing method and device, computer equipment and storage medium. BACKGROUND
[0002] With the development of computer technology, the data in related scenarios grows rapidly, so in the related technology of data processing and processing, the storage and processing of massive data become an important content of related business processing involving massive data. For example, in some business scenarios that will generate related business data in real time, business data will be generated or updated in real time, so it is necessary to store and process these business data in time. The traditional storage method of massive data has the problem of low efficiency when providing data content based on stored massive data. SUMMARY
[0003] Therefore, it is necessary to provide a data processing method, device, computer equipment and storage medium for the above technical problems.
[0004] A data processing method, the method comprising:
[0005] obtaining incremental change data in a current first time period from real-time updated business data, and generating an incremental data file of the current first time period according to the obtained incremental change data;
[0006] merging the incremental data file of each first time period before the current first time period in a current second time period with a first historical full data file of the last first time period in the last second time period to obtain a second historical full data file of the current first time period, the second time period comprising at least two first time periods; the first historical full data file is a file used to store all business data generated before the end time point of the last first time period in the last second time period; and the second historical full data file is a file used to store all business data generated before the end time point of the current first time period.
[0007] In one embodiment, before obtaining the incremental change data in the current first time period from the real-time updated business data, the method further comprises: obtaining the updated business data in real time, and storing the obtained business data in a first storage medium.
[0008] The method of obtaining the incremental change data in the current first time period from the real-time updated business data comprises: obtaining the incremental change data in the current first time period from the first storage medium.
[0009] In one embodiment, when the current first time period is the first first time period in the current second time period, the incremental data file of the current first time period is synchronized to obtain the second-period full data file.
[0010] In one embodiment, when the current first time period is the first first time period in the current second time period, the second-period full data file is merged with the first historical full data file to obtain a second historical full data file of the current first time period, including:
[0011] When the field content of the current field identifier of the current data content identifier in the second-period full data file is not empty, the field content of the current field identifier of the current data content identifier in the second-period full data file is synchronized to the second historical full data file as the content of the current field identifier of the current data content identifier.
[0012] When the field content of the current field identifier of the current data content identifier in the second-period full data file is empty, the field content of the current field identifier of the current data content identifier in the first historical full data file is synchronized to the second historical full data file as the content of the current field identifier of the current data content identifier.
[0013] In one embodiment, the first time period is 1 hour, and the second time period is 1 day.
[0014] A data processing apparatus, the apparatus comprising:
[0015] An incremental file synchronization module for obtaining incremental change data in a current first time period from real-time updated business data, and generating an incremental data file of the current first time period according to the obtained incremental change data;
[0016] A full amount synchronization module for merging, based on incremental data files of each first time period before the current first time period in a current second time period, a first historical full data file of the last first time period in the last second time period to obtain a second historical full data file of the current first time period, the second time period comprising at least two first time periods; the first historical full data file is a file for storing all business data generated before the end time point of the last first time period in the last second time period; and the second historical full data file is a file for storing all business data generated before the end time point of the current first time period.
[0017] In one embodiment, the full-amount synchronization module comprises:
[0018] The period file synchronization module is configured to merge the incremental data file of the current first time period with a first-period full-amount data file of a previous first time period in a current second time period to obtain a second-period full-amount data file of the current first time period; the first-period full-amount data file is a file configured to store all service data generated in the current second time period and before an end time point of the previous first time period in the current second time period; and the second-period full-amount data file is a file configured to store all service data generated in the current second time period and before an end time point of the current first time period.
[0019] The history file synchronization module is configured to merge the second-period full-amount data file with the first history full-amount data file to obtain a second history full-amount data file of the current first time period.
[0020] In one embodiment, the apparatus further comprises a service data updating module.
[0021] The service data updating module is configured to acquire updated service data in real time and store the acquired service data in the first storage medium.
[0022] The incremental file synchronization module is configured to acquire incremental change data in the current first time period from the first storage medium.
[0023] In one embodiment, the period file synchronization module is configured to, after obtaining the first-period full-amount data file of the previous first time period, merge the incremental data file of the current first time period with the first-period full-amount data file to obtain the second-period full-amount data file of the current first time period.
[0024] In one embodiment, the period file synchronization module is further configured to, when the current first time period is the first first time period in the current second time period, synchronize the incremental data file of the current first time period to obtain the second-period full-amount data file.
[0025] In one embodiment, the periodical file synchronization module is configured to synchronize, when the field content of the current field identification of the current data content identification in the incremental data file of the current first time period is not empty, the field content of the current field identification of the current data content identification in the incremental data file as the content of the current field identification of the current data content identification to the second periodical full data file; and synchronize, when the field content of the current field identification of the current data content identification in the incremental data file of the current first time period is empty, the field content of the current field identification of the current data content identification in the first periodical full data file as the content of the current field identification of the current data content identification to the second periodical full data file.
[0026] In one embodiment, the periodical file synchronization module is configured to merge a target file in the incremental data file of the current first time period and the first periodical full data file of the last first time period in the current second time period to obtain a merged file, the target file being a file with a file size less than or equal to a first predetermined file size; and obtain, based on the incremental data file of the current first time period and the first periodical full data file, a file with a file size greater than the first predetermined file size, and the merged file, a second periodical full data file of the current first time period.
[0027] In one embodiment, the history file synchronization module is configured to synchronize, when the field content of the current field identification of the current data content identification in the second periodical full data file is not empty, the field content of the current field identification of the current data content identification in the second periodical full data file as the content of the current field identification of the current data content identification to the second history full data file; and synchronize, when the field content of the current field identification of the current data content identification in the second periodical full data file is empty, the field content of the current field identification of the current data content identification in the first history full data file as the content of the current field identification of the current data content identification to the second history full data file.
[0028] In one embodiment, the historical file synchronization module is configured to, when the current first time period is the first first time period in a current second time period, merge the period full data file of the current first time period with the first historical full data file to obtain a second historical full data file of the current first time period; and when the current first time period is not the first first time period in the current second time period, merge the period full data file of each first time period in the current second time period with the first historical full data file in parallel to obtain a historical full data file corresponding to each first time period in the current second time period.
[0029] In one embodiment, the historical file synchronization module is configured to, after the merging of the third period full data file of the last first time period in the current second time period with the first historical full data file is completed to obtain a third historical full data file of the last first time period in the current second time period, merge the period full data file of the first first time period in the next second time period with the third historical full data file.
[0030] In one embodiment, the first time period is 1 hour, and the second time period is 1 day.
[0031] A computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the steps of the method in any of the above embodiments when executing the computer program.
[0032] A computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the method in any of the above embodiments.
[0033] The above data processing method and device, computer device and storage medium divide two different time periods, periodically generate an incremental data file of each first time period, and generate and obtain a second historical full data file of the first time period based on the incremental data file of each first time period. For any first time period, the corresponding second historical full data file contains all business data before the end time point of the first time period. Therefore, when data content needs to be provided based on stored massive data, the corresponding second historical full data file of the first time period can be directly extracted, without a cumbersome data query process, thereby improving processing efficiency. BRIEF DESCRIPTION OF DRAWINGS
[0034] Figure 1 An application environment diagram of the data processing method in one embodiment;
[0035] Figure 2 a flowchart of data processing in one embodiment;
[0036] Figure 3 a flowchart of data processing in another embodiment;
[0037] Figure 4 a schematic diagram of the division of the first time period and the second time period in one specific example;
[0038] Figure 5 a schematic diagram of the principle of the data processing method in one specific example;
[0039] Figure 6 a schematic diagram of the principle of the generation of each data file in the data processing method in one specific example;
[0040] Figure 7 a structural block diagram of the community-based data processing device in one embodiment;
[0041] Figure 8 a structural block diagram of the community-based data processing device in another embodiment;
[0042] Figure 9 an internal structural diagram of the computer device in one embodiment. DETAILED DESCRIPTION
[0043] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not used to limit the present application.
[0044] The data processing method provided by the present application can be applied to, for example, Figure 1The application environment shown. Among them, the user terminal 10 and the server 20 can communicate with each other through the network, the server 20 and the server 30 can communicate with each other through the network, in the related business scenario, the updated business data generated will be stored in the server 20, the updated business data can be generated by the user terminal 10 in the process of carrying out the related business, the server 30 obtains the incremental change data in the current first time period from the real-time updated business data in the server 20 every first time period, generates the incremental data file of the current first time period according to the obtained incremental change data; and based on the incremental data file of each first time period before the current first time period in the current second time period, the first historical full data file of the last first time period in the last second time period is merged to obtain the second historical full data file of the current first time period, the second time period includes at least two first time periods. The server 30 generates each historical full data file (including the above-mentioned first historical full data file and the second historical full data file), which can be provided to the device 40. The device 40 can be a terminal or a server. Among them, the terminal can be but not limited to various personal computers, notebook computers, smart phones, tablet computers and portable wearable devices, and the server can be realized by an independent server or a server cluster composed of multiple servers.
[0045] In some embodiments, the server 20 and the server 30 can be servers in a distributed system, and the business data described above, such as the updated business data, can be stored in a blockchain, and the incremental data file, the periodic full data file and the historical full data file generated in the above embodiments and the following embodiments can also be stored in a blockchain.
[0046] Blockchain is a new application mode of distributed data storage, peer-to-peer transmission, consensus mechanism, encryption algorithm and other computer technologies. Blockchain, in essence, is a decentralized database, which is a series of data blocks associated using cryptographic methods, each data block contains a batch of network transaction information, used to verify the validity of the information (anti-fake) and generate the next block. Blockchain can include blockchain underlying platform, platform product service layer and application service layer.
[0047] The blockchain underlying platform can include user management, basic services, smart contracts, and operation monitoring processing modules. Among them, the user management module is responsible for the identity information management of all blockchain participants, including maintaining public and private key generation (account management), key management, and user real identity and blockchain address correspondence maintenance (permission management), etc., and under authorization, supervises and audits the transaction of certain real identities, provides risk control rule configuration (risk audit); the basic service module is deployed on all blockchain node devices to verify the validity of business requests, and record to the storage after consensus for valid requests, for a new business request, the basic service first interface adaptation analysis and authentication processing (interface adaptation), then encrypt the business information through the consensus algorithm (consensus management), after encryption, the complete and consistent transmission to the shared ledger (network communication), and record storage; the smart contract module is responsible for contract registration and issuance, contract triggering and contract execution, developers can define contract logic through a certain programming language, publish to the blockchain (contract registration), according to the logic of the contract terms, call the key or other event triggers to execute, complete the contract logic, and also provide contract upgrade and cancellation functions; the operation monitoring module is mainly responsible for the deployment, configuration modification, contract setting, cloud adaptation in the product release process, and the real-time state visualization output in the product running, such as: alarm, monitoring network situation, monitoring node device health status, etc.
[0048] The platform product service layer provides basic capabilities and implementation framework of typical applications, and developers can add business characteristics based on these basic capabilities to complete the blockchain implementation of business logic. The application service layer provides application services based on the blockchain scheme for business participants to use.
[0049] In one embodiment, as shown in Figure 2 , a data processing method is provided, which is applied to the server 30 in Figure 1 for example, and includes the following steps S21 to S22.
[0050] Step S21: obtaining incremental change data in a current first time period from real-time updated business data, and generating an incremental data file of the current first time period according to the obtained incremental change data.
[0051] The real-time updated business data is business data generated in real time in an actual business scenario. The business data can be any data content in any situation, such as a piece of access content data, a piece of content data in which a field of a certain data content is modified, etc., and the content form of the business data is not limited.
[0052] In some embodiments, before obtaining the incremental change data in the current first time period from the real-time updated service data, the updated service data can be obtained in real time first, and the obtained service data can be stored in the first storage medium. In combination Figure 1 As shown in the figure, the first storage medium can refer to the server 20. At this time, the incremental change data in the current first time period is obtained from the real-time updated service data, which can be specifically obtained from the first storage medium.
[0053] The specific duration of the first time period can be set according to actual technical needs, for example, 10 minutes, 1 hour, several days, etc. The actual duration of the first time period can be set according to the needs of the actual business scenario, for example, in some embodiments, the first time period can be set to 1 hour.
[0054] The incremental change data in the current first time period obtained from the real-time updated service data refers to the service data whose generation time is in the current first time period. Taking the current first time period of 1 hour as an example, assuming that the current first time period is from 8:00 to 9:00 on February 2, 2021, the incremental change data in the current first time period obtained is the service data whose generation time is in the time interval from 8:00 to 9:00 on February 2, 2021. Based on the obtained incremental change data, the incremental data file of the current first time period can be generated, that is, the incremental data file of the current first time period, which contains the service data whose generation time is in the time interval of the current first time period.
[0055] Step S22: based on the incremental data files of each first time period before the current first time period in the current second time period, and the first historical full data file of the last first time period in the last second time period, a second historical full data file of the current first time period is obtained, the second time period includes at least two first time periods. The first historical full data file is a file used to store all service data whose generation time is before the end time point of the last first time period in the last second time period, and the second historical full data file is a file used to store all service data whose generation time is before the end time point of the current first time period
[0056] The second time period includes at least two first time periods. The specific duration of the second time period can be set in combination with actual technical needs. For example, if the first time period is 10 minutes, the second time period can be set to 1 hour, 2 hours, etc. For example, if the first time period is 1 hour, the second time period can be set to 1 day. For example, if the first time period is 1 day, the second time period can be set to 10 days or 1 month, etc. It can be understood that the setting method herein is only an example. In other embodiments, the first time period and the second time period can also be set in other ways, as long as the second time period contains more than two first time periods.
[0057] The historical full data file is a file used to store all business data generated before the end time point of the corresponding first time period. It contains all business data generated before the end time point of the corresponding first time period. For example, if the first time period is 1 hour and the second time period is 1 day, assuming that the current first time period is 2021020203 (i.e. between 3:00 and 4:00 on February 2, 2021), the historical full data file corresponding to the first time period 2021020203 stores all business data generated before 4:00 on February 2, 2021. Assuming that the current first time period is 2021020204 (i.e. between 4:00 and 5:00 on February 2, 2021), the historical full data file corresponding to the first time period 2021020204 stores all business data generated before 5:00 on February 2, 2021.
[0058] In some embodiments, when the second historical full data file of the current first time period is obtained by merging the incremental data files of each first time period before the current first time period in the current second time period with the first historical full data file of the last first time period in the last second time period, it can be that the incremental data files of each first time period before the current first time period are directly merged with the first historical full data file to obtain the second historical full data file of the current first time period. Taking the first time period as 1 hour and the second time period as 1 day as an example, assuming that the current first time period is 2021020203 (i.e. between 3 o'clock and 4 o'clock on February 2, 2021), the incremental data file of the first time period 2021020200 (i.e. between 0 o'clock and 1 o'clock on February 2, 2021), the incremental data file of the first time period 2021020201 (i.e. between 1 o'clock and 2 o'clock on February 2, 2021), the incremental data file of the first time period 2021020202 (i.e. between 2 o'clock and 3 o'clock on February 2, 2021), and the incremental data file of the current first time period 2021020203 can be merged with the first historical full data file of the last first time period 2021020123 (i.e. between 23 o'clock on February 1, 2021 and 0 o'clock on February 2, 2021) in the last second time period (i.e. on February 1, 2021) to obtain the second historical full data file of the current first time period 2021020203.
[0059] In some embodiments, with reference to Figure 3As shown, the second historical full data file of the current first time period can also be obtained by combining the period full data file for the second time period with the first historical full data file after obtaining the period full data file for the second time period in combination with the incremental data file of the current first time period. At this time, the step S22 includes the following steps S221 and S222. The period full data file is a file used to store all service data generated in the corresponding second time period and before the period end time point of the first time period of the corresponding second time period, and contains all service data generated in the corresponding second time period and before the period end time point of the first time period of the corresponding second time period. For example, taking the first time period as 1 hour and the second time period as 1 day as an example, assuming that the current first time period is 2021020203 (i.e., between 3 o'clock and 4 o'clock on February 2, 2021), the period full data file corresponding to the first time period 2021020203 stores all service data generated between 0 o'clock and 4 o'clock on February 2, 2021, and assuming that the current first time period is 2021020204 (i.e., between 4 o'clock and 5 o'clock on February 2, 2021), the period full data file corresponding to the first time period 2021020204 stores all service data generated between 0 o'clock and 5 o'clock on February 2, 2021.
[0060] Step S221: merging the incremental data file of the current first time period with the first period full data file of the previous first time period in the current second time period to obtain the second period full data file of the current first time period.
[0061] In some embodiments, when the current first time period is the first first time period in the current second time period, the second period full data file can be obtained by synchronizing the incremental data file of the current first time period directly. Taking the first time period as 1 hour and the second time period as 1 day as an example, assuming that the current second time period is 20210202 (i.e., between 0 o'clock on February 2, 2021 and 0 o'clock on February 3, 2021), and the current first time period is 2021020200 (i.e., between 0 o'clock and 1 o'clock on February 2, 2021), since the current first time period is the first time period in the current second time period, the previous first time period in the current second time period is empty, and therefore, the second period full data file of the current first time period obtained by merging only contains the incremental data file of the current first time period.
[0062] In some embodiments, after obtaining the first-period full data file of the previous first time period in the current second time period, the incremental data file of the current first time period is merged with the first-period full data file to obtain the second-period full data file of the current first time period.
[0063] In some embodiments, after obtaining the first-period full data file of the previous first time period in the current second time period, the incremental data file of the current first time period is merged with the first-period full data file to obtain the second-period full data file of the current first time period.
[0064] That is, before obtaining the second-period full data file of the current first time period, it is necessary to obtain the first-period full data file of the previous first time period in the current second time period, that is, the process of obtaining the period full data file of each first time period in the current second time period is processed in sequence according to the order of each first time period.
[0065] In some embodiments, the incremental data file of the current first time period is merged with the first-period full data file of the previous first time period in the current second time period to obtain the second-period full data file of the current first time period, including:
[0066] When the field content of the current field identifier of the current data content identifier in the incremental data file is not empty, the field content of the current field identifier of the current data content identifier in the incremental data file is synchronized to the second-period full data file as the content of the current field identifier of the current data content identifier.
[0067] When the field content of the current field identifier of the current data content identifier in the incremental data file is empty, the field content of the current field identifier of the current data content identifier in the first-period full data file is synchronized to the second-period full data file as the content of the current field identifier of the current data content identifier.
[0068] Wherein, the current data content identifier refers to the data content being synchronized, and the current field identifier refers to a specific field of the current data content identifier being synchronized. Thus, when the relevant field content in the latest incremental data file is updated during synchronization of the data content, the relevant field content in the incremental data file is retained for storage, otherwise the relevant field content in the first-period full data file is stored, ensuring the continuity of the historical field content and the update of the latest field content.
[0069] In some embodiments, the incremental data file of the current first time period is merged with the first-period full data file of the last first time period in the current second time period to obtain a second-period full data file of the current first time period, including:
[0070] The target file in the incremental data file of the current first time period and the first-period full data file of the last first time period in the current second time period is merged to obtain a merged file, the target file being a file with a file size less than or equal to a first predetermined file size;
[0071] Based on the incremental data file of the current first time period and the file with a file size greater than the first predetermined file size in the first-period full data file, and the merged file, a second-period full data file of the current first time period is obtained.
[0072] In some embodiments, the target file is a file with a file size less than or equal to a first predetermined file size, which is also referred to as a small file in the embodiments of the present application. The first predetermined file size can be set according to actual technical needs. For example, when the default block size is 128 MB (megabytes), the first predetermined file size can be set to a value much smaller than the default block size, such as 20 MB, 15 MB, etc. For example, a file with a size of 3 MB, 7 MB or 10 MB is considered a small file. In some embodiments, the merged file obtained by merging the target file has a file size less than or equal to a second predetermined file size, and the second predetermined file size is greater than the first predetermined file size. The second predetermined file size can be set according to actual technical needs, for example, in some embodiments, it can be set to the default block size or an integer multiple of the default block size. The second predetermined file size can be set to at least twice the first predetermined file size, so that at least two small files can be merged. Thus, in the process of obtaining the second-period full data file of the current first time period, the number of small files in the final obtained second-period full data file is greatly reduced, which helps to improve the efficiency of subsequently merging the second-period full data file to obtain a second historical full data file.
[0073] In some embodiments, there can be small files with a file size less than or equal to the first predetermined file size in the incremental data file of the current first time period and the first-period full data file. In this case, the second-period full data file can be obtained by merging each small file in the incremental data file of the current first time period and the first-period full data file, obtaining a merged file, and then obtaining the second-period full data file based on the merged file and other files not merged in the incremental data file of the current first time period and the first-period full data file.
[0074] In some embodiments, since the merging of small files has been performed in the process of obtaining the first-period full data file, there can be only small files in the current first-time-period incremental data file, at this time, only the small files in the current first-time-period incremental data file can be merged to obtain a merged file, and then based on the merged file, other files in the current first-time-period incremental data file that have not been merged, and the first-period full data file, the second-period full data file can be obtained.
[0075] In some embodiments, the data file is stored in blocks, and the default block size of the storage setting can be set. The block refers to the division of data for storage. For example, in a distributed storage system, data is divided into multiple parts and stored in each block to solve the problem of large single storage files. Usually, a default block size is set, i.e., the size of each block is the same. When the data of a certain file is large and needs to be stored, the data of the file is divided based on the default block size and stored in each block, so that when the relevant data of the file is needed, the corresponding data can be obtained from each block. However, in actual technical scenarios, there will inevitably be some files whose size is smaller than the default block size. If there are multiple files whose size is smaller than the default block size, the data of these multiple files can be stored in the same block, resulting in the need to access the same block multiple times when reading the content of the file. Therefore, by merging the multiple small files whose size is smaller than the default block size into one file, the data of the multiple small files is stored in the same block as the data of one file, and when the data is actually read, the block is accessed once, and the data content that would have been obtained by accessing the block multiple times can be obtained, which helps to improve the subsequent access efficiency.
[0076] In one embodiment, when merging the small files, in the process of merging the incremental data file of the current first time period and the first period full data file of the last first time period in the current second time period, the data of the incremental data file of the current first time period and the first period full data file is sorted, and then the data is written into the merged file one by one, and after the file is full, the next file is generated to continue writing the remaining data. In this way, the number of finally generated files can be greatly reduced, which helps to improve the efficiency when merging the second historical full data file based on the second period full data file. For example, taking a distributed system as an example, the size of the merged file can be set in combination with the default block size, for example, the upper limit of the size of the merged file is set to the default block size. Thus, in the process of merging the small files, after the data in front of the sorting is written into the merged file, the merged file is full, or although the merged file is not full, if the next data is written, the data of the merged file will overflow, that is, the size of the data after adding the next data will be greater than the default block size, then the writing of the current merged file is completed, and the next folder is generated to continue writing the remaining data. Thus, in the process of merging the small files, it is ensured that the merged file only corresponds to one block, which helps to improve the processing efficiency when extracting data.
[0077] In one embodiment, the small files in the incremental data file of the current first time period can be searched first to determine the number of existing small files, and when the number of small files is greater than a predetermined number threshold, the incremental data file of the current first time period and the first period full data file of the last first time period in the current second time period are merged, and the merging of the small files is performed in the process of merging. In other embodiments, other ways of merging the small files can also be used, which are not limited in the embodiments of the present application.
[0078] Step S222: merging the second period full data file and the first historical full data file to obtain a second historical full data file of the current first time period.
[0079] In some embodiments, when the current first time period is the first first time period in the current second time period, the period full data file of the current first time period and the first historical full data file can be directly merged to obtain a second historical full data file of the current first time period.
[0080] In some embodiments, when the current first time period is not the first first time period in the current second time period, the period full data files of each first time period in the current second time period can be merged with the first historical full data file in parallel to obtain the historical full data file corresponding to each first time period in the current second time period. Wherein, merging the period full data files of each first time period in the current second time period with the first historical full data file in parallel means that during the process of merging the period full data file of one first time period with the first historical full data file, the period full data file of another first time period can be merged with the first historical full data file at the same time.
[0081] Therefore, during the process of obtaining the historical full data file, the period full data files of each first time period in the current second time period are merged with the first historical full data file of the last first time period in the last second time period, which is irrelevant to the files of other first time periods in the current second time period. Therefore, the corresponding historical full data files of each first time period can be obtained in parallel to improve the processing efficiency.
[0082] In some embodiments, the merging of the period full data files of each first time period with the first historical full data file can also be performed in parallel when a delay is monitored during the generation of the historical full data file or other sudden situations occur, so as to realize speed catching up after data delay. The monitoring of the delay can be performed in any possible way, for example, when the time of generating the historical full data file of a certain first time period is later than the predetermined time interval of the first time period, it is determined that the delay condition is reached and speed catching up is needed, so that the process of merging the period full data files of each first time period with the first historical full data file is started in parallel. In other embodiments, other ways can also be used to monitor whether the delay occurs.
[0083] After the third period full data file of the last first time period in the current second time period is merged with the first historical full data file to obtain the third historical full data file of the last first time period in the current second time period, the process of merging the period full data file of the first first time period in the next second time period with the third historical full data file is entered. Wherein, the third period full data file is used to store all business data generated in the current second time period and before the end time point of the last first time period in the current second time period, that is, a file used to store all business data generated in the current second time period.
[0084] Thus, by ensuring that the third historical full data file of the last first time period in the current second time period is obtained before merging the period full data file of the first first time period in the next second time period with the third historical full data file, the accuracy of the subsequently obtained historical full data file is ensured.
[0085] In some embodiments, when the current first time period is the first first time period in the current second time period, the second period full data file is merged with the first historical full data file to obtain a second historical full data file of the current first time period, including:
[0086] When the field content of the current field identification of the current data content identification in the second period full data file is not empty, the field content of the current field identification of the current data content identification in the second period full data file is synchronized to the second historical full data file as the content of the current field identification of the current data content identification.
[0087] When the field content of the current field identification of the current data content identification in the second period full data file is empty, the field content of the current field identification of the current data content identification in the first historical full data file is synchronized to the second historical full data file as the content of the current field identification of the current data content identification.
[0088] Wherein, the current data content identification refers to the data content being synchronized, and the current field identification refers to a specific field of the current data content identification being synchronized. Thus, when the relevant field content in the latest period full data file is updated during synchronization of the data content, the relevant field content in the period full data file is stored, otherwise the relevant field content in the first historical full data file is stored, ensuring the continuity of the historical field content and the update of the latest field content.
[0089] Based on the above-described embodiments, the following will be described in detail in combination with a specific application example. In the specific example, the first time period is 1 hour and the second time period is 1 day. Figure 4As shown, assuming that the current second time period is 20210202 (i.e. between 0:00 on February 2, 2021 and 0:00 on February 3, 2021), the current first time period as described above can be any one of the first time periods in the second time period 20210202, such as the first time period 2021020200 (i.e. between 0:00 and 1:00 on February 2, 2021), the first time period 2021020201 (i.e. between 1:00 and 2:00 on February 2, 2021), the first time period 2021020202 (i.e. between 2:00 and 3:00 on February 2, 2021), and so on. Among them, in combination with Figure 4 As shown, when the current first time period is 2021020200, the previous first time period in the current second time period is empty, when the current first time period is 2021020201, the previous first time period in the current second time period is the first time period 2021020200, and so on. The last first time period in the previous second time period of each first time period 2021020200, 2021020201, … in the current second time period refers to the first time period 2021020123 (i.e. between 23:00 on February 1, 2021 and 0:00 on February 2, 2021).
[0090] Referring to Figure 5 Figure 6 As shown, in an actual business scenario, two types of storage media can be involved to store related data content and data files, respectively, which can be referred to as a first storage medium and a second storage medium in the embodiments of the present application. Among them, the first storage medium is used to store real-time acquired and updated business data, and the second storage medium is used to store the incremental data files, the periodic full data files, and the historical full data files, etc. obtained above. The second storage medium can also include a plurality of different sub-media to store incremental data files, periodic full data files, and historical full data files, respectively. The first storage medium, the second storage medium, and the sub-media in the second storage medium can all be embodied in the form of a server or a server cluster.
[0091] The types of the first storage medium and the second storage medium can be the same or different, and some optional types can be a file system, a relational database, a columnar database, an HBase table, a Hive table, and the like. In the following description of related embodiments, the obtained business data is stored in the form of an HBase table in the first storage medium, and the incremental data file, the periodic full data file, and the historical full data file are taken as examples for description of the second storage medium in the form of a Hive table. In a specific application example, the Hive table can be a TDW data table. At this time, the incremental data table is an HBase table, the incremental data file is an incremental data table, the periodic full data file is a periodic full data table, and the historical full data file is a historical full data table.
[0092] Hbase is a short name of Hadoop database, is based on Hadoop database, is a NoSQL database, mainly suitable for random real-time query of massive detailed data, Hive is a data warehouse tool based on Hadoop, used for data extraction, transformation, and loading, which is a mechanism that can store, query, and analyze massive data stored in Hadoop. The hive data warehouse tool can map a structured data file into a database table. TDW (Tencent distributed Data Warehouse) is an offline data processing platform constructed based on open source software Hadoop and Hive, and a large number of optimizations and modifications are made for specific cases such as large data volume and complex calculation.
[0093] Based on such a setting, referring to FIG. 1, Figure 5 In the first storage medium mentioned in the above embodiments, the obtained business data can be stored in the form of an Hbase table, and the incremental data table, the periodic full data table, and the historical full data table generated in the above embodiments can be respectively referred to as a Hive hourly incremental table, a Hive daily full table, and a Hive historical full table in FIG. 2. Figure 5
[0094] Referring to FIG. 3, Figure 5 Figure 6 As shown, in the business execution process of the actual business scenario, the generated business data is continuously written into the HBase table (i.e., the above incremental data file or incremental data table), and the HBase table can store massive data; the Hive table can store massive data. In the implementation process of the embodiments of the present application, four synchronization tasks can be involved: synchronization task 1, synchronization task 2, synchronization task 3, and synchronization task 4, wherein the synchronization task 1, the synchronization task 2, and the synchronization task 3 are executed in units of a first time period, for example, every 1 hour, and the synchronization task 4 is executed in units of a second time period, for example, every day or every 24 hours. The task scheduling system can be set up to support the automatic execution of related tasks every hour and every day through the task scheduling system, and the dependency relationship between tasks is supported.
[0095] The synchronization task 1 refers to a synchronization task for synchronizing the business data in the current first time period from the first storage medium to the incremental data file (i.e., the Hive hourly incremental table). The function thereof is to synchronize the newly added or written data in the current first time period (for example, from the last hour to the current hour) from the HBase table to the Hive table to generate the incremental data table of the current first time period, which can also be referred to as the current hourly incremental data Hive table in the related embodiments of the present application. By synchronizing only the data in the current first time period, for example, the data from the last hour to the current hour in 1 hour, the amount of data to be synchronized and modified can be reduced, and the time spent in merging data by the downstream synchronization task can be reduced. After the business data is continuously written and then synchronized to the Hive table, the data in the Hive table can be accessed through SQL in the subsequent processing process for data processing, data warehouse construction, data analysis, and the like, so as to improve the data processing efficiency and processing performance.
[0096] In the process of synchronizing the newly added or written data from the HBase table to the Hive table to generate the incremental data table of the current first time period, in some embodiments, the Spark (a distributed computing framework supporting in-memory iterative computing) task can be used to write into the incremental data table of the current first time period in parallel.
[0097] The synchronization task 2 is a synchronization task for merging the incremental data in the incremental data file of the current first time period (i.e., the Hive hourly incremental table) with the full amount data in the full amount data file of the last first time period in the current second time period (i.e., the Hive daily full amount table, also referred to as the first period full amount data file in the above embodiments of the present application), to generate the full amount data file of the current first time period (i.e., the Hive daily full amount table, also referred to as the second period full amount data file in the above embodiments of the present application). The function of the synchronization task is to merge the incremental data of the current first time period (e.g., the current hour) with the daily full amount data of the last hour of the day, to generate the second period full amount data file of the current first time period, which can also be referred to as the current hour daily full amount data Hive table in the related embodiments of the present application.
[0098] The synchronization task is a self-dependent task, and the execution of the current first time period must wait for the execution of the task of the last first time period in the current second time period to be completed before execution. In combination with the above Figure 6 As shown in the above
[0099] If the current hour is 0 o'clock, i.e., the current first time period is 2021020200, the data of the current hour incremental table can be directly synchronized to the daily full amount data Hive table of the current hour, i.e., the data of the daily full amount data Hive table of the current hour is consistent with the data of the current hour incremental table, i.e., the data in the period full amount data table 2021020200 is consistent with the data in the incremental data table 2021020200.
[0100] If the current hour is not 0 o'clock, i.e., the current first time period is 2021020201, 2021020202, …, or 2021020223, the data of the current hour incremental table is merged with the data in the daily full amount data Hive table of the last first time period in the current second time period, to generate the daily full amount data Hive table of the current first time period, i.e., the data in the daily full amount data Hive table of the current first time period is obtained by merging the data in the current hour incremental table with the data in the daily full amount data Hive table of the last first time period in the current second time period.
[0101] For example, if the current hour is 1 o'clock, i.e., the current first time period is 2021020201, the data of the daily full amount data Hive table of the current hour is the result of merging the data of the daily full amount data Hive table of 0 o'clock with the data of the current hour data Hive table of 1 o'clock. That is, the data in the period full amount data table 2021020201 is the result of merging the data in the period full amount data table 2021020200 with the data in the incremental data table 2021020201;
[0102] If the current hour is 9 o'clock, i.e., the current first time period is 2021020209, the data of the current hour's daily full volume data Hive table is the result of merging the data of the daily full volume data Hive table at 8 o'clock and the data of the current hour data Hive table at 9 o'clock. That is, the data in the period full volume data table 2021020209 is the result of merging the data in the period full volume data table 2021020208 and the data in the incremental data table 2021020209.
[0103] In the process of merging the incremental data file of the current first time period with the full volume data of the period full volume data file of the previous first time period of the current second time period to generate the period full volume data file of the current first time period, if the same data content identifier and different field content of the same field identifier appear in the period full volume data file of the previous first time period of the current second time period (for example, the period full volume data table 2021020200) and the incremental data file of the current first time period (for example, the incremental data table 2021020201), for example, the row ID is the same and the value of the same field is different, the following method is used for processing:
[0104] If the field content of the field identifier of the data content identifier in the incremental data table 2021020201 is not empty, the field content of the field identifier of the data content identifier in the incremental data table 2021020201 is synchronized to the period full volume data table 2021020201 as the field content of the field identifier of the data content identifier in the period full volume data table 2021020201.
[0105] If the field content of the field identifier of the data content identifier in the incremental data table 2021020201 is empty, the field content of the field identifier of the data content identifier in the period full volume data table 2021020200 is synchronized to the period full volume data table 2021020201 as the field content of the field identifier of the data content identifier in the period full volume data table 2021020201.
[0106] It should be understood that when the field content of the field identifier of the data content identifier in the period full volume data table 2021020200 is synchronized to the period full volume data table 2021020201, the field content of the field identifier of the data content identifier in the period full volume data table 2021020200 with the latest generation time can be synchronized to the period full volume data table 2021020201 as the latest field data.
[0107] For example, assuming that the data content A has 2 fields, the field identifiers are A1 and A2, and is denoted as A(A1, A2), the data content A with the latest generation time stored in the periodic full data table 2021020200 is A(m1, m2), and the data content A stored in the incremental data table 2021020201 is A(m3, ). Since the field content of the field identifier A1 is not empty and the field content of the field identifier A2 is empty in the incremental data table 2021020201, after synchronization to the periodic full data table 2021020201, the data content stored in the table is denoted as A(m3, A2).
[0108] The synthesis process of the full data table of other first time periods in the current second time period, i.e., other hours, can be processed in a similar manner.
[0109] In the above process of generating the full data file of the current first time period, the small files can be processed at the same time. In the case of synchronizing the newly added or written data from the HBase table to the Hive table by using the above Spark task to generate the incremental data table of the current first time period, based on the concurrency mechanism of Spark, many small files will be generated, and thus the merging of the small files at the same time helps to improve the merging efficiency when generating the historical full data file. It should be understood that in the case of generating the incremental data table of the current first time period by using other mechanisms, if there are many small files, the small files can also be processed.
[0110] The synchronization task 3 is a synchronization task for merging the full data in the periodic full data file (i.e., the Hive daily full table) of the current first time period with the historical full data in the historical full data file (i.e., the Hive historical full table, also referred to as the first historical full data file in the above embodiments of the present application) of the last first time period in the previous second time period to generate the historical full data file (i.e., the Hive historical full table, also referred to as the second historical full data file in the above embodiments of the present application) of the current first time period. The function of the synchronization task 3 is to merge the current full data of the current first time period (e.g., the current hour) with the historical full data of the last first time period of the previous day to generate the second historical full data file of the current first time period, which can also be referred to as the second historical full data Hive table in the related embodiments of the present application.
[0111] The synchronization task 3 does not need to depend on the completion of the synchronization task 3 of the previous first time period, and the execution of the synchronization task 3 of other first time periods in the current second time period does not need to depend on the completion of the synchronization task 3 of the previous first time period. In combination with the above description of the synchronization task 2, the synchronization task 3 can be executed in parallel with the synchronization task 2. Figure 6As shown:
[0112] If the current hour is 0 o'clock, that is, the current first time period is 2021020200, the data of the historical full data Hive table of the current hour is the result of merging the data of the daily full data Hive table of 0 o'clock and the data of the historical full data Hive table of the previous day 23 o'clock. That is, the data in the historical full data table 2021020200 is the result of merging the full data table 2021020200 and the data in the historical full data table 2021020123.
[0113] If the current hour is 1 o'clock, that is, the current first time period is 2021020201, the data of the historical full data Hive table of the current hour is the result of merging the data of the daily full data Hive table of 1 o'clock and the data of the historical full data Hive table of the previous day 23 o'clock. That is, the data in the historical full data table 2021020201 is the result of merging the full data table 2021020201 and the data in the historical full data table 2021020123.
[0114] If the current hour is 9 o'clock, that is, the current first time period is 2021020209, the data of the historical full data Hive table of the current hour is the result of merging the data of the daily full data Hive table of 9 o'clock and the data of the historical full data Hive table of the previous day 23 o'clock. That is, the data in the historical full data table 2021020209 is the result of merging the full data table 2021020209 and the data in the historical full data table 2021020123.
[0115] In the process of merging the data of the daily full data Hive table of 0 o'clock and the data of the historical full data Hive table of the previous day 23 o'clock, if the field contents of the same field identification with the same data content identification appear different in the two tables, for example, the row ID is the same, but the values of the same field are different, the following method is used for processing:
[0116] If the field content of the field identification of the data content identification in the full data table 2021020200 of 0 o'clock is not empty, the field content of the field identification of the data content identification in the full data table 2021020200 is synchronized to the historical full data table 2021020200 as the field content of the field identification of the data content identification in the full data table 2021020200;
[0117] If the field content identified by the data content identifier and the field identifier in the full data table 2021020200 at 0 point is empty, then the field content identified by the data content identifier and the field identifier in the historical full data table 2021020123 is synchronized to the historical full data table 2021020200 as the field content identified by the data content identifier and the field identifier in the historical full data table 2021020200.
[0118] It should be understood that when the field content identified by the data content identifier and the field identifier in the full data table 2021020200 is synchronized to the historical full data table 2021020200, the field content identified by the data content identifier and the field identifier in the full data table 2021020200 with the latest generation time is synchronized to the historical full data table 2021020200 as the latest field data. When the field content identified by the data content identifier and the field identifier in the historical full data table 2021020123 is synchronized to the historical full data table 2021020200, the field content identified by the data content identifier and the field identifier in the historical full data table 2021020123 with the latest generation time is synchronized to the historical full data table 2021020200 as the latest field data.
[0119] For example, it is assumed that the data content B has two fields with field identifiers B1 and B2, which can be denoted as B(B1, B2). The data content B with the latest generation time stored in the full data table 2021020200 is B(0, m4), and the data content B with the latest generation time stored in the historical full data table 2021020123 is B(m5, m6). Since the field content identified by the field identifier B1 is empty and the field content identified by the field identifier B2 is not empty in the full data table 2021020200, the data content stored in the historical full data table 2021020200 after synchronization will be denoted as B(m5, m4).
[0120] The synchronization task 4 is an empty running synchronization task for determining the successful execution of the synchronization task 3 of the last first time period of the second time period. It is an empty running data synchronization task performed at the beginning or end of every second time period, mainly serving the function of task dependency control. Its function is to determine the successful execution of the synchronization task 3 of the last first time period of the second time period, i.e., the corresponding historical full data file of the last first time period of the second time period is generated, and to determine that the synchronization task 3 of the first first time period of the next second time period can be started, i.e., the corresponding historical full data file of the first first time period of the next second time period can be generated.
[0121] That is, the execution of the synchronization task 4 is mutually dependent on the synchronization task 3, the execution of the synchronization task 4 depends on the successful execution of the synchronization task 3 of the last first time period of the second time period, and the starting execution of the synchronization task 3 of the first first time period of the second time period depends on the successful execution of the synchronization task 4 of the second time period (it can be understood that if the synchronization task is divided into the initial execution of the next second time period, it can be considered as the next second time period). That is, as shown in Figure 4 , Figure 6 The task of the synchronization task 4 of 20210202 at 0 o'clock of the day can be executed only after the successful execution of the synchronization task 3 of 20210201 at 23 o'clock of the instance, that is, the successful generation of the historical full data file 2021020123; the task instance of the synchronization task 3 of 20210202 from 0 o'clock to 23 o'clock can be executed only after the successful execution of the task of the synchronization task 4 of 20210202 at 0 o'clock of the day, that is, the historical full data file 2021020200, 2021020201……can be generated. Therefore, through the setting of the mutually dependent synchronization task 3 and the synchronization task 4, the problem of data loss caused by the generation of the historical full data file of the first first time period of the current second time period before the generation of the historical full data file of the last first time period of the previous second time period is effectively avoided, and the integrity of the data is effectively protected.
[0122] Based on the manner in the embodiments of the application as described above, any piece of historical data can be updated in each first time period, and the updated historical data is updated in the full data table of the latest first time period and the historical full data, so that the massive data can be synchronized in each first time period, and the full data is the latest. Therefore, the generation and update delay of the historical full data generated in each first time period will not be large, when data access is needed to obtain related data content, the historical full data of the nearest first time period can be accessed in each first time period, and the historical full data of the first time period is the latest. Generally, the latest historical data can be accessed within the time range of 2 first time periods,
[0123] It should be understood that although the steps in each flowchart discussed above are shown in a sequential order following the arrows, the steps are not necessarily executed in the order shown by the arrows. Unless otherwise specified herein, the steps are not necessarily executed in a strict order, and the steps can be executed in other orders. Moreover, at least some of the steps in these flowcharts can include multiple steps or multiple stages, which are not necessarily executed at the same time, and which are not necessarily executed sequentially, but can be executed at different times, and which can be executed in rotation or alternation with other steps or stages in other steps.
[0124] In one embodiment, as shown in FIG. 1, Figure 7 , 8 A data processing apparatus is provided, which can be a software module or a hardware module, or a combination of both, and is part of a computer device. The apparatus specifically includes:
[0125] An incremental file synchronization module 71 is configured to obtain incremental change data in a current first time period from real-time updated business data, and generate an incremental data file of the current first time period based on the obtained incremental change data.
[0126] A full amount synchronization module 72 is configured to merge, based on incremental data files of each first time period before the current first time period in a current second time period, a first historical full amount data file of a last first time period in a previous second time period, to obtain a second historical full amount data file of the current first time period, the second time period including at least two first time periods; the first historical full amount data file is a file configured to store all business data generated before the end time point of the last first time period in the previous second time period; and the second historical full amount data file is a file configured to store all business data generated before the end time point of the current first time period.
[0127] In one embodiment, the full amount synchronization module 72 includes:
[0128] The period file synchronization module 721 is configured to merge the incremental data file of the current first time period with the first period full data file of the previous first time period in the current second time period to obtain a second period full data file of the current first time period; the first period full data file is a file configured to store all service data generated in the current second time period and before the end time point of the previous first time period in the current second time period, and the second period full data file is a file configured to store all service data generated in the current second time period and before the end time point of the current first time period.
[0129] The history file synchronization module 722 is configured to merge the second period full data file with the first history full data file to obtain a second history full data file of the current first time period.
[0130] In an embodiment, the apparatus further includes a service data updating module.
[0131] The service data updating module is configured to acquire updated service data in real time and store the acquired service data in the first storage medium.
[0132] The incremental file synchronization module 71 is configured to acquire incremental change data in the current first time period from the first storage medium.
[0133] In an embodiment, the period file synchronization module 721 is configured to, after obtaining the first period full data file of the previous first time period, merge the incremental data file of the current first time period with the first period full data file to obtain a second period full data file of the current first time period.
[0134] In an embodiment, the period file synchronization module 721 is further configured to, when the current first time period is the first first time period in the current second time period, synchronize the incremental data file of the current first time period to obtain the second period full data file.
[0135] In one embodiment, the periodical file synchronization module 721 is configured to, when the field content of the current field identification of the current data content identification in the incremental data file of the current first time period is not empty, synchronize the field content of the current field identification of the current data content identification in the incremental data file as the content of the current field identification of the current data content identification to the second periodical full data file; and when the field content of the current field identification of the current data content identification in the incremental data file is empty, synchronize the field content of the current field identification of the current data content identification in the first periodical full data file as the content of the current field identification of the current data content identification to the second periodical full data file.
[0136] In one embodiment, the periodical file synchronization module 721 is configured to merge a target file in the incremental data file of the current first time period and the first periodical full data file of the previous first time period in the current second time period to obtain a merged file, the target file being a file with a file size less than or equal to a first predetermined file size; and obtain the second periodical full data file of the current first time period based on the incremental data file of the current first time period and the first periodical full data file, the file with a file size greater than the first predetermined file size, and the merged file.
[0137] In one embodiment, the history file synchronization module 722 is configured to, when the field content of the current field identification of the current data content identification in the second periodical full data file is not empty, synchronize the field content of the current field identification of the current data content identification in the second periodical full data file as the content of the current field identification of the current data content identification to the second history full data file; and when the field content of the current field identification of the current data content identification in the second periodical full data file is empty, synchronize the field content of the current field identification of the current data content identification in the first history full data file as the content of the current field identification of the current data content identification to the second history full data file.
[0138] In one embodiment, the historical file synchronization module 722 is configured to, when the current first time period is the first first time period in the current second time period, directly merge the period full data file of the current first time period with the first historical full data file to obtain a second historical full data file of the current first time period; and when the current first time period is not the first first time period in the current second time period, merge the period full data file of each first time period in the current second time period with the first historical full data file in parallel to obtain a historical full data file corresponding to each first time period in the current second time period.
[0139] In one embodiment, the historical file synchronization module 722 is configured to, after the merging of the third period full data file of the last first time period in the current second time period with the first historical full data file is completed to obtain a third historical full data file of the last first time period in the current second time period, then merge the period full data file of the first first time period in the next second time period with the third historical full data file.
[0140] In one embodiment, the first time period is 1 hour, and the second time period is 1 day.
[0141] The specific limitations of the data processing apparatus can refer to the limitations of the data processing method described above, which will not be repeated here. Each module in the above data processing apparatus can be realized by software, hardware and their combinations in whole or in part. The above modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory in the computer device in software form, so as to be called and executed by the processor to perform the operations corresponding to each module.
[0142] In one embodiment, a computer device is provided, which can be a server, and its internal structure diagram can be as shown in FIG. 8. Figure 9As shown in the figure. The computer device 90 includes a processor 901, a memory and a network interface 902 connected through a system bus. Among them, the processor 901 of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium 903, an internal memory 904. The non-volatile storage medium 903 stores an operating system 9031, a computer program 9032 and a database 9033. The internal memory 904 provides an environment for the running of the operating system 9031 and the computer program 9032 in the non-volatile storage medium 903. The database 9033 of the computer device is used to store data such as the incremental data file, the periodic full data file, the historical full data file and the like as described above. The network interface 902 of the computer device is used to communicate with the external terminal through the network connection. The computer program is executed by the processor to implement a data processing method.
[0143] Those skilled in the art can understand that, Figure 9 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0144] In one embodiment, a computer device is also provided, including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps in each of the method embodiments described above.
[0145] In one embodiment, a computer readable storage medium is provided, storing a computer program, which is executed by a processor to implement the steps in each of the method embodiments described above.
[0146] In one embodiment, a computer program product or computer program is provided, which includes computer instructions stored in a computer readable storage medium. The processor of the computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in each of the method embodiments described above.
[0147] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, storage, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0148] Any combination of the technical features of the above embodiments can be made. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, but as long as the combination of the technical features does not exist, it should be considered as the scope of the present application.
[0149] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent. It should be pointed out that for those skilled in the art, without departing from the concept of the present application, some modifications and improvements can be made, which are all within the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A data processing method, characterized by, The method comprises: Synchronization task 1, synchronization task 2 and synchronization task 3 are executed in a first time period, and synchronization task 4 is executed in a second time period; Synchronization task 4 is a warm-up task for successful execution of synchronization task 3 of the last first time period of the second time period; In the execution of synchronization task 1, the incremental change data in the current first time period is obtained from the real-time updated business data, and the incremental data file of the current first time period is generated according to the obtained incremental change data; In the execution of synchronization task 2, the incremental data file of the current first time period is merged with the first period full data file of the last first time period in the current second time period to obtain the second period full data file of the current first time period; The first period full data file is a file for storing all business data generated in the current second time period and before the end time point of the last first time period in the current second time period, and the second period full data file is a file for storing all business data generated in the current second time period and before the end time point of the current first time period; In the execution of synchronization task 3, the second period full data file is merged with the first historical full data file to obtain the second historical full data file of the current first time period; The second time period includes at least two first time periods; The first historical full data file is a file for storing all business data generated before the end time point of the last first time period in the last second time period, and the second historical full data file is a file for storing all business data generated before the end time point of the current first time period.
2. The method of claim 1, wherein, The merging of the incremental data file of the current first time period with the first period full data file of the last first time period in the current second time period to obtain the second period full data file of the current first time period comprises: After obtaining the first period full data file of the last first time period in the current second time period, the incremental data file of the current first time period is merged with the first period full data file of the last first time period in the current second time period to obtain the second period full data file of the current first time period.
3. The method of claim 1, wherein, The merging of the incremental data file of the current first time period with the first period full data file of the last first time period in the current second time period to obtain the second period full data file of the current first time period comprises: When the field content of the current field identifier of the current data content identifier in the incremental data file of the current first time period is not empty, the field content of the current field identifier of the current data content identifier in the incremental data file of the current first time period is synchronized to the second period full data file as the content of the current field identifier of the current data content identifier. When the field content identified by the current field identifier of the current data content identifier in the incremental data file of the current first time period is empty, the field content of the current field identifier of the current data content identifier in the first-period full data file is synchronized to the second-period full data file as the content of the current field identifier of the current data content identifier.
4. The method of claim 1, wherein, The merging of the incremental data file of the current first time period and the first-period full data file of the last first time period in the current second time period to obtain the second-period full data file of the current first time period comprises: Merging target files in the incremental data file of the current first time period and the first-period full data file of the last first time period in the current second time period to obtain a merged file, the target file being a file with a file size less than or equal to a first predetermined file size; Based on the incremental data file of the current first time period, the first-period full data file, and the merged file, the second-period full data file of the current first time period is obtained.
5. The method of claim 1, wherein, The merging of the second-period full data file and the first historical full data file to obtain the second historical full data file of the current first time period comprises: When the current first time period is the first first time period in the current second time period, the period full data file of the current first time period is merged with the first historical full data file to obtain the second historical full data file of the current first time period; When the current first time period is not the first first time period in the current second time period, the period full data files of the first time periods in the current second time period are merged with the first historical full data file in parallel to obtain the historical full data files corresponding to the first time periods in the current second time period.
6. The method of claim 1, wherein, After the third-period full data file of the last first time period in the current second time period is merged with the first historical full data file to obtain the third historical full data file of the last first time period in the current second time period, a step of merging the period full data file of the first first time period in the next second time period with the third historical full data file is entered.
7. A data processing apparatus, characterized by, The device is configured to perform a synchronization task 1, a synchronization task 2, and a synchronization task 3 in a first time period, and perform a synchronization task 4 in a second time period. The synchronization task 4 is a dummy task for determining successful execution of the synchronization task 3 of the last first time period in the second time period. An incremental file synchronization module is configured to, when the synchronization task 1 is performed, acquire incremental change data in a current first time period from real-time updated business data, and generate an incremental data file of the current first time period according to the acquired incremental change data. The period file synchronization module is configured to, when the synchronization task 2 is executed, merge the incremental data file of the current first time period with the first period full data file of the previous first time period in the current second time period to obtain a second period full data file of the current first time period; The first period full data file is a file configured to store all service data generated in the current second time period and before the end time point of the previous first time period in the current second time period, and the second period full data file is a file configured to store all service data generated in the current second time period and before the end time point of the current first time period; The history file synchronization module is configured to, when the synchronization task 3 is executed, merge the second period full data file with a first history full data file to obtain a second history full data file of the current first time period; The second time period includes at least two first time periods; The first history full data file is a file configured to store all service data generated before the end time point of the last first time period in the previous second time period, and the second history full data file is a file configured to store all service data generated before the end time point of the current first time period.
8. The data processing apparatus according to claim 7, characterized in that, The period file synchronization module is further configured to, when the first period full data file of the previous first time period in the current second time period is obtained, merge the incremental data file of the current first time period with the first period full data file of the previous first time period in the current second time period to obtain a second period full data file of the current first time period.
9. The data processing apparatus according to claim 7, characterized by The period file synchronization module is further configured to, when the field content of the current field identifier of the current data content identifier in the incremental data file of the current first time period is not empty, synchronize the field content of the current field identifier of the current data content identifier in the incremental data file of the current first time period to the second period full data file as the content of the current field identifier of the current data content identifier. When the field content of the current field identifier of the current data content identifier in the incremental data file of the current first time period is empty, synchronize the field content of the current field identifier of the current data content identifier in the first period full data file to the second period full data file as the content of the current field identifier of the current data content identifier.
10. The data processing apparatus according to claim 7, characterized by, The period file synchronization module is further configured to merge target files in the incremental data file of the current first time period and the first period full data file of the previous first time period in the current second time period to obtain a merged file, the target files being files with a file size less than or equal to a first predetermined file size. Based on the incremental data file of the current first time period, the file with a file size greater than the first predetermined file size in the first period full data file, and the merged file, a second period full data file of the current first time period is obtained.
11. The data processing apparatus according to claim 7, characterized by, The historical file synchronization module is further configured to, when the current first time period is the first first time period in a current second time period, merge the period full data file of the current first time period with the first historical full data file to obtain a second historical full data file of the current first time period. When the current first time period is not the first first time period in the current second time period, the period full data files of each first time period in the current second time period are merged with the first historical full data file in parallel to obtain the historical full data file corresponding to each first time period in the current second time period.
12. The data processing apparatus according to claim 7, characterized by The historical file synchronization module is further configured to, after the third period full data file of the last first time period in the current second time period is merged with the first historical full data file to obtain a third historical full data file of the last first time period in the current second time period, enter the step of merging the period full data file of the first first time period in the next second time period with the third historical full data file.
13. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that, The processor executes the computer program to implement the method in any one of claims 1 to 6.
14. A computer readable storage medium storing a computer program, wherein the computer program comprises program instructions configured to cause a processor to perform the method according to any one of claims 1 to 13. The computer program is executed by the processor to implement the method in any one of claims 1 to 6.
15. A computer program product comprising computer instructions, characterized in that, The computer program is executed by the processor to implement the method in any one of claims 1 to 6. The computer program is executed by the processor to implement the method in any one of claims 1 to 6.
Citation Information
Patent Citations
Data processing method and system
CN103544075A
Full-volume partition view generation method and device, storage medium and electronic device
CN111274253A