Data archiving processing method and system based on batch processing and job scheduling
By introducing batch processing and job scheduling technologies into the data archiving system, the problems of inefficiency of traditional data archiving methods and difficult to ensure data integrity are solved, efficient and automated data archiving and cleaning are achieved, and operational costs are reduced.
Patent Information
- Application Number
- CN202510006405.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-13
AI Technical Summary
Traditional historical data archiving methods are inefficient, prone to errors, difficult to process large-scale data sets, and difficult to ensure the integrity and consistency of data, resulting in data migration time and resource consumption.
The automatic data archiving processing method based on batch processing and job scheduling is adopted, and automated scheduling and execution are achieved through the Quartz and Spring Batch frameworks, including scheduling configuration, task execution management, data archiving, data cleaning and data recovery steps.
It realizes efficient and automated data archiving and cleaning, improves the speed and efficiency of data processing, reduces the risk of human error, ensures data integrity and consistency, and significantly reduces operating costs.
Smart Images

Figure CN119988313A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and in particular relates to a method and system for automatic data archiving based on batch processing and job scheduling. Background Art
[0002] With the rapid development of aviation information technology, domestic freight rate publishing business has accumulated a large amount of historical data related to flights, passengers, baggage, etc. This data is crucial for business analysis, customer service, and compliance audits. However, as time goes by, the amount of data gradually increases, which puts a heavy burden on the database, not only increasing storage costs, but also reducing the overall performance of the system.
[0003] Traditional historical data archiving methods often rely on manual operations or simple scripting tools, which are not only inefficient but also prone to errors. These problems become more prominent when dealing with large-scale data sets, resulting in time-consuming data migration, high resource consumption, and difficulty in ensuring data integrity and consistency. As business needs change, the frequency of data archiving is also increasing, which requires the system to have higher flexibility and responsiveness. Summary of the invention
[0004] In view of the deficiencies in the prior art, the present application proposes a method and system for automatic data archiving based on batch processing and job scheduling. The technical problem to be solved by the present application is achieved through the following technical solutions.
[0005] The first aspect of the present application provides a data automatic archiving processing method based on batch processing and job scheduling, comprising:
[0006] Scheduling configuration steps: Use the Quartz framework to schedule and manage job tasks;
[0007] Task execution management steps: Track the progress of task execution in the metadata table of the batch processing framework Spring Batch;
[0008] Data archiving step: back up data that meets specific conditions from the original database to the archiving server;
[0009] Data cleaning step: Delete the confirmed archived data from the source system to free up storage space and optimize the performance of the source system;
[0010] Data recovery step: recovering data from the archive server to the source system.
[0011] In combination with the first aspect, the above scheduling configuration step further includes:
[0012] It provides a job history viewing function, which includes the start and end time of each execution and key information of the execution result; it provides a job control function to realize job pause, resume, and delete functions, and flexibly control job execution.
[0013] In combination with the first aspect, the task execution management step further includes:
[0014] The execution status record not only includes the success or failure status of the entire task, but also records the status information of each subtask and each data block processing in detail.
[0015] In combination with the first aspect, the above data archiving step further includes: based on the Reader, Processor and Writer components in the Spring Batch framework, the Reader component reads data that meets preset conditions from the original database, the Processor component processes the read data, and the Writer component writes the processed data to the archive server.
[0016] In combination with the first aspect, the data archiving step further includes:
[0017] Create an archive scheduling task sub-step. Based on business needs, use the module interface to create archive, clear, and restore tasks, and configure the task information.
[0018] Configure the scheduling task sub-steps, set the necessary parameters to be passed to the data archiving task, and configure the execution time of the trigger.
[0019] Scheduling task execution sub-steps, using differentiated execution strategies for different types of tasks to ensure that each task can be completed in the best way,
[0020] Monitor the task execution sub-steps, regularly query various log tables and metadata tables generated during the task execution process, and obtain detailed information and execution status of the task execution.
[0021] In combination with the first aspect, the above-mentioned scheduling task execution sub-step further includes:
[0022] The data archiving task executes the sub-steps, reading the data to be archived from the source database and distributing it to multiple execution nodes. The execution nodes independently process the distributed data and transmit it to the archiving server in parallel.
[0023] The data cleaning task execution sub-step queries the archived record log table of the source database for data identifiers that meet the cleaning conditions, deletes these data records in sequence, and records relevant information;
[0024] The data recovery task executes sub-steps, blocks the data read from the archive server, performs preliminary unpacking and format checking, and rewrites the data into the corresponding table according to the format requirements of the source database.
[0025] In combination with the first aspect, the data recovery task execution sub-step further includes:
[0026] When the trigger reaches the trigger time, the job task is automatically started to execute the corresponding batch task. The following two algorithms are used to schedule the task:
[0027] (1) Master-slave algorithm: divide the data table to be cleaned into master table and sub-table according to the business scope. When cleaning data, the order is sub-table first and then master table. When restoring data, the order is master table first and then sub-table.
[0028] (2) Allocation balancing algorithm: Evenly distribute the data among several tasks according to the data volume. For tables with too much data, the data is evenly distributed to different tasks according to the modulus strategy to ensure that the amount of data processed by each task is relatively balanced.
[0029] In combination with the first aspect, the above data recovery step further includes: the data recovery module draws on the logic of the data archiving module, reversely executes the data writing step, and reads back to the source system after locating the data on the archive server according to the recovery conditions specified by the user.
[0030] The second aspect of the present application provides a data automatic archiving processing system based on batch processing and job scheduling, comprising:
[0031] Scheduling configuration module: Use the Quartz framework to schedule and manage job tasks;
[0032] Task execution management module: Tracks the progress of task execution in the metadata table of the batch processing framework Spring Batch to understand the current stage of the task;
[0033] Data archiving module: backs up data that meets specific conditions from the original database to the archiving server;
[0034] Data cleaning module: deletes confirmed archived data from the source system to free up storage space and optimize the performance of the source system;
[0035] Data recovery module: recovers data from the archive server to the source system.
[0036] The third aspect of the present application provides a computer-readable storage medium storing one or more programs, characterized in that when the one or more programs are executed, the above-mentioned dynamic connection path building method can be implemented.
[0037] Compared with the prior art, the present invention has the following advantages:
[0038] 1. High degree of automation: By integrating Quartz and Spring Batch, the automatic scheduling and execution of data processing tasks are realized, which reduces the risks caused by human errors and greatly improves the speed and efficiency of data processing.
[0039] 2. High reliability: It has a complete error handling and retry mechanism to ensure that data processing tasks can be completed smoothly. If an exception occurs during data migration or storage, it will automatically record the error information and try to re-execute the task. At the same time, it guarantees the integrity and accuracy of data during migration and storage within the system, thereby improving the overall reliability of the system.
[0040] 3. Cost-effectiveness: Automated processing significantly reduces labor costs. By using intelligent historical data archiving methods, work efficiency is improved, so that data archiving can be completed in a shorter time, further saving operating costs.
[0041] 4. Data protection: Through regular backup and recovery functions, data loss can be effectively prevented and business continuity can be guaranteed. Regular backup can ensure that data can be quickly restored in the event of hardware failure, software error or other unexpected situations. It not only reduces the risk of data loss, but also ensures the continuous operation of the business, avoiding business interruption and economic losses caused by data loss.
[0042] 5. Easy to maintain: It provides an intuitive and friendly user interface, where operators can configure and monitor data archiving tasks through simple page operations without writing complex scripts or commands, making daily management and troubleshooting easier and faster.
[0043] In summary, this application can automate tasks, reduce human errors, and achieve efficient archiving of data from databases to non-database media, automatic cleaning of historical data, and rapid recovery of data by defining workflows and executing batch archiving operations of data at regular intervals. At the same time, a data consistency guarantee strategy is adopted to ensure the integrity and accuracy of data during the archiving process, significantly reducing data management costs and improving data processing efficiency.
[0044] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0046] Figure 1 The overall architecture diagram of the data automatic archiving processing system based on batch processing and job scheduling is shown;
[0047] Figure 2 The overall flow chart of the data automatic archiving processing method based on batch processing and job scheduling is shown;
[0048] Figure 3 A data flow chart of a data automatic archiving processing method based on batch processing and job scheduling is shown;
[0049] Figure 4 A flow chart of data cleaning steps is shown;
[0050] Figure 5 A flow chart of data recovery steps is shown;
[0051] Figure 6 A data archiving task component diagram is shown. DETAILED DESCRIPTION
[0052] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0053] Before explaining the present application in detail, the terms used in the embodiments of the present application are explained.
[0054]
[0055] Since traditional manual or semi-automatic data processing methods can no longer meet the requirements of high efficiency and high reliability, the embodiment of the present application proposes a solution based on the Java programming language, which integrates two powerful open source frameworks: batch processing framework Spring Batch and Quartz.
[0056] As the core batch processing engine, Spring Batch provides a wealth of functions to support complex data processing tasks. It supports block processing, transaction management, and parallel processing, and can flexibly build various data processing processes, such as data backup, migration, and recovery. By defining jobs and steps, the system can effectively execute large-scale data processing tasks, ensuring the accuracy and reliability of each link.
[0057] Quartz, as a job scheduler, is responsible for triggering Spring Batch jobs according to a predetermined schedule. Quartz provides a wealth of scheduling configuration options, including simple scheduled tasks, complex Cron expressions, and cluster support, to ensure that data processing tasks can be executed on time and accurately.
[0058] This combination enables data archiving to run automatically without human intervention, greatly reducing the need for manual operation and improving the automation level of the system.
[0059] 1. Overall technical solution
[0060] like Figure 1 As shown, the data automatic archiving processing system embodiment based on batch processing and job scheduling of the present application mainly includes the following modules: scheduling configuration module, task execution management module, data archiving module, data cleaning module, data recovery module, etc.
[0061] (1) Scheduling configuration module:
[0062] This module uses the Quartz framework to implement the scheduling and management of job tasks. Users can flexibly specify the task execution time according to actual needs, which can be a specific date, a fixed interval (such as every hour, every day), or a complex Cron expression. The module also provides a job history viewing function, including key information such as the start and end time of each execution, the execution result, etc. In terms of job management, users can dynamically adjust the execution time of jobs, such as temporarily postponing the execution time of some non-critical jobs due to business peaks, or starting emergency tasks in advance. For the monitoring of job status, it can reflect in real time whether the job is currently in the waiting, executing, suspended, completed or failed states. Detailed log information records the key steps and events in the job execution process, including system initialization, parameter loading, task startup, data processing, and possible error details. In addition, this module also has a job control function, which can realize the job suspension, resumption, and deletion functions, and flexibly control the execution of jobs.
[0063] (2) Task execution management module:
[0064] This module tracks the progress of task execution in the Spring Batch metadata table, allowing users to clearly understand the current stage of the task, such as how many records have been completed in the data reading stage, the progress of the data processing stage, and the start of the data writing stage. At the same time, the execution status record not only includes the success or failure status of the overall task, but also records the status information of each subtask and each data block processing in detail.
[0065] (3) Data archiving module:
[0066] This module is responsible for backing up data that meets specific conditions from the original database to the archive server. Its core technology implementation relies on the Reader, Processor, and Writer components in the Spring Batch framework. The Reader component is responsible for accurately reading data that meets the preset conditions from the original database. During the data reading process, a parallel reading strategy is adopted to improve the efficiency of data reading. The Processor component performs necessary processing on the read data and unifies the data in different data sources into the format required by the archive server. The Writer component finally writes the processed data to the archive server.
[0067] (4) Data cleaning module:
[0068] This module is responsible for deleting confirmed archived data from the source system to free up storage space and optimize the performance of the source system. Similar to the data archiving module, Spring Batch is used to implement data cleansing operations, but it focuses more on data filtering and deletion.
[0069] (5) Data recovery module:
[0070] This module is responsible for recovering data from the archive server to the source system. The data recovery module draws on the logic of the data archive module, reverses the data writing steps, and reads the data back to the source system after locating it on the archive server according to the recovery conditions specified by the user.
[0071] 2. Specific implementation principle
[0072] In order to implement the historical data archiving system and solve the problems of high data migration cost, performance issues, data integrity and consistency, the following strategies are adopted:
[0073] (1) Adapting to different flight requirements: The system can support the operation of a single flight through different settings or parameter configurations, and can also support the simultaneous operation of multiple flights, effectively manage resources, simplify operation management, and reduce system complexity.
[0074] (2) Automated scheduling tasks: Data archiving tasks are preset through the scheduling configuration module, and tasks are automatically executed according to preset conditions. The scheduling center is responsible for monitoring the execution of scheduling, and once the conditions are met, the corresponding batch processing tasks of data archiving, data cleaning or data recovery are immediately triggered.
[0075] (3) Data consistency assurance: The system can ensure that the integrity and consistency of data are reliably guaranteed in a complex environment with large amounts of data, thereby effectively preventing the risk of misoperation caused by data inconsistency.
[0076] (4) Reasonable task scheduling and parallel processing: Through the dynamic resource scheduling mechanism, cluster resources can be fully utilized to achieve parallel execution of tasks, improve resource utilization, optimize task execution efficiency, and improve processing performance.
[0077] (5) Visual log: It can conveniently and quickly monitor the details of scheduling execution, which helps to improve system maintenance efficiency and reduce operation and maintenance costs.
[0078] (6) Execution history is traceable: The batch processing engine will output execution information to the metadata table during the execution of the job task. We can quickly query the task execution status and trace the historical execution information of the task through the interface.
[0079] The flowchart of the data automatic archiving processing method based on batch processing and job scheduling of the present application is as follows Figure 2 As shown, including:
[0080] Scheduling configuration steps: Use the Quartz framework to schedule and manage job tasks;
[0081] Task execution management steps: Track the progress of task execution in the metadata table of the batch processing framework Spring Batch;
[0082] Data archiving step: back up data that meets specific conditions from the original database to the archiving server;
[0083] Data cleaning step: Delete the confirmed archived data from the source system to free up storage space and optimize the performance of the source system;
[0084] Data recovery step: recovering data from the archive server to the source system.
[0085] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0086] For the data flow of historical data archiving, see the attached Figure 3 :
[0087] Step 1, create an archiving scheduling task. Specify the tasks to be performed, such as data archiving task 1, data cleaning task, data recovery task 2, etc. In this initial stage, the system uses the scheduling configuration module to start the creation process of data archiving system tasks. Operators use the module interface to create archiving, clearing and recovery tasks according to business needs, and configure the task information. Take the creation of a data archiving task to be executed at 3 am on the 1st of each month as an example. At this time, the system will add a new record in the task configuration table, including the task name (such as "data archiving task"), execution cycle (03:00:00 on the 1st of each month) and related business parameters (such as airline code "CA", etc.). These data will serve as the basis for the execution of subsequent tasks, clarify the execution rules and scope of the tasks, ensure that the tasks are started at the scheduled time and conditions, and are associated with specific business data.
[0088] Step 2: Configure the scheduling task. Set the necessary parameters to be passed to tasks such as data archiving, such as airline, processing date, number of batches to be processed, etc. Configure the execution time of the trigger, such as starting at 1 a.m. every day.
[0089] Step 3: Schedule task execution. When the system environment meets the pre-set task execution conditions, the scheduled task will be triggered for execution. This process uses differentiated execution strategies for different types of tasks to ensure that each task can be completed in the best way.
[0090] (I) Data archiving task execution
[0091] When the system time reaches 3:00 a.m. on the 1st of each month, the execution conditions of the "data archiving task" are met and the task is triggered for execution. The system first reads the data to be archived from the source database according to the preset query conditions (such as filtering by time range and airline code), and distributes it to multiple execution nodes according to certain data block division rules. Each node independently processes the assigned data block and transmits it to the archive server in parallel. During the transmission process, the data will be packaged into a specific data packet format with verification information attached. At the same time, the detailed information of this archiving task, including the source, quantity, timestamp, and storage path of the archived data, is recorded in the task execution log table of the archive server for subsequent query and audit. During the whole process, the data is read from the original table structure of the source database, processed, and then flows to the storage location of the archive server, realizing data migration from the production environment to the archive environment, and completing the long-term preservation and historical traceability preparation of the data.
[0092] (II) Data cleaning task execution
[0093] Suppose that at a certain point in time, the system triggers a data cleanup task. For example, data that has been successfully archived and exceeds a certain retention period (3 years) needs to be deleted from the source database to free up storage space. The system first queries the archive record log table of the source database for data identifiers that meet the cleanup conditions, and then deletes these data records in sequence. The deletion time and related information of the record will be recorded in the execution log table of the cleanup task to facilitate subsequent audits and troubleshooting. Due to the sensitivity of the data cleanup task, the entire process is executed serially on a single node to ensure data consistency and integrity.
[0094] (III) Data recovery task execution
[0095] When data loss or data corruption requires recovery, assume that a data recovery task for a specific date range and airline is triggered. The system will locate the corresponding archived data file or data block from the storage location of the archive server according to the conditions specified by the task (such as time range and airline code), and transfer it back to the source system. During the data recovery process, multi-node parallelism is also used to improve efficiency. After the data read from the archive server is divided into blocks, the data will first be preliminarily unpacked and format checked, and then the data will be rewritten to the corresponding table according to the table structure and data format requirements of the source database. The execution log table will record the recovery time, source, and related information. The entire process enables data to flow back from the storage location of the archive server to the source system, restores data availability, and ensures the normal operation of the business system.
[0096] After the system is started, when the trigger reaches the trigger time, the job task will be automatically started to execute the corresponding batch task. Here, the following two principles need to be followed to arrange the tasks to achieve the best efficiency of data archiving.
[0097] (1) Master-slave principle. The data table to be cleaned is divided into master table and sub-table according to the business scope. The order of data cleaning is sub-table first and master table second. Figure 4 When restoring data, the arrangement order is the main table first and then the sub-table. Figure 5 .
[0098] Specifically, during the data cleaning process, the system follows the master-slave principle such as Figure 4 Here, "rule slave table 1" to "rule slave table n" represent the data of multiple different slave tables that need to be cleaned. When the data cleaning of all slave tables is completed, the system will enter the data cleaning stage of the master table in sequence. "Rule master table 1" to "rule master table n" represent the data of multiple different master tables that need to be cleaned. This cleaning order of slave tables first and master tables later helps to ensure the integrity and consistency of data and avoid data conflicts or errors caused by the existence of master table related data when cleaning slave table data.
[0099] Specifically, if Figure 4 As shown in the figure, in Spring Batch, data processing is usually done in blocks. For slave table data cleaning, each "rule slave table" can be regarded as a processing unit of a data block. After the processing of the slave table data block is completed, the data flows to the master table processing stage. Each "rule master table" is also a processing unit. Slave table cleaning and master table cleaning can be regarded as two different steps in a complete job. The job execution mechanism of Spring Batch ensures that these steps are executed in a specific order, that is, the slave table data cleaning step is executed first, and then the master table data cleaning step is executed, ensuring the orderliness and data consistency of the data cleaning job. Slave table data cleaning is a pre-step of the job because slave table data often depends on the associated data of the master table. Such a master-slave relationship design can ensure that when cleaning slave table data, there will be no data inconsistency or errors due to the existence of master table data. Master table data cleaning is performed after slave table cleaning, which ensures the logical coherence and data integrity of the entire data cleaning job.
[0100] During data recovery, the system follows the master-slave principle. Figure 5 . In contrast to the data cleansing process, the system processes the master table data first. "Rule Master Table 1" to "Rule Master Table n" represent multiple different master table data that need to be restored. The recovery process of each master table involves reading data from the archive server (TFTP), processing it through the processor, and then writing the recovered data to the database. When the data recovery of the master table is completed, the system will enter the data recovery phase of the slave table in sequence. "Rule Slave Table 1" to "Rule Slave Table n" represent multiple different slave table data that need to be recovered. Their operation process is similar to that of the master table. The recovery of the slave table data is also completed through reading, processing and writing operations. This recovery order of the master table first and the slave table later ensures the relevance and consistency of the data during the recovery process, so that the relevant data can be correctly restored to the system.
[0101] Specifically, if Figure 5. In the "Master Table Orchestration", each "Rule Master Table" contains a "Reader (TFTP)", a "Processor" and a "Writer (DB)". The "Reader (TFTP)" is the ItemReader in Spring Batch, which is responsible for reading the master table data from the archive server (TFTP). The "Processor" corresponds to the ItemProcessor, which processes the read data, such as data format conversion, business logic verification, etc. The "Writer" corresponds to the ItemWriter, which writes the processed data to the database. When the master table data is restored, the slave table data recovery phase begins. The slave table data recovery process also involves reading data, processing data and writing to the database. As the core of the data, the recovery of the master table's data is a prerequisite for the recovery of the slave table data. This design ensures the consistency and integrity of the data and avoids conflicts and errors during the data recovery process.
[0102] (2) Balanced distribution principle. The data is evenly distributed among several tasks according to the data volume. For tables with too large data volume, they are divided according to a certain strategy and then distributed to various tasks. For example, the freight rate table for data archiving and data clearing is divided according to the modulus strategy and then distributed to various tasks. Figure 6 .
[0103] Specifically, Figure 6 It demonstrates the application of the principle of balanced distribution in data processing. When the amount of data to be processed is too large, in order to improve processing efficiency, the system will split the data table according to a certain strategy, and distribute the split data to different tasks. For example, "JOB1", "JOB2" and "JOB3" in the figure represent different tasks. In "JOB1", the data is divided into multiple parts such as "PART0" to "PART9". Similarly, "JOB2" and "JOB3" also have corresponding data parts. These data parts are obtained by splitting the original data table (such as the freight rate table) according to the modulo strategy. The formula below shows the specific modulo operation, such as "MOD
[0104] (FARE_REC_NO,15)=0------1569424" means to perform a modulo operation on the data column "FARE_REC_NO" (take the remainder after dividing by 15). When the remainder is 0, the corresponding number of records is 1569424. Through this modulo strategy, data can be evenly distributed to different tasks to ensure that the amount of data processed by each task is relatively balanced, thereby improving the efficiency of the entire data archiving and clearing process.
[0105] In Spring Batch, when the amount of data is too large, you can use a partitioning strategy to divide the data into multiple subsets (such as Figure 6"PART0"-"PART14" in JOB1), each subset can be processed independently in different threads to improve processing efficiency. The modulo strategy (such as "MOD(FARE_REC_NO,15)") is the rule for partitioning. Through this rule, the original large table data is evenly divided into different partitions to ensure that the amount of data processed by each task is relatively balanced. When the data is reasonably distributed to these tasks, Spring Batch can use the multi-threading mechanism to execute these tasks simultaneously. For example, "PART0"-"PART9" in JOB1, "PART10"-"PART13" in JOB2, and "PART14" in JOB3 can be processed in parallel in different threads. This processing method not only improves the efficiency of data processing, but also ensures the accuracy and consistency of data processing. The parallel processing mechanism allows multiple tasks to be carried out simultaneously, making full use of system resources, speeding up the execution of the entire data processing job, and meeting the needs of efficient data processing.
[0106] Step 4: Monitor task execution. You can view the status of the scheduling operation and view the logs in real time through the scheduling configuration module. You can also view the status and progress of batch task execution through the historical data processing interface. During the entire task execution process, the task execution management module continuously monitors and summarizes the task status. It obtains detailed information on task execution by regularly querying various log tables and metadata tables generated during task execution, including the start time, end time, amount of data processed, and execution status (success, failure, or executing) of each task step.
[0107] The intelligent historical data archiving method realizes an efficient, reliable and automated historical data archiving process. It can not only quickly complete the data migration and archiving tasks, but also provide transparency and traceability of the entire processing process, making the data governance process more efficient, convenient and error-free.
[0108] This application has the following advantages:
[0109] 1. High degree of automation: By integrating Quartz and Spring Batch, the automatic scheduling and execution of data processing tasks are realized. This high degree of automation reduces the need for manual intervention, reduces the risk caused by human errors, and greatly improves the speed and efficiency of data processing. The operator only needs to configure the relevant parameters, and the system can automatically complete the data archiving task without human supervision.
[0110] 2. High reliability: The system has a complete error handling and retry mechanism to ensure that data processing tasks can be completed smoothly. If an exception occurs during data migration or storage, the system will automatically record the error information and try to re-execute the task, while ensuring the integrity and accuracy of data during migration and storage within the system, thereby improving the overall reliability of the system.
[0111] 3. Cost-effectiveness: Automated processing significantly reduces labor costs. Traditional manual data archiving requires a lot of human participation, which is not only time-consuming but also prone to errors. By using intelligent historical data archiving methods, enterprises can significantly reduce their reliance on human resources, freeing employees from tedious data processing work and focusing on more valuable tasks. At the same time, automated processing also improves work efficiency, allowing data archiving to be completed in a shorter time, further saving operating costs.
[0112] 4. Data protection: The system effectively prevents data loss and ensures business continuity through regular backup and recovery functions. Regular backup can ensure that data can be quickly restored in the event of hardware failure, software error or other unexpected situations. This not only reduces the risk of data loss, but also ensures the continuous operation of the business and avoids business interruption and economic losses caused by data loss.
[0113] 5. Easy to maintain: The system provides an intuitive and friendly user interface, making it easy for non-technical personnel to manage data processing tasks. Operators can configure and monitor data archiving tasks through simple page operations without writing complex scripts or commands. This ease of use greatly reduces the difficulty of system maintenance, making daily management and troubleshooting easier and faster. In addition, the system also supports detailed logging and real-time monitoring to help administrators discover and solve problems in a timely manner, further enhancing the maintainability of the system.
[0114] In summary, the present application provides an efficient, secure and economical data archiving solution, which not only improves the efficiency and quality of data processing, but also significantly improves data security.
[0115] Based on the same inventive concept, the present application also provides a computer-readable storage medium storing one or more programs, which can implement the aforementioned method when the one or more programs are executed.
[0116] Although the present invention is described in detail with reference to the foregoing embodiments, the present application uses Oracle as a database management system. Other relational databases may also be used to replace the Oracle database, such as MySQL. The present application uses Quartz as a job scheduler. Other alternative job scheduling tools or job scheduling platforms may also be used. The present application uses Spring Batch as a batch processing engine. Other alternative batch processing engines may also be used.
[0117] Those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent substitutions for some of the technical features therein; and these modifications or substitutions do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for automatic data archiving based on batch processing and job scheduling, characterized in that: include: Scheduling configuration steps: Use the Quartz framework to schedule and manage job tasks; Task execution management steps: Track the progress of task execution in the metadata table of the batch processing framework Spring Batch; Data archiving step: back up data that meets specific conditions from the original database to the archiving server; Data cleaning step: Delete the confirmed archived data from the source system to free up storage space and optimize the performance of the source system; Data recovery step: recovering data from the archive server to the source system.
2. The method according to claim 1, characterized in that The scheduling configuration step further includes: It provides a job history viewing function, which includes the start and end time of each execution and key information of the execution result; it provides a job control function to realize job pause, resume, and delete functions, and flexibly control job execution.
3. The method according to claim 2, characterized in that The task execution management step further includes: The execution status record not only includes the success or failure status of the entire task, but also records the status information of each subtask and each data block processing in detail.
4. The method according to claim 3, characterized in that The data archiving step further includes: based on the Reader, Processor and Writer components in the Spring Batch framework, the Reader component reads data that meets preset conditions from the original database, the Processor component processes the read data, and the Writer component writes the processed data to the archiving server.
5. The method according to claim 4, characterized in that The data archiving step further comprises: Create an archive scheduling task sub-step. Based on business needs, use the module interface to create archive, clear, and restore tasks, and configure the task information. Configure the scheduling task sub-steps, set the necessary parameters to be passed to the data archiving task, and configure the execution time of the trigger. Scheduling task execution sub-steps, using differentiated execution strategies for different types of tasks to ensure that each task can be completed in the best way, Monitor the task execution sub-steps, regularly query various log tables and metadata tables generated during the task execution process, and obtain detailed information and execution status of the task execution.
6. The method according to claim 5, characterized in that The scheduling task execution sub-step further includes: The data archiving task executes the sub-steps, reading the data to be archived from the source database and distributing it to multiple execution nodes. The execution nodes independently process the distributed data and transmit it to the archiving server in parallel. The data cleaning task execution sub-step queries the archived record log table of the source database for data identifiers that meet the cleaning conditions, deletes these data records in sequence, and records relevant information; The data recovery task executes sub-steps, blocks the data read from the archive server, performs preliminary unpacking and format checking, and rewrites the data into the corresponding table according to the format requirements of the source database.
7. The method according to claim 6, characterized in that The data recovery task execution sub-step further includes: When the trigger reaches the trigger time, the job task is automatically started to execute the corresponding batch task. The following two algorithms are used to schedule the task: (1) Master-slave algorithm: divide the data table to be cleaned into master table and sub-table according to the business scope. When cleaning data, the order is sub-table first and then master table. When restoring data, the order is master table first and then sub-table. (2) Allocation balancing algorithm: Evenly distribute the data among several tasks according to the data volume. For tables with too much data, the data is evenly distributed to different tasks according to the modulus strategy to ensure that the amount of data processed by each task is relatively balanced.
8. The method according to claim 7, characterized in that The data recovery step further includes: the data recovery module uses the logic of the data archiving module to reversely execute the data writing step, and reads the data back to the source system after locating the data on the archive server according to the recovery conditions specified by the user.
9. A data automatic archiving processing system based on batch processing and job scheduling, characterized in that: include: Scheduling configuration module: Use the Quartz framework to schedule and manage job tasks; Task execution management module: Tracks the progress of task execution in the metadata table of the batch processing framework Spring Batch to understand the current stage of the task; Data archiving module: backs up data that meets specific conditions from the original database to the archiving server; Data cleaning module: deletes confirmed archived data from the source system to free up storage space and optimize the performance of the source system; Data recovery module: recovers data from the archive server to the source system.
10. A computer-readable storage medium storing one or more programs, characterized in that: When the one or more programs are executed, the dynamic connection path building method described in any one of claims 1 to 8 can be implemented.