Third-party cloud service data synchronous processing method and system based on timed task

Through timing tasks and asynchronous processing mechanisms, the task time and frequency are dynamically adjusted, and incremental synchronization and conflict resolution strategies are adopted to solve the problem of resource waste and data delay in data synchronization of third-party cloud services, achieving efficient and reliable data processing and system stability.

CN120336429APending Publication Date: 2025-07-18SHANGHAI QUZHI NETWORK TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510427046.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, the data synchronization processing of third-party cloud service has problems such as resource waste, data delay, system performance degradation and data conflict loss, especially when server resources are limited and interfaces are unstable, it affects the normal progress of the business.

Method used

The asynchronous data processing method based on timing tasks is adopted, the task attributes are configured through the management interface, the task time and frequency are dynamically adjusted using the ARIMA model, incremental synchronization and conflict detection and resolution mechanisms are adopted, and data interaction and synchronization are combined with asynchronous connection pooling and hash functions are used to monitor and automatically retry.

Benefits of technology

Optimize server resource utilization, improve data processing efficiency and accuracy, enhance system stability and reliability, and reduce data backlog and business interruption risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336429A_ABST
    Figure CN120336429A_ABST
Patent Text Reader

Abstract

The invention discloses a third-party cloud service data synchronization processing method and system based on a timed task, and the method carries out the attribute configuration of the timed task through a management interface, and the objects of the attribute configuration comprise execution time, frequency, a third-party cloud service data source and target system information. Storing the object of the attribute configuration in an attribute database; when the timed task is triggered, acquiring data from a third-party cloud service in an asynchronous mode, storing the data to a local cache, and recording a data identifier; synchronizing the processed data to a target system by adopting an incremental synchronization strategy according to the data identifier, and performing conflict detection and conflict resolution in the synchronization process; the execution state of the timed task is monitored in real time, and if the execution state of the timed task is abnormal, a report is generated and automatic retry is tried. According to the method, the problems of resource waste, data delay, system performance reduction, data conflict loss and the like confronted by third-party cloud service data synchronous processing in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of cloud computing and data processing, and particularly relates to a method and system for third-party cloud service data synchronization processing based on scheduled tasks. Background Art

[0002] In today's information technology environment, enterprises and organizations often need to integrate data from multiple different systems to achieve goals such as coordinated operation of business processes and data analysis. However, in the case of limited server resources, traditional real-time data synchronization and processing methods face many challenges.

[0003] On the one hand, there are often differences in the data formats, interface specifications, and data update frequencies of different systems, which makes real-time data synchronization complex and error-prone. For example, System A may push updated data in XML format every 5 minutes, while System B requires data in JSON format and can only receive data updates within a specific time period. Real-time processing of such data requires a large amount of conversion and adaptation work, consuming a large amount of server resources.

[0004] On the other hand, real-time data processing requires the server to continuously monitor and process data changes from different systems, which can lead to a decline in system performance and even cause the system to freeze or crash when server resources are scarce. For example, in a high-concurrency data update scenario, the server may not be able to respond to new data requests in a timely manner, resulting in data backlog and processing delays, affecting the normal operation of the business. Especially when the purchased third-party service is not within the enterprise intranet, the instability of interface communication will affect the normal operation of the business and even cause serious consequences of data inconsistency. Summary of the Invention

[0005] To this end, the present invention provides a method and system for third-party cloud service data synchronization processing based on scheduled tasks, which solves problems such as resource waste, data delay, system performance decline, and data conflict and loss faced by third-party cloud service data synchronization processing in the prior art.

[0006] To achieve the above object, the present invention provides the following technical solution: A method for third-party cloud service data synchronization processing based on scheduled tasks, including:

[0007] Configuring the attributes of the scheduled task through the management interface, where the objects of the attribute configuration include the execution time, frequency, third-party cloud service data source, and target system information, and storing the objects of the attribute configuration in the attribute database;

[0008] When the scheduled task is triggered, asynchronously obtain data from the third-party cloud service, store it in the local cache, and record the data identifier;

[0009] Perform format conversion, data cleaning, and verification operations on the data in the cache according to preset rules;

[0010] According to the data identifier, adopt an incremental synchronization strategy to synchronize the processed data to the target system, and perform conflict detection and conflict resolution during the synchronization process;

[0011] Monitor the execution status of the scheduled task in real time. If the execution status of the scheduled task is abnormal, generate a report and attempt to automatically retry.

[0012] As an optimal solution for the data synchronization processing method of the third-party cloud service based on the scheduled task, dynamically adjust the execution time and frequency of the scheduled task through the management interface. The dynamic adjustment adopts a scheduling model based on the time series prediction algorithm. The scheduling model updates the time series T = {t1, t2, … t n} according to the historical data of the third-party cloud service, uses the ARIMA(p, d, q) model, and determines the optimal model parameters by minimizing the Akaike information criterion or the Bayesian information criterion to predict the next data update time of the third-party cloud service, so as to dynamically adjust the execution time and frequency of the scheduled task; where p is the autoregressive order, d is the differencing order, and q is the moving average order.

[0013] As an optimal solution for the data synchronization processing method of the third-party cloud service based on the scheduled task, when the scheduled task is triggered, during the process of obtaining data from the third-party cloud service asynchronously:

[0014] Adopt the asynchronous connection pool technology to interact with the third-party cloud service. In the connection pool management, adopt a traffic control mechanism based on the token bucket algorithm, set the token generation rate r and the capacity b of the bucket. Each time you request data from the third-party cloud service, try to obtain a token. If there are tokens meeting the set quantity in the bucket, the request is allowed; if there is no token in the bucket, the data request waits or is rejected.

[0015] As an optimal solution for the data synchronization processing method of the third-party cloud service based on the scheduled task, during the process of synchronizing the processed data to the target system according to the data identifier:

[0016] By comparing the data identifiers, only synchronize the data that has changed since the last synchronization; when comparing the data identifiers, map the data identifiers to a hash ring with a fixed range through a hash function. If the positions of the data identifiers on the hash ring change during two synchronizations, or the hash values of the data contents corresponding to the data identifiers are different, it is determined that the data has changed.

[0017] As an optimal solution for the data synchronization processing method of the third-party cloud service based on the scheduled task, the conflict detection is to compare the target system and the data to be synchronized during the data synchronization process;

[0018] The conflict resolution strategies include overwriting the target data with the source data as the standard, selecting the latest data according to the timestamp, or performing data merging; when performing data merging, for numerical data, a weighted average algorithm is used for merging. Assume the data in the target system is A, the data to be synchronized is B, and the weights are w1 and w2 respectively. The merged data C = w1A + w2B;

[0019] For text data, the longest common subsequence algorithm is used to find the longest common subsequence of the two texts, and the different parts are spliced and merged according to preset rules.

[0020] The present invention also provides a third-party cloud service data synchronization processing system based on a timing task, including:

[0021] A task configuration module for configuring the attributes of the timing task through a management interface. The objects of the attribute configuration include the execution time, frequency, third-party cloud service data source, and target system information, and the objects of the attribute configuration are stored in an attribute database;

[0022] A data acquisition module for asynchronously acquiring data from a third-party cloud service when the timing task is triggered, storing it in a local cache, and recording the data identifier;

[0023] A data processing module for performing format conversion, data cleaning, and verification operations on the data in the cache according to preset rules;

[0024] A data synchronization module for synchronizing the processed data to the target system using an incremental synchronization strategy based on the data identifier, and performing conflict detection and conflict resolution during the synchronization process;

[0025] A monitoring and feedback module for real-time monitoring of the execution status of the timing task. If the execution status of the timing task is abnormal, a report is generated and an automatic retry is attempted.

[0026] As a preferred solution of the third-party cloud service data synchronization processing system based on a timing task, in the task configuration module:

[0027] The execution time and frequency of the timing task are dynamically adjusted through the management interface. The dynamic adjustment uses a scheduling model based on a time series prediction algorithm, and the scheduling model updates the time series T = {t1, t2,... t n}, using the ARIMA(p,d,q) model, determine the optimal model parameters by minimizing the Akaike Information Criterion or the Bayesian Information Criterion, and predict the next data update time of the third-party cloud service to dynamically adjust the execution time and frequency of the timing task; where p is the autoregressive order, d is the differencing order, and q is the moving average order.

[0028] As an optimal solution for the third-party cloud service data synchronization processing system based on timing tasks, in the data acquisition module:

[0029] Adopt the asynchronous connection pool technology to interact with the third-party cloud service. In the connection pool management, adopt a traffic control mechanism based on the token bucket algorithm, set the token generation rate r and the capacity b of the bucket. Each time requesting data from the third-party cloud service, try to obtain a token. If there are tokens in the bucket that meet the set quantity, the request is allowed; if there is no token in the bucket, the data request waits or is rejected.

[0030] As an optimal solution for the third-party cloud service data synchronization processing system based on timing tasks, in the data synchronization module:

[0031] By comparing the data identifiers, only synchronize the data that has changed since the last synchronization; when comparing the data identifiers, map the data identifiers to a hash ring with a fixed range through a hash function. If the positions of the data identifiers on the hash ring change during two synchronizations, or the hash values of the data contents corresponding to the data identifiers are different, it is determined that the data has changed.

[0032] As an optimal solution for the third-party cloud service data synchronization processing system based on timing tasks, in the data synchronization module:

[0033] The conflict detection is to compare the target system and the data to be synchronized during the data synchronization process;

[0034] The conflict resolution strategies include overwriting the target data with the source data as the standard, selecting the latest data according to the timestamp, or performing data merging; when performing data merging, for numerical data, use the weighted average algorithm for merging. Assume the data in the target system is A, the data to be synchronized is B, and the weights are w1 and w2 respectively. The merged data C = w1A + w2B;

[0035] For text data, use the longest common subsequence algorithm to find the longest common subsequence of the two texts, and splice and merge the different parts according to the preset rules.

[0036] The beneficial effects of the present invention are as follows:

[0037] First, optimize the utilization of server resources

[0038] Through the intelligent scheduling of timed tasks and the asynchronous data processing mechanism, the idle resources of the server are fully utilized, avoiding performance issues caused by real-time data processing when the server resources are strained. For example, data conversion and synchronization work are concentrated during periods of low server load, improving the overall utilization rate of server resources. The incremental synchronization strategy reduces the data transmission volume and processing time, further reducing the consumption of server resources and enabling the server to handle more inter-system data interaction tasks with limited resources.

[0039] Second, improve data processing efficiency and accuracy

[0040] Flexible task configuration and intelligent task allocation ensure that data can be processed in a timely and orderly manner, reducing data backlogs and processing delays. At the same time, the multi-format data caching and asynchronous conversion mechanism improve the speed and efficiency of data processing. The conflict detection and resolution mechanism guarantees the consistency and accuracy of data between different systems, avoiding business errors and data chaos caused by data conflicts, and improving the quality and reliability of data.

[0041] Third, enhance the stability and reliability of the system

[0042] The real-time monitoring and exception feedback processing mechanism enables operation and maintenance personnel to promptly discover and solve problems in system operation, improving the maintainability and stability of the system. Even in the event of an abnormal situation, the system can automatically attempt to recover, reducing the risk of data loss and service interruption caused by failures. Description of the Drawings

[0043] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are merely exemplary, and for those of ordinary skill in the art, without creative efforts, other implementation drawings can also be obtained based on the provided drawings.

[0044] The structures, ratios, sizes, etc. depicted in this specification are only used to cooperate with the content disclosed in the specification for those familiar with this technology to understand and read, and are not used to limit the limited conditions under which the present invention can be implemented. Therefore, they do not have a substantial technical meaning. Any modification of the structure, change in the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention.

[0045] Figure 1 It is a schematic flow diagram of the third-party cloud service data synchronization processing method based on timed tasks provided by the embodiments of the present invention;

[0046] Figure 2Data interaction flowchart of the third - party cloud service data synchronization processing method based on scheduled tasks provided by the embodiments of the present invention;

[0047] Figure 3 Schematic diagram of the system architecture of the third - party cloud service data synchronization processing system based on scheduled tasks provided by the embodiments of the present invention. Detailed implementation manners

[0048] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0049] Embodiment 1

[0050] Refer to Figure 1 , Embodiment 1 of the present invention provides a third - party cloud service data synchronization processing method based on scheduled tasks, including the following steps:

[0051] S1. Configure the attributes of the scheduled task through the management interface. The objects of the attribute configuration include the execution time, frequency, third - party cloud service data source, and target system information, and store the objects of the attribute configuration in the attribute database;

[0052] S2. When the scheduled task is triggered, asynchronously obtain data from the third - party cloud service, store it in the local cache, and record the data identifier;

[0053] S3. Perform format conversion, data cleaning, and verification operations on the data in the cache according to preset rules;

[0054] S4. Synchronize the processed data to the target system using the incremental synchronization strategy based on the data identifier, and perform conflict detection and conflict resolution during the synchronization process;

[0055] S5. Monitor the execution status of the scheduled task in real - time. If the execution status of the scheduled task is abnormal, generate a report and try to automatically retry.

[0056] In this embodiment, in step S1, the execution time and frequency of the scheduled task are dynamically adjusted through the management interface. The dynamic adjustment adopts a scheduling model based on the time - series prediction algorithm, and the scheduling model updates the time series T = {t1, t2,... t n}, using the ARIMA(p,d,q) model, the optimal model parameters are determined by minimizing the Akaike Information Criterion or the Bayesian Information Criterion to predict the next data update time of the third-party cloud service, so as to dynamically adjust the execution time and frequency of the timing task; where p is the autoregressive order, d is the differencing order, and q is the moving average order.

[0057] Specifically, in step S1, the attributes of the timing task are configured through the management interface, and the execution time, frequency, third-party cloud service data source, and target system information are stored in the attribute database. This is the basic setting of the entire data synchronization process, providing key parameters for subsequent data acquisition, processing, and synchronization.

[0058] Among them, a scheduling model based on the time series prediction algorithm is used to dynamically adjust the execution time and frequency of the timing task. The scheduling model is based on the historical data update time series of the third-party cloud service, and the ARIMA(p,d,q) model is used to predict the next data update time. The ARIMA model is a commonly used time series prediction model, where p is the autoregressive order, which measures the influence degree of past observations on the current value; d is the differencing order, which is used to make the time series stationary; q is the moving average order, which measures the influence of past prediction errors on the current prediction. The optimal model parameters are determined by minimizing the Akaike Information Criterion (AIC) or the Bayesian Information Criterion (BIC). AIC and BIC are important criteria for model selection, which balance the goodness of fit and complexity of the model, and select the model with the smallest AIC or BIC value as the optimal model, so as to more accurately predict the data update time of the third-party cloud service, and then reasonably and dynamically adjust the execution time and frequency of the timing task to ensure the timeliness of data acquisition and the high efficiency of server resource utilization.

[0059] In this embodiment, in step S2, when the timing task is triggered, during the process of asynchronously obtaining data from the third-party cloud service:

[0060] The asynchronous connection pool technology is used to interact with the third-party cloud service. In the connection pool management, a traffic control mechanism based on the token bucket algorithm is adopted, and the token generation rate r and the bucket capacity b are set. Each time data from the third-party cloud service is requested, an attempt is made to obtain a token. If there are tokens in the bucket that meet the set quantity, the request is allowed; if there are no tokens in the bucket, the data request waits or is rejected.

[0061] Specifically, in step S2, when the timing task is triggered, data is obtained from the third-party cloud service asynchronously. The advantage of asynchronous operation is that it does not block the main thread, allowing the server to process other tasks while waiting for the data to return, improving the utilization rate of server resources and the response speed of the system.

[0062] Among them, the asynchronous connection pool technology is adopted to interact with the third-party cloud service, and a traffic control mechanism based on the token bucket algorithm is used in the connection pool management. The principle of the token bucket algorithm is to set a token generation rate r (tokens / second) and the capacity b (tokens) of the bucket. Each time when requesting data from the third-party cloud service, an attempt is made to obtain a token. If there are tokens in the bucket that meet the set quantity, the request is allowed to access the third-party cloud service to obtain data; if there are no tokens in the bucket, the data request needs to wait until there are tokens available, or it is directly rejected. This mechanism can effectively control the traffic of data acquisition, avoid imposing excessive pressure on the third-party cloud service, ensure the stability and reliability of the data acquisition process, and prevent the third-party cloud service from responding slowly or denying service due to overly frequent requests.

[0063] In this embodiment, in step S3, different third-party cloud services and the target system may adopt different data formats, so format conversion needs to be performed according to preset rules. For example, the data provided by the third-party cloud service may be in XML format, while the target system requires data in JSON format. The preset rules can be predefined conversion scripts or mapping relationships. Through these rules, the system parses and reorganizes the XML data into JSON format to ensure that the data can be correctly used in the target system.

[0064] Among them, the purpose of data cleaning is to remove noise and errors in the data and improve data quality, including removing duplicate data. For example, by comparing the unique identifiers or specific fields of the data, duplicate records are found and deleted; correcting incorrect data, according to business logic and data verification rules, correcting incorrect data formats and data with incorrect value ranges; supplementing missing data, which can be filled with the mean, median, or inferred based on relevant data to supplement the missing values. Data verification is to check whether the processed data meets the requirements and business rules of the target system. For example, checking whether the field types of the data are correct, whether the values of the data are within a reasonable range, and whether the logical relationships between the data are correct, etc. Only the data that passes the verification will enter the next synchronization process to ensure that the data synchronized to the target system is accurate and available.

[0065] In this embodiment, in step S4, during the process of synchronizing the processed data to the target system using the incremental synchronization strategy based on the data identifier:

[0066] By comparing the data identifier, only the data that has changed since the last synchronization is synchronized; when comparing the data identifier, the data identifier is mapped to a hash ring with a fixed range through a hash function. If the position of the data identifier on the hash ring changes during two synchronizations, or the hash values of the data content corresponding to the data identifier are different, it is determined that the data has changed.

[0067] Specifically, in step S4, an incremental synchronization strategy is adopted based on data identifiers, that is, only the data that has changed since the last synchronization is synchronized. By comparing the identifiers of the currently obtained data with the identifiers recorded in the last synchronization, the system can determine which data is newly added, modified, or deleted. This strategy greatly reduces the data transmission volume and synchronization time, and reduces the consumption of system resources. For example, in a system with a large amount of user data, if only a small amount of user information changes every day, incremental synchronization only needs to transmit the changed data instead of all user data.

[0068] During the synchronization process, conflict detection is performed by comparing the existing data in the target system with the data to be synchronized to determine whether there are conflicts, such as duplicate data, data inconsistency, etc. For conflict data, different strategies are used for processing. Overwriting the target data with the source data is applicable to scenarios where the source data has higher credibility or timeliness; selecting the latest data based on the timestamp can ensure the use of the latest valid data; for data merging, different algorithms are used for numerical and text data. For numerical data, a weighted average algorithm is adopted, and the values of the two data are comprehensively calculated according to the set weights; for text data, the longest common subsequence algorithm is adopted, and the same parts are retained and the different parts are spliced according to rules to ensure the integrity and accuracy of the data.

[0069] In this embodiment, in step S4, the conflict detection is to compare the target system and the data to be synchronized during the data synchronization process;

[0070] The strategies for conflict resolution include overwriting the target data with the source data, selecting the latest data according to the timestamp, or performing data merging; when performing data merging, for numerical data, a weighted average algorithm is used for merging. Assume that the data in the target system is A, the data to be synchronized is B, and the weights are w1 and w2 respectively. The merged data C = w1A + w2B;

[0071] For text data, the longest common subsequence algorithm is adopted to find the longest common subsequence of the two texts, and the different parts are spliced and merged according to preset rules.

[0072] Specifically, in step S4, the target system and the data to be synchronized are compared during the data synchronization process to perform conflict detection. Since there may be differences in the update time, update content, etc. between the third-party cloud service and the target system, it is necessary to detect whether there are conflicts, such as duplicate data, data inconsistency, etc. during the synchronization process.

[0073] Among them, for numerical data, a weighted average algorithm is used for merging. Suppose the data in the target system is A, the data to be synchronized is B, and the weights are w1 and w2 respectively. The merged data C = w1A + w2B. This algorithm comprehensively calculates the two data according to the set weights. The setting of the weights can be determined according to factors such as the credibility and importance of the data, so as to obtain a result that comprehensively considers the data of both parties, ensuring the accuracy and rationality of the data.

[0074] For text data, the longest common subsequence algorithm (LCS) is used. This algorithm finds the longest common subsequence of two texts and then splices and merges the different parts according to preset rules. The longest common subsequence reflects the similar parts of the two texts. By retaining this part of the content and reasonably splicing the different parts, it is possible to fuse new content while retaining the key information of the text, ensuring that important information is not lost during the merging process of text data and effectively integrating new changes.

[0075] In this embodiment, in step S5, the execution status of the scheduled task is monitored in real time, and information such as the start time, execution progress, data processing volume, and synchronization status of the task is collected. This information can help the operation and maintenance personnel timely understand the operation status of the system and discover potential problems. For example, if it is found that the execution time of a certain scheduled task is too long or the data processing volume is abnormal, it may mean that there are performance problems or data anomalies in the system, and the operation and maintenance personnel can take timely measures for optimization or repair.

[0076] Among them, if the execution status of the scheduled task is abnormal, such as data acquisition failure, format conversion error, synchronization interruption, etc., the system will generate a detailed report, including error information, occurrence time, related tasks and data, etc. At the same time, the system automatically attempts to perform interval retries. This is to enable the system to automatically resume normal synchronization in the case of some temporary failures, reducing manual intervention and business interruption time. The interval time and number of retries can be configured according to the actual situation. For example, the first retry interval is 1 minute, and if it still fails, the interval time is doubled, and the maximum number of retries is 3 times, to balance the consumption of system resources and the timeliness of task recovery.

[0077] See Figure 2 , which shows the application process of the third-party cloud service data synchronization processing method based on scheduled tasks among different systems, involving four main bodies: third-party services, synchronization systems, internal systems, and other systems. Third-party services provide business data, which is the source of data; the synchronization system is responsible for the configuration and triggering of scheduled tasks, acting as a data transmission bridge; the internal system conducts business processing and stage advancement; other systems receive data, which may be used for further analysis or other business scenarios.

[0078] First, configure the scheduled tasks in the synchronization system, setting attributes such as execution time and frequency. After the configuration is completed, the scheduled tasks are triggered to obtain business data from a third-party service. The obtained data enters the internal system for business processing, and at the same time, other systems can also receive this data. After the internal system completes the business processing, it will generate phased result data. These data may undergo some data operations, and then trigger the scheduled tasks in the synchronization system again to continue transmitting the processed data to the internal system for a new round of business processing. At the same time, other systems can continuously receive data. This process continuously cycles to promote the business stage, ensuring the continuous flow and processing of data among systems. Whether it is the phased result data generated by the third-party service or the data processed by the internal system, they will ultimately be archived for long-term storage and subsequent possible query and analysis to ensure the integrity and traceability of the data. The entire process reflects the process of orderly synchronization and processing of data based on scheduled tasks among different systems. Combining the technical means such as asynchronous acquisition and incremental synchronization mentioned above, it can efficiently and accurately achieve the transfer and application of third-party cloud service data among systems.

[0079] The following gives an example scenario and experimental data from three aspects: data synchronization efficiency, accuracy, and resource consumption to prove the effectiveness of the invention:

[0080] Data Synchronization Efficiency

[0081] Example Scenario 1:

[0082] An e-commerce enterprise uses a third-party cloud service to store product information, including 100,000 pieces of product data. The internal system needs to update the product information in real time for display and sales. The traditional synchronization method is to perform a full synchronization once an hour; the method of the present invention uses scheduled tasks to dynamically adjust the synchronization frequency according to the prediction model, and performs an incremental synchronization once every 2 hours on average.

[0083] The experimental data is shown in Table 1:

[0084] Synchronization method Synchronization period Amount of data synchronized each time Synchronization time consumption Traditional full-volume synchronization 1 hour 100,000 records Approximately 300 seconds Method of the present invention Average 2 hours Approximately 500 records (incremental data) Approximately 10 seconds

[0085] Table 1 Comparison Data of Data Synchronization Efficiency

[0086] As can be seen from Table 1, the method of the present invention has greatly improved in data synchronization efficiency, and the advantage is more obvious as the data volume increases.

[0087] Example Scenario 2:

[0088] A financial institution obtains customer transaction data from a third-party cloud service. During the data synchronization process, data errors are likely to occur due to network fluctuations, data update conflicts, etc. When using the traditional method, conflict handling relies on manual intervention; the method of the present invention automatically processes through a built-in conflict detection and resolution mechanism. Suppose 1000 transaction data are synchronized, and there are 50 cases of data conflicts among them.

[0089] The test data is shown in Table 2:

[0090] Synchronization method Amount of conflicting data Number of data errors Data accuracy rate Traditional method 50 records 10 records 90% Method of the present invention 50 records 1 record 99.9%

[0091] Table 2 Comparison data of data synchronization accuracy

[0092] As can be seen from Table 2, the method of the present invention can effectively improve the accuracy of data synchronization, reduce labor costs and error risks.

[0093] Example scenario three: The data analysis system of an Internet enterprise obtains user behavior data from a third-party cloud service for analysis. The system server is configured with an 8-core CPU and 16GB of memory. The traditional synchronization method continuously occupies resources for real-time synchronization; the method of the present invention uses an asynchronous connection pool and traffic control, and obtains resources as needed when a timed task is triggered. Within a day, the third-party cloud service generates 10GB of user behavior data.

[0094] The test data is shown in Table 3:

[0095]

[0096] Table 3 Comparison data of resource consumption

[0097] The test data in Table 3 shows that the method of the present invention can significantly reduce resource consumption, ensure the stable operation of the system, and does not affect the development of other services.

[0098] Embodiment 2

[0099] See Figure 3 , Embodiment 2 of the present invention further provides a third-party cloud service data synchronization processing system based on a timed task, including:

[0100] A task configuration module 001, used to configure the attributes of the timed task through a management interface, where the objects of the attribute configuration include execution time, frequency, third-party cloud service data source, and target system information, and store the objects of the attribute configuration in an attribute database;

[0101] A data acquisition module 002, used to asynchronously acquire data from a third-party cloud service when the timed task is triggered, store it in a local cache, and record a data identifier;

[0102] The data processing module 003 is used to perform format conversion, data cleaning, and verification operations on the data in the cache according to preset rules;

[0103] The data synchronization module 004 is used to synchronize the processed data to the target system using an incremental synchronization strategy based on the data identifier, and perform conflict detection and conflict resolution during the synchronization process;

[0104] The monitoring and feedback module 005 is used to monitor the execution status of the timing task in real time. If the execution status of the timing task is abnormal, a report is generated and an automatic retry is attempted.

[0105] In this embodiment, in the task configuration module 001:

[0106] The execution time and frequency of the timing task are dynamically adjusted through the management interface. The dynamic adjustment adopts a scheduling model based on a time series prediction algorithm. The scheduling model updates the time series T = {t1, t2,... t n} according to the historical data of the third-party cloud service, uses the ARIMA(p, d, q) model, and determines the optimal model parameters by minimizing the Akaike information criterion or the Bayesian information criterion to predict the next data update time of the third-party cloud service, so as to dynamically adjust the execution time and frequency of the timing task; where p is the autoregressive order, d is the differencing order, and q is the moving average order.

[0107] In this embodiment, in the data acquisition module 002:

[0108] Asynchronous connection pool technology is used to interact with the third-party cloud service. In connection pool management, a traffic control mechanism based on the token bucket algorithm is adopted. The token generation rate r and the bucket capacity b are set. Each time data from the third-party cloud service is requested, an attempt is made to obtain a token. If there are enough tokens in the bucket, the request is allowed; if there are no tokens in the bucket, the data request waits or is rejected.

[0109] In this embodiment, in the data synchronization module 004:

[0110] By comparing the data identifiers, only the data that has changed since the last synchronization is synchronized; when comparing the data identifiers, the data identifiers are mapped to a hash ring with a fixed range through a hash function. If the positions of the data identifiers on the hash ring change during two synchronizations, or the hash values of the data contents corresponding to the data identifiers are different, it is determined that the data has changed.

[0111] In this embodiment, in the data synchronization module 004:

[0112] The conflict detection is to compare the target system and the data to be synchronized during the data synchronization process;

[0113] The conflict resolution strategies include overwriting the target data with the source data as the standard, selecting the latest data according to the timestamp, or performing data merging; when performing data merging, for numerical data, a weighted average algorithm is used for merging. Assuming the data in the target system is A, and the data to be synchronized is B, with weights w1 and w2 respectively, the merged data C = w1A + w2B;

[0114] For text data, the longest common subsequence algorithm is used to find the longest common subsequence of the two texts, and the different parts are spliced and merged according to preset rules.

[0115] It should be noted that the information interaction, execution process, etc. between the above-mentioned system modules, due to being based on the same concept as the method embodiment in Embodiment 1 of the present application, have the same technical effects as the method embodiment of the present application. For specific content, reference can be made to the description in the method embodiment shown above in the present application, and details will not be elaborated here.

[0116] Embodiment 3

[0117] Embodiment 3 of the present invention provides a non-transitory computer-readable storage medium, in which program codes for a third-party cloud service data synchronization processing method based on a timing task are stored, and the program codes include instructions for executing the third-party cloud service data synchronization processing method based on a timing task in Embodiment 1 or any possible implementation manner thereof.

[0118] The computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state drive (SSD)).

[0119] Embodiment 4

[0120] Embodiment 4 of the present invention provides an electronic device, including: a memory and a processor;

[0121] The processor and the memory communicate with each other through a bus; the memory stores program instructions executable by the processor, and the processor can execute the third-party cloud service data synchronization processing method based on a timing task in Embodiment 1 or any possible implementation manner thereof by calling the program instructions.

[0122] Specifically, the processor can be implemented by hardware or software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor can be a general-purpose processor that realizes its functions by reading software code stored in a memory. The memory can be integrated in the processor or exist independently outside the processor.

[0123] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means.

[0124] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program code executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the present invention is not limited to any specific combination of hardware and software.

[0125] Although the present invention has been described in detail above with general descriptions and specific embodiments, based on the present invention, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of protection required by the present invention.

Claims

1. A method for processing third-party cloud service data synchronization based on scheduled tasks, characterized in that Including: Configuring the attributes of the scheduled task through the management interface. The objects of the attribute configuration include the execution time, frequency, third-party cloud service data source, and target system information, and storing the objects of the attribute configuration in the attribute database; When the scheduled task is triggered, asynchronously obtaining data from the third-party cloud service, storing it in the local cache, and recording the data identifier; Performing format conversion, data cleaning, and verification operations on the data in the cache according to preset rules; Synchronizing the processed data to the target system using an incremental synchronization strategy based on the data identifier, and performing conflict detection and conflict resolution during the synchronization process; Real-time monitoring the execution status of the scheduled task. If the execution status of the scheduled task is abnormal, generating a report and attempting to automatically retry.

2. The method for processing third-party cloud service data synchronization based on scheduled tasks according to claim 1, wherein Dynamically adjust the execution time and frequency of the scheduled task through the management interface. The dynamic adjustment adopts a scheduling model based on a time series prediction algorithm. The scheduling model updates the time series T = {t1, t2, … t n} according to the historical data of the third-party cloud service, uses the ARIMA(p, d, q) model, determines the optimal model parameters by minimizing the Akaike information criterion or the Bayesian information criterion, predicts the next data update time of the third-party cloud service, and dynamically adjusts the execution time and frequency of the scheduled task; where p is the autoregressive order, d is the differencing order, and q is the moving average order.

3. The method for processing third-party cloud service data synchronization based on timed tasks according to claim 1, wherein, When the scheduled task is triggered, during the process of asynchronously obtaining data from the third-party cloud service: Using an asynchronous connection pool technology to interact with the third-party cloud service. In the connection pool management, adopting a traffic control mechanism based on the token bucket algorithm, setting the token generation rate r and the capacity b of the bucket. Each time requesting data from the third-party cloud service, attempting to obtain a token. If there are tokens in the bucket that meet the set quantity, the request is allowed; if there are no tokens in the bucket, the data request waits or is rejected.

4. The method for processing third-party cloud service data synchronization based on a scheduled task according to claim 1, wherein During the process of synchronizing the processed data to the target system using an incremental synchronization strategy based on the data identifier: By comparing the data identifiers, only synchronizing the data that has changed since the last synchronization; when comparing the data identifiers, mapping the data identifiers to a hash ring with a fixed range through a hash function. If the positions of the data identifiers on the hash ring change during two synchronizations, or the hash values of the data contents corresponding to the data identifiers are different, it is determined that the data has changed.

5. The method for processing third-party cloud service data synchronization based on scheduled tasks according to claim 1, wherein, The conflict detection is to compare the target system and the data to be synchronized during the data synchronization process; The conflict resolution strategies include overwriting the target data with the source data, selecting the latest data according to the timestamp, or performing data merging; when performing data merging, for numerical data, using a weighted average algorithm for merging. Assuming the data in the target system is A, the data to be synchronized is B, and the weights are w1 and w2 respectively, the merged data C = w1A + w2B; For text data, using the longest common subsequence algorithm to find the longest common subsequence of the two texts, and splicing and merging the different parts according to preset rules.

6. A third-party cloud service data synchronization processing system based on scheduled tasks, characterized in that, Including: A task configuration module for configuring the attributes of the scheduled task through the management interface. The objects of the attribute configuration include the execution time, frequency, third-party cloud service data source, and target system information, and storing the objects of the attribute configuration in the attribute database; A data acquisition module for asynchronously obtaining data from the third-party cloud service when the scheduled task is triggered, storing it in the local cache, and recording the data identifier; A data processing module for performing format conversion, data cleaning, and verification operations on the data in the cache according to preset rules; A data synchronization module for synchronizing the processed data to the target system using an incremental synchronization strategy based on the data identifier, and performing conflict detection and conflict resolution during the synchronization process; The monitoring feedback module is used to monitor the execution status of the timing task in real time. If the execution status of the timing task is abnormal, a report is generated and an automatic retry is attempted.

7. The third-party cloud service data synchronization processing system based on a timed task according to claim 6, wherein In the task configuration module: Dynamically adjust the execution time and frequency of the scheduled task through the management interface. The dynamic adjustment adopts a scheduling model based on a time series prediction algorithm. The scheduling model updates the time series T = {t1, t2, … t n} according to the historical data of the third-party cloud service, uses the ARIMA(p, d, q) model, determines the optimal model parameters by minimizing the Akaike information criterion or the Bayesian information criterion, predicts the next data update time of the third-party cloud service, and dynamically adjusts the execution time and frequency of the scheduled task; where p is the autoregressive order, d is the difference order, and q is the moving average order.

8. The third-party cloud service data synchronization processing system based on scheduled tasks according to claim 6, wherein In the data acquisition module: Asynchronous connection pool technology is used to interact with the third-party cloud service. In the connection pool management, a traffic control mechanism based on the token bucket algorithm is adopted. The token generation rate r and the bucket capacity b are set. Each time data from the third-party cloud service is requested, an attempt is made to obtain a token. If there are tokens in the bucket that meet the set quantity, the request is allowed; if there are no tokens in the bucket, the data request waits or is rejected.

9. The third-party cloud service data synchronization processing system based on a timed task according to claim 6, wherein, In the data synchronization module: By comparing the data identifiers, only the data that has changed since the last synchronization is synchronized; when comparing the data identifiers, the data identifiers are mapped to a hash ring with a fixed range through a hash function. If the positions of the data identifiers on the hash ring change during two synchronizations, or the hash values of the data contents corresponding to the data identifiers are different, it is determined that the data has changed.

10. The third-party cloud service data synchronization processing system based on a scheduled task according to claim 6, wherein In the data synchronization module: The conflict detection is to compare the target system and the data to be synchronized during the data synchronization process; The conflict resolution strategies include overwriting the target data with the source data as the standard, selecting the latest data according to the timestamp, or performing data merging; when performing data merging, for numerical data, a weighted average algorithm is used for merging. Assume that the data in the target system is A, the data to be synchronized is B, and the weights are w1 and w2 respectively. The merged data C = w1A + w2B; For text data, the longest common subsequence algorithm is used to find the longest common subsequence of the two texts, and the different parts are spliced and merged according to the preset rules.

Citation Information

Cited By

  • Secure storage synchronization method for cloud edge fusion

    CN121644576A