Data synchronization processing method and system based on data synchronization model

Through data synchronization model analysis and neural network optimization, the data mapping speed and mapping coefficient are dynamically adjusted, which solves the problem that data synchronization systems in the existing technology cannot adapt to real-time loads, and achieves efficient and accurate data synchronization.

CN119988504AInactive Publication Date: 2025-05-13BEIJING LIUJINSUIYUE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510473063.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing data synchronization systems cannot dynamically adjust the mapping speed based on real-time data volume and system load, resulting in backlog of synchronization tasks or failures.

Method used

By performing mapping analysis based on the data synchronization model, the synchronization data link of the data to be synchronized is determined, the characteristics of the header and adjacent data to be synchronized are analyzed, the data correlation is judged, the data mapping speed is dynamically adjusted, the data format is optimized using the neural network model, and the mapping coefficient is adjusted according to the load relationship of the data synchronized in the content and structure.

Benefits of technology

Improve the accuracy and stability of data synchronization, reduce data loss and error mapping, avoid resource waste and system overload, and improve overall synchronization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988504A_ABST
    Figure CN119988504A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses a data synchronization processing method and system based on a data synchronization model, and the method comprises the steps: carrying out the mapping analysis of all to-be-synchronized data based on the data synchronization model, extracting header to-be-synchronized data on the synchronous data chain and adjacent to-be-synchronized data adjacent to the header to-be-synchronized data, judging whether the header to-be-synchronized data and the adjacent to-be-synchronized data are associated or not based on the synchronous association factor, determining all content synchronous data and all structure synchronous data, and according to a data mapping loss value, performing synchronous synchronization on the header to-be-synchronized data and the adjacent to-be-synchronized data. And determining the data mapping speed of the source end and the target end, determining a load relation based on the content synchronization data and the structure synchronization data, obtaining a mapping coefficient of the data mapping speed, and adjusting the data mapping speed according to the mapping coefficient. According to the method, the data mapping speed is determined and adjusted through the relevance and the load relation between the data, and the reliability of data processing is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a data synchronization processing method and system based on a data synchronization model. Background Art

[0002] In the digital age, data in business links such as media asset management and content production is growing, and the demand for data synchronization is becoming more frequent. Data synchronization aims to ensure data consistency between the source and target ends, and provide support for the stable operation of the business. In the process of data synchronization, mapping synchronization is responsible for accurately mapping the source data to the target end according to established rules. However, the existing data synchronization system has the problem of being unable to adjust the mapping speed during the mapping synchronization process. Due to the complex and changeable business scenarios, the data volume and data processing requirements of different time periods and different businesses are different. In the daily stage, the data volume is relatively stable. In the operation stage, the data synchronization system cannot dynamically adjust the mapping speed according to the real-time data volume and system load, which will lead to a backlog of synchronization tasks or even failure.

[0003] Therefore, it is necessary to design a data synchronization processing method and system based on a data synchronization model to solve the problems existing in the current technology. Summary of the invention

[0004] In view of this, the present invention proposes a data synchronization processing method and system based on a data synchronization model, aiming to solve the problem that the data synchronization system cannot dynamically adjust the mapping speed according to the real-time data volume and system load, which will lead to synchronization task backlog or even failure.

[0005] In one aspect, a data synchronization processing method based on a data synchronization model includes:

[0006] Obtain all the data to be synchronized on the source side, perform mapping analysis on all the data to be synchronized based on the data synchronization model, and determine the synchronization data chain of the data to be synchronized;

[0007] Extracting the header data to be synchronized and the adjacent data to be synchronized adjacent to the header data to be synchronized on the synchronization data chain, determining a synchronization association factor according to data features of the header data to be synchronized and the adjacent data to be synchronized, judging whether the header data to be synchronized and the adjacent data to be synchronized are associated based on the synchronization association factor, judging whether the remaining data to be synchronized on the synchronization data chain are associated with the header data to be synchronized, and determining whether to merge the data chain and split the data chain according to the judgment result;

[0008] Analyze the merged data chain and the split data chain to determine all content synchronization data and all structure synchronization data, determine the data mapping loss value of the source end according to all content synchronization data and all structure synchronization data, and determine the data mapping speed of the source end and the target end according to the data mapping loss value;

[0009] Based on the load relationship determined by the content synchronization data and the structure synchronization data, a mapping coefficient of the data mapping speed is obtained, and the data mapping speed is adjusted according to the mapping coefficient.

[0010] Furthermore, before obtaining all the data to be synchronized from the source, the following steps are performed:

[0011] Acquire initial data, perform data cleaning on the initial data, wherein the data cleaning includes processing missing values, outliers and duplicate values, and perform data conversion on the cleaned initial data, wherein the data conversion includes data standardization and data normalization;

[0012] Comparing the data format of the initial data after data conversion with the mapping format of the target end, and determining the data to be synchronized according to the comparison result;

[0013] When the data format conforms to the mapping format of the target end, the initial data after the data format corresponding to the data is converted is used as the data to be synchronized;

[0014] When the data format does not conform to the mapping format of the target end, the initial data after the data format is converted to the corresponding data is deleted.

[0015] Furthermore, when all the data to be synchronized are mapped and analyzed based on the data synchronization model to determine the synchronization data chain of the data to be synchronized, it includes:

[0016] Obtain a historical data set to be synchronized, and divide the historical data set to be synchronized into a training set and a test set, use grid search to find hyperparameters of a neural network model, and establish a neural network model;

[0017] The training set is used to fit the neural network model, the test set is substituted into the neural network model and the neural network model is evaluated, when the evaluation value reaches a preset evaluation threshold, the neural network model is used as the data synchronization model, and all the data to be synchronized are substituted into the data synchronization model, and the data synchronization value of each data to be synchronized is output;

[0018] All different data synchronization values ​​are used to construct a sequence to be synchronized, and all identical data synchronization values ​​are used to construct several data synchronization value sequences, and one data synchronization value is obtained from each of the data synchronization value sequences, and the remaining data synchronization values ​​are removed;

[0019] The acquired data synchronization value is added to the to-be-synchronized sequence, and the sequences are arranged in descending order of the data synchronization value to determine the synchronization data chain.

[0020] Further, when determining a synchronization association factor according to data features of the header data to be synchronized and the adjacent data to be synchronized, and judging whether the header data to be synchronized and the adjacent data to be synchronized are associated based on the synchronization association factor, the method includes:

[0021] Analyze the data features of the header data to be synchronized, determine the header data value corresponding to each data feature and determine the header data mean;

[0022] Presetting a first synchronization set, a second synchronization set, and a third synchronization set;

[0023] The header data values ​​greater than the header data mean value are divided into the first synchronization set, the header data values ​​equal to the header data mean value are divided into the second synchronization set, and the header data values ​​less than the header data mean value are divided into the third synchronization set;

[0024] Analyze the data features of the adjacent data to be synchronized, determine the adjacent data value corresponding to each data feature and determine the adjacent data mean;

[0025] The adjacent data values ​​greater than the adjacent data mean are divided into the first synchronization set, the adjacent data values ​​equal to the adjacent data mean are divided into the second synchronization set, and the adjacent data values ​​less than the adjacent data mean are divided into the third synchronization set;

[0026] Counting the first synchronization number of all header data values ​​and all adjacent data values ​​in the first synchronization set, counting the second synchronization number of all header data values ​​and all adjacent data values ​​in the second synchronization set, and counting the third synchronization number of all header data values ​​and all adjacent data values ​​in the third synchronization set;

[0027] The synchronization association factor is determined according to the first synchronization number, the second synchronization number and the third synchronization number, and the synchronization association factor is compared with a preset synchronization association factor, and whether there is an association between the header data to be synchronized and the adjacent data to be synchronized is determined according to the comparison result.

[0028] Further, when determining the synchronization association factor according to the first synchronization number, the second synchronization number, and the third synchronization number, and comparing the synchronization association factor with a preset synchronization association factor, and determining whether the header data to be synchronized and the adjacent data to be synchronized are associated according to the comparison result, it includes:

[0029]

[0030] Wherein, S represents the synchronization correlation factor, A1 represents the first synchronization number, A2 represents the second synchronization number, A3 represents the third synchronization number, max(A1, A3) represents the maximum value between the first synchronization number and the third synchronization number, and max(A1, A2) represents the maximum value between the first synchronization number and the second synchronization number;

[0031] When the synchronization association factor is greater than or equal to the preset synchronization association factor, it is determined that the header data to be synchronized and the adjacent data to be synchronized are associated;

[0032] When the synchronization association factor is less than the preset synchronization association factor, it is determined that there is no association between the head data to be synchronized and the adjacent data to be synchronized.

[0033] Further, when determining the data mapping loss value of the source end according to all content synchronization data and all structure synchronization data, and determining the data mapping speeds of the source end and the target end according to the data mapping loss value, it includes:

[0034] Substituting all content synchronization data and all structure synchronization data into a pre-trained data mapping model, and determining a content mapping value for each content synchronization data and a structure mapping value for each structure synchronization data;

[0035] Determine a content mapping and value of all content mapping values, and take a natural logarithm of the content mapping and value to determine a first mapping value, determine a structure mapping and value of the structure mapping value, and take a natural logarithm of the structure mapping and value to determine a second mapping value;

[0036] The data mapping loss value is a product value of the first mapping value and the second mapping value;

[0037] A preset first data mapping loss value L1 and a preset second data mapping loss value L2 are preset in advance, and L1<L2;

[0038] A preset first mapping speed V1, a preset second mapping speed V2 and a preset third mapping speed V3 are preset, and V1<V2<V3;

[0039] When the data mapping loss value is less than the preset first data mapping loss value, using the preset third mapping speed as the data mapping speed;

[0040] When the data mapping loss value is greater than or equal to the preset first data mapping loss value and less than the preset second data mapping loss value, using the preset second mapping speed as the data mapping speed;

[0041] When the data mapping loss value is greater than or equal to the preset second data mapping loss value, the preset first mapping speed is used as the data mapping speed.

[0042] Further, when obtaining the mapping coefficient of the data mapping speed based on the load relationship determined by the content synchronization data and the structure synchronization data, it includes:

[0043] Determine the application load of the source end and the target end within a preset time period according to the content synchronization data and the structure synchronization data;

[0044] Determining an application load diagram according to the application load, and determining a minimum load point on the application load diagram;

[0045] Determining mapping loads of the source end and the target end within the preset time period according to the content synchronization data and the structure synchronization data;

[0046] Calculating a mapping difference between the minimum load point and the mapping load, and when the mapping difference is greater than or equal to zero, determining a mapping coefficient of the data mapping speed to be 1;

[0047] When the mapping difference is less than zero, a mapping coefficient of the data mapping speed is determined according to the application load map.

[0048] Further, when determining the mapping coefficient of the data mapping speed according to the application load diagram, it includes:

[0049] Extracting all load slopes in the applied load graph, and determining the variance of all load slopes and the standard deviation of all load slopes;

[0050] The ratio of the standard deviation to the variance is used as a mapping coefficient of the data mapping speed.

[0051] Further, when the data mapping speed is adjusted according to the mapping coefficient, it includes:

[0052] The data mapping speed is proportional to the mapping coefficient.

[0053] Compared with the prior art, the beneficial effects of the present invention are: mapping analysis is performed according to the data synchronization model to determine the synchronization data chain of the data to be synchronized, and the data characteristics of the header data to be synchronized and the adjacent data to be synchronized are analyzed to determine the synchronization correlation factor, and then the correlation between the data is judged, and the internal connection between the data can be accurately identified, and the data with correlation is integrated together for synchronization, thereby improving the accuracy of data synchronization, dynamically determining the data mapping speed according to the data mapping loss value, reducing the loss or wrong mapping of data during the synchronization process, ensuring that the data at the source and target ends are highly consistent, avoiding business logic errors caused by improper data mapping, determining the mapping coefficient according to the load relationship between the content synchronization data and the structure synchronization data, and adjusting the data mapping speed accordingly. The stability of data synchronization is guaranteed, thereby effectively utilizing resources, avoiding resource waste or system overload caused by mismatch between mapping speed and load requirements, and thus improving the overall synchronization efficiency.

[0054] On the other hand, the present application also provides a data synchronization processing system based on a data synchronization model, which is applied to the above-mentioned data synchronization processing method based on a data synchronization model, including:

[0055] The acquisition module is configured to obtain all the data to be synchronized at the source end, perform mapping analysis on all the data to be synchronized based on the data synchronization model, and determine the synchronization data chain of the data to be synchronized;

[0056] A judgment module is configured to extract the header data to be synchronized and the adjacent data to be synchronized adjacent to the header data to be synchronized on the synchronization data chain, determine a synchronization association factor according to data features of the header data to be synchronized and the adjacent data to be synchronized, judge whether the header data to be synchronized and the adjacent data to be synchronized are associated based on the synchronization association factor, judge whether the remaining data to be synchronized on the synchronization data chain are associated with the header data to be synchronized, and determine whether to merge the data chain and split the data chain according to the judgment result;

[0057] a processing module configured to analyze the merged data chain and the split data chain, determine all content synchronization data and all structure synchronization data, determine a data mapping loss value of the source end according to the content synchronization data and all structure synchronization data, and determine data mapping speeds of the source end and the target end according to the data mapping loss value;

[0058] The mapping module is configured to obtain a mapping coefficient of the data mapping speed based on a load relationship determined by the content synchronization data and the structure synchronization data, and adjust the data mapping speed according to the mapping coefficient.

[0059] It is understandable that the above-mentioned data synchronization processing method and system based on the data synchronization model have the same beneficial effects, which will not be described in detail here. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:

[0061] Figure 1 A flowchart of a data synchronization processing method based on a data synchronization model provided by an embodiment of the present invention;

[0062] Figure 2 A functional block diagram of a data synchronization processing system based on a data synchronization model provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0063] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art. It should be noted that, in the absence of conflict, the embodiments of the present invention and the features described in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0064] In some embodiments of the present application, see Figure 1 As shown, a data synchronization processing method based on a data synchronization model includes:

[0065] S100: Acquire all data to be synchronized from the source end, perform mapping analysis on all data to be synchronized based on the data synchronization model, and determine the synchronization data chain of the data to be synchronized.

[0066] S200: extracting the header data to be synchronized and the adjacent data to be synchronized that are adjacent to the header data to be synchronized on the synchronization data chain, determining the synchronization association factor according to the data characteristics of the header data to be synchronized and the adjacent data to be synchronized, judging whether the header data to be synchronized and the adjacent data to be synchronized are associated based on the synchronization association factor, judging whether the remaining data to be synchronized on the synchronization data chain are associated with the header data to be synchronized, and determining whether to merge the data chain and split the data chain according to the judgment result;

[0067] S300: Analyze the merged data chain and the split data chain to determine all content synchronization data and all structure synchronization data, determine the data mapping loss value of the source end according to the content synchronization data and all structure synchronization data, and determine the data mapping speed of the source end and the target end according to the data mapping loss value;

[0068] S400: Obtaining a mapping coefficient of a data mapping speed based on a load relationship determined by content synchronization data and structure synchronization data, and adjusting the data mapping speed according to the mapping coefficient.

[0069] Specifically, all the data to be synchronized are obtained from the source. These data to be synchronized have different types and attributes. The data synchronization model is used to map and analyze these data to be synchronized. The purpose of the mapping analysis is to find the priority between the data, thereby avoiding the mapping work and organizing them into an orderly synchronization data chain. This synchronization data chain can be regarded as a logical order of data in the synchronization process, which provides a basis for subsequent data processing. After the synchronization data chain is determined, the first data to be synchronized and the adjacent data to be synchronized are extracted from the chain. By analyzing the data characteristics of the two data, such as data type, data generation time, etc., the synchronization correlation factor is calculated. The synchronization correlation factor reflects the degree of correlation between the two data. According to the synchronization correlation factor, it can be judged whether the first data to be synchronized and the adjacent data to be synchronized are associated. Then, the same judgment is made on the remaining data to be synchronized on the synchronization data chain, so that the associated data to be synchronized are divided into a merged data chain, and the remaining data to be synchronized are divided into a split data chain.

[0070] Specifically, the data in the merged data chain has a strong correlation, while the data in the split data chain has a weak correlation. By mapping and analyzing the data and judging the correlation, the data to be synchronized is divided into the merged data chain and the split data chain, and the mapping synchronization can be performed finely. For the data to be synchronized with strong correlation, it is mapped and synchronized as a whole, which reduces the error and loss of data during the synchronization process. For example: in a data set of video image data and character details, there is a strong correlation between the video image data and the character details. Mapping and synchronizing them as a whole can ensure the consistency of the data. The merged data chain and the split data chain are deeply analyzed to determine all the content synchronization data and structure synchronization data. Content synchronization data refers to data on specific content, while structure synchronization data involves the organization and architecture of data. In the process of mapping synchronization, it may be inaccessible, damaged or lost due to some reasons. By analyzing these data to determine the data mapping loss value, the process of mapping synchronization is quantified. According to the data mapping loss value, the data mapping speed of the source and target ends can be determined to ensure that the data can be synchronized efficiently and accurately. Based on the load relationship between content synchronization data and structure synchronization data, the mapping coefficient of data mapping speed is obtained. The mapping coefficient reflects the complexity of the data and the resources required for mapping synchronization. The determined data mapping speed is adjusted according to the mapping coefficient to ensure the mapping efficiency between the source and target ends and improve the reliability and consistency of synchronization.

[0071] It can be understood that whether it is a small data set at the source (less data to be synchronized) or a large data set (more data to be synchronized), it can be divided into appropriate merged data chains and split data chains through mapping analysis and association judgment, thereby achieving efficient and accurate data synchronization.

[0072] In some embodiments of the present application, before obtaining all the data to be synchronized from the source end, it includes: obtaining initial data, performing data cleaning on the initial data, data cleaning includes processing missing values, abnormal values ​​and duplicate values, and performing data conversion on the cleaned initial data, data conversion includes data standardization and data normalization, comparing the data format of the initial data after data conversion with the mapping format of the target end, determining the data to be synchronized based on the comparison result, when the data format conforms to the mapping format of the target end, the initial data after the data format corresponding to the data conversion is used as the data to be synchronized, and when the data format does not conform to the mapping format of the target end, the initial data after the data format corresponding to the data conversion is deleted.

[0073] Specifically, through data cleaning (processing missing values, outliers, and duplicate values), noise and errors in the original data can be filtered out to ensure the data integrity and accuracy of the data to be synchronized at the source end. For example, in the media asset management scenario, the cleaned video metadata (such as resolution and duration) can avoid data errors caused by missing fields. Data standardization and normalization processing (such as unified timestamp format and value range compression) eliminates mapping conflicts caused by data format differences between the source and target ends. For example, converting the "YYYY / MM / DD" date format of the source end to the "YYYY-MM-DD" required by the target end can avoid synchronization failures caused by format incompatibility. Then, the data format of the initial data after data conversion is compared with the mapping format of the target end, including field length, data type, range limit, etc., to verify whether it complies with the mapping format of the target end, ensure that no damaged or inconsistent data is introduced, and further improve the data integrity and accuracy of the data to be synchronized.

[0074] In some embodiments of the present application, when mapping and analyzing all data to be synchronized based on a data synchronization model to determine the synchronization data chain of the data to be synchronized, it includes: obtaining a historical data set to be synchronized, and dividing the historical data set to be synchronized into a training set and a test set, using grid search to find hyperparameters of the neural network model, establishing a neural network model, using the training set to fit the neural network model, substituting the test set into the neural network model and evaluating the neural network model, when the evaluation value reaches a preset evaluation threshold, the neural network model is used as the data synchronization model, and all data to be synchronized are substituted into the data synchronization model, the data synchronization value of each data to be synchronized is output, all different data synchronization values ​​are constructed into a sequence to be synchronized, and all the same data synchronization values ​​are constructed into several data synchronization value sequences, one data synchronization value is obtained from each data synchronization value sequence, and the remaining data synchronization values ​​are removed, the obtained data synchronization values ​​are added to the sequence to be synchronized, and the data synchronization values ​​are arranged in descending order to determine the synchronization data chain.

[0075] Specifically, the historical data set to be synchronized contains data samples of various data to be synchronized in different periods, as well as the data synchronization values ​​corresponding to the data samples. The historical data set to be synchronized is divided into a training set and a test set. Usually 70%-80% of the data is used as the training set, and the rest is used as the test set. Ensure that both the training set and the test set contain data samples of various data to be synchronized and the corresponding data synchronization values ​​to improve the generalization ability of the model. The grid search establishes a neural network model by exhaustively searching the hyperparameters of the neural network model in the parameter space. The training set is used to fit the neural network model, while the test set is used to evaluate the model performance. The neural network model contains multiple levels, different types of neurons and activation functions, and is designed to capture complex relationships in the data. When fitting the neural network model with the training set, the model will try to learn the patterns and relationships in the data to improve its prediction or classification capabilities. The data in the test set is used to evaluate the model. The evaluation indicators include loss function value, recall rate, etc., which are used to measure the performance of the model, which helps the model to stably approach the global optimal solution so that it can accurately output data synchronization values.

[0076] It can be understood that the preset evaluation threshold is preferably 0.8. Through continuous training and verification of model performance, when the evaluation value reaches the preset evaluation threshold, it indicates that the neural network model at this time can accurately output the data synchronization value. It is then used as a data synchronization model. The data synchronization model outputs the data synchronization value for the synchronized data, and builds a synchronized data chain. It can remove duplicate data to be synchronized and avoid repetitive mapping. At the same time, it arranges the data in descending order based on the data synchronization value to ensure the priority of mapping processing.

[0077] In some embodiments of the present application, when determining a synchronization association factor based on data features of the header data to be synchronized and the adjacent data to be synchronized, and judging whether there is an association between the header data to be synchronized and the adjacent data to be synchronized based on the synchronization association factor, the method includes: analyzing the data features of the header data to be synchronized, determining the header data value corresponding to each data feature and determining the mean of the header data, pre-setting a first synchronization set, a second synchronization set, and a third synchronization set, dividing the header data values ​​greater than the mean of the header data into the first synchronization set, dividing the header data values ​​equal to the mean of the header data into the second synchronization set, and dividing the header data values ​​less than the mean of the header data into the third synchronization set, analyzing the data features of the adjacent data to be synchronized, determining the adjacent data values ​​corresponding to each data feature and determining Determine the mean of adjacent data, divide adjacent data values ​​greater than the mean of adjacent data into a first synchronization set, divide adjacent data values ​​equal to the mean of adjacent data into a second synchronization set, divide adjacent data values ​​less than the mean of adjacent data into a third synchronization set, count the first synchronization number of all header data values ​​and all adjacent data values ​​in the first synchronization set, count the second synchronization number of all header data values ​​and all adjacent data values ​​in the second synchronization set, count the third synchronization number of all header data values ​​and all adjacent data values ​​in the third synchronization set, determine a synchronization association factor according to the first synchronization number, the second synchronization number, and the third synchronization number, compare the synchronization association factor with a preset synchronization association factor, and determine whether there is an association between the header data to be synchronized and the adjacent data to be synchronized according to the comparison result.

[0078] Specifically, when analyzing the data features of the header data to be synchronized and the data features of the adjacent data to be synchronized, the data features include data type, data generation time, data attributes and data size, etc. The header data value and the corresponding header data mean, as well as the adjacent data value and the corresponding adjacent data mean are determined according to the data features. The header data mean and the adjacent data value are used as the basis for division, and the header data value and the adjacent data value are divided into the first synchronization set, the second synchronization set and the third synchronization set. If the distribution patterns of the header data value and the adjacent data value in each synchronization set are similar, it means that they have similar performance. For example: the number of header data values ​​and adjacent data values ​​in the first synchronization set is relatively large, the number in the second synchronization set is similar, and the number in the third synchronization set is also similar, which means that their data features have a high correlation or the same internal law as a whole. Determining the synchronization correlation factor according to the first synchronization number, the second synchronization number and the third synchronization number lays the foundation for determining whether there is a correlation between the header data to be synchronized and the adjacent data to be synchronized, and provides data support.

[0079] In some embodiments of the present application, when determining a synchronization association factor according to the first synchronization number, the second synchronization number, and the third synchronization number, and comparing the synchronization association factor with a preset synchronization association factor, and determining whether there is an association between the header data to be synchronized and the adjacent data to be synchronized according to the comparison result, it includes:

[0080]

[0081] Among them, S represents the synchronization association factor, A1 is the first synchronization number, A2 is the second synchronization number, A3 is the third synchronization number, max(A1, A3) represents selecting the maximum value between the first synchronization number and the third synchronization number, max(A1, A2) represents selecting the maximum value between the first synchronization number and the second synchronization number, when the synchronization association factor is greater than or equal to the preset synchronization association factor, it is judged that the header data to be synchronized and the adjacent data to be synchronized are associated, and when the synchronization association factor is less than the preset synchronization association factor, it is judged that the header data to be synchronized and the adjacent data to be synchronized are not associated.

[0082] Specifically, the synchronization correlation factor is preset to 0.6. By comparing the synchronization correlation factor with the preset synchronization correlation factor, it is determined whether the header data to be synchronized and the adjacent data to be synchronized are associated, which ensures the efficiency of mapping synchronization. When determining whether the remaining data to be synchronized on the synchronization data chain is associated with the header data to be synchronized, and determining whether to merge the data chain and split the data chain according to the judgment result, the process of whether the remaining data to be synchronized is associated with the header data to be synchronized is consistent with the process of determining whether the header data to be synchronized and the adjacent data to be synchronized are associated, and will not be repeated here. All the data to be synchronized associated with the header data to be synchronized and the header data to be synchronized are constructed as a merged data chain, and the remaining data to be synchronized are used as split data chains, which improves the reliability and stability of the synchronization process.

[0083] In some embodiments of the present application, when determining the data mapping loss value of the source end according to all content synchronization data and all structure synchronization data, and determining the data mapping speed of the source end and the target end according to the data mapping loss value, it includes: substituting all content synchronization data and all structure synchronization data into a pre-trained data mapping model, determining the content mapping value of each content synchronization data and the structure mapping value of each structure synchronization data, determining the content mapping and value of all content mapping values, and taking the natural logarithm of the content mapping and value to determine a first mapping value, determining the structure mapping and value of the structure mapping value, and taking the natural logarithm of the structure mapping and value to determine a second mapping value, the data mapping loss value is the first mapping value and the second mapping value The product value of the preset first data mapping loss value L1 and the preset second data mapping loss value L2 are preset, and L1<L2, the preset first mapping speed V1, the preset second mapping speed V2 and the preset third mapping speed V3 are preset, and V1<V2<V3, when the data mapping loss value is less than the preset first data mapping loss value, the preset third mapping speed is used as the data mapping speed, when the data mapping loss value is greater than or equal to the preset first data mapping loss value and less than the preset second data mapping loss value, the preset second mapping speed is used as the data mapping speed, and when the data mapping loss value is greater than or equal to the preset second data mapping loss value, the preset first mapping speed is used as the data mapping speed.

[0084] Specifically, the data to be synchronized on the merged data chain and the split data chain include content synchronization data and structure synchronization data. The content synchronization data includes field content, document content, configuration files, log information, etc., and the structure synchronization data refers to data structure, arrangement structure, etc. The data mapping model is pre-trained and is based on different data samples and the mapping values ​​matched by the data samples. The training method is consistent with the training method of the data synchronization model, and will not be repeated here. The data mapping loss value is a value that reflects the data synchronization error during the mapping synchronization process. It is used to quantify the mapping synchronization process. The larger the data mapping loss value, the more likely it is that the data will have a data synchronization error during the mapping process. The data synchronization error refers to the situation that the data may be inaccessible, damaged or lost due to some reason during the mapping process. It provides data support for determining the data mapping speed and improves the flexibility and stability of the mapping synchronization process.

[0085] In some embodiments of the present application, when obtaining a mapping coefficient of a data mapping speed based on a load relationship determined by content synchronization data and structure synchronization data, it includes: determining the application load of a source terminal and a target terminal within a preset time period according to the content synchronization data and the structure synchronization data, determining an application load graph according to the application load, and determining a minimum load point on the application load graph, determining the mapping load of a source terminal and a target terminal within a preset time period according to the content synchronization data and the structure synchronization data, calculating a mapping difference between the minimum load point and the mapping load, and when the mapping difference is greater than or equal to zero, determining the mapping coefficient of the data mapping speed to be 1, and when the mapping difference is less than zero, determining the mapping coefficient of the data mapping speed according to the application load graph.

[0086] Specifically, the application load refers to the total load caused by the mapping synchronization between the source and target ends within a preset time period, and the mapping load refers to the load caused by the mapping between the source and target ends within a preset time period. By judging whether the mapping difference is greater than or equal to zero, the stability and reliability of the mapping process can be ensured. When the mapping difference is greater than or equal to zero, the mapping coefficient of the data mapping speed is 1, indicating that the load required by the mapping process can be borne and no adjustment is required. When the mapping difference is less than zero, it indicates that the load of the mapping process exceeds the application load required at this time point. The mapping coefficient of the data mapping speed is determined according to the application load diagram, thereby adjusting the data mapping speed, thereby avoiding the mismatch between the data mapping speed and the required application load, and improving the stability and reliability of the mapping synchronization process.

[0087] In some embodiments of the present application, when determining the mapping coefficient of the data mapping speed based on the application load diagram, it includes: extracting all load slopes in the application load diagram, and determining the variance of all load slopes and the standard deviation of all load slopes, and using the ratio of the standard deviation to the variance as the mapping coefficient of the data mapping speed.

[0088] In some embodiments of the present application, when adjusting the data mapping speed according to the mapping coefficient, it includes:

[0089] The data mapping speed is proportional to the mapping coefficient.

[0090] Specifically, all load slopes in the application load graph are extracted. The load slope represents the rate of change of the load over time. By calculating the variance and standard deviation of these load slopes, the fluctuation of the application load can be accurately measured. The variance reflects the degree to which the load slope deviates from the average value, and the standard deviation further standardizes this deviation. By capturing the subtle fluctuations of the load slope, the dynamic change characteristics of the application load can be fully and accurately reflected, providing an accurate basis for the adjustment of the data mapping speed. For example, when the data mapping speed is fast, the mapping load of the mapping process will be instantly overloaded, thereby causing the mapping synchronization to fail. Assuming that the data mapping speed is V3 and the mapping coefficient is H, the adjusted data mapping speed is determined to be V3*H. By establishing a proportional relationship between the data mapping speed and the mapping coefficient, the data mapping speed can be accurately controlled to avoid the situation where the data mapping speed is too fast and does not match the required application load, thereby ensuring the synchronization stability during the mapping synchronization process.

[0091] In summary, the beneficial effects of the present invention are: mapping analysis is performed according to the data synchronization model to determine the synchronization data chain of the data to be synchronized, and the data characteristics of the header data to be synchronized and the adjacent data to be synchronized are analyzed to determine the synchronization correlation factor, and then the correlation between the data is judged, and the internal connection between the data can be accurately identified, and the data with correlation is integrated together for synchronization, thereby improving the accuracy of data synchronization, dynamically determining the data mapping speed according to the data mapping loss value, reducing the loss or incorrect mapping of data during the synchronization process, ensuring that the data at the source and target ends are highly consistent, avoiding business logic errors caused by improper data mapping, determining the mapping coefficient according to the load relationship between the content synchronization data and the structure synchronization data, and adjusting the data mapping speed accordingly. The stability of data synchronization is guaranteed, thereby effectively utilizing resources, avoiding resource waste or system overload due to mismatch between mapping speed and load requirements, and thus improving the overall synchronization efficiency.

[0092] In another preferred embodiment based on the above embodiment, refer to Figure 2 As shown, this embodiment provides a data synchronization processing system based on a data synchronization model, which is applied to the above-mentioned data synchronization processing method based on a data synchronization model, including:

[0093] The acquisition module is configured to obtain all the data to be synchronized at the source end, perform mapping analysis on all the data to be synchronized based on the data synchronization model, and determine the synchronization data chain of the data to be synchronized;

[0094] A judgment module is configured to extract the header data to be synchronized and the adjacent data to be synchronized that are adjacent to the header data to be synchronized on the synchronization data chain, determine the synchronization association factor according to the data characteristics of the header data to be synchronized and the adjacent data to be synchronized, judge whether the header data to be synchronized and the adjacent data to be synchronized are associated based on the synchronization association factor, judge whether the remaining data to be synchronized on the synchronization data chain are associated with the header data to be synchronized, and determine whether to merge the data chain and split the data chain according to the judgment result;

[0095] The processing module is configured to analyze the merged data chain and the split data chain, determine all content synchronization data and all structure synchronization data, determine a data mapping loss value of the source end according to the content synchronization data and all structure synchronization data, and determine a data mapping speed of the source end and the target end according to the data mapping loss value;

[0096] The mapping module is configured to obtain a mapping coefficient of a data mapping speed based on a load relationship determined by the content synchronization data and the structure synchronization data, and adjust the data mapping speed according to the mapping coefficient.

[0097] Specifically, mapping analysis is performed according to the data synchronization model to determine the synchronization data chain of the data to be synchronized, and the data characteristics of the first data to be synchronized and the adjacent data to be synchronized are analyzed to determine the synchronization correlation factor, and then the correlation between the data is judged, which can accurately identify the internal connection between the data, integrate the related data together for synchronization, improve the accuracy of data synchronization, dynamically determine the data mapping speed according to the data mapping loss value, reduce the loss or wrong mapping of data during the synchronization process, ensure that the data on the source and target ends are highly consistent, avoid business logic errors caused by improper data mapping, determine the mapping coefficient according to the load relationship between the content synchronization data and the structure synchronization data, and adjust the data mapping speed accordingly. The stability of data synchronization is guaranteed, so as to effectively utilize resources, avoid resource waste or system overload caused by mismatch between mapping speed and load requirements, and thus improve the overall synchronization efficiency.

[0098] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0099] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0100] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0101] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0102] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A data synchronization processing method based on a data synchronization model, characterized in that: include: Obtain all the data to be synchronized on the source side, perform mapping analysis on all the data to be synchronized based on the data synchronization model, and determine the synchronization data chain of the data to be synchronized; Extracting the header data to be synchronized and the adjacent data to be synchronized adjacent to the header data to be synchronized on the synchronization data chain, determining a synchronization association factor according to data features of the header data to be synchronized and the adjacent data to be synchronized, judging whether the header data to be synchronized and the adjacent data to be synchronized are associated based on the synchronization association factor, judging whether the remaining data to be synchronized on the synchronization data chain are associated with the header data to be synchronized, and determining whether to merge the data chain and split the data chain according to the judgment result; Analyze the merged data chain and the split data chain to determine all content synchronization data and all structure synchronization data, determine the data mapping loss value of the source end according to all content synchronization data and all structure synchronization data, and determine the data mapping speed of the source end and the target end according to the data mapping loss value; Based on the load relationship determined by the content synchronization data and the structure synchronization data, a mapping coefficient of the data mapping speed is obtained, and the data mapping speed is adjusted according to the mapping coefficient.

2. The data synchronization processing method based on the data synchronization model according to claim 1 is characterized in that: Before obtaining all the data to be synchronized from the source, including: Acquire initial data, perform data cleaning on the initial data, wherein the data cleaning includes processing missing values, outliers and duplicate values, and perform data conversion on the cleaned initial data, wherein the data conversion includes data standardization and data normalization; Comparing the data format of the initial data after data conversion with the mapping format of the target end, and determining the data to be synchronized according to the comparison result; When the data format conforms to the mapping format of the target end, the initial data after the data format corresponding to the data is converted is used as the data to be synchronized; When the data format does not conform to the mapping format of the target end, the initial data after the data format is converted to the corresponding data is deleted.

3. The data synchronization processing method based on the data synchronization model according to claim 2 is characterized in that: When mapping and analyzing all data to be synchronized based on the data synchronization model and determining the synchronization data chain of the data to be synchronized, it includes: Obtain a historical data set to be synchronized, and divide the historical data set to be synchronized into a training set and a test set, use grid search to find hyperparameters of a neural network model, and establish a neural network model; The training set is used to fit the neural network model, the test set is substituted into the neural network model and the neural network model is evaluated, when the evaluation value reaches a preset evaluation threshold, the neural network model is used as the data synchronization model, and all the data to be synchronized are substituted into the data synchronization model, and the data synchronization value of each data to be synchronized is output; All different data synchronization values ​​are used to construct a sequence to be synchronized, and all identical data synchronization values ​​are used to construct several data synchronization value sequences, and one data synchronization value is obtained from each of the data synchronization value sequences, and the remaining data synchronization values ​​are removed; The acquired data synchronization value is added to the to-be-synchronized sequence, and the sequences are arranged in descending order of the data synchronization value to determine the synchronization data chain.

4. The data synchronization processing method based on the data synchronization model according to claim 3 is characterized in that: When determining a synchronization association factor according to data features of the header data to be synchronized and the adjacent data to be synchronized, and judging whether the header data to be synchronized and the adjacent data to be synchronized are associated based on the synchronization association factor, the method includes: Analyze the data features of the header data to be synchronized, determine the header data value corresponding to each data feature and determine the header data mean; Presetting a first synchronization set, a second synchronization set, and a third synchronization set; The header data values ​​greater than the header data mean value are divided into the first synchronization set, the header data values ​​equal to the header data mean value are divided into the second synchronization set, and the header data values ​​less than the header data mean value are divided into the third synchronization set; Analyze the data features of the adjacent data to be synchronized, determine the adjacent data value corresponding to each data feature and determine the adjacent data mean; The adjacent data values ​​greater than the adjacent data mean are divided into the first synchronization set, the adjacent data values ​​equal to the adjacent data mean are divided into the second synchronization set, and the adjacent data values ​​less than the adjacent data mean are divided into the third synchronization set; Counting the first synchronization number of all header data values ​​and all adjacent data values ​​in the first synchronization set, counting the second synchronization number of all header data values ​​and all adjacent data values ​​in the second synchronization set, and counting the third synchronization number of all header data values ​​and all adjacent data values ​​in the third synchronization set; The synchronization association factor is determined according to the first synchronization number, the second synchronization number and the third synchronization number, and the synchronization association factor is compared with a preset synchronization association factor, and whether there is an association between the header data to be synchronized and the adjacent data to be synchronized is determined according to the comparison result.

5. The data synchronization processing method based on the data synchronization model according to claim 4 is characterized in that: When the synchronization association factor is determined according to the first synchronization number, the second synchronization number, and the third synchronization number, and the synchronization association factor is compared with a preset synchronization association factor, and whether the header data to be synchronized and the adjacent data to be synchronized are associated according to the comparison result, the method includes: Wherein, S represents the synchronization correlation factor, A1 represents the first synchronization number, A2 represents the second synchronization number, A3 represents the third synchronization number, max(A1, A3) represents the maximum value between the first synchronization number and the third synchronization number, and max(A1, A2) represents the maximum value between the first synchronization number and the second synchronization number; When the synchronization association factor is greater than or equal to the preset synchronization association factor, it is determined that the header data to be synchronized and the adjacent data to be synchronized are associated; When the synchronization association factor is less than the preset synchronization association factor, it is determined that there is no association between the head data to be synchronized and the adjacent data to be synchronized.

6. The data synchronization processing method based on the data synchronization model according to claim 5 is characterized in that: When determining the data mapping loss value of the source end according to all content synchronization data and all structure synchronization data, and determining the data mapping speeds of the source end and the target end according to the data mapping loss value, the method includes: Substituting all content synchronization data and all structure synchronization data into a pre-trained data mapping model, and determining a content mapping value for each content synchronization data and a structure mapping value for each structure synchronization data; Determine a content mapping and value of all content mapping values, and take a natural logarithm of the content mapping and value to determine a first mapping value, determine a structure mapping and value of the structure mapping value, and take a natural logarithm of the structure mapping and value to determine a second mapping value; The data mapping loss value is a product value of the first mapping value and the second mapping value; A preset first data mapping loss value L1 and a preset second data mapping loss value L2 are preset in advance, and L1<L2; A preset first mapping speed V1, a preset second mapping speed V2 and a preset third mapping speed V3 are preset, and V1<V2<V3; When the data mapping loss value is less than the preset first data mapping loss value, using the preset third mapping speed as the data mapping speed; When the data mapping loss value is greater than or equal to the preset first data mapping loss value and less than the preset second data mapping loss value, using the preset second mapping speed as the data mapping speed; When the data mapping loss value is greater than or equal to the preset second data mapping loss value, the preset first mapping speed is used as the data mapping speed.

7. The data synchronization processing method based on the data synchronization model according to claim 6 is characterized in that: When obtaining the mapping coefficient of the data mapping speed based on the load relationship determined by the content synchronization data and the structure synchronization data, it includes: Determine the application load of the source end and the target end within a preset time period according to the content synchronization data and the structure synchronization data; Determining an application load diagram according to the application load, and determining a minimum load point on the application load diagram; Determine the mapping load of the source end and the target end within the preset time period according to the content synchronization data and the structure synchronization data; Calculating a mapping difference between the minimum load point and the mapping load, and when the mapping difference is greater than or equal to zero, determining a mapping coefficient of the data mapping speed to be 1; When the mapping difference is less than zero, a mapping coefficient of the data mapping speed is determined according to the application load map.

8. The data synchronization processing method based on the data synchronization model according to claim 7 is characterized in that: When determining the mapping coefficient of the data mapping speed according to the application load diagram, it includes: Extracting all load slopes in the applied load graph, and determining the variance of all load slopes and the standard deviation of all load slopes; The ratio of the standard deviation to the variance is used as a mapping coefficient of the data mapping speed.

9. The data synchronization processing method based on the data synchronization model according to claim 8, characterized in that: When the data mapping speed is adjusted according to the mapping coefficient, it includes: The data mapping speed is proportional to the mapping coefficient.

10. A data synchronization processing system based on a data synchronization model, applied to the data synchronization processing method based on a data synchronization model as claimed in any one of claims 1 to 9, characterized in that: include: The acquisition module is configured to obtain all the data to be synchronized at the source end, perform mapping analysis on all the data to be synchronized based on the data synchronization model, and determine the synchronization data chain of the data to be synchronized; A judgment module is configured to extract the header data to be synchronized and the adjacent data to be synchronized adjacent to the header data to be synchronized on the synchronization data chain, determine a synchronization association factor according to data features of the header data to be synchronized and the adjacent data to be synchronized, judge whether the header data to be synchronized and the adjacent data to be synchronized are associated based on the synchronization association factor, judge whether the remaining data to be synchronized on the synchronization data chain are associated with the header data to be synchronized, and determine whether to merge the data chain and split the data chain according to the judgment result; a processing module configured to analyze the merged data chain and the split data chain, determine all content synchronization data and all structure synchronization data, determine a data mapping loss value of the source end according to the content synchronization data and all structure synchronization data, and determine data mapping speeds of the source end and the target end according to the data mapping loss value; The mapping module is configured to obtain a mapping coefficient of the data mapping speed based on a load relationship determined by the content synchronization data and the structure synchronization data, and adjust the data mapping speed according to the mapping coefficient.