Industrial data online migration method, medium and system of productivity middle platform

Through deep neural network predicting data change rate and optimizing equations, dynamically adjusting migration strategies, solving the problem of inefficient industrial data migration and achieving an efficient and adaptive data migration process.

CN120256408AInactive Publication Date: 2025-07-04BEIJING NANCAL RUIYUAN DIGITAL TECH CO LTD
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510286921.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology cannot effectively respond to the dynamic changes in industrial data change rates, resulting in inefficient data migration and difficulty adapting the migration strategy to meet the performance needs of the productivity middle platform.

Method used

The deep neural network model is used to predict the data change rate, combine the data dependency matrix and optimize the equation system, and dynamically adjust the migration strategy, and adopt batch transmission, incremental migration and real-time synchronization mechanisms to monitor and update the migration strategy in real time to match the data change characteristics.

Benefits of technology

It realizes efficient data migration under dynamic changes in data change rates, improves migration efficiency and data consistency, and meets the performance requirements of the productivity middle platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256408A_ABST
    Figure CN120256408A_ABST
Patent Text Reader

Abstract

The invention provides an industrial data online migration method, medium and system of a productivity middle platform, and belongs to the technical field of electrical digital data processing.The method comprises the steps that firstly, historical access records of an industrial data source are collected, and a deep neural network model comprising a multi-head attention layer, a time sequence feature extraction layer and a residual connection layer is constructed; and the prediction of the data change rate is realized. And based on a prediction result, dynamically calculating a data change rate demarcation point by utilizing an optimization equation set, and dividing a data table into different change frequency categories. And then constructing a data dependency relationship matrix, and determining a migration sequence by adopting an improved topological sorting algorithm. Batch migration, incremental migration and real-time synchronization mechanisms are adopted for different types of data, and model updating is triggered through real-time monitoring, so that dynamic optimization of a migration strategy is ensured. Finally, data consistency verification is carried out, migration accuracy is guaranteed, and the technical problem that the online migration efficiency of industrial data is low due to real-time dynamic change of the data change rate is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of electric digital data processing. Specifically, it relates to an industrial data online migration method, medium and system for a productivity middleware platform. Background Art

[0002] Industrial data online migration is an important part of the construction of a productivity middleware platform. Traditional industrial data migration methods mainly include three modes: batch migration, incremental migration, and real-time synchronization. Batch migration realizes the rapid migration of large-scale data through parallel transmission of data block sharding. Incremental migration realizes the synchronous update of incremental data by capturing data change records. Real-time synchronization uses a two-way mirroring mechanism to ensure data consistency between the source end and the target end. In practical applications, different migration modes are applicable to different data change scenarios. For example, in the production management of an intelligent factory, equipment basic data is suitable for batch migration, production plan data is suitable for incremental migration, and real-time monitoring data requires real-time synchronization. In the field of power dispatching, grid topology data is suitable for batch migration, load forecasting data is suitable for incremental migration, and power consumption data requires real-time synchronization. In the traffic management of a smart city, road basic data is suitable for batch migration, traffic flow statistics data is suitable for incremental migration, and real-time traffic condition data requires real-time synchronization.

[0003] However, the change characteristics of industrial data often change dynamically with the change of business load. For example, in the scenario of an intelligent factory, when the production task is adjusted temporarily, the originally stable production plan data will change frequently; in the power dispatching scenario, when extreme weather occurs, the update frequency of load forecasting data will increase significantly; in the traffic management scenario, during major events, the usually stable traffic flow data will fluctuate violently. Traditional industrial data migration methods adopt a static data classification strategy and cannot adjust the migration strategy in a timely manner according to the dynamic change of data change characteristics. This results in that when the data change rate changes significantly, the original migration mode may no longer be applicable, causing excessive consumption of system resources or delay in data consistency. In addition, there are generally complex dependency relationships among industrial data, and data changes often cause chain reactions. It is difficult for traditional methods to effectively handle the impact of data dependency relationships on the migration order.

[0004] Facing the dynamic characteristics of industrial data, it is difficult for existing technologies to solve the efficiency problems caused by the mismatch between the data change rate and the migration strategy. First, there is a lack of an accurate prediction mechanism for the data change rate, and it is impossible to adjust the migration strategy in advance to cope with the changes in the data change characteristics. Second, most of the existing data classification methods are based on empirical thresholds, and it is difficult to adaptively adjust the classification boundary according to the dynamic characteristics of the actual business scenario. Finally, there are complex coupling relationships among multiple factors such as resource consumption and latency impact during the data migration process, and it is difficult for traditional methods to achieve dynamic balance of multiple objectives. These technical difficulties ultimately lead to the efficiency of online migration of industrial data being difficult to meet the performance requirements of the productivity middle platform. That is to say, there are technical problems in the existing technology that the real-time dynamic change of the data change rate leads to low efficiency of online migration of industrial data. Summary of the Invention

[0005] In view of this, the present invention provides an online migration method, medium and system for industrial data of a productivity middle platform, which can solve the technical problem that the real-time dynamic change of the data change rate in the existing technology leads to low efficiency of online migration of industrial data.

[0006] The present invention is implemented as follows: In the first aspect of the present invention, an online migration method for industrial data of a productivity middle platform includes the following steps: collecting historical access records of each data table in the industrial data source; constructing a deep neural network model including a multi-head attention layer and a time series feature extraction layer based on the historical access records, calculating the forgetting gate threshold in the gated recurrent unit structure by logarithmically transforming the periodic intensity, data flow volatility and data table size, and calculating the skip connection weight in the residual connection layer by linearly combining the time series gradient value, predicted data change rate and network transmission delay; converting the historical access records into time series feature vectors and inputting them into the deep neural network model, and training the model using a custom loss function; using the trained deep neural network model to predict the data change rate of each data table in the industrial data source within the next 24 hours; constructing a data dependency matrix; determining two break points of the data change rate by solving an optimization equation set including an entropy value equation, an information gain equation, a resource consumption equation and a latency impact equation; dividing the data tables into different change frequency categories according to the data change rate break points; determining the data migration order by using an improved topological sorting algorithm based on the data dependency matrix and the predicted data change rate; performing data migration on the data tables of different change frequency categories by using batch transmission, incremental migration and two-way real-time synchronization methods respectively; optionally, it also includes updating the prediction model in real time and performing data consistency verification during the data migration process.

[0007] Among them, the step of collecting the historical access records is specifically to collect the data access timestamp, data table identifier, data operation type, data update frequency, business importance level, data table size, data complexity, service response time, data synchronization delay, system throughput, network bandwidth occupancy, storage space requirement, processor load, and memory usage by deploying data collection probes.

[0008] Among them, the step of constructing the deep neural network model is specifically to set the number of attention heads to 8, the dimension of each attention head to 64, and use the scaled dot-product attention mechanism to calculate the correlation weights between features; the reference range of the forgetting gate threshold is from 0.1 to 0.9, and the reference range of the skip connection weight is from 0.2 to 0.8.

[0009] Among them, in the step of training the deep neural network model, the historical access records are segmented by time window, the window size is set to 1 hour, and the sliding step is 5 minutes; the zero-mean normalization method is used for feature normalization; the learning rate is set to 0.001, and the number of training rounds is 1000.

[0010] Among them, in the step of predicting the data change rate, the confidence level of the prediction result is evaluated, the confidence threshold is set to 0.85, and the prediction results with a confidence level lower than the threshold are smoothed.

[0011] Among them, in the step of determining the data change rate breakpoint, the base of entropy value calculation is set to 2, and the normalization coefficient of information gain is 0.5; the quasi-Newton method is used to solve the optimization equation system, the maximum number of iterations is set to 1000, and the convergence threshold is 0.0001; the reference value of the first breakpoint is 0.1, and the reference value of the second breakpoint is 0.5.

[0012] Among them, for the data tables with a predicted data change rate less than the first breakpoint, a batch data transfer channel is used for data migration; for the data tables with a predicted data change rate between the two breakpoints, a remote direct memory access channel is used for incremental migration; for the data tables with a predicted data change rate greater than the second breakpoint, a two-way real-time synchronization mechanism is used for data migration.

[0013] Among them, the real-time change rate of the data table is calculated in real time during the data migration process. When the difference between the real-time change rate and the predicted data change rate exceeds 20%, the online update of the deep neural network model is triggered.

[0014] The second aspect of the present invention provides a computer-readable storage medium, in which program instructions are stored. When the program instructions run on a computer, they are used to execute the above-mentioned industrial data online migration method of a productivity middle platform.

[0015] The third aspect of the present invention provides an industrial data online migration system for a productivity middle platform, which includes the above-mentioned computer-readable storage medium. The system can be any one of a computer, a server, and a single-chip microcomputer. The computer-readable storage medium is disposed within the system, and a microprocessor for executing program instructions stored within the computer-readable storage medium is provided within the system.

[0016] Compared with the prior art, the present invention provides an industrial data online migration method, medium, and system for a productivity middle platform. The present invention proposes an industrial data online migration method based on a deep neural network. By constructing a deep neural network including a multi-head attention layer, a temporal feature extraction layer, and a residual connection layer, accurate prediction of the data change rate is achieved. This method first collects comprehensive historical access records, including multi-dimensional information such as data access patterns and business load characteristics, and establishes an accurate description of data change behavior. Through the training of the deep neural network model, the system can capture the temporal patterns and long-term dependencies of data changes, providing a reliable decision basis for the dynamic adjustment of migration strategies.

[0017] The method of the present invention uses an optimized equation set to dynamically calculate the break point of the data change rate, overcoming the limitations of the traditional fixed threshold method. By incorporating factors such as the uncertainty measure of data changes, migration resource consumption, and system performance impact into a unified optimization framework, adaptive adjustment of the classification boundary is achieved. At the same time, this method introduces a data dependency matrix and determines the optimal migration order through an improved topological sorting algorithm, effectively solving the impact of data dependency relationships on migration efficiency. During the migration execution process, the system monitors the data change characteristics in real time and triggers model updates in a timely manner, ensuring the continuous matching of migration strategies with data change characteristics. In addition, at the system implementation level, the present invention adopts technologies such as block parallel transmission, remote direct memory access, and two-way real-time synchronization, providing an efficient execution mechanism for different migration modes.

[0018] The present invention successfully solves the efficiency problem in industrial data online migration, which is mainly reflected in three aspects: First, accurate prediction of the data change rate is achieved through a deep neural network, enabling the system to adjust migration strategies in advance and avoiding efficiency losses caused by strategy lag in traditional methods. Second, the dynamic solution of the optimized equation set ensures that the classification boundary is always in an optimal state, improving the accuracy of migration mode selection. Third, the improved topological sorting algorithm takes into account the impact of data dependency relationships, optimizes the data migration order, and reduces unnecessary data synchronization operations. These technological innovations enable the present invention to always maintain a high migration efficiency under the condition of dynamic changes in the data change rate, meeting the requirements of the productivity middle platform for data migration performance and solving the technical problem of low efficiency in industrial data online migration caused by real-time dynamic changes in the data change rate in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 This is a flowchart of the method of the present invention.

[0020] Figure 2 It is a data access feature diagram of different service modules in Embodiment 2;

[0021] Figure 3 It is a change trend diagram of the loss function, training accuracy, and validation accuracy during the model training process in Embodiment 2;

[0022] Figure 4 It is a migration performance index diagram of different categories of data in Embodiment 2;

[0023] Figure 5 It is a dynamic change comparison diagram of the system load within 24 hours in Embodiment 2. Detailed implementation manners

[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0025] As Figure 1 shown, it is a flowchart of an industrial data online migration method for a productivity middle platform provided in the first aspect of the present invention. This method includes the following steps:

[0026] S01. Collect historical access records of each data table in the industrial data source. The historical access records include data access timestamps, data table identifiers, data operation types, data update frequencies, business importance levels, data table sizes, data complexities, service response times, data synchronization delays, system throughputs, network bandwidth occupations, storage space requirements, processor loads, and memory usages;

[0027] S02. Construct a deep neural network model. The deep neural network model includes an input layer, a multi-head attention layer, a temporal feature extraction layer, a residual connection layer, and an output layer. Among them, the multi-head attention layer is used to capture the long-term dependence relationships between data tables; the temporal feature extraction layer adopts a gated recurrent unit structure, and the forgetting gate threshold in the gated recurrent unit structure is calculated by function F1, and the skip connection weight in the residual connection layer is calculated by function F2;

[0028] S03. Convert the historical access records into temporal feature vectors and input them into the deep neural network model, and use a custom loss function for model training. The custom loss function combines the mean square error loss and the data table association strength weight to obtain a data change prediction model;

[0029] S04. Use the data change prediction model to predict the data change rate of each data table in the industrial data source within the next 24 hours, and obtain the predicted data change rate;

[0030] S05. Construct a data dependency matrix, where the element values in the data dependency matrix represent the degree of dependency between data tables, and the degree of dependency is calculated through the reference frequency and data flow direction between data tables;

[0031] S06. Solve according to the optimization equation system to obtain the first data change rate breakpoint and the second data change rate breakpoint;

[0032] S07. Divide the data tables with the predicted data change rate less than the first data change rate breakpoint into low-frequency change data, divide the data tables with the predicted data change rate greater than the second data change rate breakpoint into high-frequency change data, and divide the remaining data tables into medium-frequency change data;

[0033] S08. Based on the data dependency matrix, combined with the predicted data change rate, use an improved topological sorting algorithm to determine the data migration order;

[0034] S09. For the low-frequency change data, establish a batch data transmission channel and perform data migration in a block parallel manner;

[0035] S10. For the medium-frequency change data, establish a remote direct memory access channel and perform data migration in an incremental migration manner, and the incremental migration manner includes data block splitting, parallel transmission, and breakpoint resumption mechanism;

[0036] S11. For the high-frequency change data, establish a two-way real-time synchronization mechanism on the basis of the remote direct memory access channel, and the two-way real-time synchronization mechanism includes data change capture, change log transmission, and real-time synchronization processing;

[0037] S12. During the data migration process, calculate the real-time change rate of the data table in real time. When the difference between the real-time change rate and the predicted data change rate exceeds the preset deviation threshold, trigger the online update of the data change prediction model;

[0038] S13. Perform data consistency verification on the data tables that have completed migration, and the data consistency verification includes data content verification, structural consistency verification, and association integrity verification;

[0039] The function F1 in the deep neural network model is used to calculate the forgetting gate threshold in the gated recurrent unit structure. The inputs include the periodic intensity, data traffic volatility, and data table size, and the output is a forgetting gate threshold. This function performs logarithmic transformation on the input features and combines them with weights, enabling data with strong periodicity to retain more historical information, while data with large fluctuations or large volume tend to forget more historical states, thereby achieving adaptive memory control for different feature data.

[0040] The function F2 in the deep neural network model is used to calculate the skip connection weight in the residual connection layer. The inputs include the temporal gradient value, prediction data change rate, and network transmission delay, and the output is a skip connection weight value. This function linearly combines the temporal gradient information and the data change prediction result, and takes into account the negative impact of network transmission delay, realizing dynamic adjustment of the feature propagation path, enabling the model to adaptively select the information transmission method according to the temporal characteristics of the data and the network environment.

[0041] Among them, the optimization equation set includes an entropy value equation, an information gain equation, a resource consumption equation, and a time delay impact equation:

[0042] The entropy value equation is used to calculate the uncertainty measure of data change. The inputs include the data update frequency, the data table size, the data complexity, the prediction data change rate, the temporal gradient value, and the periodic intensity, and the output is the information entropy value;

[0043] The information gain equation is used to evaluate the benefit of data partitioning. The inputs include the dependence degree, the data consistency check result, the business importance level, the service response time, the data synchronization delay, and the system throughput, and the output is the partitioning benefit value;

[0044] The resource consumption equation is used to estimate the resource cost of data migration. The inputs include the network bandwidth occupancy, the storage space requirement, the processor load, the memory usage, the prediction data change rate, and the dependence degree, and the output is the resource consumption vector;

[0045] The time delay impact equation is used to evaluate the impact of data migration on system performance. The inputs include the network transmission delay, the data synchronization delay, the service response time, the system throughput, the resource consumption vector, and the partitioning benefit value, and the output is the time delay impact degree;

[0046] Among them, the term explanations are as follows:

[0047] Data complexity is a comprehensive metric value that refers to the degree of diversity of field types in the data table and the complexity of the association relationship between fields;

[0048] The temporal gradient value refers to the rate of change of the data access pattern in the time dimension;

[0049] The periodic intensity refers to the degree of significance of the periodic pattern presented by the data access behavior in the time series;

[0050] The data traffic volatility refers to the amplitude of change in the data transmission volume per unit time;

[0051] The improved topological sorting algorithm is an improved algorithm that introduces the weights of data change rate and dependence degree on the basis of the traditional topological sorting algorithm.

[0052] The following is a detailed description of the specific implementation manners of the above steps. The specific implementation manner of step S01 is to comprehensively collect the historical access records of each data table in the industrial data source by using a distributed data collection framework. This step is implemented through multiple sub-steps such as real-time data monitoring, log parsing, and performance metric collection. First, deploy data collection probes, install lightweight monitoring components on the database server and application server. The collection probes are implemented based on the lock-free queue technology to achieve high-concurrency data collection. Reduce the data sampling frequency through time window aggregation, set the sampling interval to 1 second, regularly scan the database system tables and performance counters, collect information such as data access timestamps, data table identifiers, and data operation types. Then, use a log parser to perform real-time parsing on the database log files to extract information such as data update frequency and business importance level. Next, use the system monitoring module to collect static features such as data table size and data complexity. Finally, collect dynamic metrics such as service response time and data synchronization delay through the performance monitoring component. The purpose of this step is to establish a complete data access behavior portrait and provide data support for subsequent data migration decisions.

[0053] The specific implementation manner of step S02 is to construct a multi-layer deep neural network model. This step is implemented through multiple sub-steps such as neural network structure design, parameter initialization, and model assembly. First, design the input layer structure, use a fully connected layer to encode multi-dimensional features. Then, construct a multi-head attention layer, set the number of attention heads to 8, and the dimension of each attention head to 64. The attention layer calculates the correlation weights between features through the scaled dot-product attention mechanism. Next, implement the temporal feature extraction layer, using a gated recurrent unit structure, which includes a reset gate and an update gate. The forgetting gate threshold is calculated by function F1, and the threshold reference range is from 0.1 to 0.9. Then, design the residual connection layer, and the skip connection weight is calculated by function F2, and the weight reference range is from 0.2 to 0.8. Finally, construct the output layer, and use a linear activation function to output the prediction result. The purpose of this step is to construct a deep learning model that can effectively capture the data table access pattern and dependence relationship.

[0054] The specific implementation of step S03 is to convert historical access records into time-series feature vectors that can be processed by the model and perform model training. This step is achieved through multiple sub-steps such as data preprocessing, feature engineering, and model optimization. First, the historical access records are segmented by time window, with the window size set to 1 hour and the sliding step size to 5 minutes. Then, feature standardization is performed, and the zero-mean normalization method is used to normalize the feature values to the same scale. Next, a time-series feature vector is constructed, including dimensions such as access frequency, update volume, and access pattern. Then, a custom loss function is designed, which combines the mean squared error loss and the data table association strength weight, and the weight coefficient is determined through grid search. Finally, a gradient descent optimizer is used for model training, with the learning rate set to 0.001 and the number of training epochs to 1000. The purpose of this step is to obtain an accurate data change prediction model.

[0055] The specific implementation of step S04 is to predict the data change rate using the trained deep neural network model. This step is achieved through multiple sub-steps such as feature extraction, model prediction, and result post-processing. First, the data features at the current moment are extracted and standardized to construct the model input vector. Then, the feature vector is input into the trained deep neural network model, and the predicted value of the data change rate within the next 24 hours is obtained through forward propagation calculation. Next, the confidence level of the prediction result is evaluated, with the confidence level threshold set to 0.85. The prediction results with a confidence level lower than the threshold are smoothed. Finally, the prediction result is converted into a standardized data change rate index. The purpose of this step is to provide a reliable prediction basis for subsequent data classification.

[0056] The specific implementation of step S05 is to construct a matrix reflecting the dependency relationship between data tables. This step is achieved through multiple sub-steps such as dependency relationship analysis, weight calculation, and matrix construction. First, the reference relationship between data tables is analyzed through the data dictionary and business logs, and the reference frequency is counted. Then, based on data flow analysis technology, the data flow direction is identified, and a directed graph is used to represent the data flow path. Next, the dependency degree between data tables is calculated, and factors such as reference frequency, flow direction weight, and business importance are considered in the calculation. Finally, a dependency relationship matrix is constructed, and the matrix element values range from 0 to 1. The purpose of this step is to establish a quantitative representation of the dependency relationship between data tables.

[0057] The specific implementation of step S06 is to solve the optimization equation system to obtain the demarcation points for data classification. This step is achieved through multiple sub-steps such as parameter initialization, iterative solution, and convergence judgment. First, initialize the parameters of the equation system, set the base of entropy value calculation to 2, and the normalization coefficient of information gain to 0.5. Then, use the quasi-Newton method to solve the optimization equation system, set the maximum number of iterations to 1000, and the convergence threshold to 0.0001. Next, calculate the uncertainty of data change through the entropy equation, evaluate the division benefit using the information gain equation, estimate the migration cost based on the resource consumption equation, and evaluate the performance impact through the delay impact equation. Finally, obtain two demarcation points for the data change rate, with the reference value of the first demarcation point being 0.1 and the reference value of the second demarcation point being 0.5. The purpose of this step is to determine the optimal threshold for data classification.

[0058] The specific implementation of step S07 is to classify the data table according to the predicted data change rate and demarcation points. This step is achieved through multiple sub-steps such as threshold comparison, category division, and result verification. First, compare the predicted data change rate with the first and second demarcation points. Then, divide the data table into three categories: low-frequency change, medium-frequency change, and high-frequency change according to the comparison results. Next, verify the rationality of the division results, check the consistency of the division of data tables within the same business domain. Finally, handle special cases, such as adopting consistent classification categories for strongly correlated data table groups. The purpose of this step is to achieve a reasonable classification of the data table.

[0059] The specific implementation of step S08 is to determine the execution order of data migration. This step is achieved through multiple sub-steps such as dependency analysis, weight calculation, and sorting optimization. First, construct a directed acyclic graph based on the dependency relationship matrix, where the nodes in the graph represent data tables and the edges represent dependency relationships. Then, calculate the node weights, and the weight calculation considers the data change rate, dependency degree, and business importance. Next, use an improved topological sorting algorithm for sorting. The algorithm introduces node weights and edge weights on the basis of the traditional topological sorting, and maintains the nodes to be processed through a priority queue. Finally, obtain the migration order considering the dependency relationship and change characteristics. The purpose of this step is to ensure the rationality of the order in the data migration process.

[0060] The specific implementation of step S09 is to perform batch migration on low-frequency change data. This step is achieved through multiple sub-steps such as channel establishment, data transmission, and progress management. First, establish a batch data transmission channel, and set the number of channels to twice the number of processor cores. Then, divide the data table into fixed-size data blocks according to the size, and set the block size to 64 megabytes. Next, use the multi-thread parallel method to transmit the data blocks, set the transmission buffer size to 16 megabytes, and manage the transmission tasks through a thread pool. Finally, implement the transmission progress tracking and failure retry mechanism. The purpose of this step is to efficiently complete the migration of low-frequency change data.

[0061] The specific implementation of step S10 is to perform incremental migration on intermediate-frequency change data. This step is achieved through multiple sub-steps such as channel configuration, data synchronization, and task scheduling. First, a Remote Direct Memory Access (RDMA) channel is established, and the size of the memory mapping area is configured to 20% of the total memory. Then, a data block segmentation mechanism is implemented to slice the incremental data according to the timestamp and primary key range. Next, a parallel transmission pipeline is established, and the producer-consumer pattern is adopted to handle data transmission. Finally, a resume-from-breakpoint mechanism is implemented to record transmission checkpoint information and support recovery after task interruption. The purpose of this step is to achieve reliable migration of intermediate-frequency change data.

[0062] The specific implementation of step S11 is to achieve bidirectional synchronization of high-frequency change data. This step is achieved through multiple sub-steps such as monitoring configuration, log transmission, and synchronization processing. First, change data capture is configured based on the RDMA channel, and triggers and listeners are set to achieve real-time capture of data changes. Then, a change log transmission channel is established, and an asynchronous queue is used to cache change logs, with the queue depth set to 1000. Next, real-time synchronization processing is implemented, and the synchronization performance is optimized through multi-level caching, with the synchronization delay threshold set to 100 milliseconds. Finally, a conflict detection and resolution mechanism is implemented. The purpose of this step is to ensure real-time synchronization of high-frequency change data.

[0063] The specific implementation of step S12 is to achieve online update of the prediction model. This step is achieved through multiple sub-steps such as monitoring and analysis, trigger judgment, and model update. First, the change rate of the data table is calculated in real time, and the sliding window method is used to count the number of changes within a unit time. Then, the real-time change rate is compared with the predicted value, the difference is calculated and compared with the preset deviation threshold, with the deviation threshold set to 20%. Next, when the difference exceeds the threshold, the model update is triggered, and the model parameters are updated using the incremental learning method. Finally, the performance of the updated model is verified. The purpose of this step is to maintain the accuracy of the prediction model.

[0064] The specific implementation of step S13 is to perform data consistency verification. This step is achieved through multiple sub-steps such as content verification, structure comparison, and integrity check. First, data content verification is performed, and methods such as hash verification and cyclic redundancy check are used to verify the consistency of data content. Then, structural consistency verification is executed to compare information such as table structures, indexes, and constraints. Next, associated integrity verification is performed to check foreign key relationships and referential integrity. Finally, a verification report is generated to record the verification results and exception information. The purpose of this step is to ensure the correctness and integrity of data migration.

[0065] The forgetting gate threshold calculation function F1 in the data change prediction model is specifically expressed as follows:

[0066] F forget = k1ln(P s ) - k2ln(Vf ) - k3ln(S d ) + ∈1;

[0067] In the formula, F forget is the forgetting gate threshold; P s is the periodic intensity; V f is the data traffic volatility; S d is the data table size; k1, k2, k3 are weight coefficients; ∈1 is the error term.

[0068] The calculation formula for the periodic intensity P s is as follows:

[0069]

[0070] In the formula, X i is the spectral component after performing Fourier transform on the data access time series; α i is the spectral component weight; n is the number of spectral components; β is the baseline parameter.

[0071] The calculation formula for the data traffic volatility V f is as follows:

[0072]

[0073] In the formula, f t is the data traffic at time t; T is the size of the statistical time window; γ is the smoothing factor.

[0074] The jump connection weight calculation function F2 in the residual connection layer is specifically expressed as follows:

[0075]

[0076] In the formula, F skip is the jump connection weight; G t is the time series gradient value; R p is the predicted data change rate; D n is the network transmission delay; m1, m2, m3 are weight coefficients; ∈2 is the error term.

[0077] The calculation formula for the time series gradient value G t is as follows:

[0078]

[0079] In the formula, f is the data access frequency function; t i is the time point; N is the number of sampling points; δ is the smoothing term.

[0080] The data dependency matrix D is specifically expressed as follows:

[0081]

[0082] where d ij represents the degree of dependence between data table i and data table j, and the calculation formula is:

[0083]

[0084] where c ij is the citation frequency; w ij is the flow weight.

[0085] The entropy value equation of the optimized equation set is specifically expressed as follows:

[0086]

[0087] where H is the information entropy value; p i is the data change probability; n is the number of data tables; ∈3 is the error term.

[0088] The information gain equation is specifically expressed as follows:

[0089]

[0090] where G is the information gain value; H0 is the entropy value before partitioning; H j is the entropy value of subset j; S is the data set; S j is the subset; m is the number of partitions; λ is the regularization coefficient; R is the regularization term.

[0091] The resource consumption equation is specifically expressed as follows:

[0092] C = ω1B + ω2M + ω3P + ω4S + ∈4;

[0093] where C is the resource consumption value; B is the bandwidth occupancy rate; M is the memory usage rate; P is the processor load; S is the storage space occupancy rate; ω1, ω2, ω3, ω4 are the weight coefficients; ∈4 is the error term.

[0094] The delay impact equation is specifically expressed as follows:

[0095] L = θ1T n + θ2T s + θ3T r + ∈5;

[0096] where L is the delay impact degree; T n is the network transmission delay; T s is the data synchronization delay; T r is the service response time; θ1, θ2, θ3 are the weight coefficients; ∈5 is the error term.

[0097] The parameter descriptions are as follows:

[0098] 1. The value ranges of k1, k2, and k3 are [0.3, 0.5], [0.2, 0.4], and [0.1, 0.3] respectively, and the optimal values are determined through grid search;

[0099] 2. α i It is determined through spectrum analysis, and the range is [0, 1]; β is fixed at 0.1;

[0100] 3. γ takes the value of 0.01 to prevent the denominator from being 0;

[0101] 4. The value ranges of m1, m2, and m3 are [0.4, 0.6], [0.3, 0.5], and [0.2, 0.4] respectively, and are optimized through gradient descent;

[0102] 5. δ takes the value of 0.001 for numerical stability;

[0103] 6. w ij It is determined according to the data flow analysis, and the range is [0, 1];

[0104] 7. λ takes the value of 0.1 to control the regularization strength;

[0105] 8. The value ranges of ω1, ω2, ω3, and ω4 are all [0.1, 0.4], and are dynamically adjusted according to the importance of system resources;

[0106] 9. The value ranges of θ1, θ2, and θ3 are all [0.2, 0.5], and are fitted through historical performance data;

[0107] 10. The value ranges of all ∈ terms are [-0.1, 0.1].

[0108] Principle description of equation construction:

[0109] 1. The logarithmic transformation is adopted for calculating the forgetting gate threshold because there are dimensional differences among data features. The logarithmic transformation can convert the multiplication relationship into an addition relationship, simplifying the calculation while maintaining monotonicity;

[0110] 2. The periodic intensity calculation uses the Fourier transform, which can effectively capture the periodic patterns in the time series. The weight α i reflects the importance of different frequency components;

[0111] 3. The data flow volatility adopts the average value of the relative change rate, which can better reflect the dynamic characteristics of the data flow. A smoothing factor is added to prevent the denominator from being 0;

[0112] 4. The skip connection weights adopt a linear combination form, and the reciprocal term of the predicted change rate is introduced, reflecting the characteristic that high-change-rate data requires more short-circuit connections;

[0113] 5. The calculation of the temporal gradient value considers local derivatives and can reflect the changing trend of the data access pattern. The square root term is used to reduce the influence of outliers;

[0114] 6. The dependency matrix is normalized to ensure the comparability of the dependency degree, and the directionality of data flow is reflected through the flow weights;

[0115] 7. The entropy value equation is based on information theory and is used to measure the uncertainty of data changes, providing a theoretical basis for data classification;

[0116] 8. The information gain equation combines the entropy value change and the regularization term to avoid overfitting problems;

[0117] 9. The resource consumption equation and the delay impact equation adopt a weighted summation form, and the relative importance of different resource metrics is reflected through dynamic weights.

[0118] The derivation process of the following relevant equations or calculation processes is described in detail below.

[0119] Derivation process of the forgetting gate threshold calculation function F1:

[0120] First, construct the feature vector X:

[0121] X = [x1, x2,..., x n T ;

[0122] where x i are the original features, including temporal metrics such as access frequency and data volume;

[0123] Then, standardize the features:

[0124]

[0125] where μ is the mean vector; σ is the standard deviation vector;

[0126] Next, extract the periodic features:

[0127] F Freq = FFT(X norm );

[0128]

[0129] where FFT is the fast Fourier transform; F freq,i is the i-th frequency component; α i is determined by the maximum entropy principle.​

[0130] Finally, optimize the weight coefficients through a multi-layer perceptron:

[0131]

[0132] In the formula, y j is the actual forgetting gate value; is the predicted value; m is the number of samples.

[0133] Derivation process of the skip connection weight function F2:

[0134] First, construct the time series feature matrix T:

[0135]

[0136] In the formula, t ij represents the j-th eigenvalue at the i-th time point.

[0137] Then calculate the time series gradient:

[0138]

[0139] Construction process of the data dependency relationship matrix:

[0140] First, construct the access frequency matrix C:

[0141]

[0142] In the formula, c ij represents the number of times data table i references data table j.

[0143] Then construct the flow direction weight matrix W:

[0144]

[0145] In the formula, w ij is determined through data flow analysis.

[0146] Solution process of the optimization equation system:

[0147] First, construct the objective function:

[0148] J = α1H + α2G + α3C + α4L;

[0149] In the formula, α1, α2, α3, α4 are weight coefficients.

[0150] Then use the quasi-Newton method to solve:

[0151]

[0152] In the formula, x kis the variable value for the k-th iteration; η k is the step size; H k is the approximation of the Hessian matrix; g k is the gradient vector.

[0153] Explanation of the equation effect:

[0154] 1. The forgetting gate threshold calculation function F1 realizes the adaptive memory control of different data features through logarithmic transformation and weight optimization, and can automatically adjust the memory depth according to the periodicity, volatility and scale of the data;

[0155] 2. The skip connection weight function F2 combines the temporal gradient and network delay to realize the dynamic adjustment of the residual connection, and improves the model's ability to capture temporal features;

[0156] 3. The data dependency matrix accurately depicts the association relationship between data tables through access frequency and flow analysis, providing a basis for optimizing the migration order;

[0157] 4. The optimization equation system balances the entropy value, gain, resource consumption and delay impact through multi-objective optimization, and realizes the optimal partition of data classification.

[0158] Explanation of parameter acquisition:

[0159] 1. The periodic feature weight α i is calculated by the principle of maximum entropy:

[0160]

[0161] where p i is the probability distribution of the frequency component.

[0162] 2. The flow weight w ij is obtained through data flow graph analysis:

[0163]

[0164] where f ij is the data flow.

[0165] 3. The weight coefficients of the optimization equation system are determined by cross-validation:

[0166]

[0167] where K is the number of folds of cross-validation; L k is the loss value of the k-th fold.

[0168] In a second aspect of the present invention, a computer-readable storage medium is provided. Program instructions are stored in the computer-readable storage medium, and when the program instructions run on a computer, they are used to execute the industrial data online migration method of a productivity middle platform described above.

[0169] In a third aspect of the present invention, an industrial data online migration system for a productivity middle platform is provided, which includes the above-mentioned computer-readable storage medium. The system can be any one of a computer, a server, and a single-chip microcomputer. The computer-readable storage medium is arranged inside the system, and a microprocessor for executing the program instructions stored in the computer-readable storage medium is arranged inside the system.

[0170] Specifically, the principle of the present invention is as follows: The present invention proposes an industrial data online migration method based on deep learning. By constructing a deep neural network that integrates a multi-head attention mechanism and a gated recurrent unit, accurate modeling and prediction of data change behaviors are realized. The core innovation point of this method is the design of an adaptive forgetting gate threshold calculation mechanism. Through the weighted combination of logarithmic transformation of the periodic intensity of data access, traffic volatility, and data table size, differential memory control of different feature data is achieved. At the same time, a dynamic weight calculation method based on time series gradient values, predicted change rates, and network latency is introduced in the residual connection layer, enhancing the model's ability to capture time series features.

[0171] In terms of data classification, the present invention innovatively constructs a multi-objective optimization equation set including entropy value, information gain, resource consumption, and delay impact. By solving with the quasi-Newton method, the optimal demarcation point of the data change rate is obtained, realizing the adaptive classification of data. Based on the classification results, the method adopts three strategies of batch transmission, incremental migration, and two-way real-time synchronization respectively, and improves the transmission efficiency through the remote direct memory access technology.

[0172] The technical solution of the present invention combines deep learning and optimization theory to solve the problems of empirical threshold setting, fixed classification strategies, and low transmission efficiency in traditional data migration methods. This method shows good adaptability and scalability in practical applications, can automatically adjust the migration strategy according to the dynamic changes of data access patterns, and at the same time ensures the stability and reliability of the large-scale data migration process through distributed data collection and multi-level cache optimization.

[0173] A specific embodiment 1 of the present invention is provided below. The specific implementation manners of each step in this embodiment 1 are described in detail as follows.

[0174] The specific implementation of step S01 is to comprehensively collect and monitor the historical access records of each data table in the industrial data source using a distributed data collection framework. First, data collection probes are deployed on the database server and the application server. The collection probe is implemented based on the lock-free queue technology to achieve high-concurrency data collection, and the collection interval is set to 1 second. The specific collection content includes: collecting the data access timestamp, data table identifier, and data operation type through the database system table; collecting the data update frequency and business importance level through the log parser; collecting the data table size and data complexity through system monitoring; collecting the service response time, data synchronization delay, and system throughput through performance monitoring; collecting the network bandwidth occupancy, storage space requirement, processor load, and memory usage through resource monitoring. The collected data is stored in a distributed time-series database after being standardized, providing a data basis for subsequent deep learning model training. This step establishes a complete data access behavior profile to support subsequent change prediction and migration decision-making.

[0175] The specific implementation of step S02 is to construct a deep neural network model. First, the network structure is designed. The input layer uses a fully connected layer to encode multi-dimensional features; the multi-head attention layer sets 8 attention heads, and the dimension of each head is 64. The attention calculation uses the scaled dot-product mechanism:

[0176]

[0177] In the formula, Q is the query matrix, K is the key matrix, V is the value matrix, and d k is the dimension of the key vector. The time-series feature extraction layer uses a gated recurrent unit structure, which includes a reset gate and an update gate. The update gate state is calculated as follows:

[0178] z t = σ(W z x t + U z h t-1 + b z );

[0179] In the formula, z t is the update gate state, x t is the input vector, h t-1 is the previous hidden state, W z , U z are weight matrices, and b z is the bias term. The forgetting gate threshold is calculated through the function F1, and the specific calculation formula is:

[0180] F forget = k1ln(P s ) - k2ln(V f ) - k3ln(S d ) + ∈1;

[0181] The meanings of the parameters in the formula are the same as those described above. The skip connection weights of the residual connection layer are calculated by the function F2:

[0182]

[0183] The meanings of the parameters in the formula are the same as those described above. The output layer uses a linear activation function to output the prediction result. This step realizes the modeling of data change behavior through deep learning technology.

[0184] The specific implementation of step S03 is to perform model training. First, the historical access records are segmented by time window, the window size is 1 hour, and the sliding step is 5 minutes to construct training samples. Then, feature engineering is carried out, and the feature vectors are standardized:

[0185]

[0186] In the formula, X is the original feature vector, μ is the mean value, and σ is the standard deviation. Then, a custom loss function is designed:

[0187] Loss = MSE + λ∑ i,j w ij (y i -y j ) 2 ;

[0188] In the formula, MSE is the mean square error, w ij is the association strength of the data table, y i , y j are the predicted values, and λ is the weight coefficient. The Adam optimizer is used for model training, the learning rate is set to 0.001, and the number of training rounds is 1000. This step realizes the optimization of model parameters and improves the prediction accuracy.

[0189] The specific implementation of step S04 is to perform data change prediction. First, a prediction input vector is constructed, which includes all monitoring indicators at the current moment. Then, the trained deep neural network model is used for prediction to obtain the data change rate in the next 24 hours. The confidence level of the prediction result is evaluated:

[0190]

[0191] In the formula, σ pred is the standard deviation of the predicted value, μ pred is the mean value of the predicted value. When the confidence level is lower than 0.85, exponential smoothing processing is adopted:

[0192]

[0193] In the formula, α is the smoothing coefficient, and its value is 0.3. This step realizes the accurate prediction of data change behavior.

[0194] The specific implementation of step S05 is to construct a data dependency matrix. First, analyze the reference relationships between data tables through a data dictionary and business logs, and count the reference frequency matrix C:

[0195]

[0196] In the formula, c ij represents the number of times data table i references data table j. Then, construct a flow weight matrix W based on data flow analysis. Finally, calculate the degree of dependence:

[0197]

[0198] In the formula, d ij is the degree of dependence of data table i on j. This step quantifies the dependence relationship between data tables and provides a basis for optimizing the migration order.

[0199] The specific implementation of step S06 is to solve the optimization equations to obtain the breakpoints of data classification. First, initialize the parameters of the optimization equations, including setting the base of the entropy equation to 2 and the normalization coefficient of information gain to 0.5. Then, construct the entropy equation:

[0200]

[0201] In the formula, p i is the data change probability, which is obtained by counting the change frequency. Then, construct the information gain equation:

[0202]

[0203] In the formula, H0 is the entropy value before partitioning, H j is the subset entropy value, and R is the regularization term. Then, construct the resource consumption equation:

[0204] C = ω1B + ω2M + ω3P + ω4S + ∈4;

[0205] In the formula, the meanings of the parameters are as described above. Finally, construct the delay impact equation:

[0206] L = θ1T n + θ2T s + θ3T r + ∈5. Use the quasi-Newton method to solve the optimization equations:

[0207]

[0208] In the formula, x k is the value of the k-th iteration variable, and H kis the Hessian matrix approximation value, the maximum number of iterations is set to 1000, and the convergence threshold is 0.0001. This step determines the optimal break point through multi-objective optimization.

[0209] The specific implementation of step S07 is to classify data according to the data change rate break point. First, compare the predicted data change rate with the first break point 0.1 and the second break point 0.5. For data tables with a change rate less than 0.1, they are classified as low-frequency change data; for data tables with a change rate greater than 0.5, they are classified as high-frequency change data; the remaining data tables are classified as medium-frequency change data. Then, perform classification adjustment. When it is found that data tables within the same business domain are classified into different categories, calculate the business relevance weight:

[0210]

[0211] where r ij is the business relevance between table i and table j, and n is the number of tables. When the business relevance weight is greater than 0.8, these data tables are classified into the same category. This step realizes the reasonable classification of data tables.

[0212] The specific implementation of step S08 is to determine the data migration order. First, construct a directed acyclic graph G=(V, E) based on the dependency matrix, where V is the set of data table nodes and E is the set of dependency edges. Then, calculate the node weight:

[0213] w v =α·R p +β·D s +γ·I b ;

[0214] where R p is the predicted change rate, D s is the dependency score, I b is the business importance, and α, β, γ are weight coefficients. Then, use the improved topological sorting algorithm for sorting. The algorithm maintains a priority queue Q:

[0215] Q={v|indegree(v)=0};

[0216] Each time, select the node with the largest weight from the queue and update the in-degree of its successor nodes. Finally, obtain the migration order considering the dependency relationship and change characteristics. This step ensures the rationality of the migration process order.

[0217] The specific implementation of step S09 is to perform batch migration on low-frequency change data. First, establish a batch data transmission channel, and the number of channels is set to twice the number of processor cores. Then, perform data block division:

[0218] B i =[di·s , d (i+1)·s ;

[0219] Wherein, B i is the i-th data block, s is the block size, set to 64 megabytes. Then design a parallel transmission algorithm:

[0220]

[0221] Wherein, t i is the transmission time of block i, and Δt i is the transmission delay. Use a thread pool to manage the transmission tasks, and the number of threads is set to the number of processor cores. Finally, implement a failure retry mechanism, with the number of retries set to 3 times, and the retry interval adopts an exponential backoff strategy:

[0222] t retry = t base · 2 n ;

[0223] Wherein, t base is the base interval time, set to 1 second, and n is the number of retries. This step realizes the efficient migration of low-frequency changed data.

[0224] The specific implementation of step S10 is to perform incremental migration on medium-frequency changed data. First, establish a remote direct memory access channel, and configure the size of the memory mapping area to 20% of the total memory. Then perform data block splitting, adopting a dual sharding strategy of time window and primary key range:

[0225] S time = [t0 + i·Δt, t0 + (i + 1)·Δt];

[0226] S key = [k min + j·Δk, k min + (j + 1)·Δk];

[0227] Wherein, Δt is the time window size, set to 5 minutes, and Δk is the primary key step size, dynamically calculated:

[0228]

[0229] Wherein, n thread is the number of threads. Then implement a parallel transmission pipeline, adopting the producer-consumer mode:

[0230]

[0231] Wherein, B mem is the memory bandwidth, B net is the network bandwidth, n conis the number of concurrent connections. Finally, implement the breakpoint resumption mechanism and record checkpoint information:

[0232] CP i = {t last , k last , status, retry};

[0233] In the formula, t last is the last transmission timestamp, and k last is the last transmission primary key value. This step ensures the reliable transmission of intermediate frequency changed data.

[0234] The specific implementation of step S11 is to achieve two-way synchronization of high-frequency changed data. First, configure the change data capture mechanism and set up triggers and listeners:

[0235] Trigger CDC = {op type , table, time, user, data};

[0236] In the formula, op type is the operation type, and data is the changed data. Then, establish a change log transmission channel and use an asynchronous queue for caching:

[0237] Q size = max(Rate change ·t delay , 1000);

[0238] In the formula, Rate change is the change rate, and t delay is the allowed delay time. Next, implement multi-level caching:

[0239] Cache1 = LRU(size1);

[0240] Cache2 = HashTable(size2);

[0241] Among them, size1 is set to 10% of the changed data volume, and size2 is set to 30% of the changed data volume. Finally, implement conflict detection:

[0242]

[0243] In the formula, conf i is the score of the i-th type of conflict, and w i is the weight. This step realizes the real-time synchronization of high-frequency changed data.

[0244] The specific implementation of step S12 is to realize the online update of the prediction model. First, calculate the real-time change rate:

[0245]

[0246] where Δd i is the i-th change amount. Then calculate the prediction deviation:

[0247]

[0248] When the deviation exceeds 20%, trigger model update. Update the parameters using incremental learning:

[0249]

[0250] where θ is the model parameter, η is the learning rate, and L is the loss function. Finally, verify the update effect:

[0251]

[0252] It is required that the accuracy improvement is greater than 5% to accept the update. This step maintains the accuracy of the prediction model.

[0253] The specific implementation of step S13 is to perform data consistency verification. First, perform data content verification and calculate the data block hash value:

[0254]

[0255] Compare the hash values of the source side and the target side. Then perform structural consistency verification and compare the table structures:

[0256]

[0257] where A and B are the structural feature sets of the source table and the target table respectively. Then perform associated integrity verification:

[0258]

[0259] where FK valid is the number of valid foreign keys, and FK total is the total number of foreign keys. Finally, generate a verification report:

[0260] Report = {content_check, struct_check, ref_check, error_list};

[0261] Contains the detailed results of all verification items. This step ensures the correctness of data migration.

[0262] To better understand and implement the present invention, the following provides Example 2 of a specific application scenario of the present invention: A large automobile manufacturing enterprise plans to migrate its existing production management system to a new productivity middle platform, involving multiple business modules such as equipment management, material management, quality management, and production planning. The total number of data tables exceeds 2,000, and the total data volume is approximately 50 TB. Traditional data migration solutions cannot meet the requirements of business continuity during the migration process, and there are problems such as low migration efficiency and high resource consumption. The R & D team uses the data online migration method of the present invention for implementation.

[0263] First, the R & D team deploys data collection probes to collect 3 months of historical access records. The data characteristics of the collection results are shown in Table 1:

[0264] Table 1 Statistical Table of Data Access Characteristics

[0265]

[0266]

[0267] Figure 2 It shows the data access characteristics of different business modules, including comparisons in three dimensions: access frequency, response time, and data volume. Based on the collected historical data, a deep neural network model is constructed. The model uses 8 attention heads, and the dimension of each head is 64. The model training parameters are set as shown in Table 2:

[0268] Table 2 Model Training Parameter Configuration Table

[0269] Parameter item Parameter value Batch size 128 Learning rate 0.001 Number of training epochs 1000 Window size 1 hour Sliding step size 5 minutes Proportion of validation set 0.2

[0270] Figure 3 It shows the changing trends of the loss function, training accuracy, and validation accuracy during the model training process. During the training process, the following parameters are used for calculating the forgetting gate threshold: k1 = 0.45, k2 = 0.35, k3 = 0.25. The calculation results of key indicators are shown in Table 3:

[0271] Table 3 Calculation Results Table of Key Indicators

[0272] Index type Device management module Material management module Quality management module Production planning module Periodic intensity 0.82 0.75 0.78 0.88 Data flow volatility 0.15 0.28 0.18 0.32 Time series gradient value 0.42 0.56 0.45 0.62

[0273] Based on the trained model, the data change rate is predicted. The confidence distribution of the prediction results is shown in Table 4:

[0274] Table 4 Prediction Confidence Distribution Table

[0275] Confidence interval Proportion of data table Above 0.95 45% 0.90~0.95 32% 0.85~0.90 18% Below 0.85 5%

[0276] The data change rate breakpoints are obtained by optimizing the solution of the equations: the first breakpoint is 0.12, and the second breakpoint is 0.48. The data table is classified based on these two breakpoints, and the classification results are shown in Table 5:

[0277] Table 5 Statistical Table of Data Table Classification Results

[0278] Change frequency category Number of data tables Proportion of data volume Proportion of access frequency Low-frequency change 865 35% 15% Medium-frequency change 724 42% 45% High-frequency change 411 23% 40%

[0279] Figure 4 The migration performance metrics of different categories of data are shown, including the comparison of transmission rate, transmission delay, and success rate. The R & D team adopts corresponding migration strategies for different categories of data: for low-frequency changed data, 32 parallel transmission channels are used, and the data block size is set to 64MB. The transmission performance metrics are shown in Table 6:

[0280] Table 6 Transmission Performance Table of Low-Frequency Changed Data

[0281] Performance index Index value Average transmission rate 850MB / s CPU utilization rate 45% Memory usage rate 35% Transmission success rate 99.9%

[0282] For medium-frequency changed data, remote direct memory access channels are configured, and the memory mapping area is set to 128GB. The incremental migration performance metrics are shown in Table 7:

[0283] Table 7 Migration Performance Table of Medium-Frequency Changed Data

[0284]

[0285]

[0286] For high-frequency changed data, a two-way real-time synchronization mechanism is implemented, and the key configuration parameters are shown in Table 8:

[0287] Table 8 Synchronization Configuration Table of High-Frequency Changed Data

[0288] Configuration item Configuration value Queue depth 1000 Size of level 1 cache 64GB Size of level 2 cache 192GB Maximum synchronization delay 100ms

[0289] During the migration process, the system calculates the deviation between the real-time change rate and the predicted change rate every hour. When the deviation exceeds 20%, the model is updated. During the 3-month migration process, the model update situation is shown in Table 9:

[0290] Table 9 Statistical Table of Model Updates

[0291] Update index Statistical value Number of updates 15 Average update time consumption 8 minutes Accuracy improvement 8.5% Prediction deviation decrease 12.3%

[0292] Finally, the consistency check of the migration results is performed, and the check results are shown in Table 10:

[0293] Table 10 Data Consistency Check Result Table

[0294] Verification item Verification result Data content consistency 99.999% Table structure consistency 100% Foreign key integrity 100% Index consistency 100%

[0295] Figure 5 It shows the comparison of the dynamic changes in the system load within 24 hours. Compared with the traditional data migration scheme, the present invention shows significant advantages in this embodiment: The traditional scheme adopts a fixed migration strategy and cannot adapt to the change characteristics of different data tables, resulting in data inconsistency often occurring in the migration of frequently changed data, and repeated verification and correction are required. The migration cycle usually takes 6 to 8 months. However, the present invention predicts data change behaviors through deep learning and adopts a differential migration strategy, shortening the migration cycle to 3 months while ensuring that the data consistency reaches 99.999%. The traditional scheme adopts a static allocation method in resource use, often resulting in unbalanced resource utilization. The system load during peak periods exceeds 85%, while the present invention controls the system load below 65% through dynamic resource scheduling. In addition, the traditional scheme lacks self-adaptive ability and requires manual intervention to adjust parameters when the business model changes, while the present invention can automatically adapt to business changes through an online update mechanism, and the prediction accuracy rate always remains above 90%. The implementation of the present invention significantly improves the efficiency and reliability of data migration and provides strong support for the digital transformation of enterprises.

[0296] It should be noted that the detailed explanations of the variables involved in the present invention are shown in Table 11 below.

[0297] Table 11 Variable Explanation Table

[0298]

[0299] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention.

Claims

1. An industrial data online migration method for a productivity middle platform, characterized in that, Including: Collect the historical access records of each data table in the industrial data source, and construct a deep neural network model including a multi-head attention layer and a time series feature extraction layer; convert the historical access records into time series feature vectors and input them into the deep neural network model for model training; use the trained deep neural network model to predict the data change rate of each data table in the industrial data source within the next 24 hours; construct a data dependency matrix; Determine two demarcation points of the data change rate by solving an optimization equation set including an entropy value equation, an information gain equation, a resource consumption equation and a delay impact equation; divide the data tables into different change frequency categories according to the data change rate demarcation points; Based on the data dependency matrix and combined with the predicted data change rate, use an improved topological sorting algorithm to determine the data migration order; adopt batch transmission, incremental migration and two-way real-time synchronization methods for data migration of data tables with different change frequency categories respectively.

2. The industrial data online migration method of the productivity middle platform according to claim 1, characterized in that The step of collecting the historical access records is specifically to collect the data access timestamp, data table identifier, data operation type, data update frequency, business importance level, data table size, data complexity, service response time, data synchronization delay, system throughput, network bandwidth occupancy, storage space requirement, processor load and memory usage by deploying data collection probes.

3. The industrial data online migration method of the productivity middle platform according to claim 1, characterized in that The step of constructing the deep neural network model is specifically to set the number of attention heads to 8, the dimension of each attention head to 64, and use the scaled dot-product attention mechanism to calculate the correlation weights between features; the reference interval of the forgetting gate threshold is from 0.1 to 0.9, and the reference interval of the skip connection weight is from 0.2 to 0.

8.

4. The industrial data online migration method of the productivity middleware according to claim 1, characterized in that, In the step of training the deep neural network model, segment the historical access records by time window, set the window size to 1 hour, and the sliding step size to 5 minutes; adopt the zero-mean normalization method for feature normalization; Set the learning rate to 0.001 and the number of training rounds to 1000.

5. The industrial data online migration method of the productivity middle platform according to claim 1, characterized in that, In the step of predicting the data change rate, conduct a confidence evaluation on the prediction result, set the confidence threshold to 0.85, and smooth the prediction results with a confidence lower than the threshold.

6. The industrial data online migration method of the productivity middle platform according to claim 1, wherein In the step of determining the data change rate demarcation points, set the base number of entropy value calculation to 2 and the normalization coefficient of information gain to 0.5; use the quasi-Newton method to solve the optimization equation set, set the maximum number of iterations to 1000, and the convergence threshold to 0.0001; the reference value of the first demarcation point is 0.1, and the reference value of the second demarcation point is 0.

5.

7. The industrial data online migration method of the productivity middle platform according to claim 1, characterized in that For data tables with a predicted data change rate less than the first demarcation point, use a batch data transmission channel for data migration; for data tables with a predicted data change rate between the two demarcation points, use a remote direct memory access channel for incremental migration; for data tables with a predicted data change rate greater than the second demarcation point, use a two-way real-time synchronization mechanism for data migration.

8. The industrial data online migration method of the productivity middle platform according to claim 1, characterized in that During the data migration process, calculate the real-time change rate of the data table in real time. When the difference between the real-time change rate and the predicted data change rate exceeds 20%, trigger the online update of the deep neural network model.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions that, when running on a computer, are used to execute the industrial data online migration method of a productivity middle platform according to any one of claims 1-8.

10. An industrial data online migration system for a productivity middle platform, characterized in that, It includes the computer-readable storage medium according to claim 9. The system is any one of a computer, a server, and a single-chip microcomputer. The computer-readable storage medium is disposed within the system, and a microprocessor for executing the program instructions stored in the computer-readable storage medium is disposed within the system.

Citation Information

Cited By

  • Variable address offset management method of flight simulator

    CN120560895A

  • A variable address offset management method for a flight simulator

    CN120560895B

  • Industrial data migration dynamic priority adjustment method for productivity platform

    CN121434188A

  • An industrial data migration state priority adjustment method of a productivity middle platform

    CN121434188B

  • Urban population migration prediction method and device based on power consumption data, and medium

    CN121599241A