Data compliance migration method
By monitoring regulatory changes in real time and dynamically assessing migration paths, and by employing fragmented transmission and multi-channel synchronization technologies, the problem of insufficient regulatory adaptability in cross-border e-commerce data migration has been solved, achieving efficient and secure data compliance migration and business continuity.
Patent Information
- Application Number
- CN202510942757.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-11-07
AI Technical Summary
Existing data protection solutions are not adaptable enough to dynamic changes in regulations, resulting in business interruption risks and security vulnerabilities during the migration of encrypted data in cross-border e-commerce, making it difficult to achieve efficient cross-border transmission and real-time risk assessment and scheduling.
The regulatory monitoring module scans for changes in data protection regulations in real time, triggers data classification and risk level assessment, dynamically evaluates migration paths and storage nodes, and uses fragmented transmission and multi-channel synchronization technologies to monitor the migration process in real time and adjust the migration progress and service switching timing.
It has enabled the automation and intelligentization of cross-border data compliance migration, ensuring data security compliance and business continuity, and improving transmission efficiency and data integrity.
Smart Images

Figure CN120910016A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of information technology, and in particular to a data compliance migration method. BACKGROUND
[0002] The current mainstream data protection scheme mainly relies on static regional deployment and traditional encryption technology, and these schemes show obvious lack of adaptability when facing dynamic changes in regulations. Existing systems often use fixed data storage architecture, lack the ability to adjust in real time according to changes in regulations, and there are risks of business interruption and security vulnerabilities in the process of data migration.
[0003] The fundamental challenge faced by cross-border e-commerce encrypted data processing stems from the contradiction between the dynamic nature of the regulatory environment and the static nature of the data processing architecture. When the data protection regulations of a certain region change, the enterprise needs to quickly adjust the geographical location of data processing to maintain compliance, but this adjustment will inevitably trigger the need for large-scale cross-border migration of encrypted data. During the data migration process, the mixed storage state of sensitive data and non-sensitive data makes the traditional batch migration method face huge security risks, especially the complexity of encryption key management and the difficulty of data integrity verification that may occur during data transmission. More complex is the huge difference in data sovereignty requirements between different countries, which requires encrypted data to meet the dual compliance standards of the source region and the target region during migration, while existing encryption transmission technologies often sacrifice transmission efficiency in the name of security, resulting in serious impact on business continuity.
[0004] Therefore, how to build a dynamic processing architecture that can automatically migrate encrypted data according to global regulatory changes, ensure data security compliance, and achieve efficient cross-border transmission and real-time risk assessment scheduling, has become a key problem that needs to be broken through in the field of cross-border e-commerce. SUMMARY
[0005] The present application provides a data compliance migration method, mainly comprising: The data protection regulation change information is obtained by the regulation monitoring module, key clauses and geographic location parameters are extracted, and a compliance parameter set is generated; data classification identification is triggered according to the regulation change information, the content characteristics and sensitive level of the stored data are analyzed, and the risk classification of the data block is determined; the data block is allocated to different transmission queues according to the risk classification, and a data migration task list is generated; the network delay and storage capacity indicators of the target geographic location are analyzed by the evaluation module, the comprehensive score of the candidate node is calculated, the optimal migration path and target storage node are determined; the data migration time window and transmission amount are dynamically scheduled according to the risk classification and the optimal migration path; the parallel data transmission process is started, and the data migration operation is performed; the integrity of the transmitted data is checked by the verification algorithm, and the transmission state is confirmed; the service response time in the migration process is detected by the monitoring module, and the migration progress and switching time are adjusted.
[0006] Further, the data protection regulation change information is obtained by the regulation monitoring module, key clauses and geographic location parameters are extracted, and a compliance parameter set is generated, including: obtaining data protection related texts from a global regulation database at regular intervals to form a regulation original data set; the regulation original data set is processed using named entity recognition technology to extract the effective time and geographic location information in the clauses, and a structured clause data set is generated; according to the information in the structured clause data set, the pre-established country code table is associated to determine the compliance parameter range; the compliance parameter range is matched with the preset compliance requirement table in the field, and the classified compliance parameter set is output for subsequent data classification and migration processing.
[0007] Further, the data classification identification is triggered according to the regulation change information, the content characteristics and sensitive level of the stored data are analyzed, and the risk classification of the data block is determined, including: obtaining the change clause number according to the regulation change signal analysis result, and starting a hybrid storage data processing flow; the content characteristic data set is generated using the processing flow, input into an optical character recognition component, and regular matching is performed with a pre-set sensitive keyword list to obtain a label identification result of each data block; if the label identification result hits a specific keyword, it is marked as a high sensitive category; the corresponding code identifier is determined by querying the pre-established compliance clause mapping table through the high sensitive category data, and is transmitted to the downstream archiving module.
[0008] Further, the data block is allocated to different transmission queues according to the risk classification, and a data migration task list is generated, including: obtaining a high-risk and low-risk data block set from the initial classification result, sorting the high-risk data block by using a weighted scoring rule to obtain a scoring sequence; if the data block in the scoring sequence exceeds a preset threshold, it is inserted into the tail of the encrypted transmission queue and compared with a preset hash mapping table to determine the abnormal data block; for the low-risk data block set, a batch transmission verification tool is used for processing to obtain a verification result; according to the verification result, the batch distribution is adjusted, the partition load is dynamically balanced, and it is judged whether the load difference meets the preset condition.
[0009] Further, the network delay and storage capacity indicators of the target geographic location are analyzed by the evaluation module, the comprehensive score of the candidate node is calculated, and the optimal migration path and target storage node are determined, including: obtaining network delay data of the target geographic location, comparing with a preset delay threshold, if the threshold is exceeded, marking as a high-delay location, obtaining a high-delay location list; according to the list, obtain the storage capacity indicator data of the corresponding node, use a distributed computing framework to classify, if the capacity is lower than the preset standard, mark as a low-capacity node; obtain the regulatory compliance indicator data from the node meta database, match through a preset compliance verification table, if the match fails, mark as a non-compliant node; filter the standby nodes in the same region that meet the requirements as candidate nodes, generate a priority sequence by using a weighted scoring mechanism, calculate the optimal migration path and target storage node set.
[0010] Further, the data migration time window and transmission volume are dynamically scheduled according to the risk classification and the optimal migration path, including: obtaining migration path data in the time window allocation scheme, converting to a specified format, inputting into a path calculation tool, obtaining transmission delay and bandwidth occupancy; if the transmission delay is lower than a preset threshold and the bandwidth occupancy is higher than a predetermined proportion, the path is added to a priority migration list; extracting the path identifier from the list, associating with the infrastructure management database to obtain processor and memory parameters; using a clustering algorithm to divide levels according to risk values, adjusting the time window for the high-risk data set; according to the resource identifier of the optimized window, obtaining network performance indicators, using a weighted round-robin algorithm to allocate transmission volume, generating a transmission configuration table.
[0011] Further, the starting parallel data transmission process and executing data migration operation comprises: decomposing tasks through a task allocation scheme analysis tool, obtaining a task unit list with weights, and determining priority groups according to the weight values by using a clustering algorithm; initializing container configuration according to the priority groups, and adjusting container parameters; obtaining data block information from a distributed file system according to the container parameters, and splitting into sub-unit groups; scheduling transmission tasks according to the sub-unit groups and a channel allocation scheme, and obtaining task execution results; managing multi-channel synchronous transmission through a parallel transmission controller, and ensuring that data blocks are migrated to target storage nodes in priority order.
[0012] Further, the integrity check of the transmission data through the check algorithm and the confirmation of the transmission state comprise: calculating a check value by using a hash algorithm for a data segment in a transmission process, obtaining a check identifier and storing the check identifier into a cache database; comparing the check value after transmission with an original check value through a comparison tool, outputting a difference report, and determining the transmission state; if the difference report is empty, it is determined that the transmission state is normal; according to the transmission state and the priority score, the bandwidth allocation weight of the transmission channel is adjusted, a strategy configuration file is generated, and the subsequent transmission process is optimized, so that the balance between data integrity verification and transmission efficiency is ensured.
[0013] The technical scheme provided by the embodiment of the application can include the following beneficial effects: The application discloses a cross-country data compliance migration method, which comprises the following steps: a regulation monitoring module is used to scan country data protection regulation change information in real time, a data classification identification algorithm is triggered to analyze the sensitivity of data in mixed storage and divide the risk level, an optimal migration path and a storage node are dynamically evaluated according to the network environment and regulation requirements of a target geographical position, a migration scheduling scheme is formulated according to the principle of high-risk data priority, and in the data transmission process, the transmission efficiency is improved by using the sharding transmission and multi-channel synchronous technology, and the data integrity is ensured by using the hash check algorithm. The application also monitors the business influence in the migration process in real time, adjusts the migration progress and service switching time in a timely manner, effectively balances the data compliance and business continuity requirements, and provides a kind of automatic and intelligent data compliance migration solution for cross-country enterprises. BRIEF DESCRIPTION OF DRAWINGS
[0014] Figure 1 A flowchart of a data compliance migration method of the application.
[0015] Figure 2 A schematic diagram of a data compliance migration method of the application.
[0016] Figure 3 Another schematic diagram of a data compliance migration method of the application.
[0017] Figure 4Another schematic diagram of the data compliance migration method of the present application. DETAILED DESCRIPTION
[0018] For a further understanding of the present application, reference will be made to the following description of embodiments taken in conjunction with the accompanying drawings. The following detailed description of the application is provided as an example to explain the present application and is not intended to limit the present application. It should also be noted that only parts related to the present application are shown in the drawings for the convenience of description.
[0019] As Figures 1-4 The data compliance migration method of the present application can specifically include: Step S101, using a regulation monitoring module to continuously scan the data protection regulation change information of each country, obtaining a regulation change event trigger signal and a corresponding geographical location compliance requirement parameter by analyzing the key clauses and effective time in the regulation text.
[0020] The regulation monitoring module obtains data protection related text from a global regulation database at regular intervals to obtain a regulation original data set. The named entity recognition technology is used to process the regulation original data set to extract the effective time and geographical location entities in the clauses to generate a structured clause data set. According to the effective time and geographical location information in the structured clause data set, the country code table established in advance is associated to determine the compliance parameter range. The field matching is performed on the compliance parameter range and the preset compliance requirement table to output the classified compliance parameter set for subsequent processing.
[0021] For example, in the design of the regulation monitoring module, data protection related updates can be grabbed from the global regulation database every morning through a timing task. Assuming that the original text data set of a new regulation of the European Union is grabbed, which contains 100,000 words of regulation content. Natural language processing technology can be used for segmentation processing to split the text into clause units, extract core clauses such as "data processing needs to obtain explicit consent of the user" and effective time such as "January 1, 2024", and form a structured data set. This process relies on word segmentation and semantic analysis to ensure accurate extraction of clauses, reduce manual intervention, and improve efficiency.
[0022] For example, when detecting compliance requirements, for the structured clause data set, preset keyword matching rules such as "data protection" and "user consent" are matched. If these keywords appear in the core clauses, a regulation change signal is triggered. Assuming that the European Union regulation clearly requires "data storage not to exceed 5 years", the system identifies and marks it as a potential change signal. This mechanism can quickly screen key information, reduce the risk of omission, and ensure the timeliness of compliance monitoring.
[0023] For example, combined with the effective time and geographical location data, the system can further analyze the impact range of the change signal. Assuming that the aforementioned EU regulation applies to the entire EU, after extracting the geographical location data, the system associates the compliance requirements with specific countries through mapping rules to form a classified compliance parameter set, such as "Germany: data storage limit for 5 years". This association facilitates targeted adjustment of compliance strategies by enterprises, reducing unnecessary resource waste.
[0024] For example, when comparing historical regulation data, if it is found that the original German regulation allows storage for 10 years, and the new regulation shortens it to 5 years, the system determines that there is a dynamic update requirement and generates an update marker. This comparative analysis can help enterprises identify regulation change trends and plan compliance adjustments in advance to reduce the risk of violations.
[0025] For example, in the automated push mechanism, the system transmits the compliance parameter set and update marker to the downstream processing module, such as the enterprise's internal compliance management system, through the interface. Assuming that 10 updates are pushed daily, it ensures real-time information sharing. This automated approach reduces human intervention, improves information transmission efficiency, and supports enterprises to quickly respond to regulation changes. Through the above multi-link cooperation, the system forms a closed loop from data capture to information distribution, ensuring the comprehensiveness and accuracy of regulation monitoring, ultimately providing efficient compliance support for enterprises and reducing operational risks.
[0026] Step S102, according to the regulation change event trigger signal, start the data classification identification algorithm to scan and analyze the data in the mixed storage, through data content feature matching and label identification mechanism, judge the sensitive level attribute and data type identification of each data block.
[0027] According to the signal analysis result triggered by the regulation change, the change clause number is obtained, and the mixed storage data processing flow is started. The content feature data set generated by the processing flow is input into the optical character recognition component, and regular matching is performed with the preset sensitive keyword list to obtain the label identification result of each data block. If the label identification result hits a specific keyword, it is marked as a high-sensitive category. Through the high-sensitive category data query the pre-established compliance clause mapping table, determine the corresponding code identification, and transmit to the downstream archiving module.
[0028] For example, when starting the processing flow of mixed storage data through the signal triggered by the regulation change, it can be assumed that an enterprise's internal storage environment contains multiple data types, such as user personal information and transaction records. Assuming that a new EU regulation requires limiting the storage time limit for personal information, the system receives the change signal and immediately starts the processing flow. When initially dividing the storage environment, it can be divided into local servers and cloud storage based on data sources and storage locations to form a classified data area set.
[0029] For example, when activating the scanning analysis mechanism, the content feature extraction of each data block in the region can identify whether the data block contains specific fields such as name and ID number through preset rules. Assuming that there are 1000 data blocks in the local server region, and after scanning, 300 data blocks containing user identity information are extracted, forming a preliminary content feature data set. This process relies on layer-by-layer parsing of data content to ensure the comprehensiveness of feature extraction.
[0030] For example, when combining the extracted content feature data set with label recognition technology, the data block can be compared with the sensitive information label through preset matching rules. Assuming that the rule is set to mark the data block containing the ID number as a sensitive label, and after comparison, 200 of the 300 data blocks are marked as sensitive categories. This identification method helps quickly filter out data that needs to be focused on.
[0031] For example, after determining the high-sensitive data block set, if the label recognition result meets the sensitive level standard, the high-sensitive category can be further screened out. Assuming that among the 200 sensitive data blocks, 150 involve core user privacy, the system marks them as high-sensitive categories, forming a high-sensitive data block set. This hierarchical method facilitates subsequent targeted processing.
[0032] For example, when obtaining data type information and subdividing the high-sensitive data block set, a classification model can be used to classify data types into two categories: personal identity data and financial data. Assuming that among the 150 high-sensitive data blocks, 100 belong to personal identity data and 50 belong to financial data, the refined classification results provide a basis for subsequent compliance processing.
[0033] For example, when generating identification information in combination with attribute determination rules, if the data type classification result matches the compliance requirements, such as personal identity data requiring encrypted storage, the system will generate encrypted identification for the 100 data blocks. This identification generation process ensures that data processing complies with regulations.
[0034] For example, when the generated identification information is transmitted to the downstream module through an automated distribution mechanism, assuming that 100 identification information is transmitted to the enterprise internal archiving system every day, forming the final classification and archiving record. This automated method ensures the timeliness of information transmission, helping enterprises quickly adjust storage strategies and reduce the risk of violations.
[0035] Step S103, if the data block is judged to be sensitive data, it is marked as high-risk level and assigned to the encrypted transmission queue, if the data block is non-sensitive data, it is marked as low-risk level and assigned to the ordinary transmission queue, obtaining the data migration task list classified by risk level.
[0036] The high-risk and low-risk data block sets are obtained from the initial classification results, the high-risk data blocks are sorted using a weighted scoring rule to obtain a scoring sequence. If the data blocks in the scoring sequence exceed a preset threshold, they are inserted into the encrypted transmission queue tail and compared with a preset hash mapping table to determine the abnormal data blocks. For the low-risk data block set, a batch transmission verification tool is used for processing to obtain a verification result. According to the verification result, the batch distribution is adjusted, the partition load is dynamically balanced, and it is judged whether the load difference meets the preset condition.
[0037] For example, in processing the initial classification results of the data blocks, the division of high-risk and low-risk categories can be started. For the priority sorting of high-risk category data blocks, it is assumed that there are 500 data blocks in an enterprise internal storage system, of which 200 are marked as high-risk category. The system sorts these 200 data blocks into a sequence according to the sensitive attributes of the data blocks, such as data blocks involving user core information, with higher priority, and the first 50 data blocks are placed at the front due to involving critical privacy.
[0038] For example, when activating the scheduling mechanism of the encrypted transmission queue for the sorted high-risk data block sequence, it is assumed that of the 200 data blocks, 180 meet the transmission conditions, such as moderate data block size and not being marked as damaged. These data blocks are assigned to different positions in the encrypted transmission queue to form a to-be-processed list, with the first 50 highest-priority data blocks arranged at the head of the queue to ensure priority transmission.
[0039] For example, in verifying the identification information of the data blocks in the encrypted transmission queue, it is assumed that through the mapping table, it is found that 5 of the 180 data blocks have mismatched identification information, which are marked as abnormal data blocks. Tracking processing finds that 3 of these data blocks come from unauthorized external access points, which do not meet the security standards, and are therefore removed, leaving 177 data blocks in the updated queue list.
[0040] For example, for the extraction of metadata information of the updated queue list, it is assumed that the metadata of the 177 data blocks, such as storage time and access frequency, are analyzed using a support vector machine model, and it is found that the characteristics of 160 of these data blocks meet the transmission requirements. These data blocks are confirmed to have the highest transmission priority, with the first 50 maintaining the highest priority to form the final list.
[0041] For example, for the low-risk category data block set in the ordinary transmission queue, it is assumed that of the remaining 300 low-risk data blocks, 250 meet the batch transmission conditions, such as small data block size. The system assigns them to 5 batches of the ordinary transmission queue through comparative analysis tools, with 50 data blocks in each batch, forming a batch processing list.
[0042] For example, when optimizing the distribution information of the general transmission queue batch processing list, assuming that the initial distribution shows that a batch load is too high, the system adjusts according to the scheduling rules, and reassigns part of the data blocks to other batches. Finally, the load of the 5 batches is balanced, meeting the transmission standard, and completing the generation of the final transmission task list.
[0043] For example, the processing mode of each link described above.
[0044] It can be understood that priority sorting and queue scheduling ensure timely processing of high-risk data blocks, removal of abnormal data blocks ensures transmission safety, and load balancing optimization improves overall transmission efficiency. These measures jointly support the compliance and stability of enterprise internal data storage and transmission.
[0045] In step S104, the real-time evaluation module analyzes the network delay, storage capacity and regulatory compliance indicators of the target geographic location, calculates the comprehensive score value of each candidate migration node, and determines the optimal migration path and target storage node.
[0046] Obtain the network delay data of the target geographic location, compare it with the preset delay threshold, if the delay data exceeds the threshold, mark it as a high delay location, and get a list of high delay locations. According to the list of high delay locations, obtain the storage capacity indicator data of the corresponding nodes, use a distributed computing framework to classify the storage capacity data, if the storage capacity is lower than the preset standard, mark it as a low capacity node, and determine a set of low capacity nodes. For the set of low capacity nodes, obtain the regulatory compliance indicator data from the node metadata database, match the indicator data with the pre-established compliance verification table, if the matching fails, mark it as a non-compliant node, and get a list of non-compliant nodes. According to the list of non-compliant nodes, filter the standby nodes in the same region that meet the capacity and delay requirements as candidate nodes, use a weighted scoring mechanism to evaluate the candidate nodes, and generate a priority sequence. For the priority sequence, obtain the inter-node transmission data from the network topology library, use the shortest path algorithm to calculate the preferred migration path, filter the paths that meet the delay requirements, and get a list of preferred migration paths. According to the list of preferred migration paths, obtain the real-time load data of the target nodes from the resource monitoring platform, use a polling algorithm to distribute the data to be migrated to the nodes that meet the load requirements, and determine the final target storage node set.
[0047] For example, when collecting network delay data of the target geographic location through the real-time evaluation tool.
[0048] It can be understood that network latency directly affects the efficiency of data transmission. Assuming that an enterprise's internal data storage system covers multiple geographic locations, the system will collect real-time network latency data from each location. For example, if the latency data for a certain location is 200 milliseconds and the preset threshold is 150 milliseconds, that location is marked as a high-latency location. This marking method helps quickly identify areas that may affect transmission efficiency.
[0049] For example, when obtaining storage capacity indicator information for the high-latency location list, the data can be classified using a batch processing tool. Assuming that the storage capacity of a certain high-latency location is 2TB and the preset standard is 5TB, it is marked as a low-capacity node. This classification method helps the system identify nodes with insufficient resources, providing a basis for subsequent optimization.
[0050] For example, when obtaining regulatory compliance indicator data for the low-capacity node set, the system will match it through a compliance verification table. Assuming that the data storage method of a certain node does not meet the enterprise's internal privacy protection standards, such as lacking necessary access permission control, it is marked as a non-compliant node. This verification mechanism ensures that data storage complies with relevant requirements.
[0051] For example, when using a comprehensive scoring calculation model to weight the candidate nodes for the non-compliant node list, the scoring results can be classified using a support vector machine model. Assuming there are 10 candidate nodes, the system scores them based on indicators such as latency, capacity, and compliance, and the top 3 nodes with the highest scores are prioritized in the sequence. This classification method helps to select the most suitable nodes.
[0052] For example, when obtaining migration path-related data for the priority sequence, the path optimization tool analyzes the transmission efficiency of each path. Assuming that the estimated transmission time for a certain path is 2 hours, which meets the preset standard, it is included in the preferred migration path list. This analysis ensures the efficiency of data migration.
[0053] For example, when obtaining target node distribution information for the preferred migration path list, the scheduling and distribution mechanism adjusts the distribution. Assuming that a certain node in the initial distribution has a high load, the system distributes part of its tasks to other nodes to achieve load balancing standards. This adjustment method helps improve the stability of the overall system.
[0054] For example, when obtaining node state monitoring data for the final target storage node set, the state verification tool compares the data in real time. Assuming that the state of a certain node is normal and meets the stability requirements, it is confirmed for task allocation, and finally forms a migration task execution list. This monitoring mechanism ensures the reliability of task execution.
[0055] In step S105, according to the optimal migration path and risk level classification result, the dynamic scheduling algorithm allocates the migration time window according to the high risk level data priority principle, and determines the data transmission volume and concurrent connection number of each time period through the load balancing mechanism.
[0056] The migration path data in the time window allocation scheme is obtained, the path data is converted into a specified format, input into a path calculation tool, and the transmission delay and bandwidth occupation rate of each path are obtained. If the transmission delay is lower than a preset threshold and the bandwidth occupation rate is higher than a predetermined proportion, the path is added to a priority migration list. The path identifier is extracted from the priority migration list, and the corresponding processor core number, memory occupation rate and disk throughput parameters are obtained by associating with an infrastructure management database. The parameters are divided into multiple levels according to the risk value by using a clustering algorithm, and a high-risk data set is obtained. For each entry in the high-risk data set, the processor core number, memory occupation rate and disk throughput data are extracted, and a priority scheduling algorithm is used to re-allocate the time window. If the processor utilization of the adjusted time window reaches a preset upper limit and the memory occupation is lower than a predetermined standard, the time window is recorded as an optimized window and a resource identifier is attached. According to the resource identifier of the optimized window, the network performance indicators of the corresponding time period are obtained from a monitoring system. A weighted round robin algorithm is used to allocate the transmission volume according to the resource remaining proportion of the optimized window, and a transmission configuration table is generated. The historical connection data of each window is extracted from the transmission configuration table to determine the connection pool upper limit. If the connection pool upper limit meets the preset failure rate condition, the final concurrent scheme is written.
[0057] For example, when formulating the time window allocation scheme, the role of the time window can be understood in principle, that is, by reasonably dividing the time period of data migration, resource competition and system overload can be avoided.
[0058] In a possible implementation, for the acquisition of migration path data, the transmission efficiency of multiple paths can be compared by using a path analysis tool. Assuming that the expected transmission time of a path is 3 hours, and the preset standard is within 4 hours, the path is listed in the priority migration path list.
[0059] For example, when obtaining risk level information for the priority migration path list, a classification tool can be used to stratify the interruption or delay risks that the path may face. Assuming that a path is special due to its geographical location, and historical data shows that the interruption probability is 30%, which exceeds the preset threshold of 20%, the path is classified into a high-risk data set. Such stratification helps to identify potential problem paths.
[0060] For example, when processing a high-risk data set, the acquisition of dynamic scheduling parameters can be based on historical traffic data and current network status. Assuming that traffic peaks are expected to occur at 8 pm within a certain time window, the system can adjust the migration task to the low-peak period of 2 am. If the resource utilization rate decreases from 80% to 50% after adjustment, which meets the preset standard, it is determined as the optimized time window distribution.
[0061] For example, when obtaining data transmission indicators for the optimized time window distribution, the load balancing tool can allocate according to the transmission volume of each time period. Assuming that the total transmission volume in a certain time period is 10 TB, the system will evenly distribute it to 5 sub-periods, each with 2 TB, ensuring balanced distribution of transmission pressure.
[0062] For example, in the transmission volume configuration of each time period, when obtaining concurrent connection data, the connection management tool can limit the number of connections. Assuming that the initial number of concurrent connections in a certain time period is 1000, which exceeds the preset threshold of 800, it is limited to 750 to ensure system stability, and the final scheme is determined after meeting the preset standard.
[0063] For example, when obtaining running state data for the final concurrent connection scheme, the state monitoring tool can compare the running situation of each time period in real time. Assuming that the system response time in a certain time period is 100 milliseconds, which does not exceed the preset threshold of 150 milliseconds, it is determined that the task allocation in this time period is completed and included in the final task execution list.
[0064] For example, when obtaining real-time feedback data of the migration path in the final task execution list, the feedback processing tool can analyze the actual transmission efficiency of the path. Assuming that the actual transmission time of a certain path is 2.5 hours, which is better than the expected 3 hours and meets the optimization standard, it is arranged in the execution order in priority. Such analysis helps to improve the execution efficiency of the overall migration task.
[0065] Step S106, obtain the task allocation scheme output by the dynamic scheduling algorithm, start the parallel data transmission process, and use the sharding transmission technology to split the large data block into multiple sub-segments, and execute the data migration operation through the multi-channel synchronous transmission mechanism.
[0066] The task allocation scheme is decomposed by the Apache Mesos scheme analysis tool to obtain a list of task units with weights, and the K-means clustering algorithm is used to determine the high, medium and low priority groups of the task units according to the weight values. According to the priority groups, the Kubernetes parallel transmission controller is used to initialize the container configuration, and the Ansible transmission initialization tool is used to adjust the container parameters. For the container parameters, the data block information is obtained from the HDFS, and the HadoopFileChunkSplitter is used to split the data block into sub-unit groups. According to the sub-unit groups and the channel allocation scheme, the transmission task is scheduled, and the task execution result is obtained.
[0067] For example, when analyzing the allocation scheme generated by dynamic scheduling, the role of the scheme analysis tool can be understood in principle first, that is, by decomposing the task allocation scheme, the complex overall task is divided into manageable units for subsequent classification and processing.
[0068] In one possible implementation, assume that a certain allocation scheme contains 100 migration tasks, which are decomposed into 10 task units by the scheme analysis tool, each containing 10 tasks, for subsequent grouping management.
[0069] For example, when classifying and processing the decomposed task units, they can be grouped according to the priority and resource requirements of the tasks. Assume that a certain task unit involves high-priority data migration, and another unit involves regular data migration, then they are classified into high-priority and ordinary groups respectively, forming a list of classified task groups, which provides a basis for subsequent parallel transmission.
[0070] For example, when initializing the configuration of the task groups using the parallel transmission mechanism, the transmission initialization tool will adapt to the environment according to the characteristics of the groups. Assume that the high-priority group needs higher bandwidth support, and the tool allocates double the resources to it, while the ordinary group is allocated a standard channel to ensure that the initial parameter configuration is reasonable.
[0071] For example, for large data block splitting, the data is split into sub-unit groups by the data sharding tool. Assume that a certain large data block is 50TB, which is split into 5 sub-unit groups of 10TB each. If each sub-unit group meets the preset transmission unit standard, such as single transmission not exceeding 15TB, then the splitting is determined to be complete.
[0072] For example, in the multi-channel flow setting, the channel allocation tool divides the synchronous transmission resources. Assume that there are 10 channels available, the tool allocates 6 channels to the high-priority group and 4 channels to the ordinary group. If the bandwidth allocation meets the preset standard, then the multi-channel flow allocation scheme is determined.
[0073] For example, for synchronous transmission data management, the transmission management tool batch schedules sub-unit groups. Assuming that a batch contains 5 sub-unit groups, the transmission state is stable after scheduling, and there is no delay or interruption, it meets the preset stability standard, forming an execution plan.
[0074] For example, when monitoring the progress of data migration, the progress monitoring tool compares the task execution in real time. Assuming that a task is scheduled to be completed in 5 hours, the actual progress shows that 80% has been completed, which meets the preset completion standard, and the task execution status is determined to be normal.
[0075] For example, for transmission process feedback information processing, the feedback analysis tool analyzes the data. Assuming that the task feedback shows that the transmission efficiency is high, the actual time consumption is 2 hours less than the expected time, and it meets the optimization standard, it is determined that the migration task is completed well and the process is optimized.
[0076] Step S107, for the data integrity verification requirement in the transmission process, a hash check algorithm is used to check the integrity of each transmitted data segment, and if the check value matches, the transmission is confirmed to be successful, and if the check fails, a retransmission mechanism is triggered.
[0077] For data segments in the transmission process, SHA-256 algorithm is used to calculate the hash value for integrity detection, and the check identifier is obtained and stored in the Redis cache database. By comparing the hash value calculated after transmission with the original hash value using the diff tool, a difference report is output to determine the transmission status. If the difference report is empty, it is determined that the transmission status is normal. According to the transmission status and priority score, the bandwidth allocation weight of the TCP transmission channel is adjusted, and a QoS policy configuration file is generated to optimize the transmission process.
[0078] For example, in the field of data transmission, for the integrity detection of data segments in the transmission process, the hash check algorithm is a commonly used means. Its principle is to calculate a unique hash value for each data segment as its identity identifier, which is used for subsequent comparison to determine whether the data has been tampered with or lost during transmission. Assuming that a data block is divided into 10 segments, each with a size of 5TB, the system will generate a hash value for each segment to form a check identifier list for subsequent verification.
[0079] For example, when obtaining the check result in the transmission process, the calculated hash value can be matched with the original hash value one by one through the comparison tool. Assuming that the original hash value of a segment is H1 and the calculated hash value after transmission is H2, if they are consistent, it is considered that the segment transmission status is normal; if they are not consistent, it is marked as an abnormal segment. This way can quickly locate the problem data and ensure the reliability of the transmission.
[0080] For example, for fragments with a normal transmission state, the recording management tool archives their transmission records. Assuming that 8 fragments pass the verification, the system generates a transmission log containing information such as time and path, facilitating subsequent tracing. This archiving method contributes to the standardization of data management.
[0081] For example, for fragments that do not match the verification, the screening tool classifies and labels them. Assuming that 2 fragments fail the verification, the system labels them as pending retransmission and generates a list. This classification and labeling provide clear grounds for subsequent processing.
[0082] For example, when processing fragments that require retransmission, the resource allocation tool adjusts the transmission channel priority. Assuming that the system has 5 channels, the tool allocates 3 channels to the fragments pending retransmission to accelerate processing. This adjustment effectively improves retransmission efficiency.
[0083] For example, for batch scheduling of retransmission tasks, the transmission scheduling tool determines the execution order based on fragment size and channel status. Assuming that a fragment is relatively large, the system prioritizes it for a high-bandwidth channel to ensure task balance. This scheduling method helps optimize resource utilization.
[0084] For example, during the retransmission process, the monitoring tool detects the integrity of data fragments in real time. Assuming that the hash value of a fragment after retransmission matches the original value, the system confirms task completion. This re-comparison mechanism ensures the accuracy of retransmitted data and reduces potential risks.
[0085] In step S108, the service continuity monitoring module detects the service response time and error rate during the migration process in real time, judges the business impact degree according to the preset threshold, and obtains migration progress adjustment instructions and service switching timing determination results.
[0086] The response time and error rate data during the migration process are obtained from the business monitoring module, the original monitoring data set is generated and cleaned, and the cleaned monitoring data set is obtained. According to the cleaned monitoring data set, the monitoring data set is aggregated according to the time window, and compared with the preset threshold table to determine the index list and amplitude that exceed the threshold. For high-priority indicators in the index list, calculate the fluctuation range, generate a key indicator list, and adjust the task execution order table according to the list to optimize the migration process parameters.
[0087] For example, in the field of business monitoring during data migration, the business monitoring module can continuously track key data streams such as service response time and error rate for each indicator collection and analysis during the migration process. Assuming that a migration task involves 100TB of data, the system collects response time and error rate every minute to form a preliminary monitoring data set. This real-time recording method helps to discover potential problems in a timely manner.
[0088] For example, for the hierarchical evaluation of the preliminary monitoring data set, the system can preset the response time threshold to be 500 milliseconds and the error rate threshold to be 0.5%. Through the data comparison tool, the collected indicators are matched with the thresholds one by one. Assuming that the response time of a certain collection is 600 milliseconds and the error rate is 0.8%, it is determined that the current migration process stability state is abnormal. This hierarchical evaluation can quickly locate the problem level.
[0089] For example, in the classification labeling of abnormal fluctuation data, the data screening tool can mark the data with excessive response time and error rate as high-priority problems. Assuming that in a monitoring, the proportion of data with abnormal response time is 30% and the proportion of data with abnormal error rate is 20%, the system will generate a list of key indicators, listing the indicators that need to be prioritized. This classification provides a clear basis for subsequent adjustments.
[0090] For example, for the priority configuration of the key indicator list, the resource scheduling tool can dynamically adjust task allocation. Assuming that there are currently 10 migration tasks, of which 3 tasks involve high-priority problems, the system will allocate more computing resources to these 3 tasks and adjust the execution order to ensure priority processing. This dynamic adjustment helps improve overall migration efficiency.
[0091] For example, in the condition comparison of switching timing, the condition matching tool checks whether the current state meets the preset conditions. Assuming that the switching conditions are an error rate below 0.2% and a response time below 400 milliseconds, if the current data meets the conditions, it is determined that the switching timing has arrived. This determination ensures smooth transition of the migration process.
[0092] For example, for the optimization of subsequent process parameters, a support vector machine algorithm can be used to analyze the correlation between parameters. Assuming that the parameters include bandwidth allocation and task queue length, the system will optimize these parameters based on historical data to form a more efficient execution plan. This optimization process can effectively improve the smoothness of process execution.
[0093] For example, in real-time feedback tracking of the execution plan, the data verification tool continuously monitors the latest data. Assuming that the response time after optimization is stable at 300 milliseconds and the error rate is reduced to 0.1%, the system will confirm that the migration state is normal. This continuous tracking ensures the reliability of the final completion state.
[0094] The above only describes the preferred embodiments of one or more embodiments of the present specification, and does not limit one or more embodiments of the present specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of one or more embodiments of the present specification shall be included in the protection scope of one or more embodiments of the present specification.
Claims
1. A data compliance migration method, characterized by, The method comprises the following steps: Obtain data protection regulation change information through a regulation monitoring module, extract key provisions and geographic location parameters, and generate a compliance parameter set; Trigger data classification identification according to the regulation change information, analyze the content characteristics and sensitivity level of the stored data, and determine the risk classification of the data block; According to the risk classification, the data block is assigned to different transmission queues, and a data migration task list is generated; Through the evaluation module, analyze the network delay and storage capacity indicators of the target geographic location, calculate the comprehensive score of the candidate node, determine the optimal migration path and target storage node; According to the risk classification and the optimal migration path, dynamically schedule the data migration time window and the transmission amount; Start the parallel data transmission process and execute the data migration operation; Through the verification algorithm, the integrity of the transmitted data is checked to confirm the transmission status; Through the monitoring module, detect the service response time during the migration process, and adjust the migration progress and switching time.
2. The data compliance migration method of claim 1, wherein, The method comprises the following steps: Obtain data protection related texts from a global regulation database at regular intervals to form a regulation raw data set; Use named entity recognition technology to process the regulation raw data set, extract the effective time and geographic location information in the provisions, and generate a structured provision data set; According to the information in the structured provision data set, associate the pre-established country code table to determine the compliance parameter range; Match the compliance parameter range with the preset compliance requirement table by field, output the classified compliance parameter set for subsequent data classification and migration processing.
3. The data compliance migration method of claim 1, wherein, The method comprises the following steps: According to the regulation change signal analysis result, obtain the change provision number, and start the mixed storage data processing flow; Use the processing flow to generate a content feature data set, input it into an optical character recognition component, and perform regular matching with a preset sensitive keyword list to obtain the label identification result of each data block; If the label identification result hits a specific keyword, it is marked as a high sensitivity category; Through the high sensitivity category data, query the pre-established compliance provision mapping table to determine the corresponding code identifier and transmit it to the downstream archiving module.
4. The data compliance migration method of claim 1, wherein, The method comprises the following steps: From the initial classification result, obtain the high-risk and low-risk data block set, use a weighted scoring rule to sort the high-risk data block, and obtain a scoring sequence; If the data block in the scoring sequence exceeds the preset threshold, it is inserted into the tail of the encrypted transmission queue and compared with the preset hash mapping table to determine the abnormal data block; For the low-risk data block set, use a batch transmission verification tool to process and obtain a verification result; According to the verification result, adjust the batch distribution, dynamically balance the partition load, and judge whether the load difference meets the preset condition.
5. The data compliance migration method of claim 1, wherein, The network delay and storage capacity indicators of the target geographic location are analyzed by the evaluation module, the comprehensive score of the candidate node is calculated, and the optimal migration path and target storage node are determined, comprising: Obtain network delay data of target geographic location, compare with preset delay threshold, if exceed the threshold, mark as high delay location, get high delay location list; According to the list, obtain the storage capacity indicator data of the corresponding node, and use the distributed computing framework to classify, if the capacity is lower than the preset standard, mark as low capacity node; Get the regulatory compliance indicator data from the node meta database, match through the preset compliance verification table, if the match fails, mark as non-compliant node; Screen the standby nodes in the same region that meet the requirements as candidate nodes, use the weighted scoring mechanism to generate a priority sequence, calculate the optimal migration path and target storage node set.
6. The data compliance migration method of claim 1, wherein, According to the risk classification and the optimal migration path, dynamically schedule the data migration time window and the transmission amount, comprising: Obtain migration path data in the time window allocation scheme, convert to specified format, input into path calculation tool, get transmission delay and bandwidth occupancy rate; If the transmission delay is lower than the preset threshold and the bandwidth occupancy rate is higher than the predetermined proportion, add the path to the priority migration list; Extract the path identifier from the list, associate the infrastructure management database to obtain the processor and memory parameters; Use clustering algorithm to divide levels according to risk value, adjust time window for high-risk data set; According to the resource identifier of the optimized window, obtain the network performance indicators, use the weighted round-robin algorithm to allocate the transmission amount, and generate the transmission configuration table.
7. The data compliance migration method of claim 1, wherein, The parallel data transmission process is started, and the data migration operation is executed, comprising: Decompose the task through the task allocation scheme analysis tool, obtain the weighted task unit list, and use the clustering algorithm to determine the priority grouping according to the weight value; According to the priority grouping, initialize the container configuration and adjust the container parameters; According to the container parameters, obtain the data block information from the distributed file system, and split it into sub-unit groups; According to the sub-unit group and the channel allocation scheme, schedule the transmission task to get the task execution result; Through the parallel transmission controller, manage the multi-channel synchronous transmission to ensure that the data blocks are migrated to the target storage node in priority order.
8. The data compliance migration method of claim 1, wherein, The transmission data is checked for integrity by the verification algorithm, and the transmission state is confirmed, comprising: For data segments in the transmission process, calculate the check value using the hash algorithm, get the check identifier and store it in the cache database; Compare the check value after transmission with the original check value through the comparison tool, output the difference report, and determine the transmission state; If the difference report is empty, it is determined that the transmission state is normal; According to the transmission state and the priority score, adjust the bandwidth allocation weight of the transmission channel, generate the strategy configuration file, optimize the subsequent transmission process, and ensure the balance between data integrity verification and transmission efficiency.
Citation Information
Cited By
Finance and tax data migration storage management method and system based on encrypted transmission
CN121525072A
Tax data migration storage management method and system based on encrypted transmission
CN121525072B