Object storage data migration method and system based on large model

By using large-model-driven intelligent decision-making and dynamic adaptive migration execution, the problems of intelligence and efficiency in object storage data migration are solved, achieving efficient and low-cost data migration, adapting to complex environments and improving resource utilization.

CN121365050APending Publication Date: 2026-01-20SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511379085.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-25
Publication Date
2026-01-20

AI Technical Summary

Technical Problem

Existing technologies for object storage data migration suffer from low intelligence, insufficient migration efficiency, low resource utilization, and a lack of semantic understanding capabilities, making it difficult to meet the efficient, intelligent, and reliable data migration needs of modern enterprises.

Method used

The system employs a large-model-driven intelligent decision-making module that combines multi-source data acquisition, semantic feature extraction, and dynamic adaptive transfer execution. It generates the optimal transfer scheme through a multi-layer attention mechanism and Monte Carlo tree search, and continuously improves model performance through a closed-loop optimization mechanism.

Benefits of technology

It achieves intelligent, efficient, and low-cost object storage data migration, adapts to dynamic network environments and business needs, and improves anomaly recovery efficiency and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121365050A_ABST
    Figure CN121365050A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of large models, in particular to an object storage data migration method and system based on a large model, and the method comprises the steps of migration preparation, intelligent decision making, execution and closed-loop optimization. The method has the beneficial effects that multi-dimensional information such as storage indexes, network states and data features is collected, deep semantic analysis is carried out by utilizing a fine-tuned large language model, and an optimal migration path and parameter combination are generated. Meanwhile, the fragment size, the concurrent number and the compression algorithm are dynamically adjusted in combination with reinforcement learning so as to adapt to the real-time changing network environment and service requirements. In addition, a block-level breakpoint resuming mechanism is designed, and the exception recovery efficiency is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of large models, in particular to an object storage data migration method and system based on a large model. BACKGROUND

[0002] Object storage has become the core solution for massive unstructured data storage in the cloud computing era. With the explosive growth of data volume and the popularity of multi-cloud strategies, the demand for cross-cloud and cross-region object storage data migration is growing exponentially. However, traditional data migration techniques have significant shortcomings in terms of intelligence, migration efficiency, and exception handling, making it difficult to meet the needs of modern enterprises for efficient, intelligent, and reliable data migration.

[0003] In existing technologies, object storage data migration mainly relies on static migration strategies based on rule engines, incremental synchronization tools, and sharding parallel transmission methods. While these methods have solved basic data migration needs to some extent, they generally have problems such as rigid strategies and low resource utilization. For example, static migration strategies cannot adapt to dynamic changes in network environments and business needs, incremental synchronization tools have high metadata operation overhead when dealing with massive small files, and fixed sharding size parallel transmission performs poorly in weak network environments. More importantly, existing technologies lack semantic understanding of data content, making it impossible to achieve intelligent migration decisions based on business logic.

[0004] In recent years, with the application of machine learning technology, some cloud service providers have attempted to optimize the migration process through prediction models, but these solutions often focus only on single-dimensional indicators and fail to fundamentally solve multi-objective optimization problems. At the same time, large language models have shown strong semantic understanding and decision-making capabilities in natural language processing, but their application in the storage system vertical still faces challenges such as lack of domain knowledge and real-time requirements. In particular, when dealing with data distribution, network topology, and hardware characteristics specific to storage systems, general large models often struggle to be directly applicable.

[0005] Therefore, there is an urgent need in the industry for a new data migration method that can deeply integrate the intelligent decision-making capabilities of large models with object storage professional technology, to break through the bottlenecks of existing technologies in multi-objective optimization, semantic perception, and dynamic adaptation, and to achieve truly intelligent, efficient, and low-cost object storage data migration. SUMMARY

[0006] The purpose of the present application is to provide an object storage data migration method and system based on a large model, to achieve truly intelligent, efficient, and low-cost object storage data migration, and to solve the problems raised in the background technology.

[0007] To achieve the above object, the present application provides the following technical scheme: a large model-based object storage data migration method, comprising:

[0008] Migration preparation stage: start the data perception and understanding module, collect multi-source data, capture the access mode, metadata and business association information of the object to be migrated; process the unstructured data through semantic feature extraction technology, and complete automatic data classification and grading by combining artificial annotation samples with a hybrid classification system;

[0009] Intelligent decision-making stage: the intelligent decision-making module driven by the large model receives the structured feature data and semantic labels of the data perception module, analyzes the data correlation and business requirements based on the pre-trained object storage domain knowledge graph through a multi-layer attention mechanism, completes multi-objective optimization parameter setting, generates candidate strategies through Monte Carlo tree search, and selects the final migration scheme;

[0010] Execution stage: the dynamic adaptive migration execution module starts the migration task according to the decision-making scheme, monitors the network bandwidth, target end storage utilization and data verification results in real time, dynamically adjusts the migration parameters and takes corresponding measures to ensure the smooth progress of the migration;

[0011] Closed-loop optimization stage: after the migration is completed, the model continues to learn and optimize the module to collect full-process performance indicators and user feedback data, which are stored in the feedback database after cleaning for incremental training of the large model, the model decision accuracy is evaluated regularly and the model is fine-tuned as needed to form a closed-loop optimization mechanism.

[0012] Preferably, in the migration preparation stage, the multi-source data collection specifically includes: deploying collection agents on each storage node to capture the access mode of the object to be migrated in real time, such as the time access frequency of e-commerce platform product pictures and the request IP distribution of financial transaction data; collecting the metadata of the object to be migrated, including file size, type and hash value; collecting the business association information of the object to be migrated, such as the data increment features related to manufacturing enterprise production plans.

[0013] Preferably, in the intelligent decision-making stage, the multi-objective optimization parameter setting and the final migration scheme determination are as follows: the large model is based on the pre-trained object storage domain knowledge graph, combined with data correlation and business requirements, to complete multi-objective optimization parameter setting, such as balancing efficiency, cost and business continuity in manufacturing enterprise migration, setting efficiency to complete core data migration within 4 hours, cost to bandwidth fee less than 5000 yuan, and business continuity to production system IO occupancy rate < 15%; generate candidate strategies through Monte Carlo tree search, and determine the final scheme through Pareto optimal solution screening, such as selecting 2-4 a.m. migration for multinational enterprises, using the transmission strategy of "high-priority data first + dynamic sharding", and selecting Asia-Pacific region compliant storage nodes as the target end.

[0014] Preferably, in the execution stage, dynamically adjusting the migration parameters and taking corresponding measures specifically include: the real-time monitoring module collects network bandwidth, target end storage utilization and data verification results through probes; when a weak network environment is detected, the system automatically reduces the number of shards, adjusts the shard size, and triggers the breakpoint resume transmission mechanism; if a storage node fails, immediately switch to the standby node and use three-copy redundant data to resume transmission; the elastic resource scheduling platform automatically applies for temporary bandwidth expansion to the cloud service provider to ensure that high-priority data transmission is not blocked.

[0015] Preferably, in the closed-loop optimization stage, continuous learning and optimization of the model specifically include: after migration is completed, the model continuous learning and optimization module automatically collects full-process data, including actual migration time, bandwidth utilization, error rate performance indicators, and user feedback; after cleaning, these data are stored in the feedback database for large model incremental training; the system evaluates the model decision accuracy rate every month, and when it is found that the adaptability of the migration strategy decreases in a specific environment, the model is fine-tuned again with related data to form a closed-loop optimization mechanism of "perception-decision-execution-feedback".

[0016] A system of an object storage data migration method based on a large model includes a data perception and understanding module, an intelligent decision module driven by a large model, a dynamic adaptive migration execution module, and a model continuous learning and optimization module.

[0017] The data perception and understanding module is used for multi-source data collection, semantic feature extraction, and data labeling and classification.

[0018] The intelligent decision module driven by the large model is used for model customization and optimization, migration strategy generation, and multi-objective optimization.

[0019] The dynamic adaptive migration execution module is used for real-time monitoring and feedback, elastic resource scheduling, and exception handling and fault tolerance.

[0020] The model continuous learning and optimization module is used for feedback data collection and model updating and optimization; the modules work cooperatively to achieve efficient and intelligent migration of object storage data.

[0021] Preferably, the data perception and understanding module specifically includes:

[0022] A multi-source data collection unit is used for real-time or periodic collection of access patterns in an object storage system, including access frequency, timestamp; metadata, including file size, type; and business association information, including workload fluctuation.

[0023] A semantic feature extraction unit uses natural language processing and computer vision technology to extract semantic features of unstructured data such as text, images, and videos, and converts text into a high-dimensional vector and uses a convolutional neural network to extract an image feature vector.

[0024] The data labeling and classification unit, combined with business rules and machine learning algorithms, automatically labels and classifies data according to semantic features, labels medical images with "X-ray" categories, and classifies them according to importance and sensitivity.

[0025] Preferably, the large model-driven intelligent decision module specifically includes:

[0026] The model customization and optimization unit, based on the characteristics of object storage, fine-tunes the general large language model with the data distribution and network topology domain knowledge specific to the storage system, improving its understanding and processing ability of storage-related problems.

[0027] The migration strategy generation unit inputs the information from the data perception module into the optimized large model to generate personalized strategies including migration timing, sequence, and target storage location, such as selecting a network low period for migration and prioritizing the migration of critical business data.

[0028] The multi-objective optimization unit constructs an optimization function targeting migration efficiency and cost, and the large model simulates and evaluates different strategies to find the optimal solution that balances all objectives, such as reducing speed to ensure data integrity and reducing cost during weak network conditions.

[0029] Preferably, the dynamic adaptive migration execution module specifically includes:

[0030] The real-time monitoring and feedback unit monitors key network bandwidth indicators in real-time during the migration process and feeds back the monitoring results to the intelligent decision module to adjust the strategy, such as reducing the number of parallel transmission shards when bandwidth decreases.

[0031] The elastic resource scheduling unit dynamically adjusts computing, storage, and network resources based on real-time migration task requirements, allocating more bandwidth when bandwidth demand increases and adjusting some data storage paths when target storage is tight.

[0032] The exception handling and fault tolerance unit establishes an exception detection and handling mechanism, automatically suspends migration, records progress, and attempts to recover when network interruption occurs, and if recovery is not possible, it alerts and provides fault information.

[0033] Preferably, the model continuous learning and optimization module specifically includes:

[0034] The feedback data collection unit collects actual performance indicators during the migration process, including migration time, strategy execution results, and user feedback data.

[0035] The model update and optimization unit continuously trains and optimizes the large model using feedback data, regularly evaluates performance, and adjusts model structure, parameters, or training data based on evaluation results to adapt to new business and technical environments, such as incorporating new storage device information to improve model adaptability.

[0036] Compared with the prior art, the present application has the beneficial effects that:

[0037] The object storage data migration method and system based on a large model proposed by the present application collect multi-dimensional information such as storage indicators, network states and data characteristics, perform deep semantic analysis using a fine-tuned large language model, and generate an optimal migration path and parameter combination. At the same time, the shard size, concurrency number and compression algorithm are dynamically adjusted by combining reinforcement learning to adapt to the real-time changes in the network environment and business requirements. In addition, the present application designs a block-level breakpoint resume mechanism, which significantly improves the abnormal recovery efficiency. BRIEF DESCRIPTION OF DRAWINGS

[0038] Figure 1 The method flowchart of the present application. DETAILED DESCRIPTION

[0039] In order to make the purpose, technical solution of the present application clear, complete description, and the advantages are more clear and obvious, the following will be further described in detail by combining the embodiments of the present application with the drawings. It should be understood that the specific embodiments described here are part of the embodiments of the present application, not all embodiments, and are only used to explain the embodiments of the present application, and do not limit the embodiments of the present application. All other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0040] Embodiment one, the present application provides a technical solution: an object storage data migration method based on a large model, comprising:

[0041] 1) Migration preparation stage: data perception and understanding start

[0042] Firstly, the data perception and understanding module starts the multi-source data collection work comprehensively. The collection agent deployed on each storage node real-time captures the access mode (such as the time-sharing access frequency of e-commerce platform commodity pictures, the request IP distribution of financial transaction data) of the object to be migrated, the metadata (including file size, type, hash value, etc.) and the business associated information (such as the data increment characteristics related to the production plan of manufacturing enterprises).

[0043] Then, the semantic feature extraction technology is used to process unstructured data, such as extracting lesion feature vectors from medical images and performing topic and entity recognition on text type contract files. Subsequently, the hybrid classification system combines with artificial annotation samples to complete automatic classification and grading of data, such as marking the exam preparation videos of educational institutions as high priority and marking the ordinary courseware as medium priority, providing basic data tags for subsequent decision-making.

[0044] 2) Intelligent decision-making stage: large model driven strategy generation

[0045] The large model-driven intelligent decision module receives structured feature data and semantic labels from the data perception module. Based on the pre-trained object storage domain knowledge graph (including storage protocols, network topology, etc. professional knowledge), the fine-tuned large model analyzes data correlation and business needs through multi-layer attention mechanism.

[0046] The model first completes multi-objective optimization parameter setting, such as balancing efficiency (core data migration completed within 4 hours), cost (bandwidth cost less than 5000 yuan), and business continuity (production system IO occupancy rate < 15%) in manufacturing enterprise migration. Then through Monte Carlo tree search to generate candidate strategies, and through Pareto optimal solution screening to determine the final scheme: for example, a multinational enterprise chooses to migrate at 2-4 am (network idle period), uses the transmission strategy of "high-priority data first + dynamic sharding", and selects a compliant storage node in the Asia-Pacific region as the target.

[0047] 3) Execution phase: dynamic adaptive migration and real-time control

[0048] The dynamic adaptive migration execution module starts the migration task according to the decision scheme. Initially, 50 shards are transmitted in parallel, and the real-time monitoring module collects network bandwidth (such as from 100 Mbps to 30 Mbps), target storage utilization (when reaching 85% threshold), and data verification results through probes.

[0049] When a weak network environment is detected, the system automatically reduces the number of shards to 20, adjusts the shard size from 100 MB to 50 MB, and triggers the breakpoint resume mechanism; if a storage node fails, it immediately switches to a backup node and uses three-copy redundant data to resume transmission. At the same time, the elastic resource scheduling platform automatically applies for temporary bandwidth expansion to cloud service providers to ensure that high-priority data transmission is not blocked.

[0050] 4) Closed-loop optimization phase: model continuous learning and iteration

[0051] After migration is completed, the model continuous learning and optimization module automatically collects full-process data: including actual migration time (such as 20 minutes ahead of schedule), bandwidth utilization (average 65%), error rate (0.003%), and user feedback (such as "financial data migration did not affect daytime reconciliation").

[0052] These data are stored in the feedback database after cleaning for large model incremental training. The system evaluates the model decision accuracy every month, and when it finds that the migration strategy adaptability decreases in the 5G environment, it supplements the edge node performance data to re-tune the model, and finally forms a "perception-decision-execution-feedback" closed-loop optimization mechanism to continuously improve the migration efficiency in complex scenarios.

[0053] Embodiment two, on the basis of embodiment one, proposes an object storage data migration system based on large model, comprising:

[0054] 1) Data perception and understanding module

[0055] Multi-source data collection: Real-time or periodic capture of access patterns (access frequency, timestamp, etc.), metadata (file size, type, etc.), and business-related information (workload fluctuations, etc.) in the object storage system.

[0056] Semantic feature extraction: Use natural language processing, computer vision, and other technologies to extract semantic features of unstructured data such as text, images, and videos, such as converting text into high-dimensional vectors and extracting image feature vectors using convolutional neural networks.

[0057] Data labeling and classification: Combine business rules and machine learning algorithms to automatically label and classify data based on semantic features, such as labeling medical images with categories such as "X-ray" and classifying them by importance and sensitivity.

[0058] 2) Intelligent decision-making module driven by large model

[0059] Model customization and optimization: Fine-tune general large language models with domain knowledge such as data distribution and network topology specific to the object storage field to improve their understanding and processing capabilities for storage-related issues.

[0060] Migration strategy generation: Input the information from the data perception module into the optimized large model to generate personalized strategies that include migration timing, sequence, target storage location, etc., such as migrating during network low periods and prioritizing the migration of critical business data.

[0061] Multi-objective optimization: Construct an optimization function targeting migration efficiency, cost, and other objectives, and use the large model to simulate and evaluate different strategies to find the optimal solution that balances all objectives, such as reducing speed to preserve data integrity and reduce costs during weak network periods.

[0062] 3) Dynamic adaptive migration execution module

[0063] Real-time monitoring and feedback: Monitor key indicators such as network bandwidth in real-time during migration and feed them back to the intelligent decision-making module to adjust strategies, such as reducing the number of parallel transmission shards when bandwidth decreases.

[0064] Elastic resource scheduling: Dynamically adjust computing, storage, and network resources based on real-time migration task requirements, such as allocating more bandwidth when bandwidth demand increases and adjusting some data storage paths when target storage is tight.

[0065] Exception handling and fault tolerance: Establish an exception detection and handling mechanism that automatically suspends migration, records progress, and attempts to recover when encountering network interruptions or other exceptions, and alerts and provides fault information if recovery is not possible.

[0066] 4) Model continuous learning and optimization module

[0067] Feedback data collection: Collect data such as actual performance indicators in migration (migration time, etc.), policy execution results, and user feedback.

[0068] Model updating and optimization: Continuously train and optimize the large model with feedback data, regularly evaluate performance, and adjust model structure, parameters, or training data to adapt to new business and technical environments, such as incorporating new storage device information to improve model adaptability.

[0069] Although embodiments of the present application have been shown and described, it will be understood by those having ordinary skill in the art that various changes, modifications, substitutions and alterations can be made therein without departing from the principles and spirit of the application, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A large model-based object storage data migration method, characterized in that: Comprise: Migration preparation stage: start data perception and understanding module, collect multi-source data, capture access mode, metadata and business association information of objects to be migrated; For unstructured data, semantic feature extraction technology is used to process the data, and a hybrid classification system is used to complete automatic classification and grading of the data based on manually annotated samples; Intelligent decision-making stage: the intelligent decision-making module driven by the large model receives the structured feature data and semantic labels from the data perception module, analyzes the data correlation and business requirements based on the pre-trained object storage domain knowledge graph, sets multi-objective optimization parameters, generates candidate strategies through Monte Carlo tree search, and selects the final migration scheme; Execution stage: the dynamic adaptive migration execution module starts the migration task according to the decision scheme, monitors network bandwidth, target storage utilization and data verification results in real time, dynamically adjusts migration parameters and takes corresponding measures to ensure smooth migration; Closed-loop optimization stage: after migration, the model continues to learn and optimize the module to collect performance indicators and user feedback data throughout the process, and store them in the feedback database for incremental training of the large model. The model decision accuracy is evaluated regularly and the model is fine-tuned as needed to form a closed-loop optimization mechanism.

2. The object storage data migration method based on a large model according to claim 1, characterized in that: In the migration preparation stage, multi-source data collection specifically includes: deployment of collection agents on each storage node to capture access patterns of objects to be migrated in real time, such as time-based access frequency of e-commerce platform product pictures and request IP distribution of financial transaction data; collect metadata of objects to be migrated, including file size, type and hash value; collect business association information of objects to be migrated, such as data increment features related to manufacturing enterprise production plan.

3. The object storage data migration method based on a large model according to claim 2, characterized in that: In the intelligent decision-making stage, multi-objective optimization parameter setting and final migration scheme determination are as follows: the large model based on the pre-trained object storage domain knowledge graph, combined with data correlation and business requirements, completes multi-objective optimization parameter setting, such as balancing efficiency, cost and business continuity in manufacturing enterprise migration, setting efficiency to complete core data migration within 4 hours, cost to bandwidth cost less than 5000 yuan, and business continuity to production system IO occupancy rate < 15%; generate candidate strategies through Monte Carlo tree search, and determine the final scheme through Pareto optimal solution screening, such as selecting 2-4 a.m. for cross-border enterprise migration, using "high priority data first + dynamic sharding" transmission strategy, and selecting Asia-Pacific region compliant storage nodes as target.

4. The object storage data migration method based on a large model according to claim 3, characterized in that: In the execution stage, dynamic adjustment of migration parameters and corresponding measures are as follows: the real-time monitoring module collects network bandwidth, target storage utilization and data verification results through probes; when weak network environment is detected, the system automatically reduces the number of shards, adjusts the size of the shards, and triggers the breakpoint resume mechanism; if a storage node fails, it immediately switches to a backup node and uses three-copy redundant data to resume transmission; the elastic resource scheduling platform automatically applies for temporary bandwidth expansion to cloud service providers to ensure that high-priority data transmission is not blocked.

5. The large model-based object storage data migration method of claim 4, wherein: In the closed-loop optimization phase, the model continuous learning and optimization is specifically: after migration, the model continuous learning and optimization module automatically collects the whole process data, including actual migration time, bandwidth utilization, error rate performance indicators, and user feedback; these data are stored in the feedback database after cleaning for large model incremental training; the system evaluates the model decision accuracy rate every month, and when it finds that the adaptability of the migration strategy in a specific environment decreases, it supplements related data to fine-tune the model, forming a closed-loop optimization mechanism of "perception-decision-execution-feedback".

6. The system for large model-based object storage data migration method according to claim 5, characterized in that: It includes a data perception and understanding module, a large model driven intelligent decision module, a dynamic adaptive migration execution module, and a model continuous learning and optimization module. The data perception and understanding module is used for multi-source data collection, semantic feature extraction, and data labeling and classification. The large model driven intelligent decision module is used for model customization and optimization, migration strategy generation, and multi-objective optimization. The dynamic adaptive migration execution module is used for real-time monitoring and feedback, elastic resource scheduling, and exception handling and fault tolerance. The model continuous learning and optimization module is used for feedback data collection and model updating and optimization; each module works together to achieve efficient and intelligent migration of object storage data.

7. The system of claim 6, wherein the method further comprises: The data perception and understanding module specifically includes: A multi-source data collection unit is used to collect access patterns in the object storage system in real time or periodically, including access frequency, timestamp; metadata, including file size, type; and business-related information, including workload fluctuations; A semantic feature extraction unit uses natural language processing and computer vision technology to extract semantic features of unstructured data such as text, images, and videos, converts text into high-dimensional vectors, and extracts image feature vectors using convolutional neural networks; A data labeling and classification unit automatically labels and classifies data based on semantic features, labels medical images as "X-ray" categories, and classifies them by importance and sensitivity.

8. The system of claim 7, wherein: The large model driven intelligent decision module specifically includes: A model customization and optimization unit fine-tunes a general large language model using data distribution and network topology domain knowledge specific to the storage system to improve its understanding and processing capabilities for storage-related issues; A migration strategy generation unit inputs the information from the data perception module into the optimized large model to generate personalized strategies including migration timing, sequence, and target storage location, such as selecting a network low period for migration and prioritizing the migration of critical business data; A multi-objective optimization unit constructs an optimization function targeting migration efficiency and cost, and the large model simulates and evaluates different strategies to find the optimal solution that balances all objectives, such as reducing speed to ensure data integrity and reduce costs in weak networks.

9. The system of claim 8, wherein: The dynamic adaptive migration execution module specifically includes: A real-time monitoring and feedback unit monitors network bandwidth key indicators in real time during the migration process and feeds back the monitoring results to the intelligent decision module to adjust the strategy, such as reducing the number of parallel transmission shards when bandwidth decreases; The elastic resource scheduling unit dynamically adjusts the computing, storage and network resources according to the real-time requirements of the migration task, allocates more bandwidth when the bandwidth demand increases, and adjusts part of the data storage path when the target end storage is tight. The abnormality processing and fault tolerance unit establishes an abnormality detection and processing mechanism, automatically suspends migration when a network interruption abnormality is encountered, records the progress and attempts to recover, and if recovery is not possible, an alarm is given and fault information is provided.

10. The system of claim 9, wherein: The model continuous learning and optimization module specifically includes: A feedback data collection unit collects actual performance indicators during the migration process, including migration time, strategy execution results and user feedback data. A model updating and optimization unit continuously trains and optimizes the large model using feedback data, regularly evaluates performance, and adjusts the model structure, parameters or training data according to the evaluation results to adapt to new business and technical environments, such as incorporating new storage device information to improve model adaptability.