Cold and hot data storage optimization method based on large model

Through a hot and cold data storage optimization system based on large models, the hot and cold data labels and storage strategies are dynamically adjusted, and the problem of resource waste and small file management in traditional solutions is solved, and efficient storage resource utilization and performance improvement is achieved.

CN120596490AActive Publication Date: 2025-09-05北京科杰科技有限公司

Patent Information

Application Number
CN202511099572.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-09-05
Estimated Expiration
2045-08-07

AI Technical Summary

Technical Problem

Traditional hot and cold data storage solutions are difficult to adapt to the dynamic changes of complex business systems, resulting in waste of resources and inefficient storage management. Especially in log storage and streaming computing scenarios, frequent operation of small files increases management difficulty, and existing algorithms cannot optimize storage resources according to access mode.

Method used

A hot and cold data storage optimization system based on large models is adopted, including business integration modules, data dependency modules, pattern update modules, hot and cold partition modules and file merging modules. Through large models, analyzing business semantics, capturing data pattern characteristics, dynamically adjusting hot and cold labels and storage strategies, merging small files, and optimizing storage resource utilization.

Benefits of technology

It realizes more accurate classification of hot and cold data, reduces storage costs, improves storage efficiency and performance, adapts to the needs of different business scenarios, and optimizes storage resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120596490A_ABST
    Figure CN120596490A_ABST
Patent Text Reader

Abstract

The invention relates to the field of cold and hot data storage, in particular to a cold and hot data storage optimization method based on a large model, which comprises a business integration module, a data dependence module, a mode updating module, a cold and hot partitioning module and a file merging module, and is characterized in that the business integration module is used for collecting business data to form a training data set; the data dependency module is used for fitting business data to obtain a data consanguinity model, the mode updating module is used for predicting an access mode, the cold and hot partitioning module is used for calculating the heat of the data and performing dynamic scheduling, and the file merging module is used for automatically merging small files. A data storage structure is standardized, storage space is reduced, more efficient and more accurate cold and hot data classification can be achieved, the resource utilization rate of data storage hardware is improved, hot data response delay is reduced, the system storage utilization rate is remarkably improved, and the storage efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of cold and hot data storage, and in particular to a cold and hot data storage optimization method based on a large model. Background Art

[0002] Hot and cold data storage is a data management approach that uses different storage strategies based on data access frequency and importance. By separating data into frequently accessed hot data and infrequently accessed cold data, data access characteristics are matched to hardware storage performance, maximizing data storage efficiency. Currently, this storage and computing separation architecture for hot and cold data has become a mainstream solution in cloud computing and big data processing.

[0003] Traditional hot and cold data storage solutions classify data based on access frequency or time windows. In complex computer business systems, this approach struggles to accurately adapt to diverse business scenarios. Furthermore, data dependencies between classification tasks are difficult to track, leading to wasted storage resources and difficulties in dynamically adjusting optimization strategies. Frequent business data changes shorten the hot and cold data switching cycle, and data dependencies constantly shift. Traditional hot and cold data classification algorithms are ill-suited for highly concurrent business data environments.

[0004] In addition, in business processing scenarios such as log storage, streaming computing, and transaction data, a large number of small files such as log data and transaction data are often generated. These small files not only increase the difficulty of storage management, but also lead to frequent disk read and write operations, increasing storage management and query overhead. However, simple compression algorithms cannot be effectively adjusted according to actual access patterns, resulting in suboptimal utilization of storage resources and affecting the normal operation of the database. Summary of the Invention

[0005] The purpose of the present invention is to provide a cold and hot data storage optimization method based on a large model to solve the problems raised in the above background technology.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: a large model-based cold and hot data storage optimization system, comprising: a business integration module, a data dependency module, a schema update module, a cold and hot partitioning module, and a file merging module; The business integration module is used to integrate with various business management systems through API plug-ins. The business management system includes log analysis, AI training, and data lake systems. It collects business data, including access logs, storage modes, and computing task execution records, defines a unified log format, and tracks the data generation, storage, calculation, and consumption processes. The tracked data is categorized and stored in the database to form a training data set. The data dependency module is used to use a large language model or multimodal model with a GNN+Transformer hybrid architecture to parse the business semantics in the SQL database, train the model using a training data set within a historical period, fit the association relationship between each business data in the business system, the data call chain and upstream and downstream dependency information, perform cross-level classification according to the data call or dependency relationship, and perform same-level classification according to the data association relationship to obtain a data lineage model; The pattern update module is used to embed and encode data access behavior using the trained large model to capture data pattern features, including data time series features, periodic features, and burst features, and predict the user's access pattern in the next cycle. The access pattern includes the number of accesses, access frequency, and operation type. After the next cycle ends, the collected data changes are input into the large model to automatically update the hot and cold labels and data storage strategy. The hot and cold partitioning module is used to calculate the heat of each group of data based on the current time, data creation time, user access pattern and data lineage model using the Bayesian algorithm, and divide the data into hot data, warm data and cold data. Each data unit is labeled with a hot, cold and warm label. The data unit includes: data table, data block, file and object storage unit. At the same time, the data migration mechanism is activated. Hot data is stored in high-performance storage media, cold data is stored in low-cost storage media, and warm data is dynamically scheduled according to the predicted user access pattern. The file merging module is used to periodically detect small files generated in the business system. The small files include log data and transaction data. The large model is used to analyze their access patterns, and the small files are automatically merged according to the data lineage model. After the merger, the virtual reference layer of the original file is retained so that the unupdated business code can still be accessed.

[0007] Furthermore, the business integration module includes: a data collection unit and a flow tracking unit; The data collection unit is used to perform data collection and adaptation through Kafka, CDC or ETL tools, obtain business data and clean, standardize, time-align and structure the original data; The flow tracking unit is used to parse SQL scripts through a large model, track implicit dependencies in database JOIN or WHERE operations, and build a training data set.

[0008] Furthermore, the data dependency module includes: a large model unit, a training data unit and a lineage model unit; The large model unit is used to build or call a pre-trained large model and use the large model to analyze the collected data; The training data unit is used to pre-train the large model through data entity mask prediction or relationship path prediction, and to perform secondary training on the large model through the training data set; The lineage model unit is used to establish the dependency relationship between data and business modules and the access heat propagation path through the embedding analysis of business data, and generate a data lineage model.

[0009] Furthermore, the pattern updating module includes: a feature capturing unit and a strategy adjusting unit; The feature capture unit is used to capture pattern features of data and predict user access patterns; The strategy adjustment unit is used to periodically re-collect business data and input it into the large model for re-training and inference.

[0010] Furthermore, the hot and cold partition module includes: a heat calculation unit, a data migration unit and a dynamic scheduling unit; The heat calculation unit is used to output a temperature status label and temperature trend prediction for each piece of data from the large model; The data migration unit is used to label the data as hot, cold, or warm based on the heat calculation result and distribute the data to different storage media; The dynamic scheduling unit is used to enable a data migration mechanism, including hierarchical storage, periodic preheating and delayed loading, to ensure data consistency.

[0011] Furthermore, the file merging module includes: a small file retrieval unit and an associated merging unit; The small file retrieval unit is used to retrieve small files and output the file size, level in the data lineage model and local access times of the small files; The associated merging unit is used to calculate the percentage of NameNode pressure reduction before and after the merger, and make intelligent decisions on the merger of small files.

[0012] The cold and hot data storage optimization method based on a large model includes the following steps: Step S1. Integrate the API plug-in in the business management system to track the data generation, storage, calculation and consumption process, collect business data logs, classify and store the business data in a unified log format in the database to obtain a training data set; Step S2. Use the training dataset to train the large model. The large model parses the business semantics in the SQL database, fits the association relationship between each business data, the data call chain, and the upstream and downstream dependency information, sorts the data across levels based on the call or dependency relationship, and sorts the data at the same level based on the data association relationship, to obtain the data lineage model; Step S3. Use the trained big model to embed and encode data access behavior. By capturing the pattern characteristics of business data, predict the user access pattern in the next cycle and update the big model after the next cycle. Step S4. Based on the current time, data creation time, user access patterns, and data lineage model, determine the popularity of each data group, label it as hot data, warm data, or cold data, and initiate a data migration mechanism to store hot data on high-performance storage media, cold data on low-cost storage media, and dynamically schedule warm data. Step S5. Periodically detect small files generated in the business system, use the large model to analyze their access patterns, automatically merge excessive small files based on the data lineage model, and retain the virtual reference layer of the original files after the merger.

[0013] Furthermore, step S1 includes: Step S11. Connect the cloud to various business systems via API plug-ins to collect business data. The business management system includes log analysis, AI training, and data lake systems. Business data includes access logs, storage modes, and computing task execution records. A low-overhead collection agent plug-in is installed in the business system to increase the stability of the data collection process. Step S12. Use Kafka, CDC, or ETL tools to perform data collection and adaptation, obtain business data, and clean, standardize, time-align, and structure the original data. Upload the business data to the SQL database, classify and store the business data in a unified log format, and obtain a training data set.

[0014] Furthermore, step S2 includes: Step S21. Use a large language model or multimodal model in a GNN+Transformer hybrid architecture to parse the business semantics in the SQL database. Use a historical training dataset to train the model, fit the associations between various business data in the business system, the data call chain, and upstream and downstream dependency information, and obtain a trained large model. Step S22. Parse the SQL database script through the trained large model, track the implicit dependencies in the database JOIN or WHERE operations, perform lineage analysis on the entire life cycle of the data, perform cross-level classification according to the call or dependency relationship of the data, perform same-level classification according to the data association relationship, build a hierarchical data relationship map starting from the basic data, model the strong and weak dependencies between the data and the business modules and the access heat propagation path, and obtain the data lineage model.

[0015] Furthermore, step S3 includes: Step S31. Use the trained large model to embed and encode data access behavior. By capturing the pattern characteristics of business data, including data time series characteristics, periodic characteristics, and burst characteristics, predict the user's access pattern in the next period. The access pattern includes the number of visits, access frequency, and operation type. Step S32. After the next cycle ends, business data is collected again, and the collected data changes are input into the large model for retraining and reasoning, and the hot and cold labels and data storage strategies are automatically updated.

[0016] Furthermore, step S4 includes: Step S41. Determine the heat of each group of data, satisfying H=c·e^[-a·(t1-t2)]+m·Hc+n·Hf, where H is the heat of the data, a is a coefficient related to the historical access frequency, t1 and t2 are the current time and the last access time of the data, respectively, m and n are the preset cross-level proportional coefficient and the same-level proportional coefficient, respectively, Hc is the average heat of the data with a call or dependency relationship in the data lineage model, and Hf is the average heat of the data with an association relationship in the data lineage model. If there is no call, dependency, or association data, Hc and Hf are both 0; Step S42. According to the preset heat range, the data is divided into hot data, warm data and cold data, and each data unit is labeled with a hot, cold and warm label. Hot data is stored in high-performance storage media, including SSD and cache, and cold data is stored in low-cost storage media, including cold backup repositories and archive repositories. Warm data is dynamically scheduled based on the predicted user access pattern by introducing hierarchical storage, periodic preheating and delayed loading mechanisms.

[0017] Furthermore, step S5 includes: Step S51. Periodically detect small files generated in the business system and output the file size, level in the data lineage model, and local access count of the small files, wherein the small files include log data and transaction data; Step S52: Calculate the percentage of NameNode pressure reduction before and after the merge, make intelligent decisions on merging small files, and retain the virtual reference layer of the original file after the merge so that the business code that has not been updated can still be accessed.

[0018] Compared with the prior art, the present invention has the following beneficial effects: The present invention collects access logs, storage modes and computing task execution records of business data, tracks the generation, storage, calculation and consumption processes of data, thereby performing lineage analysis on business data and determining the dependencies between data. This can reduce duplicate storage, ensure the uniqueness of data sources, standardize data storage structures, reduce storage space, and achieve more efficient and accurate classification of hot and cold data.

[0019] The present invention embeds and encodes data access behavior through a large model, captures the pattern characteristics of data, predicts user access patterns, calculates the popularity of each group of data, divides the data into hot data, warm data and cold data and stores them in categories, thereby reducing data storage costs, optimizing database access performance, improving the resource utilization of data storage hardware, and reducing hot data response delays.

[0020] The present invention uses a large model to analyze the access patterns of small files generated in the business system, automatically merges excessive small files according to data lineage, updates hot and cold labels and data storage strategies, and automatically adjusts storage strategies. It can intelligently optimize the data storage layout according to the needs of different business scenarios, and significantly improve the system storage utilization and storage efficiency through precise tiered storage of hot and cold data, small file merging optimization, and adaptive storage adjustment strategies. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1 2. It is a structural diagram of a large-scale model-based cold and hot data storage optimization system of the present invention; Figure 2 It is a schematic diagram of the steps of the cold and hot data storage optimization method based on a large model of the present invention. DETAILED DESCRIPTION

[0022] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0023] See also Figure 1 ,The present invention provides a technical solution: a large model-based cold and hot data storage optimization system, including: a business integration module, a data dependency module, a schema update module, a cold and hot partition module and a file merging module; The business integration module is used to integrate with various business management systems through API plug-ins. The business management system includes log analysis, AI training, and data lake systems. It collects business data, including access logs, storage modes, and computing task execution records, defines a unified log format, and tracks the data generation, storage, calculation, and consumption processes. The tracked data is categorized and stored in the database to form a training data set. The business integration module includes: a data collection unit and a flow tracking unit; The data collection unit is used to perform data collection and adaptation through Kafka, CDC or ETL tools, obtain business data and clean, standardize, time-align and structure the original data; The flow tracking unit is used to parse SQL scripts through a large model, track implicit dependencies in database JOIN or WHERE operations, and build a training data set.

[0024] The data dependency module is used to use a large language model or multimodal model with a GNN+Transformer hybrid architecture to parse the business semantics in the SQL database, train the model using a training data set within a historical period, fit the association relationship between each business data in the business system, the data call chain and upstream and downstream dependency information, perform cross-level classification according to the data call or dependency relationship, and perform same-level classification according to the data association relationship to obtain a data lineage model; The data dependency module includes: a large model unit, a training data unit and a blood relationship model unit; The large model unit is used to build or call a pre-trained large model and use the large model to analyze the collected data; The training data unit is used to pre-train the large model through data entity mask prediction or relationship path prediction, and to perform secondary training on the large model through the training data set; The lineage model unit is used to establish the dependency relationship between data and business modules and the access heat propagation path through the embedding analysis of business data, and generate a data lineage model.

[0025] The pattern update module is used to embed and encode data access behavior using the trained large model to capture data pattern features, including data time series features, periodic features, and burst features, and predict the user's access pattern in the next cycle. The access pattern includes the number of accesses, access frequency, and operation type. After the next cycle ends, the collected data changes are input into the large model to automatically update the hot and cold labels and data storage strategy. The pattern updating module includes: a feature capturing unit and a strategy adjusting unit; The feature capture unit is used to capture pattern features of data and predict user access patterns; The strategy adjustment unit is used to periodically re-collect business data and input it into the large model for re-training and inference.

[0026] The hot and cold partitioning module is used to calculate the heat of each group of data based on the current time, data creation time, user access pattern and data lineage model using the Bayesian algorithm, and divide the data into hot data, warm data and cold data. Each data unit is labeled with a hot, cold and warm label. The data unit includes: data table, data block, file and object storage unit. At the same time, the data migration mechanism is activated. Hot data is stored in high-performance storage media, cold data is stored in low-cost storage media, and warm data is dynamically scheduled according to the predicted user access pattern. The hot and cold partition module includes: a heat calculation unit, a data migration unit and a dynamic scheduling unit; The heat calculation unit is used to output a temperature status label and temperature trend prediction for each piece of data from the large model; The data migration unit is used to label the data as hot, cold, or warm based on the heat calculation result and distribute the data to different storage media; The dynamic scheduling unit is used to enable a data migration mechanism, including hierarchical storage, periodic preheating and delayed loading, to ensure data consistency.

[0027] The file merging module is used to periodically detect small files generated in the business system. The small files include log data and transaction data. The large model is used to analyze their access patterns, and the small files are automatically merged according to the data lineage model. After the merger, the virtual reference layer of the original file is retained so that the unupdated business code can still be accessed.

[0028] The file merging module includes: a small file retrieval unit and an associated merging unit; The small file retrieval unit is used to retrieve small files and output the file size, level in the data lineage model and local access times of the small files; The associated merging unit is used to calculate the percentage of NameNode pressure reduction before and after the merger, and make intelligent decisions on the merger of small files.

[0029] like Figure 2 As shown, the hot and cold data storage optimization method based on the large model includes the following steps: Step S1. Integrate the API plug-in in the business management system to track the data generation, storage, calculation and consumption process, collect business data logs, classify and store the business data in a unified log format in the database to obtain a training data set; Step S1 includes: Step S11. Connect the cloud to various business systems via API plug-ins to collect business data. The business management system includes log analysis, AI training, and data lake systems. Business data includes access logs, storage modes, and computing task execution records. A low-overhead collection agent plug-in is installed in the business system to increase the stability of the data collection process. Step S12. Use Kafka, CDC, or ETL tools to perform data collection and adaptation, obtain business data, and clean, standardize, time-align, and structure the original data. Upload the business data to the SQL database, classify and store the business data in a unified log format, and obtain a training data set.

[0030] Step S2. Use the training dataset to train the large model. The large model parses the business semantics in the SQL database, fits the association relationship between each business data, the data call chain, and the upstream and downstream dependency information, sorts the data across levels based on the call or dependency relationship, and sorts the data at the same level based on the data association relationship, to obtain the data lineage model; Step S2 includes: Step S21. Use a large language model or multimodal model in a GNN+Transformer hybrid architecture to parse the business semantics in the SQL database. Use a historical training dataset to train the model, fit the associations between various business data in the business system, the data call chain, and upstream and downstream dependency information, and obtain a trained large model. Step S22. Parse the SQL database script through the trained large model, track the implicit dependencies in the database JOIN or WHERE operations, perform lineage analysis on the entire life cycle of the data, perform cross-level classification according to the call or dependency relationship of the data, perform same-level classification according to the data association relationship, build a hierarchical data relationship map starting from the basic data, model the strong and weak dependencies between the data and the business modules and the access heat propagation path, and obtain the data lineage model.

[0031] Step S3. Use the trained big model to embed and encode data access behavior. By capturing the pattern characteristics of business data, predict the user access pattern in the next cycle and update the big model after the next cycle. Step S3 includes: Step S31. Use the trained large model to embed and encode data access behavior. By capturing the pattern characteristics of business data, including data time series characteristics, periodic characteristics, and burst characteristics, predict the user's access pattern in the next period. The access pattern includes the number of visits, access frequency, and operation type. Step S32. After the next cycle ends, business data is collected again, and the collected data changes are input into the large model for retraining and reasoning, and the hot and cold labels and data storage strategies are automatically updated.

[0032] Step S4. Based on the current time, data creation time, user access patterns, and data lineage model, determine the popularity of each data group, label it as hot data, warm data, or cold data, and initiate a data migration mechanism to store hot data on high-performance storage media, cold data on low-cost storage media, and dynamically schedule warm data. Step S4 includes: Step S41. Determine the heat of each group of data, satisfying H=c·e^[-a·(t1-t2)]+m·Hc+n·Hf, where H is the heat of the data, a is a coefficient related to the historical access frequency, t1 and t2 are the current time and the last access time of the data, respectively, m and n are the preset cross-level proportional coefficient and the same-level proportional coefficient, respectively, Hc is the average heat of the data with a call or dependency relationship in the data lineage model, and Hf is the average heat of the data with an association relationship in the data lineage model. If there is no call, dependency, or association data, Hc and Hf are both 0; Step S42. According to the preset heat range, the data is divided into hot data, warm data and cold data, and each data unit is labeled with a hot, cold and warm label. Hot data is stored in high-performance storage media, including SSD and cache, and cold data is stored in low-cost storage media, including cold backup repositories and archive repositories. Warm data is dynamically scheduled based on the predicted user access pattern by introducing hierarchical storage, periodic preheating and delayed loading mechanisms.

[0033] Step S5. Periodically detect small files generated in the business system, use the large model to analyze their access patterns, automatically merge excessive small files based on the data lineage model, and retain the virtual reference layer of the original files after the merger.

[0034] Step S5 includes: Step S51. Periodically detect small files generated in the business system and output the file size, level in the data lineage model, and local access count of the small files, wherein the small files include log data and transaction data; Step S52: Calculate the percentage of NameNode pressure reduction before and after the merge, make intelligent decisions on merging small files, and retain the virtual reference layer of the original file after the merge so that the business code that has not been updated can still be accessed.

[0035] Example: There are 4 data units in the business system. The large model is used to track access to the data units. It is found that data unit 1 depends on data unit 2, data unit 2 is associated with data unit 3, and data unit 4 has no dependency or association relationship. A three-layer data lineage model is constructed, with data unit 1 as the top layer, data units 2 and 3 as the lower layer, and data unit 4 as the bottom layer. The heat of the four data blocks is 0.8, 0.5, 0.7 and 0.2 respectively. Data units 1 and 3 are stored as hot data in the cache, data unit 2 is scheduled as warm data, and data unit 4 is stored as cold data on the hard disk.

[0036] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0037] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A cold and hot data storage optimization method based on a large model, characterized in that: The method comprises the following steps: Step S1. Integrate the API plug-in in the business management system to track the data generation, storage, calculation and consumption process, collect business data logs, classify and store the business data in a unified log format in the database to obtain a training data set; Step S2. Use the training dataset to train the large model. The large model parses the business semantics in the SQL database, fits the association relationship between each business data, the data call chain, and the upstream and downstream dependency information, sorts the data across levels based on the call or dependency relationship, and sorts the data at the same level based on the data association relationship, to obtain the data lineage model; Step S3. Use the trained big model to embed and encode data access behavior. By capturing the pattern characteristics of business data, predict the user access pattern in the next cycle and update the big model after the next cycle. Step S4. Based on the current time, data creation time, user access patterns, and data lineage model, determine the popularity of each data group, label it as hot data, warm data, or cold data, and initiate a data migration mechanism to store hot data on high-performance storage media, cold data on low-cost storage media, and dynamically schedule warm data. Step S5. Periodically detect small files generated in the business system, use the large model to analyze their access patterns, automatically merge excessive small files based on the data lineage model, and retain the virtual reference layer of the original files after the merger.

2. The large model-based hot and cold data storage optimization method according to claim 1, characterized in that: Step S1 includes: Step S11. Connect the cloud to various business systems via API plug-ins to collect business data. The business management system includes log analysis, AI training, and data lake systems. Business data includes access logs, storage modes, and computing task execution records. A low-overhead collection agent plug-in is installed in the business system to increase the stability of the data collection process. Step S12. Use Kafka, CDC, or ETL tools to perform data collection and adaptation, obtain business data, and clean, standardize, time-align, and structure the original data. Upload the business data to the SQL database, classify and store the business data in a unified log format, and obtain a training data set.

3. The large-model-based cold and hot data storage optimization method according to claim 2, characterized in that: Step S2 includes: Step S21. Use a large language model or multimodal model in a GNN+Transformer hybrid architecture to parse the business semantics in the SQL database. Use a historical training dataset to train the model, fit the associations between various business data in the business system, the data call chain, and upstream and downstream dependency information, and obtain a trained large model. Step S22. Parse the SQL database script through the trained large model, track the implicit dependencies in the database JOIN or WHERE operations, perform lineage analysis on the entire life cycle of the data, perform cross-level classification according to the call or dependency relationship of the data, perform same-level classification according to the data association relationship, build a hierarchical data relationship map starting from the basic data, model the strong and weak dependencies between the data and the business modules and the access heat propagation path, and obtain the data lineage model.

4. The large model-based hot and cold data storage optimization method according to claim 3, characterized in that: Step S3 includes: Step S31. Use the trained large model to embed and encode data access behavior. By capturing the pattern characteristics of business data, including data time series characteristics, periodic characteristics, and burst characteristics, predict the user's access pattern in the next period. The access pattern includes the number of visits, access frequency, and operation type. Step S32. After the next cycle ends, business data is collected again, and the collected data changes are input into the large model for retraining and reasoning, and the hot and cold labels and data storage strategies are automatically updated.

5. The large model-based hot and cold data storage optimization method according to claim 4, characterized in that: Step S4 includes: Step S41. Determine the heat of each group of data, satisfying H=c·e^[-a·(t1-t2)]+m·Hc+n·Hf, where H is the heat of the data, a is a coefficient related to the historical access frequency, t1 and t2 are the current time and the last access time of the data, m and n are the preset cross-level proportional coefficient and the same-level proportional coefficient, respectively, Hc is the average heat of the data with a call or dependency relationship in the data lineage model, and Hf is the average heat of the data with an association relationship in the data lineage model; Step S42: Data is divided into hot, warm, and cold data according to preset heat intervals. Each data unit is labeled hot, cold, and warm. Hot data is stored in high-performance storage media, including SSDs and caches, while cold data is stored in low-cost storage media, including cold backup and archive repositories. Warm data is dynamically scheduled based on predicted user access patterns using hierarchical storage, periodic preheating, and delayed loading mechanisms. Step S5 includes: Step S51. Periodically detect small files generated in the business system and output the file size, level in the data lineage model, and local access count of the small files, wherein the small files include log data and transaction data; Step S52: Calculate the percentage of NameNode pressure reduction before and after the merge, make intelligent decisions on merging small files, and retain the virtual reference layer of the original file after the merge so that the business code that has not been updated can still be accessed.

6. A large-model-based cold and hot data storage optimization system, wherein the system executes the large-model-based cold and hot data storage optimization method described in claim 1, characterized in that: Includes the following modules: business integration module, data dependency module, schema update module, hot and cold partition module and file merging module; The business integration module is used to integrate with various business management systems through API plug-ins. The business management system includes log analysis, AI training, and data lake systems. It collects business data, including access logs, storage modes, and computing task execution records, defines a unified log format, and tracks the data generation, storage, calculation, and consumption processes. The tracked data is categorized and stored in the database to form a training data set. The data dependency module is used to use a large language model or multimodal model with a GNN+Transformer hybrid architecture to parse the business semantics in the SQL database, train the model using a training data set within a historical period, fit the association relationship between each business data in the business system, the data call chain and upstream and downstream dependency information, perform cross-level classification according to the data call or dependency relationship, and perform same-level classification according to the data association relationship to obtain a data lineage model; The pattern update module is used to embed and encode data access behavior using the trained large model to capture data pattern features, including data time series features, periodic features, and burst features, and predict the user's access pattern in the next cycle. The access pattern includes the number of accesses, access frequency, and operation type. After the next cycle ends, the collected data changes are input into the large model to automatically update the hot and cold labels and data storage strategy. The hot and cold partitioning module is used to calculate the heat of each group of data based on the current time, data creation time, user access pattern and data lineage model using the Bayesian algorithm, and divide the data into hot data, warm data and cold data. Each data unit is labeled with a hot, cold and warm label. The data unit includes: data table, data block, file and object storage unit. At the same time, the data migration mechanism is activated. Hot data is stored in high-performance storage media, cold data is stored in low-cost storage media, and warm data is dynamically scheduled according to the predicted user access pattern. The file merging module is used to periodically detect small files generated in the business system. The small files include log data and transaction data. The large model is used to analyze their access patterns, and the small files are automatically merged according to the data lineage model. After the merger, the virtual reference layer of the original file is retained so that the unupdated business code can still be accessed.

7. The large model-based hot and cold data storage optimization system according to claim 6, characterized in that: The business integration module includes: a data collection unit and a flow tracking unit; The data collection unit is used to perform data collection and adaptation through Kafka, CDC or ETL tools, obtain business data and clean, standardize, time-align and structure the original data; The flow tracking unit is used to parse SQL scripts through a large model, track implicit dependencies in database JOIN or WHERE operations, and build a training data set.

8. The large model-based hot and cold data storage optimization system according to claim 7, characterized in that: The data dependency module includes: a large model unit, a training data unit and a blood relationship model unit; The large model unit is used to build or call a pre-trained large model and use the large model to analyze the collected data; The training data unit is used to pre-train the large model through data entity mask prediction or relationship path prediction, and to perform secondary training on the large model through the training data set; The lineage model unit is used to establish the dependency relationship between data and business modules and the access heat propagation path through the embedding analysis of business data, and generate a data lineage model.

9. The large model-based hot and cold data storage optimization system according to claim 8, characterized in that: The pattern updating module includes: a feature capturing unit and a strategy adjusting unit; The feature capture unit is used to capture pattern features of data and predict user access patterns; The strategy adjustment unit is used to periodically re-collect business data and input it into the large model for re-training and reasoning; The hot and cold partition module includes: a heat calculation unit, a data migration unit and a dynamic scheduling unit; The heat calculation unit is used to output a temperature status label and temperature trend prediction for each piece of data from the large model; The data migration unit is used to label the data as hot, cold, or warm based on the heat calculation result and distribute the data to different storage media; The dynamic scheduling unit is used to enable a data migration mechanism, including hierarchical storage, periodic preheating and delayed loading, to ensure data consistency.

10. The large model-based hot and cold data storage optimization system according to claim 9, characterized in that: The file merging module includes: a small file retrieval unit and an associated merging unit; The small file retrieval unit is used to retrieve small files and output the file size, level in the data lineage model and local access times of the small files; The associated merging unit is used to calculate the percentage of NameNode pressure reduction before and after the merger, and make intelligent decisions on the merger of small files.

Citation Information

Patent Citations

  • Cold data storage and regenerative analysis method and system based on Hadoop cluster

    CN117008839A

  • Method and system for realizing self-adaptive object storage data life cycle management based on deep learning large model

    CN118820200A

  • Data consanguinity full-link traceability method and system based on multi-source heterogeneous metadata and pre-training large model

    CN119917814A

  • Intelligent decision-making system and method based on enterprise life index large model

    CN119990833A

  • Data-based blood relationship analysis method, apparatus, and device and computer-readable storage medium

    WO2021218021A1

Cited By

  • Service resource dynamic optimization configuration method and system based on deep learning prediction

    CN120822787A

  • Business resource dynamic optimization configuration method and system based on deep learning prediction

    CN120822787B

  • Data optimization storage method and system based on artificial intelligence

    CN122044476A