Multi-source heterogeneous data hierarchical compression storage system driven by AI large model
The multi-source heterogeneous data hierarchical compression storage system driven by AI large models solves the problems of insufficient access compatibility, feature recognition accuracy and adaptive optimization capability of multi-source heterogeneous data storage systems, and achieves efficient storage and low-cost data management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-22
- Publication Date
- 2026-04-03
AI Technical Summary
Existing multi-source heterogeneous data hierarchical compression storage systems suffer from problems such as insufficient access compatibility, low feature recognition accuracy, poor coordination between hierarchical and compression, lack of adaptive optimization capabilities, and imperfect end-to-end feedback, resulting in low storage efficiency, difficulty in controlling data distortion rate, and high storage costs.
The multi-source heterogeneous data hierarchical compression and storage system driven by AI large model achieves accurate feature identification, dynamic hierarchical decision-making, adaptive compression, and intelligent storage management of multi-source heterogeneous data through full-link collaboration between modules. It includes a multi-source heterogeneous data access module, an AI large model-driven data feature identification module, a hierarchical decision-making module, a multi-level compression module, and a collaborative storage module, forming a closed-loop feedback mechanism.
It enables efficient access and comprehensive and accurate feature representation of multi-source heterogeneous data, dynamically adapts storage, significantly improves storage efficiency and reduces storage costs, and ensures the long-term stability of the system and its dynamic adaptability to multiple scenarios.
Smart Images

Figure CN121785534A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data processing technology, specifically an AI large-scale model-driven multi-source heterogeneous data hierarchical compression and storage system. Background Technology
[0002] In the digital economy era, the emergence of massive amounts of heterogeneous data from multiple sources, such as IoT devices and industrial sensors, has led to a variety of data formats, uneven value density, and significant differences in access frequency and real-time requirements. This places high demands on the adaptability, compression efficiency, and cost control of storage systems. Existing technologies mostly employ a combination of "tiered storage + data compression," which involves dividing data into tiers based on access frequency and pairing them with different storage media, while combining compression algorithms to reduce overhead. However, key technical shortcomings still exist: 1. Insufficient data access compatibility, limited protocol adaptation range, format conversion not adapted to subsequent feature extraction, and noise filtering unable to effectively remove redundant and hidden abnormal data. 2. The data feature recognition dimension is singular, focusing only on surface features. The value density assessment is detached from the business scenario, lacks multimodal fusion capability, and is difficult to adapt to the feature extraction of heterogeneous data. 3. The hierarchical decision-making mechanism is static, with fixed indicator weights, and it does not establish a synergistic relationship with the compression strategy, making it impossible to achieve a precise match between storage level, compression intensity, and data value. 4. The compression strategy is fixed and does not dynamically adjust based on data hierarchical attributes and feature distribution, lacking an adaptive optimization mechanism and a closed-loop feedback mechanism. 5. The rigid allocation of storage media and the reliance on fixed rules for data lifecycle management result in a lack of closed-loop feedback throughout the entire process, leading to insufficient dynamic adaptability and long-term stability of the system. In summary, the aforementioned shortcomings result in low storage efficiency, difficulty in controlling data distortion, and high storage costs in existing systems, failing to meet the high-efficiency storage requirements of large-scale, multi-source, heterogeneous data. Therefore, there is an urgent need to develop storage technology solutions with accurate feature recognition, dynamic hierarchical decision-making, and collaborative optimization capabilities. Summary of the Invention
[0003] To address the shortcomings of existing multi-source heterogeneous data hierarchical compression and storage technologies, such as insufficient access compatibility, low feature recognition accuracy, poor coordination between hierarchical and compression, lack of adaptive optimization capabilities, and imperfect end-to-end feedback, this invention provides an AI-driven multi-source heterogeneous data hierarchical compression and storage system. Through end-to-end collaboration between modules and deep empowerment by AI models, it achieves accurate feature recognition, dynamic hierarchical decision-making, adaptive compression, and intelligent storage management of multi-source heterogeneous data, thereby achieving synergistic optimization of storage efficiency, data quality, and storage cost.
[0004] The technical solution adopted by this invention to solve its technical problem is as follows: This invention proposes a power energy question-answering system based on multi-data fusion and semantic parsing, including a multi-source heterogeneous data access module that is sequentially linked and forms a closed-loop feedback, an AI large model-driven data feature recognition module, a hierarchical decision-making module, a multi-level compression module, and a collaborative storage module. The multi-source heterogeneous data access module is used to collect multi-source heterogeneous raw data of different types and formats across the entire domain and output standardized preprocessed data. The AI-driven big data feature recognition module receives the preprocessed data and extracts the semantic features, structural features, access frequency features, and value density features of the data through multimodal fusion AI big data model to generate a unified multidimensional feature vector. The hierarchical decision-making module constructs a dynamic hierarchical evaluation index system based on the multidimensional feature vector, and generates decision results that are associated with hierarchical identifiers and compression strategy guidance. The multi-level compression module matches a differentiated compression strategy based on the decision result, performs adaptive compression processing on data at different levels, and outputs compressed data and compression effect feedback information. The collaborative storage module receives the compressed data, allocates storage resources according to the decision results and establishes an association mapping, and at the same time collects data access logs and feeds them back to the feature recognition module. Through end-to-end collaboration between modules—"data access, feature extraction, hierarchical decision-making, differentiated compression, collaborative storage, and feedback optimization"—dynamic adaptation and storage of multi-source heterogeneous data can be achieved, resulting in synergistic optimization of storage efficiency, data quality, and storage cost.
[0005] Furthermore, the multi-source heterogeneous data access module includes a protocol adaptation unit, a format standardization unit, and a noise filtering unit; The protocol adaptation unit is compatible with multiple types of communication protocols and is used for seamless access to heterogeneous data sources; The format standardization unit performs format conversion on unstructured data, semi-structured data, and structured data respectively, and outputs intermediate format data that is adapted to the feature extraction requirements of AI large model; The noise filtering unit uses a multi-dimensional screening and anomaly detection mechanism to remove invalid and redundant data, thereby improving the purity of the preprocessed data.
[0006] Furthermore, the multimodal fusion AI model in the AI large-scale model-driven data feature recognition module is built on the Transformer architecture. After self-supervised learning pre-training, it possesses cross-type data feature extraction capabilities, fusing text semantic understanding branches, image feature extraction branches, time-series data analysis branches, and numerical data statistics branches. These branches strengthen key feature weights through an attention mechanism. The single-modal features output from these branches are weighted and integrated by the feature fusion layer and undergo cross-modal calibration to generate the... A unified multidimensional feature vector is used; the value density feature is calculated based on standardized preprocessed data output by the multi-source heterogeneous data access module. This standardized preprocessed data originates from the results of multi-source heterogeneous data sources after collection and preprocessing by the access module. The business scenario is the target business scenario adapted to the system, which originates from user configuration or scenario requirement information associated with the business system. The value density feature is calculated through a data-business scenario correlation quantification model. The quantification model combines the degree of support of preprocessed data for the core objectives of the business scenario, the indispensability of data in the business process, and the matching degree between data timeliness and business requirements. A multi-dimensional weighted quantification method is used to quantitatively represent the value density.
[0007] Furthermore, the dynamic hierarchical evaluation index system of the hierarchical decision-making module includes primary and secondary indicators. The primary indicators cover data value, access priority, and real-time requirements, while the secondary indicators cover value density, access frequency, update frequency, and latency tolerance. The initial indicator weights are determined by the analytic hierarchy process (AHP), and the weight coefficients are dynamically adjusted based on the compression effect feedback information. The hierarchical identifier in the decision result corresponds to at least four storage layers, including a high-frequency access hot data layer, a medium-frequency access warm data layer, a low-frequency access cold data layer, and an archived data layer, and each layer is associated with a unique compression strategy type guide.
[0008] Furthermore, the differentiated compression strategies of the multi-level compression module correspond one-to-one with the storage layers: the hot data layer adopts a lightweight lossless compression strategy, the warm data layer adopts a hybrid compression strategy, the cold data layer adopts a deep compression strategy, and the archived data layer adopts an extreme compression strategy. The lightweight lossless compression strategy is based on efficient lossless coding, the hybrid compression strategy is a collaborative application of lossless coding and low-proportion lossy compression, the deep compression strategy is a combination of deep learning quantization compression and entropy coding, and the extreme compression strategy is a linkage between AI model feature reconstruction compression and distributed compression. The matching of the compression strategy is controlled by a trade-off model between compression efficiency and distortion rate. The trade-off model receives the feature distribution information output by the large AI model and dynamically optimizes the compression parameters.
[0009] Furthermore, the collaborative storage module includes a storage media scheduling unit, a hierarchical indexing unit, a disaster recovery backup unit, and an access log collection unit; The storage medium scheduling unit allocates suitable storage media for different storage levels: the hot data layer is adapted to high-speed storage media, the warm data layer is adapted to hybrid storage media, and the cold data layer and archived data layer are adapted to low-cost storage media. The hierarchical indexing unit constructs a multi-level index structure associated with the hierarchical identifier, and associates compressed data storage address, original data characteristics, and access permission information. The disaster recovery backup unit sets differentiated backup strategies according to storage level: hot data is backed up in real time, warm data is backed up incrementally on a regular schedule, and cold data and archived data are backed up in full on a regular schedule. The access log collection unit continuously collects data access status information to provide data support for closed-loop feedback.
[0010] Furthermore, the closed-loop feedback mechanism includes: The first feedback link is for the access log collection unit of the collaborative storage module to transmit access status information to the incremental training unit of the AI large model, and update the feature extraction weights through parameter fine-tuning to optimize the accuracy of multi-dimensional feature vector generation. The second feedback link is where the compression effect monitoring unit of the multi-level compression module transmits feedback information such as compression ratio and data recovery accuracy to the indicator optimization unit of the hierarchical decision module, dynamically adjusting the weight coefficients of the hierarchical evaluation indicators. The feedback data from both the first and second feedback links are transmitted through a standardized interactive interface to ensure the consistency and real-time nature of data interaction between modules.
[0011] Furthermore, the incremental training of the multimodal fusion AI model employs a gradient descent algorithm combined with a momentum optimization strategy, updating only the model parameters corresponding to newly added data features to avoid wasting resources on full training. The sub-models of the text semantic understanding branch, image feature extraction branch, time-series data analysis branch, and numerical data statistics branch are all adapted to multi-source heterogeneous data features through transfer learning. Specifically, the text semantic understanding branch uses a semantic understanding sub-model, the image feature extraction branch uses a visual feature extraction sub-model, the time-series data analysis branch uses a time-series analysis sub-model, and the numerical data statistics branch uses a numerical statistics sub-model. Cross-branch feature fusion is achieved through a cross-modal attention mechanism, strengthening the correlation between different types of data features.
[0012] Furthermore, the deep learning quantization compression adopts an AI model-driven adaptive quantization mechanism, which dynamically determines the quantization bit specification by analyzing the distribution of data features; the AI model feature reconstruction compression is based on an autoencoder, which shares the core logic of feature extraction with the multimodal fusion AI model, and achieves efficient compression by learning the low-dimensional feature representation of the data, while ensuring the quality of data recovery; the distributed compression improves compression efficiency through multi-node parallel processing, adapting to the needs of large-scale archived data processing.
[0013] Furthermore, the collaborative storage module also includes a data lifecycle management unit, which automatically executes a data hierarchical migration strategy based on the hierarchical decision results and access log information as follows: Hot data that has not been accessed for a long time is migrated to the warm data layer; warm data that has not been accessed for a long time is migrated to the cold data layer; and cold data that has not been accessed for a long time is migrated to the archived data layer. During the data migration process, the compression strategy and storage media configuration are updated synchronously, and the amount of data transferred during migration is reduced through differential backup technology; The data lifecycle management unit maintains real-time interaction with the hierarchical decision-making module, dynamically calibrates migration triggering conditions, and ensures the rationality of hierarchical migration.
[0014] The beneficial effects of this invention are as follows: 1. The AI-driven multi-source heterogeneous data hierarchical compression and storage system described in this invention achieves efficient access and comprehensive and accurate feature representation of multiple types of heterogeneous data by combining the multi-protocol adaptation and customized format conversion of the multi-source heterogeneous data access module with a multi-dimensional noise filtering mechanism, along with the multi-modal fusion extraction and value density quantification related to business scenarios by the AI-driven data feature recognition module. This provides reliable data support for subsequent hierarchical decision-making. 2. The AI-driven multi-source heterogeneous data hierarchical compression and storage system described in this invention achieves precise matching between storage hierarchical compression strategies and storage resources through the dynamic evaluation index system and compression strategy guidance of the hierarchical decision module, which links and coordinates the hierarchical differentiated compression and adaptive parameter optimization of multi-level compression modules. This is achieved in conjunction with the hierarchical adaptation of storage media allocation and intelligent lifecycle management by the collaborative storage module. This significantly improves storage efficiency and reduces storage costs. 3. The AI large model-driven multi-source heterogeneous data hierarchical compression and storage system described in this invention uses a dual-loop closed-loop feedback mechanism throughout the entire chain to feed back storage access logs and compression effect data to the incremental training of the AI large model and the optimization of hierarchical decision indicators in real time. This achieves continuous iterative improvement in the system's feature extraction accuracy, hierarchical decision adaptability, and compression storage performance, ensuring the long-term stability of the system and its dynamic adaptability to multiple scenarios. Attached Figure Description
[0015] The invention will now be further described with reference to the accompanying drawings.
[0016] Figure 1 This is a structural diagram of the multi-source heterogeneous data access module of the present invention; Figure 2 This is a structural diagram of the AI large model-driven data feature recognition module of this invention; Figure 3 This is a structural diagram of the hierarchical decision-making module of the present invention; Figure 4 This is a structural diagram of the multi-level compression module of the present invention; Figure 5 This is a diagram of the collaborative storage module and closed-loop feedback mechanism of the present invention. Detailed Implementation
[0017] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments and accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] like Figures 1-5 As shown, the AI large-scale model-driven multi-source heterogeneous data hierarchical compression and storage system of the present invention includes a multi-source heterogeneous data access module, an AI large-scale model-driven data feature recognition module, a hierarchical decision module, a multi-level compression module, and a collaborative storage module, which are linked sequentially and form a closed-loop feedback, to collaboratively complete the entire process of data processing from access to storage and feedback optimization. The specific technical solution is as follows: The multi-source heterogeneous data access module, serving as the system data entry point, integrates a protocol adaptation unit, a format standardization unit, and a noise filtering unit. Its operation is as follows: The protocol adaptation unit is compatible with multiple types of communication protocols, such as industrial bus protocols, IoT-specific protocols, and general Internet protocols. It achieves seamless access to multiple heterogeneous data sources through a dynamic protocol parsing mechanism, solving the problem of limited protocol adaptation range in existing technologies and improving access compatibility. The format standardization unit performs customized format conversions for the incoming structured, semi-structured, and unstructured data, and outputs intermediate format data that is precisely adapted to the feature extraction requirements of the AI large model, avoiding format adaptation errors in subsequent feature extraction stages and improving feature recognition efficiency. The noise filtering unit adopts a mechanism that combines multi-dimensional screening and anomaly detection. First, it filters out explicit invalid data such as null values and format errors through rule-based screening. Then, it identifies and removes redundant and duplicate data as well as implicit abnormal data such as logical contradictions and extreme deviations through density clustering algorithms, which effectively improves the purity of preprocessed data and reduces the waste of storage resources. The AI large-scale model-driven data feature recognition module, with a multimodal fusion AI large-scale model based on the Transformer architecture as its core, works as follows: The model receives standardized preprocessed data from the multi-source heterogeneous data access module. It simultaneously extracts the semantic features of the data through four branches: text semantic understanding, image feature extraction, time series data analysis, and numerical data statistics. These features include topic relevance, content completeness, structural features (field correlation), format complexity, access frequency features (historical access count), and time distribution. Based on the target business scenario requirements configured by users or input by business systems, and using a correlation quantification model, the system combines the degree of support of data for core business objectives, its indispensability and timeliness in business processes, and its matching degree with business requirements. It adopts a multi-dimensional weighted quantification method to complete accurate calculations, thus solving the shortcomings of traditional assessments that are divorced from business scenarios. After the single-modal features output from the four branches are enhanced with key feature weights through an attention mechanism, they are weighted and integrated and cross-modal calibrated through a feature fusion layer to generate a unified multi-dimensional feature vector. This ensures the consistency and comprehensiveness of feature representation for multi-type heterogeneous data, and the feature recognition accuracy is effectively improved compared to traditional solutions.
[0019] The hierarchical decision-making module constructs a dynamic hierarchical evaluation index system based on multi-dimensional feature vectors. The working process is as follows: The indicator system includes primary indicators such as data value, access priority, and real-time requirements, and secondary indicators such as value density, access frequency, update frequency, and latency tolerance. The initial weights are determined through the analytic hierarchy process to avoid the one-sidedness of single-dimensional decision-making. The system receives multi-dimensional feature vectors output by large AI models and uses a weighted summation algorithm to generate hierarchical decision results. These results not only include hierarchical identifiers for high-frequency access hot data layers, medium-frequency access warm data layers, low-frequency access cold data layers, and archived data layers, but also associate them with corresponding compression strategy type guidelines, achieving deep collaboration between hierarchical and compression. It receives compression performance data from multi-level compression modules in real time, including compression ratio and data recovery accuracy. Through the indicator optimization unit, it dynamically adjusts the weight coefficients of secondary indicators, solving the problem of disconnect between static decision-making and data feature changes in existing technologies, and effectively improving the adaptability of hierarchical decision-making. The multi-level compression module performs differentiated compression based on the hierarchical decision results, and the working process is as follows: The compression strategy is matched according to the storage level: the hot data layer adopts a lightweight lossless compression strategy to adapt to the high-frequency access and low latency requirements; the warm data layer adopts a hybrid compression strategy of lossless + low proportion of lossy to balance efficiency and quality; the cold data layer adopts a deep compression strategy of deep learning quantization compression + entropy coding; and the archived data layer adopts an extreme compression strategy of AI model feature reconstruction compression and distributed compression. Compression strategy matching is controlled by a trade-off model between compression efficiency and distortion rate. This model receives feature distribution information output by a large AI model and dynamically optimizes compression parameters, such as quantization bit size and coding strength, to achieve precise matching of data features, hierarchical attributes, and compression strategies. Among them, the AI model feature reconstruction compression is based on an autoencoder. This autoencoder shares the core logic of feature extraction with the multimodal fusion AI model, which not only improves compression efficiency but also ensures data recovery quality. Meanwhile, distributed compression adapts to the needs of large-scale archived data processing through parallel processing of multiple nodes. The collaborative storage module enables intelligent storage and lifecycle management of compressed data, and its operation process is as follows: The storage media scheduling unit allocates appropriate media to different data levels: high-speed storage media are allocated to the hot data layer to ensure low-latency access; hybrid storage media are allocated to the warm data layer to balance performance and cost; and low-cost storage media are allocated to the cold data layer and archived data layer to reduce long-term storage overhead. The hierarchical index unit constructs a multi-level index structure associated with the hierarchical identifier, and associates compressed data storage address, original data characteristics and access permission information, supporting millisecond-level retrieval and recovery of compressed data; The disaster recovery backup unit is configured with differentiated backup strategies at different levels: real-time backup of hot data, timed incremental backup of warm data, and regular full backup of cold data and archived data, which reduces backup overhead while ensuring data security. Based on the hierarchical decision results and access logs, the data lifecycle management unit automatically performs data migration: hot data that has not been accessed for a long time is migrated to the warm data layer, warm data that has not been accessed for a long time is migrated to the cold data layer, and cold data that has not been accessed for a long time is migrated to the archived data layer. During the migration process, the compression strategy and storage media are updated synchronously, and the amount of data transmission is reduced through differential backup technology. The closed-loop feedback mechanism includes dual feedback links to achieve dynamic system optimization. The first feedback link: The access log collection unit of the collaborative storage module regularly collects information such as the number of data accesses, response time, and modification frequency, and transmits it to the incremental training unit of the AI large model through a standardized interface. The gradient descent + momentum optimization strategy is used to update the model parameters in a targeted manner, and continuously improve the accuracy of feature extraction. The second feedback link: The compression effect monitoring unit of the multi-level compression module collects information such as compression ratio and data recovery accuracy in real time, and feeds it back to the indicator optimization unit of the hierarchical decision module to dynamically adjust the weight coefficient of the evaluation indicators and optimize the hierarchical decision results. Dual-path feedback enables the system to dynamically adapt to changes in data characteristics and adjustments in business needs, effectively improving long-term operational stability.
[0020] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.
Claims
1. An AI-driven multi-source heterogeneous data hierarchical compression and storage system, characterized in that, This includes a multi-source heterogeneous data access module that links sequentially and forms a closed-loop feedback loop, an AI large-model-driven data feature recognition module, a hierarchical decision-making module, a multi-level compression module, and a collaborative storage module: The multi-source heterogeneous data access module is used to collect multi-source heterogeneous raw data of different types and formats across the entire domain and output standardized preprocessed data. The AI-driven big data feature recognition module receives the preprocessed data and extracts the semantic features, structural features, access frequency features, and value density features of the data through multimodal fusion AI big data model to generate a unified multidimensional feature vector. The hierarchical decision-making module constructs a dynamic hierarchical evaluation index system based on the multidimensional feature vector, and generates decision results that are associated with hierarchical identifiers and compression strategy guidance. The multi-level compression module matches a differentiated compression strategy based on the decision result, performs adaptive compression processing on data at different levels, and outputs compressed data and compression effect feedback information. The collaborative storage module receives the compressed data, allocates storage resources according to the decision results and establishes an association mapping, and simultaneously collects data access logs and feeds them back to the feature recognition module.
2. The AI large model-driven multi-source heterogeneous data hierarchical compression and storage system according to claim 1, characterized in that, The multi-source heterogeneous data access module includes a protocol adaptation unit, a format standardization unit, and a noise filtering unit. The protocol adaptation unit is compatible with communication protocols and is used for seamless access to heterogeneous data sources; The format standardization unit performs format conversion on unstructured data, semi-structured data, and structured data respectively, and outputs intermediate format data that is adapted to the feature extraction requirements of AI large model; The noise filtering unit eliminates invalid and redundant data through multi-dimensional screening and anomaly detection mechanisms.
3. The AI large model-driven multi-source heterogeneous data hierarchical compression and storage system according to claim 1, characterized in that, The multimodal fusion AI model in the data feature recognition module driven by the AI model is built on the Transformer architecture. After self-supervised learning pre-training, it has the ability to extract features across different data types. It integrates text semantic understanding branches, image feature extraction branches, time series data analysis branches, and numerical data statistics branches. The text semantic understanding branches, image feature extraction branches, time series data analysis branches, and numerical data statistics branches strengthen the weight of key features through an attention mechanism. The single-modal features output by the text semantic understanding branches, image feature extraction branches, time series data analysis branches, and numerical data statistics branches are weighted and integrated and cross-modal calibrated by the feature fusion layer to generate the unified multidimensional feature vector. The value density feature is calculated based on the standardized preprocessed data output by the multi-source heterogeneous data access module. The standardized preprocessed data comes from the results of multi-source heterogeneous data sources after being collected and preprocessed by the access module. The business scenario is the target business scenario adapted to the system, which comes from the scenario requirement information associated with user configuration or business system input. The value density feature is calculated through a data-business scenario correlation quantification model.
4. The AI large model-driven multi-source heterogeneous data hierarchical compression and storage system according to claim 1, characterized in that, The dynamic hierarchical evaluation index system of the hierarchical decision-making module includes primary and secondary indicators. The primary indicators cover data value, access priority, and real-time requirements, while the secondary indicators cover value density, access frequency, update frequency, and latency tolerance. The initial indicator weights are determined by the analytic hierarchy process (AHP), and the weight coefficients are dynamically adjusted based on the compression effect feedback. The hierarchical identifier in the decision result corresponds to at least four storage layers, including a high-frequency access hot data layer, a medium-frequency access warm data layer, a low-frequency access cold data layer, and an archived data layer, and each layer is associated with a unique compression strategy type guide.
5. The AI large model-driven multi-source heterogeneous data hierarchical compression and storage system according to claim 1, characterized in that, The differentiated compression strategies of the multi-level compression module correspond one-to-one with the storage layers: the hot data layer adopts a lightweight lossless compression strategy, the warm data layer adopts a hybrid compression strategy, the cold data layer adopts a deep compression strategy, and the archived data layer adopts an extreme compression strategy. The lightweight lossless compression strategy is based on efficient lossless coding. The hybrid compression strategy is a collaborative application of lossless coding and low-proportion lossy compression. The deep compression strategy is a combination of deep learning quantization compression and entropy coding. The extreme compression strategy is a linkage between AI model feature reconstruction compression and distributed compression. The matching of the compression strategy is controlled by a trade-off model between compression efficiency and distortion rate. The trade-off model receives feature distribution information output by a large AI model and dynamically optimizes the compression parameters.
6. The AI large model-driven multi-source heterogeneous data hierarchical compression and storage system according to claim 1, characterized in that, The collaborative storage module includes a storage media scheduling unit, a hierarchical indexing unit, a disaster recovery backup unit, and an access log collection unit. The storage medium scheduling unit allocates suitable storage media for different storage levels: the hot data layer is adapted to high-speed storage media, the warm data layer is adapted to hybrid storage media, and the cold data layer and archived data layer are adapted to low-cost storage media. The hierarchical indexing unit constructs a multi-level index structure associated with the hierarchical identifier, and associates compressed data storage address, original data characteristics, and access permission information. The disaster recovery backup unit sets differentiated backup strategies according to storage level: hot data is backed up in real time, warm data is backed up incrementally on a regular schedule, and cold data and archived data are backed up in full on a regular schedule. The access log collection unit continuously collects data access status information to provide data support for closed-loop feedback.
7. The AI large model-driven multi-source heterogeneous data hierarchical compression and storage system according to claim 1, characterized in that, The closed-loop feedback mechanism includes: The first feedback link is for the access log collection unit of the collaborative storage module to transmit access status information to the incremental training unit of the AI large model, and update the feature extraction weights through parameter fine-tuning to optimize the accuracy of multi-dimensional feature vector generation. The second feedback link is where the compression effect monitoring unit of the multi-level compression module transmits feedback information such as compression ratio and data recovery accuracy to the indicator optimization unit of the hierarchical decision module, dynamically adjusting the weight coefficients of the hierarchical evaluation indicators; the feedback data from both the first and second feedback links are transmitted through a standardized interactive interface.
8. The AI large model-driven multi-source heterogeneous data hierarchical compression and storage system according to claim 3, characterized in that, The incremental training of the multimodal fusion AI model adopts a gradient descent algorithm combined with a momentum optimization strategy, and only updates the model parameters corresponding to newly added data features. The sub-models of the text semantic understanding branch, image feature extraction branch, time series data analysis branch, and numerical data statistics branch are all adapted to multi-source heterogeneous data features through transfer learning. Specifically, the text semantic understanding branch adopts a semantic understanding sub-model, the image feature extraction branch adopts a visual feature extraction sub-model, the time series data analysis branch adopts a time series analysis sub-model, and the numerical data statistics branch adopts a numerical statistics sub-model. Cross-branch feature fusion is achieved through a cross-modal attention mechanism.
9. The AI large model-driven multi-source heterogeneous data hierarchical compression and storage system according to claim 5, characterized in that, The deep learning quantization compression adopts an AI model-driven adaptive quantization mechanism, which dynamically determines the quantization bit specification by analyzing the data feature distribution; the AI model feature reconstruction compression is based on an autoencoder, which shares the core feature extraction logic with the multimodal fusion AI large model; the distributed compression improves compression efficiency through multi-node parallel processing.
10. The AI large model-driven multi-source heterogeneous data hierarchical compression and storage system according to claim 6, characterized in that, The collaborative storage module also includes a data lifecycle management unit, which automatically executes a data hierarchical migration strategy based on the hierarchical decision results and access log information as follows: Hot data that has not been accessed for a long time is migrated to the warm data layer; warm data that has not been accessed for a long time is migrated to the cold data layer; and cold data that has not been accessed for a long time is migrated to the archived data layer. During the data migration process, the compression strategy and storage media configuration are updated synchronously, and the amount of data transferred during migration is reduced through differential backup technology; The data lifecycle management unit maintains real-time interaction with the hierarchical decision-making module to dynamically calibrate migration triggering conditions.