Data archiving processing method and system based on shutdown system
By applying deep learning classification model, neural network compression encryption model, graph database index construction and blockchain archive verification technology in the shutdown system, the problem of inefficient data management is solved, intelligent classification, effective compression and secure storage of data are realized, and data integrity and traceability are ensured.
Patent Information
- Application Number
- CN202510584865.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-08
AI Technical Summary
The prior art is difficult to efficiently and securely classify, effectively compress and store data in shutdown systems, resulting in inefficient data management and difficult to adapt to rapidly changing business needs.
The data is intelligently classified by a multi-dimensional classification model based on deep learning, combined with the compression and encryption of neural networks for data compression and encryption, and an index construction algorithm based on graph database is used to generate data storage path indexes, and finally data integrity is verified through blockchain-based archive verification technology.
It realizes intelligent classification, effective compression and secure storage of data in the shutdown system, ensures data integrity and traceability, and improves the efficiency and security of data management.
Smart Images

Figure CN120104569A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data archiving, and in particular to a data archiving processing method and system based on a shutdown system. Background Art
[0002] With the rapid development of information technology, a large amount of data is widely generated and used in various industries. Many companies have accumulated a large amount of historical data in the course of their operations, and the management and storage of this data has gradually become a major challenge facing enterprises. In industries such as finance, medical care, and energy, the data is not only huge in volume, but also covers sensitive information, which is related to the security and business compliance of the enterprise. Therefore, how to efficiently and securely archive the data in the shutdown system has become an important issue that needs to be solved urgently. Traditional data archiving methods usually use static classification and simple storage solutions, which often cannot meet the requirements of modern enterprises for efficiency, security, and flexibility of data storage. In these methods, the classification and processing of data often rely on manual intervention, resulting in inefficient data archiving and difficulty in adapting to rapidly changing business needs. Summary of the invention
[0003] The purpose of the present invention is to provide a data archiving processing method and system based on a shutdown system to address the deficiencies in the prior art, and to achieve intelligent classification, effective compression and secure storage of data in the shutdown system, ultimately ensuring the integrity and traceability of the data.
[0004] An embodiment of the present application provides a data archiving processing method based on shutting down a system, the method comprising: According to the data type, access frequency and business importance of the shut-down system, a multi-dimensional classification model based on deep learning is used to classify the data. The multi-dimensional classification model dynamically divides the archiving priority of the data through attention technology and adaptive weight allocation algorithm to obtain the classified data set and its priority label; Inputting the classified data set into a compression and encryption joint processing model based on a neural network to perform data compression and encryption, wherein the joint processing model combines a quantized compression algorithm and homomorphic encryption technology to achieve efficient compression while ensuring data security, thereby obtaining a compressed and encrypted data packet; Distribute the compressed and encrypted data packets to a distributed storage system, and generate a data storage path index using an index construction algorithm based on a graph database, wherein the index construction algorithm optimizes data retrieval efficiency and storage load balancing through dynamic hash mapping and a distributed consistency protocol to obtain a distributed storage index table; According to the distributed storage index table, the data integrity is verified using the blockchain-based archiving verification technology, wherein the archiving verification technology combines smart contracts and distributed ledger technology to monitor the data storage status and access records in real time, generate an archiving verification report, and ensure the reliability and traceability of data archiving.
[0005] Optionally, the data is classified using a multi-dimensional classification model based on deep learning according to the data type, access frequency and business importance of the shutdown system, wherein the multi-dimensional classification model dynamically divides the archiving priority of the data through attention technology and an adaptive weight allocation algorithm to obtain a classified data set and its priority label, including: According to the data type in the shutdown system, a data collection framework based on edge computing is used to obtain multi-source data in real time. The data is filtered for noise and filled with missing values through an adaptive data cleaning algorithm to generate a preliminary standardized data set. For the preliminary standardized data set, a multi-dimensional classification model based on deep learning is used to extract the multi-dimensional features of the data by combining data type, access frequency and business importance. The multi-head attention technology is used to capture the correlation between features of different dimensions and generate preliminary feature representations. For the preliminary feature representation, a priority division method based on an adaptive weight allocation algorithm is used to dynamically allocate the archiving priority of the data in combination with the access frequency and business importance of the data. The accuracy and rationality of the weight allocation are optimized through attention technology to generate preliminary priority labels. For the preliminary priority labels, a data classification method based on clustering algorithm is adopted. The data is divided into different categories according to the data type and priority label. The accuracy and consistency of classification are ensured through dynamic threshold adjustment technology to generate the classified data set and its priority label.
[0006] Optionally, the classified data set is input into a compression and encryption joint processing model based on a neural network to perform data compression and encryption, wherein the joint processing model combines a quantized compression algorithm and homomorphic encryption technology to achieve efficient compression while ensuring data security, and obtains a compressed and encrypted data packet, including: For the classified data set, a quantization compression algorithm based on a neural network is used to convert high-precision data into low-precision representation. The quantization accuracy is dynamically adjusted through adaptive quantization threshold adjustment technology to generate preliminary compressed data while minimizing data information loss. For the preliminary compressed data, an encryption method based on homomorphic encryption technology is used. Combined with the data priority label and security requirements, the data is encrypted and processed. Through lightweight key management technology, the efficiency and security of the encryption process are ensured to generate preliminary encrypted data. For the preliminary encrypted data, a joint optimization method based on neural network is used, combined with quantitative compression algorithm and homomorphic encryption technology, to dynamically adjust compression and encryption parameters, and through a multi-objective optimization algorithm, balance compression efficiency and encryption security to generate preliminary compressed and encrypted data packets; For the preliminary compressed and encrypted data packets, a data integrity verification method based on hash verification is used to ensure that the data is not damaged during the compression and encryption process. Through feedback correction technology, the compression and encryption parameters are dynamically adjusted to generate the final compressed and encrypted data packet.
[0007] Optionally, the compressed and encrypted data packets are distributed to a distributed storage system, and a data storage path index is generated using an index construction algorithm based on a graph database, wherein the index construction algorithm optimizes data retrieval efficiency and storage load balancing through dynamic hash mapping and a distributed consistency protocol to obtain a distributed storage index table, including: For compressed and encrypted data packets, a distribution method based on a distributed storage system is adopted. The data storage path is planned in combination with the data priority label and the load status of the storage node. The dynamic hash mapping algorithm is used to ensure the balance of data distribution and generate a preliminary storage path plan. For the preliminary storage path planning, we use the index construction method based on the graph database to abstract the data storage path into nodes and edges in the graph structure. Through the distributed consistency protocol, we ensure the consistency and reliability of the index and generate the preliminary index structure. For the preliminary index structure, an index optimization method based on dynamic load balancing is adopted. The index structure is dynamically adjusted in combination with the real-time load status of the storage node and the data access frequency. The adaptive hash mapping technology is used to optimize data retrieval efficiency and storage load balancing, and an optimized index structure is generated. For the optimized index structure, an index table generation method based on visualization technology is adopted to map the index structure into a distributed storage index table. Through real-time monitoring technology, the accuracy and consistency of the index table are ensured to generate the final distributed storage index table.
[0008] Optionally, the data integrity is verified by using a blockchain-based archiving verification technology according to the distributed storage index table, wherein the archiving verification technology combines smart contracts and distributed ledger technology to monitor data storage status and access records in real time, generate an archiving verification report, and ensure the reliability and traceability of data archiving, including: According to the distributed storage index table, the blockchain-based archiving verification technology is used in combination with smart contracts to verify the data storage status. The hash verification algorithm is used to ensure that the data is not damaged or tampered with during the storage process, and a preliminary integrity verification result is generated; For the preliminary integrity verification results, a recording method based on distributed ledger technology is used to write the data storage status and access records into the blockchain. Through consensus technology, the immutability and traceability of the records are ensured, and preliminary distributed ledger records are generated; For distributed ledger records, a real-time monitoring method based on smart contracts is used to detect abnormal behavior in combination with data access frequency and storage status, and a preliminary monitoring report is generated through anomaly detection algorithms; For the preliminary monitoring report, a report generation method based on natural language generation technology is used to integrate the integrity verification results, distributed ledger records and monitoring reports into an archive verification report, and the final archive verification report is generated through visualization technology.
[0009] Another embodiment of the present application provides a data archiving processing system based on a shutdown system, the system comprising: A classification module is used to classify data according to the data type, access frequency and business importance of the shutdown system using a multi-dimensional classification model based on deep learning, wherein the multi-dimensional classification model dynamically divides the archiving priority of the data through attention technology and an adaptive weight allocation algorithm to obtain a classified data set and its priority label; A processing module, used for inputting the classified data set into a compression and encryption joint processing model based on a neural network to perform data compression and encryption, wherein the joint processing model combines a quantized compression algorithm and homomorphic encryption technology to achieve efficient compression while ensuring data security, thereby obtaining a compressed and encrypted data packet; An index module is used to distribute the compressed and encrypted data packets to a distributed storage system, and generate a data storage path index using an index construction algorithm based on a graph database, wherein the index construction algorithm optimizes data retrieval efficiency and storage load balancing through dynamic hash mapping and a distributed consistency protocol to obtain a distributed storage index table; The archiving module is used to verify the integrity of the data according to the distributed storage index table using the archiving verification technology based on the blockchain, wherein the archiving verification technology combines smart contracts and distributed ledger technology to monitor the data storage status and access records in real time, generate an archiving verification report, and ensure the reliability and traceability of data archiving.
[0010] Yet another embodiment of the present application provides a storage medium, wherein the storage medium stores a computer program, wherein the computer program is configured to execute any of the above methods when running.
[0011] Yet another embodiment of the present application provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute any of the methods described above.
[0012] Compared with the prior art, the present invention provides a data archiving and processing method based on a shutdown system. According to the data type, access frequency and business importance of the shutdown system, a multi-dimensional classification model based on deep learning is used to classify the data to obtain a classified data set and its priority label; the classified data set is input into a compression and encryption joint processing model based on a neural network to obtain a compressed and encrypted data packet; the compressed and encrypted data packet is distributed to a distributed storage system, and an index construction algorithm based on a graph database is used to generate a data storage path index to obtain a distributed storage index table; according to the distributed storage index table, the data integrity is verified using a blockchain-based archiving verification technology to generate an archiving verification report, thereby realizing intelligent classification, effective compression and secure storage of data in the shutdown system, and ultimately ensuring the integrity and traceability of the data. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 A hardware structure block diagram of a computer terminal for a data archiving processing method based on a shutdown system provided in an embodiment of the present invention; Figure 2 A schematic flow chart of a data archiving processing method based on shutting down a system provided in an embodiment of the present invention; Figure 3 A schematic structural diagram of a data archiving processing system based on a shutdown system provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0014] The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, but should not be construed as limiting the present invention.
[0015] The embodiment of the present invention first provides a data archiving processing method based on a shutdown system. The method can be applied to electronic devices, such as computer terminals, specifically ordinary computers, etc.
[0016] The following describes it in detail by taking running on a computer terminal as an example. Figure 1 The hardware structure block diagram of a computer terminal of a data archiving processing method based on a shutdown system provided by an embodiment of the present invention. Figure 1 As shown, the computer device includes a processor, a memory, and a network interface connected via a system bus, wherein the memory may include a non-volatile storage medium and an internal memory.
[0017] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute any data archiving processing method based on shutting down the system.
[0018] The processor is used to provide computing and control capabilities and support the operation of the entire computer equipment.
[0019] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any data archiving processing method based on the shutdown system.
[0020] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art will understand that Figure 1 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0021] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0022] See also Figure 2 The embodiment of the present invention provides a data archiving processing method based on shutting down the system, which may include the following steps: S201, classifying the data using a multi-dimensional classification model based on deep learning according to the data type, access frequency and business importance of the shutdown system, wherein the multi-dimensional classification model dynamically divides the archiving priority of the data through attention technology and an adaptive weight allocation algorithm to obtain a classified data set and its priority label; In this process, the system uses deep learning algorithms to identify different types of data by analyzing the characteristics of the data, including the structure, content, and frequency of use of the data. The multi-dimensional classification model uses attention technology to focus on the key characteristics of the data, such as the access frequency of the data in historical use and its importance to the business process. At the same time, the adaptive weight allocation algorithm dynamically adjusts the weights of these characteristics based on real-time feedback, so that the model can pay more attention to high-value or frequently accessed data, thereby automatically forming the archiving priority of the data.
[0023] This deep learning-based classification method can significantly improve the efficiency and accuracy of data archiving. By dynamically prioritizing data archiving, the system ensures that important and high-frequency data can be processed and stored safely, avoiding the loss of important information. In addition, this intelligent classification helps to rationally utilize storage resources, reduce unnecessary storage overhead, and thus improve the flexibility and responsiveness of the entire data management. Combined with the application of attention mechanisms and adaptive algorithms, the system can continuously learn and adapt to new data features and business needs to achieve long-term optimization and improvement.
[0024] Specifically, according to the data type in the shutdown system, a data collection framework based on edge computing can be used to obtain multi-source data in real time, and the data can be filtered for noise and filled with missing values through an adaptive data cleaning algorithm to generate a preliminary standardized data set; In this step, the system will adopt an edge computing-based data collection framework to achieve real-time capture of various types of data in the shutdown system. The introduction of edge computing brings data processing closer to the data source, reduces latency, and improves collection efficiency. The system automatically identifies and filters noise in the data, such as outliers and erroneous inputs, through an adaptive data cleaning algorithm, and implements a missing value filling strategy to fill missing data by using the mean or median of the surrounding data, thereby ensuring that the generated data set has high completeness and accuracy.
[0025] This method of data collection and cleaning can maximize data quality and provide a solid foundation for subsequent data analysis and classification. By effectively filtering and cleaning data during the collection phase, the system can reduce data redundancy and improve the efficiency of subsequent processing, thereby effectively shortening the overall data processing cycle. High-quality standardized data sets are more conducive to the training of deep learning models, thereby improving the classification accuracy of multi-dimensional classification models.
[0026] In this step, the system first deploys edge computing nodes to obtain multi-source data in the shutdown system. The advantage of edge computing is that it can place data processing tasks as close to the source of data generation as possible, thereby reducing network latency and improving the real-time nature of data collection. The system designs a data collection framework that supports access to different data sources, including sensor data, user input, and historical records. For example, when the system collects real-time monitoring data (such as temperature, humidity, flow, etc.) from multiple sensors, it can aggregate these data to edge nodes in a stream for aggregation and preliminary processing.
[0027] After data is collected, the system will use an adaptive data cleaning algorithm to process the acquired data. The algorithm automatically identifies and filters noise by analyzing the characteristics of the data set. For example, during the collection process, if the system finds that some sensor data is abnormal (such as higher or lower than the preset reasonable range), the data will be marked as noise and filtered out. At the same time, for missing values, the algorithm will fill in the missing data through interpolation methods or statistical methods (such as mean filling, median filling, etc.) to ensure the integrity and reliability of the generated preliminary standardized data set.
[0028] After data cleaning, the generated preliminary standardized data set will be formatted into a consistent structure for subsequent processing. This structure usually uses a common data format (such as CSV or JSON) to ensure that each record contains necessary fields, such as data type, timestamp, and acquisition source. This data standardization process not only improves data consistency, but also lays a good foundation for subsequent analysis. High-quality data sets are the key to the effectiveness of subsequent deep learning models.
[0029] For the preliminary standardized data set, a multi-dimensional classification model based on deep learning is used to extract the multi-dimensional features of the data by combining data type, access frequency and business importance. The multi-head attention technology is used to capture the correlation between features of different dimensions and generate preliminary feature representations. In this step, the system will apply a multi-dimensional classification model based on deep learning to extract features for the preliminary standardized data set. Combining data type, access frequency, and business importance, the model will analyze the multi-dimensional features of each piece of data and capture the complex relationship between features through a multi-head attention mechanism. For example, some data may appear important in high-frequency access situations, while other data has its unique value in key business links. Multi-head attention can ensure that the model fully considers this information and generates more expressive feature representations.
[0030] The main significance of this step is to improve the model's ability to understand data features. By using deep learning models and multi-head attention mechanisms, the system can better identify potential patterns and relationships in the data, which is crucial for subsequent classification decisions. The generated preliminary feature representation reflects the true value of the data, allowing the model to perform more reasonable archiving priority division based on multiple dimensions, improving the accuracy and reliability of the overall classification.
[0031] After inputting the preliminary standardized data set into the multi-dimensional classification model based on deep learning, the system will first embed the data and convert it into a tensor form suitable for model processing. At this point, the model will analyze and extract potential multi-dimensional features based on data type, access frequency, and business importance. For example, the system may analyze the user's access behavior data, identify the usage habits of different user groups for specific data, and thus extract features related to user behavior.
[0032] Next, the system uses a multi-head attention mechanism to process the extracted multi-dimensional features. The multi-head attention mechanism can focus on the relationship between data features from different perspectives and improve the model's expressiveness. For example, for a set of features, some of which are directly related to business importance, while others may better reflect the frequency of data use. This mechanism allows the model to process this information in parallel in multiple "heads" and gradually learn the importance of each feature in different contexts.
[0033] Finally, after multi-head attention processing, the model will generate preliminary feature representations. These feature representations can be regarded as a multi-dimensional comprehensive description of each piece of data, including the data representation information and the evaluation of its importance. These feature descriptions will provide the necessary information basis for the subsequent dynamic division of priorities, ensuring the rationality and accuracy of the final classification.
[0034] For the preliminary feature representation, a priority division method based on an adaptive weight allocation algorithm is used to dynamically allocate the archiving priority of the data in combination with the access frequency and business importance of the data. The accuracy and rationality of the weight allocation are optimized through attention technology to generate preliminary priority labels. In this step, the system will apply an adaptive weighting algorithm-based method to the initially generated feature representation to prioritize the data. The system dynamically adjusts the weight of each data by analyzing its access frequency and business importance, and prioritizes the data based on these weights. The application of attention technology enables the system to focus on the features most relevant to data archiving, thereby achieving more reasonable archiving priorities and ensuring that high-value data is processed quickly.
[0035] Through this dynamic priority division, the system significantly improves the decision-making efficiency of data archiving, ensures that important information can be processed and stored in a timely manner, and reduces the business losses that may be caused by information delays. At the same time, accurate priority allocation can not only optimize the use of storage resources, but also improve the system's responsiveness to user needs and increase business flexibility and adaptability.
[0036] Based on the preliminary feature representation, the system will use an adaptive weight assignment algorithm to dynamically calculate the archiving priority of each data item. This process first combines the access frequency and business importance of the data and passes its respective weights as input to the algorithm. For example, some data may have a higher priority because it is frequently accessed, while other data may also be given a higher weight because it plays an important role in key decisions.
[0037] In the process of dynamically calculating weights, the system also uses attention technology to optimize the weight distribution of features. By paying attention to the influence of different features in data archiving, the model can assign priorities more accurately. The system will traverse all features, identify the features that are directly related to the archiving decision, and assign them corresponding attention scores, and finally calculate the comprehensive priority label for each data. For example, if the access frequency of a piece of data surges within a specific period, the model will automatically detect this change and dynamically increase its priority.
[0038] Finally, the generated preliminary priority labels will serve as a guide for subsequent classification, ensuring that high-importance and high-frequency data are processed in a timely manner. This process not only ensures the flexibility of the algorithm, but also ensures the scientific nature of priority allocation, laying a guiding foundation for the effectiveness of the entire data management system.
[0039] For the preliminary priority labels, a data classification method based on clustering algorithm is adopted. The data is divided into different categories according to the data type and priority label. The accuracy and consistency of classification are ensured through dynamic threshold adjustment technology to generate the classified data set and its priority label.
[0040] In this step, the system will use a data classification method based on clustering algorithms to divide data into different categories based on preliminary priority labels and data types. Clustering algorithms can automatically group similar data based on data characteristics and priorities. This classification not only takes into account the business importance of the data, but also affects subsequent data management strategies. Dynamic threshold adjustment technology ensures that during the classification process, appropriate thresholds can be adaptively set for different categories to optimize the classification effect.
[0041] The key to this step is to ensure the accuracy and consistency of data classification, thus laying a good foundation for the subsequent data archiving and processing. By rationally classifying the data, the system can archive and store the data in a targeted manner, avoiding complicated data management and processing processes and improving the efficiency of data processing. This classification method also facilitates subsequent analysis and retrieval, making the use of data more flexible and efficient.
[0042] In this step, the system will use a clustering algorithm-based method to combine the adopted preliminary priority labels with the data type to automatically classify the data. The system first selects a clustering algorithm suitable for the current data set (such as K-Means, hierarchical clustering, etc.) and applies it to a data set that meets certain conditions to discover the inherent connections and similarities between the data. For example, the system may cluster all data marked as high priority, identify the commonalities between them, and thus divide the data into different groups.
[0043] As clustering progresses, the system also uses dynamic threshold adjustment technology to ensure the accuracy and consistency of classification results. The dynamic threshold is set based on the classification effect of real-time evaluation. The system will continuously monitor the clustering results and dynamically adjust the threshold according to the amount of data and feature similarity in different categories to optimize the classification results. For example, if the amount of data in a category is too small, the system may lower the clustering threshold of this category to include more data and ensure that important information is not cut off.
[0044] Finally, after processing by clustering algorithm and dynamic threshold adjustment, the system will generate a classified data set and its priority label. This result not only provides a basis for subsequent data archiving and storage, but also provides a good structure for future data retrieval and analysis, making overall data management more efficient and systematic.
[0045] S202, inputting the classified data set into a compression and encryption joint processing model based on a neural network to perform data compression and encryption, wherein the joint processing model combines a quantized compression algorithm and homomorphic encryption technology to achieve efficient compression while ensuring data security, thereby obtaining a compressed and encrypted data packet; In this step, the classified data set is input into the neural network-based compression and encryption joint processing model for data compression and encryption. The joint processing model will use a combination of quantization compression algorithm and homomorphic encryption technology to ensure that the integrity and security of the data are not lost while compressing the data volume. The quantization compression algorithm converts high-precision data into low-precision representation by reducing the precision of the data, thereby reducing the storage space occupied by the data. Homomorphic encryption technology allows the encrypted data to remain operational during processing, that is, calculations can still be performed in the encrypted state, avoiding the risk of exposing the original data during the exchange of sensitive data. This allows the data to be effectively compressed without loss, accelerating the efficiency of data transmission and storage.
[0046] This joint processing mechanism of compression and encryption has important practical significance. On the one hand, efficient data compression can significantly reduce storage costs and bandwidth requirements, especially when processing large-scale data, which can quickly improve data transmission efficiency and storage capacity; on the other hand, homomorphic encryption can be used to ensure data privacy and security, effectively preventing data from being stolen or tampered with during processing, and ensuring the security of data during storage and transmission. The implementation of this mechanism not only ensures the security of data when enterprises process data in shutdown systems, but also improves the efficiency of data processing and enhances the overall performance and availability of the system.
[0047] Specifically, a quantization compression algorithm based on a neural network can be used for the classified data set to convert high-precision data into low-precision representation, and the quantization accuracy can be dynamically adjusted through the adaptive quantization threshold adjustment technology to generate preliminary compressed data while minimizing the loss of data information. In this step, the system quantizes and compresses the classified data set. The core of the quantization compression algorithm is to convert high-precision data (such as floating-point numbers) into low-precision representations (such as fixed-point numbers), thereby effectively reducing the storage space of the data. For example, a high-precision measurement value may need to be stored using a 32-bit floating-point number, but after quantization, it can be compressed into a 16-bit or 8-bit fixed-point number, thereby saving storage overhead. This compression method is not limited to numerical data, but can also be applied to data in multiple fields such as images and audio, ensuring that data processing is not limited by high storage costs.
[0048] In this process, the system uses adaptive quantization threshold adjustment technology to dynamically adjust the quantization accuracy according to the characteristics of each data feature and its distribution in the data set. For example, for some important features, the system may choose a higher quantization accuracy, while for features with less impact on the business, a lower quantization accuracy is used. In this way, the system can effectively find the best balance between compression rate and information loss, and maximize the preservation of data information integrity.
[0049] Through quantitative compression, the system can significantly reduce the storage space occupied by data, thereby improving processing efficiency when facing massive data. Especially in scenarios where long-term data storage or real-time data analysis is required, quantitative technology can reduce the storage burden and improve the system's response speed. At the same time, the strategy of dynamically adjusting the quantization accuracy ensures that the basic validity of the data is not affected while compressing, which is crucial for subsequent data analysis and processing. It can maintain the availability and accuracy of the data, and ensure the flexibility and efficiency of the entire system in data management.
[0050] In this step, the system first receives the classified data set and applies a neural network-based quantization compression algorithm to each data feature. The algorithm aims to convert high-precision data (such as floating-point numbers) into low-precision representations (such as fixed-point numbers) to reduce the storage space occupied by the data. For example, assuming that the temperature data recorded by a sensor is in the form of floating-point numbers (such as 23.456°C), after quantization compression, the data can be converted into a lower-precision fixed-point number form (such as 23°C), which not only reduces storage requirements but also improves the efficiency of subsequent data processing.
[0051] In order to ensure that information loss is minimized during the compression process, the system uses adaptive quantization threshold adjustment technology, which monitors changes in data features in real time and automatically adjusts the quantization accuracy based on the characteristics of different data. For example, for some feature data with large changes, the system may choose to retain a higher quantization accuracy, while for data with small changes, the quantization accuracy can be reduced. This dynamic adjustment mechanism enables the system to flexibly optimize the compression strategy according to actual conditions, ensuring that the generated compressed data can achieve efficient storage while retaining key information.
[0052] Finally, the quantized and compressed data is aggregated to form a preliminary compressed data set. These compressed data not only excel in saving storage space, but also lay the foundation for subsequent encryption steps. The low-precision representation of compressed data ensures efficient storage and transmission, allowing subsequent processing steps to proceed more quickly, and overall improving the system's response speed and processing capabilities.
[0053] For the preliminary compressed data, an encryption method based on homomorphic encryption technology is used. Combined with the data priority label and security requirements, the data is encrypted and processed. Through lightweight key management technology, the efficiency and security of the encryption process are ensured to generate preliminary encrypted data. In this step, the system will perform homomorphic encryption on the preliminary data after quantization and compression. Unlike traditional encryption methods, homomorphic encryption allows operations to be performed on encrypted data without first decrypting it, which keeps the data secure during processing. During implementation, the system will apply different encryption strategies to different types of data based on the data's priority label and security requirements. For example, for highly sensitive data, the system may choose a stronger encryption algorithm, while for less sensitive data, a more basic encryption method is used to optimize processing efficiency while meeting security requirements.
[0054] In addition, lightweight key management technology plays an important role in this process. The system generates and manages the keys required for the encryption process, ensuring that the distribution and storage of keys meet security standards while providing efficient key operations. For example, a hash algorithm can be used to generate a corresponding hash value for a key, thereby verifying its validity without exposing the key. This process not only protects the security of data, but also improves the flexibility and response speed of the system in data encryption processing.
[0055] Homomorphic encryption builds a bridge between data security and processing power, allowing users to perform necessary calculations while maintaining data privacy. Such technical applications are particularly suitable for use in cloud computing or multi-party collaboration environments, and can effectively avoid security risks caused by data leakage. In addition, through lightweight key management, the encryption process of the system can be carried out efficiently, and users can enjoy a smooth operation experience while enjoying security protection. Ultimately, the generated preliminary encrypted data provides security for subsequent data storage and retrieval, ensuring the integrity and reliability of the data archiving process.
[0056] In this step, the system will apply homomorphic encryption technology to the preliminary compressed data for encryption. The core advantage of this technology is that it allows calculations to be performed on encrypted data without decryption, thereby protecting the privacy of the data. For example, for a set of data containing user sensitive information, the use of homomorphic encryption can ensure that the data will not leak its original content during processing, while still being able to perform necessary calculations, such as statistics or analysis. This makes this technology particularly suitable for scenarios where sensitive data needs to be securely processed in a cloud computing environment.
[0057] During the encryption process, the system will dynamically select the appropriate homomorphic encryption strategy based on the priority label and security requirements of each piece of data. For example, for highly sensitive financial transaction data, the system may choose to use a complex encryption process based on encryption algorithms (such as Paillier or RGH) to ensure the security of the data after encryption. At the same time, lightweight key management technology can improve the efficiency of the encryption process and ensure the efficiency of key generation, distribution, and use. For example, the system can use a random number generator to generate encryption keys and use hash technology to verify the validity and security of the keys, thereby greatly reducing the complexity of operation and maintenance.
[0058] Ultimately, the homomorphically encrypted data will form a preliminary encrypted data packet, ensuring security and privacy during storage and transmission. The encryption processing at this stage provides a secure and reliable data foundation for the subsequent joint optimization stage, ensuring that the data will not be threatened by damage or leakage during processing, thereby improving the overall security architecture of the system.
[0059] For the preliminary encrypted data, a joint optimization method based on neural network is used, combined with quantitative compression algorithm and homomorphic encryption technology, to dynamically adjust compression and encryption parameters, and through a multi-objective optimization algorithm, balance compression efficiency and encryption security to generate preliminary compressed and encrypted data packets; In this step, the system will jointly optimize the preliminary encrypted data, with the goal of balancing the compression efficiency of the data with the security of encryption. The joint optimization method based on neural networks can dynamically adjust the parameters of the quantitative compression and encryption algorithms by learning the characteristics of historical data. For example, by analyzing the characteristics of the current data and its access patterns, the system can identify which data should use a higher compression rate to save storage, and which data requires more sophisticated encryption to ensure security.
[0060] By introducing a multi-objective optimization algorithm, the system can balance compression efficiency and encryption strength and select the parameter settings that best suit the current data characteristics. This algorithm usually considers multiple target outputs at the same time, such as achieving the best compression rate while ensuring that the data is not distorted, and dynamically enhancing encryption strength when necessary to adapt to changes in external risks. This intelligent optimization strategy ensures flexibility in data processing and can adapt to various operating conditions in real time.
[0061] Through this joint optimization, the system can ensure the security of information and reduce the potential risk of data leakage while ensuring efficient data compression. This flexible adjustment capability not only adapts to various types of data needs, but also provides scalability for future data processing. Therefore, the generated preliminary compressed and encrypted data packets not only save storage space, but also improve the security and reliability of data during transmission and storage, which plays an important role in promoting the efficiency of the entire data processing system.
[0062] In this step, the system applies a neural network-based joint optimization method to the preliminary encrypted data. This method aims to dynamically adjust the quantitative compression and homomorphic encryption parameters by learning and analyzing data characteristics to achieve a balance between compression efficiency and encryption security. Specifically, the system can use the trained neural network model to evaluate the characteristics of the current data and determine the most appropriate compression and encryption parameters based on this.
[0063] For example, the system may analyze the access frequency, sensitivity, and compression effect of different data types. For frequently accessed low-sensitivity data, the system can choose a higher compression rate to improve storage efficiency, while for sensitive data, it gives priority to enhancing encryption strength. In this process, combined with the application of multi-objective optimization algorithms, the system can simultaneously process multiple optimization goals, such as compression rate, processing speed, and encryption strength. Through this multi-level optimization, the system can find the best balance between different processing requirements, making the processing process more efficient and secure.
[0064] Finally, the generated preliminary compressed and encrypted data packet achieves excellent storage performance on the basis of ensuring data integrity and security. This data packet provides an effective basis for the subsequent integrity verification steps, ensuring that the entire data processing process achieves the optimal solution between efficiency and security, and improving the system's adaptability in data management.
[0065] For the preliminary compressed and encrypted data packets, a data integrity verification method based on hash verification is used to ensure that the data is not damaged during the compression and encryption process. Through feedback correction technology, the compression and encryption parameters are dynamically adjusted to generate the final compressed and encrypted data packet.
[0066] In this step, the system will verify the data integrity of the preliminary compressed and encrypted data packet to ensure that the data has not been damaged or tampered with during the entire compression and encryption process. To this end, the system will use a hash-based verification method to first perform a hash operation on the preliminary data packet to generate a fixed-length hash value to represent the integrity of the data. When the data packet is created, the system will compare the hash value with the uncompressed and unencrypted original data to verify that the data remains intact and secure during the processing. This process is critical, especially when dealing with sensitive data, and can effectively prevent the risks caused by data loss or tampering.
[0067] At the same time, the system will also implement feedback correction technology to dynamically adjust compression and encryption parameters. By analyzing the results of hash verification, the system can identify potential problems in data processing. For example, if the hash value does not match, the system will trigger a feedback mechanism to automatically adjust the quantitative compression ratio or encryption strength to find a solution. This mechanism allows the system to perform self-correction and optimization during data processing, thereby ensuring that the final generated data packet is both secure and efficient.
[0068] This integrity verification process provides a strong layer of security for data processing, ensuring that data will not be damaged or tampered with during the compression and encryption process, laying a solid foundation for subsequent data transmission and storage. By implementing hash verification and dynamic feedback correction technology, the system can achieve efficient self-adjustment, making the data processing process more flexible and adaptable. This not only enhances data security, but also improves the reliability of the entire data archiving process, providing strong support for the subsequent management and use of data.
[0069] In this step, the system uses a hash-based data integrity verification method to ensure that the initial compressed and encrypted data packet has not been damaged during processing. Hash verification is performed by generating a fixed-length hash value for the data content, which can uniquely represent the data. When the data packet is created, the system calculates its hash value and compares it with the hash value of the original data during subsequent processing to confirm its integrity and consistency. This process can effectively prevent data from being damaged or tampered with due to operational errors or external attacks.
[0070] In addition, if the result of the hash check shows that the data is abnormal or inconsistent, the system will immediately enable feedback correction technology. This technology will analyze the possible causes of hash mismatch and correct the compression and encryption parameters based on the results of dynamic adjustment. For example, assuming that the hash value of a data packet does not match the original value, the system can reduce the compression ratio to retain more information, or increase the encryption strength to enhance data security. This feedback mechanism ensures that the integrity and security of the data are maintained under any circumstances, improving the robustness of the system.
[0071] Finally, after integrity verification and necessary adjustments, the system will generate the final compressed and encrypted data package. This data package can ensure stability and security in storage and transmission, laying a solid foundation for subsequent data management and use. This step not only enhances data security, but also improves the reliability of the entire data processing process, ensuring efficient connection of all links, allowing users to perform data operations more confidently.
[0072] S203, distributing the compressed and encrypted data packets to a distributed storage system, and using an index construction algorithm based on a graph database to generate a data storage path index, wherein the index construction algorithm optimizes data retrieval efficiency and storage load balancing through dynamic hash mapping and a distributed consistency protocol to obtain a distributed storage index table; In this method, the compressed and encrypted data packets are first transmitted to the distributed storage system to ensure the high availability and reliability of the data. At this point, the system generates a data storage path index through an index construction algorithm based on a graph database. Specifically, the algorithm implements dynamic hash mapping technology to distribute data on different storage nodes, thereby preventing a node from becoming a bottleneck due to excessive load. At the same time, using a distributed consistency protocol, the system ensures data consistency in all storage nodes to avoid data redundancy or inconsistency problems, and ensures efficient and stable data access. The resulting storage path index not only covers the location information of the data storage, but also provides efficient support for subsequent data retrieval.
[0073] By distributing compressed and encrypted data packets to the distributed storage system and generating corresponding storage path indexes, the access efficiency and availability of data are greatly improved. In summary, the index construction algorithm based on the graph database optimizes the data retrieval efficiency, can quickly and accurately locate frequently read data, and improves the user experience. In addition, the combination of dynamic hash mapping and distributed consistency protocol ensures the performance of the overall system in load balancing, so that the security and effectiveness of data during storage can be guaranteed. This method is suitable for large-scale data storage and management scenarios, especially enterprise information systems involving sensitive information.
[0074] Specifically, a distribution method based on a distributed storage system can be used for compressed and encrypted data packets, and the data storage path can be planned in combination with the data priority label and the load status of the storage node. The dynamic hash mapping algorithm can be used to ensure the balance of data distribution and generate a preliminary storage path plan. In this step, the system plans the data storage path through a distributed storage-based distribution method, combining the data priority label and the storage node load status. First, the system analyzes each compressed and encrypted data packet and evaluates its priority label to ensure that important data can be stored on storage nodes with stronger performance. When determining the selection of storage nodes, the system also monitors the load status of each storage node in real time to avoid performance degradation due to overloading of a node.
[0075] This dynamic planning of storage paths ensures balanced distribution of data on different storage nodes, thereby optimizing data access performance. At the same time, comprehensive consideration of priority tags enables faster response speeds for key business data, improving overall system performance and reliability. With this approach, enterprises can effectively manage and store large-scale data, ensure that important data is available at any time, and improve data processing efficiency and scalability.
[0076] In this step, the system first analyzes the received compressed and encrypted data packets and extracts their priority tags. Priority tags are usually based on the importance, access frequency, and security requirements of the data. For example, financial transaction data may have a higher priority, while log files may have a lower priority. The system assigns a storage node to each data packet. The selection of storage nodes takes into account not only the priority of the data, but also the current load of each node. In this way, the system can prioritize sending important data packets to nodes with lower loads to ensure fast access and processing of data.
[0077] To achieve balanced distribution of data storage paths, the system uses a dynamic hash mapping algorithm. This algorithm dynamically calculates hash values based on the characteristics of data packets and distributes the data packets to the corresponding storage nodes. For example, suppose there are three storage nodes A, B, and C, and data packets X, Y, and Z to be stored. The system calculates the hash values of these data packets and selects the most appropriate node for storage based on the load. This dynamic calculation and selection mechanism ensures uniform distribution of data in the storage system and avoids performance bottlenecks caused by overloading of certain nodes.
[0078] Finally, the generated preliminary storage path plan will include the target storage node and storage path information of each data packet. This information will be recorded for subsequent data retrieval and management. In this way, the system not only achieves efficient data storage, but also improves the overall system performance, ensuring scalability and reliability in large-scale data processing scenarios.
[0079] For the preliminary storage path planning, we use the index construction method based on the graph database to abstract the data storage path into nodes and edges in the graph structure. Through the distributed consistency protocol, we ensure the consistency and reliability of the index and generate the preliminary index structure. In this step, the system builds a graph database index for the preliminary storage path. Specifically, the system treats the storage path as a graph structure containing multiple nodes and edges to facilitate real-time access and query of data. Each node represents a storage location, while the edge represents the connection relationship between the storage nodes. By dividing these structures, the system can effectively organize and manage storage paths and improve data retrieval speed. At the same time, combined with the distributed consistency protocol, the reliability of the index is guaranteed to avoid data inconsistencies that may occur in a distributed environment.
[0080] Through the index construction method of the graph database, the management of data storage paths becomes more efficient and flexible. The formation of the graph structure enables data queries to be performed through an effective graph traversal algorithm, thereby significantly improving the speed of data retrieval. In addition, the implementation of the distributed consistency protocol ensures the reliability of the index structure, prevents index inconsistencies caused by node failures, and improves the fault tolerance and availability of the system. This method provides an innovative solution for data management and retrieval in a dynamic storage environment.
[0081] In this step, the system converts the preliminary storage path planning into an index structure in the graph database. Specifically, each storage node is regarded as a node in the graph, and the connection relationship between nodes (i.e., data storage path) is regarded as an edge in the graph. The advantage of this graph structure is that it can intuitively display the relationship between storage paths and data, making storage management clearer and more efficient.
[0082] During the implementation process, the system will extract the information of the storage nodes and build connections between the nodes. For example, if data packet X is stored in node A, and data packet Y needs to be read from node B, the system will form an edge from A to B in the graph. This structure allows you to quickly find the path to the data storage when querying data, significantly reducing the time for data retrieval. In addition, the system will use distributed consistency protocols (such as Paxos or Raft) to manage indexes to ensure the consistency of index information between nodes. Each time the index is updated, the protocol will ensure the coordination and consistency of the modification operations, thereby avoiding data inconsistencies caused by node failures or network delays.
[0083] As a result, the generated preliminary index structure can ensure that the overall index information is reliable and consistent. The implementation of this graph structure not only improves the visualization of the data storage path, but also provides important support for subsequent data access, greatly improving the efficiency and accuracy of data retrieval.
[0084] For the preliminary index structure, an index optimization method based on dynamic load balancing is adopted. The index structure is dynamically adjusted in combination with the real-time load status of the storage node and the data access frequency. The adaptive hash mapping technology is used to optimize data retrieval efficiency and storage load balancing, and an optimized index structure is generated. In this step, the system uses an index optimization method based on dynamic load balancing to adjust the preliminary index structure. First, the system monitors the real-time load status of each storage node and the access frequency of each data to determine which nodes are currently overloaded and which data are frequently accessed. Based on this information, the system will dynamically optimize the index structure so that frequently accessed data can be preferentially located on nodes with lower loads, in order to achieve storage load balancing and improve data retrieval efficiency.
[0085] This method of dynamically adjusting the index structure can effectively disperse storage pressure and prevent a few storage nodes from affecting the performance of the entire system due to overload. In addition, through adaptive hash mapping technology, data retrieval becomes faster and more efficient, especially when facing frequently accessed large data sets, which greatly improves the user experience. In summary, this optimization process not only improves storage performance, but also ensures the rational use of resources, providing strong support for the sustainable development of the system.
[0086] In this step, the system implements a dynamic load balancing optimization method for the preliminary index structure. The system first monitors the real-time load status of the storage node, including CPU usage, memory usage, and the current amount of data access requests. For example, when the CPU load of node A is high and node B is relatively idle, the system will recognize that node B should receive more data storage requests. Through such dynamic adjustments, the system can effectively avoid performance degradation caused by overload on some nodes.
[0087] At the same time, the system will also optimize the index structure based on the frequency of data access. For frequently accessed data, the system will adjust its storage path so that it is stored on nodes with lighter loads, thereby improving access speed. For example, if a specific data file has been read many times recently, the system can choose to migrate it from a node with higher load to a node with lower load, in order to balance the overall load while ensuring access speed.
[0088] Finally, the index structure was optimized through adaptive hash mapping technology. This process ensures fast data retrieval and consistency, and achieves significant results in balancing storage load. The optimized index structure enables data to be retrieved quickly when accessed, thereby improving the user experience, especially in scenarios with high concurrent access.
[0089] For the optimized index structure, an index table generation method based on visualization technology is adopted to map the index structure into a distributed storage index table. Through real-time monitoring technology, the accuracy and consistency of the index table are ensured to generate the final distributed storage index table.
[0090] In this step, the system maps the optimized index structure to a distributed storage index table for further data management and retrieval. Through visualization technology, users can view and manage the index structure in a graphical way, which is convenient for understanding the storage path and data distribution. The system will present the visualization results as an easy-to-read and understand table or graph, showing the status of each storage node, the stored data items, and the access frequency. At the same time, through real-time monitoring technology, the system can continuously detect the accuracy and consistency of the index table to ensure that there will be no errors or omissions in the data access process.
[0091] By displaying the index structure in a visual way, users can not only intuitively understand the status of the storage system, but also quickly identify potential problems, such as excessive load on a storage node or abnormal data access. In addition, the introduction of real-time monitoring technology ensures the continuous update and accuracy of the index table, thereby improving the efficiency and reliability of data retrieval. This method improves the manageability and ease of use of the storage system, promotes the security and efficiency of data management, and enables users to use storage resources more conveniently.
[0092] In this step, the system converts the optimized index structure into a visual distributed storage index table. Visualization graphically displays the status of storage nodes and data storage paths, allowing managers to easily understand and monitor the overall operation of the storage system. Specifically, the system may use some visualization tools to present information such as node load, stored data and its priority as clear charts or graphs to facilitate user decision-making and operations.
[0093] During the implementation process, the system will integrate real-time monitoring technology to continuously track the status of each storage node. This includes not only the node load, but also the access frequency and storage status of the data. For example, if the load of a node is continuously too high, the system will quickly issue an alarm and mark the node in the visual interface for the administrator to pay attention. This real-time monitoring enables the administrator to adjust the storage strategy immediately, thereby improving the stability and effectiveness of data storage.
[0094] Ultimately, the generated distributed storage index table will provide a solid foundation for system management and efficient data access. Its accuracy and consistency ensure that there are no delays or errors when performing data retrieval operations, improving the user experience. Through this method, users can easily monitor the usage of storage resources and make timely decisions, greatly enhancing the manageability and efficiency of the system.
[0095] S204, according to the distributed storage index table, the data integrity is verified by using the blockchain-based archiving verification technology, wherein the archiving verification technology combines smart contracts and distributed ledger technology to monitor the data storage status and access records in real time, generate an archiving verification report, and ensure the reliability and traceability of data archiving.
[0096] In this method, based on the distributed storage index table, the system uses blockchain-based archiving verification technology to verify the integrity of the data. Specifically, the system transmits the information in the distributed storage index table to the blockchain network, and verifies the storage status of the data in real time in combination with smart contracts. Smart contracts are automatically executed contract codes that can automatically execute corresponding operations when specific conditions are met. Through smart contracts, the system can check the integrity of each data entry during the storage process to ensure that the data has not been tampered with or damaged. At the same time, using the distributed ledger technology of blockchain, the system can also record all data storage status and access records to ensure that these records cannot be tampered with, providing a reliable basis for subsequent auditing and tracing.
[0097] By using blockchain technology to verify data integrity, the security and reliability of data during storage and access are ensured. This method greatly enhances the traceability of data and can quickly trace back to the true state of data when data problems occur, providing strong support for timely problem handling. In addition, this technical architecture that combines smart contracts and distributed ledgers makes the data verification process efficient, transparent and automated, reduces the need for manual operations, and improves the security and management efficiency of the overall system. This mechanism is particularly important for industries with large amounts of sensitive data, such as finance and healthcare.
[0098] Specifically, according to the distributed storage index table, the blockchain-based archiving verification technology can be used in combination with smart contracts to verify the data storage status. Through the hash verification algorithm, it can be ensured that the data has not been damaged or tampered with during the storage process, and a preliminary integrity verification result can be generated; In this step, the system first extracts relevant data storage information from the distributed storage index table and generates a unique hash value for each piece of data. The hash verification algorithm converts the content of the data into a hash value of fixed length. Any change in the data will cause a change in the hash value, so this method can effectively detect the consistency and integrity of the data. The system will use smart contracts to compare these hash values with the storage status to determine whether the data has been tampered with or damaged during the storage process.
[0099] The implementation of this step ensures the integrity and security of the data. Once a hash value mismatch is found, the system will immediately trigger an alarm and record the specific abnormal event, providing clues for subsequent tracing and processing. This method makes data storage management more transparent and reliable, providing a strong guarantee for data protection, especially for business scenarios that require a high level of security, such as financial transactions and storage of personal privacy data.
[0100] In this step, the system first extracts detailed information about each piece of data from the distributed storage index table, including the data's identity, storage location, and current hash value. This process is usually performed while the data is being uploaded, and the system records the hash value of each data object for subsequent integrity verification. Hash value generation algorithms, such as SHA-256, calculate the data content and return a fixed-length string to ensure that even minor modifications will result in significant changes in the hash value. This enables the system to monitor the integrity of the data in real time during storage, and if the data is found to have been tampered with, inconsistencies can be immediately identified.
[0101] Next, the system will compare the hash value of each piece of data with its storage record. Through smart contracts, this verification process can be seamlessly and automatically executed, following pre-defined contract rules. For example, the smart contract will set that every time the data is accessed, the system automatically calculates the current hash value of the data and compares it with the original hash value stored on the blockchain. If the two are consistent, the system will record the verification result and generate a preliminary integrity verification result; if they are inconsistent, an alarm will be triggered, marking the data as a potential risk and requiring further review.
[0102] Through this process, the system ensures the integrity and security of the data since its creation. This not only improves the transparency of data management, but also enhances users' trust in the security of data storage. Especially when dealing with sensitive information (such as medical data or financial records), the use of blockchain technology for verification can effectively prevent data tampering and achieve a high degree of security.
[0103] For the preliminary integrity verification results, a recording method based on distributed ledger technology is used to write the data storage status and access records into the blockchain. Through consensus technology, the immutability and traceability of the records are ensured, and preliminary distributed ledger records are generated; In this step, the system stores the preliminary integrity verification results and related data storage status and access records in the blockchain. In particular, the system will call on the distributed ledger technology to record each data access and its status on the blockchain. The consensus mechanism (such as PoW, PoS, etc.) is used to ensure the validity and consistency of each record, thereby preventing data inconsistency caused by malicious tampering. In this way, the storage status and access records of all data will form a complete and tamper-proof historical track.
[0104] Through this recording method, the system provides strong traceability, and any access and change to data can be subsequently traced and audited. This is especially important for industries with high legal, compliance and audit requirements (such as banking, insurance, etc.). At the same time, due to the decentralized nature of blockchain, data storage and management become more transparent, which improves users' trust in the system. At the same time, this method also helps to achieve efficient data governance and security compliance.
[0105] In this step, the system integrates the information related to the preliminary integrity verification results into the blockchain to record the data storage status and access behavior. Specifically, the system uses distributed ledger technology to write each data verification, storage status and its corresponding hash value, timestamp and visitor information into the blockchain. Through such records, all access and verification information will be retained on the blockchain in a transparent and tamper-proof form, ensuring that future audits and traceability work can rely on reliable evidence.
[0106] By using consensus mechanisms (such as Proof of Work or Proof of Stake), the validity and integrity of each record will be agreed upon among network nodes. When a data access request occurs, the smart contract will automatically trigger the recording mechanism and add the newly generated status data to the transaction pool to be confirmed. Only after verification and confirmation by multiple nodes will these records be added to the blockchain. In this way, the system ensures the authenticity and credibility of data records, and any administrator or relevant personnel can view and verify the data access and storage status in any time period at any time, thereby ensuring the transparency of the entire system.
[0107] For example, suppose a financial transaction system processes user transaction data. When each transaction is stored, the system will record the specific content of the transaction (such as the amount, sender and receiver accounts, etc.), calculate the hash value of the transaction through a hash algorithm, and then write the information into the blockchain. This not only ensures that the transaction record is accurate and cannot be tampered with, but also provides a complete chain of evidence for subsequent audits, so that any disputes can be quickly verified.
[0108] For distributed ledger records, a real-time monitoring method based on smart contracts is used to detect abnormal behavior in combination with data access frequency and storage status, and a preliminary monitoring report is generated through anomaly detection algorithms; In this step, the system uses smart contracts to achieve real-time monitoring, combining data access frequency and storage status to dynamically detect abnormal behavior. Smart contracts can set conditional trigger mechanisms. For example, if the access frequency of a certain data increases suddenly, or the load status of a storage node is abnormal, the system will automatically execute the preset detection task and conduct in-depth analysis. In this way, the system can capture potential security threats or abnormal operations in a timely manner.
[0109] The implementation of this process improves the security of the system and enables it to respond quickly to possible security risks. By timely detecting and handling abnormal behaviors, the risk of data leakage or tampering can be effectively reduced. In addition, monitoring reports can provide system managers with valuable information, which is helpful for subsequent system optimization and security management decisions. This intelligent monitoring method not only effectively improves data security, but also provides enterprises with a complete data protection solution.
[0110] In this step, the system uses the real-time monitoring capabilities of smart contracts to analyze records in the distributed ledger to identify potential abnormal behavior. This process involves monitoring data access frequency and real-time checking of storage status, for example, whether the number of accesses to a specific data entry within a certain period of time exceeds a preset threshold, or whether the load of a data storage node is abnormal. These monitoring indicators can be dynamically established based on historical data to form a comprehensive set of monitoring standards.
[0111] Smart contracts automatically run on the blockchain network and trigger corresponding monitoring behaviors through set conditions. For example, if the access frequency of a storage node is observed to increase significantly in a short period of time, the smart contract will immediately trigger the predefined anomaly detection algorithm to conduct an in-depth analysis of the node to determine whether the access is normal. If suspicious behavior is detected, the system will generate a preliminary monitoring report, detailing the nature, time, and possible impact of the anomaly.
[0112] Such a monitoring mechanism not only improves the security of the system, but also provides a real-time basis for management decisions. For example, when a storage node is accessed multiple times in a short period of time, and the access behavior is significantly different from the historical access pattern, the system can immediately feedback this information to the administrator. The administrator can quickly take response measures based on the preliminary monitoring report, such as further investigation or temporary freezing of related data, to prevent potential security risks from expanding.
[0113] For the preliminary monitoring report, a report generation method based on natural language generation technology is used to integrate the integrity verification results, distributed ledger records and monitoring reports into an archive verification report, and the final archive verification report is generated through visualization technology.
[0114] In this step, the system integrates the preliminary monitoring report, integrity verification results and distributed ledger records to generate a comprehensive archive verification report. The generation of this report relies on natural language generation (NLG) technology. The system will automatically write logical and readable report content based on the existing data information and monitoring results. In the process of generating the report, the NLG algorithm will convert complex technical information into easy-to-understand language, so that non-technical personnel can also clearly understand the storage status and security of the data. At the same time, the system uses visualization technology to present the key data and analysis results in the report in charts or other visual forms, enhancing the readability and information transmission effect of the report.
[0115] This automated report generation mechanism greatly improves the efficiency of data management and reduces the workload of manual report writing. By integrating information from multiple sources into a clear archive verification report, decision makers can quickly grasp the status, integrity and security of data storage, so as to make timely responses and adjustments. This approach not only improves the transparency of reports, but also promotes information sharing within the organization, helping various departments to better collaborate and communicate. In addition, reports with visual elements are more attractive and can intuitively display data anomalies or potential risks, providing managers with intuitive decision support.
[0116] In this step, the system combines the preliminary monitoring report, integrity verification results and distributed ledger records to create a comprehensive archive verification report using natural language generation (NLG) technology. This process involves extracting relevant information from different data sources, integrating it and converting it into easy-to-read natural language text. NLG technology can analyze data and automatically generate structured content, such as providing management with a comprehensive situation report by describing the storage status of data, the results of integrity verification and anomalies found in monitoring.
[0117] In the process of generating reports, the system also uses visualization technology to enhance the effect of information transmission. Data can be presented in the form of charts, images or dashboards, such as using a bar chart to show the change in data access frequency in the past month, or using a pie chart to show the status distribution of different data items. This visual presentation method not only makes the report content richer and easier to understand, but also helps decision makers quickly obtain key data.
[0118] The final generated archive verification report will be output in PDF or web page format and stored in a secure storage location for relevant departments to review. For example, in a hospital management system, when relevant departments need to view the storage and access records of patient data, they can quickly access the report to find the integrity verification results, storage status, and any abnormal behavior found in monitoring of patient data. This automated report generation method greatly improves work efficiency and transparency, making data management more scientific and efficient.
[0119] It can be seen that according to the data type, access frequency and business importance of the shutdown system, a multi-dimensional classification model based on deep learning is used to classify the data to obtain a classified data set and its priority label; the classified data set is input into a compression and encryption joint processing model based on a neural network to obtain a compressed and encrypted data packet; the compressed and encrypted data packet is distributed to a distributed storage system, and an index construction algorithm based on a graph database is used to generate a data storage path index to obtain a distributed storage index table; according to the distributed storage index table, the data integrity is verified using the blockchain-based archiving verification technology, and an archiving verification report is generated, thereby realizing intelligent classification, effective compression and secure storage of data in the shutdown system, and ultimately ensuring the integrity and traceability of the data.
[0120] Another embodiment of the present invention provides a data archiving processing system based on a shutdown system, see Figure 3 , the system may include: A classification module 301 is used to classify data according to the data type, access frequency and business importance of the shutdown system using a multi-dimensional classification model based on deep learning, wherein the multi-dimensional classification model dynamically divides the archiving priority of the data through attention technology and an adaptive weight allocation algorithm to obtain a classified data set and its priority label; The processing module 302 is used to input the classified data set into a compression and encryption joint processing model based on a neural network to perform data compression and encryption, wherein the joint processing model combines a quantized compression algorithm and a homomorphic encryption technology to achieve efficient compression while ensuring data security, thereby obtaining a compressed and encrypted data packet; An index module 303 is used to distribute the compressed and encrypted data packets to a distributed storage system, and generate a data storage path index using an index construction algorithm based on a graph database, wherein the index construction algorithm optimizes data retrieval efficiency and storage load balancing through dynamic hash mapping and a distributed consistency protocol to obtain a distributed storage index table; The archiving module 304 is used to verify the data integrity according to the distributed storage index table using the blockchain-based archiving verification technology, wherein the archiving verification technology combines smart contracts and distributed ledger technology to monitor the data storage status and access records in real time, generate an archiving verification report, and ensure the reliability and traceability of data archiving.
[0121] It can be seen that according to the data type, access frequency and business importance of the shutdown system, a multi-dimensional classification model based on deep learning is used to classify the data to obtain a classified data set and its priority label; the classified data set is input into a compression and encryption joint processing model based on a neural network to obtain a compressed and encrypted data packet; the compressed and encrypted data packet is distributed to a distributed storage system, and an index construction algorithm based on a graph database is used to generate a data storage path index to obtain a distributed storage index table; according to the distributed storage index table, the data integrity is verified using the blockchain-based archiving verification technology, and an archiving verification report is generated, thereby realizing intelligent classification, effective compression and secure storage of data in the shutdown system, and ultimately ensuring the integrity and traceability of the data.
[0122] An embodiment of the present invention further provides a storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above method embodiments when running.
[0123] Specifically, in this embodiment, the above storage medium may be configured to store a computer program for performing the following steps: S201, classifying the data using a multi-dimensional classification model based on deep learning according to the data type, access frequency and business importance of the shutdown system, wherein the multi-dimensional classification model dynamically divides the archiving priority of the data through attention technology and an adaptive weight allocation algorithm to obtain a classified data set and its priority label; S202, inputting the classified data set into a compression and encryption joint processing model based on a neural network to perform data compression and encryption, wherein the joint processing model combines a quantized compression algorithm and homomorphic encryption technology to achieve efficient compression while ensuring data security, thereby obtaining a compressed and encrypted data packet; S203, distributing the compressed and encrypted data packets to a distributed storage system, and using an index construction algorithm based on a graph database to generate a data storage path index, wherein the index construction algorithm optimizes data retrieval efficiency and storage load balancing through dynamic hash mapping and a distributed consistency protocol to obtain a distributed storage index table; S204, according to the distributed storage index table, the data integrity is verified by using the blockchain-based archiving verification technology, wherein the archiving verification technology combines smart contracts and distributed ledger technology to monitor the data storage status and access records in real time, generate an archiving verification report, and ensure the reliability and traceability of data archiving.
[0124] It can be seen that according to the data type, access frequency and business importance of the shutdown system, a multi-dimensional classification model based on deep learning is used to classify the data to obtain a classified data set and its priority label; the classified data set is input into a compression and encryption joint processing model based on a neural network to obtain a compressed and encrypted data packet; the compressed and encrypted data packet is distributed to a distributed storage system, and an index construction algorithm based on a graph database is used to generate a data storage path index to obtain a distributed storage index table; according to the distributed storage index table, the data integrity is verified using the blockchain-based archiving verification technology, and an archiving verification report is generated, thereby realizing intelligent classification, effective compression and secure storage of data in the shutdown system, and ultimately ensuring the integrity and traceability of the data.
[0125] An embodiment of the present invention further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0126] Specifically, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0127] Specifically, in this embodiment, the processor may be configured to perform the following steps through a computer program: S201, classifying the data using a multi-dimensional classification model based on deep learning according to the data type, access frequency and business importance of the shutdown system, wherein the multi-dimensional classification model dynamically divides the archiving priority of the data through attention technology and an adaptive weight allocation algorithm to obtain a classified data set and its priority label; S202, inputting the classified data set into a compression and encryption joint processing model based on a neural network to perform data compression and encryption, wherein the joint processing model combines a quantized compression algorithm and homomorphic encryption technology to achieve efficient compression while ensuring data security, thereby obtaining a compressed and encrypted data packet; S203, distributing the compressed and encrypted data packets to a distributed storage system, and using an index construction algorithm based on a graph database to generate a data storage path index, wherein the index construction algorithm optimizes data retrieval efficiency and storage load balancing through dynamic hash mapping and a distributed consistency protocol to obtain a distributed storage index table; S204, according to the distributed storage index table, the data integrity is verified by using the blockchain-based archiving verification technology, wherein the archiving verification technology combines smart contracts and distributed ledger technology to monitor the data storage status and access records in real time, generate an archiving verification report, and ensure the reliability and traceability of data archiving.
[0128] It can be seen that according to the data type, access frequency and business importance of the shutdown system, a multi-dimensional classification model based on deep learning is used to classify the data to obtain a classified data set and its priority label; the classified data set is input into a compression and encryption joint processing model based on a neural network to obtain a compressed and encrypted data packet; the compressed and encrypted data packet is distributed to a distributed storage system, and an index construction algorithm based on a graph database is used to generate a data storage path index to obtain a distributed storage index table; according to the distributed storage index table, the data integrity is verified using the blockchain-based archiving verification technology, and an archiving verification report is generated, thereby realizing intelligent classification, effective compression and secure storage of data in the shutdown system, and ultimately ensuring the integrity and traceability of the data.
[0129] The above describes in detail the structure, features and effects of the present invention based on the embodiments shown in the drawings. The above is only a preferred embodiment of the present invention, but the present invention is not limited to the scope of implementation shown in the drawings. Any changes made according to the concept of the present invention, or modifications to equivalent embodiments with equivalent changes, which still do not exceed the spirit covered by the description and drawings, should be within the protection scope of the present invention.
Claims
1. A data archiving processing method based on a shutdown system, characterized in that: The method comprises: According to the data type, access frequency and business importance of the shut-down system, a multi-dimensional classification model based on deep learning is used to classify the data. The multi-dimensional classification model dynamically divides the archiving priority of the data through attention technology and adaptive weight allocation algorithm to obtain the classified data set and its priority label; Inputting the classified data set into a compression and encryption joint processing model based on a neural network to perform data compression and encryption, wherein the joint processing model combines a quantized compression algorithm and homomorphic encryption technology to achieve efficient compression while ensuring data security, thereby obtaining a compressed and encrypted data packet; Distribute the compressed and encrypted data packets to a distributed storage system, and generate a data storage path index using an index construction algorithm based on a graph database, wherein the index construction algorithm optimizes data retrieval efficiency and storage load balancing through dynamic hash mapping and a distributed consistency protocol to obtain a distributed storage index table; According to the distributed storage index table, the data integrity is verified using the blockchain-based archiving verification technology, wherein the archiving verification technology combines smart contracts and distributed ledger technology to monitor the data storage status and access records in real time, generate an archiving verification report, and ensure the reliability and traceability of data archiving.
2. The method according to claim 1, characterized in that According to the data type, access frequency and business importance of the shutdown system, a multi-dimensional classification model based on deep learning is used to classify the data, wherein the multi-dimensional classification model dynamically divides the archiving priority of the data through attention technology and adaptive weight allocation algorithm, and obtains the classified data set and its priority label, including: According to the data type in the shutdown system, a data collection framework based on edge computing is used to obtain multi-source data in real time. The data is filtered for noise and filled with missing values through an adaptive data cleaning algorithm to generate a preliminary standardized data set. For the preliminary standardized data set, a multi-dimensional classification model based on deep learning is used to extract the multi-dimensional features of the data by combining data type, access frequency and business importance. The multi-head attention technology is used to capture the correlation between features of different dimensions and generate preliminary feature representations. For the preliminary feature representation, a priority division method based on an adaptive weight allocation algorithm is used to dynamically allocate the archiving priority of the data in combination with the access frequency and business importance of the data. The accuracy and rationality of the weight allocation are optimized through attention technology to generate preliminary priority labels. For the preliminary priority labels, a data classification method based on clustering algorithm is adopted. The data is divided into different categories according to the data type and priority label. The accuracy and consistency of classification are ensured through dynamic threshold adjustment technology to generate the classified data set and its priority label.
3. The method according to claim 2, characterized in that The classified data set is input into a compression and encryption joint processing model based on a neural network to perform data compression and encryption, wherein the joint processing model combines a quantized compression algorithm and a homomorphic encryption technology to achieve efficient compression while ensuring data security, and obtains a compressed and encrypted data packet, including: For the classified data set, a quantization compression algorithm based on a neural network is used to convert high-precision data into low-precision representation. The quantization accuracy is dynamically adjusted through adaptive quantization threshold adjustment technology to generate preliminary compressed data while minimizing data information loss. For the preliminary compressed data, an encryption method based on homomorphic encryption technology is used. Combined with the data priority label and security requirements, the data is encrypted and processed. Through lightweight key management technology, the efficiency and security of the encryption process are ensured to generate preliminary encrypted data. For the preliminary encrypted data, a joint optimization method based on neural network is used, combined with quantitative compression algorithm and homomorphic encryption technology, to dynamically adjust compression and encryption parameters, and through a multi-objective optimization algorithm, balance compression efficiency and encryption security to generate preliminary compressed and encrypted data packets; For the preliminary compressed and encrypted data packets, a data integrity verification method based on hash verification is used to ensure that the data is not damaged during the compression and encryption process. Through feedback correction technology, the compression and encryption parameters are dynamically adjusted to generate the final compressed and encrypted data packet.
4. The method according to claim 3, characterized in that The compressed and encrypted data packets are distributed to a distributed storage system, and an index construction algorithm based on a graph database is used to generate a data storage path index, wherein the index construction algorithm optimizes data retrieval efficiency and storage load balancing through dynamic hash mapping and a distributed consistency protocol to obtain a distributed storage index table, including: For compressed and encrypted data packets, a distribution method based on a distributed storage system is adopted. The data storage path is planned in combination with the data priority label and the load status of the storage node. The dynamic hash mapping algorithm is used to ensure the balance of data distribution and generate a preliminary storage path plan. For the preliminary storage path planning, we use the index construction method based on the graph database to abstract the data storage path into nodes and edges in the graph structure. Through the distributed consistency protocol, we ensure the consistency and reliability of the index and generate the preliminary index structure. For the preliminary index structure, an index optimization method based on dynamic load balancing is adopted. The index structure is dynamically adjusted in combination with the real-time load status of the storage node and the data access frequency. The adaptive hash mapping technology is used to optimize data retrieval efficiency and storage load balancing, and generate an optimized index structure. For the optimized index structure, an index table generation method based on visualization technology is adopted to map the index structure into a distributed storage index table. Through real-time monitoring technology, the accuracy and consistency of the index table are ensured to generate the final distributed storage index table.
5. The method according to claim 4, characterized in that According to the distributed storage index table, the data integrity is verified by using the blockchain-based archiving verification technology, wherein the archiving verification technology combines smart contracts and distributed ledger technology to monitor the data storage status and access records in real time, generate an archiving verification report, and ensure the reliability and traceability of data archiving, including: According to the distributed storage index table, the blockchain-based archiving verification technology is used in combination with smart contracts to verify the data storage status. The hash verification algorithm is used to ensure that the data is not damaged or tampered with during the storage process, and a preliminary integrity verification result is generated; For the preliminary integrity verification results, a recording method based on distributed ledger technology is used to write the data storage status and access records into the blockchain. Through consensus technology, the immutability and traceability of the records are ensured, and preliminary distributed ledger records are generated; For distributed ledger records, a real-time monitoring method based on smart contracts is used to detect abnormal behavior in combination with data access frequency and storage status, and a preliminary monitoring report is generated through anomaly detection algorithms; For the preliminary monitoring report, a report generation method based on natural language generation technology is used to integrate the integrity verification results, distributed ledger records and monitoring reports into an archive verification report, and the final archive verification report is generated through visualization technology.
6. A data archiving processing system based on a shutdown system, characterized in that: The system comprises: A classification module is used to classify data according to the data type, access frequency and business importance of the shutdown system using a multi-dimensional classification model based on deep learning, wherein the multi-dimensional classification model dynamically divides the archiving priority of the data through attention technology and an adaptive weight allocation algorithm to obtain a classified data set and its priority label; A processing module, used for inputting the classified data set into a compression and encryption joint processing model based on a neural network to perform data compression and encryption, wherein the joint processing model combines a quantized compression algorithm and homomorphic encryption technology to achieve efficient compression while ensuring data security, thereby obtaining a compressed and encrypted data packet; An index module is used to distribute the compressed and encrypted data packets to a distributed storage system, and generate a data storage path index using an index construction algorithm based on a graph database, wherein the index construction algorithm optimizes data retrieval efficiency and storage load balancing through dynamic hash mapping and a distributed consistency protocol to obtain a distributed storage index table; The archiving module is used to verify the integrity of the data according to the distributed storage index table using the archiving verification technology based on the blockchain, wherein the archiving verification technology combines smart contracts and distributed ledger technology to monitor the data storage status and access records in real time, generate an archiving verification report, and ensure the reliability and traceability of data archiving.
7. The system according to claim 6, characterized in that The classification module is specifically used for: According to the data type in the shutdown system, a data collection framework based on edge computing is used to obtain multi-source data in real time. The data is filtered for noise and filled with missing values through an adaptive data cleaning algorithm to generate a preliminary standardized data set. For the preliminary standardized data set, a multi-dimensional classification model based on deep learning is used to extract the multi-dimensional features of the data by combining data type, access frequency and business importance. The multi-head attention technology is used to capture the correlation between features of different dimensions and generate preliminary feature representations. For the preliminary feature representation, a priority division method based on an adaptive weight allocation algorithm is used to dynamically allocate the archiving priority of the data in combination with the access frequency and business importance of the data. The accuracy and rationality of the weight allocation are optimized through attention technology to generate preliminary priority labels. For the preliminary priority labels, a data classification method based on clustering algorithm is adopted. The data is divided into different categories according to the data type and priority label. The accuracy and consistency of classification are ensured through dynamic threshold adjustment technology to generate the classified data set and its priority label.
8. The system according to claim 7, characterized in that The processing module is specifically used for: For the classified data set, a quantization compression algorithm based on a neural network is used to convert high-precision data into low-precision representation. The quantization accuracy is dynamically adjusted through adaptive quantization threshold adjustment technology to generate preliminary compressed data while minimizing data information loss. For the preliminary compressed data, an encryption method based on homomorphic encryption technology is used. Combined with the data priority label and security requirements, the data is encrypted and processed. Through lightweight key management technology, the efficiency and security of the encryption process are ensured to generate preliminary encrypted data. For the preliminary encrypted data, a joint optimization method based on neural network is used, combined with quantitative compression algorithm and homomorphic encryption technology, to dynamically adjust compression and encryption parameters, and through a multi-objective optimization algorithm, balance compression efficiency and encryption security to generate preliminary compressed and encrypted data packets; For the preliminary compressed and encrypted data packets, a data integrity verification method based on hash verification is used to ensure that the data is not damaged during the compression and encryption process. Through feedback correction technology, the compression and encryption parameters are dynamically adjusted to generate the final compressed and encrypted data packet.
9. A storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program is configured to execute the method according to any one of claims 1 to 5 when executed.
10. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Method and system for safely and rapidly storing image data of imaging department
CN119517325A
Tumor early screening data sharing platform construction method and system based on cloud computing
CN119920488A
System and method for dynamic adaptive user-based prioritization and display of electronic messages
US20060010217A1
Cited By
Storage management method and system based on data encryption
CN120724465A
A data encryption-based storage management method and system
CN120724465B
Financial transaction historical data storage optimization method
CN121365043A