Efficient data governance method for encrypted database based on key management system
By using multi-dimensional feature vectors to generate dynamic encryption keys and access control vectors in a distributed storage system, the security risks and low resource utilization efficiency brought by static keys are solved, and efficient and secure dynamic management of data is achieved.
Patent Information
- Application Number
- CN202510313216.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-03-17
AI Technical Summary
In existing distributed storage systems, the use of static keys leads to high security risks, and the encryption strategy is single and cannot adapt to the multi-dimensional attribute changes of data in complex business scenarios, resulting in inefficient resource utilization.
A key management system-based approach is adopted to dynamically generate encryption keys by generating multi-dimensional feature vectors. Combined with access control vectors and load balancing factors, encryption, storage, and access strategies are dynamically adjusted to ensure data integrity and consistency and optimize system performance.
It ensures data security and integrity in a high-concurrency environment, improves resource utilization efficiency, reduces the risk of performance bottlenecks, and adapts to complex business needs.
Smart Images

Figure CN119884252B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of database management and processing, and in particular to an efficient data governance method for an encrypted database based on a key management system. Background Art
[0002] Data encryption technology is currently widely used in distributed storage systems to protect the security of sensitive data. Traditional encryption methods primarily include symmetric algorithms (such as AES and DES) and asymmetric algorithms (such as RSA and ECC). These methods provide fundamental security for data transmission and storage. However, in practical applications, these encryption methods often face the following limitations: Most existing data encryption systems use static keys for encryption. This means that a fixed key is generated during system initialization and used for all data encryption and decryption operations. While simple to implement, this approach presents significant security risks. Once a key is compromised, an attacker can decrypt all encrypted data associated with that key, resulting in serious data breaches. Furthermore, static keys lack a dynamic update mechanism, making them incapable of adapting to the dynamic data changes in complex business scenarios. Many existing technologies design encryption strategies based solely on a single data characteristic (such as sensitivity or size), failing to fully consider the multidimensional data attributes (such as access frequency, real-time requirements, and encryption level). This single-dimensional encryption strategy struggles to meet complex business needs, especially in distributed storage environments, where different types of data have significantly different security and performance requirements. Single-dimensional strategies often lead to inefficient resource utilization. Summary of the Invention
[0003] The purpose of this invention is to provide an efficient data governance method for encrypted databases based on a key management system, comprehensively improving data security, resource utilization efficiency, and access performance in distributed storage systems. This method overcomes the shortcomings of traditional methods, which often involve static keys, a single security policy, and insufficient resource allocation. It dynamically adjusts encryption, storage, and access policies based on data characteristics and system load, ensuring data integrity and consistency in highly concurrent environments while optimizing system performance.
[0004] To solve the above technical problems, the present invention provides an efficient data management method for an encrypted database based on a key management system, the method comprising:
[0005] Step 1: The database is stored in a distributed storage system. Based on the business attributes of the data in the database, a multi-dimensional feature vector is created for each type of data, and a unique encryption key is generated from this vector. Each type of data is encrypted using the encryption key to generate the corresponding encrypted data block. Based on the operating parameters of each storage node, the comprehensive evaluation value of each storage node in the distributed storage system is calculated to determine the resource capacity of each storage node, and the preliminary average optimal number of shards for all encrypted data blocks in the distributed storage system is determined.
[0006] Step 2: An access control vector is generated by combining the operating parameters of each storage node and the permission level allowed for the user's role in the distributed storage system. This vector is then constrained by the time dimension to prevent timeout or unauthorized access, thereby dynamically strengthening the access control vector. The access control vector is dynamically refreshed when the time or key is updated.
[0007] Step 3: For each encrypted data block, calculate the hash value, cyclic redundancy check value, error correction code, and message authentication code, and normalize them accordingly in combination with the access control vector to complete data integrity verification and consistency check; calculate a load balancing factor based on the real-time total data access load and the total number of concurrent requests, and combine it with the preliminary average optimal number of shards to guide the average optimal number of shards for all encrypted data blocks in the distributed storage system to evolve towards the optimal state.
[0008] Furthermore, the operating parameters of the storage node include: bandwidth capacity, CPU usage, memory usage, disk input and output operations per second, average response time and concurrency.
[0009] Furthermore, the business attributes include: sensitivity score, data access frequency, data type identifier, response time requirement, data size, number of concurrent access users, encryption level and timestamp; the sensitivity score ranges from 1 to 10, indicating the security sensitivity level of this type of data, and the higher the sensitivity score, the greater the contribution to the key; the data access frequency indicates the access frequency of this type of data; the data type identifier takes a value of 1, 2 or 3, when the value is 1, it indicates structured data, when the value is 2, it indicates semi-structured data, and when the value is 3, it indicates unstructured data; the response time requirement indicates the response time constraint on this type of data, and the smaller the value, the higher the real-time requirement; the encryption level ranges from 1, 2 or 3, and the higher the value, the higher the corresponding encryption level.
[0010] Furthermore, let The multidimensional feature vector of the class data is:
[0011] ;
[0012] in, For the Sensitivity score of class data; For the The frequency with which the data of the class data is accessed; For the Data type identifier of class data; For the Response time requirements for class data; For the Data size of class data; For the The number of concurrent users accessing class data; For the The encryption level of the class data; For the The timestamp of the class data; a unique encryption key is generated using the following formula :
[0013] ;
[0014] in, is the number of categories of data; Represents the L1 norm of the multidimensional feature vector; is the set encrypted large prime number; is the modulo operation; express The hash value of .
[0015] Furthermore, each type of data is encrypted using the CBC mode of the AES encryption algorithm using an encryption key.
[0016] Furthermore, the initial average optimal number of fragments for all encrypted data blocks is Calculated by the following formula:
[0017] ;
[0018] in, For the Memory usage of storage nodes; For the The number of disk input and output operations per second for each storage node; is the determinant operator; is the number of storage nodes; For the The bandwidth capacity of each storage node; For the The CPU usage of each storage node.
[0019] Furthermore, access control vector Calculated by the following formula:
[0020] ;
[0021] in, The user's permission level is an integer from 1 to 5. The CPU usage allowed for the access corresponding to the user, The number of concurrent accesses allowed for this user. The number of disk input and output operations per second allowed for the access corresponding to this user; The access bandwidth allowed for the user; The average response time allowed for the access corresponding to the user; The memory usage allowed for access by the user; No. A hash value of the timestamp of the class data; is the minimum function; For the The number of concurrent storage nodes; For the The average response time of each storage node.
[0022] Furthermore, the verification value is calculated using the following formula to complete data integrity verification and consistency check:
[0023] ;
[0024] in, No. The verification value of the encrypted data block corresponding to the class data. When the verification value is within the set threshold range, it means that the encrypted data block corresponding to the class data has passed the verification and is allowed to be accessed; otherwise, access is denied; for The hash value of the encrypted data block corresponding to the class data; for The error correction code of the encrypted data block corresponding to the class data; for The message authentication code of the encrypted data block corresponding to the class data; for The cyclic redundancy check code of the encrypted data block corresponding to the class data.
[0025] Furthermore, the following formula is used to guide the average optimal number of shards for all encrypted data blocks in the distributed storage system toward the optimal state:
[0026] ;
[0027] in, For time The load balancing factor is calculated as follows:
[0028] ;
[0029] in, For time The total number of concurrent requests at time; For time Total data access load at time ; For time The total number of concurrent requests at the time.
[0030] The present invention provides an efficient data management method for encrypted databases based on a key management system, which has the following beneficial effects:
[0031] The present invention employs a dynamic key generation method based on the multidimensional characteristics of data in key management, breaking through the security limitations of traditional static keys. The encryption key for each type of data is determined by multiple characteristics such as sensitivity score, access frequency, data type, response time requirement, encryption level, and timestamp. The introduction of timestamp hashing and dynamic adjustment parameters into the formula ensures that the key has a high degree of timeliness and uniqueness. This design ensures that even if the same type of data is encrypted at different times or under different circumstances, its key will be completely different, effectively resisting replay attacks and brute force cracking. Through dynamic access control vectors, the present invention implements real-time access permission management based on user permission level, storage node resource status, and data characteristics.
[0032] Compared with traditional static access control lists, the present invention can dynamically adjust permission configurations based on the user's role level, system load status, and concurrent requests. For example, in high-concurrency scenarios, the system will prioritize resource allocation for high-privilege users and high-priority tasks, while limiting the access frequency of low-privilege users, thereby achieving efficient resource utilization and priority management. The dynamic access control mechanism also combines timestamp hashing and key strength to constrain permissions through the time dimension, preventing the abuse of permissions that have been idle for a long time. At the same time, the combination of access control vectors with data characteristics and node status enables the system to automatically adapt to load changes in permission management, avoiding resource conflicts and security vulnerabilities.
[0033] By combining dynamic parameters such as storage node bandwidth, CPU utilization, memory usage, disk IOPS, and average response time, this invention achieves dynamic sharding optimization for encrypted data blocks in distributed storage systems. This optimization method balances node load in high-concurrency access scenarios, preventing overload on some nodes while ensuring overall stable performance. Compared to traditional static sharding strategies, this dynamic optimization mechanism significantly improves storage system resource utilization and reduces the risk of performance bottlenecks caused by hotspot data. Furthermore, the load balancing factor accurately models system load characteristics through a combination of Poisson and gamma distributions, effectively guiding the evolution of the sharding strategy towards its optimal state. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0035] Figure 1 A flowchart of an efficient data governance method for an encrypted database based on a key management system provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0036] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0037] Example 1, reference Figure 1 : An efficient data management method for an encrypted database based on a key management system, the method comprising:
[0038] Step 1: The database is stored in a distributed storage system. Based on the business attributes of the data in the database, a multi-dimensional feature vector is created for each type of data, and a unique encryption key is generated from this vector. Each type of data is encrypted using the encryption key to generate the corresponding encrypted data block. Based on the operating parameters of each storage node, the comprehensive evaluation value of each storage node in the distributed storage system is calculated to determine the resource capacity of each storage node, and the preliminary average optimal number of shards for all encrypted data blocks in the distributed storage system is determined.
[0039] First, the system categorizes all data in the database and extracts multiple dimensions that reflect the data's characteristics based on business attributes such as importance, sensitivity, structuredness, access frequency, and concurrent access. These dimensions are quantified into numerical values and integrated into a multidimensional feature vector. The construction of this feature vector is not simply a collection of parameters; it incorporates a variety of dynamic analysis methods. For example, data sensitivity is quantified by assessing the level of risk the data poses to the business (e.g., the confidentiality of financial data, the privacy of medical data). This sensitivity score directly influences the subsequent selection of key strength and encryption strategies. Quantifying access frequency requires reference to historical access logs and predictive models. This involves not only calculating the average number of accesses over a period of time but also incorporating dynamic weighting within time windows. For example, predicting peak access times in the future will give higher weight. Data type identification distinguishes between structured data (e.g., relational database tables), semi-structured data (e.g., JSON files), and unstructured data (e.g., video files), as these different types of data have distinct storage requirements and encryption complexity. Once the feature vectors are constructed, the key management system generates a unique encryption key based on these vectors. During key generation, the timestamp is introduced as a key dynamic variable, closely linking the key to the point in time when the data is stored, thereby improving the key's timeliness and security. This temporal correlation not only prevents key reuse but also provides additional security constraints during subsequent dynamic updates or access control of data. The key generation formula comprehensively considers data characteristics (such as sensitivity, size, and concurrency) and system status (such as storage node resources and encryption level). It uses the hash function in the encryption algorithm to process the timestamp and eigenvalue, and then generates the final key through modular operations. In this way, even for the same type of data, the keys will be significantly different due to differences in timestamps and dynamic changes in encryption levels, greatly improving resistance to attacks.
[0040] After key generation, data is partitioned to accommodate the distributed storage system's storage mechanism. The partitioning rules are adjusted based on data size and node storage performance. For example, large data files are split into smaller chunks to distribute across multiple storage nodes. This not only improves storage efficiency but also reduces single-node load and the risk of data loss. Each data chunk is encrypted using a distributed key using a symmetric encryption algorithm (such as AES) and an initialization vector (IV) to ensure randomness and independence of the encryption results. After encryption, the resulting encrypted data chunk is accompanied by metadata such as the chunk ID, encryption key ID, initialization vector, and encryption algorithm type. This information facilitates subsequent data recovery and decryption. Simultaneously, the system evaluates the resource capabilities of each distributed storage node based on its real-time operating parameters (such as bandwidth, CPU utilization, memory utilization, and disk IOPS). This evaluation utilizes a comprehensive performance scoring approach that considers the weighted distribution of resources across various dimensions and dynamically generates a score for each node. This score determines the initial allocation strategy for encrypted data chunks, prioritizing them to nodes with high performance and high availability. The calculation of the preliminary average optimal number of shards is an overall optimization process based on the global node resource capabilities. By evenly distributing data blocks, the system can effectively reduce the resource bottleneck problem caused by the centralized storage of hot data sets, while ensuring that highly sensitive data blocks are preferentially allocated to nodes with higher security levels.
[0041] Step 2: An access control vector is generated by combining the operating parameters of each storage node and the permission level allowed for the user's role in the distributed storage system. This vector is then constrained by the time dimension to prevent timeout or unauthorized access, thereby dynamically strengthening the access control vector. The access control vector is dynamically refreshed when the time or key is updated.
[0042] To ensure access security and compliance for encrypted data, the system constructs a multidimensional access control vector based on each user's role and permission level. This vector describes the user's access capabilities to specific data blocks. The user's role level (e.g., administrator, user, guest) and permission level (e.g., read, write, delete) together determine the user's basic permissions. This basic permission definition utilizes a hierarchical model, preset based on business needs during system initialization. However, the specific permissions assigned are dynamically adjusted over time, data status, and user behavior. In addition to roles and permission levels, the access control vector incorporates other dynamic parameters, such as the user's historical access history, the security assessment of the current session, and priority access requirements for specific sensitive data. These parameters collectively contribute to the dynamic nature of access control, enabling the system to automatically adjust permission configurations in complex access scenarios and avoid unnecessary security vulnerabilities. To further enhance access control security, this invention introduces a time dimension constraint on top of the traditional access control model. The inclusion of the time dimension serves two primary purposes: preventing the long-term abuse of permissions and enabling real-time updates of the access control vector to adapt to changes in data and the environment. In practice, the system generates a unique timestamp for each access request and binds it to the user's permission level, creating a time-sensitive access token. This token is time-sensitive and valid only within a specific time window, preventing malicious exploitation of long-unused permissions. Furthermore, the timestamp serves as a trigger for dynamic updates to the access control vector. When the timestamp expires or the key is updated, the system reevaluates the user's access rights and refreshes the access control vector. This real-time refresh mechanism ensures that access control can accurately adapt to new security requirements even when data status or keys change. Furthermore, step 2 introduces a dynamic access behavior enforcement mechanism to further enhance the flexibility and security of access control. The key to dynamic enforcement is real-time analysis and adjustment of user access requests based on the operating parameters of the storage node and the user's permission level. For example, when a storage node is highly loaded, the system dynamically offloads access requests based on the node's operating parameters (such as bandwidth, CPU usage, and memory usage), prioritizing resources for requests from high-privilege users or high-priority data. At the same time, for users with lower permissions, the system may reduce their access frequency or extend the queue time of their requests to ensure the stability of the overall system performance. This dynamic reinforcement mechanism based on operating parameters can not only effectively avoid resource bottlenecks caused by high concurrent access, but also ensure the access security of highly sensitive data.
[0043] It's worth noting that in step 2, the dynamic refresh of the access control vector not only relies on the timestamp but is also directly related to the key update frequency. When the encryption key is updated, the system automatically triggers the reconstruction of the access control vector. At this point, the new encryption key is combined with user permissions and access behavior to regenerate an access control policy that matches the current environment. For example, when the key level of a certain type of data is upgraded from medium to high, the system automatically adjusts the user permissions associated with that data in the access control vector. This may temporarily restrict access for some low-privileged users until they pass higher-level authentication or the system load decreases. This key-driven dynamic refresh mechanism not only enhances data security but also enables access control to quickly respond to changes in keys and the environment, thereby maintaining system flexibility and stability. By constructing dynamic access control vectors and introducing time-dependent constraints, real-time and refined permission allocation is achieved, in stark contrast to existing static access control mechanisms. In traditional distributed storage systems, access permissions are typically set during system initialization and are relatively fixed, lacking the ability to dynamically adjust. By incorporating multiple factors—such as user permissions, data characteristics, storage node status, and time—into the access control model, the present invention enables the system to dynamically optimize access policies based on actual conditions. For example, when access requests for highly sensitive data surge, the system can instantly increase the security level of relevant access controls, prioritizing the normal operation of critical business operations without requiring downtime or manual adjustments to permissions. This dynamic optimization capability significantly enhances the system's adaptability to complex business scenarios.
[0044] Step 3: For each encrypted data block, calculate the hash value, cyclic redundancy check value, error correction code, and message authentication code, and normalize them accordingly in combination with the access control vector to complete data integrity verification and consistency check; calculate a load balancing factor based on the real-time total data access load and the total number of concurrent requests, and combine it with the preliminary average optimal number of shards to guide the average optimal number of shards for all encrypted data blocks in the distributed storage system to evolve towards the optimal state.
[0045] To ensure the integrity of encrypted data blocks during storage and transmission, the system generates multiple checksums for each data block, including hash values, cyclic redundancy check values (CRC), error correction codes (ECC), and message authentication codes (MAC). These checksums not only verify the encrypted content of the data block but also serve as an important foundation for subsequent consistency checks. The hash value quickly determines whether a data block has been tampered with. By comparing the hash value generated during storage with the recalculated hash value during access, abnormal data blocks can be quickly located. The cyclic redundancy check provides efficient verification by detecting low-level errors in the data block, while the error correction code provides automatic repair capabilities for even minor data corruption. The message authentication code further ensures that data is not maliciously tampered with during transmission between multiple nodes. These multi-layered verification mechanisms complement each other to form a secure and stable integrity verification framework. After completing data integrity verification, the system normalizes the data based on the access control vector to facilitate further consistency checks. The core of normalization is to unify various checksum results, access rights levels, and data block priority information into a standardized metric. This not only simplifies the processing of multidimensional information but also provides a common quantitative foundation for subsequent optimization calculations. Access control vectors play a crucial role in this process, not only determining whether a user has permission to access certain encrypted data blocks but also providing a reference for access behavior during consistency checks. For example, if the system detects an abnormal frequency of access requests or an unexpected permission level for a user, it can dynamically adjust the weight parameters in the normalization process, further strengthening the constraints and optimization of data access behavior.
[0046] After data integrity verification and consistency checks are complete, the system calculates a global load balancing factor based on the real-time total data access load and the total number of concurrent requests. The calculation of this factor is crucial to step 3, as it determines how the distributed storage system dynamically adjusts its data block sharding strategy to adapt to the current load. The core logic of the load balancing factor combines storage node resource usage (such as bandwidth, CPU utilization, memory utilization, and IOPS) with data block characteristics (such as access frequency, sensitivity, and data size) to calculate a comprehensive optimized value using a mathematical model. This optimized value not only reflects the current system load but also guides the system toward the optimal load through a dynamic weighting mechanism. For example, when certain storage nodes are overloaded, the load balancing factor dynamically offloads data blocks from these nodes, migrating some high-priority blocks to less-loaded nodes to achieve balanced load distribution. It is worth noting that the initial average optimal number of shards plays a crucial role in calculating the load balancing factor. The preliminary average optimal number of shards is the initial sharding strategy calculated in step 1 based on the resource capabilities of distributed storage nodes. It provides a basic reference value for the optimization calculation in step 3. In actual applications, the system will combine the preliminary average optimal number of shards with the real-time load status to adjust the number and distribution of shards for the current data block. Through this dynamic iterative approach, the system can gradually approach the global optimal sharding state, maximizing storage efficiency and access performance. Furthermore, the sharding optimization in step 3 fully considers data security requirements. In traditional distributed storage systems, data sharding is often limited to performance optimization, while security is neglected. However, by incorporating access control vectors and data integrity verification results into the sharding optimization calculation process, the present invention ensures that highly sensitive data blocks are preferentially allocated to high-security nodes for storage. For example, when a certain type of data block has a high access frequency and sensitivity score, the system automatically allocates it to nodes with higher bandwidth, lower load, and higher security for storage, while lower-priority data blocks are allocated to nodes with relatively less resources, thus achieving dual optimization of performance and security. Unlike traditional static sharding strategies, this invention introduces load balancing factors and dynamic sharding optimization mechanisms, enabling the system to automatically adjust sharding strategies based on real-time changes in load conditions and access requirements. This dynamic adjustment not only improves the system's resource utilization efficiency but also greatly enhances the distributed storage system's adaptability in high-concurrency and complex business scenarios. For example, in certain peak access scenarios, the system can allocate more resources to accessing critical data blocks by optimizing the load balancing factors in real time, thereby ensuring business continuity and stability without the need for human intervention or system restarts.
[0047] Example 2: The operating parameters of the storage node include: bandwidth capacity, CPU usage, memory usage, disk input and output operations per second, average response time and concurrency.
[0048] Specifically, in practice, the bandwidth capacity of a storage node determines its ability to handle large-scale data transfers. Especially in distributed storage environments, bandwidth limitations can directly become a bottleneck for system performance. Encrypted databases require frequent transmission of encrypted data blocks between nodes, for example, for multi-replica synchronization, large-scale data queries, or backup and recovery operations, all of which rely on sufficient bandwidth. Therefore, the system monitors each node's bandwidth utilization in real time and dynamically adjusts data allocation based on the node's available bandwidth. For example, when a node's bandwidth is nearing saturation, the system prioritizes nodes with greater bandwidth margins to store frequently accessed data blocks, thereby avoiding data access delays caused by network congestion. This dynamic bandwidth allocation mechanism not only improves system throughput but also maintains stable performance in high-concurrency scenarios. In addition to bandwidth, CPU utilization is a key indicator of storage node performance, reflecting the pressure on the node when processing compute-intensive tasks such as encryption, decryption, and data verification. In encrypted databases, data block encryption and integrity verification are common high-load computing tasks. Excessive CPU utilization on a node can lead to decreased data processing efficiency and even task timeouts. Therefore, the system dynamically evaluates CPU usage and appropriately allocates computing tasks to balance the load. For example, intensive encryption operations are prioritized for nodes with lower CPU loads, while high-load nodes are primarily used for lightweight tasks such as data transfer or temporary storage. This CPU usage control mechanism enables the system to find the optimal balance between encryption computing tasks and storage operations, thereby improving overall resource utilization efficiency.
[0049] Memory utilization, a key parameter reflecting a node's storage buffering capacity, is closely related to the caching and read performance of encrypted data blocks. In high-access scenarios, encrypted databases frequently read and cache data blocks, making sufficient memory resources crucial to system performance. By monitoring node memory usage in real time, the system dynamically adjusts the distribution strategy for encrypted data blocks, prioritizing the storage of frequently accessed blocks on nodes with ample memory, thereby reducing disk I / O pressure and improving data access speed. Furthermore, the system optimizes access control and permission allocation based on memory usage. For example, it can restrict access requests from low-priority users on memory-constrained nodes to ensure quality of service for high-priority users. Disk Input / Output Operations Per Second (IOPS) is a core metric for measuring storage node disk performance and is particularly important in the distributed environment of encrypted databases. The storage and read of each encrypted data block relies on disk read and write capabilities. Especially when handling large-scale concurrent access, IOPS determines the system's responsiveness to high-load requests. When the IOPS of certain nodes approaches its maximum, the system dynamically adjusts the sharding strategy to reduce read and write operations on these nodes and allocate newly generated encrypted data blocks to nodes with more sufficient IOPS resources. Furthermore, for frequently accessed hot data blocks, the system prioritizes nodes with higher IOPS performance for storage, thus avoiding data access delays caused by disk performance bottlenecks. This dynamic optimization mechanism fully utilizes node storage resources and improves the robustness of the distributed storage system under high-load scenarios.
[0050] Average response time is another key parameter used by the system to evaluate node performance. It comprehensively reflects the node's efficiency in performing data storage, transmission, and computation tasks. For encrypted databases, each access request may involve data exchange between multiple nodes. Nodes with longer response times directly impact the overall request completion speed. Therefore, the system dynamically monitors the response time of each node and prioritizes the allocation of highly sensitive or real-time-critical encrypted data blocks to nodes with shorter response times. This response-time-based allocation strategy not only improves access efficiency for critical data but also gradually reduces global system latency by optimizing resource distribution. Finally, concurrency, as an indicator of storage node load, reflects a node's ability to handle multiple requests simultaneously. High concurrency is common in encrypted databases, such as multiple users simultaneously querying or modifying the same dataset. Excessive concurrency can lead to overloaded nodes, degrading service quality and even causing node unavailability. Therefore, by monitoring concurrency in real time, the system implements traffic throttling or traffic migration measures for high-concurrency nodes within its sharding strategy and load balancing mechanisms. For example, when a node's concurrency exceeds a threshold, the system will distribute some newly generated encrypted data blocks to other nodes for storage. It will also dynamically adjust access control vectors to restrict access requests from low-priority users. This dynamic concurrency control mechanism effectively avoids performance bottlenecks caused by single-node overload while ensuring overall system stability.
[0051] Example 3: The business attributes include: sensitivity score, data access frequency, data type identifier, response time requirement, data size, number of concurrent access users, encryption level and timestamp; the sensitivity score ranges from 1 to 10, indicating the security sensitivity level of this type of data, and the higher the sensitivity score, the greater the contribution to the key; the data access frequency indicates the access frequency of this type of data; the data type identifier takes a value of 1, 2 or 3, when the value is 1, it indicates structured data, when the value is 2, it indicates semi-structured data, and when the value is 3, it indicates unstructured data; the response time requirement indicates the response time constraint for this type of data, and the smaller the value, the higher the real-time requirement; the encryption level ranges from 1, 2 or 3, and the higher the value, the higher the corresponding encryption level.
[0052] Specifically, the sensitivity score is a crucial dimension of data business attributes, directly reflecting the data's security level. The score ranges from 1 to 10, with higher values indicating more sensitive data. For example, financial account information, medical records, or confidential business documents are typically assigned higher sensitivity scores. The sensitivity score acts as a weighting factor in the key generation process, with highly sensitive data contributing more to the key's strength. This design enables the system to generate keys with differentiated strengths based on the importance of the data, enabling targeted protection within limited resources and avoiding the waste of computing resources caused by over-encrypting less sensitive data. Data access frequency is a key indicator of data activity, reflecting how frequently a particular type of data is accessed over a period of time. Frequently accessed data, such as inventory information for popular products or real-time transaction logs, is typically encrypted and stored on high-performance nodes to reduce access latency and improve user experience. The system monitors access frequency in real time and incorporates this attribute into the feature vector, ensuring a more balanced distribution of high-frequency data within the distributed storage system. This also assigns higher weight to high-frequency data during key generation, enhancing security.
[0053] The data type identifier distinguishes between structured, semi-structured, and unstructured data. Specifically, structured data (value 1) is typically stored in relational databases and has clearly defined fields. Semi-structured data (value 2), such as JSON files or XML documents, has some structure but a more flexible format. Unstructured data (value 3), such as video, image, or audio files, has no fixed structure at all. The system uses this identifier to implement different encryption and storage strategies for different types of data. For example, structured data typically requires fast query speed and efficient decryption, so lightweight encryption algorithms are preferred. Unstructured data, due to its large size and relatively low access frequency, may require block storage and shard encryption. The response time requirement represents the real-time constraints on data processing; smaller values indicate higher response speed requirements. In real-world applications, transactional data or user authentication data typically require millisecond responses, so the response time requirement is set lower. Access to historical archived data, on the other hand, may tolerate higher latency and therefore require a higher response time. The system takes this attribute into consideration when allocating storage nodes, assigning data with low response time requirements to higher-performance nodes. It also weights this attribute during key generation, prioritizing the protection of high-real-time data in encryption and access control. Data size is a key indicator of data storage and transmission costs. When distributing data, the system needs to adjust the sharding strategy based on the size of the data blocks. For example, large data blocks can be split into smaller units for distributed storage and faster access. Furthermore, data size also affects key generation. Larger data blocks require stronger keys, thereby increasing encryption security and cracking resistance.
[0054] The number of concurrent access users measures the number of requests a system must process within a specific time period. This attribute is crucial for load balancing and access control. High-concurrency data, such as shared documents or public service interfaces, is typically prioritized for storage nodes that support high concurrency. The system also dynamically adjusts access control policies based on the number of concurrent users to limit access by low-priority users to ensure overall system stability. The encryption level directly describes data security requirements, with values ranging from 1, 2, or 3, corresponding to low, medium, and high levels, respectively. Data with a high encryption level is typically used in sensitive scenarios, such as finance and healthcare. The system assigns stronger keys, employs more complex encryption algorithms, and increases storage replication or redundancy to improve fault tolerance. Data with a low encryption level, on the other hand, may require only basic encryption protection to reduce the system's computational burden and resource consumption. The timestamp is a key attribute that introduces dynamicity, marking the time when data is generated or modified. During key generation, the timestamp is used as a dynamic variable, working alongside other attributes to determine the uniqueness and timeliness of the key. This design can effectively prevent the reuse of keys and ensure that the regeneration of keys can be automatically triggered when data is updated, thereby further improving the security of the system. Through the comprehensive definition and combination of the above business attributes, the present invention constructs a flexible multi-dimensional feature system that can accurately describe the business characteristics of data and play an important role in encryption key generation, data storage optimization and dynamic access control. Compared with traditional technologies, the innovation of the present invention is to integrate these attributes in the form of quantization and feature vectors, and give them a dynamic weight control mechanism in the key management system, thereby realizing differentiated management of different scenarios and different types of data. This differentiated management not only improves the utilization efficiency of system resources, but also significantly enhances the security and controllability of data, providing unprecedented technical support for efficient data governance in large-scale distributed storage systems.
[0055] Example 4: Assume The multidimensional feature vector of the class data is:
[0056] ;
[0057] in, For the Sensitivity score of class data; For the The frequency with which the data of the class data is accessed; For the Data type identifier of class data; For the Response time requirements for class data; For the Data size of class data; For the The number of concurrent users accessing class data; For the The encryption level of the class data; For the The timestamp of the class data; a unique encryption key is generated using the following formula :
[0058] ;
[0059] in, is the number of categories of data; Represents the L1 norm of the multidimensional feature vector; is the set encrypted large prime number; It is a modulo operation.
[0060] Specifically, the principle of the formula proposed in Example 4 is to comprehensively utilize the multi-dimensional characteristics of the data, and use nonlinear transformation and dynamic parameter coupling to generate an encryption key that can fully reflect the security requirements, access requirements and real-time status of each type of data. This section explains how it constructs a core sub-expression that numerically reflects both the data security level and the access load characteristics by multiplying the sensitivity score, access frequency, and data size, and combining the data type and exponential factor. This sub-expression amplifies the nonlinear impact of the sensitivity score through exponential operations, and takes the square root of the product of the response time requirement and the number of concurrent access users in the denominator, thereby limiting the security risks that may be brought about by high concurrency and low latency scenarios. In business practice, highly sensitive and frequently accessed data often require stricter encryption strength, and this formula cleverly reflects this requirement in the exponential processing, achieving accurate weighting of sensitive data. At the same time, for the case of large amounts of data, the formula introduces The cumulative sum is combined in the denominator and The square root operation can prevent the system performance bottleneck caused by excessive expansion of encryption strength. On the other hand, This structure performs a logarithmic transformation on the product of the data size and the encryption level, and combines it with a timestamp hash function to constrain the denominator. In this way, the system can dynamically adjust the impact of the logarithmic result according to the timestamp when constructing the encryption key. On the one hand, it improves the randomness of the key changes over time, and on the other hand, it brings the convenience of traceability and expiration update mechanism to the key management system. After the logarithmic operation, data with a high encryption level will obtain additional numerical gain, thereby improving the strength of the final key in combination with the hash value. At the same time, if a certain type of data has a high encryption level, but the data block is not large, or the timestamp has changed recently, then in The contribution to the overall key can also be moderately controlled in the term to avoid excessive accumulation of encryption complexity that affects actual usage efficiency. This flexible mechanism ensures that the system can adaptively find a balance between security and performance when facing data of different types and sizes. In the overall structure of this formula, the multi-dimensional features of each type of data are quantized and then summed, and the final modulus is a large prime number. , which can not only ensure that the generated key is controlled in the numerical range, but also enhance the resistance to reverse inference by attackers. Through the modulo operation, the key management system can fix the effective length of the encryption key and avoid transmitting too long key values between distributed storage nodes. Furthermore, by adding the vector norm The purpose is to introduce a measure of the overall scale of multidimensional features into the formula, so that when the eigenvalues are distributed in different orders of magnitude, their normalized differences within the same data category can be reflected through the norm. In this way, if the sensitivity score and access frequency of the same type of data are particularly high, then The weight of this type of data in key generation will also be larger, which will correspondingly increase the proportion of this type of data in key generation, thereby achieving focused protection of key data. If the overall indicators of each dimension of a certain type of data are low, even if the value in a specific dimension is more prominent, it will not excessively increase the overall key generation amount, maintaining the balance and robustness of the formula. At the implementation level, when the distributed storage system receives various types of data in the database, it will generate a feature vector for each category, and then calculate the contribution value of each category one by one through this formula and accumulate them. During this period, the timestamp hash function It will change continuously as the system clock advances. During periodic key updates or event-driven updates, different hash values will be generated for the new round of calculations, causing the final key to exhibit significant time-varying characteristics. The key management system can thus flexibly version keys, regularly abolish expired keys and activate new keys, and prevent the same key from being brute-forced or leaked due to long-term use. At the same time, the nonlinear factors and logarithmic operations in the formula ensure accurate weighting of highly sensitive, highly concurrent, and highly accessed data, enabling it to achieve higher encryption strength than other ordinary data during the key generation process. Data that requires a higher level of security protection will also be numerically fully weighted. This design not only meets the needs for refined control of large-scale data in a distributed environment, but also enables the key management process to have higher dynamic adaptability.
[0061] Example 5: Each type of data is encrypted using an encryption key using the CBC mode of the AES encryption algorithm.
[0062] Specifically, during the implementation process, each type of data first dynamically generates a unique encryption key based on its business attributes. . The generation of this key comprehensively considers multi-dimensional features such as the sensitivity score, access frequency, encryption level and timestamp of the data, so that the key can not only match the security requirements of the data in a targeted manner, but also be dynamically adjusted as time or the environment changes. When this dynamically generated key is input into the AES encryption algorithm, it provides a highly unique and unpredictable encryption parameter for each type of data, greatly improving the security of the encryption process. The core of the AES algorithm is to convert plaintext data into ciphertext through a series of complex operations such as substitution, shift and column confusion. This encryption process is extremely sensitive to the input data, and any slight change in the key will result in completely different encryption results. Therefore, the dynamic key generation mechanism in the present invention is the cornerstone of ensuring the overall security of the system. The introduction of CBC mode further enhances the randomness and anti-attack capabilities of encryption. In CBC mode, plaintext data is divided into blocks of fixed size (for example, 128 bits), and each data block is XORed with the previous ciphertext block before being input into the AES algorithm for encryption. In order to initialize the encryption process, the CBC mode requires an initial vector IV as input before the encryption of the first data block. The design of IV needs to meet the requirements of randomness and uniqueness. It ensures that even if the same plaintext and key are encrypted at different times, completely different ciphertexts will be generated. This mechanism effectively prevents the possibility of repeated data patterns being exploited by attackers. In this invention, IV is generated based on the timestamp of the data. By performing hashing or pseudo-random number transformation on these dynamic parameters, the IV is ensured to be unpredictable. This approach not only increases the dynamic nature of encryption, but also makes data more secure during storage and transmission.
[0063] The core advantage of the AES-CBC mode lies in the chain dependency between data blocks. The generation of each ciphertext block depends on the output of the previous ciphertext block. This chain structure prevents an attacker from independently decrypting the original plaintext, even if they gain access to a ciphertext block, because the input of the previous ciphertext block is a prerequisite for the decryption process. Furthermore, if a ciphertext block is tampered with during the decryption process, the chain structure prevents all subsequent data blocks from that block from being correctly decrypted, thereby protecting data integrity. In the context of the present invention, this feature is particularly suitable for data management in distributed storage environments, as each encrypted data block may be distributed across different storage nodes. The chain dependency of CBC provides additional guarantees for data consistency and tamper resistance in this distributed structure. Furthermore, the performance of the AES-CBC mode is well suited to the requirements of distributed encrypted databases. The encryption and decryption computational complexity is relatively low with modern hardware support, enabling efficient processing of large-scale data. Furthermore, since distributed storage systems typically split large files into multiple small data blocks for storage, the AES-CBC mode is naturally well-suited for this block-based encryption scenario. The encryption of each data block is completed independently. Although they are logically linked through chain dependencies, they can be physically distributed to different nodes for parallel processing, significantly improving the system's encryption and decryption efficiency. In a distributed storage system, the distribution of data blocks is closely related to the management of encryption keys. The present invention uses a key management system to bind dynamically generated keys to encrypted data blocks and record metadata for the encryption process of each data block. This metadata includes information such as the encryption key identifier, IV value, and encryption algorithm mode. This information is securely stored in the key management system and dynamically retrieved during data decryption. In this way, the system can quickly locate encryption keys and complete decryption operations in highly concurrent access scenarios. In addition, the key management system also supports periodic key rotation and expiration management. When a key reaches its expiration date, the system triggers a re-encryption operation, generates a new key, and encrypts the relevant data blocks, further enhancing data security. Compared with traditional static keys or simple encryption modes, the present invention achieves comprehensive improvements in dynamism, attack resistance, and flexibility. The disadvantage of traditional static keys is that once the key is leaked, all related data will be at risk. However, the present invention combines dynamic key generation with the AES-CBC mode, so that even if the key of a certain time period is leaked, the attacker cannot crack the data of other time periods. In addition, the random IV design of the AES-CBC mode ensures that the ciphertext of the same data is completely different when encrypted at different times, thereby preventing known plaintext attacks and pattern recognition attacks. Given the high concurrency and multi-node characteristics of distributed storage systems, the encryption scheme of the present invention can efficiently coordinate between nodes and be uniformly controlled by the key management system, achieving a perfect balance between security and performance.
[0064] Example 6: Preliminary average optimal number of fragments for all encrypted data blocks Calculated by the following formula:
[0065] ;
[0066] in, For the Memory usage of storage nodes; For the The number of disk input and output operations per second for each storage node; is the determinant operator; is the number of storage nodes; For the The bandwidth capacity of each storage node; For the The CPU usage of each storage node.
[0067] Specifically, the core idea of the formula is to The resource capacity of each storage node is mathematically modeled. The resource status of each node is represented in matrix form, which includes memory usage , disk IOPS , bandwidth capacity , and CPU usage Through determinant calculation, these indicators are combined into a scalar to reflect the overall performance of a single node. The introduction of the determinant can not only capture the interaction between resource indicators, but also give certain key parameters (such as bandwidth or IOPS) a higher weight through the diagonal advantage of the matrix, thereby enhancing the sensitivity to resource bottlenecks in the system. For example, when the bandwidth capacity of a node is When the CPU usage is high, the determinant value will increase significantly, indicating that the node has an advantage in supporting high-throughput data transmission tasks; When it is lower, it will also have a positive impact on the determinant value, indicating that the node has greater potential in processing encryption or decryption tasks. It is one of the important parameters in the formula, which directly affects the node's cache capacity and storage performance. Nodes with lower memory usage generally have higher resource margins, can support frequent access requests, and improve data access response speed. Therefore, during the sharding optimization process, memory usage becomes one of the key criteria for evaluating node priority. Disk IOPS It measures the read and write performance of the storage device and is an important indicator for handling high-frequency access tasks. For encrypted data blocks that need to be read or written frequently, the priority of high IOPS nodes will be significantly increased. Bandwidth capacity The formula reflects the node's network transmission capacity. This is especially true in distributed storage systems, where synchronization and user access across multiple nodes place significant demands on bandwidth. When nodes have high bandwidth capacity, the sharding strategy prioritizes frequently transmitted or accessed data blocks to these nodes, reducing network latency and improving access performance.
[0068] In addition, the CPU usage in the formula Represents the computing resource load of the node. For encrypted database systems, operations such as encryption, decryption, and integrity verification of data blocks all consume a lot of CPU resources. Therefore, nodes with low CPU usage can better support these computationally intensive tasks and thus receive higher priority in the sharding strategy. The formula combines these resource parameters and quantifies the comprehensive performance score of each node through determinant calculation, which can fully reflect the node's capabilities in storage, computing, and transmission. By normalizing the scores of all nodes, the formula finally calculates the preliminary average optimal number of shards. This value represents the global optimization result of the distribution of encrypted data blocks across the entire distributed storage system, enabling it to assign a reasonable number of shards and distribution strategy to each type of data. For example, for frequently accessed data blocks, the system prioritizes assigning them to nodes with higher performance scores; whereas for less frequently accessed cold data, the system selects nodes with lower performance scores but sufficient storage capacity. This strategy not only maximizes the utilization efficiency of system resources but also avoids single-point overload or performance bottlenecks by balancing the load. The determinant operation in the formula also has the characteristic of dynamic adaptation. When the resource status of a node changes, such as an increase in memory usage or a decrease in bandwidth capacity, its determinant value adjusts accordingly, affecting the global optimization result. This dynamic adjustment mechanism enables the distributed storage system to achieve a real-time balance between resource utilization and data distribution, ensuring that the system can continue to operate efficiently in high-concurrency and complex business scenarios. In addition, through a comprehensive evaluation of node resources, the formula integrates previously dispersed resource information into a unified performance metric, simplifying the design and implementation of sharding strategies and providing high practicality. Compared to traditional sharding strategies, the formula proposed in this invention establishes a strong correlation between data characteristics and node performance. Traditional strategies typically use data size or simple random distribution as the basis for sharding, ignoring differences in storage node resource status and easily leading to resource waste or performance bottlenecks. This invention, however, dynamically adjusts data distribution and the number of shards by comprehensively evaluating multiple dimensions, including memory, IOPS, bandwidth, and CPU utilization, achieving dual optimization of resource utilization and data access performance. Furthermore, the formula design fully considers the dynamic nature of resources in distributed storage systems, enabling the system to adapt to load fluctuations during operation, further improving its robustness and stability.
[0069] Example 7: Access Control Vector Calculated by the following formula:
[0070] ;
[0071] in, The user's permission level is an integer from 1 to 5. The CPU usage allowed for the access corresponding to the user, The number of concurrent accesses allowed for this user. The number of disk input and output operations per second allowed for the access corresponding to this user; The access bandwidth allowed for the user; The average response time allowed for the access corresponding to the user; The memory usage allowed for access by the user; No. A hash value of the timestamp of the class data; is the minimum function; For the The number of concurrent storage nodes; For the The average response time of each storage node.
[0072] The access control vector in the specific formula It is a six-dimensional vector, each dimension of which represents the upper limit of the resources that the user is allowed to access in the system, including CPU usage , concurrent number , disk IOPS ,bandwidth , average response time and memory usage These dimensions comprehensively describe the dynamic relationship between user access rights and system resources, which can effectively limit the excessive consumption of system resources by user access behaviors, while providing priority protection for access by high-priority users and highly sensitive data. It is one of the core variables in the calculation of access control vector. The value range of permission level is an integer from 1 to 5. The higher the value, the greater the user permission. This design provides the basis for the system to implement multi-level access control. The system can adjust resource allocation priorities for different users globally. For example, high-authority users (such as system administrators or key business users) will receive greater power in formula calculations. Weights are assigned to privileged users, allowing them to occupy more system resources when accessing data, such as higher CPU usage, greater bandwidth, and lower response times. Conversely, resource allocation for low-privilege users is restricted, preventing them from affecting system performance and security.
[0073] The operating parameters of the storage node are calculated by the minimum function The introduction of formulas provides dynamic constraints for access control vectors. Specifically, the resource upper limit of each dimension is determined by the minimum value of the corresponding parameters of all storage nodes in the system. For example, CPU usage The minimum CPU usage of all storage nodes Determined by bandwidth The minimum bandwidth capacity of the node is This minimum constraint ensures that the user access rights set will not exceed the current minimum resource level of the system, avoiding the problem of global resource over-allocation caused by insufficient performance of a single node. Data characteristics are determined by exponential decay function and timestamp hash value. The calculation of access control vector is introduced to reflect the impact of data dynamic attributes on user access rights. Describes the key strength , data size , and response time requirements A comprehensive effect on access rights. Data with stronger key strength usually has higher security requirements, so the constraints on user permissions in the access control vector are more stringent; data with larger data size or lower response time requirements will further increase the difficulty of access rights to prevent large-scale data access from impacting system resources. Timestamp hash value The introduction of gives the access control vector dynamics and randomness, allowing user permissions to be automatically updated over time, further improving the security of the system. The formula also includes the accumulation and normalization of the timestamp hash values of all data categories. Specifically, by By summing and normalizing the timestamp hash values of the data, the system can globally balance the access rights of different data categories. This global processing method helps to ensure the access priority of highly sensitive data in multi-user concurrent scenarios, while limiting the excessive access of low-priority users to non-critical data. The calculation result is output in the form of a six-dimensional vector, which contains the specific limit value of each resource dimension. Possibly , indicating that the user can occupy a maximum of 20% of the CPU usage, allow 100 concurrent requests, use a maximum of 500 disk I / O operations, occupy 1Gbps of bandwidth, ensure the average response time is less than 100ms, and limit the memory usage to below 50%. This quantitative access restriction can be directly applied to the access control module of the distributed storage system to make real-time judgments and restrictions on user access requests. Compared with the traditional static access control model, the access control vector formula in the present invention has significant dynamics and adaptability. Traditional methods often set user permissions based on predefined fixed rules and cannot be dynamically adjusted according to system status and data characteristics. This formula combines user permission levels, node operating parameters and data dynamic characteristics to enable access control to be automatically updated as system resources change, ensuring that the system can still maintain stability and efficiency in high concurrency and complex business scenarios. At the same time, the introduction of timestamps makes the calculation results of access rights timely, effectively preventing the abuse of long-term idle permissions.
[0074] Example 8: Calculate the verification value using the following formula to complete data integrity verification and consistency check:
[0075] ;
[0076] in, No. The verification value of the encrypted data block corresponding to the class data. When the verification value is within the set threshold range, it means that the encrypted data block corresponding to the class data has passed the verification and is allowed to be accessed; otherwise, access is denied; for The hash value of the encrypted data block corresponding to the class data; for The error correction code of the encrypted data block corresponding to the class data; for The message authentication code of the encrypted data block corresponding to the class data.
[0077] Specifically, verify the value It is a comprehensive indicator that dynamically evaluates whether the data block meets the requirements of integrity and consistency by performing nonlinear calculations on multiple verification parameters of the data block, combining the business attributes of the data and the optimization results of the storage system. When the verification value is within the set security threshold range, it means that the data block has not been tampered with or damaged, and the system allows access to the data block; otherwise, the access request will be rejected. This verification mechanism can not only efficiently identify data damage or illegal tampering, but also adapt to different data security requirements and system performance requirements through threshold adjustment. In the formula, the hash value It is the basis for data integrity verification. Each encrypted data block generates a unique hash value when it is stored. This value reflects the fixed fingerprint of the data content. When the data is accessed, the system recalculates the hash value of the current data block and compares it with the hash value recorded when it was stored. A comparison is performed. If the two are consistent, it indicates that the data block has not been tampered with or damaged during storage and transmission. Hash value comparison, as a core part of the formula, can quickly and efficiently complete basic integrity verification.
[0078] Error Correcting Code The robustness of data verification is further enhanced in the formula. Unlike hash values, error-correcting codes can not only detect errors in data blocks, but can also automatically repair errors to a certain extent. For example, if certain bits are flipped during data transmission, error-correcting codes can restore the original data through redundant information. This feature allows the system to automatically complete repairs even when faced with relatively small-scale data corruption without having to deny access requests, thereby improving data availability and the system's fault tolerance. Message Authentication Code The authentication function is introduced for data verification. It can not only verify the integrity of the data, but also ensure the credibility of the source of the data. The message authentication code is composed of the encryption key It is generated together with the data block content, and its value is closely related to the encryption key. When an attacker tries to tamper with the data, even if he forges a new hash value or error correction code, he cannot access or reproduce it. , generated The addition of message authentication codes greatly improves the system's ability to resist malicious tampering and forgery, providing additional protection for data security in distributed storage systems. The key strength, data size and response time requirements are combined to dynamically adjust the weight distribution of verification values. The larger the size, the higher the security requirement of the data and the stricter the verification requirement. and response time It regulates the dynamic characteristics of the formula, ensuring that the calculation of the verification value can balance the relationship between security and performance in large data blocks or high real-time scenarios. In addition, the timestamp hash The introduction of gives the formula dynamics and timeliness, ensuring that the verification value can be automatically adjusted over time to prevent long-term static verification parameters from being exploited by attackers. When calculating the verification value, the formula also introduces a global optimization parameter and access control vectors . The optimal number of shards is obtained through global storage node resource evaluation, reflecting the current storage status and performance indicators of the system. Including it in the verification value calculation can ensure the consistency of the data verification process and system performance optimization, and avoid excessive verification load affecting the overall performance of the system. Access Control Vector The combination of user permissions and system resource allocation provides dynamic constraints for data validation, further improving the adaptability and refinement of the formula. This is a comprehensive value generated by the combined effects of the above parameters. When the verification value falls within the preset security threshold, it indicates that the data block has passed integrity and consistency checks, and the access request is allowed. Otherwise, the system denies access and may trigger the corresponding security incident handling process. Through this mechanism, the system can efficiently and reliably maintain data security and consistency in a distributed storage environment.
[0079] Example 9: The following formula is used to guide the average optimal number of shards for all encrypted data blocks in the distributed storage system to evolve toward the optimal state:
[0080] ;
[0081] in, For time The load balancing factor is calculated as follows:
[0082] ;
[0083] in, For time The total number of concurrent requests at time; For time Total data access load at time ; For time The total number of concurrent requests at the time.
[0084] Specifically, the formula It shows that the optimal number of fragments for the encrypted data block is the average optimal number of fragments calculated initially. and time Load balancing factor when Joint decision. Here, It is the static optimization result calculated in Example 6, which provides the basic sharding solution of the system based on the resource status of the storage node (such as bandwidth, CPU usage, memory usage, etc.). It is a dynamic adjustment factor that can dynamically modify the sharding strategy according to the real-time load and concurrent requests of the system at different time points, so that it can gradually evolve to the optimal state. The calculation formula of is complex and sophisticated, reflecting an in-depth analysis of system resource utilization and load distribution characteristics. The formula is divided into two main parts: the first part Describes a probability density function of the Poisson distribution, which is used to model the randomness and volatility of data access requests in the system. The parameters of the Poisson distribution are Indicates time The greater the total load, the higher the system load pressure. Is the total number of concurrent requests, used to describe the number of requests the system handles at the same time. and When the ratio is close to 1, the probability density value of the Poisson distribution is the largest, indicating that the system is in a state of balanced resource utilization; when the ratio deviates from 1, the efficiency of load distribution decreases. The second part of the formula is a normalized gamma distribution function used to describe the system's responsiveness and resource allocation efficiency under high-load scenarios. The shape parameter and scale parameter of the gamma distribution are respectively given by and Decision, among which Describes the extent to which concurrent requests consume system resources. It represents the overall ability of the system to handle these requests. The exponential decay term in the distribution function High load scenarios are constrained to ensure that When the number of shards for a single request is reduced, the system can maintain overall performance. Dynamically balances system load and concurrent requests The relationship between It can be adjusted dynamically according to the actual operation situation. For example, when the total system load Higher but concurrent requests When the value is moderate, the weight of the Poisson distribution part is large, and the system will give priority to increasing the number of shards to relieve node pressure; Increase to nearly When the limit is reached, the weight of the gamma distribution part increases, and the system will gradually reduce the number of shards to ensure overall load balance. In addition, the integral term in the formula is a specific form of the gamma function, which is used to normalize the distribution results. Through this integral term, the formula can generate a comparable result under different load and concurrency conditions. This allows distributed storage systems to quickly adapt to new sharding strategies in highly dynamic scenarios without having to recalculate the entire storage layout.
[0085] The present invention has been described in detail above. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only intended to help understand the method and core ideas of the present invention. It should be noted that, for those skilled in the art, without departing from the principles of the present invention, several improvements and modifications may be made to the present invention, and such improvements and modifications also fall within the scope of protection of the claims of the present invention.
Claims
1. An efficient data management method for encrypted databases based on a key management system, characterized in that: The method comprises: Step 1: The database is stored in a distributed storage system. Based on the business attributes of the data in the database, a multi-dimensional feature vector is created for each type of data, and a unique encryption key is generated from this vector. Each type of data is encrypted using the encryption key to generate the corresponding encrypted data block. Based on the operating parameters of each storage node, the comprehensive evaluation value of each storage node in the distributed storage system is calculated to determine the resource capacity of each storage node, and the preliminary average optimal number of shards for all encrypted data blocks in the distributed storage system is determined. Step 2: An access control vector is generated by combining the operating parameters of each storage node and the permission level allowed for the user's role in the distributed storage system. This vector is then constrained by the time dimension to prevent timeout or unauthorized access, thereby dynamically strengthening the access control vector. The access control vector is dynamically refreshed when the time or key is updated. Step 3: For each encrypted data block, the hash value, cyclic redundancy check value, error correction code, and message authentication code are calculated and normalized accordingly, combined with the access control vector, to complete data integrity verification and consistency checks. Based on the real-time total data access load and the total number of concurrent requests, a load balancing factor is calculated. Combined with the preliminary average optimal sharding number, the average optimal sharding number of all encrypted data blocks in the distributed storage system is guided towards the optimal state. Set up the first The multidimensional feature vector of the class data is: ; in, For the Sensitivity score of class data; For the The frequency with which the data of the class data is accessed; For the Data type identifier of class data; For the Response time requirements for class data; For the Data size of class data; For the The number of concurrent users accessing class data; For the The encryption level of the class data; For the The timestamp of the class data; a unique encryption key is generated using the following formula : ; in, is the number of categories of data; Represents the L1 norm of the multidimensional feature vector; is the set encrypted large prime number; is the modulo operation; express The hash value of The initial average optimal number of fragments for all encrypted data blocks Calculated by the following formula: ; in, For the Memory usage of storage nodes; For the The number of disk input and output operations per second for each storage node; is the determinant operator; is the number of storage nodes; For the The bandwidth capacity of each storage node; For the The CPU usage of each storage node.
2. The method for efficient data management of encrypted database based on key management system according to claim 1, characterized in that: The operating parameters of the storage node include: bandwidth capacity, CPU usage, memory usage, disk input and output operations per second, average response time, and concurrency.
3. The method for efficient data management of encrypted database based on key management system according to claim 2, characterized in that: The business attributes include: sensitivity score, data access frequency, data type identifier, response time requirement, data size, number of concurrent access users, encryption level and timestamp; the sensitivity score ranges from 1 to 10, indicating the security sensitivity level of this type of data, and the higher the sensitivity score, the greater the contribution to the key; the data access frequency indicates the access frequency of this type of data; the data type identifier takes a value of 1, 2 or 3, when the value is 1, it indicates structured data, when the value is 2, it indicates semi-structured data, and when the value is 3, it indicates unstructured data; the response time requirement indicates the response time constraint for this type of data, and the smaller the value, the higher the real-time requirement; the encryption level ranges from 1, 2 or 3, and the higher the value, the higher the corresponding encryption level.
4. The method for efficient data management of encrypted database based on key management system according to claim 3, characterized in that: Each type of data is encrypted using the encryption key using the CBC mode of the AES encryption algorithm.
5. The method for efficient data management of encrypted database based on key management system according to claim 4, characterized in that: Access Control Vector Calculated by the following formula: ; in, The user's permission level is an integer from 1 to 5. The CPU usage allowed for the access corresponding to the user, The number of concurrent accesses allowed for this user. The number of disk input and output operations per second allowed for the access corresponding to this user; The access bandwidth allowed for the user; The average response time allowed for the access corresponding to the user; The memory usage allowed for access by the user; No. A hash value of the timestamp of the class data; is the minimum function; For the The number of concurrent storage nodes; For the The average response time of each storage node.
6. The method for efficient data management of encrypted database based on key management system according to claim 5, characterized in that: The verification value is calculated using the following formula to complete data integrity verification and consistency check: ; in, No. The verification value of the encrypted data block corresponding to the class data. When the verification value is within the set threshold range, it means that the encrypted data block corresponding to the class data has passed the verification and is allowed to be accessed; otherwise, access is denied; for The hash value of the encrypted data block corresponding to the class data; for The error correction code of the encrypted data block corresponding to the class data; for The message authentication code of the encrypted data block corresponding to the class data; for The cyclic redundancy check code of the encrypted data block corresponding to the class data.
7. The method for efficient data management of encrypted database based on key management system according to claim 6, characterized in that: The following formula is used to guide the average optimal number of shards for all encrypted data blocks in the distributed storage system toward the optimal state: ; in, For time The load balancing factor is calculated as follows: ; in, For time The total number of concurrent requests at time; For time Total data access load at time ; For time The total number of concurrent requests at the time.
Citation Information
Patent Citations
Data encryption management system and data encryption method
CN118153081A
Method and device for ensuring data security of distributed storage system
CN118862170A