Cloud storage data encryption and authority management platform

The cloud storage data encryption and permission management platform addresses semantic unawareness and coarse-grained control by employing semantic-aware segmentation and dynamic key generation, enhancing security and efficiency through fragmented and disguised content management.

CN120321023APending Publication Date: 2025-07-15SHANGHAI XINGYI ANYI TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510658604.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

When dealing with massive heterogeneous data, traditional cloud storage data encryption and permission management platforms have problems such as lack of semantic perception, coarse security granularity, rigid permissions, and insufficient storage elasticity, resulting in lack of semantic perception of data segmentation, coarse permission management cannot meet the needs of differentiated authorization, high risk of data exposure, and inefficient resource scheduling.

Method used

The file is dynamically segmented by semantic perception segmentation technology, sub-file units are generated and encrypted, and semantic expansion is combined with the distributed processing module. Through the multi-level key separation and dynamic combination mechanism, fine-grained permission management and dynamic node allocation are realized, hardware security modules are integrated for self-destruct protection, and a dual-key protection system of "content encryption-storage location association" is built.

Benefits of technology

Significantly reduce the risk of data being detected, improve system flexibility and processing efficiency, enhance attack resistance, and ensure data security and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120321023A_ABST
    Figure CN120321023A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer communication and data security, in particular to a cloud storage data encryption and authority management platform which comprises a user side, a transmission channel module, a distributed processing module, a cloud server and a key management module. A user side divides a file to generate sub-file units and encrypts the sub-file units, and a secret key is stored in a hardware security module; the transmission channel module transmits the sub-file units to the distributed storage nodes; the distributed processing module generates extended data and a storage position key; the cloud server manages and ranks the keys; and the key management module generates a combined key and distributes the combined key according to authority. Through the block chain technology and the semantic hiding rule, the original content is disguised in a fragmented manner, and the detection risk is reduced; and a multi-stage key separation and dynamic combination mechanism is adopted, so that full-disk data exposure caused by leakage of a single key is avoided, and the data security is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer communication and data security, and specifically, to a cloud storage data encryption and permission management platform. Background Art

[0002] In recent years, the rapid development of the digital economy has promoted the leap of data from the "resource" to the core position of "production factor", and the storage security and efficient management of data have become the key supports for the high-quality development of the digital economy. However, when dealing with massive heterogeneous data, traditional cloud storage data encryption and permission management platforms expose core pain points such as lack of semantic perception, coarse security granularity, rigid permissions, and insufficient storage elasticity. At the same time, blockchain technology, with its characteristics of decentralization, immutability, intelligent contract automation, and distributed collaboration, provides an innovative path for the trusted and fine-grained management of the entire data life cycle, promoting the evolution of cloud storage technology towards the coordinated optimization of "semantics-security-efficiency".

[0003] The core defects of traditional solutions can be summarized into three aspects:

[0004] First, data segmentation and single-key encryption based on fixed rules, mechanically segmenting using a preset shard size (such as 1GB / shard) or file format identifiers (such as document page breaks, video timestamps), and only applying a single key to the shard data during the encryption process, with the key stored in the cloud server or the user-side software environment;

[0005] Second, static role-based permission control, allocating permissions through access control lists (ACLs) or preset roles (such as administrators, users), with the permission rule being a binary mode of "entire file allowed / denied", and unable to perform differential authorization for the content within the file;

[0006] Third, semantic-irrelevant data augmentation and static storage, where data augmentation is only used to generate retrieval tags (such as keyword extraction, hash value calculation), without involving content hiding or semantic confusion; distributed storage nodes use a fixed resource pool, and node allocation has nothing to do with the data content, resulting in low resource utilization;

[0007] However, the above-mentioned existing technologies have the following problems:

[0008] First, data segmentation and encryption lack semantic perception. Fixed segmentation may cut off the semantic coherence of text (such as splitting a complete paragraph into two pieces) or damage the integrity of the image main body, and if the single key is leaked, all shard data will get out of control, with a coarse security granularity;

[0009] Second, the risk of data exposure is high and resource scheduling is inefficient. The original data features directly exist in the shard content, and attackers can locate sensitive information through content analysis; the static allocation of storage nodes leads to resource idleness when small-scale data is segmented and node performance bottlenecks when large-scale data is segmented, resulting in insufficient system elasticity;

[0010] In view of this, a cloud storage data encryption and permission management platform is proposed. Summary of the Invention

[0011] The purpose of the present invention is to provide a cloud storage data encryption and permission management platform, which is used to solve the problems of "in the prior art, the segmentation of cloud storage data lacks semantic awareness, resulting in incomplete content, the coarse-grained permission management cannot meet the differentiated authorization requirements, the risk of data exposure is high, and the resource scheduling is inefficient".

[0012] To solve the above technical problems, the present invention provides a cloud storage data encryption and permission management platform, including:

[0013] A user terminal, which is used to receive multiple files uploaded by the user, perform segmentation processing on the content of each file through a segmentation module, generate multiple sub-file units arranged in the original order of each file, and encrypt each sub-file unit to generate a corresponding sub-file key sequence, and the sub-file key sequence is stored in the hardware security module of the user terminal;

[0014] A transmission channel module, which is connected to the user terminal and is used to transmit the sub-file units to the distributed storage nodes one by one;

[0015] A distributed processing module, which is deployed in the distributed storage nodes and is used to perform semantic expansion processing on each sub-file unit to generate extended data, and generate a storage location key according to the storage location of the extended data;

[0016] A cloud server, which is connected to the transmission channel module and is used to receive and manage the storage location keys, and form a storage location key sequence in order;

[0017] A key management module, which is deployed on the user terminal and is used to correspond the sub-file key sequence with the storage location key sequence issued by the cloud server to generate a combined key, and dynamically allocate the proportion of the combined key according to the preset permission level to generate a differentiated access key corresponding to the permission level.

[0018] As a further improvement of this technical solution, the segmentation module includes:

[0019] A format parsing and preprocessing unit, which is used to obtain multiple files uploaded by the user, perform format parsing on the content of the multiple files, and at the same time, define the part carrying information in the multiple files as the entity part area, and the rest as the non-entity part area;

[0020] The segmentation position determination unit is connected to the format parsing and preprocessing unit. It determines the number of segments according to the size of each file, preliminarily determines the segmentation positions based on the number of segments, detects each segmentation position, judges whether the segmentation position exists in the entity part area. When it exists in the entity part area, it finds the nearest non-entity part area to the segmentation position as the new segmentation position;

[0021] The segmentation unit is connected to the segmentation position determination unit, obtains the segmentation positions and segments the files.

[0022] As a further improvement of this technical solution, the hardware security module includes an encryption chip and a secure storage circuit, which are used to generate and store the sub-file key sequence. When illegal disassembly, abnormal voltage or abnormal temperature is detected, the secure storage circuit triggers an automatic destruction mechanism to destroy the stored sub-file key sequence.

[0023] As a further improvement of this technical solution, the transmission channel module dynamically activates the same number of distributed storage nodes according to the number of sub-file units generated by the user side for segmenting each file; each distributed storage node corresponds to a sub-file unit one by one.

[0024] As a further improvement of this technical solution, the user side further includes an identification module, which is used to generate a unique increasing number identification for the sub-file units in the original order of the files when the sub-file units are transmitted.

[0025] As a further improvement of this technical solution, the distributed processing module includes a resource library and a storage library, where:

[0026] The resource library has a preset theme, crawls network public data, domain knowledge bases or historical stored data around the preset theme to form a content collection related to the theme; when the sub-file unit is transmitted to the distributed processing module, it identifies the content of the sub-file unit, performs semantic transformation through hiding rules, and based on the content collection of the resource library, performs semantic extension on the semantically transformed content through a generative AI model to generate extended data related to the preset theme;

[0027] The storage library is connected to the resource library, stores the extended data, and generates a corresponding file storage location key through the locality-sensitive hashing algorithm according to the storage location of the extended data. After the file storage location key is generated, the same number identification as the sub-file unit is added.

[0028] As a further improvement of this technical solution, the hiding rules include:

[0029] The semantic segmentation unit is used to semantically segment the content of the sub-file unit into multiple fragments and generate at least two alternative semantic interpretations for each fragment;

[0030] An adversarial generative expansion unit, connected to the semantic segmentation unit, generates an extended paragraph containing different alternative semantics for each fragment based on the content set of the resource library through a generative adversarial network;

[0031] A context recombination unit, connected to the adversarial generative expansion unit, inserts the extended paragraphs into the base text according to discontinuous logic, where the discontinuous logic includes timestamp jumps, role permutations, and injection of irrelevant logical chains, and dynamically switches the extended paragraphs according to the real-time feedback of the full-text semantic discriminator; the base text is generated by the adversarial generative expansion unit;

[0032] A fragment mapping key management unit, connected to the context recombination unit, generates a unique mapping key slice for each fragment, records its alternative semantic version and the position offset of the fragment in the extended data, and combines multiple mapping key slices in the order in the extended data to form a mapping key, where the position offset of the fragment in the extended data is the position of the extended paragraph generated by the fragment in the extended data;

[0033] A dynamic self-destruction unit, in response to an abnormal access warning from the cloud server module, implements a hierarchical self-destruction strategy.

[0034] As a further improvement of this technical solution, the cloud server sorts the storage location keys according to the number identifiers, and removes the number identifiers after sorting.

[0035] As a further improvement of this technical solution, the key management module is communicatively connected to the hardware security module; the sub-file key sequence and the storage location key sequence are matched and corresponding according to the arrangement order, and the encrypted algorithm is used to splice the matched sub-file key and the storage location key to generate a combined key sequence consistent with the order of the sub-file units.

[0036] As a further improvement of this technical solution, the preset permission level is the user role level, and the user role level includes administrator, editor, and visitor;

[0037] Denote the total number of key fragments in the combined key sequence as N, the preset permission level as K, the permission of the administrator is K = N×100%, the permission of the editor is K = N×A%, and the permission of the visitor is K = N×B%, where 100% > A% > B% > 0.

[0038] Compared with the prior art, the beneficial effects of the present invention:

[0039] In this cloud storage data encryption and permission management platform, through the semantic hiding rule, the original content is fragmented and disguised as irrelevant information, significantly reducing the risk of data detection; the multi-level key separation and dynamic combination mechanism avoid the exposure of all data caused by the leakage of a single key, thereby improving data security.

[0040] In this cloud storage data encryption and permission management platform, the dynamic node allocation mechanism accurately matches resources according to the number of sub-files, improving the system elasticity and processing efficiency. At the same time, the one-to-one correspondence between nodes and sub-files simplifies the scheduling logic, reduces system overhead, and improves resource utilization.

[0041] In this cloud storage data encryption and permission management platform, the discontinuous logical recombination and adversarial generative expansion make the data content confusing and increase the analysis difficulty for attackers. At the same time, the dynamic self-destruction mechanism automatically erases key information when detecting anomalies, protects the security of the remaining data, and enhances the anti-attack ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 It is a structural schematic diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0044] In recent years, the rapid development of the digital economy has promoted the leap of data from the "resource" to the core position of "production factor", and the storage security and efficient management of data have become the key support for the high-quality development of the digital economy. However, when dealing with massive heterogeneous data, traditional cloud storage data encryption and permission management platforms expose core pain points such as lack of semantic perception, coarse security granularity, rigid permissions, and insufficient storage elasticity. At the same time, blockchain technology, with its characteristics of decentralization, immutability, intelligent contract automation, and distributed collaboration, provides an innovative path for the whole life cycle management of data, promoting the evolution of cloud storage technology towards the coordinated optimization direction of "semantics-security-efficiency";

[0045] Currently, existing cloud storage data encryption and permission management platforms generally have problems such as lack of semantic perception in segmentation resulting in incomplete content, coarse-grained permission management unable to meet differentiated authorization requirements, high data exposure risk, and inefficient resource scheduling;

[0046] In view of this, please refer to Figure 1 As shown, the purpose of the present invention is to provide a cloud storage data encryption and permission management platform, which includes:

[0047] The client is used to receive multiple files uploaded by users, perform segmentation processing on the content of each file through a segmentation module, generate multiple sub-file units arranged in the original order of each file, and encrypt each sub-file unit to generate a corresponding sub-file key sequence, which is stored in the hardware security module of the client;

[0048] The transmission channel module is connected to the client and is used to transmit the sub-file units to the distributed storage nodes one by one;

[0049] The distributed processing module is deployed in the distributed storage nodes and is used to perform semantic expansion processing on each sub-file unit to generate extended data, and generate a storage location key according to the storage location of the extended data;

[0050] The cloud server is connected to the transmission channel module and is used to receive and manage the storage location keys, and form a storage location key sequence in order;

[0051] The key management module is deployed on the client and is used to correspond the sub-file key sequence with the storage location key sequence issued by the cloud server to generate a combined key, and dynamically allocate the proportion of the combined key according to the preset permission level to generate a differentiated access key corresponding to the permission level.

[0052] In this cloud storage data encryption and permission management platform, through a multi-module collaborative architecture and innovative technology mechanisms, such as storing the sub-file key sequence in the client hardware security module, realizing hardware-level protection of key generation, storage, and use through an encryption chip and an anti-tampering circuit, avoiding the risk of key leakage in the software environment, and ensuring the physical and logical security of core encryption information from the source; at the same time, performing semantic-aware segmentation and confusion, the segmentation module intelligently identifies the "entity part area" of the file, dynamically adjusts the segmentation position to preserve the content integrity (such as avoiding cutting off text paragraphs and image main bodies); the distributed processing module generates extended data through semantic expansion, fragments the original content and embeds associated topic information in combination with the hiding rules, realizing semantic-level content camouflage, and significantly reducing the possibility of sensitive data being detected; moreover, the sub-file key and the storage location key are generated separately and combined synergistically to form a dual key protection system of "content encryption - storage location association". Even if a single key is leaked, the complete data cannot be restored, breaking through the security bottleneck of traditional single keys.

[0053] Considering that traditional cloud storage segmentation technologies use fixed sizes or format identifiers for mechanical segmentation, often resulting in broken text paragraphs, incomplete image main bodies, or cut-off audio semantic units, seriously damaging content integrity and increasing the complexity of subsequent encryption and permission management. Therefore, a semantic-aware dynamic segmentation technology is adopted in the segmentation module. The specific technical path and advantages are as follows:

[0054] Considering the differences in the effective information distribution of different types of files (documents, images, audio) (such as text paragraphs in documents, pixel matrices in images, and peak regions of sound waves in audio), the traditional "one-size-fits-all" segmentation method cannot recognize these semantic units and is prone to fragmenting key information. Therefore, a format parsing and preprocessing unit is used to obtain multiple files uploaded by the user, parse the formats of the contents of the multiple files, and at the same time, define the part carrying information in multiple files as the entity part region and the rest as the non-entity part region; specifically, the paragraph structure and logical chapters of text documents are parsed through natural language processing (NLP), the main body regions in images (such as people, object contours) are recognized using computer vision (CV), and the effective sound wave bands in audio are detected based on digital signal processing (DSP) (eliminating silent or noisy intervals), and all are uniformly defined as the "entity part region", and the remaining irrelevant parts (such as blank pages in documents, solid-color backgrounds in images, silent segments in audio) are defined as the "non-entity part region"; through the above technical means, semantic-level understanding of the file content is achieved, providing "information-sensitive region" annotations for subsequent segmentation, and avoiding the destruction of core data units by the segmentation operation from the source;

[0055] Since directly dividing evenly according to the file size may forcibly truncate in the middle of the entity part region (such as splitting a complete paragraph into two pieces), resulting in the loss of the independent semantic value of the sub-file unit and increasing the difficulty and risk of subsequent semantic expansion, a segmentation position determination unit is set up. The segmentation position determination unit is connected to the format parsing and preprocessing unit, determines the number of segments according to the size of each file, initially determines the segmentation positions based on the number of segments, detects each segmentation position, judges whether the segmentation position exists in the entity part region, and when it exists in the entity part region, finds the non-entity part region closest to the segmentation position as the new segmentation position; enables each sub-file unit to contain complete semantic fragments (such as complete paragraphs, independent image main bodies, continuous audio semantic segments), improves the logical integrity of the segmented data, provides high-quality input for the semantic expansion of the subsequent distributed processing module, and reduces the generation of invalid data fragments;

[0056] Even if the splitting position of the non-entity part area is determined, if the cutting boundary is blurred or contains mixed information (such as containing both text and blanks), it may still cause some sensitive features to be carried by the sub-file units, increasing the risk of data exposure. Therefore, a splitting unit is set up. The splitting unit is connected to the splitting position determination unit to obtain the splitting position and split the file. Specifically: for document files, use the format parsing result to locate logical boundaries such as paragraph separators and chapter titles; for image files, determine the clear boundary between the main body and the background based on the edge detection algorithm; for audio / video files, precisely truncate through key frames or silent segments to ensure that the boundaries of the split sub-file units are pure non-entity part content, thereby generating "clean" sub-file units, avoiding cross-slice leakage of sensitive information, and at the same time reducing the semantic expansion difficulty of the distributed processing module (such as not needing to handle the ambiguity of boundary mixed content), improving the generation quality and efficiency of the extended data.

[0057] Considering that the encryption algorithms implemented by software rely on CPU operations, the encryption keys are easily read by debugging tools in memory, and the operation process may be subject to side-channel attacks (such as power consumption analysis, electromagnetic radiation monitoring), which cannot meet the high-security requirements of sensitive fields such as finance and healthcare. Therefore, an integrated dedicated encryption chip (such as a hardware module meeting the FIPS140-3 security level) is adopted, with an independent operation unit and security firmware built-in to achieve hardware-level encapsulation of operations such as key generation, encryption / decryption, and signature verification; through memory isolation technology, ensure that the encryption key is only generated and used in the internal registers of the chip, and the external interface cannot directly access the original key data, thereby constructing an "operation security island" to completely isolate the key processing process from the operating system and application programs, preventing logical attacks (such as buffer overflow, malicious driver injection) at the bottom layer, and at the same time using chip-level anti-side-channel design (such as random power consumption noise generation, electromagnetic shielding) to enhance the anti-analysis ability of key operations;

[0058] Considering that traditional storage media (such as EEPROM, Flash) lack physical layer protection, attackers can directly read the stored data through chip probing, and even restore the erased encryption key through cryogenic power-off technology. Therefore, metal packaging and potting processes are used, with voltage / temperature sensors, pressure sensing circuits, and anti-disassembly microswitches built-in to form a three-dimensional protective shell; at the same time, one-time programmable (OTP) memory is used to store the encryption key. Once illegal physical contact or abnormal environmental parameters (such as voltage drop, temperature exceeding the threshold) are detected, the fuse circuit is triggered to permanently destroy the key storage area, and it cannot be restored; through the above technical solutions, the "physical invisibility" of the encryption key is achieved. Even if the hardware security module is violently disassembled, the attacker can only obtain the destroyed key fragments; the sensor network monitors the working environment in real time to form an active defense system, meeting the high-security requirements of "the device is the security boundary".

[0059] Considering that static storage protection cannot cope with progressive physical attacks (such as slow temperature rise and voltage glitch injection), and traditional security modules lack real-time risk assessment and response capabilities and may continue to run after the key is leaked, an acceleration sensor (to detect severe vibration), a photosensitive element (to detect shell opening), and a voltage fluctuation monitoring module are integrated to build an abnormal behavior recognition model. When illegal disassembly, voltage abnormality or temperature abnormality is detected, a temporary freeze of the key is triggered when a mild abnormality (such as a short voltage fluctuation) is detected, and any operation is prohibited; when a severe abnormality (such as shell disassembly, sensor signal interruption) is detected, a strong current pulse is used to burn the key storage circuit within microseconds, and an unalterable self-destruction log is generated, thereby realizing an upgrade from "passive protection" to "active defense", dynamically responding to different threat levels, completing self-destruction before the key is at risk of leakage, and retaining audit evidence (such as self-destruction timestamp, triggering reason) to meet compliance audit requirements.

[0060] Considering that traditional distributed storage systems use fixed node pools or dynamic allocation strategies based on load balancing, there is a core problem of "mismatch between resource allocation and processing tasks", that is, the idle rate of node resources is high when small-scale file segmentation occurs, and the insufficient concurrent processing capacity of nodes leads to increased latency when large-scale segmentation occurs. In addition, the stateless mapping of nodes and data shards increases routing complexity. Therefore, dynamic node matching and one-to-one binding technology are used in the transmission channel module to achieve precise coupling of "task-node" through intelligent resource scheduling. The specific technical path and advantages are as follows:

[0061] First, considering that the traditional fixed node mode cannot perceive the real-time scale of file segmentation (for example, when a user uploads a single 10MB file and uploads 10,000 10MB files in batches, the node resource usage differs by a thousand times), resulting in the polarization problem of "large tasks stuck and small tasks wasted", especially in the scenario of bursty data upload, the system throughput fluctuates greatly. Therefore, through a lightweight metadata scanner, the number of sub-file units N is obtained immediately after the user-side segmentation is completed. Based on Kubernetes or DockerSwarm technology, N independent container nodes are started on demand in a distributed cluster. Each container encapsulates a complete distributed processing module (resource library, storage repository), and a closed-loop process of "start-process-destroy" is set. After the processing is completed, the node resources are automatically released to avoid memory leaks or zombie processes, thereby achieving a linear match between resource usage and task scale, improving resource utilization, and reducing processing delays. It is especially suitable for high-frequency small file upload scenarios such as short videos and medical images;

[0062] Secondly, traditional node load balancing algorithms (such as round-robin and least connections) allocate multiple shards to the same node, resulting in increased context switching overhead, and a single node failure may affect the processing progress of multiple shards, which does not conform to the principle of "fault isolation". Therefore, through the consistent hashing algorithm, the unique number identifier (such as UUID) of each sub-file unit is mapped to a specific node to ensure that "one node only processes one shard"; the metadata (number, semantic expansion status, key generation progress) of the corresponding sub-file unit is temporarily stored in the local cache of the bound node to avoid network latency caused by cross-node data interaction; each node regularly sends heartbeats to the transmission channel module. If no response is received after the timeout, the standby node is immediately started to take over the task to control the fault recovery time; by eliminating the context interference of shard processing, the processing efficiency of a single node is improved, and at the same time, the accurate positioning of faults and the minimization of impacts (only a single shard processing is interrupted) are achieved, meeting the requirements of scenarios with extremely high reliability requirements such as financial transactions and autonomous driving data uploads;

[0063] Finally, since the dynamically started nodes may have inconsistent processing progress due to network latency and hardware performance differences, if the storage location key is directly uploaded, it is easy to cause disorder in the cloud server side, affecting subsequent key combination and permission verification. Therefore, a global sequence number generator is implemented based on ZooKeeper or etcd to ensure that each node obtains an increasing global timestamp when generating the storage location key. A memory-level FIFO queue is set in the transmission channel module. After the node processing is completed, the keys are sorted and queued according to the number identifier, and the queue uniformly sends them to the cloud server in order. The duplicate keys are detected through the Bloom filter, and the missing shards are confirmed by combining the hash checksum, and the retransmission request is automatically triggered to ensure the consistency of the key sequence received by the cloud server and the original order of the sub-file units, eliminating the risk of disorder in the distributed environment, providing a reliable data basis for the accurate matching of the subsequent key management module, and avoiding permission verification failures or data decryption errors caused by incorrect order.

[0064] Since traditional cloud storage only generates bare metadata or simple associated information during data expansion, the original semantic features of sub-file units are directly exposed, and attackers can quickly locate sensitive data through content analysis, and the existing technology lacks the ability to actively disguise data. Therefore, the topic-associated semantic disguise and dynamic key binding technology are adopted in the distributed processing module to achieve data concealment and orderly management through the collaborative design of the resource library and the storage library. The specific technical path and advantages are as follows:

[0065] Since traditional semantic expansion technologies (such as keyword extraction, hash digest) are only used for retrieval or verification and do not change the original semantic features of the data, attackers can directly associate sub-file units with the original file through means such as text similarity calculation and image feature matching. Especially in scenarios such as medical images and financial statements, sensitive information is easily reverse-inferred by features. Therefore, each resource library dynamically crawls multi-source data (public web pages, professional knowledge bases, historical storage records) around a preset theme (such as "environmental protection policies", "market analysis"), constructs a content collection related to the theme, and forms a "material library" with semantic camouflage; parses the core semantics of the sub-file content through natural language processing (NLP), and uses a generative adversarial network (GAN) to map it to an expression space related to the theme but with a relatively large semantic distance (such as transforming "financial data" into a metaphorical expression in "industry trend analysis"), generating "camouflaged semantics" that contains the original information but whose features are confused; based on models such as GPT and T5, combined with the content of the resource library, fills in the context of the camouflaged semantics to generate logically coherent extended data (such as supplementary paragraphs, associated charts, knowledge graph nodes), so that the original content is fully integrated into the theme context, forming a "semantic bunker", thereby achieving active protection of "content is camouflage". Even if an attacker obtains the extended data, they will be guided to the semantic space of the preset theme and cannot recognize the true meaning of the original content; at the same time, the theme consistency of the extended data improves the subsequent storage and retrieval efficiency and enhances the retrieval accuracy;

[0066] Traditional storage location keys (such as hash values, storage addresses) only identify the storage location of the data, lack association with the original order and content features of the sub-file units, easily lead to key disorder and matching errors in a distributed environment, and cannot provide fine-grained location association information for permission management. Therefore, according to the storage location of the extended data (physical node ID, logical partition address, distributed hash table key value), combined with the content fingerprint (such as the feature vector of the camouflaged semantics), a storage location key is generated to ensure that "keys for the same theme content are similar, and keys for different theme content are significantly different", improving the anti-collision ability; uses the unique incremental number of the sub-file unit generated by the user side (such as 001, 002) as the prefix of the key metadata, forming a "number-key value" structure (such as 001-abc123), so that the key is strongly bound to the original order of the sub-file unit; maintains a three-way index table of "number identifier-storage location-extended data hash" inside the storage library to support quickly locating the corresponding relationship between the key and the data through the number, and at the same time provides underlying data support for the order verification of the cloud server; thus constructs a three-dimensional key system of "content-location-order". The location-sensitive hash combined with the content fingerprint after semantic camouflage makes it difficult to reverse-infer the original data through the storage location, improving security; the number identifier ensures that the key maintains the original order during distributed transmission, improves the sorting efficiency of the cloud server, and avoids verification failures caused by disorder in the traditional scheme;

[0067] If the resource library and the storage library are designed independently, it may lead to a logical break in semantic expansion and key generation (e.g., the key cannot sense when the extended data is tampered with), or the mismatch between the key order and the content order, affecting the generation accuracy of the subsequent combined key. Therefore, after the sub-file unit generates extended data through semantic transformation in the resource library, the key generation process of the storage library is immediately triggered to ensure that "the time difference between data generation and key generation < preset value (e.g., 10 ms)", avoiding being tampered with in the intermediate link. Before generating the key, the storage library synchronously verifies the topic consistency of the extended data (through the content hash of the resource library) and the legality of the sub-file number (through the number verification code of the transmission channel). Only when the two factors match can a valid key be generated. If the resource library detects an update of the topic content (such as adding a domain knowledge base), it automatically triggers the key regeneration process of the storage library to ensure that the semantic relevance between the extended data and the key is always in the latest state, thus forming a security closed-loop of "semantic disguise - key binding - collaborative verification" and eliminating the risks of "naked data" and "key out of control" from the source.

[0068] Considering that traditional data hiding technologies only perform information hiding through simple encryption or format transformation, lacking in-depth confusion at the semantic level, attackers can restore the original content through context analysis and logical reasoning, and existing solutions cannot dynamically respond to attack behaviors, resulting in a decline in the hiding effect over time. Therefore, multi-level semantic fragmentation and dynamic adversarial disguise technologies are adopted in the hiding rules, and the security goals of "unreadable data, non-linkable fragments, and traceable self-destruction" are achieved through the collaborative design of five core units. The specific technical paths and advantages are as follows:

[0069] First, the semantic segmentation unit: fragmented semantics and polysemous mapping:

[0070] Since traditional holistic hiding is easily cracked by holistic semantic analysis (such as extracting core words through topic models), and the single semantic interpretation makes the fragments spliceable, attackers can restore the original content through fragment recombination (such as a jigsaw puzzle-style attack). Therefore, pre-trained models such as BERT are used to perform syntactic analysis, entity recognition, and logical relationship extraction on the atomic file unit, and it is segmented into 5 - 10 semantic fragments according to semantic boundaries (such as subject-predicate-object structure, image object contour, audio semantic frame). At the same time, at least two alternative semantic interpretations are generated for each fragment (e.g., converting "sales increased by 20%" into "market share increased" or "cost optimization effect"). Based on semantic networks such as WordNet, it is ensured that the alternative interpretations have a reasonable logical relationship with the original semantics but significant feature differences. By disassembling the complete semantics into "unindependently understandable" fragments, the polysemous interpretations of each fragment form a "semantic fog", and attackers cannot determine its true meaning even if they obtain some fragments. The cracking difficulty increases exponentially with the number of fragments (e.g., the combined interpretations of 10 fragments reach 2^10 = 1024).

[0071] II. Adversarial Generative Expansion Unit: Generation of Realistic Interference Content:

[0072] Since traditional expansion data generation relies on fixed templates or rules, there are pattern differences between the generated content and real scenarios (such as repeated words and logical discontinuities), which are easily recognized by adversarial sample detection algorithms, leading to the failure of hiding. Therefore, using the content in the resource library as real samples, a "generator-discriminator" adversarial system is constructed. The generator learns the language style, image texture, or audio waveform features related to the theme, and the discriminator distinguishes between real content and generated content. Each semantic fragment is attached with a theme label (such as "financial analysis", "market trend") and sentiment tendency (positive / neutral / negative) to control the generated extended paragraphs to strictly conform to the content distribution characteristics of the resource library (such as the density of professional terms and the complexity of sentence patterns), so as to generate extended paragraphs that are indistinguishable from real content in terms of "vision / semantics / logic", forming an adversarial interference layer that "looks like the real thing", and fundamentally eliminating the "machine-generated traces" of traditional expansion data.

[0073] III. Context Reorganization Unit: Discontinuous Logic Disruption and Dynamic Switching:

[0074] Even if there is interference content, if the expansion data maintains continuous logic (such as timeline, causal relationship), attackers can still reverse-infer the original semantics through the narrative structure (such as arranging fragments in chronological order to restore the whole picture of the event). Therefore, the logic chain break technology is adopted, that is, three types of discontinuous logics are introduced when inserting extended paragraphs: timestamp jump technology, such as mixing the description of an event that occurred in "2023" with the prediction content in "2025"; role replacement technology, swapping the main characters in the text, such as rewriting "Company A acquires Company B" as "Company B invests in Company A" while maintaining grammatical correctness; irrelevant logic chain injection technology, inserting paragraphs that are not directly related to the theme but conform to language logic (such as inserting a section on industry policy interpretation in a technical document). Then, by deploying a full-text semantic discriminator based on Transformer, the "logical coherence score" of the expansion data is calculated in real time. When the score exceeds the threshold (such as >0.7), it automatically switches to the backup extended paragraph to ensure that the logical coherence is always maintained at a state of "understandable but incomplete" (score 0.3 - 0.6). Thus, a "chaotic but self-consistent" semantic space is constructed, making the expansion data exhibit the characteristics of "jigsaw narrative" - the local paragraph logic is reasonable, but the whole cannot form a complete semantic chain. Compared with traditional solutions, attackers need to consume more computing resources for logical reconstruction, and the restoration accuracy is low.

[0075] IV. Fragment Mapping Key Management Unit: Precise Restoration and Permission Anchoring:

[0076] In the absence of records of the mapping relationship between fragments and extended data, legitimate users will be unable to decrypt and restore the original content, and traditional key management cannot support "permission allocation at the fragment granularity" (such as allowing viewing of the extended data corresponding to some fragments); therefore, the key slice generation technology is adopted to generate a unique 128-bit mapping key slice for each semantic fragment, including the following information: alternative semantic version number (such as V1.0, V2.3); position offset of the fragment in the extended data (accurate to characters / pixels / sampling points, that is, determining the position of the fragment-generated extended segment with different alternative semantics in the extended data); permission association label (such as "visible only to administrators", "modifiable by editors"); dynamic key combination: according to the arrangement order of the fragments in the extended data, the key slices are combined into a mapping key (such as generating a 1280-bit key for 10 fragments), and the final storage location key is generated through exclusive OR operation with the storage location key, forming a triangular association of "fragment-level permission - key - location", so that legitimate users can accurately restore the original content with the complete mapping key;

[0077] V. Dynamic self-destruction unit: Abnormal response and local fusing:

[0078] Since traditional security mechanisms usually adopt the "global locking" strategy when detecting abnormal access, resulting in restricted synchronization for legitimate users and being unable to prevent the spread of leaked fragments, lacking the ability of "precision strike"; therefore, the abnormal detection linkage technology is adopted to interface with the access log analysis module of the cloud server to monitor the following abnormal behaviors in real time: high-frequency key requests within a short period of time (>100 times / second); cross-permission-level fragment access (such as a visitor attempting to obtain an administrator-level key slice); abnormal geographical location access (such as requests initiated from an unauthorized area). Then, a hierarchical self-destruction strategy is adopted, that is, mild abnormality: freeze the mapping key slice of the corresponding fragment for a preset time (such as 10 minutes), during which only basic metadata can be read; severe abnormality: send a fusing instruction through the hardware security module to erase the key slice and position offset record of the fragment within a preset time (such as 50 microseconds), making the corresponding extended data permanently irrecoverable. At the same time, an immutable self-destruction log (including timestamp, triggering reason, affected fragment ID) is generated. By constructing an "attack immune" system, the abnormal response time is shortened from the minute level of the traditional solution to the millisecond level, and the attack impact range is controlled within a single fragment (<5% of the extended data). At the same time, the self-destruction operation does not affect the normal use of other fragments, ensuring that "local fusing" does not cause system-level failures and conforming to the principle of least privilege of "zero trust security".

[0079] Considering that if the numbered identifier is retained in the storage location key for a long time, an attacker can infer metadata such as the file segmentation scale and the location of sensitive data by analyzing the numbering rules (such as the increment step and the number range) (for example, numbers 001 - 1000 may correspond to a 1GB file), forming a security vulnerability of "reverse engineering the original file structure through numbers". Therefore, after sorting is completed, a cryptographic hash function (such as SHA-256) is used to perform a digest process on the numbered identifier to generate a hash value of a fixed length (such as 32 bytes), which replaces the original number and is embedded in the key metadata to ensure that the hash value cannot be reversely restored to the number. Also, a "temporary number - hash" mapping table is maintained in the cloud server memory, and this table is immediately destroyed after the key sorting is completed and synchronized to the key management module, without storing any number information on disk; at the same time, a CRC-64 checksum is added to the key after removing the number, and the calculation of the checksum includes the key value and the original number hash value to ensure that the key has not been tampered with during subsequent use; after removing the number, the attacker cannot infer the file structure through the key sequence, reducing the risk of metadata leakage; and removing the number reduces the size of the key metadata, reducing the storage overhead and network transmission load of the cloud server.

[0080] In a distributed storage scenario, the sub-file key sequence (generated by the user side) and the storage location key sequence (issued by the cloud server) may have a sequence deviation due to reasons such as network latency and node failures, resulting in incorrect combined key generation, leading to data decryption failures or permission verification errors. Therefore, based on the unique incrementing numbers of the sub-file units generated by the user side (such as 001, 002), a two-way mapping table of "number - sub-file key" and "number - storage location key" is established to achieve linear alignment of the two sequences through the number primary key; at the same time, for keys that time out due to abnormal delays (such as the storage location key of number 005 arriving 500ms late), a temporary combined key is pre-generated using the cached sub-file key and automatically updated after the complete sequence is received to ensure business continuity; and before matching, the SHA-512 hash values of each key in the two sequences are compared. If the hash values are found to be inconsistent (indicating that the key has been tampered with), the key retransmission mechanism is immediately triggered and the attack log is recorded, thus achieving "physical distribution - logical unity" key coordination, reducing the error rate of combined key generation, eliminating system-level failures caused by sequence errors, and ensuring the stability of real-time data processing scenarios (such as online transaction encryption);

[0081] Considering that if the single-key encryption mode is cracked, everything will be lost, and simply concatenating keys (such as string concatenation) may introduce redundant features, increasing the risk of key analysis and not supporting fine-grained permission control (such as assigning different permissions by sub-file units). Therefore, an appropriate concatenation method is selected according to the content type of the sub-file. For example, in the symmetric encryption scenario (such as AES): The key derivation function (KDF) is used to mix the sub-file key and the storage location key to generate a new key, and the entropy value is increased to more than 256 bits; in the asymmetric encryption scenario (such as RSA): The Chinese Remainder Theorem (CRT) is used to decompose the dual keys into private key fragments, and the complete private key is restored after combination to achieve "fragmented storage and collaborative decryption".

[0082] Considering that traditional permission descriptions (such as "edit" and "visitor") are qualitative labels and lack precise permission boundary definitions (such as exactly which content can be accessed and how much can be modified for "edit"), resulting in permission allocation relying on manual experience and prone to management loopholes such as "over-permission" or "under-permission". Therefore, three basic roles of administrator (100%), editor (A%), and visitor (B%) are defined, supporting custom extensions (such as "reviewer" and "temporary user"). They are divided into public (30%), internal (70%), and confidential (100%) according to content sensitivity, and cross with the role axis to form a permission matrix; set temporary permissions (such as B% within 72 hours) and long-term permissions (permanent A%), and dynamically adjust the effective permission value through the product of the timestamp and the ratio parameter; moreover, the permission level is quantified by percentage, stipulating that the administrator permission K = 100% (full access), the editor permission K = A% (adjustable from 50% to 80%), and the visitor permission K = B% (adjustable from 10% to 30%). The parameters can be visually configured through the management interface. By converting the fuzzy permission description into a computable mathematical parameter, the permission allocation accuracy is improved from the "file level" to the "key fragment level", and the administrator can precisely control the modification range of the edit through the slider (such as "allow editing 75% of the document paragraphs"), avoiding the "feeling-based" problem of permission allocation in the traditional solution;

[0083] Considering that directly cutting the complete key in proportion may damage the mathematical integrity of the key (for example, the RSA private key cannot be correctly decrypted after cutting), and there may be a security risk that low-privilege users may restore the complete key through a combination attack after obtaining some key fragments. Therefore, the Shamir secret sharing algorithm is adopted. The combined key sequence is regarded as "secret S" and split into N key fragments by using the Lagrange interpolation method. It is stipulated that at least K = N × P% fragments (P is the privilege ratio) need to be obtained to reconstruct the secret. For example: the administrator needs N × 100% = N fragments (complete reconstruction); the editor needs N × 70% fragments (70% fragments can be partially decrypted); the visitor needs N × 30% fragments (only the summary information can be decrypted). Moreover, higher weights are assigned to the key fragments corresponding to the core data (for example, the weight of the financial data fragment is 20% per piece, and the weight of the ordinary metadata fragment is 5% per piece). When the editor obtains 70% of the weighted fragments, it actually corresponds to 60% of the core content + 10% of the non-sensitive content, realizing "prior protection of important content". By constructing a two-factor control system of "privilege ratio = number of key fragments / weight", the fragment combination below the privilege ratio cannot restore the key through the Lagrange interpolation method, improving the anti-combination attack ability. At the same time, it supports dynamic adjustments such as "the administrator temporarily grants the editor 80% privilege" and "the visitor obtains 40% of the fragments within a limited time". The privilege change delay < 50ms, and there is no need to re-encrypt the original data.

[0084] Considering that the traditional access key only contains the privilege level identifier (such as "editor") and is not associated with the specific key fragment range and valid time, resulting in the possibility that low-privilege users may access high-privilege content through vulnerabilities and the inability to trace the usage boundary of the key; through the key fragment screening algorithm technology, according to the privilege ratio K = N × P%, the top P% of the high-weight fragments are screened from the combined key sequence (or non-core fragments are randomly selected according to the scenario configuration). For example, when the visitor privilege B% = 30%, 70% of the core fragments marked as "confidential" are automatically excluded. Moreover, through the privilege label encryption and embedding technology, three tags are added to the generated access key: ratio tag: clearly mark the privilege ratio in plain text (such as "30% Access") for quick verification by the client; time tag: bind the validity period stamp and the key fragment hash through the HMAC algorithm, and the tag automatically becomes invalid after expiration; content tag: record the range of sub-file unit numbers that can be accessed by this key (such as "only allowing decryption of fragments numbered 001 - 010"), preventing unauthorized access, so as to achieve access control of "visible privilege, clear boundary, and controllable timeliness".

[0085] The basic principles, main features and advantages of the present invention have been shown and described above. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.

Claims

1. Cloud storage data encryption and permission management platform, characterized in that, Comprising: A client for receiving multiple files uploaded by a user, splitting the content of each file through a splitting module to generate multiple sub-file units arranged in the original order of each file, and encrypting each sub-file unit to generate a corresponding sub-file key sequence, which is stored in the hardware security module of the client; A transmission channel module connected to the client for transmitting the sub-file units one by one to the distributed storage nodes; A distributed processing module deployed in the distributed storage nodes for performing semantic expansion processing on each sub-file unit to generate extended data and generating a storage location key according to the storage location of the extended data; A cloud server connected to the transmission channel module for receiving and managing the storage location keys and forming a storage location key sequence in order; A key management module deployed in the client for corresponding the sub-file key sequence with the storage location key sequence issued by the cloud server to generate a combined key, and dynamically allocating the proportion of the combined key according to a preset permission level to generate a differentiated access key corresponding to the permission level.

2. The cloud storage data encryption and permission management platform according to claim 1, characterized in that: The splitting module includes: A format parsing and preprocessing unit for obtaining multiple files uploaded by a user, parsing the format of the content of the multiple files, and at the same time, defining the part carrying information in the multiple files as the entity part area and the rest as the non-entity part area; A splitting position determination unit connected to the format parsing and preprocessing unit for determining the splitting quantity according to the size of each file, initially determining the splitting position from the splitting quantity, detecting each splitting position, judging whether the splitting position exists in the entity part area, and when it exists in the entity part area, finding the nearest non-entity part area to the splitting position as the new splitting position; A splitting unit connected to the splitting position determination unit for obtaining the splitting position and splitting the file.

3. The cloud storage data encryption and permission management platform according to claim 1, characterized in that: The hardware security module includes an encryption chip and a secure storage circuit for generating and storing the sub-file key sequence. When detecting illegal disassembly, abnormal voltage or abnormal temperature, the secure storage circuit triggers an automatic destruction mechanism to destroy the stored sub-file key sequence.

4. The cloud storage data encryption and permission management platform according to claim 1, characterized in that, The transmission channel module dynamically starts the same number of distributed storage nodes according to the number of sub-file units generated by splitting each file at the client; each distributed storage node corresponds to one sub-file unit one by one.

5. The cloud storage data encryption and permission management platform according to claim 1, characterized in that: The client further includes an identification module for generating a unique increasing number identification for the sub-file units in the original order of the files during the transmission of the sub-file units.

6. The cloud storage data encryption and permission management platform according to claim 5, characterized in that, The distributed processing module includes a resource library and a storage library, where: The resource library has a preset theme and crawls publicly available network data, domain knowledge bases or historical storage data around the preset theme to form a content collection related to the theme; when the sub-file unit is transmitted to the distributed processing module, the content of the sub-file unit is identified, semantic transformation is performed through a hiding rule, and based on the content collection of the resource library, semantic expansion is performed on the semantically transformed content through a generative AI model to generate extended data related to the preset theme; A repository, connected to a resource library, stores extended data and generates a corresponding file storage location key through a location-sensitive hashing algorithm according to the storage location of the extended data. After the file storage location key is generated, the same number identifier as the sub-file unit is added.

7. The cloud storage data encryption and permission management platform according to claim 6, characterized in that: The hiding rules include: A semantic segmentation unit, which is used to semantically segment the content of the sub-file unit into multiple fragments and generate at least two alternative semantic interpretations for each fragment; An adversarial generative extension unit, connected to the semantic segmentation unit, generates extended paragraphs containing different alternative semantics for each fragment based on the content set of the resource library through a generative adversarial network; A context recombination unit, connected to the adversarial generative extension unit, inserts the extended paragraphs into the base text according to discontinuous logic, where the discontinuous logic includes timestamp jumps, role permutations, and injection of irrelevant logical chains, and dynamically switches the extended paragraphs according to the real-time feedback of the full-text semantic discriminator to form extended data; the base text is generated by the adversarial generative extension unit; A fragment mapping key management unit, connected to the context recombination unit, generates a unique mapping key slice for each fragment, records its alternative semantic version and the position offset of the fragment in the extended data, and combines multiple mapping key slices in the order in the extended data to form a mapping key, where the position offset of the fragment in the extended data is the position of the extended paragraph generated by the fragment in the extended data; A dynamic self-destruction unit, which responds to the abnormal access warning of the cloud server module and implements a hierarchical self-destruction strategy.

8. The cloud storage data encryption and permission management platform according to claim 6, characterized in that: The cloud server sorts the storage location keys according to the number identifier, and removes the number identifier after sorting.

9. The cloud storage data encryption and permission management platform according to claim 1, characterized in that: The key management module is communicatively connected to the hardware security module; the sub-file key sequence and the storage location key sequence are matched and corresponding according to the arrangement order, and the encrypted algorithm is used to splice the matched sub-file key and storage location key to generate a combined key sequence consistent with the order of the sub-file units.

10. The cloud storage data encryption and permission management platform according to claim 1, wherein: The preset permission level is the user role level, and the user role level includes administrator, editor, and visitor; Denote the total number of key fragments in the combined key sequence as N, the preset permission level as K, the administrator's permission is K = N×100%, the editor's permission is K = N×A%, and the visitor's permission is K = N×B%, where 100% > A% > B% > 0.

Citation Information

Cited By

  • Data synchronization method and device for cloud document management system

    CN121478738A

  • Safe starting system and method based on dynamic environment binding and heterogeneous inspection

    CN121658088A