A cloud computing-based big data distributed storage method and system

By combining manifold space modeling and quantum dynamic sharding with Lie group cryptographic scheduling and cross-modal cooperative transmission, the problems of data distribution mismatch and mixed data transmission conflicts in quantum storage systems are solved, realizing efficient, secure and scalable data management in quantum storage systems.

CN120499204BActive Publication Date: 2026-04-10BEIJING NANSHAN TONGXING TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING NANSHAN TONGXING TECHNOLOGY CO LTD
Filing Date
2025-06-17
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Traditional storage architectures struggle to capture the inherent manifold structure of high-dimensional unstructured quantum data, leading to a mismatch between storage resources and the spatial geometry of data distribution. Furthermore, traditional scheduling algorithms lack orthogonal group space constraints on quantum storage nodes, resulting in dimensional collapse and orthogonality destruction during resource mapping, which affects system scalability and security.

Method used

An integrated approach combining manifold space modeling, quantum dynamic sharding, Lie group cryptographic scheduling, and cross-modal cooperative transmission, including manifold embedding, quantum Ising models, Lie group optimization, and cross-modal bus protocols, is used to construct a dynamic data organization architecture that enables adaptive matching and secure transmission of data sharding and resource mapping.

Benefits of technology

It significantly improves the quantum storage system's ability to process complex data patterns, ensures data security and system scalability, solves the problems of storage resource and data distribution mismatch and mixed data transmission conflicts in the quantum storage environment, and forms an intelligent storage infrastructure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120499204B_ABST
    Figure CN120499204B_ABST
Patent Text Reader

Abstract

The application relates to the field of quantum storage and cloud computing fusion, and discloses a big data distributed storage method and system based on cloud computing. First, the data intrinsic geometric structure is revealed through manifold embedding, and manifold coordinates are generated; second, an Ising physical model is constructed based on the geometric structure, and the model ground state is solved through quantum annealing to determine an optimal data dynamic fragmentation scheme; then, the data fragments are homomorphically encrypted, and a Lie group optimization algorithm is used to solve on a special orthogonal group manifold to generate a resource mapping matrix that takes into account safety and efficiency; finally, the fragmentation scheme is instructed, transmission tensor loads of fusion manifold information, data content and security mapping are constructed for each fragment, and transmission configuration is optimized. The application combines multiple frontier technologies, realizes data structure-aware dynamic fragmentation and safe and efficient resource mapping, and significantly improves the overall performance and safety of the system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of quantum storage and cloud computing fusion, in particular to a big data distributed storage method and system based on cloud computing. BACKGROUND

[0002] With the rapid development of quantum computing technology, quantum storage systems face the dual challenges of explosive growth of data dimensions and large-scale access of heterogeneous devices. When dealing with high-dimensional unstructured data, traditional storage architectures usually use Euclidean space clustering algorithms for resource allocation, but such methods are difficult to capture the intrinsic manifold structure characteristics of data, resulting in mismatch between storage resources and spatial geometric characteristics of data distribution, causing storage fragmentation and access delay accumulation problems. Especially in the quantum computing environment, the entanglement characteristics between quantum bits cause the data correlation dimension to grow exponentially, and the traditional sharding strategy based on scalar distance measurement cannot meet the storage needs of quantum state data.

[0003] Existing encryption storage schemes mostly use homomorphic encryption combined with static resource mapping, which can ensure data privacy, but the fixed topology of resource allocation mode is difficult to adapt to dynamic changes in storage load. In particular, in the distributed quantum storage scenario, the quantum characteristics of storage nodes are fundamentally different from classical hardware, and traditional scheduling algorithms lack mathematical modeling of orthogonal group space constraints, resulting in dimension collapse and destruction of orthogonality during resource mapping, which seriously restricts the scalability and security of the storage system.

[0004] Current cross-modal data transmission mechanisms generally use a combination of time division multiplexing and priority queue arbitration strategies, but they cannot effectively solve the space-time conflict problem between quantum instruction streams and classical data packets. Real-time control instructions generated during quantum computing processes have strong timeliness and correlation, while the transmission model of classical storage systems based on statistical multiplexing lacks the ability to model the space-time characteristics of quantum states, resulting in low bus bandwidth utilization and increased transmission jitter, which is a key bottleneck restricting the performance improvement of quantum storage systems.

[0005] Therefore, the present application provides a big data distributed storage method and system based on cloud computing to solve the deficiencies of the prior art. SUMMARY

[0006] In view of the deficiencies of the prior art, the present application provides a big data distributed storage method and system based on cloud computing, which effectively solves the three core problems of high-dimensional data storage mismatch, dynamic resource scheduling imbalance and mixed data transmission conflict through the integration of manifold space modeling, quantum dynamic sharding, Lie group encryption scheduling and cross-modal collaborative transmission.

[0007] To achieve the above object, the present application is implemented by the following technical solutions: a big data distributed storage method and system based on cloud computing, comprising the following steps:

[0008] Step S1: manifold embedding is performed on the original data blocks to generate a manifold space coordinate matrix;

[0009] Step S2: a quantum Ising model is constructed based on the geometric features of the manifold space coordinate matrix, spin configuration is solved through a quantum annealing algorithm, and data shards are dynamically divided;

[0010] Step S3: based on the results of the data shards, storage resources are allocated in the orthogonal group space through a Lie group optimization algorithm to generate a resource mapping matrix;

[0011] Step S4: according to the resource mapping matrix, data shards and quantum instructions are cooperatively transmitted to target storage nodes through a cross-modal bus protocol.

[0012] Preferably, the step S1 comprises

[0013] The step S1 comprises:

[0014] S1.1: for each of the data blocks , its -dimensional feature vector is extracted, wherein , wherein represents the th predefined feature extraction function used to capture the key attributes of the original data block;

[0015] S1.2: a similarity matrix is constructed based on the extracted feature vectors of each data block, and the elements of the similarity matrix are calculated by the formula , wherein is the square of the Euclidean distance between the feature vectors and , used to quantify the similarity between the data blocks, is the total number of data blocks, is the kernel width parameter, adjusting the rate of similarity decay with distance;

[0016] S1.3: a diagonal matrix is calculated, wherein the elements on the diagonal of the diagonal matrix represent the total similarity of the data blocks , and all non-diagonal elements of the diagonal matrix are zero, and the similarity matrix and the diagonal matrix​ constructing a transition matrix , the transition matrix is the probability transition of random walk between data blocks;

[0017] S1.4: performing eigen decomposition on the transition matrix , obtaining its first largest eigenvalues , and the eigenvector set corresponding to the largest eigenvalues , where the largest eigenvalues reflect the importance of the data manifold in the corresponding direction, is the th eigenvector corresponding to the eigenvalue , which represents a principal direction of the data in the manifold space;

[0018] Based on the largest eigenvalues and the components in the eigenvector set, the low-dimensional coordinates of each data block in the low-dimensional manifold space are constructed , the low-dimensional coordinates of each data block in the low-dimensional manifold space are constructed , the low-dimensional coordinates of each data block in the low-dimensional manifold space are constructed , the low-dimensional coordinates of each data block in the low-dimensional manifold space are constructed , the low-dimensional coordinates of each data block in the low-dimensional manifold space are constructed

[0019] Preferably, the value of the kernel width parameter used in step S1.2 for calculating the similarity matrix is determined by calculating the average value of the square of the Euclidean distance between the eigenvectors of all data block pairs, and the specific calculation formula is as follows:

[0020] ;

[0021] wherein is the total number of data blocks in step S1.1, and are the eigenvectors extracted in step S1.1 of the th and the th data block, denotes the square of the Euclidean distance between the eigenvectors and .

[0022] Preferably, to realize dynamic quantum fragmentation based on the geometric characteristics of the data manifold, the step S2 specifically comprises:

[0023] S2.1: obtaining a manifold space coordinate matrix based on the step S1 , calculating a Riemannian metric tensor describing local intrinsic geometric properties of the data manifold , wherein, and are indices of a local coordinate system of the manifold;

[0024] S2.2: deriving coupling coefficients in the Ising model for any pair of data blocks and the manifold space coordinate matrix , based on the Riemannian metric tensor and , the coupling coefficients quantitatively describe geometric correlation strength or interaction strength of data blocks and on the data manifold;

[0025] S2.3: constructing a Hamiltonian of the target Ising model, wherein, is a spin variable assigned to data block , the value of indicates the shard to which the data block belongs, is an external local field parameter applied to spin ;

[0026] optimizing and solving the Hamiltonian to obtain a set of optimal spin configurations , the optimal spin configurations determine the dynamic shard results of all data blocks.

[0027] Preferably, in step S2.3, the step of optimizing and solving the Hamiltonian includes a step of solving the optimal spin configuration of the Ising model Hamiltonian by using quantum annealing technology, specifically:

[0028] a) initializing the system as a system containing quantum bits, wherein each quantum bit corresponds to a spin variable of a data block, and the initial quantum state of all possible computational ground states is a uniform superposition: The initial state is the ground state of the initial Hamiltonian ;

[0029] b) the state of the quantum system evolves over time in accordance with the instantaneous Schrödinger equation:​​ where h is the Planck constant is taken to be 1;

[0030] c) the instantaneous Hamiltonian from the initial Hamiltonian slowly, continuously adiabatically evolves to a final target Hamiltonian ;

[0031] the final target Hamiltonian is the Ising model Hamiltonian of the optimal data sharding scheme after ground state encoding , so that the system at the end of the evolution is in the ground state of the final target Hamiltonian with high probability, and the optimal spin configuration is obtained by measuring the ground state of the final target Hamiltonian .

[0032] Preferably, to achieve secure, optimized and structure-preserving mapping of encrypted data shards to storage resources, the step S3 comprises:

[0033] S3.1: constructing an aggregated algebraic space for describing the transformation characteristics of the storage resources, where is the Lie algebra of the special orthogonal group composed of dimensional real anti-symmetric matrices, is the number of aggregated Lie algebra components;

[0034] S3.2: applying a fully homomorphic encryption algorithm to the data shards obtained from step S2 to encrypt them, obtaining ciphertext data shards that can be calculated in an encrypted state, where the encryption process can be represented as Enc , and (mod , where is the plaintext unit, is a public matrix, s is a private key, is an error term, is a modulus;

[0035] S3.3: constructing a cost matrix in the encrypted domain, where is the data shard weight, is the encrypted data shard, is the resource request vector, and obtaining an optimal special orthogonal resource mapping matrix by solving an optimization problem Tr The Special orthogonal resource mapping matrix Used to align encrypted data fragments with storage resources;

[0036] S3.4: Based on the optimal n×n special orthogonal resource mapping matrix U obtained from the optimization solution, generate the final... Resource Mapping Matrix The final Resource Mapping Matrix Special orthogonal group The elements satisfy the orthogonal constraint. , It is an identity matrix and can be solved by Lie algebra. To Liqun The exponential mapping is represented as ,in, for antisymmetric basis matrix, To optimize the determined parameters.

[0037] Preferably, the resource mapping matrix in step S3.4 Through a set of scalar parameters and scalar parameters Corresponding, predefined Lie algebras of Real antisymmetric matrix During parameterization generation, step S3.3 is used to solve the optimization problem. The rules for iteratively updating the resource mapping matrix include the following operations:

[0038] In the In the next iteration, if the parameters To make adjustments, the resource mapping matrix is ​​changed from its current state using the resource state mapping formula. Update to the next state ,in, For parameter index, The resource state mapping formula is:

[0039] ;

[0040] in,

[0041] For the first At the start of the next iteration Special orthogonal resource mapping matrix;

[0042] This is the updated version after this parameter adjustment. Special orthogonal resource mapping matrix;

[0043] (eta) is a positive scalar learning rate, controlling the iteration step size;

[0044] is a matrix exponential function;

[0045] is a real anti-symmetric matrix belonging to the Lie algebra constructed according to the gradient information of the parameter , and the specific calculation steps are as follows: First, the cost function is calculated when the current resource mapping matrix is

[0046] , and the partial derivative value of the th scalar parameter is recorded as the scalar gradient value , , wherein is the cost matrix; Then, the real anti-symmetric matrix is constructed using the scalar gradient value

[0047] and the Lie algebra base matrix corresponding to the scalar parameter , . .

[0048] Preferably, for constructing a cross-modal transmission channel based on a tensor manifold encapsulation, the step S4 further comprises:

[0049] S4.1: encoding the optimal spin configuration obtained by quantum annealing in the step S2.3 into a quantum instruction sequence , which is aggregated by quantum labels corresponding to each data block and is denoted as , wherein is the total number of data blocks, is the final spin state of the th data block, denotes an aggregation operation, and the data structure of each quantum label is defined as: ;

[0050] ;

[0051] wherein

[0052] is an 8-bit opcode field used to indicate the type of the label or subsequent operation;​​​​ a 16-bit qubit address field for uniquely identifying a qubit corresponding to the i-th data block;

[0053] a 32-bit field for storing or representing the final spin state of the i-th data block;

[0054] S4.2: For each data block index set determined by the optimal spin configuration, a tensorized vector representation of the data slice is first constructed, and then the manifold coordinates of all data blocks within the data slice are accumulated with the tensor product of the data blocks themselves by a tensor accumulation model, the formula of which is:

[0055]

[0056] wherein,

[0057] is the manifold coordinate vector of the i-th data block;

[0058] is the i-th original data block or its predetermined representation;

[0059] denotes the tensor product operation;

[0060] Then, the tensorized vector representation is transformed and encapsulated by using the transpose of the resource mapping matrix combined with the manifold basis space tensor generated by the manifold feature vector, to construct a transmission tensor load wherein, denotes the multiplication operation of matrix and vector or the generalized linear transformation.

[0061] Preferably, a cloud computing-based big data distributed storage system comprises:

[0062] a data manifold feature extraction module for extracting a feature vector from a data block, calculating a similarity matrix between data blocks, and performing feature decomposition on the similarity matrix to obtain manifold coordinates and a manifold feature vector of a data manifold;

[0063] ​​​​​​​​​​​​​​​A dynamic quantum fragmentation module is configured to construct an Ising physical model according to geometric characteristics of the data manifold, and to solve the ground state of the model by quantum annealing technology to determine an optimal fragmentation scheme for allocating data blocks to different data fragments.

[0064] A secure resource mapping module is configured to homomorphically encrypt the data fragments, and to construct and solve an optimization problem defined on a special orthogonal group manifold to generate an optimal special orthogonal mapping matrix for securely mapping the encrypted data fragments to storage resources.

[0065] A transmission packaging module is configured to encode the optimal fragmentation scheme into an instruction sequence, and to construct a transmission tensor load for each data fragment, which encapsulates manifold geometric information, data content, and is transformed by the optimal special orthogonal mapping matrix.

[0066] The application provides a cloud computing-based big data distributed storage method and system.

[0067] 1. The application constructs a dynamic data organization architecture for quantum storage optimization through deep coupling of manifold embedding and quantum fragmentation.

[0068] 2. The application innovatively designs a Lie group space encryption scheduling mechanism to achieve secure and efficient dynamic allocation of storage resources under orthogonal group constraints.

[0069] 3. The application proposes a cross-modal spatiotemporal collaborative transmission model to achieve efficient transmission of quantum-classical hybrid data through tensor manifold packaging and variational temporal arbitration. BRIEF DESCRIPTION OF DRAWINGS

[0070] Figure 1 The flowchart of the cloud computing-based big data distributed storage method;

[0071] Figure 2This is a diagram of a cloud-based big data distributed storage system architecture. Detailed Implementation

[0072] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0073] Please see the appendix Figure 1 This invention provides a cloud computing-based distributed storage method for big data. The steps of the method are described in detail below.

[0074] Step S1: Perform manifold embedding on the original data block to generate a manifold space coordinate matrix.

[0075] In the big data distributed storage method based on cloud computing provided in this embodiment, the core task of the first step (S1) is to reveal the underlying nonlinear intrinsic geometric structure within the original data block set and express this structure in a numerical form. The final output of this step is a manifold spatial coordinate matrix, which provides a key geometric basis for subsequent dynamic quantum partitioning.

[0076] Specifically, for a data set containing multiple original data blocks, the present invention first targets each data block. Using a pre-defined feature extraction algorithm, it is transformed into a multi-dimensional feature vector that can describe its core characteristics. These eigenvectors constitute a numerical representation of the original dataset and form the basis for subsequent geometric analysis.

[0077] After obtaining the feature vectors of all data blocks, this invention further constructs a similarity matrix. This is used to quantify the local similarity between any two data blocks in the feature space. Similarity matrix Each element The calculation formula is derived from the Gaussian kernel function:

[0078] ;

[0079] in, Indicates the first The and the first Similarity between data blocks; and The first The and the first Feature vectors of each data block; This represents the square of the Euclidean distance between the two eigenvectors; is the kernel width parameter of the Gaussian kernel function, which controls the rate at which similarity decays with distance.

[0080] To ensure that the similarity matrix accurately reflects the inherent scale of the dataset and avoids distortion of manifold structure information due to inappropriate parameter selection, the kernel width parameter... The value of is specifically defined in this invention. Its value is determined by calculating the average of the squared Euclidean distances between the feature vectors of all data block pairs, using the following formula:

[0081] ;

[0082] in, This represents the total number of data blocks. This method provides a deterministic and adaptive calculation basis for selecting the kernel width parameter.

[0083] To perform subsequent spectral analysis to reveal the manifold structure, it is necessary to use the calculated similarity matrix. To construct a graph Laplacian matrix First, construct a diagonal matrix, i.e., the metric matrix. The elements on its diagonal Defined as a similarity matrix No. The sum of all elements in the row, i.e. Subsequently, the Thulaplace matrix It can then be defined as .

[0084] Thulaplatz matrix The spectrum (i.e., its eigenvalues ​​and eigenvectors) contains the intrinsic geometric information of the data manifold.

[0085] This step extracts this information by solving the following generalized eigenvalue problem: in, For eigenvalues, This corresponds to its eigenvector. Solving this problem will yield a set of eigenvalue-eigenvector pairs. And sort them according to the size of the eigenvalues, that is

[0086] Finally, to achieve the embedding of data from the high-dimensional original feature space to the low-dimensional intrinsic manifold space, this invention discards the method corresponding to the minimum eigenvalue. The trivial solution (whose corresponding eigenvector is a constant vector) is obtained, and the subsequent solution is selected. The smallest non-zero eigenvalues The corresponding eigenvector This Combining the feature vectors column by column forms a To conduct subsequent spectral analysis to reveal the manifold structure, it is necessary to use the calculated similarity matrix... To construct a graph Laplacian matrix .

[0087] First, construct a diagonal matrix, i.e., the metric matrix. The elements on its diagonal Defined as a similarity matrix No. The sum of all elements in the row, i.e. .

[0088] Subsequently, the TuLaplace matrix It can then be defined as 1-dimensional manifold space coordinate matrix .

[0089] The first of the matrix row vector , that is, the first Data blocks in the embedded The new coordinates in the 3D manifold space accurately reflect the geometric location of the data block within the entire dataset.

[0090] At this point, step S1 is complete, and the resulting manifold space coordinate matrix is... This will serve as the basis for the next step of dynamic quantum fragmentation.

[0091] Step S2: Construct a quantum Ising model based on the geometric features of the manifold space coordinate matrix, solve the spin configuration using the quantum annealing algorithm, and dynamically divide the data into slices.

[0092] Specifically, the present invention first uses the manifold space coordinate matrix obtained in step S1. Zhang Xing, Calculate the metric of the data manifold This metric tensor It precisely describes how distances and angles are measured within an infinitesimal neighborhood of any point in a low-dimensional space defined by manifold coordinates, thus capturing the inherent local geometric properties of the dataset.

[0093] in, and An index for the local coordinate system of the manifold, its value ranges from 1 to the dimension of the manifold space. The calculation of this metric tensor can be obtained through numerical methods such as differencing or differentiating the manifold coordinates, and it is the fundamental basis for establishing geometric relationships between data blocks.

[0094] Obtaining the metric tensor Then, the present application further converts the geometric information into interaction parameters of the Ising physical model. For any two data blocks and a coupling coefficient is defined, which quantifies the interaction strength between the two data blocks in the Ising model.

[0095] The value of the coupling coefficient is directly determined by their geometric proximity on the data manifold. A specific implementation is which is defined as a function related to the geodesic distance between the data blocks and on the manifold.

[0096] When the geometric relationship between two data blocks on the manifold is closer, the corresponding coupling coefficient is set to a larger positive value, and vice versa, thus providing physical constraints for the subsequent energy minimization process.

[0097] Subsequently, the present application constructs an Ising model Hamiltonian that describes the total energy of the system, whose mathematical expression is

[0098] ;

[0099] where is a spin variable associated with each data block , whose value is , which directly determines the data shard to which the data block belongs, for example, indicates that the data block is divided into the first shard, and indicates that it is divided into the second shard; is the coupling coefficient defined above, which reflects the geometric relationship between the data blocks and ;

[0100] is an external local field parameter, which is a bias parameter that can be set according to specific business requirements or system constraints, used to exert preferential guidance on the division of specific data blocks. Finding a spin configuration that minimizes the above Hamiltonian is a combinatorial optimization problem. The present application uses quantum annealing technology to efficiently solve this problem. The process first maps the Ising model to a quantum system containing quantum bits, where the state of the 60th quantum bit corresponds to the spin variable .

[0101] In the initial stage of quantum annealing, the system is prepared in the ground state of an initial Hamiltonian , which is a uniform superposition of all possible classical spin configurations, expressed as:

[0102] ;

[0103] where represents a specific classical spin configuration (i.e., a complete data partitioning scheme), denotes the summation over all possible configurations, is the total number of qubits, i.e., the total number of data chunks. This initial state implies that at the beginning of the solution, all possible partitioning schemes are treated equally.

[0104] Then, the Hamiltonian that the system follows will slowly and continuously evolve adiabatically from the initial Hamiltonian to the final target Hamiltonian over time . Here is the quantum mechanical representation of the aforementioned Ising model that needs to be solved, which is specifically:

[0105] ;

[0106] where is the Pauli Z operator acting on the th qubit, whose eigenvalues are the classical spin values ±1. The evolution of the quantum state of the entire system follows the instantaneous Schrödinger equation , where is the reduced Planck constant.

[0107] According to the quantum adiabatic theorem, under the condition that the evolution is slow enough, the system will always remain in the ground state of the instantaneous Hamiltonian with a very high probability.

[0108] Therefore, at the end of the evolution, the quantum state of the system is the ground state of the target Hamiltonian . At this point, measuring all the qubits in the computational basis will yield a unique spin configuration .

[0109] This configuration is the Hamiltonian ​The global optimal solution or high-quality approximate solution is obtained, and finally the dynamic division scheme of all data blocks is determined.

[0110] Step S3: Based on the result of data fragmentation, the storage resources are allocated in the orthogonal group space by the Lie group optimization algorithm to generate a resource mapping matrix.

[0111] In the embodiment of the application, after the dynamic division of the data blocks is completed in step S2, step S3 is entered. This step aims to allocate specific storage resources to each data fragment obtained in a manner that takes into account data privacy security and resource utilization efficiency. This step converts the resource mapping problem into an optimization problem on a special orthogonal group manifold by introducing Lie group and Lie algebra theory, and solves to obtain a resource mapping matrix that can maintain the integrity of the data structure.

[0112] To accurately describe the transformation characteristics of the storage resources, the application first constructs an aggregate algebraic space The space is composed of multiple Lie algebra components through tensor product operation, and its mathematical expression is , wherein is the Lie algebra corresponding to the special orthogonal group , which is composed of all real anti-symmetric matrices of dimension ; is the specific number of Lie algebra components that constitute this aggregate space; represents the tensor product operation.

[0113] The establishment of this algebraic structure provides a rigorous mathematical framework for subsequent resource mapping. To protect the confidentiality of the data content in the resource allocation process, the application applies a fully homomorphic encryption algorithm to the data fragments obtained in step S2.

[0114] Through this encryption operation, each plaintext data fragment is converted into a ciphertext data fragment , which can directly perform subsequent mathematical operations without decryption. This encryption process can be abstractly represented as Enc , and one specific implementation is (mod .

[0115] In this expression, is the plaintext unit to be encrypted; is a public matrix in the encryption scheme; is a private key used in the encryption scheme; is a random error term introduced in the encryption process; is a large integer modulus relied on by the encryption operation.

[0116] With the data already encrypted, this invention constructs and solves an optimization problem within the encrypted domain. First, a cost matrix is ​​constructed. Its expression is In this definition, For the first Preset weighting coefficients for each data shard; For the first homomorphic encryption Data shards; For a representation of the first The class can store vectors of resource characteristics or specific resource requests. Subsequently, this invention proposes the following optimization objectives:

[0117] Tr in, It is a problem to be solved. A special orthogonal matrix representing a rotational transformation from data sharding to storage resource allocation. For all The group space formed by special orthogonal matrices of dimension Tr; This is the trace operation of a matrix.

[0118] The goal of this optimization problem is to find an optimal rotation transformation. This minimizes the cost of any association between encrypted data fragments and storage resources.

[0119] To solve the above in For optimization problems defined on manifolds, this invention employs an iterative optimization algorithm. In the algorithm's... In the next iteration, the current estimate of the optimal matrix... Updated according to the following update rules In this iterative formula, and The first At the start of the next iteration and after the update Special orthogonal matrices; It is a positive scalar learning rate used to control the step size of each iteration; Represents the matrix exponentiation function, which ensures that the updated matrix... It remains a special orthogonal matrix, thus maintaining the validity of the solution throughout the optimization process.

[0120] matrix It is A real antisymmetric matrix of dimension 1, i.e. It represents the current point. Cost function The optimal descent direction is represented in the Lie algebra space. This matrix... The value is based on the cost function. Regarding matrices At point The gradient is calculated at a given point, ensuring that the iteration proceeds along the geodesic direction on the manifold, approaching the optimal solution in the most efficient way. By repeatedly executing the above iterative update steps until the preset convergence condition is met, an optimal special orthogonal transformation matrix is ​​finally obtained. .

[0121] This matrix is ​​then determined as the final resource mapping matrix. This matrix As a special orthogonal group An element that naturally satisfies orthogonal constraints (in for (Identity matrix). Furthermore, this matrix... It can also be parameterized using the exponential mapping from Lie algebras to Lie groups, in the form of: ,in Lie algebra A set of basis matrices, and These are the scalar parameters ultimately determined through the optimization process described above. This process ensures that the generated resource mapping is both optimal and structure-preserving.

[0122] Step S4: Construct the transport encapsulation structure and optimize the transport configuration.

[0123] In this embodiment of the invention, after steps S1 to S3 complete the disclosure of the inherent geometric structure of the data, the optimal data fragmentation based on this structure, and the secure mapping for storage resources, the process proceeds to step S4. This step aims to finally encapsulate and integrate the results of the preceding steps, construct a structured data payload suitable for cross-modal transmission, and determine the optimal operating parameters for the transmission channel carrying this payload.

[0124] Specifically, this invention first solidifies and instructs the optimal data partitioning scheme determined by quantum annealing in step S2. The optimal partitioning scheme is embodied in a set of spin configurations. , each of them The value represents the first The final ownership of each data block. To facilitate subsequent processing and recording, this invention encodes this configuration into a quantum instruction sequence. .

[0125] This sequence is achieved by assigning a quantum tag to each data block. Formed by aggregation operations, represented as In this aggregation formula, The total number of data blocks. This represents a predefined aggregation operation, such as concatenating the data structures of individual quantum tags in sequence. Each quantum tag It is itself a unit with a specific data structure, defined as: opcode(8), qubit_addr(16), spin_val(32)], where opcode(8) is an 8-bit opcode field used to indicate the type of the tag or related subsequent operation instructions; qubit addr(16) is a 16-bit address field whose value uniquely identifies the address relative to the first address. The data blocks are associated with qubits during the quantum annealing process; while spinval(32) is a 32-bit data field used to store the qubits corresponding to the data blocks measured in step S2.3. spin states of data blocks .

[0126] While instructing the fragmentation scheme, this invention provides each data fragment determined by the optimal spin configuration. Construct a transport tensor payload that encapsulates its core information. First, for a specific data fragment... Construct its tensorized vector representation. This construction process is achieved by accumulating the tensor product of the manifold coordinates of all data blocks within the partition and their own data. The calculation formula is as follows:

[0127] ;

[0128] In this formula, the summation iterates through the partitions. All data block indexes ; For the first The manifold coordinate vectors of each data block; For the first A raw data block or its pre-extracted numerical representation; This represents the tensor product operation. This operation converts the geometric position information of the data block... With data content They are organically integrated into a unified tensor structure.

[0129] Subsequently, the finalized resource mapping matrix was used. and manifold eigenvectors For the above tensor vector representation The final transformation and encapsulation are performed to generate the transport tensor payload. The calculation formula for this process is:

[0130] ;

[0131] in, Resource mapping matrix The transpose of , which acts on It achieves secure and structure-preserving data-to-resource mapping within the encrypted domain; it represents matrix-vector multiplication or generalized linear transformation; For the obtained manifold eigenvectors The generated manifold basis space tensor provides a reference basis for transporting loads, describing the global geometric background.

[0132] Finally, the present invention provides a means to carry the transport tensor load constructed above. The bus library or transmission channel determines an optimal operating parameter. This parameter It is also quantified as an improvement level in data processing. Optimal operating parameters This is obtained by solving a specific maximization problem:

[0133] ;

[0134] in, These are the runtime parameters to be optimized, and T is the predefined valid range of values ​​for them. Objective function Defined as The potential energy function in this function The value is proportional to the aggregation strength of the transport tensor payloads of all data fragments and is affected by the data boosting level. The dynamic adjustment is specifically expressed as follows:

[0135] ;

[0136] In the definition of this potential energy function It is based on Dynamically adjusted predefined nonnegative scalar functions; This is the total number of data shards; Is assigned to the first Predefined positive real weights for each data shard; It corresponds to the first The transmission tensor payload of each data fragment; This represents the square norm of the transport tensor load. By finding the optimal... This invention configures an optimal working state that balances operational performance and load cost for the entire data transmission process.

[0137] Please see the appendix Figure 2 The present invention also provides a big data distributed storage system based on cloud computing, the details of which are as follows:

[0138] A cloud computing-based big data distributed storage system, comprising:

[0139] A data manifold feature extraction module, configured to extract feature vectors from data blocks, calculate a similarity matrix between data blocks, and perform feature decomposition on the similarity matrix to obtain manifold coordinates and manifold feature vectors of the data manifold;

[0140] A dynamic quantum fragmentation module, configured to construct an Ising physical model according to the geometric characteristics of the data manifold, and solve the ground state of the model by quantum annealing technology to determine an optimal fragmentation scheme for allocating data blocks to different data fragments;

[0141] A secure resource mapping module, configured to homomorphically encrypt the data fragments, and construct and solve an optimization problem defined on a special orthogonal group manifold to generate an optimal special orthogonal mapping matrix for securely mapping the encrypted data fragments to storage resources;

[0142] A transmission packaging module, configured to encode the optimal fragmentation scheme into an instruction sequence, and construct a transmission tensor load for each data fragment, which encapsulates manifold geometric information, data content, and is transformed by the optimal special orthogonal mapping matrix.

[0143] The system of the embodiment can be used to execute the method embodiments described above, and has similar principles and technical effects, which will not be described here again.

[0144] Although the embodiments of the present application have been shown and described, it can be understood by those skilled in the art that various changes, modifications, replacements and variations can be made to the embodiments without departing from the principles and spirits of the present application, and the scope of the present application is defined by the appended claims and their equivalents.

Claims

1. A cloud computing-based big data distributed storage method, characterized in that, The method comprises the following steps: Step S1: manifold embedding is performed on the original data block to generate a manifold space coordinate matrix; Step S2: constructing a quantum Ising model based on the geometric features of the manifold space coordinate matrix, and obtaining an optimal spin configuration by a quantum annealing algorithm and dynamically dividing data shards; Step S3: based on the results of the data slices, a storage resource is allocated in the orthogonal group space through a Lie group optimization algorithm to generate a resource mapping matrix; Step S4: according to the resource mapping matrix, the data slices and quantum instructions are cooperatively transmitted to the target storage node through a cross-modal bus protocol; The step S1 comprises: S1.1: For each of the data blocks S1.2: Extracting its S1.3: Wherein, S1.4: Extracting its S1.5: Wherein, S1.6: Representing the S1.7: Wherein, the S1.8: A pre-defined feature extraction function to capture its key attributes from the original data block; S1.2: Constructing a similarity matrix based on the extracted feature vectors of each data block , whose elements are calculated by formula , where is the squared Euclidean distance between the feature vectors and , which is used to quantify the similarity between data blocks, is the total number of data blocks, is the kernel width parameter, which adjusts the rate of similarity decay with distance; S1.3: Calculate the angle matrix The angle matrix elements on the diagonal Represents data block The total similarity, and the angle matrix All off-diagonal elements are zero, and the similarity matrix is ​​used. and the angle matrix Construct the transition matrix The transition matrix represents the probability transitions of random walks between data blocks. S1.4: For the transition matrix Perform eigenvalue decomposition to obtain its frontier. The largest eigenvalues , where and the set of eigenvectors corresponding to the largest eigenvalue. Among them, the largest eigenvalue This reflects the importance of the data manifold in the corresponding direction. For corresponding eigenvalues The Each eigenvector represents a principal direction of the data in the manifold space; Based on the maximum eigenvalue and each component in the eigenvector set, each data block is represented as a low-dimensional manifold coordinate Constructing its low-dimensional manifold coordinate in the low-dimensional manifold space - dimensional coordinate , the low-dimensional manifold coordinate - dimensional coordinate The projection of the data block in each principal direction is integrated and weighted by the eigenvalue, so that the manifold coordinates of all data blocks are combined to form a manifold space coordinate matrix , the coordinate matrix The first Behavior , realizes the low-dimensional manifold embedding representation of the original high-dimensional data block; In order to construct a cross-modal transmission channel based on tensor manifold encapsulation, the step S4 further comprises: S4.1: the optimal spin configuration obtained by quantum annealing in step S2 encoding into a sequence of quantum instructions encoding into a sequence of quantum instructions corresponding to each data block are aggregated, denoted as wherein, is the total number of data blocks, is the final spin state of the th data block, denotes the aggregation operation, and each quantum label has the following data structure: ; Wherein, an 8-bit opcode field to indicate the type of the tag or a subsequent operation; a 16-bit qubit address field to uniquely identify a qubit corresponding to the i th data block; is a 32-bit field that stores or represents the final spin state of the i-th data block ;​ S4.2: For each data chunk in the data slice determined by the optimal spin configuration a data chunk index set of each data slice determined , first construct a tensorized vector representation of the data slice , accumulate the tensor product of the manifold coordinates of all data chunks in the data slice and the data chunk itself by a tensor accumulation model, and the formula of the tensor accumulation model is: ; Wherein, is the manifold coordinate vector for the th data block; for the first original data block or its predetermined representation; denotes a tensor product operation; Then, using the resource mapping matrix transpose Combining manifold eigenvectors The generated manifold basis space tensor For the tensor quantization vector representation Transformation and encapsulation are performed to construct the transport tensor payload. , ,in, It represents the multiplication operation of matrices and vectors, or a generalized linear transformation. 2.The big data distributed storage method based on cloud computing of claim 1, wherein, The kernel width parameter of the similarity matrix calculated in the step S1.2 The value of the kernel width parameter is determined by calculating the average value of the square of the Euclidean distance between the feature vectors of all data block pairs, and the specific calculation formula is as follows:​ ; wherein, is the total number of data blocks in the step S1.1, and are the feature vectors extracted in the step S1.1 for the th and the th data block, respectively, denotes the squared Euclidean distance between the feature vectors and . 3.The big data distributed storage method based on cloud computing of claim 1, wherein, In order to realize dynamic quantum slicing based on the geometric characteristics of the data manifold, the step S2 specifically comprises: S2.1: obtaining a manifold space coordinate matrix based on the manifold space coordinate matrix obtained in step S1 , calculating a Riemannian metric tensor describing the local intrinsic geometric properties of the data manifold wherein, and are indices of the local coordinate system of the manifold S2.2: deriving the Riemannian metric tensor from the data manifold and the manifold space coordinate matrix for any pair of data blocks and deriving coupling coefficients in the lsing model , the coupling coefficients quantitatively describe the geometric correlation strength or interaction strength of data blocks and on the data manifold; S2.3: Constructing the Hamiltonian of the target Ising model wherein, is a spin variable assigned to a data block , takes a value indicating a shard to which the data block belongs, is an external local field parameter applied to the spin . optimizing the hamiltonian to obtain a set of optimal spin configurations the optimal spin configurations determines the dynamic sharding result of all data blocks. 4.The big data distributed storage method based on cloud computing of claim 3, wherein, In step S2.3, the step of optimizing and solving the Hamiltonian comprises the step of solving the optimal spin configuration of the Ising model Hamiltonian using quantum annealing technology, specifically: a) Initialize the system to a state containing A system of 100 qubits, where each qubit corresponds to the spin variable of a data block. Its initial quantum state For all One possible ground state to calculate Uniform re-addition: The initial quantum state is the initial Hamiltonian. The ground state; b) state of the quantum system The evolution over time follows the instantaneous Schrödinger equation: where the Planck constant is taken as 1 ; c) the instantaneous Hamiltonian from the initial Hamiltonian slow, continuous adiabatic evolution to a final target Hamiltonian ; the final target Hamiltonian is the Ising model Hamiltonian of the optimal data shard scheme after ground state encoding , so that the system is in the ground state of the final target Hamiltonian with high probability at the end of evolution, and the optimal spin configuration is obtained by measuring the ground state of the final target Hamiltonian . ​ 5. The big data distributed storage method based on cloud computing according to claim 1, characterized in that, In order to realize secure, optimized and structure-preserving mapping of encrypted data slices to storage resources, the step S3 comprises: S3.1: Constructing an algebraic space for describing storage resource transformation characteristics , wherein is a special orthogonal group consisting of n x n real skew-symmetric matrices is the Lie algebra of the special orthogonal group , wherein is the number of Lie algebra components of the aggregation S3.2: obtaining data shards by step S2 Applying a fully homomorphic encryption algorithm to encryption, obtaining ciphertext data shards capable of computing in an encrypted state , wherein the encryption process can be represented as Enc , and (mod , wherein, is a plaintext unit, is a public matrix, s is a private key, is an error term, is a modulus; S3.3: constructing a cost matrix based on the ciphertext data shards , constructing a cost matrix , wherein, is a data shard weight, is an encrypted data shard, is a resource request vector, and by solving an optimization problem Tr an optimal special orthogonal resource mapping matrix , the special orthogonal resource mapping matrix is used to align the encrypted data shards with the storage resources; S3.4: Based on the optimal n x n special orthogonal resource mapping matrix U obtained by optimization solution, the final resource mapping matrix is generated , which is an element of the special orthogonal group and satisfies the orthogonal constraint , is the unit matrix, and can be expressed by the exponential mapping of Lie algebra to Lie group , where is the skew-symmetric basis matrix of , is the parameter determined by optimization. 6.The big data distributed storage method based on cloud computing according to claim 5, wherein, Resource mapping matrix in step S3.4 by a set of scalar parameters , and the corresponding, pre-defined Lie algebra of real anti-symmetric matrices Parameterizing the generation of the resource mapping matrix in step S3.3 for solving the optimization problem with iterative updates of the resource mapping matrix includes the following operations: In the In the next iteration, if the parameters To make adjustments, the resource mapping matrix is ​​changed from its current state using the resource state mapping formula. Update to the next state ,in, For parameter index, The resource state mapping formula is: ; Wherein, For the first iteration, the special orthogonal resource mapping matrix is is updated by the parameter adjustment special orthogonal resource mapping matrix (eta) is a positive scalar learning rate, controlling the iteration step size; is the matrix exponential function; For a parameter The gradient information is constructed and belongs to the Lie algebra. of Real antisymmetric matrix The specific calculation steps are as follows: First, the cost function is calculated When the current resource mapping matrix is , the partial derivative value of the th scalar parameter is recorded as the scalar gradient value , , where is the cost matrix; Then, using the scalar gradient values and the scalar parameter corresponding Lie algebra basis matrix to construct the real antisymmetric matrix , .

7. A system for implementing the method of any one of claims 1 to 6, characterized in that, Comprise: A data manifold feature extraction module is configured to extract feature vectors from data blocks, calculate a similarity matrix between data blocks, and perform feature decomposition on the similarity matrix to obtain manifold coordinates and manifold feature vectors of the data manifold; A dynamic quantum slicing module is configured to construct an Ising physical model according to the geometric characteristics of the data manifold, and solve the ground state of the model through quantum annealing technology to determine the optimal slicing scheme for allocating data blocks to different data slices; A secure resource mapping module is configured to homomorphically encrypt the data slices, and construct and solve an optimization problem defined on a special orthogonal group manifold to generate an optimal special orthogonal mapping matrix for securely mapping encrypted data slices to storage resources; A transmission encapsulation module is configured to encode the optimal slicing scheme into an instruction sequence, and construct a transmission tensor load for each data slice, which encapsulates manifold geometric information, data content and is transformed by the optimal special orthogonal mapping matrix.

Citation Information

Patent Citations

  • 3D graph data processing method of quantum graph neural network based on rotation arrangement equivariant

    CN119625238A

  • Equipment group collaborative fault prediction analysis method and system for complex industrial process

    CN120029230A