Distributed data storage system based on block chain
By introducing blockchain technology and smart contract modules into the distributed data storage system, data security and consistency problems are solved, efficient processing of complex data types and semantic-level data matching are achieved, and the overall performance and reliability of the system are improved.
Patent Information
- Application Number
- CN202510066970.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
AI Technical Summary
The existing distributed data storage technology has the problems of data being susceptible to tampering and leaking, difficulty in maintaining data consistency, and limited processing capabilities for complex data types.
A distributed data storage system based on blockchain is adopted, and data block processing, encryption and hash value calculation is used, combined with the improved PoS consensus mechanism, smart contract module, hash index and Bloom filter combined index structure to achieve secure storage and efficient retrieval of data.
It improves data security and consistency, enhances processing capabilities for complex data types, realizes semantic-level data matching, improves the accuracy and speed of data retrieval, and ensures high availability and anti-attackability of data storage.
Smart Images

Figure CN119995824A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of distributed storage technology, and in particular to a distributed data storage system based on blockchain. Background Art
[0002] In today's digital age, data is growing explosively, and the demands for data storage in various industries are becoming increasingly complex and diverse. Distributed data storage technology has emerged as the times require, becoming a key architecture to support large-scale data management and applications. The emergence of blockchain technology has further brought new opportunities for change in distributed data storage. The distributed data storage system based on blockchain combines the decentralized, tamper-proof, and traceable characteristics of blockchain with the efficient storage and processing capabilities of distributed storage, and has vital application value in many fields such as finance, medical care, and the Internet of Things.
[0003] However, existing distributed data storage technologies still have many shortcomings. In traditional distributed storage systems, data is vulnerable to tampering and leakage risks, and it is difficult to meet application scenarios with extremely high requirements for data integrity and confidentiality. In terms of data consistency maintenance, data synchronization between nodes is prone to delays or errors, resulting in frequent data inconsistency problems, affecting the reliability and availability of the system. The processing capabilities for complex data types are limited. For example, when facing multimedia data, image recognition-related data, etc., there is a lack of effective data conversion and feature extraction mechanisms, making it difficult to fully tap the value of data and unable to meet the growing needs of intelligent data applications. To this end, we propose a distributed data storage system based on blockchain. Summary of the invention
[0004] In view of the deficiencies in the prior art, the present invention provides a distributed data storage system based on blockchain, thereby solving the technical problems mentioned in the background technology.
[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: a distributed data storage system based on blockchain, including a data storage module, a consensus module, a smart contract module, a data retrieval module and a data update and maintenance module;
[0006] The data storage module is responsible for processing the data in blocks. It divides the original data into blocks of appropriate size according to the system performance and network conditions. Let the data block be x i (i=1,2...,n), the encryption key is k, then the encrypted data block E(x i ) can be expressed as E(x i )=AES(x i ,k), for each data block x i , calculate its hash value h(x i ), the formula is h(x i )=SHA-256(xi ), for image recognition related data management, let the data be y;
[0007] When the data y is in image format, the color value of the pixel is used to read the image data through the Pillow library, and it is converted into a pixel code p in the form of a multidimensional array according to the arrangement order of the pixels, that is, p = P(y), and then the scale space of the image is constructed. The image is convolved with Gaussian blur kernels of different scales to obtain images of different scales, and then extreme points are found in the scale space. The extreme points are representative feature points in the image, and these feature points are used as the center to calculate the gradient direction and amplitude of the surrounding neighborhood pixels to construct a descriptor, that is, the feature vector v, v = F(p);
[0008] When the data y is not in image format, it needs to go through a complex conversion process. Specifically, each character in the text data is mapped to a specific value by using ASCII code or Unicode encoding, and these values are arranged and combined to form a structure similar to the image pixel matrix, and then normalized to finally obtain the pixel code p. Then, the bag-of-words model is used to count the frequency of each word to obtain the feature vector v of the entire text, v = F(p);
[0009] When storing pixel codes and feature vectors, block processing is performed. Suppose the length of the pixel code p is L p , the block size is B p , then the number of pixel code blocks in Indicates rounding up. For the j-th pixel code block p j , perform encryption operation E(p j )=AES(p j ,k p ), where k p is the encryption key of the pixel code, and its hash value h(p j )=SHA-256(p j ) for information storage;
[0010] The consensus module adopts an improved PoS consensus mechanism. Nodes compete for the right to record accounts based on their rights. When new data is written, the node that obtains the right to record accounts generates a new block in combination with the blockchain status and broadcasts it. Other nodes are added after verification. New nodes are assigned rights and interests according to resource information after application verification. Data of exiting nodes is migrated according to the storage capacity and load of other nodes. Forks are handled according to the longest chain principle. In special cases, secondary hashing is used to determine the main chain.
[0011] The smart contract module defines rules based on user attributes and data access requirements in data access control, and determines user permissions through logical expressions; in storage resource allocation management, it intelligently allocates tasks based on node storage capacity, network bandwidth and load; it triggers automated execution through blockchain events, and ensures safety and reliability through code auditing before execution and multi-node verification during execution;
[0012] The data retrieval module constructs a combined index structure based on hash index and Bloom filter, establishes hash indexes for data blocks, pixel code blocks and feature vector sub-vectors respectively, uses Bloom filter to quickly determine whether an element exists, and first filters it through Bloom filter before searching accurately during query; verifies the location and integrity of data blocks with the help of blockchain block information, retrieves feature vectors with vector similarity algorithm and determines matching data based on threshold value;
[0013] When updating data, the data update and maintenance module generates a request containing the update content and the original data identifier at the initiating node, and sends it to the storage node to update the data block after verification and consensus. When multiple update requests conflict, they are processed first according to the timestamp; node failures are detected through the heartbeat mechanism, and data of the failed node is migrated; important data is backed up to the backup node according to the backup strategy, and restored from the backup node when data is lost or damaged.
[0014] In one possible implementation, in the consensus module, the rights and interests of a node are determined based on its resource contribution in the system, token holdings, or other factors related to system contribution.
[0015] In one possible implementation, in the smart contract module, the access control function of the smart contract uses a logical expression to determine the user's access rights to data.
[0016] In one possible implementation, the smart contract module and the storage task allocation function allocate storage tasks to different nodes according to a weighted strategy of storage capacity and network bandwidth, and the weights can be set according to system requirements.
[0017] In a possible implementation, the vector similarity calculation function of the data retrieval module uses a cosine similarity formula to calculate the similarity between the query feature vector and the stored feature vector, and determines whether the data matches based on a set threshold.
[0018] In a possible implementation, in the data update and maintenance module, the conflict handling function during the data update process determines the execution priority of the update request strictly according to the order of timestamps to ensure the consistency and timeliness of the data update.
[0019] In a possible implementation, in the heartbeat mechanism of node fault detection in the data update and maintenance module, if a node does not send a heartbeat signal within a specific time, it is determined to be a fault. The specific time can be adjusted according to system performance and network conditions.
[0020] Beneficial effects compared with the prior art:
[0021] 1. In this solution, the system converts data into pixel codes and extracts feature vectors. For different types of data, such as multimedia, text or structured data, a deeper information representation of the data is achieved through the conversion of pixel codes and feature vectors. This means that the system is no longer limited to traditional keyword-based or simple index-based search methods when searching, but can perform semantic-level matching based on feature vectors. This method greatly improves the accuracy and speed of data retrieval. For large data retrieval tasks, users can find the required data more quickly, which improves the overall information processing efficiency of the system and provides more powerful data retrieval capabilities for various data-intensive application scenarios, such as search engines, content recommendation systems, image databases, etc.
[0022] 2. In this scheme, in terms of data storage, this system demonstrates excellent security features. By using the AES encryption algorithm to encrypt data blocks, pixel code blocks, and feature vector sub-vectors, the confidentiality of data during storage and transmission is ensured to prevent the risk of data leakage. At the same time, the blockchain's hash algorithm SHA-256 is used to protect the integrity of the data and generate a unique hash value for each data block. Any tampering with the data will cause the hash value to change, making it easier to detect data anomalies in a timely manner. This not only ensures the security of data storage, but also avoids the impact of single point failures on data through distributed storage and the decentralized nature of blockchain, ensuring the high availability and anti-attack of data storage;
[0023] 3. In this solution, the consensus mechanism and resource allocation mechanism of the system significantly improve performance and resource utilization efficiency. The improved PoS consensus mechanism is adopted to reasonably allocate accounting rights according to the rights and interests of the nodes, avoiding the waste of resources and performance bottlenecks in the traditional consensus mechanism, and ensuring the consistency and efficiency of data operations. In terms of storage resource allocation, the smart contract intelligently allocates storage tasks according to factors such as the storage capacity of the node, network bandwidth and current load, and realizes the optimal configuration of resources. At the same time, the timestamp priority principle and conflict handling mechanism during data update ensure the consistency and timeliness of data updates. In distributed storage applications, the system can efficiently handle a large number of data storage and operation requests, reduce data processing delays and errors caused by system performance problems, improve the overall performance and reliability of the system, and provide a better solution for large-scale data storage and processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention and implement it according to the contents of the specification, the following is a detailed description of the preferred embodiments of the present invention in conjunction with the accompanying drawings.
[0025] Figure 1 This is a system diagram of the distributed data storage system based on blockchain of the present invention. DETAILED DESCRIPTION
[0026] The preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings. However, the present invention can be implemented in various forms, so the present invention is not limited to the embodiments described below. In addition, in order to more clearly describe the present invention, components that are not connected with the invention will be omitted from the drawings.
[0027] The technical solution in the embodiment of the present application is to solve the problems of the above-mentioned background technology, and the overall idea is as follows:
[0028] Embodiment 1:
[0029] This embodiment introduces a distributed data storage system based on blockchain, including a data storage module, a consensus module, a smart contract module, a data retrieval module, and a data update and maintenance module;
[0030] 1. Data storage module
[0031] Data storage is the foundation of this distributed system and plays a vital role in system performance and data security.
[0032] First, data needs to be divided into blocks before entering the storage link. Assuming that the size of the original data is D bytes, the data is divided into blocks of B bytes according to the system performance and network conditions. Then the number of data blocks n can be calculated by the formula Calculated, among which Indicates rounding up. This is to facilitate data management and transmission, splitting large data into multiple small blocks for storage and processing in distributed storage nodes.
[0033] For data encryption, an advanced encryption algorithm is used, specifically the AES (Advanced Encryption Standard) encryption algorithm. Suppose the data block is x i (i=1,2...,n), the encryption key is k, then the encrypted data block E(x i ) can be expressed as E(x i )=AES(x i ,k). Encryption operations ensure the security of data during storage and transmission, preventing data leakage and tampering.
[0034] The blockchain hash algorithm (such as SHA-256) is used to ensure the integrity of the data. i , calculate its hash value h(x i ), the formula is h(x i )=SHA-256(x i ). These hash values will be stored in the corresponding locations of the blockchain as the "fingerprint" of the data block. When it is necessary to verify whether the data has been tampered with, the hash value h'(x i ), and compared with the original hash value h(x i ) comparison, if h(x i )=h'(x i ), the data has not been tampered with; otherwise, it indicates that the data has been modified.
[0035] In many application scenarios, such as multimedia data storage and image recognition related data management, data is converted into pixel codes and feature vectors are extracted to better analyze and retrieve data and use some processing algorithms based on image or data features. For the process of converting data into pixel codes and extracting feature vectors, assume that the data is y.
[0036] If the data y itself is in image format, then the conversion to pixel code is relatively straightforward. An image is essentially a matrix of pixels, each of which has a specific color value (in the common RGB color mode, each pixel is represented by the values of the three channels of red, green, and blue, and the value range is usually 0-255). The image data is read out through the corresponding image reading library (such as the Pillow library in Python), and then converted into a pixel code p in the form of a multidimensional array according to the order of the pixels, that is, p = P(y).
[0037] If the data p is not in image format but other types of data, such as text, audio, etc., it needs to go through a complex conversion process. Taking text data as an example, first map each character in the text to a specific value (using ASCII code or Unicode encoding), arrange and combine these values to build a structure similar to the image pixel matrix. Taking a simple two-dimensional arrangement as an example, we can arrange it according to the length of the text and the set "matrix" dimension. If the text length is L, set the width of the two-dimensional matrix to W, and the height H can be obtained by It is calculated that Indicates rounding up, and fills the mapped values into this two-dimensional matrix in order from left to right and from top to bottom. For example, for the text "Hello", the ASCII codes corresponding to its characters are 72, 101, 108, 108, and 111. If the width W is set to 3, the height Then the matrix after arrangement is Here, due to the insufficient length of the text, the last position is filled with 0 (other filling values or methods can also be used according to specific needs). This arrangement forms a structure similar to the image pixel matrix, which is convenient for subsequent operations similar to image processing. In order to make this structure more in line with the needs of subsequent processing, normalization and other operations are required to finally obtain the pixel code p.
[0038] The process of extracting feature vectors from pixel codes requires the use of special feature extraction algorithms. For image pixel codes, commonly used algorithms include the scale-invariant feature transform (SIFT) algorithm and the speeded-up robust feature (SURF) algorithm. Taking the SIFT algorithm as an example, the scale space of the image is first constructed, and a series of images at different scales are obtained by convolving the image with Gaussian blur kernels of different scales. Then, extreme points are found in these scale spaces. These extreme points are considered to be representative feature points in the image. Then, with these feature points as the center, the gradient direction and amplitude of the surrounding neighborhood pixels are calculated, and a descriptor is constructed based on this information. This descriptor is the feature vector v extracted from the pixel code, that is, v = F(p).
[0039] For pixel codes converted from non-image data, the feature extraction algorithm will also be different. For example, for specially processed text pixel codes, the bag-of-words model or word embedding technology is used to extract feature vectors. The bag-of-words model regards all words in the text as a set, regardless of the order of the words, and only counts the frequency of each word to construct a feature vector. The word embedding technology maps each word to a low-dimensional vector space, and obtains the feature vector of the entire text by combining or operating on these vectors.
[0040] When storing these pixel codes and feature vectors, they are also processed in blocks. Suppose the length of the pixel code p is L p , the block size is B p , then the number of pixel code blocks For the jth pixel code block p j , perform encryption operation E(p j )=AES(p j ,k p ), where k p is the encryption key of the pixel code. At the same time, its hash value h(p j )=SHA-256(p j ).
[0041] For a feature vector v, it can be represented as an m-dimensional vector v = (v1, v2, ..., vm ), according to the dimension of the feature vector and the storage strategy set by the system, it is split into multiple sub-vectors v k (k=1,2,...,q), where q is the number of subvectors determined based on storage optimization. Each subvector v k Encryption operation E(v k )=AES(v k ,k v ), where k v is the encryption key of the feature vector, and its hash value h(v k )=SHA-256(v k ).
[0042] The encrypted data blocks, pixel code blocks and feature vector sub-vectors are stored in different storage nodes according to the system storage node set S = {s1, s2, ..., s N}, according to a certain storage strategy ST, the data is allocated to the storage nodes, for example, using a hash function H to determine the storage location, the i-th data block x i Stored on the node Above, the pixel code block p j Stored on the node On the eigenvector vector v k Stored on the node superior.
[0043] 2. Consensus Module
[0044] This system adopts an improved PoS (Proof of Stake) consensus mechanism to ensure the data consistency and reliability of each node in the system.
[0045] Under this mechanism, each node n has a certain equity S n , the total equity of all nodes in the system is The probability P of node n obtaining the right to record accounts n The formula Calculation. The stake of a node can be determined based on its resource contribution to the system, token holdings, or other factors related to its contribution to the system.
[0046] When new data needs to be written into the system, each node participates in the competition for the right to record according to its rights. Suppose the system generates a new transaction or data update operation T, the node will calculate its own probability of recording rights according to its rights and try to package the operation. w Combine data operation T with the previous blockchain state BC prev Combined, a new block B is generated through function M n =M(T,BC prev ).
[0047] The node will generate a new block B n Broadcast to other nodes. After receiving the new block, other nodes will verify it. Verification process V(B n ) includes checking the structure of the block, the validity of the transaction, the consistency of the data, etc. If the verification passes, the node will add the new block to its own blockchain. If the verification fails, the node will reject the new block and trigger the system's exception handling mechanism.
[0048] In terms of handling node joining and leaving, when a new node n new When joining the system, you first need to submit application information to the system new , including its resource information, identity information, etc. The verification node of the system will new Verify, the verification function is V(I new ). If V(I new ) is true, the new node is allowed to join and is assigned to the new node based on the resource information R it provides. new Allocate corresponding equity S new , according to the system's equity distribution function A(R new ), namely S new =A(R new ).
[0049] When node n exit When exiting the system, the system will start the exit process. First, the data stored in the node is redistributed. exit The stored data set is D exit , the system will be based on the storage capacity of other nodes C = {C1, C2, ..., C N} and the current load L={L1,L2,...,L N}, through the data migration function Migr(D exit ,C,L) Migrate data to other suitable nodes.
[0050] For the fork problem, suppose there are two or more possible blockchain branches BC1, BC2, ..., BC m , the system adopts the longest chain principle. By calculating the length of each chain Len(BC i ), the node will select the longest blockchain as the main chain, that is, When a fork occurs, the node will suspend operations and wait for a certain time window t. During the time window, if the length of one chain exceeds that of other chains, then that chain will become the main chain. If there are still multiple longest chains of equal length after the time window ends, an additional consensus mechanism, specifically quadratic hash voting, can be used to determine the final main chain.
[0051] 3. Smart Contract Module
[0052] Smart contracts play a key role in automated and intelligent management in this distributed data storage system.
[0053] In terms of data access control, smart contracts can define access rules. Suppose user u has attribute set A u ={a1,a2,...a k}, data d has access requirement set R d ={r1,r2,...,r l}. The access control function AC(A u ,R d ) is used to determine whether a user has the right to access data. For example, for user u to access data d, when AC(A u ,R d ) is true, the user is allowed access, otherwise access is denied. Logical expressions can be used to express access conditions, such as Where Λ represents the logical AND operation.
[0054] In the allocation and management of storage resources, the smart contract allocates storage resources according to the storage capacity of the node C = {C1, C2, ..., C N}、Network bandwidth W={W1,W2,...,W N} and the current storage load L = {L1, L2, ..., L N} to allocate storage tasks. Let the storage task be ST, and the allocation function of the smart contract be Alloc(ST,C,W,L), which will allocate storage tasks to different nodes according to a certain strategy. For example, tasks can be allocated based on the weighted sum of storage capacity and network bandwidth. The proportion of storage tasks allocated to node n is P n It can be expressed as where w1 and w2 are the corresponding weights.
[0055] The automated execution of smart contracts is achieved through the event triggering mechanism of the blockchain. When certain conditions are met, such as a data update event E update When it happens, the smart contract will automatically execute the corresponding storage resource reallocation operation. The execution of the smart contract can be expressed as Exec(E update ), it calls the corresponding allocation function Alloc and updates the storage task of the node.
[0056] In order to ensure the security and reliability of smart contracts, multiple verification and audit mechanisms are adopted. For the smart contract code SC, a code audit (SC) will be conducted before deployment to check for code vulnerabilities and potential security risks. At the same time, during the execution of the smart contract, there will be multiple verification nodes V = {v1, v2, ..., v m} Perform multiple verifications. For the execution request ER of the smart contract, the verification function Ver(ER) will check the legality and security of the execution request. The execution of the smart contract is allowed only when Audit(SC) passes and Ver(ER) is true.
[0057] 4. Data Retrieval Module
[0058] In distributed storage systems, achieving efficient data retrieval is the key to improving system performance. A combined index structure based on hash index and Bloom filter is adopted. i , pixel code block p j and the eigenvector v k , respectively construct hash indexes. Let the hash function be H, then the data block x i The index is Pixel code block p j The index is Eigenvector v k The index is
[0059] Bloom filter BF is used to quickly determine whether an element exists. For a data block set X = {x1, x2, ..., x n}, pixel code block collection and the eigenvector set V = {v1,v2,...,v q}, when the system is initialized, their indexes are added to the Bloom filter, i.e. When querying data, a quick judgment is first made through the Bloom filter. If the Bloom filter determines that the data does not exist, the result is directly returned as not found; if it may exist, an accurate search is performed based on the hash index.
[0060] For the query of data blocks, let the query condition be Q x , find the set of nodes that may store data blocks through hash index For the query of pixel code blocks, let the query condition be Q p , find the set of nodes that may store pixel code blocks For the query of feature vectors, let the query condition be Q v , find the set of nodes that may store sub-vectors of eigenvectors
[0061] In terms of using blockchain features to accelerate data retrieval, the block header of the blockchain contains information such as the hash value and Merkle root of the previous block. i Stored in a block, the block header H block The Merkle root MR in can help verify the location and integrity of the data block. Assume that the Merkle tree verification function VerifyMR(x i ,MR) can verify whether the data block belongs to the block. When VerifyMR(x i ,MR) is true, confirming the storage location and integrity of the data block.
[0062] For feature vector retrieval, its special vector characteristics can be used. Assume that the query feature vector is v q =(v q1 ,v q2 ,...,v qm ), the stored eigenvector subvector is v k =(v k1 ,v k2 ,...,v km ), through the vector similarity calculation function Sim(v q ,v k ), the cosine similarity formula is The similarity is calculated and a similarity threshold t is set. When the similarity exceeds the set threshold t, it is considered to be matching data.
[0063] 5. Data update and maintenance module
[0064] Data update is an inevitable operation in the system, and the consistency and reliability of data update need to be guaranteed.
[0065] When data needs to be updated, an update request UR is first generated at the initiating node, containing the updated data content d new and the ID of the original data old The update request will be sent to the verification node for verification V(UR), and after verification, it will be sent to the consensus node for consensus operation C(UR).
[0066] Assume that the data block involved in the data update is x i , and its updated content is The updating process can be expressed as After verification and consensus operation, the node set storing the data block The node will update the data block according to the update information
[0067] During the update process, conflicts may occur, especially in a distributed environment. When multiple update requests are made for the same data at the same time, the timestamp priority principle is adopted. Let the timestamp of update request UR1 be t1, and the timestamp of update request UR2 be t2. If t1>t2, UR1 will be executed first. The conflict handling function Conflict(UR1,UR2) will determine the priority based on the timestamp to ensure the order and consistency of the update.
[0068] For daily maintenance of the storage system, node fault detection is performed first. Through the heartbeat mechanism, node n will periodically send heartbeat signals HB n , the system maintenance node will monitor the heartbeat signal. If the heartbeat signal of node n is not received within time T, it is determined that the node may be faulty. f , the data set stored is D nf , through the fault handling function HandleFailure(n f ,D nf ) Migrate data to other normal nodes.
[0069] For data backup, according to the system backup strategy BP, important data sets D important Back up to the backup node set N backup , the backup operation can be expressed as Backup(D important ,N backup ,BP). For example, a redundant backup strategy is adopted to back up data to multiple nodes in different geographical locations to ensure data availability. When data is lost or damaged, the recovery function Recover(D lost ,N backup ) Restore data from the backup node, where D lost is the set of missing data.
[0070] This blockchain-based distributed data storage system forms a complete, secure and reliable data storage ecosystem through the collaborative work of multiple modules such as data storage, consensus mechanism, smart contracts, data retrieval, and data update and maintenance. Each module plays its own unique role in ensuring the security, consistency and efficiency of data, and cooperates with each other to provide powerful data storage and management functions for different application scenarios.
[0071] Finally, it should be noted that: Obviously, the above embodiments are only examples for clearly explaining the present invention, and are not intended to limit the implementation methods. For ordinary technicians in the relevant field, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. The obvious changes or modifications derived from this are still within the scope of protection of the present invention.
Claims
1. A distributed data storage system based on blockchain, characterized in that: It includes data storage module, consensus module, smart contract module, data retrieval module and data update and maintenance module; The data storage module is responsible for processing the data in blocks. It divides the original data into blocks of appropriate size according to the system performance and network conditions. Let the data block be x i (i=1,2...,n), the encryption key is k, then the encrypted data block E(x i ) can be expressed as E(x i )=AES(x i ,k), for each data block x i , calculate its hash value h(x i ), the formula is h(x i )=SHA-256(x i ), for image recognition related data management, let the data be y; When the data y is in image format, the color value of the pixel is used to read the image data through the Pillow library, and it is converted into a pixel code p in the form of a multidimensional array according to the arrangement order of the pixels, that is, p = P(y), and then the scale space of the image is constructed. The image is convolved with Gaussian blur kernels of different scales to obtain images of different scales, and then extreme points are found in the scale space. The extreme points are representative feature points in the image, and these feature points are used as the center to calculate the gradient direction and amplitude of the surrounding neighborhood pixels to construct a descriptor, that is, the feature vector v, v = F(p); When the data y is not in image format, it needs to go through a complex conversion process. Specifically, each character in the text data is mapped to a specific value by using ASCII code or Unicode encoding, and these values are arranged and combined to form a structure similar to the image pixel matrix, and then normalized to finally obtain the pixel code p. Then, the bag-of-words model is used to count the frequency of each word to obtain the feature vector v of the entire text, v = F(p); When storing pixel codes and feature vectors, block processing is performed. Suppose the length of the pixel code p is L p , the block size is B p , then the number of pixel code blocks in Indicates rounding up. For the j-th pixel code block p j , perform encryption operation E(p j )=AES(p j ,k p ), where k p is the encryption key of the pixel code, and its hash value h(p j )=SHA-256(p j ) for information storage; The consensus module adopts an improved PoS consensus mechanism. Nodes compete for the right to record accounts based on their rights. When new data is written, the node that obtains the right to record accounts generates a new block in combination with the blockchain status and broadcasts it. Other nodes are added after verification. New nodes are assigned rights and interests according to resource information after application verification. Data of exiting nodes is migrated according to the storage capacity and load of other nodes. Forks are handled according to the longest chain principle. In special cases, secondary hashing is used to determine the main chain. The smart contract module defines rules for data access control based on user attributes and data access requirements, and determines user permissions through logical expressions; In terms of storage resource allocation management, tasks are intelligently allocated based on node storage capacity, network bandwidth, and load; automated execution is triggered by blockchain events, and code auditing before execution and multi-node verification during execution ensure safety and reliability; The data retrieval module constructs a combined index structure based on hash index and Bloom filter, establishes hash indexes for data blocks, pixel code blocks and feature vector sub-vectors respectively, uses Bloom filter to quickly determine whether an element exists, and first filters it through Bloom filter before searching accurately during query; verifies the location and integrity of data blocks with the help of blockchain block information, retrieves feature vectors with vector similarity algorithm and determines matching data based on threshold value; When updating data, the data update and maintenance module generates a request containing the update content and the original data identifier at the initiating node, and sends it to the storage node to update the data block after verification and consensus. When multiple update requests conflict, they are processed first according to the timestamp; Detect node failures through the heartbeat mechanism and migrate data from failed nodes; Back up important data to the backup node according to the backup strategy, and restore from the backup node when the data is lost or damaged.
2. The distributed data storage system based on blockchain as claimed in claim 1, characterized in that: In the consensus module, the rights and interests of a node are determined based on its resource contribution in the system, token holdings, or other factors related to its contribution to the system.
3. The distributed data storage system based on blockchain as claimed in claim 1, characterized in that: The smart contract module, the access control function of the smart contract uses a logical expression to determine the user's access rights to the data.
4. The distributed data storage system based on blockchain as claimed in claim 1, characterized in that: The smart contract module and the storage task allocation function allocate storage tasks to different nodes according to the weighted and strategy of storage capacity and network bandwidth, and the weights can be set according to system requirements.
5. The distributed data storage system based on blockchain as claimed in claim 1, characterized in that: In the data retrieval module, the vector similarity calculation function uses the cosine similarity formula to calculate the similarity between the query feature vector and the stored feature vector, and determines whether the data matches according to a set threshold.
6. The distributed data storage system based on blockchain as claimed in claim 1, characterized in that: In the data updating and maintenance module, the conflict handling function in the data updating process strictly determines the execution priority of the update request according to the sequence of timestamps to ensure the consistency and timeliness of the data updating.
7. The distributed data storage system based on blockchain as claimed in claim 1, characterized in that: In the data updating and maintenance module, in the heartbeat mechanism of node fault detection, if a node does not send a heartbeat signal within a specific time, it is determined to be a fault. The specific time can be adjusted according to system performance and network conditions.