Large-scale data distributed storage method

By adopting distributed storage architecture and blockchain authentication in the data storage system, the problems of insufficient performance and low security in traditional data storage under massive data are solved, and efficient and secure data transmission is achieved.

CN120050294APending Publication Date: 2025-05-27SHANDONG ENERGY GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510051286.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When facing massive data, traditional data storage methods are difficult to meet the requirements of capacity, performance and bandwidth, resulting in data storage systems that are prone to excessive delays and process blockages in high concurrency, and even system crashes. At the same time, the prior art cannot guarantee the security and coordination of data during data transmission.

Method used

The large-scale distributed storage method is adopted to classify the data sets and store them on different storage nodes using a distributed storage architecture. Users are authenticated through blockchain and smart contracts are used to determine the user's read permissions. During the data transmission process, a direct-connected transmission method is used to verify the communication addresses of the data receiving and sending ends through multiple encryption to ensure the security of data transmission.

Benefits of technology

By moving data to the storage node closest to the requested location, the data transmission rate is increased and the load is reduced. The use of blockchain and smart contracts ensures the transparency and traceability of data authentication, reduces data transmission latency and improves the speed, while enhancing the security of data transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050294A_ABST
    Figure CN120050294A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of big data, and relates to a large-scale data distributed storage method. In order to solve the problems that in the prior art, the storage position of data is not associated with the request position of the data, the collaboration of data access records cannot be ensured, and third-party injection and network attacks are easily suffered in the data transmission process, the method comprises the following steps of: classifying request frequencies of any data at different request places; the data is moved to a storage node closest to a request site in spatial distance, a block chain method is used for bearing data request records of a user, and a direct connection type transmission method is adopted in the data transmission process. According to the invention, the data security and the data transmission rate are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of big data and relates to a method for distributed storage of large-scale data. Background Art

[0002] With the booming development of various applications such as mobile devices, social networks, and the Internet of Things, the data generated by human society has grown explosively. The traditional data storage method is usually disk storage. Users store all the data to be stored on the disk so that they can view the data anytime and anywhere. However, as the amount of data to be stored increases, it becomes increasingly difficult for traditional disks to meet the storage needs of users based on massive data in terms of capacity, performance, and bandwidth. Therefore, a data storage system supported by a cloud platform has emerged. A data center is deployed in the data storage system, and users can upload the data to be stored to the data center, and the data center stores the data to be stored.

[0003] In related technologies, a storage cluster for storing data is set in the data center. When a user uploads the data to be stored to the data storage system, the data center will receive the data to be stored, and the data center adds the received data to be stored to the storage cluster for storage. However, the storage location of the data cannot be associated with the request location of the data. During the data extraction process, it is easy to cause excessive delay. In the case of multiple data requests occurring simultaneously, it is also easy to cause process blockage, resulting in the collapse of the storage system.

[0004] At the same time, in the prior art, identity verification is required when accessing data, and the identity verification method cannot be synchronized. When a user accesses a single storage node, the identity of the current requesting user is not recorded in other storage nodes. The coordination of data cannot be guaranteed.

[0005] And during the data transmission process, in the prior art, verification and forwarding are all performed through relay nodes, and it is easy to suffer from third-party injection and network attacks during the data transmission process, and the security of the data during the transmission process cannot be guaranteed. Summary of the Invention

[0006] To solve the problems in the background art, the present invention proposes a method for distributed storage of large-scale data.

[0007] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0008] Classify the large-scale data set and adopt a distributed storage architecture, and each data category is distributed and stored on different storage nodes;

[0009] In response to initiating a call request for data in a storage node, the user is authenticated through a blockchain. After successful authentication, access rights are automatically obtained.

[0010] After successful authentication, address verification is performed between the data sender and the data receiver. After successful address verification, the data sender sends data to the data receiver.

[0011] Furthermore, the specific method for classifying the large-scale data set is as follows:

[0012] By reading the request location of any data in the large-scale data set and the number of requests at the request location, the request frequency of data at different request locations is obtained.

[0013] Read the data with a request frequency greater than a preset threshold frequency at any request location, and extract the feature attributes of the data with a request frequency greater than the preset threshold frequency at this request location.

[0014] Classify the data in the large-scale data set according to the feature attributes to obtain the data categories corresponding to different request locations.

[0015] Furthermore, the data categories corresponding to different request locations are not unique, and any one of the data can appear repeatedly in different categories.

[0016] Furthermore, the specific method for authentication through the blockchain is as follows:

[0017] Calculate the hash value based on the user identity code to obtain the user identity password, and judge the authenticity of the user identity password. If the user authentication fails, access is refused.

[0018] If the user authentication is successful, judge the user's permission to read data based on the user identity password.

[0019] Furthermore, judging the user's permission to read data based on the user identity password is performed through a smart contract.

[0020] Furthermore, the method for address verification between the data sender and the data receiver is as follows:

[0021] The data receiver encrypts the data receiver's communication address multiple times. The character length of the data receiver's communication address gradually becomes longer after each encryption until the character length after n times of encryption is the same as the preset number of digits of the current request user identity code.

[0022] The data sender encrypts the data sender's communication address multiple times. The character length of the data sender's communication address gradually becomes longer after each encryption until the character length after m times of encryption is the same as the current request user's request time.

[0023] The data sender and the data receiver each transmit the encrypted text data and the number of times of their own encryption to each other. The data receiver and the data sender decrypt the data. If the decrypted data conforms to the communication address format, the verification passes.

[0024] Further, the method for the data sender to send data to the data receiver is as follows:

[0025] After the address verification passes, the data sender decrypts the communication address sent by the data receiver;

[0026] The data sender directly sends information to the data receiver according to the communication address of the data receiver received by the data sender;

[0027] The data receiver decrypts the communication address of the data sender according to the data received by the data receiver and receives the information sent by the data sender.

[0028] Compared with the prior art, the present invention has the following beneficial effects:

[0029] By classifying the request frequencies of any data at different request locations, the present invention can move the data to the storage node that is the closest in spatial distance to the request location, greatly improving the data transmission rate and reducing the load of data transmission.

[0030] By using the blockchain method to carry the user's data request records, the present invention ensures that the data can be traced back every time it is requested and the records are retained. At the same time, the present invention also uses smart contracts to verify the user's identity, ensuring the transparency and traceability of the user request records.

[0031] In addition, the present invention also adopts a direct connection transmission method during the data transmission process. After the data receiver and the data sender pass the verification, the data is directly transmitted between the data receiver and the data sender, reducing the data transmission delay and improving the data transmission rate at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 is the overall operation flowchart of the present invention;

[0033] Figure 2 is the data transmission schematic diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0035] As Figure 1 - Figure 2 shown, the technical solution adopted by the present invention is as follows: A large-scale data distributed storage method, including:

[0036] As Figure 1 shown, the method adopted by the present invention is divided into the following processes:

[0037] Classify the large-scale data set, adopt a distributed storage architecture, and each data category is distributed and stored on different storage nodes.

[0038] In response to a call request initiated for the data in the storage node, authenticate the user through the blockchain, and automatically obtain the access permission after successful authentication.

[0039] After the authentication is passed, perform address verification between the data sender and the data receiver. After the address verification is passed, the data sender sends data to the data receiver.

[0040] First, classify the large-scale data set, adopt a distributed storage architecture, and each data category is distributed and stored on different storage nodes.

[0041] The specific method for classifying the large-scale data set is:

[0042] Read the request location of any data in the large-scale data set and the number of requests at the request location to obtain the request frequency of data at different request locations.

[0043] Classify the data requested at a certain request location according to the correlation between multiple data requested at the request location. Enterprises at a certain request location may call the market information in the area, and it also includes users calling the information of users in the same city in the social field, etc. For such reasons, the request location can be associated with the data requested at the request location.

[0044] Read the data with a request frequency greater than the preset threshold frequency at any request location, and extract the characteristic attributes of the data with a request frequency greater than the preset threshold frequency at the request location.

[0045] Cross-validate the data requested for the area according to the deep learning model, perform cross-computation between every two pieces of data requested for the area, and calculate the correlation between all data. The above-mentioned association method is only part of the association information, and it is necessary to rely on the deep learning model to complete the association of all data with all the requested locations included in the requested area.

[0046] Classify the data in the large-scale dataset according to the feature attributes to obtain the data categories corresponding to different requested locations.

[0047] Match the feature attribute of the data category requested for the current area with all locations in the current area through the deep learning model to generate the matching degree between the location and the feature attribute. A single piece of data may have multiple attributes. Calculate the features of all the data requested for the requested area uniformly to obtain the common feature attributes of most of the data among all the requested data.

[0048] According to the matching results between each location in the current area and the feature attributes, obtain the associated attributes of any location and the feature attributes. Further refine the feature attributes, and associate them according to all the requested locations in the requested area to obtain the feature attributes corresponding to the requested locations.

[0049] Match the associated attributes of all locations in the current area with the feature attributes in the overall data to generate an associated list of data request locations and data categories. According to the frequency of the data requested by any requested location, obtain the data represented by the high request frequency, and extract the common feature attributes from the data with high request frequency to complete data classification. Obtain the data requested by all requested locations to complete the classification of the data requested by all locations.

[0050] The data categories corresponding to different requested locations are not unique, and any one of the data can appear repeatedly in different categories. For example, the market data of a certain industry in Beijing can be included in the category represented by any enterprise in the same industry at a certain location in North China, and at the same time, any enterprise represented by a requested location in South China can also request this data and can also be included in the data category requested by the requested location where the enterprise is located.

[0051] Classify the request frequencies of any data at different requested locations, which can move the data to the storage node that is the closest in spatial distance to the requested location, greatly improving the data transmission rate and reducing the load of data transmission.

[0052] In response to a call request initiated for the data in the storage node, authenticate the user through the blockchain, and automatically obtain the access permission after successful authentication.

[0053] The specific method for identity authentication through the blockchain is:

[0054] Calculate the hash value based on the user identification code to obtain the user identity password, and determine the authenticity of the user identity password. If the user identity verification fails, access is denied. Blockchain is a chain block structure based on hash proof, that is, the data structure called blockchain. Utilizing the complexity and consensus of hash value calculation, it can excellently determine the user identity and is tamper-proof.

[0055] Use the blockchain method to carry the user's data request records, ensuring that the data can be traced every time a request is made and the records are retained. At the same time, the present invention also uses smart contracts to verify the user identity, ensuring the transparency and traceability of the user request records.

[0056] If the user identity verification passes, determine the user's data reading permission based on the user identity password.

[0057] The determination of the user's data reading permission based on the user identity password is carried out through smart contracts. Smart contract is one of the four core technologies in blockchain. Blockchain is a chain storage, tamper-proof, secure and trustworthy decentralized distributed ledger, which combines technologies such as distributed storage, peer-to-peer transmission, consensus mechanism, and cryptography. It records transactions and information through an ever-growing data block chain to ensure the security and transparency of data.

[0058] Smart contracts can eliminate the middleman and allow users to autonomously establish contracts entirely relying on technology. Smart contracts are also transparent and fair. Smart contracts will write the conditions clearly in code and record them on the blockchain. The whole process is executed by the program, and even the developers who write the smart contract code cannot tamper with it. Smart contracts have sufficient flexibility to allow users to freely establish contracts. Smart contract is one of the core technologies of blockchain and plays an executive role in blockchain.

[0059] As Figure 2 shown, after the identity verification passes, address verification is carried out between the data sender and the data receiver.

[0060] In the general trend of digital transformation, data has become the basis for activities such as enterprise daily office work, production and operation, technological innovation, and strategic development. Data security has become the basic guarantee for the healthy and stable development of digital enterprises. At present, data faces security risk challenges such as diverse transmission subjects, complex processing activities, upgraded attack means, and frequent internal leaks during the transmission process. Ensuring the security, integrity, and availability of data during the transmission process is of great significance for maintaining enterprise business continuity, protecting enterprise competitiveness and economic interests, and ensuring the safe transformation and sustainable and healthy development of enterprises.

[0061] The method for address verification between the data sender and the data receiver is as follows:

[0062] The data receiving end encrypts the communication address of the data receiving end multiple times. The character length of the communication address of the data receiving end gradually becomes longer after each encryption until the character length is the same as the preset number of digits of the current requesting user identity code after n times of encryption.

[0063] The data sending end encrypts the communication address of the data sending end multiple times. The character length of the communication address of the data sending end gradually becomes longer after each encryption until the character length is the same as the current requesting user's request time after m times of encryption. The user request time is the number of seconds from a certain fixed time. For example, it can be set as the number of seconds from 0:00 of the current day, or it can be the number of seconds from the previous whole hour. The user can freely set it.

[0064] Through multiple encryptions, the security of the data can be improved, avoiding problems such as third-party injection or communication address loss during address verification, and ensuring the security of the transmission channel during data transmission.

[0065] The data sending end and the data receiving end respectively transmit the encrypted text data and the number of their own encryptions to each other. If the data receiving end encrypts the data 5 times, the data receiving end sends the data 5 times to the data sending end, and the data sending end also uses this method to send data. Since the number of data encryptions will change according to the user identity and the user request time, the difficulty of cracking the data will also increase, better ensuring the security of the data transmission channel.

[0066] The data receiving end and the data sending end decrypt the data. If the decrypted data conforms to the communication address format, the verification passes. In each system, the format of the data is different. Even if someone cracks the encryption method mentioned in the present invention, the communication address format adopted in the present invention can also serve as the last line of defense to prevent the destruction of the data transmission channel.

[0067] The data encryption and decryption method can be:

[0068] The data receiving end encrypts the first text according to a preset rule to obtain a second text, and sends the second text to the data sending end.

[0069] The preset rule is to set two or more key lists K, each key has the same encrypted character length. The first text is segmented with a 32-byte character length, and the content of multiple 32-byte character lengths is encrypted using the keys in the user-preset order. The specific formula is:

[0070]

[0071] where n is the number of text segments after the first text is segmented, T i is the i-th text segment, K lFor the l-th key, where l is determined according to the user's preset order, multiple text segments are encrypted and then concatenated to obtain the encrypted result C, and the encrypted result C is the second text. By using multiple keys in a non-fixed order, it is possible to encrypt data with a length of every 32 bytes in the first text using different keys. If brute force cracking is used, the computational complexity increases exponentially, which better ensures the security of the data.

[0072] The data encryption and decryption methods can also be implemented using other existing technologies.

[0073] If the text parsed by the data receiving end conforms to the user's permission value for the requested data and the communication address rule of the current request location, the verification passes. When the key is correct, the text before encrypting the first text and after decrypting the second text should be in the same format. If any of the intermediate keys is incorrect, the content formats obtained from the second text and the first text are different. Therefore, it is only necessary to determine whether the format of the decrypted second text conforms to the format of the first text to complete the address verification.

[0074] After the address verification passes, the data sending end sends data to the data receiving end. The specific method is as follows:

[0075] The storage node where the current user requests data is the data sending end, which is responsible for sending the data from its own storage node to the data receiving end at the user's node.

[0076] During data transmission, corresponding processing is performed based on the address verification result.

[0077] After the address verification passes, the data sending end decrypts the communication address sent by the data receiving end.

[0078] The data receiving end sends the communication address of the data receiving end to the data sending end, and the data sending end sends the communication address of the data sending end to the data receiving end.

[0079] The data sending end directly sends information to the data receiving end based on the communication address of the data receiving end received by the data sending end.

[0080] The data sending end directly sends information to the data receiving end based on the communication address of the data receiving end received by the data sending end. In this way, the data transmission can be directly completed, avoiding attacks on the intermediate node due to third-party injection and other methods, and ensuring the security of the data.

[0081] At the same time, the data receiving end decrypts the communication address of the data sending end based on the data received by the data receiving end and receives the information sent by the data sending end.

[0082] During the data transmission process, a direct connection transmission method is adopted. After passing the verification at the data receiving end and the data sending end, the verification node sends the communication address, and the data is directly transmitted between the data receiving end and the data sending end, reducing the data transmission delay and improving the data transmission rate at the same time.

[0083] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A large-scale data distributed storage method, characterized in that: Included are: Classify large-scale data sets and adopt a distributed storage architecture, with each data category stored on different storage nodes; In response to a call request for data in a storage node, the user is authenticated through the blockchain and automatically obtains access rights after successful authentication; After the identity authentication is passed, address verification is performed between the data sender and the data receiver. After the address verification is passed, the data sender sends data to the data receiver.

2. A large-scale data distributed storage method according to claim 1, characterized in that: The specific method for classifying a large-scale data set is as follows: By reading the request location and request times of any data in a large-scale data set, the request frequency of data from different request locations can be obtained; Reading data whose request frequency of any requested location is greater than a preset threshold frequency, and extracting characteristic attributes from the data whose request frequency of the requested location is greater than the preset threshold frequency; Classify the data in large-scale data sets according to feature attributes to obtain data categories corresponding to different request locations.

3. A large-scale data distributed storage method according to claim 2, characterized in that: The data categories corresponding to the different request locations are not unique, and any data therein may appear repeatedly in different categories.

4. A large-scale data distributed storage method according to claim 1, characterized in that: The specific methods of identity authentication through blockchain are: Calculate the hash value based on the user identity code to obtain the user identity password, and determine the authenticity of the user identity password. If the user identity verification fails, access is denied. If the user identity authentication is passed, the user's permission to read data is determined based on the user identity password.

5. A large-scale data distributed storage method according to claim 4, characterized in that: The determination of the user's authority to read data based on the user's identity password is carried out through a smart contract.

6. A large-scale data distributed storage method according to claim 1, characterized in that: The method for performing address verification between the data sending end and the data receiving end is: The data receiving end encrypts the communication address of the data receiving end multiple times, and the character length of the communication address of the data receiving end gradually increases after each encryption, until the character length of the communication address of the data receiving end is the same as the preset number of digits of the identity code of the current requesting user after n encryptions; The data sending end encrypts the data sending end communication address multiple times, and the character length of the data sending end communication address gradually increases after each encryption, until after m encryptions, the character length is the same as the request time of the current requesting user; The data sending end and the data receiving end each transmit the encrypted text data and the number of times they have been encrypted to each other. The data receiving end and the data sending end decrypt the data. If the decrypted data conforms to the communication address format, the verification is successful.

7. A large-scale data distributed storage method according to claim 1 or 6, characterized in that: The method for the data transmitting end to send data to the data receiving end is: After the address verification is passed, the data sending end decrypts the communication address sent by the data receiving end; The data sending end directly sends information to the data receiving end according to the communication address of the data receiving end received by the data sending end; The data receiving end decrypts the communication address of the data sending end based on the data received by the data receiving end, and receives the information sent by the data sending end.