Hybrid data center data security storage scheduling method and device
Through a hybrid data center data security storage scheduling method, combining centralized and distributed data centers, and utilizing the IPFS protocol and erasure code mechanism, the data storage scheduling problem under unstable green energy power supply is solved, and the reliability and efficiency of data storage are achieved.
Patent Information
- Application Number
- CN202411404095.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-09
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2044-10-09
AI Technical Summary
In scenarios dominated by green energy, existing technologies cannot effectively schedule and allocate data storage between centralized data centers and distributed data centers, resulting in unstable data storage and complicated computing tasks.
A hybrid data center data security storage scheduling method is adopted. Through the IPFS protocol and erasure code mechanism, the advantages of centralized and distributed data centers are combined. The supply status of green energy is used to divide the data storage ratio and back up and restore data blocks, ensuring the reliability of the data center under unstable green energy conditions.
It achieves the reliability and efficiency of data storage tasks under the condition of unstable green energy power supply, reduces the consumption of traditional energy, and improves data recovery capabilities and system stability.
Smart Images

Figure CN119473131B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data storage scheduling, and in particular to a hybrid data center data security storage scheduling method and device. Background Art
[0002] With the advancement of digitalization, data centers are expanding in size, and their functions and structures are becoming increasingly diverse. To achieve low-carbon goals, data centers are also beginning to adopt green energy for power supply. However, green energy supply is often less stable than traditional thermal power. As computing needs become more complex, ensuring that data centers efficiently complete data storage and computing tasks while relying primarily on green energy has become a challenge for existing technologies.
[0003] Centralized and distributed data center architectures each have their advantages and disadvantages. Centralized data centers are more centralized, making data storage and maintenance relatively simple, but they also carry the risk of single points of failure. A single failure in a centralized data center can potentially cause the entire system to fail. Distributed data centers can use multiple replicas, tolerate multiple node failures, and offer automatic fault tolerance and failover. They are more suitable for large-scale systems with high reliability requirements, but they are more complex in terms of data storage and operation, requiring coordination and management of resources in different locations, making operations and maintenance more challenging.
[0004] Furthermore, as more distributed data center nodes are powered by renewable energy sources (such as hydropower, photovoltaics, and wind power), energy supply instability is significantly higher compared to traditional electricity. In the event of energy supply fluctuations, the operation of distributed data center nodes may become unstable, potentially leading to node offline situations. Taking into account other failure scenarios such as server load fluctuations, the scheduling of computing tasks will also become more complex and difficult.
[0005] In summary, in scenarios dominated by green energy, existing technologies cannot effectively schedule and allocate data storage between centralized data centers and distributed data centers, which urgently needs to be solved. Summary of the Invention
[0006] The present application provides a hybrid data center data security storage scheduling method and device to solve the problem that the existing technology cannot effectively schedule and allocate data storage between centralized data centers and distributed data centers in scenarios where green energy is the main focus.
[0007] The first aspect of the present application provides a hybrid data center data security storage scheduling method, including the following steps: obtaining a file retrieval request from a target user, determining at least one distributed data center node that stores multiple data blocks and multiple verification data blocks corresponding to the target file based on the file retrieval request, and determining the number of online nodes in the at least one distributed data center node, and judging whether the number of online nodes is less than the number of data blocks of the multiple data blocks; if the number of online nodes is less than the number of data blocks, sending a data block request to a preset centralized data center through a preset IPFS protocol, so that the centralized data center recovers the target file according to the data block request, and sends the target file to the target user; if the number of online nodes is greater than or equal to the number of data blocks, judging whether the at least one distributed data center node is all online, wherein, if the at least one distributed data center node is all online, obtaining the target file to send to the target user, otherwise recovering the target file according to the multiple verification data blocks and the preset erasure code mechanism, and sending the target file to the target user.
[0008] Optionally, in one embodiment of the present application, before determining the at least one distributed data center node that stores the multiple data blocks and the multiple verification data blocks corresponding to the target file, it also includes: splitting the target file into the multiple data blocks, and processing the multiple data blocks through the erasure code mechanism to obtain the multiple verification data blocks, and backing up the multiple data blocks in the centralized data center; based on the supply status of each of the preset multiple green energy sources, determining the data storage ratio of the distributed data center nodes corresponding to each of the preset multiple green energy sources; storing the multiple data blocks and the multiple verification data blocks in the distributed data center nodes corresponding to the preset distributed data center according to the data storage ratio.
[0009] Optionally, in one embodiment of the present application, before splitting the target file into the multiple data blocks, it also includes: establishing multiple data nodes that use the multiple green energy sources, and constructing the distributed data center based on the multiple data nodes and the IPFS protocol, wherein the multiple green energy sources include wind energy, solar energy, hydropower and pumped storage.
[0010] Optionally, in one embodiment of the present application, the sending of a data block request to a preset centralized data center through a preset IPFS protocol, so that the centralized data center restores the target file according to the data block request, includes: determining at least one offline data block based on the number of online nodes and the number of data blocks, and sending the data block request to the centralized data center according to the at least one offline data block and the IPFS protocol; in response to the data block request, obtaining the at least one offline data block among the multiple data blocks pre-backed up by the centralized data center, and combining the data blocks stored in the online distributed data center nodes among the at least one distributed data center node to restore the target file.
[0011] The second aspect of the present application provides a hybrid data center data security storage scheduling device, including: a judgment module, used to obtain a file retrieval request from a target user, to determine at least one distributed data center node that stores multiple data blocks and multiple verification data blocks corresponding to the target file according to the file retrieval request, and to determine the number of online nodes in the at least one distributed data center node, and to determine whether the number of online nodes is less than the number of data blocks of the multiple data blocks; a first recovery module, used to send a data block request to a preset centralized data center through a preset IPFS protocol if the number of online nodes is less than the number of data blocks, so that the centralized data center recovers the target file according to the data block request and sends the target file to the target user; a second recovery module, used to determine whether the at least one distributed data center node is all online if the number of online nodes is greater than or equal to the number of data blocks, wherein if the at least one distributed data center node is all online, the target file is obtained to be sent to the target user, otherwise the target file is recovered according to the multiple verification data blocks and the preset erasure code mechanism, and the target file is sent to the target user.
[0012] Optionally, in one embodiment of the present application, it further includes: a backup module, used to split the target file into the multiple data blocks before determining the at least one distributed data center node that stores the multiple data blocks and the multiple verification data blocks corresponding to the target file, and process the multiple data blocks through the erasure code mechanism to obtain the multiple verification data blocks, and back up the multiple data blocks in the centralized data center; a proportion division module, used to determine the data storage ratio of the distributed data center nodes corresponding to each of the preset multiple green energy sources based on the supply status of each of the preset multiple green energy sources; a storage module, used to store the multiple data blocks and the multiple verification data blocks in the distributed data center nodes corresponding to the preset distributed data center according to the data storage ratio.
[0013] Optionally, in one embodiment of the present application, it also includes: an establishment module for establishing multiple data nodes that use the multiple green energy sources before splitting the target file into the multiple data blocks, and constructing the distributed data center based on the multiple data nodes and the IPFS protocol, wherein the multiple green energy sources include wind energy, solar energy, hydropower and pumped storage.
[0014] Optionally, in one embodiment of the present application, the first recovery module includes: a determination unit, used to determine at least one offline data block based on the number of online nodes and the number of data blocks, and send the data block request to the centralized data center according to the at least one offline data block and the IPFS protocol; a combining unit, used to obtain the at least one offline data block among the multiple data blocks pre-backed up by the centralized data center in response to the data block request, and combine the data blocks stored in the online distributed data center nodes in the at least one distributed data center node to restore the target file.
[0015] The third aspect of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the program to implement the hybrid data center data security storage scheduling method as described in the above embodiment.
[0016] A fourth aspect of the present application provides a computer-readable storage medium, which stores a computer program that, when executed by a processor, implements the above hybrid data center data security storage scheduling method.
[0017] The fifth aspect of the present application provides a computer program product, including a computer program, which is executed to implement the above-mentioned hybrid data center data security storage scheduling method.
[0018] Therefore, the embodiments of the present application have the following beneficial effects:
[0019] The embodiment of the present application can obtain a file retrieval request from a target user to determine at least one distributed data center node storing multiple data blocks and multiple check data blocks corresponding to the target file according to the file retrieval request, and determine the number of online nodes in at least one distributed data center node, and determine whether the number of online nodes is less than the number of data blocks of the multiple data blocks; if the number of online nodes is less than the number of data blocks, then send a data block request to a preset centralized data center through a preset IPFS protocol, so that the centralized data center recovers the target file according to the data block request and sends the target file to the target user; if the number of online nodes is greater than or equal to the number of data blocks, then determine whether at least one distributed data center node is online, wherein if at least one distributed data center node is online, then obtain the target file and send it to the target user, otherwise, recover the target file according to multiple check data blocks and a preset erasure code mechanism, and send the target file to the target user. The present application supplies data center energy based on the distribution of green energy, combines the characteristics of "centralized" and "distributed" data center architectures, and fully utilizes green energy while ensuring the reliable completion of data storage tasks. Thus, it solves the problem that the existing technology cannot effectively schedule and allocate data storage between centralized data centers and distributed data centers in scenarios dominated by green energy.
[0020] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0022] Figure 1 This is a flowchart of a hybrid data center data security storage scheduling method provided according to an embodiment of the present application;
[0023] Figure 2 A schematic diagram of a hybrid data center framework provided for one embodiment of the present application;
[0024] Figure 3 A schematic diagram of the logical architecture of a hybrid data center data security storage scheduling method provided by one embodiment of the present application;
[0025] Figure 4 This is an example diagram of a hybrid data center data security storage scheduling device according to an embodiment of the present application;
[0026] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0027] Among them, 10-hybrid data center data security storage scheduling device, 100-judgment module, 200-first recovery module, 300-second recovery module, 501-memory, 502-processor, 503-communication interface. DETAILED DESCRIPTION
[0028] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0029] The following describes a hybrid data center data security storage scheduling method and device according to an embodiment of the present application with reference to the accompanying drawings. In response to the problems mentioned in the above background technology, the present application provides a hybrid data center data security storage scheduling method, in which a file retrieval request of a target user is obtained to determine at least one distributed data center node storing multiple data blocks and multiple verification data blocks corresponding to the target file according to the file retrieval request, and the number of online nodes in the at least one distributed data center node is determined, and it is judged whether the number of online nodes is less than the number of data blocks of the multiple data blocks; if the number of online nodes is less than the number of data blocks, a data block request is sent to a preset centralized data center through a preset IPFS protocol, so that the centralized data center recovers the target file according to the data block request and sends the target file to the target user; if the number of online nodes is greater than or equal to the number of data blocks, it is judged whether at least one distributed data center node is online, wherein if at least one distributed data center node is online, the target file is obtained to be sent to the target user, otherwise the target file is recovered according to the multiple verification data blocks and the preset erasure code mechanism, and the target file is sent to the target user. This application uses the distribution of green energy to power data centers, combining the characteristics of both centralized and distributed data center architectures to fully utilize green energy while ensuring the reliable completion of data storage tasks. This addresses the problem of existing technologies being unable to effectively schedule and allocate data between centralized and distributed data centers in scenarios where green energy is the primary source.
[0030] Specifically, Figure 1 This is a flowchart of a hybrid data center data security storage scheduling method provided in an embodiment of the present application.
[0031] like Figure 1 As shown, the hybrid data center data security storage scheduling method includes the following steps:
[0032] In step S101, a file retrieval request of a target user is obtained to determine, based on the file retrieval request, at least one distributed data center node storing multiple data blocks and multiple verification data blocks corresponding to the target file, and the number of online nodes in the at least one distributed data center node is determined, and it is determined whether the number of online nodes is less than the number of data blocks of the multiple data blocks.
[0033] The embodiments of the present application can first obtain the user's file retrieval request, and according to the file retrieval request, find and determine the distributed data center nodes where the n data blocks and c verification data blocks corresponding to the target file N are located, and determine the number m of online nodes in these distributed data center nodes, and judge whether the number of online nodes is less than the number of data blocks of multiple data blocks, thereby providing reliable guidance and basis for subsequent file retrieval and recovery.
[0034] Optionally, in one embodiment of the present application, before determining at least one distributed data center node for storing multiple data blocks and multiple verification data blocks corresponding to the target file, it also includes: splitting the target file into multiple data blocks, and processing the multiple data blocks through an erasure code mechanism to obtain multiple verification data blocks, and backing up the multiple data blocks in a centralized data center; based on the supply status of each of the preset multiple green energy sources, determining the data storage ratio of the distributed data center nodes corresponding to each green energy; and storing the multiple data blocks and multiple verification data blocks in the distributed data center nodes corresponding to the preset distributed data center according to the data storage ratio.
[0035] In actual implementation, existing distributed data center nodes face problems such as unstable green energy supply. The embodiments of this application can build a hybrid data center framework by combining centralized and distributed architectures. Specifically, the embodiments of this application can deploy distributed nodes near geographical locations with relatively abundant green energy, making full use of green energy for power supply and completing data storage and computing tasks; at the same time, the computing and storage advantages of centralized data centers are utilized to give full play to the stability of the power supply energy of centralized data centers, thereby ensuring that even if a node fails due to unstable power supply, the data storage and computing tasks can still be effectively completed.
[0036] Optionally, in one embodiment of the present application, before splitting the target file into multiple data blocks, it also includes: establishing multiple data nodes that use multiple green energy sources, and building a distributed data center based on the multiple data nodes and the IPFS protocol, wherein the multiple green energy sources include wind energy, solar energy, hydropower and pumped storage.
[0037] Specifically, if Figure 2As shown in the figure, the hybrid data center framework includes a centralized data center, a distributed data center based on the IPFS system, the Internet, traditional energy and green energy supply. In addition, the hybrid data center framework also includes a data recovery mechanism.
[0038] The main structure of a distributed data center is composed of multiple nodes that work with green energy; in terms of data storage, the embodiment of the present application can build a distributed data center based on the IPFS protocol, and each distributed node can process the access of file data in accordance with the IPFS protocol. In terms of energy supply, the embodiment of the present application can use wind energy, hydropower and solar energy as power supply sources according to the energy distribution, thereby making full use of local renewable energy, reducing carbon emissions, and realizing the green development of hybrid data centers. In addition, since green energy is affected by factors such as weather and time, the supply is not stable, and some distributed nodes may occasionally go offline. Therefore, the embodiment of the present application can further construct a data recovery mechanism based on the mechanism of the IPFS system to ensure that data recovery is performed to the greatest extent possible when a distributed node goes offline.
[0039] It's important to note that the hybrid data center framework includes a central data center (i.e., a centralized data center). This central data center's data nodes are powered by stable traditional energy sources and have data transmission links with distributed data centers. The centralized data center stores some data blocks of source files in the distributed data centers. If too many nodes in the distributed data centers go offline, resulting in severe data failure, the centralized data center will transfer some of the database back to the distributed data centers, allowing for effective file data recovery.
[0040] The source file data blocks in the distributed data nodes refer to a portion of the source file database within the distributed data center, and these data blocks are backed up in the centralized data center. In the data center structure improved based on IPFS, files are stored in blocks, and each distributed node has only one copy of the source file. However, the online state of distributed nodes is unstable, and a certain amount of data redundancy is required to cope with data failures and facilitate data recovery. Therefore, the embodiments of the present application can improve the data recovery mechanism by backing up a certain number of source file data blocks in a centralized data center.
[0041] It should be noted that in order to reduce data storage occupancy, the centralized data center only stores part of the source file's data blocks rather than all of it. These data blocks can be used by distributed data centers for data recovery.
[0042] To summarize, in a distributed data center, when a file is uploaded, it will be split into several data blocks. After being processed by the erasure coding mechanism, it will be distributed and stored in several distributed nodes according to a certain strategy. Another part of the data blocks of the file will be uploaded to the centralized data center via network transmission for data backup, thereby ensuring the integrity of data storage to the greatest extent possible when a distributed node is offline.
[0043] Therefore, the embodiments of the present application determine the online status of nodes and allocate data storage by combining the energy supply conditions of distributed nodes, thereby fully utilizing green energy while avoiding data failure caused by node offline.
[0044] To achieve stable operation of a hybrid data center, embodiments of the present application can classify data storage based on energy supply conditions to prevent data from becoming ineffective in a short period of time due to node offline. A centralized data center primarily serves as a centralized data storage and computing platform and is connected to distributed data center nodes via data transmission links. Centralized data center nodes can centrally receive data uploaded by distributed nodes and act as a data backup. Distributed data centers utilize green energy based on geographic location and store data using the IPFS protocol.
[0045] It should be noted that the embodiments of the present application mainly consider four green energy sources, namely "wind energy", "solar energy", "hydro energy" and "pumped storage energy" as energy sources for distributed data centers.
[0046] During the dry season, hydropower generation is low, and data centers powered by this type of energy cannot operate normally. Pumped storage is also reduced accordingly. In this case, the proportion of distributed node data stored in hydropower and pumped storage power is set to 0. At this time, only the distribution ratio of wind power, photovoltaic power, and the centralized data center are considered. When the distributed system can work normally, about 30% of the data blocks are stored in the centralized data center nodes.
[0047] During the flood season, the amount of hydropower generation is relatively high, and a larger proportion of the data storage tasks of the distributed data center are undertaken by the hydropower supply nodes; the pumped storage power station will also supply power to the distributed nodes, but considering the cost, the distribution ratio is not set too large; the embodiment of the present application can be set according to general climatic conditions and power plant deployment, to set the hydropower supply distributed nodes to store more than 50% of the data blocks, and the pumped storage function distributed nodes to store about 10% of the data blocks; on this basis, the distribution ratio of wind power, photovoltaic power and the distribution ratio of the centralized data center are considered.
[0048] Therefore, the embodiments of the present application coordinate centralized nodes according to the green energy supply situation, and collaboratively perform data storage and computing tasks through centralized data center nodes and distributed data center nodes to form an efficient and highly reliable green data center development model, and divide the energy supply conditions into proportions for data storage, thereby ensuring high reliability of data storage and reducing the consumption of traditional energy by centralized data centers.
[0049] In step S102, if the number of online nodes is less than the number of data blocks, a data block request is sent to a preset centralized data center through the preset IPFS protocol, so that the centralized data center recovers the target file according to the data block request and sends the target file to the target user.
[0050] In step S103, if the number of online nodes is greater than or equal to the number of data blocks, it is determined whether at least one distributed data center node is online. If at least one distributed data center node is online, the target file is obtained and sent to the target user. Otherwise, the target file is restored according to multiple verification data blocks and a preset erasure code mechanism, and the target file is sent to the target user.
[0051] Furthermore, if the number of online nodes m is less than the number of data blocks n, the embodiment of the present application can send a data block request to the centralized data center through the IPFS system of the distributed data center and restore the file N; otherwise, the embodiment of the present application can find out whether the distributed nodes where the n data blocks storing the file N are located are all online. If they are all online, the search for file N is completed, and the file N is transmitted to the user to complete the response to the user's data request. Otherwise, the existing m data blocks are used to complete the recovery of the file N according to the erasure code mechanism, and the file N is transmitted to the user.
[0052] Those skilled in the art should understand that the erasure code mechanism is a fault-tolerant technology widely used in distributed data storage systems. In a hybrid data center framework, combining the erasure code mechanism with the distributed data center structure can ensure that file data can still be accessed even if a distributed node fails or the network is interrupted. When a file data needs to be accessed and downloaded by a user, the distributed data center will search the internal nodes to find the complete file data; if the distributed nodes are offline and the file data blocks cannot be fully obtained, the distributed data center system will first perform data recovery internally; if too many distributed nodes are offline and the distributed data center system cannot perform data recovery independently, the centralized data center will provide a certain number of file data block backups, such as Figure 3 As shown, the user's file access request is completed.
[0053] Optionally, in one embodiment of the present application, a data block request is sent to a preset centralized data center through a preset IPFS protocol, so that the centralized data center restores the target file according to the data block request, including: determining at least one offline data block based on the number of online nodes and the number of data blocks, and sending a data block request to the centralized data center based on the at least one offline data block and the IPFS protocol; in response to the data block request, obtaining at least one offline data block from multiple data blocks pre-backed up by the centralized data center, and combining the data blocks stored in the online distributed data center nodes in at least one distributed data center node to restore the target file.
[0054] In the specific implementation process, for the missing at least (nm) data blocks (i.e., offline data blocks) that can restore the complete file N, the distributed data center sends a request to the centralized data center. The centralized data center transmits these data blocks to the distributed data center IPFS system, and performs data recovery based on the erasure code mechanism to obtain the complete file N, and transmits file N to the user to complete the response to the user's data request.
[0055] Therefore, the embodiments of the present application improve the recovery capability of data storage and reduce the possibility of external attacks when distributed data center nodes are offline by integrating the IPFS protocol, erasure coding mechanism and hybrid data center architecture.
[0056] To sum up, the embodiments of the present application use a hybrid structure data center framework to supply data center energy based on the distribution of green energy, and combine the characteristics of the two data center architectures of "centralized" and "distributed" to fully utilize green energy while ensuring the reliable completion of data storage tasks; the hybrid structure data center framework proposed in the embodiments of the present application adopts a data storage architecture and data recovery method based on the IPFS protocol to ensure that data can be safely stored even if the distributed nodes are offline; in addition, taking into account the energy supply conditions of the distributed data center nodes, the embodiments of the present application use a hybrid data center data storage and scheduling strategy based on green energy to allocate the data storage ratio of different nodes under different energy conditions, so that when the green energy power supply is unstable, the data center can fully utilize green energy while reliably completing the data storage scheduling task.
[0057] According to the hybrid data center data security storage scheduling method proposed in the embodiment of the present application, by obtaining the file retrieval request of the target user, at least one distributed data center node that stores multiple data blocks and multiple verification data blocks corresponding to the target file is determined according to the file retrieval request, and the number of online nodes in at least one distributed data center node is determined, and it is judged whether the number of online nodes is less than the number of data blocks of the multiple data blocks; if the number of online nodes is less than the number of data blocks, a data block request is sent to a preset centralized data center through a preset IPFS protocol, so that the centralized data center recovers the target file according to the data block request, and sends the target file to the target user; if the number of online nodes is greater than or equal to the number of data blocks, it is judged whether at least one distributed data center node is online, wherein if at least one distributed data center node is online, the target file is obtained to be sent to the target user, otherwise the target file is recovered according to multiple verification data blocks and a preset erasure code mechanism, and the target file is sent to the target user. The present application supplies energy to the data center based on the distribution of green energy, combines the characteristics of the two data center architectures of "centralized" and "distributed", and makes full use of green energy while ensuring the reliable completion of data storage tasks.
[0058] Secondly, a hybrid data center data security storage scheduling device proposed according to an embodiment of the present application is described with reference to the accompanying drawings.
[0059] Figure 4 It is a block diagram of a hybrid data center data security storage scheduling device according to an embodiment of the present application.
[0060] like Figure 4 As shown, the hybrid data center data security storage scheduling device 10 includes: a judgment module 100, a first recovery module 200 and a second recovery module 300.
[0061] Among them, the judgment module 100 is used to obtain a file retrieval request from a target user, to determine at least one distributed data center node that stores multiple data blocks and multiple verification data blocks corresponding to the target file according to the file retrieval request, and to determine the number of online nodes in the at least one distributed data center node, and to determine whether the number of online nodes is less than the number of data blocks of the multiple data blocks.
[0062] The first recovery module 200 is used to send a data block request to a preset centralized data center through a preset IPFS protocol if the number of online nodes is less than the number of data blocks, so that the centralized data center can recover the target file according to the data block request and send the target file to the target user.
[0063] The second recovery module 300 is used to determine whether at least one distributed data center node is online if the number of online nodes is greater than or equal to the number of data blocks. If at least one distributed data center node is online, the target file is obtained and sent to the target user. Otherwise, the target file is restored based on multiple verification data blocks and a preset erasure code mechanism, and the target file is sent to the target user.
[0064] Optionally, in one embodiment of the present application, the hybrid data center data security storage scheduling device 10 of the embodiment of the present application further includes: a backup module, a proportional division module and a storage module.
[0065] Among them, the backup module is used to split the target file into multiple data blocks before determining at least one distributed data center node that stores multiple data blocks and multiple verification data blocks corresponding to the target file, and process the multiple data blocks through the erasure code mechanism to obtain multiple verification data blocks, and back up the multiple data blocks in the centralized data center.
[0066] The proportion division module is used to determine the data storage ratio of the distributed data center nodes corresponding to each green energy based on the supply status of each green energy among the preset multiple green energy sources.
[0067] The storage module is used to store multiple data blocks and multiple verification data blocks in distributed data center nodes corresponding to preset distributed data centers according to the data storage ratio.
[0068] Optionally, in one embodiment of the present application, it also includes: an establishment module for establishing multiple data nodes that use multiple green energy sources before splitting the target file into multiple data blocks, and building a distributed data center based on the multiple data nodes and the IPFS protocol, wherein the multiple green energy sources include wind energy, solar energy, hydropower and pumped storage.
[0069] Optionally, in one embodiment of the present application, the first recovery module 200 includes: a determination unit and a combination unit.
[0070] Among them, the determination unit is used to determine at least one offline data block based on the number of online nodes and the number of data blocks, and send a data block request to the centralized data center based on the at least one offline data block and the IPFS protocol.
[0071] The combining unit is used to obtain at least one offline data block from a plurality of data blocks pre-backed up by the centralized data center in response to a data block request, and combine the data blocks stored in an online distributed data center node in at least one distributed data center node to restore the target file.
[0072] It should be noted that the above explanation of the embodiment of the hybrid data center data security storage scheduling method is also applicable to the hybrid data center data security storage scheduling device of this embodiment, and will not be repeated here.
[0073] According to the hybrid data center data security storage scheduling device proposed in the embodiment of the present application, it includes a judgment module for obtaining a file retrieval request from a target user, and determining at least one distributed data center node that stores multiple data blocks and multiple verification data blocks corresponding to the target file according to the file retrieval request, and determining the number of online nodes in at least one distributed data center node, and judging whether the number of online nodes is less than the number of data blocks of the multiple data blocks; a first recovery module for sending a data block request to a preset centralized data center through a preset IPFS protocol if the number of online nodes is less than the number of data blocks, so that the centralized data center recovers the target file according to the data block request and sends the target file to the target user; a second recovery module for judging whether at least one distributed data center node is online if the number of online nodes is greater than or equal to the number of data blocks, wherein if at least one distributed data center node is online, the target file is obtained to be sent to the target user, otherwise the target file is recovered according to multiple verification data blocks and a preset erasure code mechanism, and the target file is sent to the target user. The present application performs data center energy supply based on the distribution of green energy, combines the characteristics of the two data center architectures of "centralized" and "distributed", and makes full use of green energy while ensuring the reliable completion of data storage tasks.
[0074] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include:
[0075] Memory 501 , processor 502 , and computer programs stored in the memory 501 and executable on the processor 502 .
[0076] When the processor 502 executes the program, the hybrid data center data security storage scheduling method provided in the above embodiment is implemented.
[0077] Furthermore, the electronic device further includes:
[0078] The communication interface 503 is used for communication between the memory 501 and the processor 502 .
[0079] The memory 501 is used to store computer programs that can be run on the processor 502 .
[0080] The memory 501 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0081] If the memory 501, processor 502, and communication interface 503 are implemented independently, the communication interface 503, memory 501, and processor 502 can be connected to each other via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0082] Optionally, in a specific implementation, if the memory 501, the processor 502 and the communication interface 503 are integrated on a chip, the memory 501, the processor 502 and the communication interface 503 can communicate with each other through an internal interface.
[0083] The processor 502 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0084] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above hybrid data center data security storage scheduling method.
[0085] An embodiment of the present application also provides a computer program product, including a computer program, which, when executed, is used to implement the above-mentioned hybrid data center data security storage scheduling.
[0086] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0087] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of this application, "N" means at least two, for example, two, three, etc., unless otherwise specifically defined.
[0088] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing a custom logical function or process step, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed in a different order than shown or discussed, including performing functions in a substantially simultaneous manner or in a reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application pertain.
[0089] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or N wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program can be obtained electronically by optically scanning the paper or other medium and then editing, interpreting or processing it in other suitable ways as necessary, and then storing it in a computer memory.
[0090] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiment, the N steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0091] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0092] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0093] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present application. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A hybrid data center data security storage scheduling method, characterized in that: The following steps are involved: Obtaining a file retrieval request from a target user, determining, based on the file retrieval request, at least one distributed data center node storing a plurality of data blocks and a plurality of check data blocks corresponding to the target file, determining a number of online nodes in the at least one distributed data center node, and determining whether the number of online nodes is less than a number of data blocks of the plurality of data blocks; If the number of online nodes is less than the number of data blocks, a data block request is sent to a preset centralized data center through a preset IPFS protocol, so that the centralized data center recovers the target file according to the data block request and sends the target file to the target user; If the number of online nodes is greater than or equal to the number of data blocks, determining whether all of the at least one distributed data center nodes are online, wherein if all of the at least one distributed data center nodes are online, obtaining the target file and sending it to the target user; otherwise, restoring the target file based on the multiple check data blocks and a preset erasure coding mechanism, and sending the target file to the target user; Before determining the at least one distributed data center node storing the multiple data blocks and the multiple verification data blocks corresponding to the target file, the method further includes: Splitting the target file into the plurality of data blocks, processing the plurality of data blocks using the erasure coding mechanism to obtain the plurality of check data blocks, and backing up the plurality of data blocks in the centralized data center; Determining, based on the supply status of each of the preset multiple green energy sources, a data storage ratio of the distributed data center nodes corresponding to each of the green energy sources; The plurality of data blocks and the plurality of check data blocks are stored in distributed data center nodes corresponding to a preset distributed data center according to the data storage ratio.
2. The method according to claim 1, characterized in that Before splitting the target file into the multiple data blocks, the method further includes: Establish multiple data nodes that use the multiple green energy sources, and construct the distributed data center based on the multiple data nodes and the IPFS protocol, wherein the multiple green energy sources include wind energy, solar energy, hydropower and pumped storage.
3. The method according to claim 1, characterized in that The sending of a data block request to a preset centralized data center through a preset IPFS protocol, so that the centralized data center restores the target file according to the data block request, includes: Determining at least one offline data block based on the number of online nodes and the number of data blocks, and sending the data block request to the centralized data center according to the at least one offline data block and the IPFS protocol; In response to the data block request, at least one offline data block among the multiple data blocks pre-backed up by the centralized data center is obtained, and combined with the data blocks stored in the online distributed data center nodes among the at least one distributed data center node to restore the target file.
4. A hybrid data center data security storage scheduling device, characterized in that: include: a determination module, configured to obtain a file retrieval request from a target user, determine, based on the file retrieval request, at least one distributed data center node storing a plurality of data blocks and a plurality of check data blocks corresponding to the target file, determine the number of online nodes in the at least one distributed data center node, and determine whether the number of online nodes is less than the number of data blocks of the plurality of data blocks; A first recovery module is configured to send a data block request to a preset centralized data center through a preset IPFS protocol if the number of online nodes is less than the number of data blocks, so that the centralized data center recovers the target file according to the data block request and sends the target file to the target user; a second recovery module, configured to, if the number of online nodes is greater than or equal to the number of data blocks, determine whether all of the at least one distributed data center nodes are online; wherein, if all of the at least one distributed data center nodes are online, obtain the target file and send it to the target user; otherwise, restore the target file based on the multiple check data blocks and a preset erasure coding mechanism, and send the target file to the target user; The hybrid data center data security storage scheduling device further includes: a backup module, configured to split the target file into the plurality of data blocks before determining the at least one distributed data center node storing the plurality of data blocks and the plurality of check data blocks corresponding to the target file, process the plurality of data blocks using the erasure coding mechanism to obtain the plurality of check data blocks, and back up the plurality of data blocks in the centralized data center; A proportion division module is used to determine the data storage ratio of the distributed data center nodes corresponding to each of the preset multiple green energy sources based on the supply status of each green energy source; A storage module is used to store the multiple data blocks and the multiple verification data blocks in a distributed data center node corresponding to a preset distributed data center according to the data storage ratio.
5. The device according to claim 4, characterized in that The first recovery module includes: a determining unit, configured to determine at least one offline data block based on the number of online nodes and the number of data blocks, and send the data block request to the centralized data center according to the at least one offline data block and the IPFS protocol; A combining unit is used to obtain, in response to the data block request, at least one offline data block among the multiple data blocks pre-backed up by the centralized data center, and combine the data blocks stored in the online distributed data center nodes among the at least one distributed data center node to restore the target file.
6. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the hybrid data center data security storage scheduling method according to any one of claims 1 to 3.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the hybrid data center data security storage scheduling method as described in any one of claims 1 to 3.
8. A computer program product comprising a computer program, characterized in that The computer program is executed to implement the hybrid data center data security storage scheduling method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Data center hybrid energy storage system based on layering and energy consumption control method
CN106100050A
Multi-cloud fragmentation safety storage method and multi-cloud fragmentation safety storage system based on erasure coding
CN107154945A