Method, electronic device and computer program product for backing up data

By determining data tiering information in the main storage system and dividing appropriate storage tiers in the cloud server, the high cost problem caused by mixed data backup in cloud storage is solved, and storage resources are optimized while data security and availability are improved.

CN122195729APending Publication Date: 2026-06-12DELL PROD LP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DELL PROD LP
Filing Date
2024-12-09
Publication Date
2026-06-12

AI Technical Summary

Technical Problem

In existing technologies, cloud storage cannot effectively utilize the hierarchical information of data when backing up data, resulting in the mixed backup of all data, which increases unnecessary storage costs. Furthermore, traditional hard disk storage has limited expansion and high costs, and faces the risk of single point of failure.

Method used

By determining the data tiering information in the main storage system and dividing different storage tiers in the cloud server based on this information, backup data can be stored in the appropriate tier according to its importance and access frequency, thereby optimizing storage resource utilization and reducing costs.

Benefits of technology

It enables tiered backups based on data type and importance in cloud storage, optimizing storage resource utilization, reducing storage costs, improving data security and availability, reducing hardware maintenance costs, and enhancing disaster recovery capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122195729A_ABST
    Figure CN122195729A_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, an electronic device and a computer program product for backing up data. The method comprises determining hierarchical information of data stored in a primary storage system based on storage information of the data. The method further comprises determining target data for backup and a storage location corresponding to backup data corresponding to the target data based on the hierarchical information, the storage location being located in a cloud server in communication with the primary storage system. The method further comprises controlling the cloud server to store the backup data to the corresponding level in the cloud server based on the storage location. In this way, hierarchical information can be introduced in the backup process, so that the backup storage system can directly obtain the storage information of the data and store it in the corresponding level, thereby saving storage costs for users while ensuring the security and availability of the data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of communication technology, and more specifically, to methods, electronic devices, and computer program products for backing up data. Background Technology

[0002] As users utilize electronic devices, the amount of stored data grows exponentially with the expansion of their business. Specifically, the development of technologies such as cloud computing, the Internet of Things, social networks, and mobile internet has led to an explosive growth in the types and scale of data across various fields. In the storage, transmission, and exchange of data, it is crucial to ensure the security and reliability of large-scale data storage and transmission processes to prevent data loss or corruption. Summary of the Invention

[0003] Embodiments of this disclosure provide a method, electronic device, and computer program product for backing up data.

[0004] In a first aspect of this disclosure, a method for backing up data is provided. The method includes determining data hierarchical information based on storage information of data stored in a primary storage system. The method further includes determining, based on the hierarchical information, target data for backup and the storage location of backup data corresponding to the target data, the storage location being located in a cloud server communicating with the primary storage system. The method also includes, based on the storage location, controlling the cloud server to store the backup data at the corresponding hierarchical level within the cloud server.

[0005] In a second aspect of this disclosure, an electronic device is provided, including a processor; and a memory coupled to the processor, the memory having instructions stored therein, the instructions causing the electronic device to perform actions when executed by the processor, the actions including: determining data hierarchical information based on storage information of data stored in a main storage system; determining, based on the hierarchical information, target data for backup and the storage location corresponding to the backup data of the target data, the storage location being located in a cloud server communicating with the main storage system; and controlling the cloud server to store the backup data to the corresponding hierarchical level in the cloud server based on the storage location.

[0006] In a third aspect of this disclosure, a computer program product is provided, which is tangibly stored on a computer-readable medium and includes machine-executable instructions that, when executed, implement the method described in the first aspect of this disclosure.

[0007] It should be understood that the description in the Summary of the Invention section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0008] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0009] Figure 1 A schematic diagram of an example environment in which several embodiments of the present disclosure may be implemented is shown;

[0010] Figure 2 A schematic diagram of a process for backing up data according to some embodiments of the present disclosure is shown;

[0011] Figure 3 A schematic diagram illustrating the determination of hierarchical information of data according to some embodiments of the present disclosure is shown;

[0012] Figure 4 Schematic diagrams for restoring data according to some embodiments of the present disclosure are shown;

[0013] Figure 5 Schematic diagrams for backing up files according to some embodiments of the present disclosure are shown;

[0014] Figure 6 A workflow diagram of backup data according to some embodiments of this disclosure is shown;

[0015] Figure 7 A block diagram of an apparatus for backing up data according to some embodiments of the present disclosure is shown; and

[0016] Figure 8 A block diagram of a device that can implement several embodiments of the present disclosure is shown. Detailed Implementation

[0017] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0018] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0019] As mentioned above, data storage requires a significant amount of hardware and infrastructure, resulting in high investment costs and substantial physical space consumption. Furthermore, with the continuous growth of data volume, the maintenance and upgrading of hardware becomes a problem. Related technologies include backing up data in cloud storage to save on storage expenses. It's understandable that different types of data have significantly different business values ​​and access frequencies, leading to different storage media requirements. For example, compared to "cold" data such as historical archives and record data that require long-term storage and infrequent access and processing, "hot" data—which is accessed frequently and is more critical to business and applications—usually requires fast and efficient access and processing, thus placing higher demands on storage media performance. Therefore, creating different storage areas for different types of data in cloud storage becomes crucial.

[0020] Therefore, embodiments of this disclosure propose a method for backing up data. In embodiments of this disclosure, the method includes determining hierarchical information of the data based on storage information of data stored in a primary storage system. The method further includes determining, based on the hierarchical information, target data for backup and the storage location corresponding to the backup data, wherein the storage location is located in a cloud server communicating with the primary storage system. The method also includes, based on the storage location, controlling the cloud server to store the backup data at the corresponding level within the cloud server.

[0021] This approach allows for the integration of primary and backup storage systems. Data tiering is determined based on the primary storage system's information (storage tier, temperature, storage area, and relationships with other data blocks). This tiered information is preserved when backing up data using a cloud server, enabling the development of appropriate backup strategies and selection of suitable storage locations. This ensures data security and availability while saving users unnecessary costs.

[0022] Figure 1 A schematic diagram of an example environment 100 in which various embodiments of this disclosure may be implemented is shown. For example... Figure 1As shown, in example environment 100, there is a main storage system 102 (which may include, but is not limited to, server 102-1, client 102-2, etc.) that needs to back up data, target data 104-1, target data 104-2, and backup data 106. Example environment 100 also includes a cloud server 108 and a corresponding cloud storage space 110 that are communicatively connected to the main storage system 102. The main storage system 102 is not limited to server 102-1 and client 102-2 shown in the example environment, but may also include, but is not limited to, personal computers, server computers, handheld or laptop devices, mobile devices, multiprocessor systems, consumer electronics, wearable electronic devices, smart home devices, minicomputers, mainframe computers, edge computing devices, and distributed computing environments including any of the above systems or devices.

[0023] In some embodiments, cloud server 108 may be a service delivery model, allowing users to easily and on-demand access a shared pool of configurable computing resources (e.g., network, network bandwidth, servers, processing power, memory, storage, applications, virtual machines, and services) over a network. The cloud server may include one or more cloud storage spaces 110. Cloud server 108 can be used to store data that needs to be stored on client 102-2 or server 102-1. For example, cloud server 108 can receive and store data uploaded by client 102-2. Additionally, cloud server 108 can also support client 102-2 or server 102-1 downloading data stored on cloud server 108.

[0024] In some embodiments of this disclosure, the main storage system, such as user terminal 102-2, may include memory and storage media. A memory may be divided into multiple storage areas. There may be multiple memory units and storage media, and the types of these multiple memory units or storage media may be the same or different. For example, they may include, but are not limited to, main memory such as Random-Access Memory (RAM) and Static Random-Access Memory (SRAM), auxiliary memory, or cache memory. Storage media may include Hard Disk Drive (HDD), Solid State Drive (SSD), hard disk, and optical disk. It is understood that, since the importance of data varies depending on its type, multiple storage levels can be divided in the main storage system 102 to meet the storage needs of different data. For example, it may be divided into Level 1, Level 2, Level 3, etc., and each level may include multiple storage areas. Different storage levels correspond to different storage states. The storage state refers to the storage capacity (e.g., remaining capacity), read / write speed, and storage performance of the storage level.

[0025] In some embodiments of this disclosure, the primary storage system 102 can hierarchically store data based on its business value or importance using a tiered storage architecture. For example, data can be sorted according to the frequency of access by users and applications, and data in different order can be allocated to different tiers in the primary storage system 102. In one example, Full Automated Storage Tiering for Virtual Pools (FAST VP) can be used to monitor data access patterns within the system pools and dynamically match the performance requirements of the data determined by the access patterns with the storage or storage medium providing that performance level. For example, high-capacity Serial Attach SCSI (SAS) or NL-SAS can be used as the lower-level storage to reduce the cost of the storage system, while high-speed solid-state drives can be used as the higher-level storage. In some embodiments, when creating different storage tiers, the user can specify a tiering strategy that determines which storage tier of the primary storage system different data will be placed in. The tiering strategy can be a user-defined tiering principle, such as determining which storage tier of the primary storage system the data will be stored in based on the data's category, size, access frequency, confidentiality level, etc.

[0026] It is understandable that data loss or corruption can occur during storage, transmission, and exchange due to various reasons (such as hardware failure, data corruption, or malicious attacks). On the other hand, with increasing user activity, data is growing at an explosive rate, and the storage capacity and performance of the main storage system cannot meet user demands. Some data services are of high value and frequently accessed by users, crucial for business continuity and efficient operation. However, some data needs to be archived and stored long-term, and users rarely access it. Storing this data continuously in the main storage system not only wastes storage resources and causes storage system delays but also increases data storage costs. Therefore, data backup is particularly important. In some embodiments of this disclosure, data can be backed up from the main storage system 102 to the cloud server 108. It is understood that cloud storage offers high availability and scalability; users can adjust storage capacity according to their needs without installing additional hardware or incurring additional hardware maintenance costs. Furthermore, cloud storage also offers higher security and recoverability. When data is damaged or lost due to sudden hardware failure, natural disasters, or human sabotage in the main storage system, the data can be quickly retrieved or repaired from cloud storage, improving the disaster recovery and security performance of the main storage system.

[0027] like Figure 1 As shown, to ensure the reliability, reversibility, and security of data storage, the primary storage system 102 can back up data to a cloud server. The cloud server, as an auxiliary storage system, can be deployed in a PowerProtect Data Domain Virtual Edition (DDVE). DDVE's long-term retention feature helps users move infrequently accessed data from the data center to cloud storage to reduce costs. However, since data tiering information is stored in the primary storage system 102, the cloud server 108 cannot access this tiering information or determine the importance of the data to be backed up. Therefore, during the backup process, all data, whether important or unimportant, is mixed together, making tiered backup impossible. This results in users paying the same price for all data storage, increasing unnecessary storage costs.

[0028] Based on this, in some embodiments of this disclosure, after receiving the target data 104 for backup transmitted by the main storage system 102, the cloud server 108 can determine the storage location of the corresponding backup data 106 in the cloud storage space 110 according to the hierarchical information carried by the data. For example, for data with high access frequency and critical to business and applications, the data can be stored on a high-performance, low-latency storage tier, facilitating fast and efficient access and processing of this data. For data that needs to be stored for a long time but does not require frequent access and processing, the data can be stored on a lower-cost, higher-capacity storage tier.

[0029] By incorporating data tiering information into the data backup process, the cloud server used for backup can directly obtain data storage information, determine the data's storage location, and store it in the corresponding tier. This not only optimizes storage resource utilization but also significantly reduces storage costs. Backing up data using tiered information enhances the flexibility and speed of backup, ensuring data security and availability.

[0030] Figure 2 A flowchart of a method 200 for backing up data according to some embodiments of the present disclosure is shown. Method 200 can be... Figure 1 The cloud server 108 shown in the image is executing. Now refer to... Figure 2 This disclosure describes a method 200 for backing up data according to embodiments of the present disclosure. For ease of understanding, the specific examples mentioned in the following description are exemplary and not intended to limit the scope of protection of this disclosure. Figure 2 As shown in box 202, method 200 can determine the data hierarchical information based on the storage information of the data stored in the main storage system. The main storage system is the system used to store data on the client or server side. The data can be business data such as program data, log data, etc. For example, the data can include data generated by users when using various applications (database files, log files, configuration files, etc.).

[0031] In some embodiments, the primary storage system can be divided into multiple tiers, each corresponding to a different storage state, such as varying storage performance. It can be understood that storage tiers correspond to different storage ranges within the primary storage system; for example, the start and end positions of storage tier 1 differ from those of storage tier 2. Taking storage performance as an example, tiers can be divided into high-performance, low-performance, and medium-performance tiers. Different types of data can be stored at different tiers to meet the storage needs of different data types. For example, data that users do not frequently access can be stored in the low-performance tier; while data that users need to access frequently can be stored in the high-performance tier, thereby meeting the user's need for efficient access.

[0032] It is understood that a series of storage information is generated during the process of storing data in the main storage system, including the data file itself and its associated metadata. This includes, for example, the actual stored data content (including but not limited to text, images, audio, video, etc.), the data encoding method and organizational structure, the physical location of the data on the storage device (file path, disk sector, etc.), and data access control information (including but not limited to user permissions, role permissions, etc.). In some embodiments, the storage information may also include the storage level in the main storage system where the data resides and the data type or data category corresponding to that data.

[0033] In box 204, method 200 can determine the target data for backup and the storage location of the corresponding backup data based on hierarchical information. The storage location is located on a cloud server communicating with the main storage system. As mentioned above, the main storage system can monitor the access patterns of data and allocate data with different access frequencies to different storage tiers according to a preset hierarchical strategy. With users' use of various applications, web pages, etc., a large amount of new data is generated. Some data is frequently accessed and is crucial for business continuity and efficient operation; some data has a low access frequency but needs to be stored for a long time. It is understood that the main storage system has limited storage space and performance resources. To save storage resources while meeting user storage needs, some data, such as important data, can be stored in the main storage system, while other data, such as unimportant data, can be stored on the cloud server. Alternatively, to ensure the reliability and security of important data storage, some important data can be stored on the cloud server through backup. Furthermore, data in the main storage system may be lost due to various reasons (such as hardware failure, data corruption, or malicious attacks). In this scenario, all data in the main storage system can be backed up. The backup allows for the recovery of the original data after data loss, thus ensuring business continuity and data integrity.

[0034] Understandably, the expansion of traditional hard drive storage is typically limited by physical hardware. Increasing storage capacity may require purchasing additional hardware and involves complex installation and configuration processes. Furthermore, initial investment costs can be high, and users will incur even higher storage fees when facing large-capacity storage needs. On the other hand, traditional hard drive storage also faces the risk of single points of failure. If a natural disaster or human-caused damage results in hard drive damage or loss, data may be unrecoverable, causing significant losses to user access and business continuity. In contrast, cloud servers communicating with the primary storage system can be used for data backup. Cloud storage offers high scalability, allowing users to dynamically adjust storage capacity according to their needs, without being limited by physical hardware or incurring additional hardware maintenance costs. Cloud storage also provides more advanced data management functions and different tiers of storage services, allowing users to select different storage types based on data access frequency and importance, thereby optimizing costs. Furthermore, cloud storage provides stronger disaster recovery capabilities, ensuring rapid data recovery and resumption of business operations in the event of a disaster by storing data in multiple data centers across different geographical locations.

[0035] In some embodiments of this disclosure, to ensure backup of different types of data while conserving storage resources, different backup areas or backup levels can be defined in the cloud server. For example, backup level 1, backup level 2, and backup level 3, etc. Different backup levels have different storage performance and reliability, which can meet the backup needs of different data. It should be noted that the way the cloud server divides its levels and the resulting backup levels are similar to the way the main storage system divides its levels, and the resulting storage levels correspond to the backup levels. For example, storage level 1 and backup level 1 can be high-performance levels, used to store data with high access frequency or high business value.

[0036] Based on this, when the cloud server retrieves data to be backed up from the primary storage system, it can determine the corresponding backup level or backup area based on the hierarchical information, and store the data from different levels in the primary storage system into the corresponding backup levels on the cloud server. For example, data that is frequently accessed and has high business value can be stored in a backup level with an active layer, which typically has high performance and availability to ensure that the data can be accessed quickly. Data that is not frequently accessed but needs to be stored long-term can be stored in a backup level with an archive layer, which typically has lower cost and higher capacity to save storage costs. This categorized storage of data helps users save costs and improve storage efficiency.

[0037] In box 206, method 200 can control the cloud server to store backup data at the corresponding tier within the cloud server based on storage location. As described above, the data tiering information can indicate the storage area of ​​the data in the main storage system and also the corresponding storage location in the cloud server. In some embodiments, the cloud server can quickly determine the storage location of the data and store it at the corresponding tier based on the data tiering information according to a preset backup storage strategy.

[0038] In this way, data stratification information can be queried and transferred during the backup process. Based on this stratification information, cloud servers can quickly and accurately allocate different types of data to corresponding storage areas, optimizing the utilization of storage resources, preventing unimportant data from occupying high-performance storage space for extended periods, and reducing storage costs.

[0039] In some embodiments, the primary storage system may employ a FAST VP tiered strategy to monitor data access patterns within the system pool (including but not limited to tracking key metrics such as data read, write, and access frequencies). Based on the data access patterns, data is categorized into different performance requirement levels. According to these performance requirement levels, the data is stored in different tiers. Furthermore, the tier of the data, its corresponding performance requirement level, or other data information can be determined based on its storage location.

[0040] Figure 3 A schematic diagram illustrating the determination of hierarchical information of data according to some embodiments of the present disclosure is shown. For example... Figure 3 As shown, in the example environment 300, there is a primary storage system 302 and a backup storage system 304. The primary storage system 302 is responsible for storing user data and has an automatic data tiering function. The primary storage system 302 can determine the data tiering information based on the data characteristics and resource tiering strategies, and allocate the data to different storage areas according to the storage level indicated by the tiering information. The data characteristics mainly include access frequency, importance, and business value.

[0041] In some embodiments of this disclosure, access frequency thresholds can be preset according to actual tiering strategies or specific requirements. Based on this, the temperature category of data can be determined by comparing the data's access frequency with the preset access frequency threshold. For example, data with an access frequency higher than the preset access frequency threshold can be called hot data, while data with an access frequency lower than the preset access frequency threshold can be called cold data. It can be understood that data that requires frequent access and has high business value is hot data (e.g., real-time analytics data, website access logs, payment records on e-commerce platforms, etc.). This data requires fast response and efficient access, and therefore is usually stored on high-performance storage tiers, such as solid-state drives (SSDs) or high-performance SAS hard drives. Data with low access frequency or relatively low business value is cold data (e.g., archived log files, infrequently used application data, etc.). This data typically does not require fast response and can be stored on lower-cost storage tiers, such as NL-SAS hard drives, tape libraries, or cloud storage services.

[0042] To process large amounts of data quickly, the main storage system 302 can implement a tiered storage strategy, selecting appropriate storage locations based on data type. In some embodiments of this disclosure, the main storage system can calculate the data's "temperature" based on storage pool configuration information and I / O statistics, and allocate it to corresponding storage areas (e.g., hot data is stored in the high-performance layer 306-1, and cold data in the low-cost layer 306-2). Specifically, the main storage system can determine the data's activity level by tracking and statistically analyzing read / write I / O volumes. Based on the statistically obtained information and weighting the statistical data at a certain time frequency, the data's "temperature" (e.g., whether it is cold or hot data) can be determined. In some examples, data with higher activity levels can be considered more frequently accessed data and therefore can be regarded as hot data.

[0043] In some embodiments of this disclosure, the block layer 316 in the primary storage system 302 can determine the storage level of data based on storage information such as data access frequency and read / write patterns, and then add corresponding layering information to the data using the management mechanism of the file layer 318. For example, frequently accessed data can be marked as hot data, suitable for storage in the active layer; infrequently accessed data is marked as cold data, suitable for storage in the cloud layer or other lower-performance storage layers. In some embodiments, the layering information of data can be identified by extending the file attributes corresponding to the data. The backup storage system can directly determine the storage location of the backup data in the backup storage system from the data identification information. In other embodiments, a metadata manager or database can also be used to track and record the layering information of each data block, facilitating quick retrieval and updating of layering information during data backup. During the process of data being read from the primary storage system and sent to the backup storage system, layering information can also be transmitted by embedding layering information in the data path to ensure that the layering information can be correctly transmitted to the backup storage system along with the data. To ensure that the data can be transmitted correctly, the primary storage system 302 can migrate the data and the corresponding layering information to the backup storage system 304 through an appropriate protocol path 308.

[0044] In some embodiments of this disclosure, to efficiently manage and optimize storage resources while ensuring that the backup needs of various types of data are met, different backup areas or backup levels can be divided in the backup storage system. It is understood that different backup areas have different storage performance, which can meet the backup needs of different types of data. After receiving data, the backup storage system 304 can determine the temperature category of the data based on the data's stratification information and allocate it to the corresponding backup area. For example, data with high activity levels in the main storage system can be selected from a high-performance backup area (…). Figure 3 (Taking storage bucket 314-1 corresponding to active layer 312-1 as an example). For data with low access frequency in the main storage system, a large-capacity storage area can be selected as the backup storage location. Figure 3 (Taking storage bucket 314-2 corresponding to cloud layer 312-2 as an example). Cloud storage provides scalable storage space, as well as data redundancy and backup services, making it suitable for storage scenarios requiring high availability and flexibility. Choosing cloud storage as a backup storage system can significantly reduce storage costs and alleviate the pressure on users in terms of storage resource investment while ensuring data security and integrity.

[0045] It's understandable that data attributes are not static but change based on factors such as access frequency and business needs. For example, a data block might be "hot" data when first written because it's frequently accessed, but over time, access frequency may decrease, turning it into "cold" data. Similarly, a dataset that was previously rarely accessed might become important due to new business needs or analytical tasks, thus becoming "hot" data. Hot data typically requires more frequent backups and shorter restore times, while cold data can be backed up more frequently and restored with lower priority. However, as data's "hot" or "cold" attributes change, it may be necessary to migrate data from one storage tier to another. This migration ensures that data is always stored on the storage device best suited to its access frequency and value.

[0046] In some embodiments, to address changes in the hot / cold attributes of data, data can be monitored and analyzed, and the data's tiered information can be updated periodically. Further, a data migration strategy can be developed, and based on the updated tiered information obtained from the primary storage system 302, data can be promptly migrated from its current storage location to the storage location corresponding to the updated tiered information. For example, the backup storage system 304 retrieves the updated tiered information from the primary storage system 302, indicating that the current data has changed from cold data to frequently accessed hot data. Therefore, the backup storage system 304 migrates the data from bucket 314-2 corresponding to the current cloud layer 312-2 to bucket 314-1 corresponding to the higher-performance, more readily accessible active layer 312-1. During the backup process, backup and migration strategies need to be developed based on the hot / cold attributes of the data. By implementing these strategies and methods, hot and cold data can be effectively managed, storage efficiency improved, costs reduced, and evolving business needs met.

[0047] Figure 4 Schematic diagrams for data restoration according to some embodiments of the present disclosure are shown. For example... Figure 4 As shown, in example environment 400, there are backup storage system 402 and primary storage system 404. Backup storage system 402 is located on a cloud server, and its storage resources are cloud resources purchased on demand. Data restoration refers to restoring data stored in backup storage system 402 to its original location or a specified location in primary storage system 404.

[0048] It is understandable that data may be lost, damaged, or corrupted during its generation, transmission, storage, and application due to hardware failures, software errors, malicious attacks, or human error. Data restoration can recover damaged or distorted original data, thereby ensuring business continuity and data integrity, and reducing losses caused by data loss. When data restoration is required, a data restoration request is typically triggered automatically by the user or the primary storage system. In some embodiments, the restoration request sent by the primary storage system 404 specifies the data to be restored. Similar to the backup process, the restoration process is also based on the data's hierarchical information. The backup storage system 402 can determine which data needs to be restored based on the restoration request and determine the storage information corresponding to the data to be restored, such as which backup level the data belongs to.

[0049] For example, based on the restore request, file 1 to be restored can be determined to be high-value, frequently accessed data, and its original data is stored in the active layer 406-1 of backup storage system 402. File 2 to be restored can be determined based on the restore request to be stored in the cloud layer 406-2 of the backup storage system. Figure 4 As shown, data at different levels is stored in different storage areas, allowing the restoration process to be performed in parallel, thereby improving restoration efficiency. The backup storage processor 410 reads data from the corresponding storage area upon request. For example, data corresponding to file 1 is stored in storage area 408-1 corresponding to the active layer, and data corresponding to file 2 is stored in storage area 408-2 corresponding to the cloud layer.

[0050] In some embodiments, the backup storage system 402 transmits data to the primary storage system 404 through an appropriate protocol layer 412. Choosing a suitable protocol layer is crucial to ensuring data integrity and security. For example, for data requiring encrypted transmission, secure protocols such as HTTPS can be prioritized, providing encryption protection to prevent data theft or tampering. Furthermore, to optimize transmission efficiency and performance, specific protocol layer configurations may be selected based on data characteristics and transmission requirements, such as adjusting the transmission buffer size and setting timeout periods. It is understood that the protocol layer ensures that data is not corrupted or lost during transmission through a series of error detection and correction mechanisms. For sensitive or important data, the protocol layer provides encryption and authentication functions to ensure that data can only be accessed by authorized users or systems. In other embodiments, in addition to relying on protocol layer protection, other measures can be taken to enhance data integrity and security. For example, data checksums (such as hash values) can be used to verify data integrity. For particularly important data, redundant or backup transmission methods can also be considered to ensure rapid recovery even if errors occur during transmission.

[0051] In some embodiments of this disclosure, after receiving the data to be restored, the main storage system 404 can determine the storage location of the data in the storage pool 414 based on the hierarchical information at the file layer 416, and store the data in the corresponding storage location. For example, file 1 is stored in storage pool 414-1, which has a faster read / write speed and higher performance, based on the hierarchical information. Figure 4 (Using an SSD as an example), file 2 determines the storage location as 414-2 (which has a slower read / write speed but a larger storage capacity) based on the hierarchical information. Figure 4 (Example: HDD). It's understandable that to ensure data is correctly restored to its original storage location or specified location, it's necessary to verify the accuracy of the restoration result. For example, after data restoration, the restored file content needs to be compared with the original file content. This can be done by calculating the hash value of the original file, and then calculating the hash value of the restored file. If the two hash values ​​are the same, it means the file content was not tampered with or corrupted during the restoration process. For files with metadata, it's necessary to verify whether the metadata was correctly restored during the restoration process. Metadata may include file permission settings, owner information, etc. Once the data is verified to be complete and error-free, the restoration process ends, and the data can continue to be read and used.

[0052] In some embodiments of this disclosure, the Common Block File System (CBFS) in block layer 418 is also used in some data restoration processes. CBFS is responsible for managing data blocks and can divide data into fixed-size blocks for storage. For example, when storing a large number of files, CBFS can split the files into data blocks, which can be more flexibly allocated to different storage media (such as 414-1 (SSD Slices) and 414-2 (HDD Slices)), thereby optimizing the utilization of storage resources. CBFS can also employ data protection mechanisms, such as data redundancy and error correction codes. During data block storage, by adding redundant data blocks or error correction codes, data can be recovered even if some data blocks are damaged, ensuring data integrity and reliability.

[0053] In this way, in cases of data loss, corruption, or the need for migration or upgrades, data restoration operations can quickly recover data, ensuring its integrity, continuity, and security. Simultaneously, data types requiring restoration, such as critical business data, sensitive data, historical data, and backup data, can also be properly protected and managed.

[0054] Figure 5 Schematic diagrams for backing up files according to some embodiments of the present disclosure are shown. For example... Figure 5As shown, in example environment 500, there is file 502, primary storage system 506, and backup storage system 508. It's understandable that large files often exhibit complex and diverse data access characteristics. Taking large media files as an example, during video playback, the beginning section may have a high access frequency due to frequent user previews and jumps to specific segments, making it "hot data." The large amount of regular plot content in the middle has a relatively low access frequency in daily access and can be classified as "cold data." For large enterprise-level design drawings, frequently modified areas during the design process are frequently accessed during the project, making them "hot data," while the access frequency of the established basic framework decreases sharply afterward, becoming "cold data." If the entire file is stored uniformly as hot data, storing the entire file on high-cost, high-performance storage media to meet the access speed requirements of hot data would waste storage resources and drastically increase costs. If the entire file is stored on low-cost media to save costs, it would lead to excessively high latency in accessing the hot data portion, impacting business efficiency.

[0055] In some embodiments, to achieve a balance between storage performance and cost, a preset file size threshold can be pre-set based on actual storage or backup needs. If the file size of file 502 exceeds the preset threshold, file 502 can be split into multiple sub-files (such as sub-file 504). Different sub-files are stored at different levels of the main storage system based on the access frequency or availability of the data contained in sub-file 504. For example, in the aforementioned large video file, the header portion, frequently previewed and accessed by users, is marked as hot data (stored in a high-performance storage layer), while the less frequently accessed, regular content in the middle is marked as cold data (stored in a large-capacity storage layer). Based on this, the storage information corresponding to file 502 includes the storage location and storage level of each sub-file 504. This storage information is sent to a cloud server, which can then use this information to store the file 502, which needs to be backed up, into multiple sub-files at different backup levels on the cloud server.

[0056] Figure 6 A workflow diagram of backup data according to some embodiments of this disclosure is shown. Figure 6 As shown in box 602, the cloud server can receive a backup request instruction from the primary storage system. This backup request instruction may include the data to be backed up in the primary storage system and whether the data contains tiered information. In box 604, in response to the backup request instruction, the cloud server can determine whether the primary storage system is divided into multiple storage tiers, and whether different types of data, such as hot and cold data, are stored in different storage tiers. In box 606, if it is determined that the primary storage system supports tiered data storage, it can determine whether to use a tiered backup method for the backup data.

[0057] like Figure 6 As shown in box 608, if it is determined that the primary storage system does not support tiered data storage, the data stored in the primary storage system can be backed up to the cloud server according to a preset backup frequency. For example, it can be backed up to DDVE. In box 610, if it is determined that the primary storage system supports tiered data storage and tiered backup is used for the backup data, the tiering information (i.e., the storage tier where the data is stored and the hot / cold category of the data) corresponding to the data can be obtained from the primary storage system. It can be understood that in box 612, this tiering information can be carried by the data itself; for example, the storage tier and hot / cold category of the data can be carried by extending the file attributes of the data.

[0058] In some embodiments, at block 614, a cloud server such as DDVE can determine the tier information corresponding to the data based on the file's extended attributes. For example, the data may be cold data stored in storage tier S3. Based on this information, backup data corresponding to the data can be stored in backup tier S3 of the cloud server.

[0059] In this way, different backup methods can be selected according to different situations, thus making data backup more applicable and the backup methods more flexible.

[0060] Figure 7 A block diagram of an apparatus 700 for backing up data according to some embodiments of the present disclosure is shown. Figure 7 As shown, the device includes a hierarchical information determination unit 702, configured to determine the hierarchical information of the data based on the storage information of the data stored in the main storage system. The device 700 also includes a storage location determination unit 704, configured to determine, based on the hierarchical information, the target data to be backed up and the storage location corresponding to the backup data, wherein the storage location is located in a cloud server communicating with the main storage system. The device 700 further includes a backup storage unit 706, configured to, based on the storage location, control the cloud server to store the backup data at the corresponding level within the cloud server.

[0061] In some embodiments, the apparatus 700 further includes a data restoration unit configured to restore the backup data to the corresponding level in the main storage system based on the hierarchical information of the data.

[0062] In some embodiments, the data restoration unit is further configured to: determine, based on the hierarchical information, the storage location of the backup data to be restored on the cloud server and the storage location on the main storage system; obtain the backup data from the cloud server based on the storage location of the backup data to be restored on the cloud server; and store the backup data in the main storage system based on the storage location of the backup data on the main storage system.

[0063] In some embodiments, the hierarchical information determination unit 702 is further configured to: determine the hierarchical information of the data based on the data characteristics of the data and the hierarchical strategy of the main storage system.

[0064] In some embodiments, the tiering information determination unit 702 is further configured to: determine the temperature category of the data based on storage configuration information and data access information, the temperature category being used to characterize the activity level of the data; and determine a tiering strategy for the main storage system based on the temperature category of the data, the tiering strategy including determining the storage location of the data in the main storage system based on the temperature category of the data.

[0065] In some embodiments, the hierarchical information determination unit 702 is further configured to: determine the temperature category of the data based on a comparison result between the access frequency of the data and a preset access frequency threshold; determine the hierarchical information of the data based on the temperature category of the data and allocate a corresponding storage area for the data.

[0066] In some embodiments, the storage location determination unit 704 is further configured to: allocate a corresponding storage area for the data in the cloud server based on the hierarchical information of the data, wherein the storage area includes a first storage area and a second storage area, the first storage area is used to store data with a temperature level greater than a preset temperature level, and the second storage area is used to store data with a temperature level less than a preset temperature level.

[0067] In some embodiments, the hierarchical information determination unit 702 is further configured to: determine corresponding attribute information based on the hierarchical information of the data; and expand the file attributes corresponding to the data based on the attribute information, wherein the expanded file attributes include the hierarchical information of the data.

[0068] In some embodiments, the apparatus 700 further includes a data migration unit configured to: in response to a change in the hierarchical information of the data, obtain updated hierarchical information of the data from the main storage system; and based on the updated hierarchical information, migrate the data from the current storage location to the storage location corresponding to the updated hierarchical information.

[0069] In some embodiments, the storage location determination unit 704 is further configured to: determine whether the file size corresponding to the data is greater than a preset file size threshold; in response to the file size being greater than the preset file size threshold, divide the file corresponding to the data into a preset number of sub-files based on a preset number of divisions; and determine the storage location of the sub-files in the cloud server based on the file location of the sub-files and the hierarchical information corresponding to the data.

[0070] It is understood that by utilizing the apparatus 700 of this disclosure, at least one of the many advantages that can be achieved by the methods or processes described above can be realized.

[0071] Figure 8 A schematic block diagram of an example device 800 that can be used to implement embodiments of the present disclosure is shown. As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 802 or loaded from storage unit 808 into random access memory (RAM) 803. Various programs and data required for the operation of device 800 may also be stored in RAM 803. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.

[0072] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0073] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as method 200. For example, in some embodiments, method 200 may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of method 200 described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform method 200 by any other suitable means (e.g., by means of firmware).

[0074] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload programmable logic devices (CPLDs), and so on.

[0075] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0076] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing. Furthermore, although operations are depicted in a specific order, this should be understood as requiring that such operations be performed in the specific order shown or in sequential order, or requiring that all illustrated operations be performed to achieve the desired result. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the foregoing discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented individually or in any suitable sub-combination in multiple implementations.

[0077] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A method for backing up data, comprising: Based on the storage information of the data stored in the main storage system, the hierarchical information of the data is determined; Based on the hierarchical information, the target data for backup and the storage location corresponding to the backup data are determined, wherein the storage location is located in a cloud server that communicates with the main storage system; as well as Based on the storage location, the cloud server is controlled to store the backup data in the corresponding layer of the cloud server.

2. The method according to claim 1, further comprising: Based on the hierarchical information of the data, the backup data is restored to the corresponding level in the main storage system.

3. The method according to claim 2, further comprising restoring the backup data to the corresponding level in the main storage system: Based on the hierarchical information, the storage location of the backup data to be restored on the cloud server and the storage location on the main storage system are determined. Based on the storage location of the backup data to be restored on the cloud server, the backup data is retrieved from the cloud server; Based on the storage location of the backup data in the main storage system, the backup data is stored in the main storage system.

4. The method according to claim 1, wherein determining the hierarchical information of the data includes: Based on the data characteristics of the data and the tiering strategy of the main storage system, the tiering information of the data is determined.

5. The method according to claim 4, further comprising: Based on storage configuration information and data access information, the temperature category of the data is determined, and the temperature category is used to characterize the activity level of the data. Based on the temperature category of the data, a tiering strategy for the main storage system is determined, wherein the tiering strategy includes determining the storage location of the data in the main storage system based on the temperature category of the data.

6. The method according to claim 4, wherein determining the hierarchical information of the data further comprises: Based on the comparison between the access frequency of the data and the preset access frequency threshold, the temperature category of the data is determined; Based on the temperature category of the data, the hierarchical information of the data is determined and the corresponding storage area is allocated to the data.

7. The method of claim 6, wherein determining the hierarchical information of the data and allocating corresponding storage areas for the data comprises: Based on the hierarchical information of the data, a corresponding storage area is allocated for the data in the cloud server. The storage area includes a first storage area and a second storage area. The first storage area is used to store data with a temperature level greater than a preset temperature level, and the second storage area is used to store data with a temperature level less than the preset temperature level.

8. The method according to claim 1, further comprising: Based on the hierarchical information of the data, the corresponding attribute information is determined; as well as Based on the attribute information, the file attributes corresponding to the data are extended, and the extended file attributes include the hierarchical information of the data.

9. The method according to claim 1, further comprising: In response to a change in the hierarchical information of the data, the updated hierarchical information of the data is obtained from the main storage system; Based on the updated hierarchical information, the data is migrated from the current storage location to the storage location corresponding to the updated hierarchical information.

10. The method according to claim 1, further comprising: Determine whether the file size corresponding to the data is greater than a preset file size threshold; In response to the file size being greater than a preset file size threshold, the file corresponding to the data is divided into a preset number of sub-files based on a preset number of segments; as well as Based on the file location corresponding to the sub-file and the hierarchical information corresponding to the data, the storage location of the sub-file in the cloud server is determined.

11. An electronic device, comprising: processor; as well as A memory coupled to the processor, the memory having instructions stored therein, the instructions which, when executed by the processor, cause the electronic device to perform actions, the actions including: Based on the storage information of the data stored in the main storage system, the hierarchical information of the data is determined; Based on the hierarchical information, the target data for backup and the storage location corresponding to the backup data are determined, wherein the storage location is located in a cloud server communicating with the main storage system; and Based on the storage location, the cloud server is controlled to store the backup data in the corresponding layer of the cloud server.

12. The device according to claim 11, wherein the action further includes: Based on the hierarchical information of the data, the backup data is restored to the corresponding level in the main storage system.

13. The device according to claim 12, wherein the action further includes: Based on the hierarchical information, the storage location of the backup data to be restored on the cloud server and the storage location on the main storage system are determined. Based on the storage location of the backup data to be restored on the cloud server, the backup data is retrieved from the cloud server; Based on the storage location of the backup data in the main storage system, the backup data is stored in the main storage system.

14. The device of claim 11, wherein determining the layering information of the data includes: Based on the data characteristics of the data and the tiering strategy of the main storage system, the tiering information of the data is determined.

15. The device of claim 14, wherein the layering strategy comprises: Based on storage configuration information and data access information, the temperature category of the data is determined, and the temperature category is used to characterize the activity level of the data. Based on the temperature category of the data, determine the tiering strategy for the main storage system. The tiered strategy includes determining the storage location of the data in the main storage system based on the temperature category of the data.

16. The device of claim 14, wherein determining the layering information of the data further includes: Based on the comparison between the access frequency of the data and the preset access frequency threshold, the temperature category of the data is determined; Based on the temperature category of the data, the hierarchical information of the data is determined and the corresponding storage area is allocated to the data.

17. The device of claim 16, wherein determining the hierarchical information of the data and allocating corresponding storage areas for the data comprises: Based on the hierarchical information of the data, a corresponding storage area is allocated for the data in the cloud server. The storage area includes a first storage area and a second storage area. The first storage area is used to store data with a temperature level greater than a preset temperature level, and the second storage area is used to store data with a temperature level less than the preset temperature level.

18. The apparatus of claim 11, further comprising: Based on the hierarchical information of the data, the corresponding attribute information is determined; as well as Based on the attribute information, the file attributes corresponding to the data are extended, and the extended file attributes include the hierarchical information of the data.

19. The apparatus of claim 11, further comprising: In response to a change in the hierarchical information of the data, the updated hierarchical information of the data is obtained from the main storage system; Based on the updated hierarchical information, the data is migrated from the current storage location to the storage location corresponding to the updated hierarchical information.

20. A computer program product tangibly stored on a computer-readable medium and comprising machine-executable instructions that, when executed, cause a machine to perform actions, the actions comprising: Based on the storage information of the data stored in the main storage system, the hierarchical information of the data is determined; Based on the hierarchical information, the target data for backup and the storage location corresponding to the backup data are determined, wherein the storage location is located in a cloud server that communicates with the main storage system; as well as Based on the storage location, the cloud server is controlled to store the backup data in the corresponding layer of the cloud server.