Method, device, system and storage medium for data backup and source data access

By backing up the source data of the main source station in different public cloud object storage spaces and setting source back rules in high-concurrency scenarios, the storage and bandwidth usage problems of high-availability architectures under high-concurrency access in the existing technology are solved, and the cost and operation and maintenance benefits are improved.

CN112099991BActive Publication Date: 2025-05-23BEIJING KINGSOFT CLOUD NETWORK TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010922994.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-04
Publication Date
2025-05-23
Estimated Expiration
2040-09-04

AI Technical Summary

Technical Problem

When the existing massive source data backup method maintains a high availability architecture under high concurrent access, it occupies a lot of storage resources and bandwidth resources, is costly and has difficulty in operation and maintenance.

Method used

At least two different public cloud object storage spaces are used to backup the source data in the main source station separately. Each backup source station backs up part of the source data, and the return destination address when the requested data is not included is another backup source station or main source station.

Benefits of technology

It avoids high storage and bandwidth usage caused by full backup, reduces cost and operation and maintenance difficulties, and at the same time realizes a high availability architecture under high concurrency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112099991B_ABST
    Figure CN112099991B_ABST
Patent Text Reader

Abstract

The present application relates to a method, device, system and storage medium for data backup and access to source data, the method comprising: reading source data stored in a main source station; backing up the read source data to at least two first backup source stations, each of the at least two first backup source stations using a different public cloud object storage space; wherein each of the first backup source stations backs up part of the source data in the main source station; the first backup source station's return-to-source destination address when it does not contain requested data is other backup source stations or the main source station. The present application is used to solve the problems of maintaining a high-availability architecture under high concurrent access in existing methods of backing up massive source data, occupying a lot of storage resources and bandwidth resources, high cost, and difficult operation and maintenance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer network technology, and in particular to a method, device, system, and storage medium for backing up data and accessing source data. Background Art

[0002] There are about 800 million Internet users in China. In the Internet age, everyone is a creator of information. Hundreds of PB of data are generated every hour around the world. Especially in the Internet industry, major live broadcast and on-demand services mainly used for video content sharing are developing rapidly. The massive amount of data read by users poses a huge challenge to the service capabilities of the source station. How to cope with the high-concurrency access of hundreds of millions of customers, reasonable load balancing strategies and high-availability architectures are essential. At the same time, they must cope with the peaks and troughs that are common in the business, ensuring customer experience while reducing costs as much as possible.

[0003] The current solutions to this problem are roughly as follows:

[0004] Method 1: Video on demand companies build their own source sites to store data, prepare sufficient resources for periodic business outbreaks, expand vertically, and purchase more servers and bandwidth. In order to ensure the high availability of the source site in high-concurrency scenarios, multiple sites must be built to achieve master-slave switching and perform full backup of the source site. This solution obviously wastes resources and has high construction costs. In order to achieve dynamic load balancing, it is also necessary to build a content distribution network (CDN) on its own, which poses a challenge to operation and maintenance, and the user experience cannot be guaranteed. In addition, the high-availability architecture of self-built storage for multiple source sites is too expensive, which is not conducive to the long-term development of the enterprise and is not flexible and reliable enough to cope with changes in business peaks.

[0005] Method 2: Video on demand companies can build their own primary and backup source sites for full storage, or build their own primary source sites and store all source data in public cloud objects. In addition, CDN edge node cache can be used to reduce the pressure on the source site under high concurrent access and improve customer access speed. However, the full backup method requires a lot of storage capacity and bandwidth to maintain a high-availability architecture. Summary of the invention

[0006] The present application provides a method, device, system and storage medium for data backup and source data access, which are used to solve the problems of maintaining a high-availability architecture under high concurrent access in the existing massive source data backup method, occupying a lot of storage resources and bandwidth resources, high cost and difficult operation and maintenance.

[0007] In a first aspect, an embodiment of the present application provides a data backup method, including:

[0008] Read the source data stored in the primary source station;

[0009] Backing up the read source data to at least two first backup source stations, where each of the at least two first backup source stations uses a different public cloud object storage space;

[0010] Among them, each of the first backup source stations backs up part of the source data in the main source station; the return-to-source destination address of the first backup source station when it does not contain the requested data is other backup source stations or the main source station.

[0011] Optionally, after reading the source data stored in the primary source station, the method further includes:

[0012] Backing up the read source data to at least one second backup source station, wherein the second backup source station stores source data of a volume not less than a preset proportion of data in the primary source station;

[0013] Among them, the return-to-source destination address of the second backup source station is the primary source station, and the return-to-source destination address of the first backup source station is the second backup source station.

[0014] Optionally, the sum of the amount of source data stored in each of the second backup source stations is equal to the total amount of source data in the primary source station.

[0015] Optionally, the method further comprises:

[0016] Determine a back-to-source priority for each of the first backup source stations, and configure a back-to-source rule for the first backup source station according to the back-to-source priority.

[0017] Optionally, configuring a back-to-source rule of the first backup source station according to the back-to-source priority includes:

[0018] The back-to-source rule for configuring the first backup source station is: the back-to-source destination address is the address of the first backup source station with a high priority, or the address of the primary source station.

[0019] Optionally, the types of source data stored in different first backup source stations are different, and / or the types of source data stored in different second backup source stations are different.

[0020] Optionally, the types of the source data are divided according to the access popularity of the source data.

[0021] Optionally, the correspondence between the address of the first backup source station and the type of the source data is saved in a central node of the content distribution network.

[0022] In a second aspect, an embodiment of the present application provides a method for accessing source data, wherein the source data is backed up using the data backup method described in the first aspect, and the method is applied to any of the first backup source stations, including:

[0023] Get data access request;

[0024] Determining whether to back up the source data requested by the data access request;

[0025] If so, return the source data;

[0026] Otherwise, according to the return-to-source destination address, the source data is obtained from other backup source stations or the main source station, and the source data is returned.

[0027] Optionally, obtaining the source data from another backup source station according to the return-to-source destination address includes:

[0028] According to the return-to-source destination address, the source data is obtained from the second backup source station; wherein, when the second backup source station has not backed up the source data, the source data is obtained from the primary source station and returned;

[0029] or,

[0030] According to the return-to-source destination address, the source data is obtained from the first backup source station with a high priority, wherein when the first backup source station with a high priority has not backed up the source data, the source data is obtained from the primary source station and returned.

[0031] In a third aspect, an embodiment of the present application provides a system for accessing source data, wherein the source data is backed up using the data backup method described in the first aspect, and the system includes at least two first backup source stations and a primary source station;

[0032] Any of the first backup source stations is used to obtain a data access request, determine whether to back up the source data requested by the data access request, and if so, return the source data; otherwise, obtain the source data from other backup source stations or the primary source station according to the back-to-source destination address, and return the source data;

[0033] The primary source station is used to provide source data for the first backup source station or the other backup source stations.

[0034] Optionally, the system further comprises: at least one second backup source station, wherein the second backup source station stores source data of a volume not less than a preset proportion of data in the primary source station;

[0035] The first backup source station is used to obtain the source data from the second backup source station according to the return-to-source destination address and return the source data;

[0036] The second backup source station is used to obtain the source data from the primary source station and return it to the first backup source station when the source data requested by the first backup source station is not included.

[0037] Optionally, the system further comprises a central node and an edge node of a content distribution network;

[0038] The edge node is used to obtain the data access request submitted by the client, and when the source data requested by the data access request is cached, return the source data to the client; when the source data requested by the data access request is not cached, forward the data access request to the central node;

[0039] The central node is used to obtain the data access request forwarded by the edge node, obtain the address of the first backup source station corresponding to the type of source data requested by the data access request, and forward the data access request according to the address of the first backup source station.

[0040] In a fourth aspect, an embodiment of the present application provides a data backup device, including:

[0041] A reading module, used to read source data stored in the primary source station;

[0042] A backup module, used to back up the read source data to at least two first backup source stations, wherein the at least two first backup source stations each use a different public cloud object storage space;

[0043] Among them, each of the first backup source stations backs up part of the source data in the main source station; the return-to-source destination address of the first backup source station when it does not contain the requested data is other backup source stations or the main source station.

[0044] In a fifth aspect, an embodiment of the present application provides a data backup system, including: a primary source station and at least two first backup source stations, wherein the at least two first backup source stations each use a different public cloud object storage space;

[0045] The primary source station is used to store source data;

[0046] The first backup source station is used to back up part of the source data read from the main source station. When the first backup source station does not contain the requested data, the return-to-source destination address is other backup source stations or the main source station.

[0047] In a sixth aspect, an embodiment of the present application provides an electronic device, including: a processor and a memory;

[0048] The memory is used to store computer programs;

[0049] The processor is used to execute the program stored in the memory to implement the data backup method described in the first aspect, or to implement the method for accessing source data described in the second aspect.

[0050] In a seventh aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the data backup method described in the first aspect, or implements the method for accessing source data described in the second aspect.

[0051] The above technical solution provided by the embodiment of the present application has the following advantages over the prior art: the method provided by the embodiment of the present application sets a main source station for storing source data, and uses at least two different public cloud object storage spaces to back up the source data in the main source station in different public cloud object storage spaces, respectively, to obtain at least two first backup source stations backed up using different public cloud object storage spaces, and each first backup source station backs up part of the source data in the main source station, thereby avoiding the problem of occupying a lot of storage resources and bandwidth resources and high cost by fully backing up the source data in the main source station in the public cloud, and, compared with the method of building a completely self-built source data storage architecture, it reduces the cost while also reducing the difficulty of operation and maintenance. In addition, using different public cloud object storage spaces to obtain the first backup source station can also achieve a high-availability architecture under high concurrency. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0053] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0054] Figure 1 This is a schematic diagram of the mirroring back-to-source function in an embodiment of the present application;

[0055] Figure 2 A schematic diagram of a specific process of data backup in an embodiment of the present application;

[0056] Figure 3 This is a schematic diagram of the source data storage architecture in an embodiment of the present application;

[0057] Figure 4 A schematic diagram of a method flow for accessing source data in an embodiment of the present application;

[0058] Figure 5 A schematic diagram of a system architecture for accessing source data in an embodiment of the present application;

[0059] Figure 6 This is a schematic diagram of the structure of a data backup device in an embodiment of the present application;

[0060] Figure 7 This is a schematic diagram of the data backup system architecture in an embodiment of the present application;

[0061] Figure 8 Schematic diagram of the structure of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION

[0062] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0063] Object storage is a distributed storage for the Internet. It supports reading, writing and managing files at any time and place through the Http / Https protocol. It supports the standard presentation layer state transfer - application program interface (RestAPI for short). It provides customers with unlimited storage space through a flat storage architecture. Object storage is a highly reliable, highly available, low-cost, and wirelessly scalable storage method. It is suitable for storing massive amounts of unstructured data, and resources can support elastic expansion. Object storage is mainly used in public clouds.

[0064] Most public clouds have image back-to-source functions. The main use scenario of image back-to-source is to seamlessly migrate data to cloud object storage. The so-called image back-to-source means that after configuring the back-to-source rules, when the data requested by the user does not exist in the object storage, the back-to-source rules are used to obtain the data requested by the user from the set back-to-source destination address. In the back-to-source rules, at least two source station domain names are generally configured as the back-to-source destination addresses.

[0065] For example, Figure 1 The image back-to-source function diagram shown in the figure shows that if the image back-to-source is configured on the bucket, it is specifically manifested as follows: when a user sends an access request to the object storage, if the requested data does not exist in the object storage, the requested data is pulled from the customer source station, and the pulled data is stored in the object storage, and the data is synchronously returned to the user. In other words, when a user performs a get (GET) operation on a non-existent file in the bucket, the object storage service (OSS) will request the file from the back-to-source destination address (domain name), and synchronously return the file to the user after obtaining the file, and store the file in the bucket at the same time.

[0066] In an embodiment of the present application, in order to make the backup architecture of massive data, while ensuring the high availability of massive source data, save the occupied storage capacity and bandwidth as much as possible and reduce costs, a data backup method is proposed. This method mainly relies on the mirroring and source return function of the public cloud object storage space to build a relatively low-cost, highly available and high-concurrency data service architecture to ensure the normal operation of the business.

[0067] The data backup method provided in the first embodiment of the present application can be embedded in any electronic device in the form of a software program, for example, embedded in the main source station of the data operator.

[0068] like Figure 2 As shown, in the embodiment of the present application, the specific process of data backup includes:

[0069] Step 201, read the source data stored in the primary source station.

[0070] In a specific embodiment, the primary source station can be built by the data operator through the public cloud object storage space, or it can be built by the data operator using its own equipment (referred to as self-built). The primary source station is also called a primary source station.

[0071] Among them, the data operator's self-built main source station has the advantage that data resources are the data operator's precious wealth. If it completely relies on public cloud storage and then migrates the data to the self-built source station for preservation in the future, it will be time-consuming and high-risk. For data operators at the medium-sized enterprise level and above, building a first-level source station by themselves is a way that can save time and cost and is more secure.

[0072] Step 202: back up the read source data to at least two first backup source stations, where each of the at least two first backup source stations uses a different public cloud object storage space.

[0073] Among them, each public cloud object storage space backup, that is, part of the source data in the main source station is backed up in the first backup source station. The return-to-source rule of the first backup source station is used to define the return-to-source destination address when the first backup source station does not contain the requested data. The return-to-source destination address of the first backup source station when it does not contain the requested data is other backup source stations or the main source station. For example, the return-to-source rule of the first backup source station is to send a return-to-source request to the main source station when the first backup source station does not contain the requested data.

[0074] By using multiple different public cloud object storage spaces to back up source data, it plays the role of multi-cloud backup, which can effectively avoid the problem of backup source data being unavailable due to problems with a public cloud service, and achieve high availability of backup source data. In addition, by using the feature of mirroring back to the source to support multiple source station settings, it can achieve a primary-backup or even multiple backup source station architecture to achieve high availability of source data.

[0075] In a specific embodiment, in addition to the main source station and several first backup source stations, at least one second backup source station can be deployed to back up the source data read from the main source station to at least one second backup source station, wherein the second backup source station stores source data of a data volume not less than a preset proportion in the main source station. In addition, the return-to-source destination address of the second backup source station is the main source station, and the return-to-source destination address of the first backup source station is the second backup source station. In this embodiment, by setting at least one second backup source station to store source data, the high availability of source data under high concurrency can be further guaranteed, and data security can be further guaranteed. In addition, the use of a non-full backup method avoids the problem of high storage space and high bandwidth occupied by full backup, and avoids high construction costs.

[0076] In an exemplary embodiment, the sum of the amount of source data stored in each second backup source station is equal to the total amount of source data in the primary source station. In this embodiment, the source data in the primary source station is dispersed and deployed in different second backup source stations, thereby effectively alleviating the pressure on the primary source station, ensuring load balancing under high concurrency, and being able to dynamically expand the bandwidth requirements of the public cloud during business peaks.

[0077] For example, two second backup source stations are deployed, one of which stores no less than 60% of the source data in the main source station, and the other second backup source station stores no less than 40% of the source data in the main source station. The total data stored in the two second backup source stations is equal to the total amount of source data in the main source station.

[0078] The primary source station is a primary source station, at least one second backup source station is a secondary source station, and at least two first backup source stations built using public cloud object storage space are tertiary source stations.

[0079] The second backup source station can also be directly implemented in the public cloud. In the early stage of the construction of the data operator, the sum of the second backup source station and the first backup source station can be controlled to 2-3. The second backup source station and the first backup source station can be deployed in different public clouds respectively.

[0080] In a specific embodiment, when multiple first backup source stations are configured, the back-to-source priority of each first backup source station is determined, and the back-to-source rule of the first backup source station is configured according to the back-to-source priority. Specifically, the back-to-source rule of the first backup source station is configured as follows: the back-to-source destination address is the address of the first backup source station with a high priority, or the address of the primary source station. By setting the back-to-source priority, a hierarchical data backup architecture can be implemented to improve the high availability of source data.

[0081] In a specific embodiment, the types of source data stored in different first backup source stations are different. The type of source data is divided according to the access popularity of the source data, etc. For example, in the case of setting up two first backup source stations, hot data is stored in one of the first backup source stations, and warm data is stored in the other first backup source station. It should be noted that the type division of warm data and hot data here is only for example, and the type can also be obtained by dividing the source data in other ways. By saving different types of source data to different first backup source stations for storage, load balancing under high concurrency can be achieved to effectively alleviate the pressure on the main source station.

[0082] In a specific embodiment, when multiple second backup source stations are constructed, different types of source data are stored in different second backup source stations. The types of source data are divided according to the access popularity of the source data, etc. By saving different types of source data to different second backup source stations for storage, load balancing under high concurrency can be achieved to effectively relieve the pressure on the main source station.

[0083] In a specific embodiment, it is assumed that two different cloud object storage service distribution buckets are established, that is, two first backup source stations, and the return rule with the second backup source station built in the public cloud and the self-built main source station is: return to the second backup source station first, and then return to the self-built main source station. Among them, the second backup source station backs up no less than 90% of the source data to improve availability. Of course, if it is for cost considerations, the second backup source station can also be not set. One of the business distribution buckets stores hot data, accounting for about 30% of the total source data, and the other business distribution bucket stores warm data, accounting for about 50% of the total source data. It can ensure that the self-built main source station can achieve load balancing under high concurrency, effectively alleviate the pressure of a single source station, and dynamically expand the bandwidth demand of the public cloud during business peaks, avoiding high construction costs. During business off-peak periods, as the number of visits decreases, cloud consumption will also decrease. Relying on the elastic scaling and pay-as-you-go model of the public cloud, it can effectively save the operation and maintenance costs of the data operator.

[0084] In a specific embodiment, the correspondence between the address of the first backup source station and the type of source data is saved to the central node of the content distribution network (CDN). In this way, when hot data is cached in the CDN edge node, the end user can access the hot data nearby, which reduces the delay and reduces the pressure on the source station. At the same time, the CDN itself can be combined with the data preheating function and automatic refresh function of the object storage to effectively improve the CDN hit rate and data accuracy, and improve the experience of the end user.

[0085] The so-called CDN preheating function for stored data means that any node in the CDN network actively goes to the primary source station or the backup source station to download the source data and caches it to the edge node. This process is different from the process in which CDN obtains source data from the source station based on the access request uploaded by the edge node. Instead, the business party determines the hot data within a period of time based on the business scenario. The hot data is likely to be accessed by users, and the hot data is downloaded in advance to the CDN central node and edge node for storage. This process is called data preheating, which can effectively reduce the CDN's back-to-source pressure on source stations at all levels, while allowing customers to obtain the required resources faster.

[0086] The so-called automatic refresh means that when the data in the cloud vendor's backup source site is updated, it will trigger the CDN to automatically refresh the node cache, pull the updated data from the backup source site and cache it, to ensure that the source data stored on the CDN central node and edge nodes is correct, thereby ensuring that the source data obtained by the customer from the CDN node is correct, thereby improving data accuracy.

[0087] like Figure 3 The source data storage architecture shown includes a self-built first-level source station (i.e., the main source station for full storage), public cloud vendor 1-object storage (backup source station), public cloud vendor 2-object storage (business distribution bucket), and public cloud vendor 3-object storage (business distribution bucket).

[0088] Among them, the terminal sends an access request for source data to the CDN edge node, and the edge node forwards the access request to the CDN central node. The CDN central node obtains the corresponding destination address according to the identifier of the source data carried in the access request, and forwards the access request to public cloud vendor 1-object storage, public cloud vendor 2-object storage or public cloud vendor 3-object storage according to the destination address.

[0089] Assuming that the destination address is cloud vendor 1-object storage, the CDN central node forwards the access request to public cloud vendor 1-object storage. If the source data does not exist in public cloud vendor 1-object storage, the source data is obtained from the primary source site through mirroring back to source 1, stored locally and returned to the end user.

[0090] Assuming that the destination address is public cloud vendor 2-object storage, the CDN central node forwards the access request to public cloud vendor 2-object storage. If the source data does not exist in public cloud vendor 2-object storage, the source data is obtained from cloud vendor 1-object storage through mirroring back to source 1. If the source data does not exist in public cloud vendor 1-object storage, the source data is obtained from the primary source station through mirroring back to source 2. Public cloud vendor 2-object storage saves the source data locally and returns it to the end user.

[0091] The process of obtaining source data whose destination address is public cloud vendor 3-object storage can refer to the process of obtaining source data of public cloud vendor 2-object storage.

[0092] In the embodiment of the present application, by setting up a main source station for storing source data, and using different public cloud object storage spaces, the source data in the main source station is backed up in different public cloud object storage spaces, and at least two first backup source stations backed up using different public cloud object storage spaces are obtained, and each first backup source station backs up part of the source data in the main source station, thereby avoiding the problem of full backup of the source data in the main source station in the public cloud, resulting in a large amount of storage resources and bandwidth resources occupied and high costs. In addition, compared with the method of building a source data storage architecture entirely by oneself, it reduces the cost while also reducing the difficulty of operation and maintenance. In addition, using different public cloud object storage spaces to obtain the first backup source station can also achieve a high-availability architecture under high concurrency.

[0093] The architecture of using multiple public cloud object storage spaces as the backup source station takes advantage of the flexible usage of object storage, which is ready for use and eliminates the time cost of self-construction. It can quickly meet the expansion and contraction business needs of the company's own business, so 2-3 different cloud vendors can be selected to provide object storage services at the beginning.

[0094] In the second embodiment of the present application, a method for accessing source data is also provided. The source data is backed up using the data backup method of the first embodiment. The method for accessing source data can be applied to any first backup source station. Figure 4 As shown, the specific process of accessing source data is as follows:

[0095] Step 401: Obtain a data access request.

[0096] In a specific embodiment, the correspondence between the address of the first backup source station and the type of source data is stored in the central node of the CDN. After obtaining the user's data access request, the central node of the CDN obtains the address of the first backup source station corresponding to the type of source data requested by the data access request, and forwards the data access request according to the address.

[0097] Step 402 , determining whether to back up the source data requested by the data access request, if so, executing step 403 , otherwise, executing step 404 .

[0098] Step 403, returning the source data.

[0099] Step 404, according to the back-to-source destination address, source data is obtained from other backup source stations or the main source station, and the source data is returned.

[0100] Specifically, when the first backup source station that performs the process from step 401 to step 404 configures the return-to-source destination address as the address of the second backup source station, the first backup source station obtains the source data from the second backup source station according to the return-to-source destination address; wherein, when the second backup source station does not back up the source data, it obtains the source data from the primary source station and returns it;

[0101] or,

[0102] When the first backup source station in the process of executing steps 401 to 404 configures a return-to-source destination address as the address of the first backup source station with a high priority, the first backup source station obtains source data from the first backup source station with a high priority according to the return-to-source destination address, wherein the first backup source station with a high priority obtains source data from the primary source station and returns it when the source data is not backed up, and at this time, the return-to-source destination address configured by the first backup source station with a high priority is the address of the primary source station.

[0103] In this embodiment, based on the source data backed up by the data backup method provided in the first embodiment, when the source data is accessed, if the requested source data is stored in the first backup source station, it is directly returned; if the requested source data is not stored in the first backup source station, the source data is obtained from other backup source stations or the main source station based on the return-to-source destination address configured by the first backup source station, thereby achieving high availability under high concurrency of source data access.

[0104] The third embodiment of the present application also provides a system for accessing source data, and the source data is backed up using the data backup method provided in the first embodiment, such as Figure 5 As shown, the system includes at least two first backup source stations 501 and a main source station 502;

[0105] Any first backup source station 501 is used to obtain a data access request, determine whether to back up the source data requested by the data access request, and if so, return the source data; otherwise, obtain the source data from other backup source stations or the main source station according to the back-to-source destination address, and return the source data;

[0106] The main source station 502 is used to provide source data for the first backup source station or other backup source stations.

[0107] In a specific embodiment, the system further includes: at least one second backup source station 503, in which source data of a preset proportion of data in the primary source station is stored. The first backup source station 501 is used to obtain source data from the second backup source station according to the return-to-source destination address and return the source data. The second backup source station 503 is used to obtain source data from the primary source station 502 and return it to the first backup source station 501 when the source data requested by the first backup source station is not included.

[0108] In a specific embodiment, the system further includes a central node 504 and an edge node 505 of the CDN.

[0109] The edge node 505 is used to obtain the data access request submitted by the client, and when the source data requested by the data access request is cached, the source data is returned to the client; when the source data requested by the data access request is not cached, the data access request is forwarded to the central node 504.

[0110] The central node 504 is used to obtain the data access request forwarded by the edge node 505, obtain the address of the first backup source station corresponding to the type of source data requested by the data access request, and forward the data access request to the first backup source station 501 according to the address of the first backup source station.

[0111] In this embodiment, based on the source data backed up by the data backup method provided in the first embodiment, when the source data is accessed, if the requested source data is stored in the first backup source station, it is directly returned; if the requested source data is not stored in the first backup source station, the source data is obtained from other backup source stations or the main source station based on the return-to-source destination address configured by the first backup source station, thereby achieving high availability under high concurrency of source data access.

[0112] In addition, in combination with CDN, some data is cached in CDN edge nodes, so that end users can access hot data nearby, which reduces latency and reduces the pressure on the source station. When the edge node does not cache the source data requested for access, it is forwarded to the first backup source station through the central node, and the source data is obtained through the first backup source station, thereby further reducing the access pressure of the first backup source station and improving the high availability of the first backup source station under high concurrency.

[0113] Based on the same concept, the fourth embodiment of the present application provides a data backup device. The specific implementation of the device can refer to the description of the method embodiment part, and the repeated parts will not be repeated. Figure 6 As shown, the device mainly includes:

[0114] The reading module 601 is used to read the source data stored in the primary source station;

[0115] A backup module 602 is used to back up the read source data to at least two first backup source stations, where each of the at least two first backup source stations uses a different public cloud object storage space;

[0116] Among them, each of the first backup source stations backs up part of the source data in the main source station; the return-to-source destination address of the first backup source station when it does not contain the requested data is other backup source stations or the main source station.

[0117] Based on the same concept, the fifth embodiment of the present application also provides a data backup system, such as Figure 7As shown, the system mainly includes: a main source station 701 and at least two first backup source stations 702, and the at least two first backup source stations 702 each use a different public cloud object storage space.

[0118] Specifically, the primary source station 701 is used to store source data.

[0119] The first backup source station 702 is used to back up part of the source data read from the main source station 701. The return-to-source destination address of the first backup source station when it does not contain the requested data is other backup source stations or the main source station.

[0120] Based on the same concept, the sixth embodiment of the present application also provides an electronic device, such as Figure 8 As shown, the electronic device mainly includes: a processor 801 and a memory 802, the memory 802 stores a program that can be executed by the processor 801, and the processor 801 executes the program stored in the memory 802 to implement the following steps:

[0121] Read the source data stored in the main source station; back up the read source data to at least two first backup source stations, and the at least two first backup source stations each use a different public cloud object storage space; wherein each of the first backup source stations backs up part of the source data in the main source station; the return-to-source destination address of the first backup source station when it does not contain the requested data is other backup source stations or the main source station.

[0122] or,

[0123] Obtain a data access request; determine whether to back up the source data requested by the data access request; if so, return the source data; otherwise, obtain the source data from other backup source stations or the main source station according to the back-to-source destination address, and return the source data.

[0124] The processor 801 and the memory 802 in the above electronic device can be connected via a communication bus, which can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0125] The memory 802 may include a random access memory (RAM) or a non-volatile memory, such as at least one disk memory. Alternatively, the memory may also be at least one storage device located away from the processor 801.

[0126] The above-mentioned processor 801 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc., and can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components.

[0127] In another embodiment of the present application, a computer-readable storage medium is provided, in which a computer program is stored. When the computer program runs on a computer, the computer executes the data backup method or the method for accessing source data described in the above embodiments.

[0128] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instruction is loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instruction can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instruction is transmitted from a website site, a computer, a server or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server, a data center, etc. that contains one or more available media integration. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape, etc.), an optical medium (e.g., a DVD) or a semiconductor medium (e.g., a solid-state hard disk), etc.

[0129] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.

[0130] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will conform to the broadest scope consistent with the principles and novel features claimed herein.

Claims

1. A data backup method, It is characterized in that include: Read the source data stored in the primary source station; The read source data is backed up to at least two first backup source stations, each of which uses a different public cloud object storage space; each of the first backup source stations backs up part of the source data in the primary source station; the return destination address of the first backup source station when it does not contain the requested data is other backup source stations or the primary source station; Among them, after reading the source data stored in the main source station, the method also includes: backing up the read source data to at least one second backup source station, the second backup source station storing source data of a data volume not less than a preset proportion in the main source station; the return-to-source destination address of the second backup source station is the main source station, and the return-to-source destination address of the first backup source station is the second backup source station.

2. The data backup method according to claim 1, It is characterized in that The sum of the amount of source data stored in each of the second backup source stations is equal to the total amount of source data in the primary source station.

3. The data backup method according to claim 1, It is characterized in that The method further comprises: Determine a back-to-source priority for each of the first backup source stations, and configure a back-to-source rule for the first backup source station according to the back-to-source priority.

4. The data backup method according to claim 3, It is characterized in that The configuring the return-to-source rule of the first backup source station according to the return-to-source priority includes: The back-to-source rule for configuring the first backup source station is: the back-to-source destination address is the address of the first backup source station with a high priority, or the address of the primary source station.

5. The data backup method according to claim 1, It is characterized in that The types of source data stored in different first backup source stations are different, and / or the types of source data stored in different second backup source stations are different.

6. The data backup method according to claim 5, It is characterized in that The types of the source data are divided according to the access popularity of the source data.

7. The data backup method according to claim 6, It is characterized in that The correspondence between the address of the first backup source station and the type of the source data is saved in the central node of the content distribution network.

8. A method for accessing source data, It is characterized in that The source data is backed up using the data backup method according to any one of claims 1 to 7, and the method is applied to any one of the first backup source stations, comprising: Get data access request; Determining whether to back up the source data requested by the data access request; If so, return the source data; Otherwise, according to the return-to-source destination address, the source data is obtained from other backup source stations or the main source station, and the source data is returned.

9. The method for accessing source data according to claim 8, It is characterized in that The acquiring the source data from other backup source stations according to the return-to-source destination address includes: According to the return-to-source destination address, the source data is obtained from the second backup source station; wherein, when the second backup source station has not backed up the source data, the source data is obtained from the primary source station and returned; or, According to the return-to-source destination address, the source data is obtained from the first backup source station with a high priority, wherein when the first backup source station with a high priority has not backed up the source data, the source data is obtained from the primary source station and returned.

10. A system for accessing source data, It is characterized in that The source data is backed up using the data backup method according to any one of claims 1 to 7, and the system includes at least two first backup source stations and a main source station; Any of the first backup source stations is used to obtain a data access request, determine whether to back up the source data requested by the data access request, and if so, return the source data; otherwise, obtain the source data from other backup source stations or the primary source station according to the back-to-source destination address, and return the source data; The primary source station is used to provide source data for the first backup source station or the other backup source stations.

11. The system for accessing source data according to claim 10, It is characterized in that The system further comprises: at least one second backup source station, wherein the second backup source station stores source data of a volume not less than a preset proportion of data in the primary source station; The first backup source station is used to obtain the source data from the second backup source station according to the return-to-source destination address and return the source data; The second backup source station is used to obtain the source data from the primary source station and return it to the first backup source station when the source data requested by the first backup source station is not included.

12. The system for accessing source data according to claim 11, It is characterized in that The system also includes a central node and an edge node of the content distribution network; The edge node is used to obtain the data access request submitted by the client, and when caching the source data requested by the data access request, return the source data to the client; When the source data requested by the data access request is not cached, forwarding the data access request to the central node; The central node is used to obtain the data access request forwarded by the edge node, obtain the address of the first backup source station corresponding to the type of source data requested by the data access request, and forward the data access request according to the address of the first backup source station.

13. A data backup device, It is characterized in that include: A reading module, used to read source data stored in the primary source station; A backup module, used to back up the read source data to at least two first backup source stations, each of which uses a different public cloud object storage space; each of the first backup source stations backs up part of the source data in the main source station; the first backup source station has a return destination address of other backup source stations or the main source station when it does not contain the requested data; The backup module is further used to: after reading the source data stored in the primary source station, back up the read source data to at least one second backup source station, where the second backup source station stores source data of a data volume not less than a preset proportion in the primary source station; The return-to-source destination address of the second backup source station is the primary source station, and the return-to-source destination address of the first backup source station is the second backup source station.

14. A data backup system, It is characterized in that include: A primary source station, at least two first backup source stations, and at least one second backup source station, wherein the at least two first backup source stations respectively use different public cloud object storage spaces; The primary source station is used to store source data; The first backup source station is used to back up part of the source data read from the primary source station, and the return-to-source destination address of the first backup source station when it does not contain the requested data is other backup source stations or the primary source station; The second backup source station is used to back up the source data read from the primary source station, and the second backup source station stores source data of a volume not less than a preset proportion of data in the primary source station; The return-to-source destination address of the second backup source station is the primary source station, and the return-to-source destination address of the first backup source station is the second backup source station.

15. An electronic device, It is characterized in that include: Processor and memory; The memory is used to store computer programs; The processor is used to execute the program stored in the memory to implement the data backup method described in any one of claims 1 to 7, or to implement the method for accessing source data described in any one of claims 8 to 9.

Citation Information

Patent Citations

  • Data backup and recovery method and device

    CN105740091A