Data processing method and device, electronic equipment, readable storage medium and product
By sharing the storage resources of the tape library storage system between multiple data nodes in the cluster, the problem of data storage or read queue delay caused by a single data node is solved, and data processing efficiency is improved.
Patent Information
- Application Number
- CN202411885083.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-13
AI Technical Summary
Directly mounting the tape library storage system under a single data node causes delays in queuing of data storage or reading, reducing data processing efficiency.
By connecting the tape library storage system in at least two data nodes in the cluster, creating tasks and assigning idle nodes to perform tasks according to the data node status by connecting to the tape library storage system, in response to storage or read requests for user data, the task is created and the idle nodes are assigned to perform tasks according to the data node status, thereby sharing storage resources and avoiding delays caused by node exclusiveness.
By sharing the storage resources of the tape library storage system by multiple data nodes in the cluster, data storage or reading tasks can be evenly allocated, avoiding queuing delays of individual nodes and improving data processing efficiency.
Smart Images

Figure CN119987653A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, and in particular to the field of data storage technology. Background Art
[0002] Tape storage technology has the characteristics of low power consumption, low cost, and long data retention time, but the tape itself is an offline storage medium that needs to be loaded into a drive for reading and writing operations, which also brings problems of high latency and low throughput. Traditional tape library storage systems can be directly mounted on a single data node, and the data node performs storage and reading of the tape library storage system.
[0003] However, since a single data node is directly connected to a tape library storage system, the storage resources of the tape library storage system are exclusively used, so a data node may experience queuing delays in data storage or reading, thereby reducing data processing efficiency. Summary of the invention
[0004] The present disclosure provides a data processing method, device, electronic device, readable storage medium and product.
[0005] According to one aspect of the present disclosure, there is provided a data processing method, which is applied to a cluster consisting of at least two data nodes, wherein the at least two data nodes are connected to a tape library storage system, and the method comprises:
[0006] In response to a data storage request for user data, creating a data storage task for the user data;
[0007] Allocating a data node in an idle state to the data storage task according to the states of the at least two data nodes;
[0008] The data storage task is performed by utilizing the data nodes in the idle state to store the user data in the tape library storage system.
[0009] According to another aspect of the present disclosure, another data processing method is provided, which is applied to a cluster consisting of at least two data nodes, wherein the at least two data nodes are connected to a tape library storage system, and the method comprises:
[0010] In response to a data reading request of user data, creating a data reading task of the user data;
[0011] According to the states of the at least two data nodes, allocating a data node in an idle state to the data reading task;
[0012] The data reading task is performed by utilizing the data nodes in the idle state to read the user data stored in the tape library storage system.
[0013] According to another aspect of the present disclosure, there is provided a data processing device, which is applied to a cluster consisting of at least two data nodes, wherein the at least two data nodes are connected to a tape library storage system, and the device comprises:
[0014] A task creation unit, configured to create a data storage task for the user data in response to a data storage request for the user data;
[0015] A task allocation unit, configured to allocate a data node in an idle state to the data storage task according to the states of the at least two data nodes;
[0016] The data storage unit is used to utilize the data nodes in the idle state to execute the data storage task so as to store the user data in the tape library storage system.
[0017] According to another aspect of the present disclosure, another data processing device is provided, which is applied to a cluster consisting of at least two data nodes, wherein the at least two data nodes are connected to a tape library storage system, and the device comprises:
[0018] A task creation unit, configured to create a data reading task for the user data in response to a data reading request for the user data;
[0019] A task allocation unit, configured to allocate a data node in an idle state to the data reading task according to the states of the at least two data nodes;
[0020] The data reading unit is used to use the data nodes in the idle state to execute the data reading task to read the user data stored in the tape library storage system.
[0021] According to another aspect of the present disclosure, there is provided an electronic device, comprising:
[0022] at least one processor; and
[0023] a memory communicatively connected to the at least one processor; wherein,
[0024] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any possible implementation manner and the aspects described above.
[0025] According to yet another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the method of the above-mentioned aspect and any possible implementation manner.
[0026] According to yet another aspect of the present disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the computer program implements the above-mentioned aspects and any possible implementation method.
[0027] On the one hand, it can be seen from the above technical solution that the embodiment of the present disclosure creates a data storage task for the user data in response to a data storage request for the user data, and then allocates an idle data node to the data storage task according to the status of the at least two data nodes, so that the idle data node can be used to execute the data storage task to store the user data in the tape library storage system. The storage resources of the tape library storage system can be shared by multiple data nodes in the cluster, so that the data storage task can be allocated to different data nodes in the cluster for corresponding processing, which can avoid queuing delays in data storage of individual data nodes, thereby improving data processing efficiency.
[0028] On the other hand, it can be seen from the above technical solution that the embodiment of the present disclosure creates a data reading task for the user data in response to a data reading request for the user data, and then allocates an idle data node to the data reading task according to the status of the at least two data nodes, so that the idle data node can be used to execute the data reading task to read the user data stored in the tape library storage system. By sharing the storage resources of the tape library storage system by multiple data nodes in the cluster, the data reading task can be allocated to different data nodes in the cluster for corresponding processing, which can avoid queuing delays in data reading of individual data nodes, thereby improving data processing efficiency.
[0029] In addition, by adopting the technical solution provided by the present invention, multiple copies of user data are created by utilizing the copy capability of the file system, and based on the data storage tasks of these copies of data, idle data nodes corresponding to the copy data in the cluster are allocated to the copy data, and the copy data is written to the tape library storage system, which can effectively improve the security and reliability of data processing.
[0030] In addition, by adopting the technical solution provided by the present invention, the data nodes in the cluster are logically divided into at least a first node and a second node, each corresponding to the storage tasks of different copies, so that different copy data can be written concurrently to the tape library storage system at the same time, which can effectively improve the security and reliability of data processing.
[0031] In addition, by adopting the technical solution provided by the present invention, the storage capacity of the tape library storage system is utilized to store multiple copies of data corresponding to the user data. Based on the data reading task of the user data, the idle data node corresponding to any copy in the cluster is allocated, and the data node is read from the tape medium in the corresponding tape library storage system, which can effectively improve the security and reliability of data processing.
[0032] In addition, by adopting the technical solution provided by the present invention, the data nodes in the cluster are logically divided into at least a first node and a second node, which correspond to the reading tasks of different replicas respectively, so that the reading tasks can be allocated based on different replica data, thereby reducing the reading delay of a single user data.
[0033] In addition, the technical solution provided by the present disclosure can effectively improve the user experience.
[0034] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following is a brief introduction to the drawings required for use in the embodiments or prior art descriptions. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative labor. The drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure. Among them:
[0036] Figure 1 is a schematic diagram according to a first embodiment of the present disclosure;
[0037] Figure 2 yes Figure 1 An application schematic diagram of the corresponding embodiment;
[0038] Figure 3 yes Figure 1 Another application schematic diagram of the corresponding embodiment;
[0039] Figure 4 is a schematic diagram according to a second embodiment of the present disclosure;
[0040] Figure 5 yes Figure 4 An application schematic diagram of the corresponding embodiment;
[0041] Figure 6 yes Figure 4 Another application schematic diagram of the corresponding embodiment;
[0042] Figure 7 is a schematic diagram according to a third embodiment of the present disclosure;
[0043] Figure 8 is a schematic diagram according to a fourth embodiment of the present disclosure;
[0044] Fig. 9 It is a block diagram of an electronic device used to implement the data processing method of the embodiment of the present disclosure. DETAILED DESCRIPTION
[0045] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0046] Obviously, the described embodiments are only part of the embodiments of the present disclosure, but not all of them. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present disclosure.
[0047] It should be noted that the terminal devices involved in the embodiments of the present disclosure may include but are not limited to mobile phones, personal digital assistants (PDAs), wireless handheld devices, tablet computers and other smart devices; display devices may include but are not limited to personal computers, televisions and other devices with display functions.
[0048] In addition, the term "and / or" in this article is only a description of the association relationship between the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0049] At present, tape storage technology is mainly used in archiving and backup scenarios, and has two applications: offline storage and near-line storage. The tape library storage system for offline storage can include storage devices and warehouses. When the storage medium (i.e., tape medium) is full of data, the storage device will take the storage medium out and put it in the warehouse for offline storage; for near-line storage, the storage medium (i.e., tape medium) is always kept in the storage device. Regardless of near-line storage or offline storage, there are two commonly used tape library storage systems.
[0050] A tape library storage system uses a Fibre Channel (FC) network. The data nodes and tape library storage system are connected to the FC switch. There is also a master control node and a cache system. The master control node is used for task allocation, and the cache system is used to receive data to be stored and cache data to be read. The entire system is expensive.
[0051] Another type of tape library storage system is to use data nodes directly mounted on the tape library storage system. The data node also serves as a cache system. Since it is a single node, no additional master node is required for task allocation. The disadvantage is that the data node is mounted on the storage resources of the tape library storage system. For example, the drive and tape media are exclusive. It may happen that a data node is busy with tasks and the queuing delay increases, but other data nodes are idle.
[0052] In response to the above problems, the present invention provides a data processing method, which also adopts the method of directly mounting a tape library storage system under a data node, but multiple data nodes form a cluster, and the storage resources of the tape library storage system in the cluster are shared. Storage tasks or reading tasks can be evenly distributed to the data nodes in the cluster through task scheduling, which can improve the throughput capacity of tape media storage, shorten the entire task time, and reduce user waiting time.
[0053] The technical solution provided by the present disclosure can be applied to a cluster consisting of at least two data nodes, wherein the at least two data nodes are connected to a tape library storage system, and is mainly used for storing and reading cold data. The tape library storage system is a backup system based on tape media, and can be composed of multiple tape drives, such as Linear Tape-Open (LTO), multiple slots, and mechanical arms.
[0054] Since the storage rate of the tape medium is determined by the rate of the tape drive, if the tape drive is mounted on a single data node, the read and write rate of the single data node is fixed. With the cluster composed of at least two data nodes provided by the present disclosure, since multiple data nodes form a cluster, multiple tape drives can be mounted on the cluster, and these tape drives are shared by all data nodes in the cluster, so that the maximum throughput of the cluster is the sum of the concurrent processing rates of all data nodes, thereby effectively shortening the time for data processing and reading data.
[0055] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure, such as Figure 1 shown.
[0056] 101. In response to a data storage request for user data, create a data storage task for the user data.
[0057] 102. Allocate an idle data node to the data storage task according to the states of the at least two data nodes.
[0058] 103. Utilize the idle data nodes to execute the data storage task, so as to store the user data in the tape library storage system.
[0059] At this point, all data nodes in the cluster share the storage resources of the tape library storage system, and evenly distribute data storage tasks to different data nodes in the cluster for corresponding processing, which can effectively utilize the processing resources of the data nodes in the cluster, thereby improving data processing efficiency.
[0060] It should be noted that part or all of the execution entities 101 to 103 may be a processing engine located in a network side server, or may also be a distributed system located on the network side, for example, a processing engine or distributed system in a trigger platform on the network side, or may also be located at a specific data node in the cluster, for example, a name node, a control node, etc. This embodiment does not specifically limit this.
[0061] In this way, by responding to the data storage request of the user data, a data storage task of the user data is created, and then according to the status of the at least two data nodes, an idle data node is allocated to the data storage task, so that the idle data node can be used to execute the data storage task to store the user data in the tape library storage system. By mounting the tape library storage system under multiple data nodes in the cluster, multiple data nodes in the cluster can share the storage resources of the tape library storage system, so that the data storage task can be allocated to different data nodes in the cluster for corresponding processing, which can avoid queuing delays in data storage in individual data nodes, thereby improving data processing efficiency.
[0062] The technical solution provided by the present disclosure can be applied to a cluster consisting of at least two data nodes, for example, a cluster consisting of data node 1, data node 2, ..., data node N, wherein these data nodes in the cluster are connected to a tape library storage system, for example, the data nodes are connected to a tape drive of the tape library storage system via FC, and share the storage resources of the tape library storage system, such as Figure 2 As shown. It can be understood that the storage resource of the tape library storage system can be a complete tape library storage system resource or a logical resource divided by the tape library storage system. The data nodes in the cluster share the same file system, all data nodes present a unified file system path, and multiple data nodes can concurrently write or read the file system at the same time, which can make the capacity of the file system a multiple of the bandwidth of a single data node.
[0063] The technical solution provided by the present disclosure uses a shared file system to form a cluster of multiple data nodes, and also forms a corresponding file system cluster, namely a tape library storage system, in the storage layer of the tape medium. The storage resources of the tape medium are also divided according to the cluster, and the resources are shared within the cluster.
[0064] The technical solution provided by the present invention provides a single mounting interface for upper-layer applications, simplifies the complexity of upper-layer applications to the greatest extent, and realizes cluster-wide sharing of storage resources of the back-end tape library storage system. The data access throughput capacity of the entire system is improved by several times compared with a single data node, which greatly shortens the batch access time and expands the application scenarios of tape media.
[0065] The technical solution provided by the present disclosure is that the file system provided by the tape library storage system is a tape file system. Since its file format is different from the file format of the general file system of the computer, for example, the Linear Tape File System (LTFS) of the tape file system and the general file system Linux of the computer, therefore, when executing data storage tasks, format conversion between file systems is required.
[0066] The technical solution provided by the present disclosure can be set up to execute by a single device or multiple devices, specifically executing user data storage operation-related tasks such as storage task creation, storage task allocation, and storage space allocation for user data, and each data node in the cluster, i.e., the assigned (i.e., called) data node, respectively executes user data metadata registration, metadata update, and other user data storage status-related tasks, as well as write-related tasks of writing user data into the tape library storage system.
[0067] The technical solution provided by the present disclosure can also be executed by a specific data node in the cluster, such as a control node or a name node, to specifically perform user data storage operation-related tasks such as storage task creation, storage task allocation, and storage space allocation for user data. The specific data node can further perform user data storage status-related tasks such as metadata registration and metadata update for user data, while other ordinary data nodes in the cluster, i.e., the assigned (i.e., called) data nodes, respectively perform write-related tasks of writing user data into the tape library storage system.
[0068] It should be noted that the present disclosure does not limit the specific manner of creating a data storage task for the user data in response to a data storage request for the user data, and the specific manner may be selected according to actual circumstances.
[0069] Optionally, in a possible implementation of this embodiment, a data storage request for user data transmitted from a client via Ethernet may be received, and the data storage request may be triggered by the client through an operation such as uploading a file.
[0070] In the technical solution provided by the present disclosure, there is no restriction on the format of the user data obtained from the client, which may be text data, or picture data, or audio data, or video data, and this embodiment does not specifically limit this.
[0071] After receiving the user data from the client, the user data can be preprocessed to obtain the user identification and data type. For example, the user unique identification and data label are marked, and then the user data is packaged and sliced to obtain a fixed-size file block and record the metadata of the user data. The processed user data is stored in the file system shared by the data nodes in the cluster and waits for the next step of processing.
[0072] The metadata of the user data is used to record the status information of the user data, and may include all data required for data access control. When accessing the user data, the metadata of the user data is first queried, and then subsequent I / O operations such as data reading and writing are performed through the obtained metadata.
[0073] The metadata of the user data can be considered as an electronic directory for locating the user data, and can provide shared access to any authorized system / device. Therefore, the metadata of the user data can be stored and processed separately in a metadata storage system for unified management.
[0074] It should be noted that the present disclosure does not limit the specific manner of allocating idle data nodes to the data storage task according to the states of the at least two data nodes, and the manner may be selected according to actual conditions.
[0075] Optionally, in a possible implementation of this embodiment, the busy and idle status of data nodes in the cluster can be monitored in real time, and then based on the busy and idle status of these data nodes, idle data nodes can be selected and data storage tasks can be assigned to the idle data nodes.
[0076] It should be noted that the present disclosure does not limit the specific manner of utilizing the idle data nodes to execute the data storage task so as to store the user data in the tape library storage system, and the specific manner may be selected according to actual conditions.
[0077] Optionally, in a possible implementation of this embodiment, the tape library storage system includes at least two logical tape library storage subsystems, and each of the at least two logical tape library storage subsystems is logically isolated from each other.
[0078] Specifically, the tape library storage system can be divided into multiple logical partitions, each logical partition is a logical tape library storage subsystem, and each logical partition is logically isolated from each other. Each logical tape library storage subsystem is divided into hardware resources such as one or more tape drives and tape media.
[0079] At this time, in the technical solution provided by the present invention, the cluster can call the tape drive connected to any data node to access (i.e., write or read) the tape media resources in all logical tape library storage subsystems. Concurrent access by tape drives connected to multiple data nodes can effectively improve the throughput of the entire tape library storage system, thereby solving the problem of waiting for resource release due to tape media resource occupation in a single data node, and also solving the problem of queuing for processing due to limited tape drive resources of a single data node.
[0080] In a specific implementation process, in 103, the usage status of the tape medium in each of the at least two logical tape library storage subsystems can be obtained, and then, based on the usage status of the tape medium in each logical tape library storage subsystem, the user data can be written into the tape medium in the corresponding logical tape library storage subsystem by using the idle data nodes.
[0081] Optionally, in a possible implementation of this embodiment, in order to improve data security and system robustness, multiple copies need to be saved, for example, 2 copies of data or multiple copies of data are saved, so the data storage task can be a multiple copy storage task, which can at least include but is not limited to a first copy storage task and a second copy storage task. The data nodes can be logically divided into data nodes corresponding to storage tasks of different copies at least.
[0082] Specifically, the data nodes may at least include but are not limited to at least two first nodes and at least two second nodes, the at least two first nodes correspond to the first replica storage tasks, and the at least two second nodes correspond to the second replica storage tasks.
[0083] In this implementation, multiple copies of user data are created by utilizing the copy capability of the file system. Based on the data storage tasks of these copies of data, idle data nodes corresponding to the copies in the cluster are allocated to the copy data, and the copy data is written to the tape library storage system, which can effectively improve the security and reliability of data processing.
[0084] In a specific implementation process, in 102, copy data corresponding to the user data can be created based on the multi-copy storage task, and the copy data can at least include but are not limited to a first copy and a second copy, the first copy corresponds to the first copy storage task, and the second copy corresponds to the second copy storage task. Furthermore, based on the status of the at least two first nodes, an idle first node can be allocated to the first copy storage task, and based on the status of the at least two second nodes, an idle second node can be allocated to the second copy storage task.
[0085] Specifically, the copy capability of the file system can be used to create multiple copies of the user data. A different tape media storage path can be created for each copy of the data, and the data is stored in the tape storage medium in the corresponding tape library storage system based on the path. When a read request for user data from a client is received, one or more corresponding tape media storage paths can be obtained according to the data characteristics of the user data, and the copy data of the user data can be read using any tape media storage path. If any data is read successfully, the read request is successfully completed. Furthermore, if multiple copies of data are stored, data reading tasks can be further allocated according to the copy data, thereby reducing the reading delay of a single copy of data.
[0086] Correspondingly, in 103, the first node in the idle state can be used to execute the first copy storage task to write the first copy to the tape library storage system, and the second node in the idle state can be used to execute the second copy storage task to write the second copy to the tape library storage system.
[0087] In this implementation, by logically dividing the data nodes in the cluster into at least a first node and a second node, each corresponding to the storage tasks of different copies, different copy data can be written concurrently to the tape library storage system, which can effectively improve the security and reliability of data processing.
[0088] In another specific implementation process, the tape library storage system can be logically divided into at least storage systems corresponding to data nodes of different copies. Specifically, the tape library storage system can include at least but not limited to a first storage system connected to the first node and a second storage system connected to the second node, for example, a first storage system connected to the first node 1, the first node 2, ..., the first node N and its tape medium, a second storage system connected to the second node 1, the second node 2, ..., the second node N and its tape medium, such as Figure 3 shown.
[0089] Correspondingly, in 103, the first node in the idle state can be used to execute the first copy storage task to write the first copy to the first storage system, and the second node in the idle state can be used to execute the second copy storage task to write the second copy to the second storage system.
[0090] In this implementation, by logically dividing the tape library storage system into at least a first storage system and a second storage system, each corresponding to the storage tasks of different copies, different copy data can be written concurrently to different storage systems in the tape library storage system, which can more effectively improve the security and reliability of data processing.
[0091] In this embodiment, a data storage task for the user data is created in response to a data storage request for the user data, and then an idle data node is allocated to the data storage task based on the status of the at least two data nodes, so that the idle data node can be used to execute the data storage task to store the user data in the tape library storage system. The storage resources of the tape library storage system can be shared by multiple data nodes in the cluster, so that the data storage task can be allocated to different data nodes in the cluster for corresponding processing, which can avoid queuing delays in data storage for individual data nodes, thereby improving data processing efficiency.
[0092] In addition, by adopting the technical solution provided by the present invention, multiple copies of user data are created by utilizing the copy capability of the file system. Based on the data storage tasks of these copies of data, idle data nodes corresponding to the copies in the cluster are allocated to the copy data, and the copy data is written to the tape library storage system, which can effectively improve the security and reliability of data processing.
[0093] In addition, by adopting the technical solution provided by the present invention, the data nodes in the cluster are logically divided into at least a first node and a second node, each corresponding to the storage tasks of different copies, so that different copy data can be written concurrently to the tape library storage system at the same time, which can effectively improve the security and reliability of data processing.
[0094] In addition, by adopting the technical solution provided by the present invention, the data nodes in the cluster are logically divided into at least a first node and a second node, which correspond to the reading tasks of different replicas respectively, so that the reading tasks can be allocated based on different replica data, thereby reducing the reading delay of a single user data.
[0095] In addition, the technical solution provided by the present disclosure can effectively improve the user experience.
[0096] Figure 4 is a schematic diagram according to the second embodiment of the present disclosure, such as Figure 4 shown.
[0097] 401. In response to a data reading request of user data, create a data reading task of the user data.
[0098] 402. Allocate an idle data node to the data reading task according to the status of the at least two data nodes.
[0099] 403. Utilize the data node in the idle state to execute the data reading task to read the user data stored in the tape library storage system.
[0100] At this point, all data nodes in the cluster share the storage resources of the tape library storage system, and evenly distribute data reading tasks to different data nodes in the cluster for corresponding processing, which can effectively utilize the processing resources of the data nodes in the cluster, thereby improving data processing efficiency.
[0101] It should be noted that part or all of the execution entities 401 to 403 may be a processing engine located in a network side server, or may also be a distributed system located on the network side, for example, a processing engine or distributed system in a trigger platform on the network side, or may also be located at a specific data node in the cluster, for example, a name node, a control node, etc. This embodiment does not specifically limit this.
[0102] In this way, by responding to a data read request for user data, a data read task for the user data is created, and then according to the status of the at least two data nodes, an idle data node is allocated to the data read task, so that the idle data node can be used to execute the data read task to read the user data stored in the tape library storage system. By mounting the tape library storage system under multiple data nodes in the cluster, multiple data nodes in the cluster can share the storage resources of the tape library storage system, so that data read tasks can be allocated to different data nodes in the cluster for corresponding processing, which can avoid queuing delays in data reading of individual data nodes, thereby improving data processing efficiency.
[0103] The technical solution provided by the present disclosure can be applied to a cluster consisting of at least two data nodes, for example, a cluster consisting of data node 1, data node 2, ..., data node N. These data nodes in the cluster are connected to a tape library storage system and share the storage resources of the tape library storage system, such as Figure 5 As shown. It can be understood that the storage resource of the tape library storage system can be a complete tape library storage system resource or a logical resource divided by the tape library storage system. The data nodes in the cluster share the same file system, all data nodes present a unified file system path, and multiple data nodes can concurrently write or read the file system at the same time, which can make the capacity of the file system a multiple of the bandwidth of a single data node.
[0104] The technical solution provided by the present disclosure uses a shared file system to form a cluster of multiple data nodes, and also forms a corresponding file system cluster, namely a tape library storage system, in the storage layer of the tape medium. The storage resources of the tape medium are also divided according to the cluster, and the resources are shared within the cluster.
[0105] The technical solution provided by the present invention provides a single mounting interface for upper-layer applications, simplifies the complexity of upper-layer applications to the greatest extent, and realizes cluster-wide sharing of storage resources of the back-end tape library storage system. The data access throughput capacity of the entire system is improved by several times compared with a single data node, which greatly shortens the batch access time and expands the application scenarios of tape media.
[0106] The technical solution provided by the present disclosure can be set up to execute by a single device or multiple devices, specifically executing user data reading task creation, reading task allocation and other user data reading operation-related tasks, and each data node in the cluster, that is, the assigned (i.e., called) data node, respectively executes user data data feature matching, metadata query and other user data reading status-related tasks to find the tape medium where the user data to be read is located and the position offset information, and reads the user data from the corresponding tape medium in the tape library storage system and temporarily stores it in the file system.
[0107] The technical solution provided by the present disclosure can also be executed by specific data nodes in the cluster, such as control nodes or name nodes, to specifically perform user data reading task creation, reading task allocation and other user data reading operation-related tasks, and each data node in the cluster, i.e., the assigned (i.e., called) data node, respectively performs user data reading status-related tasks such as data feature matching and metadata query of user data to find the tape medium where the user data to be read is located and the position offset information, and reads the user data from the corresponding tape medium in the tape library storage system and temporarily stores it in the file system.
[0108] It should be noted that the present disclosure does not limit the specific manner of creating a data reading task for the user data in response to a data reading request for the user data, and the specific manner may be selected according to actual conditions.
[0109] Optionally, in a possible implementation of this embodiment, a data reading request for user data transmitted from a client via Ethernet may be received, and the data reading request may be triggered by the client through an operation such as downloading a file.
[0110] It should be noted that the present disclosure does not limit the specific manner of allocating idle data nodes to the data reading task according to the states of the at least two data nodes, and the manner may be selected according to actual conditions.
[0111] Optionally, in a possible implementation of this embodiment, the busy and idle status of data nodes in the cluster can be monitored in real time, and then based on the busy and idle status of these data nodes, idle data nodes can be selected and data reading tasks can be assigned to the idle data nodes.
[0112] It should be noted that the present disclosure does not limit the specific manner of utilizing the idle data nodes to execute the data reading task to read the user data stored in the tape library storage system, and the specific manner may be selected according to actual conditions.
[0113] Optionally, in a possible implementation of this embodiment, the tape library storage system includes at least two logical tape library storage subsystems, and each of the at least two logical tape library storage subsystems is logically isolated from each other.
[0114] Specifically, the tape library storage system can be divided into multiple logical partitions, each logical partition is a logical tape library storage subsystem, and each logical partition is logically isolated from each other. Each logical tape library storage subsystem is divided into hardware resources such as one or more tape drives and tape media.
[0115] At this time, in the technical solution provided by the present invention, the cluster can call the tape drive connected to any data node to access (i.e., write or read) the tape media resources in all logical tape library storage subsystems. Concurrent access by tape drives connected to multiple data nodes can effectively improve the throughput of the entire tape library storage system, thereby solving the problem of waiting for resource release due to tape media resource occupation in a single data node, and also solving the problem of queuing for processing due to limited tape drive resources of a single data node.
[0116] In a specific implementation process, in 403, the tape medium in the logical tape library storage subsystem where the user data is stored can be specifically determined, and then the data node in the idle state can be used to execute the data reading task to read the user data stored on the tape medium in the logical tape library storage subsystem.
[0117] Optionally, in a possible implementation of this embodiment, in order to improve data security and system robustness, multiple copies need to be saved, for example, 2 copies of data or multiple copies of data are saved, therefore, the copy data corresponding to the user data stored in the tape library storage system may at least include but not limited to the first copy and the second copy. The data nodes may be logically divided into at least data nodes corresponding to different copies.
[0118] Specifically, the data nodes may at least include but are not limited to at least two first nodes and at least two second nodes, the at least two first nodes correspond to the first replica storage tasks, and the at least two second nodes correspond to the second replica storage tasks.
[0119] In this implementation, the storage capacity of the tape library storage system is utilized to store multiple copy data corresponding to the user data. Based on the data reading task of the user data, an idle data node corresponding to any copy data in the cluster is allocated, and the data node is read from the tape medium in the corresponding tape library storage system, which can effectively improve the security and reliability of data processing.
[0120] In a specific implementation process, in 403, the data reading task can be performed by using the allocated first node or second node in the idle state to read the first copy or the second copy stored in the tape library storage system.
[0121] In this implementation, the data nodes in the cluster are logically divided into at least a first node and a second node, each corresponding to the reading tasks of different replicas, so that the reading tasks can be allocated based on different replica data, thereby reducing the reading delay of a single user data.
[0122] In another specific implementation process, the tape library storage system can be logically divided into at least storage systems corresponding to data nodes of different copies. Specifically, the tape library storage system can include at least but not limited to a first storage system connected to the first node and a second storage system connected to the second node, for example, a first storage system connected to the first node 1, the first node 2, ..., the first node N and its tape medium, a second storage system connected to the second node 1, the second node 2, ..., the second node N and its tape medium, such as Figure 3 shown.
[0123] Accordingly, in 403, the data reading task can be specifically performed by utilizing the allocated first node or second node in the idle state to read the first copy stored in the first storage system or the second copy stored in the second storage system.
[0124] In this implementation, the tape library storage system is logically divided into at least a first storage system and a second storage system, each corresponding to the reading tasks of different copies, so that the reading tasks can be allocated and executed based on different copy data, thereby further reducing the reading delay of a single user data.
[0125] Specifically, the idle data node can find the data storage path according to the data characteristics indicated by the data reading task, and can read data from any replica. In actual application, the data is read according to the real-time task status in the replica. If any replica is read successfully, the data reading request is completed. Furthermore, the reading task can be assigned by replica to reduce the delay of reading a single replica.
[0126] In this embodiment, a data reading task for the user data is created in response to a data reading request for the user data, and then a data node in an idle state is allocated to the data reading task according to the state of the at least two data nodes, so that the data reading task can be executed by utilizing the data node in the idle state to read the user data stored in the tape library storage system. By sharing the storage resources of the tape library storage system by multiple data nodes in the cluster, the data reading task can be allocated to different data nodes in the cluster for corresponding processing, which can avoid queuing delays in data reading of individual data nodes, thereby improving data processing efficiency.
[0127] In addition, by adopting the technical solution provided by the present invention, the storage capacity of the tape library storage system is utilized to store multiple copy data corresponding to the user data. Based on the data reading task of the user data, an idle data node corresponding to any copy data in the cluster is allocated, and the data node is read from the tape medium in the corresponding tape library storage system, which can effectively improve the security and reliability of data processing.
[0128] In addition, by adopting the technical solution provided by the present invention, the data nodes in the cluster are logically divided into at least a first node and a second node, which correspond to the reading tasks of different replicas respectively, so that the reading tasks can be allocated based on different replica data, thereby reducing the reading delay of a single user data.
[0129] In addition, the technical solution provided by the present disclosure can effectively improve the user experience.
[0130] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the present disclosure is not limited by the order of the actions described, because according to the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present disclosure.
[0131] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0132] Figure 7 is a schematic diagram according to the third embodiment of the present disclosure, Figure 7 As shown. The data processing device 700 of this embodiment may include a task creation unit 701, a task allocation unit 702 and a data storage unit 703. The task creation unit 701 is used to create a data storage task for the user data in response to a data storage request for the user data; the task allocation unit 702 is used to allocate an idle data node to the data storage task according to the status of the at least two data nodes; the data storage unit 703 is used to use the idle data node to execute the data storage task to store the user data in the tape library storage system.
[0133] It should be noted that part or all of the data processing device of this embodiment may be a processing engine located in a network side server, or may also be a distributed system located on the network side, for example, a processing engine or distributed system in a trigger platform on the network side, etc., or may also be located at a specific data node in the cluster, for example, a name node, a control node, etc. This embodiment does not specifically limit this.
[0134] Optionally, in a possible implementation of this embodiment, the tape library storage system includes at least two logical tape library storage subsystems, and each of the at least two logical tape library storage subsystems is logically isolated from each other; the data storage unit 703 is specifically used to obtain the usage status of the tape medium in each of the at least two logical tape library storage subsystems; and according to the usage status of the tape medium in each logical tape library storage subsystem, use the idle data nodes to execute the data storage task to write the user data into the tape medium in the corresponding logical tape library storage subsystem.
[0135] Optionally, in a possible implementation of this embodiment, the data storage task is a multi-copy storage task, including at least a first copy storage task and a second copy storage task; the data nodes include at least at least two first nodes corresponding to the first copy storage task and at least two second nodes corresponding to the second copy storage task; the task allocation unit 702 is specifically used to create copy data corresponding to the user data according to the multi-copy storage tasks, and the copy data includes at least a first copy corresponding to the first copy storage task and a second copy corresponding to the second copy storage task; according to the status of the at least two first nodes, allocate a first node in an idle state to the first copy storage task; and according to the status of the at least two second nodes, allocate a second node in an idle state to the second copy storage task.
[0136] Specifically, the data storage unit 703 can be used to utilize the first node in the idle state to execute the first copy storage task to write the first copy to the tape library storage system; and utilize the second node in the idle state to execute the second copy storage task to write the second copy to the tape library storage system.
[0137] Correspondingly, the tape library storage system includes at least a first storage system connected to the first node and a second storage system connected to the second node; the data storage unit 703 is specifically used to use the first node in the idle state to execute the first copy storage task to write the first copy to the first storage system; and use the second node in the idle state to execute the second copy storage task to write the second copy to the second storage system.
[0138] It should be noted that Figure 1-Figure 3 The method in the embodiment corresponding to any of the figures can be implemented by the data processing device provided in this embodiment. Figure 1-Figure 3 The relevant contents in the embodiments corresponding to any of the accompanying drawings will not be repeated here.
[0139] In this embodiment, a task creation unit responds to a data storage request for user data to create a data storage task for the user data, and then a task allocation unit allocates idle data nodes to the data storage task based on the status of the at least two data nodes, so that the data storage unit can use the idle data nodes to execute the data storage task to store the user data in the tape library storage system. By sharing the storage resources of the tape library storage system by multiple data nodes in the cluster, the data storage tasks can be allocated to different data nodes in the cluster for corresponding processing, which can avoid queuing delays in data storage for individual data nodes, thereby improving data processing efficiency.
[0140] In addition, by adopting the technical solution provided by the present invention, multiple copies of user data are created by utilizing the copy capability of the file system. Based on the data storage tasks of these copies of data, idle data nodes corresponding to the copies in the cluster are allocated to the copy data, and the copy data is written to the tape library storage system, which can effectively improve the security and reliability of data processing.
[0141] In addition, by adopting the technical solution provided by the present invention, the data nodes in the cluster are logically divided into at least a first node and a second node, each corresponding to the storage tasks of different copies, so that different copy data can be written concurrently to the tape library storage system at the same time, which can effectively improve the security and reliability of data processing.
[0142] In addition, by adopting the technical solution provided by the present invention, the data nodes in the cluster are logically divided into at least a first node and a second node, which correspond to the reading tasks of different replicas respectively, so that the reading tasks can be allocated based on different replica data, thereby reducing the reading delay of a single user data.
[0143] In addition, the technical solution provided by the present disclosure can effectively improve the user experience.
[0144] Figure 8 is a schematic diagram according to a fourth embodiment of the present disclosure, Figure 8 As shown. The data processing device 800 of this embodiment may include a task creation unit 801, a task allocation unit 802 and a data reading unit 803. The task creation unit 801 is used to create a data reading task for the user data in response to a data reading request for the user data; the task allocation unit 802 is used to allocate a data node in an idle state to the data reading task according to the states of the at least two data nodes; the data reading unit 803 is used to use the data node in the idle state to execute the data reading task to read the user data stored in the tape library storage system.
[0145] It should be noted that part or all of the data processing device of this embodiment may be a processing engine located in a network side server, or may also be a distributed system located on the network side, for example, a processing engine or distributed system in a trigger platform on the network side, etc., or may also be located at a specific data node in the cluster, for example, a name node, a control node, etc. This embodiment does not specifically limit this.
[0146] Optionally, in a possible implementation of this embodiment, the tape library storage system includes at least two logical tape library storage subsystems, and each of the at least two logical tape library storage subsystems is logically isolated from each other; the data reading unit 803 is specifically used to determine the tape medium in the logical tape library storage subsystem where the user data is stored; and to use the data node in the idle state to execute the data reading task to read the user data stored on the tape medium in the logical tape library storage subsystem.
[0147] Optionally, in a possible implementation of this embodiment, the tape library storage system stores copy data corresponding to the user data, and the copy data includes at least a first copy and a second copy; the data nodes include at least at least two first nodes corresponding to the first copy and at least two second nodes corresponding to the second copy; the data reading unit 803 is specifically used to determine the first copy or the second copy according to the data reading task; and allocate an idle first node or a second node to the data reading task according to the status of the at least two first nodes corresponding to the first copy or the status of the at least two second nodes corresponding to the second copy.
[0148] Specifically, the data reading unit 803 is specifically configured to use the allocated first node or second node in the idle state to execute the data reading task to read the first copy or the second copy stored in the tape library storage system.
[0149] Correspondingly, the tape library storage system includes at least a first storage system connected to the first node and a second storage system connected to the second node; the data reading unit 803 is specifically used to use the allocated idle first node or second node to execute the data reading task to read the first copy stored in the first storage system or the second copy stored in the second storage system.
[0150] It should be noted that Figure 4-Figure 6 The method in the embodiment corresponding to any of the figures can be implemented by the data processing device provided in this embodiment. Figure 4-Figure 6The relevant contents in the embodiments corresponding to any of the accompanying drawings will not be repeated here.
[0151] In this embodiment, a task creation unit creates a data reading task for the user data in response to a data reading request for the user data, and then a task allocation unit allocates an idle data node to the data reading task according to the status of the at least two data nodes, so that the data reading unit can use the idle data node to execute the data reading task to read the user data stored in the tape library storage system. By sharing the storage resources of the tape library storage system by multiple data nodes in the cluster, the data reading task can be allocated to different data nodes in the cluster for corresponding processing, which can avoid queuing delays in data reading of individual data nodes, thereby improving data processing efficiency.
[0152] In addition, by adopting the technical solution provided by the present invention, the storage capacity of the tape library storage system is utilized to store multiple copy data corresponding to the user data. Based on the data reading task of the user data, an idle data node corresponding to any copy data in the cluster is allocated, and the data node is read from the tape medium in the corresponding tape library storage system, which can effectively improve the security and reliability of data processing.
[0153] In addition, by adopting the technical solution provided by the present invention, the data nodes in the cluster are logically divided into at least a first node and a second node, which correspond to the reading tasks of different replicas respectively, so that the reading tasks can be allocated based on different replica data, thereby reducing the reading delay of a single user data.
[0154] In addition, the technical solution provided by the present disclosure can effectively improve the user experience.
[0155] In the technical solution disclosed herein, the acquisition, storage and application of user data involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0156] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0157] Fig. 9A schematic block diagram of an example electronic device 900 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0158] like Fig. 9 As shown, the electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0159] Multiple components in the electronic device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the electronic device 900 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0160] The computing unit 901 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 901 performs the various methods and processes described above, such as data processing methods. For example, in some embodiments, the data processing method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the data processing method described above may be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to perform the data processing method in any other appropriate manner (e.g., by means of firmware).
[0161] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0162] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0163] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0164] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0165] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0166] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short). The server may also be a server of a distributed system, or a server combined with a blockchain.
[0167] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0168] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A data processing method, applied to a cluster consisting of at least two data nodes, wherein the at least two data nodes are connected to a tape library storage system, the method comprising: In response to a data storage request for user data, creating a data storage task for the user data; Allocating a data node in an idle state to the data storage task according to the states of the at least two data nodes; The data storage task is performed by utilizing the data nodes in the idle state to store the user data in the tape library storage system.
2. The method according to claim 1, wherein: The tape library storage system comprises at least two logical tape library storage subsystems, and each logical tape library storage subsystem in the at least two logical tape library storage subsystems is logically isolated from each other; The step of utilizing the idle data node to execute the data storage task to store the user data in the tape library storage system includes: Obtaining usage status of tape media in each of the at least two logical tape library storage subsystems; According to the usage of the tape media in each logical tape library storage subsystem, the data storage task is executed by utilizing the idle data nodes to write the user data into the tape media in the corresponding logical tape library storage subsystem.
3. The method according to claim 1 or 2, wherein: The data storage task is a multi-copy storage task, including at least a first copy storage task and a second copy storage task; the data nodes include at least two first nodes corresponding to the first copy storage task and at least two second nodes corresponding to the second copy storage task; The allocating an idle data node to the data storage task according to the states of the at least two data nodes comprises: According to the multi-copy storage task, create copy data corresponding to the user data, the copy data at least including a first copy corresponding to the first copy storage task and a second copy corresponding to the second copy storage task; According to the states of the at least two first nodes, allocating an idle first node to the first replica storage task; According to the states of the at least two second nodes, a second node in an idle state is allocated to the second replica storage task.
4. The method according to claim 3, wherein: The step of utilizing the idle data node to execute the data storage task to store the user data in the tape library storage system includes: Utilizing the first node in the idle state, executing the first copy storage task to write the first copy into the tape library storage system; The second copy storage task is executed by utilizing the second node in the idle state to write the second copy into the tape library storage system.
5. The method according to claim 4, wherein: The tape library storage system comprises at least a first storage system connected to the first node and a second storage system connected to the second node; The utilizing the first node in the idle state to execute the first copy storage task to write the first copy to the tape library storage system includes: Using the first node in the idle state, executing the first copy storage task to write the first copy to the first storage system; The utilizing the second node in the idle state to execute the second copy storage task to write the second copy to the tape library storage system comprises: The second copy storage task is performed by utilizing the second node in the idle state to write the second copy into the second storage system.
6. A data processing method, applied to a cluster consisting of at least two data nodes, wherein the at least two data nodes are connected to a tape library storage system, the method comprising: In response to a data reading request of user data, creating a data reading task of the user data; According to the states of the at least two data nodes, allocating a data node in an idle state to the data reading task; The data reading task is performed by utilizing the data nodes in the idle state to read the user data stored in the tape library storage system.
7. The method according to claim 6, wherein: The tape library storage system comprises at least two logical tape library storage subsystems, and each logical tape library storage subsystem in the at least two logical tape library storage subsystems is logically isolated from each other; The utilizing the idle data node to execute the data reading task to read the user data stored in the tape library storage system includes: Determine the tape medium in the logical tape library storage subsystem where the user data is stored; The data reading task is performed by utilizing the data node in the idle state to read the user data stored in the tape medium in the logical tape library storage subsystem.
8. The method according to claim 6 or 7, wherein: The tape library storage system stores copy data corresponding to the user data, wherein the copy data includes at least a first copy and a second copy; the data nodes include at least two first nodes corresponding to the first copy and at least two second nodes corresponding to the second copy; The allocating an idle data node to the data reading task according to the states of the at least two data nodes comprises: Determining the first copy or the second copy according to the data reading task; According to the status of at least two first nodes corresponding to the first replica or the status of at least two second nodes corresponding to the second replica, a first node or a second node in an idle state is allocated to the data reading task.
9. The method according to claim 8, wherein: The utilizing the idle data node to execute the data reading task to read the user data stored in the tape library storage system includes: The data reading task is performed by utilizing the allocated first node or second node in the idle state to read the first copy or the second copy stored in the tape library storage system.
10. The method according to claim 9, wherein: The tape library storage system at least includes a first storage system connected to the first node and a second storage system connected to the second node; the utilizing the first node or the second node in the allocated idle state to execute the data reading task to read the first copy or the second copy stored in the tape library storage system comprises: The data reading task is performed using the allocated first node or second node in the idle state to read the first copy stored in the first storage system or the second copy stored in the second storage system.
11. A data processing device, applied to a cluster consisting of at least two data nodes, wherein the at least two data nodes are connected to a tape library storage system, the device comprising: A task creation unit, configured to create a data storage task for the user data in response to a data storage request for the user data; A task allocation unit, configured to allocate a data node in an idle state to the data storage task according to the states of the at least two data nodes; The data storage unit is used to utilize the data nodes in the idle state to execute the data storage task so as to store the user data in the tape library storage system.
12. The device according to claim 11, wherein The tape library storage system comprises at least two logical tape library storage subsystems, each of which is logically isolated from the other; the data storage unit is specifically used for Obtaining usage status of tape media in each of the at least two logical tape library storage subsystems; as well as According to the usage of the tape media in each logical tape library storage subsystem, the data storage task is executed by utilizing the idle data nodes to write the user data into the tape media in the corresponding logical tape library storage subsystem.
13. The device according to claim 11 or 12, wherein: The data storage task is a multi-copy storage task, including at least a first copy storage task and a second copy storage task; the data nodes include at least two first nodes corresponding to the first copy storage task and at least two second nodes corresponding to the second copy storage task; the task allocation unit is specifically used to According to the multi-copy storage task, create copy data corresponding to the user data, the copy data at least including a first copy corresponding to the first copy storage task and a second copy corresponding to the second copy storage task; According to the states of the at least two first nodes, allocating an idle first node to the first replica storage task; as well as According to the states of the at least two second nodes, a second node in an idle state is allocated to the second replica storage task.
14. The device according to claim 13, wherein: The data storage unit is specifically used for Utilizing the first node in the idle state, executing the first copy storage task to write the first copy into the tape library storage system; as well as The second copy storage task is executed by utilizing the second node in the idle state to write the second copy into the tape library storage system.
15. The device according to claim 14, wherein: The tape library storage system at least includes a first storage system connected to the first node and a second storage system connected to the second node; the data storage unit is specifically used for Using the first node in the idle state, executing the first copy storage task to write the first copy to the first storage system; as well as The second copy storage task is performed by utilizing the second node in the idle state to write the second copy into the second storage system.
16. A data processing device, applied to a cluster consisting of at least two data nodes, wherein the at least two data nodes are connected to a tape library storage system, the device comprising: A task creation unit, configured to create a data reading task for the user data in response to a data reading request for the user data; A task allocation unit, configured to allocate a data node in an idle state to the data reading task according to the states of the at least two data nodes; The data reading unit is used to use the data nodes in the idle state to execute the data reading task to read the user data stored in the tape library storage system.
17. The device according to claim 16, wherein: The tape library storage system comprises at least two logical tape library storage subsystems, each of which is logically isolated from each other; the data reading unit is specifically used for Determine the tape medium in the logical tape library storage subsystem where the user data is stored; as well as The data reading task is performed by utilizing the data node in the idle state to read the user data stored in the tape medium in the logical tape library storage subsystem.
18. The device according to claim 16 or 17, wherein: The tape library storage system stores copy data corresponding to the user data, and the copy data includes at least a first copy and a second copy; the data nodes include at least two first nodes corresponding to the first copy and at least two second nodes corresponding to the second copy; the data reading unit is specifically used to Determining the first copy or the second copy according to the data reading task; as well as According to the status of at least two first nodes corresponding to the first replica or the status of at least two second nodes corresponding to the second replica, a first node or a second node in an idle state is allocated to the data reading task.
19. The device according to claim 18, wherein: The data reading unit is specifically configured to use the allocated first node or second node in the idle state to execute the data reading task, so as to read the first copy or the second copy stored in the tape library storage system.
20. The device according to claim 19, wherein The tape library storage system at least includes a first storage system connected to the first node and a second storage system connected to the second node; the data reading unit is specifically used to The data reading task is performed using the allocated first node or second node in the idle state to read the first copy stored in the first storage system or the second copy stored in the second storage system.
21. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1-5 or 6-10.
22. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-5 or 6-10.
23. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1-5 or 6-10.
Citation Information
Patent Citations
Parallel data transmission method and device, medium and electronic equipment
CN110677463A
Data processing method and device, electronic equipment and storage medium
CN111767169A
Data storage method and device, equipment, storage medium and program
CN113220650A
Backup performance analysis method for backup link node
CN115454801A
Data storage method and device, computer equipment and storage medium
CN117631953A