File synchronization method, storage medium and program method
By performing transactional splitting and concurrent control of large-scale file systems, the high memory usage and resource waste problems during RSYNC synchronization are solved, and the stable and efficient synchronization of the file system is achieved.
Patent Information
- Application Number
- CN202411813432.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-05-06
AI Technical Summary
When a large-scale file system is synchronized using RSYNC software, it will cause large-scale concurrent data writing, occupy high memory, cause client lag, and even system service abnormalities, and connect to the server for a long time, resulting in high resource usage and affect service capabilities.
By conducting transactional splitting of the file system, multiple file tasks are obtained, and the concurrency of file tasks is determined based on the client's data writeback speed, bandwidth and bandwidth of the server, and the scheduling task to the target client to perform file synchronization.
It realizes stable synchronization of large-scale file systems, avoids client lag and resource waste, and improves business processing efficiency and system stability.
Smart Images

Figure CN119938253A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of file synchronization, and relate to but are not limited to a file synchronization method, a storage medium, and a program method. Background Art
[0002] The synchronization of file systems is usually handled by RSYNC software, which can synchronize directories and files between different computers in a heterogeneous environment. When a large-scale file system uses RSYNC for synchronization, it will cause a large number of concurrent data writes, high memory usage, client lag, and even system service abnormalities; long-term connection to the server will result in high server resource usage, waste of server resources, and affect the ability to provide services. Therefore, the above-mentioned real-time file synchronization method cannot adapt to large-scale file system scenarios. Summary of the invention
[0003] In view of this, in order to solve the problem that the method of synchronizing files cannot adapt to large-scale file system scenarios, the embodiments of the present application provide a file synchronization method, a storage medium and a program method.
[0004] The technical solution of the embodiment of the present application is implemented as follows:
[0005] In a first aspect, an embodiment of the present application provides a file synchronization method, which is applied to a server, and the method includes:
[0006] Perform transactional splitting on a file system including multiple files to obtain at least two file tasks;
[0007] The concurrency number of the file task is determined based on at least one of the following: the data write-back speed of each client, the maximum bandwidth of each client, the maximum bandwidth of the server, and the preset file synchronization bandwidth, wherein at least two of the clients are used to synchronize the files of the server to at least one file storage node;
[0008] Based on the concurrency number of the file tasks, the target file task is scheduled to the target client to perform file synchronization, so as to synchronize the target file task to at least one of the file storage nodes.
[0009] In a second aspect, an embodiment of the present application provides a file synchronization method, which is applied to a client, and the method includes:
[0010] Recording temporary files generated during synchronization of a target file on the server to at least one file storage node;
[0011] generating a first file list based on the temporary file;
[0012] Send the first file list to the server.
[0013] In a third aspect, an embodiment of the present application provides a file synchronization method, which is applied to a file synchronization system, wherein the file synchronization system includes a server, a client, and a file storage node, and the method includes:
[0014] The server performs transactional splitting on a file system including multiple files to obtain at least two file tasks;
[0015] The server determines the number of concurrency of the file tasks based on at least one of the following: the data write-back speed of each of the clients, the maximum bandwidth of each of the clients, the maximum bandwidth of the server, and a preset file synchronization bandwidth, wherein at least two of the clients are used to synchronize the files of the server to at least one file storage node;
[0016] The server dispatches the target file task to the target client based on the concurrent number of the file task;
[0017] The target client performs file synchronization to synchronize the target file task to at least one of the file storage nodes.
[0018] In a fourth aspect, an embodiment of the present application provides a file synchronization device, the device comprising:
[0019] A splitting module, used for transactionally splitting a file system including multiple files to obtain at least two file tasks;
[0020] A determination module, configured to determine the number of concurrency of the file task based on at least one of the following: a data write-back speed of each client, a maximum bandwidth of each of the clients, a maximum bandwidth of the server, and a preset file synchronization bandwidth, wherein at least two of the clients are used to synchronize the files of the server to at least one file storage node;
[0021] The scheduling module is used to schedule the target file task to the target client to perform file synchronization based on the concurrent number of the file task, so as to synchronize the target file task to at least one of the file storage nodes.
[0022] In a fifth aspect, an embodiment of the present application provides a file synchronization device, the device comprising:
[0023] A recording module, used for recording temporary files generated in the process of synchronizing the target file of the server to at least one file storage node;
[0024] A generating module, used for generating a first file list based on the temporary file;
[0025] A sending module is used to send the first file list to the server.
[0026] In a sixth aspect, an embodiment of the present application provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, and the processor implements the above method when executing the program.
[0027] In a seventh aspect, an embodiment of the present application provides a storage medium storing executable instructions for implementing the above method when executed by a processor.
[0028] In an eighth aspect, an embodiment of the present application provides a computer program product, including a computer program or instructions, which, when executed by a processor, implement the steps in the above method.
[0029] In an embodiment of the present application, the large file system is first split according to the transactional nature of the business itself, so as to split the synchronization task of the large-scale file system into multiple file tasks that can provide atomic capabilities, at least two file tasks. Then, according to the data write-back speed of the client writing to the storage node, the client bandwidth and the server bandwidth, multiple file tasks can be concurrently controlled. The atomic splitting of tasks and concurrent control are achieved, which improves the business processing efficiency and ensures the stability of large-scale synchronization file business. Finally, based on the concurrent number of file tasks, the target file task is scheduled to the target client to perform file synchronization, so that the target file task can be synchronized to at least one file storage node. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 A schematic diagram of an implementation flow of a file synchronization method provided in an embodiment of the present application;
[0031] Figure 2 A schematic diagram of an implementation flow of a file synchronization method provided in an embodiment of the present application;
[0032] Figure 3 A schematic diagram of an implementation flow of a file synchronization method provided in an embodiment of the present application;
[0033] Figure 4A A schematic diagram of a file synchronization process provided by an embodiment of the present application;
[0034] Figure 4B A flowchart of concurrent control of file synchronization tasks provided in an embodiment of the present application;
[0035] Figure 5A A schematic diagram of the structure of a file synchronization device provided in an embodiment of the present application;
[0036] Figure 5B A schematic diagram of the structure of a file synchronization device provided in an embodiment of the present application;
[0037] Figure 6A hardware entity schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0038] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the specific technical solution of the embodiments of the present application will be further described in detail below in conjunction with the drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application, but are not used to limit the scope of the present application.
[0039] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0040] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0041] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0042] The present application embodiment provides a file synchronization method, which is applied to a server, such as Figure 1 As shown, the method includes:
[0043] Step S110: transactionally split the file system including multiple files to obtain at least two file tasks;
[0044] Here, the file system including multiple files may be a large-scale file system. In the implementation process, a threshold value of the number of files may be set to determine that the file system is a large-scale file system when the number of files is greater than the threshold value of the number of files.
[0045] Transactional splitting refers to breaking down files into multiple transactions (single atomic functions) in the file system to ensure that a business can be run or implemented independently based on the files in each transaction. After splitting, it can be ensured that a single RSYNC task can be used continuously as a single atomic function. Among them, RSYNC is a data synchronization tool. The synchronization of the file system can be processed using RSYNC software. It can synchronize directories and files between different computers in a heterogeneous environment, and effectively transfer part of the updated content to the target file system of the target machine, thereby reducing the amount of data transmission and improving synchronization efficiency. After splitting, a RSYNC synchronization task with a large amount of data, including different transactions, can be split into multiple atomic RSYNC tasks.
[0046] During the implementation process, you can put files with strong relevance and frequent operations together according to the business logic to form an independent business unit, namely, a file task. A file task can include multiple files.
[0047] Take the mirror source website as an example. The data scale of the mirror source website can vary depending on the number of data sources, but usually the data scale of a mirror source is about 2 terabytes (T), and each mirror source contains multiple subdirectories. For example: the combination of repodata and Packages is the atomic capability that a mirror source can provide services, so the file synchronization task split by RSYNC can be a combination of repodata and Packages. Among them, the Repodata directory contains the metadata index file of the software repository, and the Packages directory (or a similarly named directory) is used to store the actual software package files.
[0048] Step S120, based on at least one of the following: the data write-back speed of each client, the maximum bandwidth of each client, the maximum bandwidth of the server, and the preset file synchronization bandwidth, determining the number of concurrency of the file task, wherein at least two of the clients are used to synchronize the files of the server to at least one file storage node;
[0049] Here, the data write-back speed of the client refers to the speed at which the file storage node processes the file synchronization request and returns the result data to the client after the client sends the file synchronization request to the file storage node. The storage system including at least one file storage node may be a distributed file storage system, and the data write-back speed of the client may be the speed at which the client writes the synchronization file from the client's memory back to the disk of the distributed file storage system.
[0050] The maximum bandwidth of a client refers to the maximum data transfer rate that a client device (node) can support over a network connection. This rate can be measured in bits per second (bps). The maximum bandwidth of a client refers to the maximum rate at which a client device can transfer data over a network connection under ideal conditions.
[0051] The maximum bandwidth of the server refers to the maximum data transfer rate that the file synchronization server can support in the network connection. This rate can determine the server's ability to handle concurrent requests and data traffic. It is measured in bits per second (bps) and represents the maximum transmission capacity of the server when there are no restrictions.
[0052] The preset file synchronization bandwidth refers to the expected or maximum data transfer rate set for the system or application during the file synchronization process. It determines the amount of data that the system can handle during the synchronization process. For example, the RSYNC synchronization bandwidth refers to the data transfer rate that can be achieved when using RSYNC for file synchronization. During the implementation process, you can set the RSYNC synchronization bandwidth. The RSYNC software provides a synchronization bandwidth parameter to limit the bandwidth usage during the RSYNC synchronization process. This parameter can help users avoid RSYNC occupying all network resources when the network bandwidth is limited or when it is necessary to avoid affecting other services. When using the synchronization bandwidth parameter, you can specify a speed value, which represents the maximum bandwidth allowed for RSYNC synchronization. The speed value can be in kilobytes per second (KB / s) or other specified units.
[0053] The concurrency of file tasks refers to the number of file tasks that can be processed concurrently by the client communicating with the server at the current moment. In the implementation process, at least two clients can be set to communicate with the server to synchronize the files of the server to at least one file storage node.
[0054] During the implementation process, since a server can execute file synchronization tasks for multiple clients, the data write-back speed of each client and the maximum bandwidth of each client may affect the number of concurrent file tasks. At the same time, the maximum bandwidth of the server itself and the preset file synchronization bandwidth may also affect the number of concurrent file tasks, so the number of concurrent file tasks can be determined based on at least one of the following: the data write-back speed of each client, the maximum bandwidth of each client, the maximum bandwidth of the server, and the preset file synchronization bandwidth.
[0055] Step S130: Based on the concurrency number of the file task, the target file task is scheduled to the target client to perform file synchronization, so as to synchronize the target file task to at least one of the file storage nodes.
[0056] Here, the target file tasks can be arranged in the order of the total file tasks to be executed based on the number of concurrent file tasks. The arrangement order can be arranged based on the importance level of the file tasks or based on the task volume of the file tasks, and the basis of the arrangement order is not limited.
[0057] For example, a task queue (Qtask) can be established for each RSYNC server, the length of the task queue is equal to the number of tasks after each server is split, and all the split file tasks are put into the queue. Then the target file task is determined in the task queue.
[0058] The target client may be determined based on available resources of the client among the clients communicating with the server. For example, the target client may be determined based on at least one of the following parameters of the client: available data write-back speed, remaining bandwidth, central processing unit (CPU) utilization, memory utilization, disk utilization, etc.
[0059] During implementation, the target file task determined based on the concurrent number of file tasks can be scheduled to the target client, and the target client performs file synchronization to synchronize the target file task to the storage node on the distributed file storage system.
[0060] In an embodiment of the present application, the large file system is first split according to the transactional nature of the business itself, so as to split the synchronization task of the large-scale file system into multiple file tasks that can provide atomic capabilities, at least two file tasks. Then, according to the data write-back speed of the client writing to the storage node, the client bandwidth and the server bandwidth, multiple file tasks can be concurrently controlled. The atomic splitting of tasks and concurrent control are achieved, which improves the business processing efficiency and ensures the stability of large-scale synchronization file business. Finally, based on the concurrent number of file tasks, the target file task is scheduled to the target client to perform file synchronization, so that the target file task can be synchronized to at least one file storage node.
[0061] In some embodiments, the above step S120 "determining the number of concurrent file tasks based on at least one of the following: the data write-back speed of each client, the maximum bandwidth of each client, the maximum bandwidth of the server, and the preset file synchronization bandwidth" can be implemented by the following steps:
[0062] Step S121, determining the remaining write-back speed of each of the clients based on the data write-back speed of each of the clients and the file synchronization bandwidth;
[0063] In the implementation process, the remaining write-back speed of the client, that is, the speed Vmem(q) at which available client data is written back from the memory to the disk, can be calculated using the following formula (1):
[0064]
[0065] Where n represents a node, and t represents a file synchronization task. Rband(n,t) represents the download bandwidth occupied by file synchronization task t on node n. q represents a queue, and Vmem(q) represents the maximum disk write-back speed that can be used by queue q (the remaining write-back speed of the client).
[0066] Step S122: determining the remaining bandwidth of each client based on the maximum bandwidth of each client and the file synchronization bandwidth;
[0067] In the implementation process, the remaining bandwidth of the client, that is, the available bandwidth Vband(q), can be calculated using the following formula (2):
[0068]
[0069] Where n represents a node, and t represents a file synchronization task. Rband(n,t) represents the download bandwidth occupied by file synchronization task t on node n. q represents a queue, Vmem(q) represents the maximum disk write-back speed that queue q can use (the speed at which available client data is written back from memory to disk), and Vband(q) represents the maximum client network card bandwidth that queue q can use (the remaining bandwidth of the client).
[0070] Step S123: determining the remaining bandwidth of the server based on the maximum bandwidth of the server and the file synchronization bandwidth;
[0071] In the implementation process, the remaining bandwidth of the server, that is, the maximum available server bandwidth Vrband(q) of the queue q of the server, can be calculated using the following formula (3):
[0072]
[0073] Among them, Vrband(all) represents the total bandwidth of the server. Indicates the bandwidth used by the server.
[0074] Step S124: Determine the number of concurrency of the file tasks based on the remaining write-back speeds of all the clients, the remaining bandwidths of all the clients, and the remaining bandwidth of the server.
[0075] Here, since the remaining write-back speed of all clients, the remaining bandwidth of all clients and the remaining bandwidth of the server are the remaining resources that can be used for the file synchronization that can be called this time, the concurrency number of file tasks can be determined based on the above parameters.
[0076] In the embodiment of the present application, the remaining write-back speed of each client can be determined based on the data write-back speed and file synchronization bandwidth of each client; the remaining bandwidth of each client can be determined based on the maximum bandwidth and file synchronization bandwidth of each client; the remaining bandwidth of the server can be determined based on the maximum bandwidth and file synchronization bandwidth of the server; finally, the remaining write-back speed of all clients, the remaining bandwidth of all clients and the remaining bandwidth of the server are comprehensively considered to determine the concurrent number of file tasks that can be executed in this file synchronization.
[0077] In some embodiments, the above step S124 "determining the number of concurrent file tasks based on the remaining write-back speeds of all the clients, the remaining bandwidths of all the clients, and the remaining bandwidth of the server" can be implemented by the following steps:
[0078] Step 1241: determine the minimum value among the remaining write-back speeds of all the clients, the remaining bandwidths of all the clients, and the remaining bandwidth of the server as the maximum concurrent bandwidth;
[0079] In the implementation process, the maximum concurrent bandwidth, that is, the maximum concurrent bandwidth Bmax(q) actually available for queue q, can be calculated using the following formula (4):
[0080] Bmax(q)=MIN(Vmem(q),Vband(q),Vrband(q)) (4);
[0081] Among them, Bmax(q) is the minimum value of Vmem(q), Vband(q) and Vrband(q).
[0082] Step 1242: Determine the concurrency number of the file task based on the maximum concurrent bandwidth and the file synchronization bandwidth.
[0083] During the implementation process, the current concurrent number p(q) that can be run can be calculated using the following formula (5):
[0084] p(q)=Bmax(q) / Rband(q) / 2+c (5);
[0085] Here, c represents the incentive factor, and its initial value is 1. The reason for setting the incentive factor is that after calculating the number of concurrent tasks Bmax(q) / Rband(q), the number of concurrent tasks is divided by 2, and then a slow start method is used to ensure that the tasks with a concurrent number of p(q) can be effectively executed.
[0086] In some embodiments, after calculating the concurrent number Bmax(q) / Rband(q), the divisor can be set according to actual conditions, such as 3, 4, etc., and is not limited to 2.
[0087] In the embodiment of the present application, the minimum value among the remaining write-back speed of all clients, the remaining bandwidth of all clients and the remaining bandwidth of the server is determined as the maximum concurrent bandwidth. The concurrency number of the file task is determined based on the maximum concurrent bandwidth and the file synchronization bandwidth. In this way, the concurrency number of the file task can be obtained, which enables the server to effectively and reasonably schedule the target file synchronization task that meets the concurrency number to the target client to perform file synchronization.
[0088] In some embodiments, the above step S130 "scheduling the target file task to the target client to perform file synchronization based on the concurrent number of the file task" can be implemented by the following steps:
[0089] Step S131: When it is determined that the remaining write-back speed of the client and / or the remaining bandwidth of the client are constraints for limiting the maximum concurrent bandwidth, a target file smaller than the concurrent number is scheduled to the target client for file synchronization;
[0090] Here, the remaining write-back speed of the client and / or the remaining bandwidth of the client are determined as the constraints for limiting the maximum concurrent bandwidth, which can be understood as the bottleneck for determining the value of Bmax(q). If the bottleneck is Vmem(q) and / or Vband(q) (the client is the limiting termination condition), then a target client can be selected to complete the scheduling of target files less than the concurrent number, that is, the target files less than the concurrent number are scheduled to the target client for file synchronization.
[0091] In some embodiments, the target client, i.e., the client that meets the requirements, can select the client with the smallest speed of writing client data from memory to disk and / or the smallest available client bandwidth to maximize the utilization of the client that has already executed the file synchronization task, effectively improving the utilization rate of the client.
[0092] In some embodiments, the target client, i.e., the client that meets the requirements, can select the client with the maximum speed of writing client data from memory to disk and / or the maximum available client bandwidth to achieve load balancing, so that each client in the client cluster can achieve balanced usage.
[0093] Step S132: When it is determined that the remaining bandwidth of the server is a constraint condition for limiting the maximum concurrent bandwidth, the target file that meets the concurrent number is scheduled to the target client for file synchronization;
[0094] The target client is determined based on the remaining write-back speed and / or the remaining bandwidth.
[0095] Here, the remaining bandwidth of the server is determined as the constraint condition for limiting the maximum concurrent bandwidth, which can be understood as determining the value bottleneck of Bmax(q). If the bottleneck is Vrband(q) (the server is the limit termination condition), the target file is preferentially scheduled to a client (node) that meets the requirements to execute p(q) synchronization tasks.
[0096] In the embodiment of the present application, when it is determined that the remaining write-back speed of the client and / or the remaining bandwidth of the client are the constraints for limiting the maximum concurrent bandwidth, the target files less than the concurrent number are scheduled to the target client for file synchronization. In this way, when the client resources are limited, the appropriate number of target files can be scheduled to the target client for file synchronization, thereby improving the success rate of file synchronization.
[0097] When it is determined that the remaining bandwidth of the server is a constraint condition for limiting the maximum concurrent bandwidth, the target files that meet the concurrent number are scheduled to the target client for file synchronization. In this way, when the server resources are limited, the client resources can be fully utilized, and the target files that meet the concurrent number can be scheduled to the target client for file synchronization, thereby improving the efficiency of file synchronization.
[0098] In some embodiments, the above file synchronization method further includes the following steps:
[0099] Step S140: Obtain a first file list sent by the target client, wherein the first file list records temporary files generated in synchronizing the target file;
[0100] Here, the first list can be a file mapping table created based on the files stored in the temporary directory, and the {key: value} of the mapping content is: {NodeID of the target file: NodeID of the temporary file}. The file mapping table will be deleted after the synchronization is successful, and will be retained if the synchronization fails.
[0101] During implementation, the server may obtain the first file list when file synchronization fails.
[0102] Step S150: Compare the first file list with the second file list corresponding to the target file to generate a third file list to be synchronized;
[0103] Here, the second file list may be a source file list of the server, that is, a file list to be synchronized.
[0104] In the implementation process, when the RSYNC synchronization algorithm constructs the destination file list (filelist), it first scans the file system on the client to generate the destination file list, and then checks whether there is a file mapping table (first file list). If so, the mapped temporary file will be replaced with the original file in the destination file list based on the first file list, thereby updating the destination file list; if so, the destination file list will not be updated. Finally, the latest file list (third file list) can be generated by comparing the source file list and the destination file list on the server.
[0105] Correspondingly, the above step S130 "scheduling the target file task to the target client to perform file synchronization based on the concurrent number of the file task" can be implemented by the following process:
[0106] Based on the concurrency number of the file task and the third file list, the target file task is scheduled to the target client to perform file synchronization.
[0107] Here, the third file list lists the file tasks to be synchronized. Based on the concurrency number of the file tasks, the target files that meet the concurrency number can be determined in the third file list, and the target file tasks can be scheduled to the target client to perform file synchronization.
[0108] In the embodiment of the present application, the first file list of temporary files generated in the target file for synchronization is first obtained, and then the first file list is compared with the second file list corresponding to the target file to generate a third file list to be synchronized; finally, based on the concurrency number of the file task and the third file list, the target file task is scheduled to the target client to perform file synchronization. In this way, by comparing the first file list with the second file list, the third file list is generated, and when the synchronization file is determined based on the third file list, the temporary file will not be retransmitted. The problem of repeated transmission of temporary files is avoided, so that in the case of complex network environment and temporary interruption of business, the efficiency of file synchronization update is improved, and the consistency of the final file synchronization is guaranteed.
[0109] The present application embodiment provides a file synchronization method, which is applied to a client, such as Figure 2 As shown, this can be achieved by following the steps below:
[0110] Step S210, recording a temporary file generated in the process of synchronizing the target file of the server to at least one file storage node;
[0111] During the implementation process, a file mapping table can be created based on the files stored in the temporary directory, and the {key: value} of the mapping content is: {NodeID of the target file: NodeID of the temporary file}. The file mapping table will be deleted after the synchronization is successful, and will be retained if the synchronization fails.
[0112] Step S220: Generate a first file list based on the temporary file;
[0113] During implementation, the first file list may be generated based on the file identifier recorded in the temporary file, wherein the file identifier may be the name, identification number, etc. of the file.
[0114] Step S230: Send the first file list to the server.
[0115] During the implementation process, when the file synchronization task fails to execute and the server is reconnected to execute the file synchronization task, the first file list can be sent to the server. When the server constructs the target file list (filelist) using the RSYNC synchronization algorithm, it first scans the file system on the client to generate the target file list, and then checks whether there is a file mapping table (first file list). If so, the mapped temporary file will be replaced with the original file in the target file list based on the first file list, thereby updating the target file list; if so, the target file list will not be updated.
[0116] In an embodiment of the present application, the client first records the temporary files generated in the process of synchronizing the target files of the server to at least one file storage node; then generates a first file list based on the temporary files; and finally sends the first file list to the server. In this way, the server can determine the files to be synchronized when synchronizing files based on the first file list, without retransmitting temporary files. The problem of repeated transmission of temporary files is avoided, thereby improving the efficiency of file synchronization and updating in the case of complex network environment and temporary business interruption, and ensuring the consistency of the final file synchronization.
[0117] The present application embodiment provides a file synchronization method, which is applied to a file synchronization system, wherein the file synchronization system includes a server, a client, and a file storage node, such as Figure 3 As shown, this can be achieved by following the steps below:
[0118] Step S310: The server performs transactional splitting on a file system including multiple files to obtain at least two file tasks;
[0119] Step S320: The server determines the number of concurrency of the file task based on the data write-back speed of each of the clients, the maximum bandwidth of each of the clients, the maximum bandwidth of the server, and the preset file synchronization bandwidth, wherein at least two of the clients are used to synchronize the files of the server to at least one file storage node;
[0120] Step S330: the server dispatches the target file task to the target client based on the concurrent number of the file task;
[0121] Step S340: The target client performs file synchronization to synchronize the target file task to at least one of the file storage nodes.
[0122] In the embodiment of the present application, the server first splits the large file system according to the transactional nature of the business itself, and splits the synchronization task of the large-scale file system into multiple file tasks that can provide atomic capabilities. Then the server can perform concurrent control on multiple file tasks based on the data write-back speed of the client writing to the storage node, the client bandwidth and the server bandwidth. The atomic splitting of tasks and concurrent control are achieved, which improves the business processing efficiency and ensures the stability of the business. Finally, based on the concurrent number of file tasks, the server schedules the target file task to the target client to perform file synchronization, and can use the target client to synchronize the target file task to at least one file storage node.
[0123] In some embodiments, the above file synchronization method further includes the following steps:
[0124] Step S350: the target client records a temporary file generated in the process of synchronizing the target file of the server to at least one of the file storage nodes, generates a first file list based on the temporary file, and sends the first file list to the server;
[0125] Step S360: The server compares the first file list with the second file list corresponding to the target file to generate a third file list to be synchronized;
[0126] Correspondingly, the above step S330 "the server dispatches the target file task to the target client based on the concurrent number of the file task" can be implemented by the following process:
[0127] The server schedules the target file task to the target client based on the concurrent number of the file task and the third file list.
[0128] In the embodiment of the present application, by comparing the first file list with the second file list, a third file list is generated, and when the synchronization file is determined based on the third file list, the temporary file will not be retransmitted. The problem of repeated transmission of temporary files is avoided, thereby improving the efficiency of file synchronization and updating in a complex network environment and temporary service interruption, and ensuring the consistency of the final file synchronization.
[0129] Figure 4A A schematic diagram of a file synchronization process provided by an embodiment of the present application, such as Figure 4A As shown, this can be achieved by following the steps below:
[0130] Step S401, task splitting;
[0131] During the implementation process, the large file system can be split according to the transactional nature of the business itself. After the split, it is guaranteed that a single RSYNC task can be used continuously as a single atomic function.
[0132] Here, we take the mirror source website as an example. The data scale of the mirror source website can vary depending on the number of data sources, but usually the data scale of a mirror source is around 2T, and each mirror source contains multiple subdirectories. For example: The combination of repodata and Packages is the atomic capability that a mirror source can provide services, so the file synchronization task split by RSYNC can be a combination of repodata and Packages. Among them, the Repodata directory contains the metadata index file of the software repository, and the Packages directory (or a similarly named directory) is used to store the actual software package files.
[0133] The task (Task) records parameters such as the synchronized folder, RSYNC bandwidth (Rband), etc.
[0134] like Figure 4A As shown, task A on server A can be split into taskA1, taskA*, taskAi, etc.; task B on server B can be split into taskB1, taskB*, taskBi, etc.
[0135] Step S402: construct a task queue;
[0136] During the implementation process, a task queue (Qtask) may be established for each RSYNC server. The length of the task queue is equal to the number of tasks split from each server, and all the split tasks are put into the queue.
[0137] like Figure 4A As shown, task queue A (QtaskA) and task queue B (QtaskB) can be combined.
[0138] Step S403: concurrent control;
[0139] During the implementation process, the maximum concurrency number p(q) that can be used by the task queue q can be calculated, and the target client can be used to execute the file synchronization of the tasks that meet the maximum concurrency number.
[0140] You can perform the following steps to determine the maximum number of concurrent file synchronizations:
[0141] Step 1: For each task queue, use the following formulas (1) and (2) to calculate the speed Vmem(q) at which available client data is written back from memory to disk and the available client bandwidth Vband(q):
[0142]
[0143] Where n represents a node, and t represents a file synchronization task. Rband(n,t) represents the download bandwidth occupied by file synchronization task t on node n. q represents a queue, Vmem(q) represents the maximum disk write-back speed that queue q can use (the speed at which available client data is written back from memory to disk), and Vband(q) represents the maximum client network card bandwidth that queue q can use (client available bandwidth).
[0144] The following formula (3) is used to calculate the maximum available server bandwidth Vrband(q) of queue q:
[0145]
[0146] Among them, Vrband(all) represents the total bandwidth of the server. Indicates the bandwidth used by the server.
[0147] In the implementation process, the following formula (4) is used to calculate the maximum concurrent bandwidth Bmax(q) actually available to queue q:
[0148] Bmax(q)=MIN(Vmem(q),Vband(q),Vrband(q)) (4);
[0149] Among them, Bmax(q) is the minimum value of Vmem(q), Vband(q) and Vrband(q).
[0150] Step 2: Calculate the number of concurrent file synchronization tasks p that can currently be run;
[0151] During the implementation, the following formula (5) is used to calculate the current concurrent number p(q):
[0152] p(q)=Bmax(q) / Rband(q) / 2+c (5);
[0153] Here, c represents the incentive factor, and its initial value is 1. The reason for setting the incentive factor is that after calculating the possible concurrent number Bmax(q) / Rband(q), the possible concurrent number is divided by 2, and then the slow start method is adopted to ensure that the task with a concurrent number of p(q) can be effectively executed.
[0154] During the implementation process, there are currently p(q) parallel tasks that can be executed immediately. At this time, it is necessary to determine the bottleneck of the value of Bmax(q). If the bottleneck is Vrband(q) (the server is the limit termination condition), the file will be scheduled to a client (node) that meets the requirements to execute p(q) synchronization tasks; if the bottleneck is Vmem(q) or Vband(q) is the bottleneck (the client is the limit termination condition), a client will be selected to complete a task scheduling, and p(q) will be recalculated until the value of p(q) is less than 1, or the Vrband(q) termination condition is met.
[0155] In some embodiments, a client that meets the requirements may select a client with the smallest speed of writing client data from memory to disk and / or the smallest available client bandwidth to maximize the utilization of the client that has already executed the file synchronization task, effectively improving the utilization rate of the client.
[0156] In some embodiments, a client that meets the requirements may select a client with the maximum speed of writing client data from memory to disk and / or the maximum available client bandwidth to achieve load balancing, so that each client in the client cluster can achieve balanced usage.
[0157] Step S404: Task execution.
[0158] This can be achieved by following the steps below:
[0159] Step 3: For the concurrency value of the task queue, if the file synchronization task is executed successfully, the incentive factor is increased by 1; if the file synchronization task is executed abnormally, the concurrency number is reduced by 1;
[0160] Step 4: When any file synchronization task is completed, p(q) can be recalculated for the task queue to which the file synchronization task belongs to determine whether to continue executing concurrent tasks;
[0161] Step 5: After all tasks in a task queue are completed, they can be restarted by scheduled tasks or manual triggering.
[0162] During the synchronization process, the embodiment of the present application also updates the generation algorithm of filelist in the RSYNC tool, which can be implemented by the following steps:
[0163] Step A: The real-time access synchronization method can store the files in the synchronization process in a temporary directory (.~tmp~). Each time synchronization starts, a file mapping table is created based on the files stored in the temporary directory. The {key: value} of the mapping content is: {NodeID of the target file: NodeID of the temporary file}. The file mapping table will be deleted after the synchronization is successful, and will be retained if the synchronization fails.
[0164] Step B: When the RSYNC synchronization algorithm constructs the destination file list (filelist), it first scans the file system on the client to generate the destination file list, and then checks whether there is a file mapping table. If so, the mapped temporary file is replaced with the original file in the destination file list, thereby updating the destination file list; if so, the destination file list is not updated. Finally, the latest file list can be generated by comparing the source file list and the destination file list on the server.
[0165] Figure 4B A flowchart of concurrent control of file synchronization tasks provided in an embodiment of the present application is shown as follows: Figure 4B As shown, this can be achieved by following the steps below:
[0166] Step S411, calculating the available client write-back speed, the available client bandwidth and the available server bandwidth;
[0167] Step S412: Calculate the maximum concurrent bandwidth and the maximum number of tasks;
[0168] Step S413, determine whether the maximum number of tasks is greater than 0;
[0169] If it is determined that the maximum number of tasks is greater than 0, step S414 is executed; if it is determined that the maximum number of tasks is less than 0, the process ends.
[0170] Step S414: determine whether the server termination condition is met;
[0171] If it is determined that the server-side termination condition is met, step S415 is executed; if it is determined that the server-side termination condition is not met, that is, the client-side termination condition is met, step S416 is executed.
[0172] Step S415: Select the client with the highest score to execute the file synchronization task with the maximum number of tasks;
[0173] In some embodiments, the client with the highest score can select the client with the smallest speed of writing client data from memory to disk and / or the smallest client available bandwidth to maximize the utilization of the client that has already executed the file synchronization task, effectively improving the utilization rate of the client.
[0174] In some embodiments, the client with the highest score can select the client with the largest speed of writing client data from memory to disk and / or the largest client available bandwidth to achieve load balancing, so that each client in the client cluster can achieve balanced usage.
[0175] Here, since the server sets a termination condition that limits the maximum number of tasks, the client can determine that it can execute the maximum number of tasks, and can schedule the maximum number of file synchronization tasks to the client with the highest score for execution.
[0176] Step S416: Select the client with the highest score to execute a file synchronization task;
[0177] Here, since the termination condition for limiting the maximum number of tasks is not the server, but the client, the file synchronization task with a smaller number of tasks than the maximum number can be scheduled to the client with the highest score for execution.
[0178] In some embodiments, this process can schedule one file synchronization task to the client with the highest score to improve the synchronization success rate of the file synchronization task.
[0179] Step S417: Every time a file synchronization is completed, the excitation factor is updated.
[0180] During the implementation process, for the concurrency value of the task queue, that is, the file synchronization task executed in this round, if the file synchronization task is executed successfully, the incentive factor is increased by 1; if the file synchronization task is executed abnormally, the concurrency number is reduced by 1.
[0181] The above is based on Figure 4A and Figure 4B In the application embodiment provided, a large-scale file system synchronization method suitable for real-time access is provided. For the synchronization task of large data volume, it is split into multiple subtasks that can provide atomic capabilities. For multiple subtasks, a concurrency control algorithm is designed to perform concurrency control. In this way, the atomic splitting of tasks and concurrency control are achieved to ensure the stability of the business.
[0182] In order to solve the problem of repeated transmission of temporary files, a method of updating the source file tree is proposed to ensure that temporary files will not be retransmitted when performing file synchronization tasks based on the generated latest file list. In this way, the problem of repeated transmission of temporary files is effectively avoided, thereby improving the efficiency of file synchronization updates in complex network environments and temporary business interruptions, and ensuring the consistency of the final file synchronization.
[0183] Based on the foregoing embodiments, the embodiments of the present application provide a file synchronization device, which includes the modules included, each module includes each sub-module, each sub-module includes a unit, which can be implemented by a processor in an electronic device; of course, it can also be implemented by a specific logic circuit; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP) or a field programmable gate array (FPGA), etc.
[0184] Figure 5A A schematic diagram of the structure of the file synchronization device provided in the embodiment of the present application is shown in FIG. Figure 5A As shown, the device 500 includes:
[0185] A splitting module 501 is used to perform transactional splitting on a file system including multiple files to obtain at least two file tasks;
[0186] The determination module 502 is used to determine the number of concurrency of the file task based on at least one of the following: the data write-back speed of each client, the maximum bandwidth of each client, the maximum bandwidth of the server, and the preset file synchronization bandwidth, wherein at least two of the clients are used to synchronize the files of the server to at least one file storage node;
[0187] The scheduling module 503 is used to schedule the target file task to the target client to perform file synchronization based on the concurrent number of the file task, so as to synchronize the target file task to at least one of the file storage nodes.
[0188] In some embodiments, the determination module 502 includes a first determination submodule, a second determination submodule, a third determination submodule and a fourth determination submodule, wherein the first determination submodule is used to determine the remaining write-back speed of each client based on the data write-back speed of each client and the file synchronization bandwidth; the second determination submodule is used to determine the remaining bandwidth of each client based on the maximum bandwidth of each client and the file synchronization bandwidth; the third determination submodule is used to determine the remaining bandwidth of the server based on the maximum bandwidth of the server and the file synchronization bandwidth; the fourth determination submodule is used to determine the concurrency of the file task based on the remaining write-back speeds of all the clients, the remaining bandwidths of all the clients and the remaining bandwidth of the server.
[0189] In some embodiments, the fourth determination submodule includes a first determination unit and a second determination unit, wherein the first determination unit is used to determine the minimum value of the remaining write-back speed of all the clients, the remaining bandwidth of all the clients, and the remaining bandwidth of the server as the maximum concurrent bandwidth; and the second determination unit is used to determine the concurrency number of the file task based on the maximum concurrent bandwidth and the file synchronization bandwidth.
[0190] In some embodiments, the scheduling module 503 includes a first scheduling sub-module and a second scheduling sub-module, wherein the first scheduling sub-module is used to determine that when the remaining write-back speed of the client and / or the remaining bandwidth of the client are constraints for limiting the maximum concurrent bandwidth, the target files less than the concurrent number will be scheduled to the target client for file synchronization; the second scheduling sub-module is used to determine that when the remaining bandwidth of the server is a constraint for limiting the maximum concurrent bandwidth, the target files meeting the concurrent number will be scheduled to the target client for file synchronization; wherein the target client is determined based on the remaining write-back speed and / or the remaining bandwidth.
[0191] In some embodiments, the file synchronization device also includes an acquisition module and a comparison module, wherein the acquisition module is used to acquire a first file list sent by the target client, wherein the first file list records temporary files generated in synchronizing the target file; the comparison module is used to compare the first file list with a second file list corresponding to the target file to generate a third file list to be synchronized; correspondingly, the scheduling module 503 is also used to schedule the target file task to the target client to perform file synchronization based on the concurrency number of the file task and the third file list.
[0192] Figure 5B A schematic diagram of the structure of the file synchronization device provided in the embodiment of the present application is shown in FIG. Figure 5B As shown, the device 510 includes:
[0193] A recording module 511, used to record temporary files generated during the process of synchronizing the target file of the server to at least one file storage node;
[0194] A generating module 512, configured to generate a first file list based on the temporary file;
[0195] The sending module 513 is used to send the first file list to the server.
[0196] The description of the above device embodiment is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. For technical details not disclosed in the device embodiment of the present application, please refer to the description of the method embodiment of the present application for understanding.
[0197] It should be noted that in the embodiment of the present application, if the above method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable an electronic device (which can be a mobile phone, a tablet computer, a laptop computer, a desktop computer, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific combination of hardware and software.
[0198] Correspondingly, an embodiment of the present application provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the file synchronization method provided in the above embodiment are implemented.
[0199] Correspondingly, an embodiment of the present application provides an electronic device, Figure 6 A hardware entity diagram of an electronic device provided in an embodiment of the present application, such as Figure 6 As shown, the hardware entity of the device 600 includes: a memory 601 and a processor 602, wherein the memory 601 stores a computer program that can be run on the processor 602, and the processor 602 implements the steps in the file synchronization method provided in the above embodiment when executing the program.
[0200] The memory 601 is configured to store instructions and applications executable by the processor 602, and can also cache data to be processed or processed by the processor 602 and various modules in the electronic device 600 (for example, image data, audio data, voice communication data, and video communication data), which can be implemented through flash memory (FLASH) or random access memory (Random Access Memory, RAM).
[0201] It should be noted here that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.
[0202] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned sequence numbers of the embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.
[0203] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0204] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0205] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0206] In addition, all functional units in the embodiments of the present application may be integrated into one processing unit, or each unit may be a separate unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0207] A person skilled in the art can understand that all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, etc., various media that can store program codes.
[0208] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application can be essentially or partly reflected in the form of a software product that contributes to the relevant technology. The computer software product is stored in a storage medium, including several instructions to enable an electronic device (which can be a mobile phone, a tablet computer, a laptop computer, a desktop computer, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0209] The methods disclosed in several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0210] The features disclosed in several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0211] The features disclosed in several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0212] The above is only an implementation method of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A file synchronization method, applied to a server, characterized in that: The method comprises: Perform transactional splitting on a file system including multiple files to obtain at least two file tasks; The concurrency number of the file task is determined based on at least one of the following: the data write-back speed of each client, the maximum bandwidth of each client, the maximum bandwidth of the server, and the preset file synchronization bandwidth, wherein at least two of the clients are used to synchronize the files of the server to at least one file storage node; Based on the concurrency number of the file tasks, the target file task is scheduled to the target client to perform file synchronization, so as to synchronize the target file task to at least one of the file storage nodes.
2. The method according to claim 1, characterized in that The determining of the number of concurrency of the file tasks based on at least one of the following: the data write-back speed of each client, the maximum bandwidth of each client, the maximum bandwidth of the server, and the preset file synchronization bandwidth includes: Determine the remaining write-back speed of each client based on the data write-back speed of each client and the file synchronization bandwidth; Determine the remaining bandwidth of each of the clients based on the maximum bandwidth of each of the clients and the file synchronization bandwidth; Determining a remaining bandwidth of the server based on the maximum bandwidth of the server and the file synchronization bandwidth; The concurrency number of the file task is determined based on the remaining write-back speed of all the clients, the remaining bandwidth of all the clients and the remaining bandwidth of the server.
3. The method according to claim 2, characterized in that The determining the number of concurrency of the file task based on the remaining write-back speed of all the clients, the remaining bandwidth of all the clients, and the remaining bandwidth of the server includes: The minimum value among the remaining write-back speeds of all the clients, the remaining bandwidths of all the clients and the remaining bandwidth of the server is determined as the maximum concurrent bandwidth; The concurrency number of the file task is determined based on the maximum concurrent bandwidth and the file synchronization bandwidth.
4. The method according to claim 3, characterized in that The step of scheduling the target file task to the target client to perform file synchronization based on the concurrent number of the file task includes: When determining that the remaining write-back speed of the client and / or the remaining bandwidth of the client are constraints for limiting the maximum concurrent bandwidth, scheduling target files smaller than the concurrent number to the target client for file synchronization; When it is determined that the remaining bandwidth of the server is a constraint condition for limiting the maximum concurrent bandwidth, the target file that meets the concurrent number is scheduled to the target client for file synchronization; The target client is determined based on the remaining write-back speed and / or the remaining bandwidth.
5. The method according to claim 1, characterized in that The method further comprises: Acquire a first file list sent by the target client, wherein the first file list records temporary files generated by synchronizing the target file; Comparing the first file list with the second file list corresponding to the target file, generating a third file list to be synchronized; Correspondingly, scheduling the target file task to the target client to perform file synchronization based on the concurrent number of the file task includes: Based on the concurrency number of the file task and the third file list, the target file task is scheduled to the target client to perform file synchronization.
6. A file synchronization method, applied to a client, characterized in that: The method comprises: Recording temporary files generated during synchronization of a target file on the server to at least one file storage node; generating a first file list based on the temporary file; Send the first file list to the server.
7. A file synchronization method, applied to a file synchronization system, wherein the file synchronization system comprises a server, a client and a file storage node, characterized in that: The method comprises: The server performs transactional splitting on a file system including multiple files to obtain at least two file tasks; The server determines the number of concurrency of the file tasks based on at least one of the following: the data write-back speed of each of the clients, the maximum bandwidth of each of the clients, the maximum bandwidth of the server, and a preset file synchronization bandwidth, wherein at least two of the clients are used to synchronize the files of the server to at least one file storage node; The server dispatches the target file task to the target client based on the concurrent number of the file task; The target client performs file synchronization to synchronize the target file task to at least one of the file storage nodes.
8. The method according to claim 7, characterized in that The method further comprises: The target client records a temporary file generated in the process of synchronizing the target file of the server to at least one of the file storage nodes, generates a first file list based on the temporary file, and sends the first file list to the server; The server compares the first file list with a second file list corresponding to the target file to generate a third file list to be synchronized; Correspondingly, the server dispatches the target file task to the target client based on the concurrent number of the file task, including: The server schedules the target file task to the target client based on the concurrent number of the file task and the third file list.
9. A storage medium, characterized in that Executable instructions are stored, which are used to cause a processor to execute and implement the steps of the method described in any one of claims 1 to 5, or claim 6, or claim 7 or 8.
10. A computer program product, characterized in that The method comprises a computer program or an instruction, which, when executed by a processor, implements the steps of the method described in any one of claims 1 to 5, or claim 6, or claim 7 or 8.