Rolling upgrading method and device
By adopting the method of coexisting new processes and old processes in the distributed storage system, and using shared memory to synchronize task status information, the problem of unavailability of server resources during the upgrade is solved, resource utilization is improved, and business stability and continuity is ensured.
Patent Information
- Application Number
- CN202410094986.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-23
- Publication Date
- 2025-07-25
AI Technical Summary
The existing distributed storage system upgrade solution causes server resources to be unavailable, reduce resource utilization, and put pressure on the foreground IO performance and metadata persistence system.
By allowing new processes and old processes to coexist during the upgrade, synchronize task status information with shared memory, the new processes replace old processes to handle tasks, ensuring that the server continuously provides services.
It improves resource utilization, reduces the risk of business interruption, ensures business stability and continuity, and reduces the impact on the front-end IO performance and the pressure on the metadata persistence system.
Smart Images

Figure CN120371366A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of communications, and in particular, to a rolling upgrade method and device. Background Art
[0002] In a distributed storage system, the existing upgrade solution is to perform an upgrade in a load sharing manner. For example, at time T1, both worker nodes worker1 and worker2 are in the old version state that needs to be upgraded. The client calls the routing view v1 and sends the IO data to worker1. At time T2, the client updates the routing view to the routing view v2, and changes the target node for sending the IO data from worker1 to worker2. During the normal sending of IO data by the client, worker1 upgrades according to the normal process, first shuts down and then restarts the new version. At time T3, worker1 completes the upgrade, the client updates the routing view to the routing view v3, and changes the destination node for sending the IO data back to worker1 again; then worker2 performs the shutdown upgrade and restart operation.
[0003] During the process of stopping and restarting the process, the server cannot provide services to the outside, and the host resources cannot be used by the service either, reducing the resource utilization rate. Therefore, there is a need for an upgrade method that can ensure the continuous service provided by a single server to improve the resource utilization rate. Summary of the Invention
[0004] Embodiments of the present application provide a rolling upgrade method and device. The second process executes a target task according to the first task status information of the first process. During the upgrade, the new process and the old process coexist, solving the problem that the server resources are unavailable during the upgrade and improving the resource utilization rate.
[0005] In a first aspect, embodiments of the present application provide a signal synchronization method, which includes: calling a first process to execute a target task; writing the first task status information into a shared memory, where the first task status information is the relevant status information of the first process executing the target task; calling a second process to read the first task status information from the shared memory; calling the second process to execute the target task according to the first task status information; and stopping the first process when the first process stops processing the first task data of the target task, where the first task data is the task data received by the first process.
[0006] In this possible implementation, the second process executes the target task according to the first task status information of the first process. During the upgrade, the new process and the old process coexist, solving the problem that the server resources are unavailable during the upgrade and improving the resource utilization rate.
[0007] In one possible implementation, before invoking the first process to execute the target task, the method further includes: triggering an update of the first process to the second process.
[0008] In one possible implementation, before invoking the second process to read the first task status information from the shared memory and after writing the first task status information to the shared memory, the method further includes: creating a second process, where the task interface of the second process is the same as that of the first process.
[0009] In one possible implementation, after invoking the second process to execute the target task based on the first task status information, the method further includes: stopping, by the operating system kernel, distributing the task data of the target task to the first process.
[0010] In one possible implementation, after stopping, by the operating system kernel, distributing the task data of the target task to the first process, the method further includes: invoking the first process to process the first task data of the target task.
[0011] In one possible implementation, the method further includes: in the case of stopping, by the operating system kernel, distributing the task data of the target task to the first process, distributing, by the operating system kernel, the task data of the target task to the second process.
[0012] In this possible implementation, the operating system kernel of the network device can distribute, through the load balancing mechanism, the data received by the network device to multiple processes (the first process and the second process), thereby solving the problem that the foreground IO performance is affected during the upgrade of the storage node, reducing the risk of service interruption, and ensuring the stability and continuity of the service.
[0013] In one possible implementation, invoking the second process to execute the target task according to the first task status information includes: invoking the second process to restore the task process status according to the first task status information; invoking the second process to execute the target task according to the task process status.
[0014] In one possible implementation, the target task is a storage task or a management task.
[0015] In one possible implementation, when the target task is a management task, writing the first task status information to the shared memory includes: writing the first task status information and the metadata information to the shared memory.
[0016] In this possible implementation, for the management node, the management process can write the first task status information to the shared cache and can also write the metadata information to the shared memory, so that the second management process can obtain the metadata information from the shared cache without accessing the third-party metadata persistence system, reducing the pressure on the third-party persistence system during the upgrade.
[0017] In a second aspect, an embodiment of the present application provides a network device, including: a processor and a memory. The processor is coupled to the memory; the memory is used to store computer instructions, and the computer instructions are loaded and executed by the processor to enable the network device to implement any one of the methods provided in the first aspect.
[0018] In a third aspect, an embodiment of the present application provides a chip, which includes: a processor and an interface circuit; the interface circuit is used to receive code instructions and transmit them to the processor; the processor is used to run the code instructions to execute any one of the methods provided in the first aspect.
[0019] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which at least one computer program instruction is stored, and the computer program instruction is loaded and executed by the processor to implement any one of the methods provided in the first aspect as described above.
[0020] In a fifth aspect, an embodiment of the present application provides a computer program product, including computer execution instructions, and when the computer execution instructions run on a computer, the computer is enabled to execute any one of the methods provided in the first aspect.
[0021] For the technical effects brought by any one of the implementation manners in the second aspect to the fifth aspect, reference may be made to the technical effects brought by the corresponding implementation manners in the first aspect, which will not be elaborated here. Description of the Drawings
[0022] Figure 1 It is a schematic diagram of a scenario of a rolling upgrade method;
[0023] Figure 2 It is a schematic diagram of the architecture of a rolling upgrade method;
[0024] Figure 3 It is a schematic diagram of the process of a rolling upgrade method;
[0025] Figure 4 It is a schematic diagram of the architecture of a rolling upgrade method provided by an embodiment of the present application;
[0026] Figure 5 It is a schematic diagram of the process of a rolling upgrade method provided by an embodiment of the present application;
[0027] Figure 6 It is a schematic diagram of a scenario of a rolling upgrade method provided by an embodiment of the present application;
[0028] Figure 7 It is a schematic diagram of a scenario of another rolling upgrade method provided by an embodiment of the present application;
[0029] Figure 8This application example provides a schematic structural diagram of a network device;
[0030] Figure 9 This application example provides another schematic structural diagram of a network device. Detailed implementation manners
[0031] In the description of this application, unless otherwise specified, " / " indicates that the objects associated before and after are in an "or" relationship. For example, A / B may represent A or B; "and / or" in this application is merely a description of the association relationship of the associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B may be singular or plural.
[0032] In the description of this application, unless otherwise specified, "a plurality of" means two or more than two. "At least one (item)" or similar expressions thereof refer to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c may represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c may be single or multiple.
[0033] In addition, in order to facilitate a clear description of the technical solutions of the embodiments of this application, in the embodiments of this application, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and roles. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and "first" and "second" do not necessarily mean different.
[0034] In the embodiments of this application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific way for easy understanding.
[0035] It can be understood that the "embodiments" mentioned throughout the specification mean that specific features, structures, or characteristics related to the embodiments are included in at least one embodiment of this application. Therefore, the various embodiments mentioned throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner. It can be understood that in the various embodiments of this application, the magnitude of the serial numbers of the various processes does not mean the sequence of execution is prior or subsequent, and the execution sequence of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.
[0036] It can be understood that some optional features in the embodiments of the present application can, in some scenarios, be implemented independently without relying on other features, such as the current solution they are based on, to solve corresponding technical problems and achieve corresponding effects. In some scenarios, they can also be combined with other features according to requirements. Correspondingly, the devices given in the embodiments of the present application can also implement these features or functions accordingly, which will not be elaborated here.
[0037] In the present application, unless otherwise specified, the same or similar parts among various embodiments can be referred to each other. In the present application, if there is no special specification and logical conflict among various embodiments, the terms and / or descriptions among different embodiments are consistent and can be cited mutually. Different embodiments can be combined to form new embodiments according to their internal logical relationships. The following embodiments of the present application do not constitute a limitation on the protection scope of the present application.
[0038] To facilitate the understanding of the technical solutions of the embodiments of the present application, a brief introduction to the related technologies of the present application is given as follows:
[0039] Rolling update refers to upgrading components one by one or in batches in a certain order in a distributed IT system. During the rolling update process, the new version of the software and the old version of the software coexist in the system components. The number of old version components gradually decreases to zero, and the number of new version components gradually increases until the entire system is fully covered.
[0040] Load balance means that when multiple data packets need to be forwarded to multiple processes, it can ensure that the forwarding is definitely successful, and the data packet load received by the normal state processes is overall balanced.
[0041] As Figure 1 shown, the main architecture of the existing distributed storage system includes the client side, the server side, and the master side. The client side can be multiple client clients; the server side can be in the form of a server cluster, such as multiple worker nodes; the master side can be in the form of a master cluster, such as the main master and the standby master. The master side can communicate with the view data persistence system.
[0042] The client side close to the user side can receive the input and output (IO) data issued by the user, call the globally consistent routing view, and send the IO data to the specific worker node on the server side.
[0043] The server cluster consists of a large number of worker nodes. Each worker node can receive data from the client and persist the received data to the local persistent medium, such as storing it on a mechanical disk or a solid-state drive.
[0044] The master cluster consists of multiple nodes with management functions and achieves high availability (HA) in a primary and standby deployment manner. The primary master maintains heartbeat connections and data transfer connections with each worker node in the server cluster, and at the same time synchronously writes key cluster view data and IO view data to a third-party persistent system, such as a view data persistent system. The standby master monitors the health information of the primary master at all times. When the primary master fails, the standby master can quickly take over the management of the server cluster.
[0045] For example Figure 2 As shown, the existing server storage cluster is upgraded in a load-sharing manner. Specifically as follows:
[0046] At time T1, both worker nodes, worker1 and worker2, are in the old version state that needs to be upgraded. The client calls the routing view v1 and sends the IO data to worker1.
[0047] At time T2, the client updates the routing view to routing view v2 and changes the target node for sending the IO data from worker1 to worker2. During the normal sending of IO data by the client, worker1 is upgraded according to the normal process, first shutting down and then restarting the new version.
[0048] At time T3, worker1 has completed the upgrade. The client updates the routing view to routing view v3 and changes the destination node for sending the IO data back to worker1 again; then worker2 performs the shutdown upgrade and restart operation.
[0049] For example Figure 3 As shown, for the Master management cluster, it is usually upgraded in a primary and standby switchover manner. Specifically as follows:
[0050] At time T1, the master cluster includes a primary master and a standby master. The operation and maintenance personnel directly shut down master2 for upgrade. Since the server storage cluster has elected master1 as the leader at this time and is connected to master1, shutting down master2 for upgrade does not affect the normal functions of the management cluster. In addition, the primary master will write the key metadata information of the cluster into the persistent system in real time, such as a database or ZooKeeper, etc.
[0051] At time T2, after the new version of master2 is upgraded and ready, it reads the key metadata information written by master1 from the metadata persistent system, and then the operation and maintenance personnel shut down master1 for upgrade. After the server cluster determines that master1 fails, it elects master2 as the leader and establishes a connection with master2.
[0052] At time T3, master1 is upgraded and ready and runs as a standby master. At this time, in the management cluster, both the primary master and the standby master have completed the upgrade.
[0053] In a distributed storage system, there are multiple problems with existing upgrade solutions. For example, during the shutdown and restart period, the resources of the server are unavailable. During the process of stopping and restarting the process, the server cannot provide services externally, and the host resources cannot be used by the business, reducing the resource utilization rate. In addition, it also reduces the foreground IO performance. During the load sharing upgrade process, there is a certain delay in the client to update the routing view. When the client sends a request to the upgraded node using the unupdated view, it will fail due to the routing view error. After the failure, the client continuously tries to send requests until the client updates to the new routing view before it can perform IO normally. Therefore, the load sharing upgrade will affect the foreground IO performance.
[0054] In addition, during the upgrade process, the frequent update of metadata causes great pressure on the metadata persistent system. The master cluster will write the key cluster metadata into a third-party storage system, such as a database or a key-value storage system like ZooKeeper. During the upgrade process, the frequent update of the routing view requires a large amount of reading and writing of such third-party systems, causing certain pressure on these systems.
[0055] During the system upgrade process of a highly available distributed storage system, multiple aspects need to be considered. For example, the impact of system upgrade on the functions of foreground I / O; reducing the performance impact of system upgrade on foreground I / O; and whether the system upgrade is easy to quickly roll back in case of failure. Existing distributed storage systems (such as the Pangu file system of Alibaba Cloud and the GFS of Google) have adopted a stable storage system upgrade solution based on the above three key points. For the server storage cluster, a load sharing method is used, and for the master management cluster, a primary / backup switchover method is used.
[0056] In the existing storage system upgrade solution, the load sharing method will modify the I / O routing view information of the storage cluster, and the client side will fail and retry during the synchronization of view data, reducing the performance of the storage system. On the other hand, the primary / backup switchover method needs to communicate with the view data persistence system frequently, increasing the read / write pressure on the persistence system. Secondly, the existing upgrade method needs to stop the old version process first, and then start the new version process. During these two operations, a single server does not provide any services externally, wasting server resources.
[0057] Therefore, there is a need for an upgrade method that has little impact on the performance of foreground I / O during the upgrade process, has little pressure on the view data persistence system, and ensures the continuous service provided by a single server externally to improve resource utilization.
[0058] Based on this, the embodiment of the present application provides a rolling upgrade method, which includes: calling a first process to execute a target task; writing the first task status information into the shared memory, where the first task status information is the relevant status information of the first process executing the target task; calling a second process to read the first task status information from the shared memory; according to the first task status information, calling the second process to execute the target task; when the first process stops processing the first task data of the target task, stopping the first process, where the first task data is the task data received by calling the first process. By having the second process execute the target task according to the first task status information of the first process, the new process and the old process coexist during the upgrade, solving the problem of unavailable server resources during the upgrade and improving resource utilization.
[0059] The rolling upgrade method provided by the embodiment of the present application can be applicable to Figure 4 the distributed storage system shown. The main architecture of this distributed storage system includes a client side, a server side, and a master side. The client side can be multiple client clients; the server side can be in the form of a server cluster, such as multiple worker nodes; the master side can be in the form of a master cluster, such as a primary master and a standby master. The master side can communicate with the view data persistence system.
[0060] The client side close to the user can accept the input and output (IO) data sent by the user, call the globally consistent routing view, and send the IO data to the worker nodes on the specific server side.
[0061] The server cluster consists of a large number of worker nodes. Each worker node can accept data from the client and persist the received data to the local persistent medium, such as storing it on a mechanical disk or a solid-state drive.
[0062] The master cluster consists of multiple nodes with management functions and achieves high availability (HA) in a primary and standby deployment manner. The primary master maintains heartbeat connections and data transfer connections with each worker node in the server cluster, and simultaneously synchronously writes the key cluster view data and IO view data into a third-party persistent system, such as a view data persistent system. The standby master constantly monitors the health information of the primary master. When the primary master fails, the standby master can quickly take over the management of the server cluster.
[0063] The rolling upgrade method provided by the embodiments of this application can be applicable to Figure 4 rolling upgrade of the worker nodes as storage nodes in the distributed storage system shown, and can also be applicable to Figure 4 rolling upgrade of the master nodes as management nodes in the distributed storage system shown.
[0064] In the embodiments of this application, the rolling upgrade method can be applicable to the distributed storage system. In addition, it can also be applicable to systems providing other services, such as distributed computing systems providing computing services, distributed network systems providing networks, and database systems, etc. The embodiments of this application are mainly applicable to scenarios where the processes of nodes change but continuous services need to be provided. Therefore, it can be applicable to other situations. For example, after the first process fails, a second process with the same version as the first process is used to replace the first process, and the continuous service is provided externally through the rolling upgrade method provided by the embodiments of this application during the replacement process.
[0065] It can be understood that in the embodiments of this application, the execution subject can execute some or all of the steps in the embodiments of this application. These steps or operations are only examples. The embodiments of this application can also execute other operations or various deformations of the operations. In addition, each step can be executed in a different order presented in the embodiments of this application, and it is possible not to execute all the operations in the embodiments of this application.
[0066] It should be noted that the message names between devices or the names of parameters in the messages in the following embodiments of this application are only examples. In specific implementations, other names can also be used, and this application does not make specific limitations in this regard.
[0067] In the embodiments of this application, the rolling upgrade method can be applied to multiple scenarios. The following takes the upgrade of storage nodes and management nodes in a distributed storage system as an example for illustration:
[0068] As Figure 5 shown, it is a rolling upgrade method provided by the embodiments of this application, and this method includes the following steps:
[0069] In the offline stage, the rolling upgrade method provided by the embodiments of this application can perform the following steps:
[0070] 501. Trigger the update of the first process.
[0071] The network device is triggered to update the first process, and plans to update the first process to the second process. The first process is the process of the old version before the upgrade, and the second process is the process of the new version after the upgrade.
[0072] In the embodiments of this application, the update and upgrade of the first process can be triggered by operation and maintenance management personnel or the network device system, and specific limitations are not made here.
[0073] In the embodiments of this application, in a distributed storage system, the network device can be a storage node for performing storage tasks, and both the first process and the second process are storage processes. The network device can also be a management node for performing management tasks, and both the first process and the second process are management processes. In addition, it can also be other types of devices for implementing other corresponding functions, and specific limitations are not made here.
[0074] 502. Write the first task status information into the shared memory.
[0075] The network device calls the first process to execute the target task. After triggering the update of the first process, the network device can write the first task status information into the shared memory. The first task status information is the relevant status information of the first process for executing the target task.
[0076] In an embodiment of the present application, after triggering the update of the first process, the first process will continue to execute the target task, but will write the key runtime status information of the task execution to the shared memory. The first task status information indicates the execution status of the first process for the target task. For example, when the process is a storage process, the first task status information can indicate the execution status of the target storage task. For example, the target storage task is to store 10 data packets into the storage medium. For the 6 data packets received by the network device, the first process has stored 5 data packets into the storage medium, and there is still one data packet being processed. At this time, the network device can call the first process to store this status information into the shared memory.
[0077] In an embodiment of the present application, the shared memory is a shared memory that can be accessed and read by both the first process and the second process. The shared memory can be the internal memory of the network device, and specific details are not limited here.
[0078] In a possible implementation, the network device is a storage node for executing storage tasks, and both the first process and the second process are storage processes. The first process can write the first task status information into the shared cache.
[0079] In a possible implementation, the network device is a management node for executing management tasks, and both the first process and the second process are management processes. The first management process can write the first task status information into the shared cache. In addition, the first management process can also write metadata information into the shared memory, and the metadata information is the metadata information of the entire storage cluster.
[0080] In an embodiment of the present application, for the management node, the management process can write the first task status information into the shared cache, and can also write the metadata information into the shared memory, so that the second management process can obtain the metadata information from the shared cache without accessing the third-party metadata persistence system, reducing the pressure on the third-party persistence system during the upgrade.
[0081] In an embodiment of the present application, the third-party persistence system can be a persistence system such as a database or ZooKeeper, and the storage medium can be a storage medium such as a disk or a solid-state drive, and specific details are not limited here.
[0082] 503. Create a second process.
[0083] The network device creates (starts) the second process, and the task interface of the second process is the same as that of the first process.
[0084] Specifically, after triggering the update of the first process, the network device can create a new version of the second process. The task interface of the second process is the same as that of the first process, that is, the second process can receive the same task data as the first process.
[0085] For example, when the network device is a storage node for executing storage tasks and both the first process and the second process are storage processes, the network device can distribute the received task data to each process through the operating system (OS) kernel. When creating the second process, the OS port listened by the second process is the same as the OS port listened by the first process. And when establishing a socket connection, the so-reuseport feature is used to enable multiple processes to bind to the same port.
[0086] 504. Invoke the second process to read the first task status information from the shared memory.
[0087] The network device invokes the second process to read the first task status information from the shared memory, restore the process state, and prepare to execute the target task according to the first task status information. After the second process restores the process state according to the first task status information, the network device can consider that the second process is ready to execute the target task.
[0088] 505. Stop invoking the first process to receive the task data of the target task.
[0089] After the second process of the network device restores the process state according to the first task status information, that is, after the second process is ready to execute the target task, the network device stops invoking the first process to receive the task data of the target task, that is, the OS kernel of the network device no longer distributes the task data to the first process.
[0090] However, for the task data already received by the first process, the first process will continue to process this task data until all of it is processed. For example, for storage tasks, for the task data already received by the first storage process, the first storage process will continue to write this task data to the storage medium until all of it is written.
[0091] 506. Invoke the second process to execute the target task.
[0092] The network device invokes the second process to execute the target task.
[0093] Specifically, after the second process of the network device is ready to execute the target task, the network device stops invoking the first process to receive the task data of the target task and starts invoking the second process to execute the target task. The second process starts to process the task data corresponding to the target task according to the process state restored from the first task status information.
[0094] In a possible implementation, the network device is a storage node for performing storage tasks, and both the first process and the second process are storage processes. After the second storage process is ready, the operating system kernel of the network device distributes the received storage task data to the second storage process, and the second storage process writes the storage task data into the storage medium. The storage task data received by the storage node is IO data issued by the user, with a large packet volume and a high sending and receiving frequency.
[0095] In a possible implementation, the network device is a management node for performing management tasks, and both the first process and the second process are management processes. After the second storage process is ready, the operating system kernel of the network device distributes the received management task data to the second management process, and the second management process writes the management task data into a third-party persistent system, such as a database or ZooKeeper. The management task data received by the management node belongs to the command control type, with a small packet volume and a low sending and receiving frequency.
[0096] In the embodiments of the present application, the second process executes the target task according to the first task status information of the first process. During the upgrade, the new process and the old process coexist, solving the problem that the server resources are unavailable during the upgrade and improving the resource utilization rate.
[0097] In the embodiments of the present application, the operating system kernel of the network device can distribute the data received by the network device through the load balancing mechanism to multiple processes (the first process and the second process), thereby solving the problem that the foreground IO performance is affected during the upgrade of the storage node, reducing the risk of service interruption, and ensuring the stability and continuity of the service.
[0098] 507. Stop the first process.
[0099] In the case where the first process stops processing the first task data of the target task, the first process is stopped, and the first task data is the task data received by the first process.
[0100] In the embodiments of the present application, as described in step 505, after stopping distributing the task data to the first process, the first process still needs to continue processing the first task data that has been received. After the first process finishes processing the first task data, the network device can stop (kill) the first process.
[0101] In a possible implementation, the network device is a storage node for performing storage tasks, and both the first process and the second process are storage processes. The network device invokes the first storage process to continue processing the received first task data, which was received before the first storage process stopped receiving task data. After the first storage process finishes processing the first task data, the network device can stop the first storage process. For example, after the first storage process stores all the received first task data into the storage medium, if the network device detects that the first storage process no longer writes data into the storage medium, it can determine that the first storage process has completed all IO requests, and then the network device can stop the first storage process.
[0102] In a possible implementation, the network device is a management node for performing management tasks, and both the first process and the second process are management processes. The network device invokes the first management process to continue processing the received first task data, which was received before the first management process stopped receiving task data. After the first management process finishes processing the first task data, the network device can stop the first management process. For example, after the first management process writes all the received first task data into a third-party persistent system, if the network device detects that the first management process no longer writes data into the third-party persistent system, it can determine that the first management process has processed all the first task data, and then the network device can stop the first management process.
[0103] An embodiment of the present application provides a network device 800. In the embodiment of the present application, the functional modules of the network device 800 can be divided according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiment of the present invention is illustrative, only a logical function division, and there may be other division methods in actual implementation.
[0104] In the case of dividing each functional module corresponding to each function, Figure 8 A possible structural schematic diagram of the network device 800 involved in the above embodiment is shown. As Figure 8 shown, the network device 800 includes:
[0105] An update module 801, configured to trigger an update of the first process to the second process. For example, in step 501, trigger an update of the first process.
[0106] A first execution module 802, configured to invoke the first process to execute a target task; for example, in step 502, write the first task status information into the shared memory.
[0107] A writing module 803, configured to write first task status information into a shared memory, where the first task status information is status information related to a first process executing a target task; for example, in step 502, the first task status information is written into the shared memory.
[0108] In a possible implementation, the target task is a management task, and the writing module 803 is specifically configured to: write the first task status information and metadata information into the shared memory. For example, in step 502, the first task status information is written into the shared memory.
[0109] A creating module 804, configured to create a second process, where the task interface of the second process is the same as that of the first process. For example, in step 503, the second process is created.
[0110] A reading module 805, configured to call the second process to read the first task status information from the shared memory; for example, in step 504, the second process is called to read the first task status information from the shared memory.
[0111] A distributing module 806, configured to, when the task data of the target task is no longer distributed to the first process through the operating system kernel, distribute the task data of the target task to the second process through the operating system kernel. For example, in step 506, the second process is called to execute the target task.
[0112] A second stopping module 807, configured to stop distributing the task data of the target task to the first process through the operating system kernel. For example, in step 505, calling the first process to receive the task data of the target task is stopped.
[0113] A processing module 808, configured to call the first process to process the first task data of the target task. For example, in step 506, the second process is called to execute the target task.
[0114] A second execution module 809, configured to call the second process to execute the target task according to the first task status information; for example, in step 506, the second process is called to execute the target task.
[0115] In a possible implementation, the second execution module 809 includes:
[0116] A recovery unit 810, configured to call the second process to recover the task process status according to the first task status information; for example, in step 506, the second process is called to execute the target task.
[0117] An execution unit 811, configured to call the second process to execute the target task according to the task process status. For example, in step 506, the second process is called to execute the target task.
[0118] The first stop module 812 is configured to stop the first process when the first process stops processing the first task data of the target task, where the first task data is the task data received by the first process. For example, in step 507, the first process is stopped.
[0119] Each module of the above non-linear compensation device can also be used to perform other actions in the above method embodiments. All relevant contents of the steps involved in the above method embodiments can be cited in the function descriptions of the corresponding functional modules, and will not be elaborated here.
[0120] Figure 9 FIG. is a schematic structural diagram of a network device provided by an embodiment of the present application. The network device 900 may include one or more central processing units (CPUs) 901 and a memory 905. One or more application programs or data are stored in the memory 905.
[0121] Among them, the memory 905 may be volatile storage or persistent storage. The program stored in the memory 905 may include one or more modules, and each module may include a series of instruction operations on the network device. Further, the central processing unit 901 may be configured to communicate with the memory 905 and execute a series of instruction operations in the memory 905 on the network device 900.
[0122] Among them, the central processing unit 901 is configured to execute the computer program in the memory 905, so that the network device 900 is used to perform: calling the first process to execute the target task; writing the first task status information into the shared memory, where the first task status information is the relevant status information of the first process executing the target task; calling the second process to read the first task status information from the shared memory; according to the first task status information, calling the second process to execute the target task; when the first process stops processing the first task data of the target task, stopping the first process, where the first task data is the task data received by the first process. For the specific implementation method, please refer to Figure 5 Steps 501-507 in the illustrated embodiment are not elaborated here.
[0123] The network device 900 may further include one or more power supplies 902, one or more wired or wireless network interfaces 903, one or more input / output interfaces 904, and / or one or more operating systems, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0124] The network device 900 can execute the foregoing Figure 5The operations performed by the network device in the illustrated embodiments are not specifically elaborated herein.
[0125] The embodiments of the present application also provide a computer program product containing instructions. The computer program product can be software or a program product containing instructions that can run on the plug-in result reuse device or be stored in any available medium. When the computer program product runs on the plug-in result reuse device, it causes the network device to execute the Figure 5 rolling upgrade method performed in the illustrated embodiments.
[0126] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a cache server or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state drive). The computer-readable storage medium includes instructions that instruct the network device to execute the Figure 5 rolling upgrade method performed in the illustrated embodiments.
[0127] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes (or functions) of the embodiments of the present application are implemented. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, a computer, a server, or a data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center containing one or more available media integrated. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state disk (SSD)). In the embodiments of the present application, the computer can include the foregoing device.
[0128] Although the present application has been described in conjunction with the various embodiments, it will be understood by those skilled in the art that various changes in the disclosed embodiments can be understood and effected while practicing the claimed application. In the claims, the term "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. A single processor or other unit may implement several functions recited in the claims. Certain measures are recited in mutually different dependent claims, but this does not indicate that these measures cannot be combined to advantage.
Claims
1. A rolling upgrade method, characterized in that, The method includes: Invoking a first process to execute a target task; Writing first task status information into a shared memory, where the first task status information is the relevant status information of the first process executing the target task; Invoking a second process to read the first task status information from the shared memory; Invoking the second process to execute the target task according to the first task status information; Stopping the first process when the first process stops processing the first task data of the target task, where the first task data is the task data received by the first process.
2. The method according to claim 1, characterized in that, Before invoking the first process to execute the target task, the method further includes: Triggering to update the first process to the second process.
3. The method according to claim 2, wherein Before invoking the second process to read the first task status information from the shared memory and after writing the first task status information into the shared memory, the method further includes: Creating the second process, where the task interface of the second process is the same as that of the first process.
4. The method according to claim 3, wherein After invoking the second process to execute the target task according to the first task status information, the method further includes: Stopping, by an operating system kernel, distributing the task data of the target task to the first process.
5. The method according to claim 4, wherein After stopping, by the operating system kernel, distributing the task data of the target task to the first process, the method further includes: Invoking the first process to process the first task data of the target task.
6. The method according to claim 4 or 5, characterized in that, The method further includes: When stopping, by the operating system kernel, distributing the task data of the target task to the first process, distributing, by the operating system kernel, the task data of the target task to the second process.
7. The method according to claim 6, wherein The invoking the second process to execute the target task according to the first task status information includes: Invoking the second process to restore a task process status according to the first task status information; Invoking the second process to execute the target task according to the task process status.
8. The method according to any one of claims 1 to 7, characterized in that, The target task is a storage task or a management task.
9. The method according to claim 8, wherein When the target task is a management task, the writing the first task status information into the shared memory includes: Writing the first task status information and metadata information into the shared memory.
10. A network device, characterized in that, The network device includes: A first execution module, configured to invoke a first process to execute a target task; A writing module, configured to write first task status information into a shared memory, where the first task status information is the relevant status information of the first process executing the target task; A reading module, configured to invoke a second process to read the first task status information from the shared memory; A second execution module, configured to invoke the second process to execute the target task according to the first task status information; A first stop module, configured to stop the first process when the first process stops processing the first task data of the target task, where the first task data is the task data received by the first process.
11. The network device according to claim 10, characterized in that, The network device further includes: An update module, configured to trigger to update the first process to the second process.
12. The network device according to claim 11, characterized in that, The network device further includes: A creation module, configured to create the second process, and a task interface of the second process is the same as that of the first process.
13. The network device according to claim 12, characterized in that, The network device further includes: A second stop module, configured to stop distributing task data of the target task to the first process through an operating system kernel.
14. The network device according to claim 13, wherein The network device further includes: A processing module, configured to call the first process to process first task data of a target task.
15. The network device according to claim 13 or 14, characterized in that, The network device further includes: A distribution module, configured to, when stopping distributing task data of the target task to the first process through the operating system kernel, distribute the task data of the target task to the second process through the operating system kernel.
16. The network device according to claim 15, wherein The second execution module includes: A recovery unit, configured to call the second process to recover a task process state according to the first task state information; An execution unit, configured to call the second process to execute a target task according to the task process state.
17. The network device according to claim 16, wherein, The target task is a management task, and the writing module is specifically configured to: Write first task state information and metadata information into a shared memory.
18. A network device, characterized in that, The network device includes a processor and a memory; the processor is coupled to the memory; the memory is configured to store computer instructions, and the computer instructions are loaded and executed by the processor to enable the network device to implement the method according to any one of claims 1-9.
19. A computer-readable storage medium, characterized in that, At least one computer program instruction is stored in the computer-readable storage medium, and the computer program instruction is loaded and executed by a processor to implement the method according to any one of claims 1-9.
20. A computer program product, characterized in that, The computer program product includes computer execution instructions, and when the computer execution instructions run on a computer, the computer is configured to implement the method according to any one of claims 1-9.