Method and apparatus for upgrading software
Patent Information
- Application Number
- CN202211063592.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-01
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2042-09-01
AI Technical Summary
但对于银行核心业务系统而言,这种方式虽然稳妥,但每次投产切换的实施过程均需要申请停机4个小时以上的停机窗口,这种升级投产方案降低了系统的可用性时长,影响用户的使用体验
[0085]由以上方案可知,本申请提供一种软件的升级投产方法及装置,所述软件的升级投产方法包括:首先,按照虚机的功能类型将并行耦合系统内多个虚机进行拆分,得到第一批次虚机和第二批次虚机;其中,所述虚机的功能类型分为业务、网络以及数据;然后,对所述第一批次虚机进行旧版本软件向新版本软件的投产切换,得到新版本软件的第一批次虚机;其中,所述新版本软件的第一批次虚机和第二批次的虚机同时对外提供服务;对所述新版本软件的第一批次虚机进行功能正确性验证;若所述新版本软件的第一批次虚机进行功能正确性验证通过,对所述第二批次虚机进行旧版本软件向新版本软件的投产切换。有效减少投产切换带来的影响,以不停机方式完成投产切换工作实施,提高用户的使用体验。
Smart Images

Figure CN115408037B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method and apparatus for upgrading and putting software into production. Background Technology
[0002] Upgrading the host platform's basic software is a complex project, encompassing software installation, deployment in the test environment, functional and non-functional testing of the new basic software, deployment drills, and the switchover to the production environment. The project has a long duration, involves numerous changes, and a large number of stakeholders, placing extremely high demands on project management. Furthermore, to mitigate the risks that may affect the system's transaction delivery capabilities during the deployment of the new software version, detailed contingency plans must be developed in advance, and these plans must be validated and estimated beforehand to ensure their availability and accuracy.
[0003] During production environment switchovers, system maintainers typically choose to perform the switchover during off-peak business hours, such as the early morning. The process involves restarting virtual machines using the new software version and completing all necessary upgrade steps before services can be restored. The upgraded products involve large amounts of data and numerous update operations for each product, resulting in a relatively long switchover process, typically 4-5 hours per operation. The isolation and merging processes involved in virtual machine upgrades can impact transaction performance, including increased transaction response times leading to performance fluctuations and potential failures of transactions in transit. In traditional software upgrade deployments, implementers proactively request downtime windows for the switchover to avoid unnecessary disruptions to online transactions. However, while this approach is reliable for core banking systems, each switchover requires requesting downtime of more than 4 hours, reducing system availability and impacting user experience. Summary of the Invention
[0004] In view of this, this application provides a method and apparatus for upgrading and deploying software, which effectively reduces the impact of deployment switchover and improves the user experience.
[0005] The first aspect of this application provides a method for upgrading and deploying software, comprising:
[0006] The parallel coupled system is divided into a first batch of virtual machines and a second batch of virtual machines according to their functional types. The functional types of the virtual machines are divided into service, network, and data.
[0007] The first batch of virtual machines is switched from the old version of the software to the new version of the software for production deployment, resulting in the first batch of virtual machines with the new version of the software; wherein, the first batch of virtual machines with the new version of the software and the second batch of virtual machines provide services to the outside world simultaneously.
[0008] The first batch of virtual machines using the new version of the software were functionally corrected.
[0009] If the first batch of virtual machines using the new software version passes the functional correctness verification, the second batch of virtual machines will be switched from the old software version to the new software version for production deployment.
[0010] Optionally, the upgrade and production deployment method further includes:
[0011] The operation schemes for virtual machines of each functional type during the production switchover process are analyzed, and the optimization and adjustment results of the operation schemes for each functional type of virtual machine during the production switchover process are obtained; wherein, the operation schemes are divided into isolation operation schemes and merging operation schemes;
[0012] The gateway restart was analyzed to determine its impact, and the gateway restart method was improved.
[0013] Optionally, the virtual machine's functional type is a business virtual machine. The analysis of the operation plan during the production switchover process for each functional type of virtual machine yields optimized adjustment results for each type of virtual machine during the production switchover process, including:
[0014] If the operation scheme is an isolated operation scheme, then on the host side, the transaction middleware application controller is set to a static state, and the routing policy from the gateway to the inter-transaction middleware routing controller is turned off in the inter-transaction middleware routing controller;
[0015] On the gateway side, the logical unit of the host platform network protocol layer of the transaction routing controller that has a closed routing policy is shut down;
[0016] On the host side, shut down the process corresponding to the transaction middleware routing controller;
[0017] On the host side, shut down the transaction middleware application controller process and isolate the business middleware of the business virtual machine;
[0018] Shut down the database process and isolate the database software on the business virtual machine;
[0019] Isolate multiple business virtual machines one by one from the parallel coupled system;
[0020] If the operation plan is to merge the operation plan, the business virtual machine operating system will be started with the new version of the medium, and the host platform relational database system software and transaction middleware software will not be automatically started.
[0021] Control multiple business virtual machines to be successively integrated into a parallel coupled system;
[0022] Control the concurrent startup of database instances on multiple business virtual machines, and start them in sequence.
[0023] Optionally, the virtual machine's functional type is a network virtual machine. The analysis of the operation plan during the production switchover process for each functional type of virtual machine yields optimized adjustment results for each type of virtual machine during the production switchover process, including:
[0024] If the operation plan is an isolation operation plan, then the active dependent logical unit requester connection in the network virtual machine will be switched to run on another network virtual machine.
[0025] Switch the active inter-node control point session to run on another network virtual machine;
[0026] On the host side, the physical and logical units corresponding to each gateway node are sequentially activated and killed.
[0027] On the host side, the control points corresponding to each gateway node are deactivated.
[0028] Isolate multiple network virtual machines one by one from the parallel coupled system;
[0029] If the operation scheme is an incorporation operation scheme, control multiple network virtual machines to be incorporated into the parallel coupled system one by one.
[0030] Optionally, the virtual machine's functional type is a data virtual machine. The analysis of the operation plan during the production switchover process for each functional type of virtual machine yields optimized adjustment results for each type of virtual machine during the production switchover process, including:
[0031] If the operation scheme is an isolation operation scheme, then multiple data virtual machines will be isolated one by one from the parallel coupled system.
[0032] If the operation scheme is a merge operation scheme, control multiple data virtual machines to merge into the parallel coupled system one by one.
[0033] Optionally, the analysis of the gateway restart to obtain the degree of impact of the gateway restart and the improvement of the gateway restart method include:
[0034] To ensure service availability, m gateways are reserved without restarting, while the remaining nm gateways are shut down. A total of n gateway devices are connected to the gateway host via the SNA protocol, where n and m are positive integers, and m... <n / 2;
[0035] Start m gateway hosts that have been shut down, and rebuild the logical unit on the newly started m gateway hosts;
[0036] Start up n-2m gateway hosts that are in a shutdown state, wherein the n-2m gateway hosts will send uniform logical unit creation requests to all transaction routing controllers during the logical unit reconstruction process;
[0037] Reboot m gateways that have never been rebooted before;
[0038] The second restart involves m gateway hosts that have already been started once.
[0039] Optionally, the upgrade and production deployment method further includes:
[0040] If, during the first batch of virtual machine deployment and switchover, a problem is encountered that cannot be immediately identified as the root cause and cannot be resolved in a short period of time, causing the virtual machines to fail to start normally, the upgrade deployment should be stopped immediately, and the upgraded virtual machines should be rolled back.
[0041] Optionally, the upgrade and production deployment method further includes:
[0042] After all virtual machines in the first batch have been put into production and switched over, if the parallel coupled system encounters problems that cannot be immediately located and cannot be resolved in a short period of time while running externally in a mixed storage state, and these problems affect transactions, then the virtual machines that have been put into production will be immediately isolated and the first batch of upgraded virtual machines will be rolled back.
[0043] Optionally, the upgrade and production deployment method further includes:
[0044] If all virtual machines have been put into production and switched over, and problems are encountered during the trial operation phase of the parallel coupled system in full upgrade state that cannot be immediately located and cannot be resolved in a short period of time, then all virtual machines that have been put into production should be immediately isolated and all upgraded virtual machines should be rolled back.
[0045] A second aspect of this application provides a software upgrade and production apparatus, comprising:
[0046] The splitting unit is used to split multiple virtual machines in the parallel coupled system according to their functional types to obtain a first batch of virtual machines and a second batch of virtual machines; wherein, the functional types of the virtual machines are divided into service, network and data.
[0047] The first production switching unit is used to switch the first batch of virtual machines from the old version of the software to the new version of the software, so as to obtain the first batch of virtual machines with the new version of the software; wherein, the first batch of virtual machines with the new version of the software and the second batch of virtual machines provide services to the outside world at the same time.
[0048] The verification unit is used to verify the functional correctness of the first batch of virtual machines in the new version of the software.
[0049] The second production switching unit is used to switch the second batch of virtual machines from the old version of the software to the new version of the software if the first batch of virtual machines passes the functional correctness verification.
[0050] Optionally, the software upgrade and production deployment device further includes:
[0051] The first analysis unit is used to analyze the operation plan of virtual machines of each functional type during the production switchover process, and obtain the optimization and adjustment results of the operation plan of virtual machines of each functional type during the production switchover process; wherein, the operation plan is divided into isolation operation plan and merging operation plan;
[0052] The second analysis unit is used to analyze the gateway restart, obtain the impact of the gateway restart, and improve the gateway restart method.
[0053] Optionally, the virtual machine's functional type is a business virtual machine, and the first analysis unit includes:
[0054] The first shutdown unit is used to set the transaction middleware application controller to a static state on the host side and to shut down the routing policy from the gateway to the inter-transaction middleware routing controller in the inter-transaction middleware routing controller if the operation scheme is an isolation operation scheme.
[0055] The second shutdown unit is a logical unit at the host platform network protocol layer of the transaction routing controller with a closed routing policy, used on the gateway side to shut down the transaction routing controller with a closed routing policy.
[0056] The third shutdown unit is used to shut down the process corresponding to the transaction middleware routing controller on the host side;
[0057] The fourth shutdown unit is used to shut down the transaction middleware application controller process on the host side, thereby isolating the business middleware of the business virtual machine;
[0058] The fifth shutdown unit is used to shut down the database process and isolate the database software of the business virtual machine;
[0059] The first isolation unit is used to isolate multiple business virtual machines one by one from the parallel coupled system.
[0060] The startup unit is used to start the business virtual machine operating system with the new version media if the operation plan is a merge operation plan, without automatically starting the host platform relational database system software and transaction middleware software;
[0061] The first control unit is used to control multiple business virtual machines to be successively integrated into the parallel coupled system;
[0062] The second control unit is used to control the concurrent and sequential startup of database instances on multiple business virtual machines.
[0063] Optionally, the virtual machine's functional type is a network virtual machine, and the first analysis unit includes:
[0064] The first switching unit is used to switch the connection of the active dependent logical unit requester in the network virtual machine to run on another network virtual machine if the operation scheme is an isolation operation scheme.
[0065] The second switching unit is used to switch an active inter-node control point session to run on another network virtual machine.
[0066] The first kill unit is used to kill the physical and logical units corresponding to each gateway node in sequence on the host side.
[0067] The second kill unit is used to kill the control points corresponding to each gateway node on the host side;
[0068] The second isolation unit is used to isolate multiple network virtual machines from the parallel coupled system one by one in sequence;
[0069] The third control unit is used to control multiple network virtual machines to be successively merged into the parallel coupled system if the operation scheme is an incorporation operation scheme.
[0070] Optionally, the virtual machine's functional type is a data virtual machine, and the first analysis unit includes:
[0071] The third isolation unit is used to isolate multiple data virtual machines from the parallel coupled system one by one if the operation scheme is an isolation operation scheme.
[0072] The fourth control unit is used to control multiple data virtual machines to be successively incorporated into the parallel coupled system if the operation scheme is an incorporation operation scheme.
[0073] Optionally, the second analysis unit includes:
[0074] The shutdown unit is used to reserve m gateways without restarting to ensure service availability, and to shut down the remaining nm gateways. There are n gateway devices connected to the gateway host via the SNA protocol, where n and m are positive integers, and m... <n / 2;
[0075] The first startup unit is used to start m gateway hosts that have been shut down, and the newly started m gateway hosts are used to rebuild the logic unit.
[0076] The second startup unit is used to start n-2m gateway hosts that are in a shutdown state; wherein, during the process of rebuilding the logical unit, the n-2m gateway hosts will send a uniform logical unit creation request to all transaction routing controllers.
[0077] The first restart unit is used to restart m gateways that have never been restarted before;
[0078] The second restart unit is used to restart m gateway hosts that have already been started once.
[0079] Optionally, the method for upgrading and deploying the software further includes:
[0080] The first rollback unit is used to immediately stop the upgrade and roll back some upgraded virtual machines if, during the first batch of virtual machine commissioning and switching, a problem is encountered that cannot be immediately located and cannot be resolved in a short period of time, causing the virtual machines to fail to start normally.
[0081] Optionally, the method for upgrading and deploying the software further includes:
[0082] The second rollback unit is used to immediately isolate the deployed virtual machines and roll back the first batch of upgraded virtual machines when the parallel coupled system is in a mixed storage state and encounters problems that cannot be immediately located and cannot be resolved in a short period of time, which affect transactions, after all virtual machines in the first batch have been put into production and switched over.
[0083] Optionally, the method for upgrading and deploying the software further includes:
[0084] The third rollback unit is used to immediately isolate all deployed virtual machines and roll back all upgraded virtual machines if, during the trial operation phase of the parallel coupled system in full-scale upgrade state, a problem is encountered that cannot be immediately located and cannot be resolved in a short period of time. This is because all virtual machines have been put into production and the system is in rollback.
[0085] As can be seen from the above scheme, this application provides a software upgrade and deployment method and apparatus. The software upgrade and deployment method includes: first, splitting multiple virtual machines in a parallel coupled system according to their functional types to obtain a first batch of virtual machines and a second batch of virtual machines; wherein the functional types of the virtual machines are divided into service, network, and data; then, performing a deployment switch from the old version of the software to the new version of the software on the first batch of virtual machines to obtain a first batch of virtual machines with the new version of the software; wherein the first batch of virtual machines with the new version of the software and the second batch of virtual machines provide services to the outside world simultaneously; performing functional correctness verification on the first batch of virtual machines with the new version of the software; if the functional correctness verification of the first batch of virtual machines with the new version of the software passes, performing a deployment switch from the old version of the software to the new version of the software on the second batch of virtual machines. This effectively reduces the impact of deployment switching, completes the deployment switching work without downtime, and improves the user experience. Attached Figure Description
[0086] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0087] Figure 1 A detailed flowchart of a software upgrade and deployment method provided in this application embodiment;
[0088] Figure 2 A flowchart illustrating a software upgrade and deployment method according to another embodiment of this application;
[0089] Figure 3 A flowchart illustrating a software upgrade and deployment method according to another embodiment of this application;
[0090] Figure 4 A flowchart illustrating a software upgrade and deployment method according to another embodiment of this application;
[0091] Figure 5 A flowchart illustrating a software upgrade and deployment method according to another embodiment of this application;
[0092] Figure 6 A flowchart illustrating a software upgrade and deployment method according to another embodiment of this application;
[0093] Figure 7 This is a schematic diagram of a software upgrade and production apparatus provided in another embodiment of this application. Detailed Implementation
[0094] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0095] It should be noted that the concepts of "first," "second," etc., mentioned in this application are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies. The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0096] This application provides a method for upgrading and deploying software, such as... Figure 1 As shown, the specific steps include:
[0097] S101. According to the functional type of the virtual machine, the multiple virtual machines in the parallel coupled system are split into the first batch of virtual machines and the second batch of virtual machines.
[0098] The virtual machine's functional types are divided into service, network, and data. The implementation windows for the two batches can be selected during the lowest peak business hours, such as 2:30 AM to 6:00 AM.
[0099] Specifically, virtual machines that need to be upgraded and put into production are selected according to their functional type. In each batch, 50% of each type of virtual machine are selected to complete the production switchover. Specifically, in each batch of production switchover, 50% of the business virtual machines, 50% of the network virtual machines, and 50% of the data virtual machines are selected.
[0100] S102. Switch the first batch of virtual machines from the old version of the software to the new version of the software for production deployment, and obtain the first batch of virtual machines with the new version of the software.
[0101] The first and second batches of virtual machines in the new version of the software will provide services to the public simultaneously.
[0102] It should be noted that after the first batch of virtual machines completes the production switch from the old version to the new version, the current parallel coupled system contains a coexistence of the first batch of virtual machines with the new version and the second batch of virtual machines still using the old version. Both the old and new versions of virtual machines provide services to the outside world simultaneously. Validating the service capabilities of the newly upgraded virtual machines in this coexisting state allows for the early detection of potential problems caused by the basic software upgrade, providing a basis for rapid rollback decisions. Furthermore, since the number of virtual machines to be rolled back in this coexisting state is relatively small, rollback time is saved, and the system can quickly restore all service capabilities. This two-step production switchover approach reduces the risk factors associated with the implementation of the new version of the basic software upgrade, minimizing the impact of the production switchover. Performing a full upgrade and switchover after thorough validation in the coexisting version avoids the possibility of a low-probability, high-impact overall rollback during project implementation.
[0103] In the specific implementation of this application, the virtual machine selection method for each batch of production switching can be as follows:
[0104] 1. Selection of virtual machines to be switched to production: For business virtual machines, one method is to select one business virtual machine on each of the four physical hosts during the first batch of production switching. This can achieve the task of upgrading and switching 50% of business virtual machines in the first batch. At the same time, if the new version virtual machine fails to be deployed, the old version virtual machines distributed on the four physical hosts can provide business support, making full use of the computing resources of the four physical hosts.
[0105] 2. Selection of virtual machines to be put into production and switched over: For business virtual machines, when switching over in the first batch of production, Method 2 is that if the 4 physical hosts are installed and deployed in different data center modules or in two data center modules in different buildings, the virtual machines can be selected according to the data center module or from the perspective of building deployment. The implementation of this project can provide a rare downtime window for other work such as data center management.
[0106] 3. For network virtual machines, it is preferable to upgrade and switch to backup network virtual machines in the first batch to reduce the risk of new software versions affecting network functions.
[0107] 4. For data virtual machines, it is also recommended to prioritize the deployment and switching to backup data virtual machines to resolve the potential impact of new software product issues on the data replication function.
[0108] After verifying the new version of the virtual machines under the mixed storage state, the second batch of virtual machine production switching was carried out to complete the switching of the remaining 50% of the virtual machines.
[0109] Optionally, in another embodiment of this application, one implementation of the software upgrade and deployment method is as follows: Figure 2 As shown, it also includes:
[0110] S201. Analyze the operation plan for each type of virtual machine during the production switchover process, and obtain the optimization and adjustment results of the operation plan for each type of virtual machine during the production switchover process.
[0111] The operational plan is divided into an isolation operational plan and an integration operational plan.
[0112] Specifically, for business virtual machines that need to be upgraded and put into production, isolation is required. Special attention needs to be paid to the shutdown process of database processes to ensure that multiple database instance processes in multiple business virtual machines are not isolated at the same time. Simultaneous isolation will generate a large number of global database lock transfer operations, which will cause a surge in the consumption of host computing resources and have a significant impact on transaction performance.
[0113] For business virtual machines that need to be upgraded and put into production, isolate them, shut down the business virtual machine operating system, and isolate the business virtual machine from the parallel coupling system (cluster). In order to ensure that the impact of the virtual machine isolation process on the business is minimized, multiple virtual machine isolation actions need to be executed sequentially. Because multiple virtual machines leaving the parallel coupling system at the same time can easily increase the pressure of the global lock GRS of the parallel coupling system, which may affect the business response time.
[0114] For business virtual machines that need to be upgraded and put into production, isolate them, shut down the business virtual machine operating system, and isolate the business virtual machine from the parallel coupling system (cluster). In order to ensure that the impact of the virtual machine isolation process on the business is minimized, multiple virtual machine isolation actions need to be executed sequentially. Because multiple virtual machines leaving the parallel coupling system at the same time can easily increase the pressure of the global lock GRS of the parallel coupling system, which may affect the business response time.
[0115] To isolate network virtual machines that are undergoing software upgrades and deployment, the network support services currently being handled by the virtual machine need to be switched to another network virtual machine.
[0116] Because shutting down the transaction middleware routing controller (CICS TOR) process will affect transactions in transit; simultaneously isolating multiple database instance processes will generate a large number of global database locks for normal handover, which will impact transaction performance; after shutting down the business virtual machine operating system, isolating multiple virtual machines from the parallel coupling system (cluster) can easily lead to increased pressure on the global lock (GRS) of the parallel coupling system, which may affect business response time. Therefore, in another embodiment of this application, if the virtual machine's function type is a business virtual machine and the operation scheme is an isolation operation scheme, then one implementation of step S201 is as follows: Figure 3 As shown, it includes:
[0117] The specific isolation process and method between virtual machine services are as follows:
[0118] S301. On the host side, set the transaction middleware application controller to a static state and disable the routing policy from the gateway to the transaction middleware inter-routing controller in the transaction middleware inter-routing controller.
[0119] It should be noted that after a transaction middleware application controller is placed into a dormant state, it will no longer accept new transaction requests. Other active transaction middleware application controllers will continue transaction execution. This process will have no impact on transactions. The transaction middleware application controller will only enter the expected dormant state after all transactions are completed. Similarly, disabling the routing policy from the gateway to the inter-transaction middleware routing controller in the inter-transaction middleware routing controller will not affect in-transit transactions. The final disabling operation of the routing policy will only be completed after all in-transit transactions are finished. After disabling, the gateway will send new transactions to the routing controller with the routing policy enabled.
[0120] S302. On the gateway side, close the logical unit of the host platform network protocol layer of the transaction routing controller with the closed routing policy.
[0121] It should be noted that, on the gateway side, the transaction routing controller that closes or has closed routing policies is a logical unit at the network protocol layer of the host platform. Since no transactions use this logical unit during the closure process, it does not affect the transactions.
[0122] S303. On the host side, shut down the process corresponding to the transaction middleware routing controller.
[0123] It should be noted that since no transactions have passed through this route, closing the process will not affect transactions.
[0124] S304. On the host side, shut down the transaction middleware application controller process and isolate the business middleware of the business virtual machine.
[0125] It should be noted that since the middleware application controller is in a static state and has not processed any transaction requests, the process of shutting down the process has no impact on transactions.
[0126] The specific isolation process and method for business virtual machine database software are as follows:
[0127] S305. Shut down the database process and isolate the database software of the business virtual machine.
[0128] Specifically, the shutdown process of the database process is manually intervened to ensure that the database instance processes in multiple business virtual machines do not become isolated at the same time.
[0129] The specific process for isolating business virtual machines from the parallel coupled system (cluster) is as follows:
[0130] S306. Isolate multiple service virtual machines one by one from the parallel coupled system.
[0131] Because database software or transaction middleware software starts providing external services before it is fully ready during startup, it affects the overall service capability of the parallel coupling system (cluster); when multiple virtual machine operating systems are merged into the parallel coupling system, the resulting contention for the global lock of the parallel coupling system affects the performance of online transactions; during the simultaneous startup of multiple database instances of multiple business virtual machines, there is an interactive handover process of the global shared lock, and the database instances are also in a busy state. This process consumes a lot of host computing resources (CPU), affecting the performance of online transactions; therefore, in another embodiment of this application, if the function type of the virtual machine is a business virtual machine and the operation scheme is a merge operation scheme, then one implementation of step S201 is as follows: Figure 4 As shown, it includes:
[0132] S401. When starting the business virtual machine operating system with the new version of the media, the host platform relational database system software and transaction middleware software will not be automatically started.
[0133] It should be noted that starting the business virtual machine operating system with the new version of the medium does not automatically start the host platform's relational database system software and transaction middleware software's host platform customer information control system. This is to prevent the database software or transaction middleware software from starting external services before it is fully ready, thereby affecting the overall service capability of the parallel coupling system (cluster). At the same time, when multiple virtual machine operating systems are merged into the parallel coupling system, the concurrency of the merger must be controlled to avoid contention for the global lock of the parallel coupling system caused by multiple virtual machines joining the parallel coupling system (cluster) at the same time, so as to avoid affecting the performance of online transactions.
[0134] For business virtual machines that are integrated into the parallel coupled system with the new version of the basic software, the database software is started manually. During the startup of multiple database instances on multiple virtual machines, it is necessary to control the concurrency of the number of instances started and start them in a sequential order. Because multiple database instances start simultaneously, there will be an interactive handover process of the global shared lock, and the database instances will be in a busy state. This process will consume a lot of host computing resources (CPU). Therefore, the concurrency of the database instance integration process must be strictly controlled to avoid affecting online transactions.
[0135] S402 controls multiple business virtual machines to be successively integrated into a parallel coupled system.
[0136] This avoids contention for the global lock of the parallel coupled system when multiple virtual machines join the parallel coupled system (cluster) at the same time.
[0137] S403: Control the concurrent startup of database instances on multiple business virtual machines, and start them in sequence.
[0138] This reduces database contention for global shared locks.
[0139] For business virtual machines that are integrated into the parallel coupled system using the new version of the basic software, the specific integration process during the startup of the transaction middleware software is as follows:
[0140] 1) Perform the actions required to upgrade the trading middleware, enabling the new version of the trading middleware to start;
[0141] 2) Manually start the transaction middleware application controller and set it to a dormant state to prevent the new version of the application controller from responding to transaction requests and causing transaction anomalies if the middleware application controller is not started properly;
[0142] 3) Check the startup status of each process of the middleware launched by the new version to confirm that the process starts normally and meets the upgrade expectations;
[0143] 4) Change the middleware application controller state from static to external service state, waiting for transaction requests to be received;
[0144] 5) Start the transaction middleware routing controller to receive transaction requests from the gateway.
[0145] Because if the control point is not actively triggered, a network switch will occur during the process of shutting down the virtual remote communication access method process under the operating system of the network virtual machine, causing transaction jitter and affecting business response time. After shutting down the operating system of the network virtual machine, isolating multiple virtual machines from the parallel coupling system (cluster) can easily lead to an increase in global lock pressure on the parallel coupling system, which may affect business response time. Therefore, in another embodiment of this application, if the function type of the virtual machine is a network virtual machine and the operation scheme is an isolation operation scheme, then one implementation of step S201 is as follows: Figure 5 As shown, it includes:
[0146] S501. Switch the active dependent logical unit requester connection in the network virtual machine to run on another network virtual machine.
[0147] It should be noted that this operation has no impact on business operations.
[0148] S502. Switch the active inter-node control point session to another network virtual machine to run.
[0149] It should be noted that this operation has no impact on business operations.
[0150] S503. On the host side, the physical units and logical units corresponding to each gateway node are activated sequentially.
[0151] It should be noted that this operation has no impact on business operations.
[0152] S504. On the host side, activate the control points corresponding to each gateway node.
[0153] It should be noted that this operation has no impact on business operations. However, if the control point is not actively shut down, the process of shutting down the operating system of the network virtual machine and removing the Virtual Telecommunication Access Method (VTAM) process will cause a network protocol switch within the host platform, resulting in transaction jitter and affecting business response time.
[0154] S505: Isolate multiple network virtual machines one by one from the parallel coupled system.
[0155] It should be noted that, in order to minimize the impact of isolating virtual machines on business operations, the isolation process of multiple virtual machines needs to be executed sequentially. If multiple virtual machines leave the parallel coupled system at the same time, it may increase the global lock pressure of the parallel coupled system, which may affect the business response time.
[0156] Because the merging of multiple virtual machine operating systems into a parallel coupled system causes contention for the global lock of the parallel coupled system, affecting the performance of online transactions, in another embodiment of this application, if the virtual machine's function type is a network virtual machine and the operation scheme is an merging operation scheme, then one implementation of step S201 includes:
[0157] Control multiple network virtual machines to be successively incorporated into a parallel coupled system.
[0158] This avoids contention for the global lock of the parallel coupled system when multiple virtual machines join the parallel coupled system (cluster) at the same time.
[0159] In the specific implementation of this application, the network virtual machines of the parallel coupling system are integrated with the new version of the basic software, and the operating system is started with the new version of the medium. The network function is automatically restored without any additional manual intervention. At the same time, when multiple virtual machines are integrated into the parallel coupling system, the concurrency of integration must be controlled to avoid contention for the global lock of the parallel coupling system caused by multiple virtual machines joining the parallel coupling system (cluster), so as to avoid affecting the performance of online transactions.
[0160] Because shutting down the data virtual machine operating system and isolating multiple virtual machines from the parallel coupling system (cluster) can easily increase the global lock pressure on the parallel coupling system, potentially affecting service response time. Therefore, in another embodiment of this application, if the virtual machine's function type is a data virtual machine and the operation scheme is an isolation operation scheme, then one implementation of step S201 includes:
[0161] Multiple virtual machines are isolated one by one from the parallel coupled system.
[0162] It should be noted that isolating a data virtual machine can directly shut down the data virtual machine's operating system. The data replication software (Geographically Dispersed Parallel Sysplex peer-to-peer remote copy, GDPS) will be managed by the automation tool (System Automation, SA) and shut down along with the operating system. To ensure that the impact of isolating virtual machines on business is minimized, the isolation process of multiple virtual machines needs to be executed sequentially. If multiple virtual machines leave the parallel coupled system at the same time, it may increase the global lock pressure of the parallel coupled system, which may affect the business response time.
[0163] Because the merging of multiple virtual machine operating systems into a parallel coupled system triggers contention for the global lock of the parallel coupled system, it affects the performance of online transactions. Therefore, in another embodiment of this application, if the virtual machine's function type is a data virtual machine and the operation scheme is an isolated operation scheme, then one implementation of step S201 includes:
[0164] Control multiple virtual data machines to be successively incorporated into a parallel coupled system.
[0165] This avoids contention for the global lock of the parallel coupled system when multiple virtual machines join the parallel coupled system (cluster) at the same time.
[0166] In the specific implementation of this application, for data virtual machines that are integrated into the parallel coupling system with the new version of the basic software, the operating system is started with the new version of the medium, the data replication software GDPS is started along with the automation tool software, and the data replication management function is automatically restored without the need for additional manual intervention. At the same time, when multiple virtual machine operating systems are integrated into the parallel coupling system, the concurrency of integration must be controlled to avoid contention for the global lock of the parallel coupling system caused by multiple virtual machines joining the parallel coupling system (cluster), so as to avoid affecting the performance of online transactions.
[0167] S202. Analyze the gateway restart, obtain the impact of the gateway restart, and improve the gateway restart method.
[0168] It should be noted that if the transaction middleware routing controller process restarts, each gateway needs to restart to restore its logical units with the host transaction middleware routing controller. During the reconstruction of logical units, the gateway will try to make the number of logical units on the gateway and each transaction middleware routing controller as similar as possible to achieve load balancing. For newly started middleware routing controllers, since there are no logical units distributed on them yet, the gateway restart operation will cause all newly established logical units to be established on the newly started transaction routing controller. After the gateway restarts, although the number of logical units distributed on all middleware routing controllers is the same, the goal of full connectivity between each gateway and all middleware routing controllers has not been achieved.
[0169] If the logical unit does not maintain a full connection with all middleware routing controllers, it will cause differences in the pressure borne by each business virtual machine on the backend host, and the computing resources of the physical host cannot be redistributed, ultimately leading to a decrease in transaction performance and an increase in response time.
[0170] The existing solution is to perform a full restart of all gateways after all middleware routing controllers have been restarted. However, this method prevents transactions from being sent to the host through the gateway during the gateway restart phase. During this period, the system cannot provide services to the outside world, causing business interruption and downtime.
[0171] To avoid downtime issues caused by fully enabling all gateways, and to achieve the goal of full connectivity between each gateway and all middleware routing controllers on the host, with an equal and evenly distributed number of logical units in each middleware routing controller, in another embodiment of this application, one implementation of step S202 is as follows: Figure 6 As shown, it includes:
[0172] S601. Reserve m gateways without restarting to ensure service availability, and shut down the remaining nm gateways.
[0173] There are n gateway devices connected to the gateway host via the SNA protocol, where n and m are positive integers, and m <n / 2。
[0174] S602. Start m gateway hosts in the gateways that have been shut down, and rebuild the logical unit on the newly started m gateway hosts.
[0175] It should be noted that this will make the number of logical units on all transaction routing controllers on the host the same. At this time, each gateway is only connected to half of the transaction routing controllers on the host, and the goal of full connectivity has not yet been met.
[0176] S603. Start n-2m gateway hosts that are in a powered-off state.
[0177] During the process of rebuilding logical units, n-2m gateway hosts will send uniform logical unit creation requests to all transaction routing controllers.
[0178] Since the number of logical units on each transaction routing controller is already the same, the n-2m gateway hosts will send uniform logical unit creation requests to all transaction routing controllers during the logical unit reconstruction process. After the newly started n-2m gateways start, they will achieve a fully connected state with all transaction routing controllers, and the number of logical units distributed on each transaction routing controller will be balanced.
[0179] S604. Reboot m gateways that have never been rebooted.
[0180] It should be noted that the newly created logical units of the m gateways that have never been restarted also reach a fully connected state after the restart, and the number of logical units on each transaction routing controller is the same.
[0181] S605, the second restart of m gateway hosts that have already been started once.
[0182] It should be noted that after the second restart of the m gateway hosts that have been started once, the newly created logical units of these m gateways also reach a fully connected state, and the number of logical units on each transaction routing controller is the same.
[0183] Once all gateways have restarted, without employing a full shutdown and full restart approach, all gateway hosts have rebuilt their logical units with all transaction routing controllers in a fully connected manner. At the same time, the number of gateway hosts distributed across the logical units of each transaction routing controller has been balanced, and the business continuity requirement has been met without a full shutdown and full restart of the gateways. To ensure that the gateway restart operation has minimal impact on transactions, the shutdown process of multiple gateways can optionally be executed sequentially.
[0184] S103. Verify the functionality of the first batch of virtual machines in the new version of the software.
[0185] S104. If the first batch of virtual machines using the new version of the software passes the functional correctness verification, the second batch of virtual machines will be switched from the old version of the software to the new version for production deployment.
[0186] It should be noted that in the actual application process, this application will also formulate flexible decision-making plans and targeted emergency rollback plans at each stage of production to ensure the normal implementation of the upgrade and production project, so that the decision-making is based on evidence and the emergency plan is accurate and effective.
[0187] If, during the first batch of virtual machine deployment and switchover, a problem is encountered that cannot be immediately identified as the root cause and cannot be resolved in a short period of time, causing the virtual machines to fail to start normally, the upgrade deployment should be stopped immediately, and the upgraded virtual machines should be rolled back.
[0188] After all virtual machines in the first batch have been put into production and switched over, if the parallel coupled system encounters problems that cannot be immediately located and cannot be resolved in a short period of time while running externally in a mixed storage state, and these problems affect transactions, then the virtual machines that have been put into production will be immediately isolated and the first batch of upgraded virtual machines will be rolled back.
[0189] If all virtual machines have been put into production and switched over, and problems are encountered during the trial operation phase of the parallel coupled system in full upgrade state that cannot be immediately located and cannot be resolved in a short period of time, then all virtual machines that have been put into production should be immediately isolated and all upgraded virtual machines should be rolled back.
[0190] It should be noted that, in the specific implementation of this application, the rollback scheme includes, but is not limited to, the following:
[0191] Rollback Option 1:
[0192] 1) Shut down any virtual machines (VMs) that have been started with the new version;
[0193] 2) Start the virtual machine with the old version of the software and merge it into the newly started virtual machine;
[0194] 3) Restart the gateway.
[0195] Rollback Option 2:
[0196] 1) Isolate the first batch of business virtual machines, network virtual machines, and data virtual machines that have been deployed using the new version of the software;
[0197] 2) Start all virtual machines in the first batch of production using the old version of the software and merge them into the newly started virtual machines;
[0198] 3) Restart the gateway.
[0199] Rollback Option 3:
[0200] 1) Stop the gateway;
[0201] 2) Shut down the first network virtual machine;
[0202] 3) Start the first network virtual machine with the old version of the software to complete the merging of the first network virtual machine;
[0203] 4) Shut down all remaining virtual machines;
[0204] 5) Start all virtual machines except the first network virtual machine using the old version of the software to complete the virtual machine merging;
[0205] 6) Restart the gateway.
[0206] As can be seen from the above scheme, this application provides a method for upgrading and deploying software: First, multiple virtual machines in a parallel coupled system are split according to their functional types to obtain a first batch of virtual machines and a second batch of virtual machines; the functional types of the virtual machines are divided into service, network, and data; then, the first batch of virtual machines is switched from the old version of the software to the new version for deployment, resulting in a first batch of virtual machines with the new version of the software; both the first batch of virtual machines with the new version of the software and the second batch of virtual machines provide services simultaneously; the functional correctness of the first batch of virtual machines with the new version of the software is verified; if the functional correctness verification of the first batch of virtual machines with the new version of the software passes, the second batch of virtual machines is switched from the old version of the software to the new version for deployment. This effectively reduces the impact of deployment switching and improves the user experience.
[0207] In another embodiment of this application, one implementation of the software upgrade and production deployment apparatus is as follows: Figure 7 As shown, it includes:
[0208] Splitting unit 701 is used to split multiple virtual machines in a parallel coupled system according to the functional type of the virtual machines, to obtain the first batch of virtual machines and the second batch of virtual machines.
[0209] Virtual machines are categorized into three functional types: business, network, and data.
[0210] The first production switching unit 702 is used to switch the first batch of virtual machines from the old version of the software to the new version of the software, so as to obtain the first batch of virtual machines with the new version of the software.
[0211] The first and second batches of virtual machines in the new version of the software will provide services to the public simultaneously.
[0212] Verification unit 703 is used to verify the functional correctness of the first batch of virtual machines in the new version of the software.
[0213] The second production switching unit 704 is used to switch the second batch of virtual machines from the old version of the software to the new version of the software if the first batch of virtual machines passes the functional correctness verification.
[0214] For details on the specific working process of the units disclosed in the above embodiments of this application, please refer to the corresponding method embodiments, such as... Figure 1 As shown, it will not be elaborated further here.
[0215] Optionally, in another embodiment of this application, one implementation of the software upgrade and production deployment apparatus further includes:
[0216] The first analysis unit is used to analyze the operation plan of virtual machines of each functional type during the production switchover process, and obtain the optimization and adjustment results of the operation plan of virtual machines of each functional type during the production switchover process.
[0217] The operational plan is divided into an isolation operational plan and an integration operational plan.
[0218] The second analysis unit is used to analyze the gateway restart, obtain the impact of the gateway restart, and improve the gateway restart method.
[0219] For details on the specific working process of the units disclosed in the above embodiments of this application, please refer to the corresponding method embodiments, such as... Figure 2 As shown, it will not be elaborated further here.
[0220] Optionally, in another embodiment of this application, the virtual machine's functional type is a business virtual machine, and one implementation of the first analysis unit further includes:
[0221] The first shutdown unit is used to set the transaction middleware application controller to a static state on the host side and to shut down the routing policy from the gateway to the inter-transaction middleware routing controller in the inter-transaction middleware routing controller if the operation scheme is an isolated operation scheme.
[0222] The second shutdown unit is a logical unit at the host platform network protocol layer used on the gateway side to shut down the transaction routing controller with a closed routing policy.
[0223] The third shutdown unit is used on the host side to shut down the process corresponding to the transaction middleware routing controller.
[0224] The fourth shutdown unit is used on the host side to shut down the transaction middleware application controller process and isolate the business middleware of the business virtual machine.
[0225] The fifth shutdown unit is used to shut down the database process and isolate the database software of the business virtual machine.
[0226] The first isolation unit is used to isolate multiple business virtual machines one by one from the parallel coupled system.
[0227] For details on the specific working process of the units disclosed in the above embodiments of this application, please refer to the corresponding method embodiments, such as... Figure 3 As shown, it will not be elaborated further here.
[0228] Optionally, in another embodiment of this application, one implementation of the first analysis unit further includes:
[0229] The "Start-up" unit is used to start the business virtual machine operating system with the new version media if the operation plan is a merge operation plan, without automatically starting the host platform relational database system software and transaction middleware software.
[0230] The first control unit is used to control multiple business virtual machines to be successively integrated into the parallel coupling system.
[0231] The second control unit is used to control the concurrent and sequential startup of database instances on multiple business virtual machines.
[0232] For details on the specific working process of the units disclosed in the above embodiments of this application, please refer to the corresponding method embodiments, such as... Figure 4 As shown, it will not be elaborated further here.
[0233] Optionally, in another embodiment of this application, the virtual machine's functional type is a network virtual machine, and one implementation of the first analysis unit further includes:
[0234] The first switching unit is used to switch the connection of an active dependent logical unit requester in a network virtual machine to another network virtual machine if the operation scheme is an isolation operation scheme.
[0235] The second switching unit is used to switch an active inter-node control point session to run on another network virtual machine.
[0236] The first kill unit is used to kill the physical and logical units corresponding to each gateway node in sequence on the host side.
[0237] The second kill unit is used to kill the control points corresponding to each gateway node on the host side.
[0238] The second isolation unit is used to isolate multiple network virtual machines one by one from the parallel coupled system.
[0239] For details on the specific working process of the units disclosed in the above embodiments of this application, please refer to the corresponding method embodiments, such as... Figure 5 As shown, it will not be elaborated further here.
[0240] Optionally, in another embodiment of this application, the virtual machine's functional type is a network virtual machine, and one implementation of the first analysis unit further includes:
[0241] The third control unit is used to control multiple network virtual machines to be successively merged into the parallel coupled system if the operation scheme is an incorporation operation scheme.
[0242] For details on the specific working process of the units disclosed in the above embodiments of this application, please refer to the corresponding method embodiments, which will not be repeated here.
[0243] Optionally, in another embodiment of this application, the virtual machine's functional type is a data virtual machine, and one implementation of the first analysis unit further includes:
[0244] The third isolation unit is used to isolate multiple virtual data machines from the parallel coupled system one by one if the operation scheme is an isolation operation scheme.
[0245] For details on the specific working process of the units disclosed in the above embodiments of this application, please refer to the corresponding method embodiments, which will not be repeated here.
[0246] Optionally, in another embodiment of this application, the virtual machine's functional type is a data virtual machine, and one implementation of the first analysis unit further includes:
[0247] The fourth control unit is used to control multiple data virtual machines to be successively merged into the parallel coupled system if the operation scheme is an incorporation operation scheme.
[0248] For details on the specific working process of the units disclosed in the above embodiments of this application, please refer to the corresponding method embodiments, which will not be repeated here.
[0249] Optionally, in another embodiment of this application, one implementation of the second analysis unit further includes:
[0250] The shutdown unit is used to reserve m gateways without restarting to ensure service availability, and to shut down the remaining nm gateways.
[0251] There are n gateway devices connected to the gateway host via the SNA protocol, where n and m are positive integers, and m <n / 2。
[0252] The first startup unit is used to start m gateway hosts that have been shut down, and the newly started m gateway hosts are used to rebuild the logic unit.
[0253] The second startup unit is used to start n-2m gateway hosts that are in a powered-off state.
[0254] During the process of rebuilding logical units, n-2m gateway hosts will send uniform logical unit creation requests to all transaction routing controllers.
[0255] The first restart unit is used to restart m gateways that have never been restarted before.
[0256] The second restart unit is used to restart m gateway hosts that have already been started once.
[0257] For details on the specific working process of the units disclosed in the above embodiments of this application, please refer to the corresponding method embodiments, such as... Figure 6 As shown, it will not be elaborated further here.
[0258] Optionally, in another embodiment of this application, one implementation of the software upgrade and production deployment apparatus further includes:
[0259] The first rollback unit is used to immediately stop the upgrade and roll back some upgraded virtual machines if, during the first batch of virtual machine commissioning and switching, a problem is encountered that cannot be immediately located and cannot be resolved in a short period of time, causing the virtual machines to fail to start normally.
[0260] For details on the specific working process of the units disclosed in the above embodiments of this application, please refer to the corresponding method embodiments, which will not be repeated here.
[0261] Optionally, in another embodiment of this application, one implementation of the software upgrade and production deployment apparatus further includes:
[0262] The second rollback unit is used to immediately isolate the deployed virtual machines and roll back the first batch of upgraded virtual machines when the parallel coupled system is in a mixed storage state and encounters problems that cannot be immediately located and cannot be resolved in a short period of time, which affect transactions, after all virtual machines in the first batch have been put into production and switched over.
[0263] For details on the specific working process of the units disclosed in the above embodiments of this application, please refer to the corresponding method embodiments, which will not be repeated here.
[0264] Optionally, in another embodiment of this application, one implementation of the software upgrade and production deployment apparatus further includes:
[0265] The third rollback unit is used to immediately isolate all deployed virtual machines and roll back all upgraded virtual machines if, during the trial operation phase of the parallel coupled system in full-scale upgrade state, a problem is encountered that cannot be immediately located and cannot be resolved in a short period of time. This is because all virtual machines have been put into production and the system is in rollback.
[0266] For details on the specific working process of the units disclosed in the above embodiments of this application, please refer to the corresponding method embodiments, which will not be repeated here.
[0267] As can be seen from the above scheme, this application provides a software upgrade and deployment device: First, the splitting unit 701 splits multiple virtual machines in the parallel coupled system according to their functional types, obtaining a first batch of virtual machines and a second batch of virtual machines; wherein, the functional types of the virtual machines are divided into service, network, and data; then, the first deployment switching unit 702 performs deployment switching from the old version of the software to the new version of the software on the first batch of virtual machines, obtaining a first batch of virtual machines with the new version of the software; wherein, the first batch of virtual machines with the new version of the software and the second batch of virtual machines provide services to the outside world simultaneously; the verification unit 703 performs functional correctness verification on the first batch of virtual machines with the new version of the software; if the functional correctness verification of the first batch of virtual machines with the new version of the software passes, the second deployment switching unit 704 performs deployment switching from the old version of the software to the new version of the software on the second batch of virtual machines. This effectively reduces the impact of deployment switching and improves the user experience.
[0268] In the embodiments disclosed in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus and method embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0269] Furthermore, the functional modules in the various embodiments of this disclosure can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part. If the functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a live streaming device, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing program code.
[0270] Those skilled in the art will be able to implement or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for upgrading and deploying software, characterized in that, include: The parallel coupled system is divided into a first batch of virtual machines and a second batch of virtual machines according to their functional types. The functional types of the virtual machines are divided into service, network, and data. The first batch of virtual machines is switched from the old version of the software to the new version of the software for production deployment, resulting in the first batch of virtual machines with the new version of the software; wherein, the first batch of virtual machines with the new version of the software and the second batch of virtual machines provide services to the outside world simultaneously. The first batch of virtual machines using the new version of the software were functionally corrected. If the first batch of virtual machines using the new version of the software passes the functional correctness verification, the second batch of virtual machines will be switched from the old version of the software to the new version for production deployment. The operation schemes for virtual machines of each functional type during the production switchover process are analyzed, and the optimization and adjustment results of the operation schemes for each functional type of virtual machine during the production switchover process are obtained; wherein, the operation schemes are divided into isolation operation schemes and merging operation schemes; The gateway restart was analyzed to determine its impact, and the gateway restart method was improved. The analysis of gateway restarts to determine their impact and the improvement of the gateway restart method include: To ensure service availability, m gateways are reserved without restarting, while the remaining nm gateways are shut down. A total of n gateway devices are connected to the gateway host via the SNA protocol, where n and m are positive integers, and m... <n / 2; Start m gateway hosts that have been shut down, and rebuild the logical unit on the newly started m gateway hosts; Start up n-2m gateway hosts that are in a shutdown state, wherein the n-2m gateway hosts will send uniform logical unit creation requests to all transaction routing controllers during the logical unit reconstruction process; Reboot m gateways that have never been rebooted before; The second restart involves m gateway hosts that have already been started once.
2. The upgrade and commissioning method according to claim 1, characterized in that, The virtual machine's functional type is a business virtual machine. The analysis of the operation plan during the production switchover process for each functional type of virtual machine yields optimized adjustment results for each type of virtual machine during the production switchover process, including: If the operation scheme is an isolated operation scheme, then on the host side, the transaction middleware application controller is set to a static state, and the routing policy from the gateway to the inter-transaction middleware routing controller is turned off in the inter-transaction middleware routing controller; On the gateway side, the logical unit of the host platform network protocol layer of the transaction routing controller that has a closed routing policy is shut down; On the host side, shut down the process corresponding to the transaction middleware routing controller; On the host side, shut down the transaction middleware application controller process and isolate the business middleware of the business virtual machine; Shut down the database process and isolate the database software on the business virtual machine; Isolate multiple business virtual machines one by one from the parallel coupled system; If the operation plan is to merge the operation plan, the business virtual machine operating system will be started with the new version of the medium, and the host platform relational database system software and transaction middleware software will not be automatically started. Control multiple business virtual machines to be successively integrated into a parallel coupled system; Control the concurrent startup of database instances on multiple business virtual machines, and start them in sequence.
3. The upgrade and production commissioning method according to claim 1, characterized in that, The virtual machine's functional type is a network virtual machine. The operational plan for each functional type of virtual machine during the production switchover process is analyzed, and the optimized adjustment results of the operational plan for each functional type of virtual machine during the production switchover process are obtained, including: If the operation plan is an isolation operation plan, then the active dependent logical unit requester connection in the network virtual machine will be switched to run on another network virtual machine. Switch the active inter-node control point session to run on another network virtual machine; On the host side, the physical and logical units corresponding to each gateway node are sequentially activated and killed. On the host side, the control points corresponding to each gateway node are deactivated. Isolate multiple network virtual machines one by one from the parallel coupled system; If the operation scheme is an incorporation operation scheme, control multiple network virtual machines to be incorporated into the parallel coupled system one by one.
4. The upgrade and production commissioning method according to claim 1, characterized in that, The virtual machine's functional type is a data virtual machine. The analysis of the operation plan during the production switchover process for each functional type of virtual machine yields optimized adjustment results for each type, including: If the operation scheme is an isolation operation scheme, then multiple data virtual machines will be isolated one by one from the parallel coupled system. If the operation scheme is a merge operation scheme, control multiple data virtual machines to merge into the parallel coupled system one by one.
5. The upgrade and commissioning method according to claim 1, characterized in that, include: If, during the first batch of virtual machine deployment and switchover, a problem is encountered that cannot be immediately identified as the root cause and cannot be resolved within a preset time, causing the virtual machine to fail to start normally, the upgrade deployment will be immediately stopped and the upgraded virtual machines will be rolled back.
6. The upgrade and production commissioning method according to claim 1, characterized in that, include: After all virtual machines in the first batch have been put into production and switched over, if the parallel coupled system encounters problems that cannot be immediately located and cannot be resolved within a preset time when running externally in a mixed storage state, and these problems affect transactions, then the virtual machines that have been put into production will be immediately isolated and the first batch of upgraded virtual machines will be rolled back.
7. The upgrade and commissioning method according to claim 1, characterized in that, include: If all virtual machines have been put into production and switched over, and problems are encountered during the trial operation phase of the parallel coupled system in full upgrade state that cannot be immediately located and cannot be resolved within a preset time, then all virtual machines that have been put into production should be immediately isolated and all upgraded virtual machines should be rolled back.
8. A software upgrade and production deployment device, characterized in that, include: The splitting unit is used to split multiple virtual machines in the parallel coupled system according to their functional types to obtain a first batch of virtual machines and a second batch of virtual machines; wherein, the functional types of the virtual machines are divided into service, network and data. The first production switching unit is used to switch the first batch of virtual machines from the old version of the software to the new version of the software, so as to obtain the first batch of virtual machines with the new version of the software; wherein, the first batch of virtual machines with the new version of the software and the second batch of virtual machines provide services to the outside world at the same time. The verification unit is used to verify the functional correctness of the first batch of virtual machines in the new version of the software. The second production switching unit is used to switch the second batch of virtual machines from the old version of the software to the new version of the software if the first batch of virtual machines passes the functional correctness verification. The first analysis unit is used to analyze the operation plan of virtual machines of each functional type during the production switchover process, and obtain the optimization and adjustment results of the operation plan of virtual machines of each functional type during the production switchover process; wherein, the operation plan is divided into isolation operation plan and merging operation plan; The second analysis unit is used to analyze the gateway restart, obtain the degree of impact of the gateway restart, and improve the gateway restart method; The second analysis unit includes: The shutdown unit is used to reserve m gateways without restarting to ensure service availability, and to shut down the remaining nm gateways. There are n gateway devices connected to the gateway host via the SNA protocol, where n and m are positive integers, and m... <n / 2; The first startup unit is used to start m gateway hosts that have been shut down, and the newly started m gateway hosts are used to rebuild the logic unit. The second startup unit is used to start n-2m gateway hosts that are in a shutdown state; wherein, during the process of rebuilding the logical unit, the n-2m gateway hosts will send a uniform logical unit creation request to all transaction routing controllers. The first restart unit is used to restart m gateways that have never been restarted before; The second restart unit is used to restart m gateway hosts that have already been started once.
Citation Information
Patent Citations
Software defined automation system and architecture
CN108513655A
Method for application and practice of software-defined data center in operator network
CN112039682A