Method for maintaining edge cloud computing node, operation and maintenance platform and cloud computing system
By generating automated workflows for taking faulty devices offline and deploying new devices, the maintenance issues of edge cloud computing nodes were resolved, achieving end-to-end automated maintenance of edge cloud computing nodes and ensuring the continuity of service capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ALIBABA CLOUD COMPUTING CO LTD
- Filing Date
- 2023-03-21
- Publication Date
- 2026-07-14
AI Technical Summary
Existing technologies lack automated maintenance solutions for edge cloud computing nodes, and maintenance solutions for central clouds are not applicable to edge cloud computing nodes deployed in customer data centers.
A maintenance method for edge cloud computing nodes is provided, including generating workflows for taking faulty devices offline and deploying new devices. By automating tasks, the method completes the offline process for faulty devices and the deployment of new devices, ensuring the continuity of service capabilities.
It enables fully automated maintenance of edge cloud computing nodes, ensuring the service capabilities of edge cloud computing nodes with hardware failures in the production environment and addressing the shortcomings of existing maintenance solutions.
Smart Images

Figure CN116319279B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cloud computing, and more specifically, to maintenance methods, operation and maintenance platforms, and cloud computing systems for edge cloud computing nodes. Background Technology
[0002] Distributed cloud is a cloud computing model that deploys cloud services on demand to different geographical locations and provides unified management capabilities. Distributed cloud can take the form of central cloud, regional cloud, and edge cloud. Among them, edge cloud is located at the network edge as close as possible to the source of things and data, providing elastically scalable cloud service capabilities, with features such as fast response, low latency, and lightweight computing, and can support collaboration with central cloud or regional cloud.
[0003] In some scenarios, edge computing nodes extend the computing, storage, and network infrastructure of public clouds to the customer's local data center in a hardware-software integrated manner to meet the customer's needs for data security, local data processing, and low latency. Unlike centralized clouds, these edge computing nodes are deployed in the customer's data center, and existing maintenance solutions for centralized clouds are not applicable to these edge computing nodes. Currently, there are no automated maintenance solutions for these edge computing nodes. Summary of the Invention
[0004] This application provides a maintenance method, operation and maintenance platform, and cloud computing system for edge cloud computing nodes, with the aim of realizing an automated maintenance solution for edge cloud computing nodes.
[0005] Firstly, this application provides a method for repairing an edge cloud computing node, comprising:
[0006] In response to edge cloud computing node failure and maintenance events, generate workflows for taking faulty equipment offline and deploying new equipment.
[0007] The new equipment deployment workflow identifies target standby machines in the standby pool as new equipment to be deployed.
[0008] The status of the faulty equipment is changed to a non-production status through the faulty equipment offline workflow.
[0009] The target standby machine is installed using the new equipment deployment workflow, and the target standby machine is added to the cloud computing node cluster where the faulty equipment is located, and the status of the target standby machine is marked as non-production status.
[0010] After removing the faulty device from the user's computer room and installing the target backup device in the user's computer room, the status of the target backup device is changed to production status through the new device deployment workflow to end the new device deployment workflow;
[0011] After the faulty equipment is repaired, the repaired faulty equipment is tested through the faulty equipment offline workflow, and the faulty equipment that passes the test is added to the standby pool to end the faulty equipment offline workflow.
[0012] In one implementation, after modifying the state of the faulty device to a non-production state, the method further includes:
[0013] The faulty equipment offline workflow generates a faulty equipment logistics order, which is used to instruct the faulty equipment to be transported from the user's computer room to the repair location.
[0014] After marking the target standby machine as non-production, the method further includes:
[0015] A new equipment logistics order is generated through the new equipment deployment workflow. The new equipment logistics order is used to instruct the delivery of the target backup equipment to the user's data center. The departure date of the faulty equipment logistics order is the arrival date of the new equipment logistics order.
[0016] In one implementation, after modifying the state of the faulty device to a non-production state, the method further includes:
[0017] Disable monitoring and alarms for the faulty device.
[0018] In one implementation, the step of installing the target standby machine through the new device deployment workflow and adding the target standby machine to the cloud computing node cluster where the faulty device resides includes:
[0019] The logical configuration information of the target standby device is modified to the logical configuration information of the faulty device through the new device deployment workflow;
[0020] Perform switch port allocation and production route configuration on the target standby machine;
[0021] The target standby machine is then installed with an operating system and service components.
[0022] Configure the network for the target standby machine;
[0023] The target standby machine is added to the cloud computing node cluster in an expansion manner, using the same computing service component version as the cloud computing node cluster.
[0024] In one implementation, after marking the target standby machine as non-production, the method further includes:
[0025] Add a preset tag to the target standby machine, the preset tag being used to indicate that the target standby machine is the user's edge cloud computing node.
[0026] In one implementation, after marking the target standby machine as non-production, the method further includes:
[0027] Scan the target backup device to determine if preset information exists; if it does, delete the preset information.
[0028] In one implementation, before marking the target standby machine's state as non-production, the method further includes:
[0029] Lock the target backup device;
[0030] After marking the target standby machine as non-production, the method further includes:
[0031] The monitoring and alarms for the target backup machine are turned off, and the target backup machine is powered off.
[0032] In one implementation, after the target standby machine is racked in the user's computer room, the method further includes:
[0033] Change the network mode of the target standby machine to field mode.
[0034] In one implementation, after modifying the network mode of the target standby machine to the field mode, the method further includes:
[0035] The target standby machine and central cloud are subject to full lifecycle management, network and operation and maintenance collaborative verification.
[0036] In one implementation, before modifying the state of the target standby machine to production state via the new equipment deployment workflow, the method further includes:
[0037] Enable monitoring and alarms for the target backup machine, and unlock the target backup machine.
[0038] In one implementation, the edge cloud computing node failure repair event is confirmed and triggered after the user migrates or releases services on the faulty device.
[0039] Secondly, this application provides an operation and maintenance platform for an edge cloud computing node, including: a memory and a processor;
[0040] The memory is used to store computer programs;
[0041] The processor is configured to execute a computer program stored in the memory, wherein the computer program, when executed, causes the processor to perform the method described in the first aspect.
[0042] Thirdly, this application provides a cloud computing system, including edge cloud computing nodes and an operation and maintenance platform for the cloud computing nodes as described in the second aspect.
[0043] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the method described in the first aspect.
[0044] The edge cloud computing node maintenance method, operation and maintenance platform and cloud computing system provided in this application provide a fully automated maintenance solution for edge cloud computing nodes that experience hardware failures in the production environment. This solution enables the removal of faulty devices from the network and the replacement of faulty devices with new devices, thus ensuring the service capabilities of edge cloud computing nodes. Attached Figure Description
[0045] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0046] Figure 1 This is a flowchart illustrating a maintenance method for an edge computing node provided in an embodiment of this application;
[0047] Figure 2 This is a flowchart illustrating another method for maintaining an edge cloud computing node provided in an embodiment of this application;
[0048] Figure 3 This is a flowchart illustrating another edge computing node maintenance method provided in an embodiment of this application;
[0049] Figure 4 This is a schematic block diagram of the operation and maintenance platform for edge cloud computing nodes provided in the embodiments of this application. Detailed Implementation
[0050] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0051] First, the terminology used in the embodiments of this application will be introduced.
[0052] 1. CloudBox: A fully managed cloud service that extends and deploys the computing, storage, and networking infrastructure of public cloud to the customer's local data center in a hardware and software integrated manner to meet the needs of data security, local data processing, and low latency.
[0053] 2. Distributed cloud: Public cloud services (usually including the necessary hardware and software) are distributed to different physical locations (i.e., the edge), while the ownership, operation, governance, updates and development of the services remain the responsibility of the original public cloud provider.
[0054] 3. Edge computing nodes: Nodes deployed at the network edge where things and data sources originate, including the necessary hardware and software, to provide cloud computing services nearby.
[0055] 4. Cloudbot: A comprehensive operation and maintenance platform for managing cloud box deployment, maintenance, and expansion.
[0056] 5. Workflow: A series of automated tasks arranged in a certain order to complete a specific function.
[0057] 6. Cloud Box Backup Server Pool: A unified resource pool for managing the inventory, scheduling, and hardware checks of cloud box backup servers.
[0058] 7. Elastic Compute Service (ECS) Event System: An event management platform that integrates all ECS event sources, including daily operations and maintenance, service failures, alarms, and routine management.
[0059] 8. Cloud Enterprise Network (CEN): A global network service that enables the rapid construction of hybrid cloud and distributed business systems.
[0060] 9. Virtual Private Cloud (VPC): A logically isolated private cloud.
[0061] 10. Out-of-Band Management System (OOB): A server remote control and management system independent of the production network. This system uses a separate network unrelated to the production network, so it can serve as an emergency channel to remotely repair and handle the production network when it fails and becomes inaccessible, in order to achieve the goal of recovering losses as quickly as possible.
[0062] 11. Tianji: A distributed data center hardware and software management platform that provides lifecycle management capabilities for software systems such as operating systems and application services on top of hardware such as servers / network devices.
[0063] 12. staragent: Server operation and maintenance infrastructure, which is the channel for interaction between the operation and maintenance system and the server.
[0064] 13. Configuration Management Database (CMDB): A platform used to store a series of related information (often called configuration items) about hardware and software assets and the relationships between them.
[0065] As edge computing nodes, cloud box fault repair involves a relatively complex process. With users' increasing demand for cloud box repair capabilities and the operational challenges of hardware failures actually occurring in cloud box computing nodes in production environments, this application provides a repair solution for edge computing nodes. It solves the problems in the entire process of edge computing node repair from end to end, provides automated repair methods for edge computing nodes that experience hardware failures in production environments, and builds edge computing node repair capabilities from scratch.
[0066] Figure 1 This is a flowchart illustrating a maintenance method for an edge cloud computing node provided in an embodiment of this application. The method of this embodiment is applied to an operation and maintenance platform for edge cloud computing nodes, such as the Cloudbot platform, and will be described using the Cloudbot platform as an example in subsequent embodiments. Figure 1 As shown, the method includes:
[0067] S101. In response to edge cloud computing node failure and maintenance events, generate a workflow for taking the faulty equipment offline and a workflow for deploying new equipment.
[0068] In this embodiment, the edge cloud computing node refers to an edge cloud computing node located in the user's data center. The faulty device decommissioning workflow includes a series of automated tasks arranged in a specific order. The completion of the faulty device decommissioning workflow indicates that the faulty device has been decommissioned from the user's data center and subsequent maintenance has been completed. The new device deployment workflow also includes a series of automated tasks arranged in a specific order. The completion of the new device deployment workflow indicates that the new device has been deployed in the user's data center and has replaced the faulty device to perform the functions of an edge cloud computing node.
[0069] Typically, cloud computing systems have a fault monitoring platform to monitor the operation of each cloud computing node. When the fault monitoring platform detects a hardware failure in an edge cloud computing node, it can trigger an edge cloud computing node fault repair event. For example, the fault monitoring platform can generate an edge cloud computing node fault repair event through the ECS event system. In response to the edge cloud computing node fault repair event, the cloudbot platform generates a faulty device offline workflow and a new device deployment workflow to automatically complete the edge cloud computing node repair through these two workflows.
[0070] S102. The target standby machine in the standby pool is identified as the new equipment to be deployed through the new equipment deployment workflow.
[0071] The new device deployment workflow first initiates a standby machine scheduling task, selecting a target standby machine from the cloud box standby machine pool that is the same or similar in model as the faulty device as the new device to be deployed. Then, in the subsequent tasks of the new device deployment workflow, the target standby machine is installed and deployed to replace the faulty device.
[0072] S103. Modify the status of the faulty equipment to non-production status through the faulty equipment offline workflow.
[0073] Before taking a faulty device offline, its status is modified from production to non-production. For example, the relevant information for the faulty device is stored in a CMDB (Configuration Management Database). The Cloudbot platform uses a faulty device offline workflow to change the device's status in the CMDB to non-production.
[0074] S104. Install the target standby machine through the new equipment deployment workflow, add the target standby machine to the cloud computing node cluster where the faulty equipment is located, and mark the status of the target standby machine as non-production status.
[0075] After selecting a target standby machine as the new device to be deployed, it needs to be installed and configured accordingly, including installing the operating system, service components, and performing network-related configurations. It also needs to be added to the cloud computing node cluster where the faulty device resides, so that the target standby machine can replace the faulty device as an edge cloud computing node in the cluster. Furthermore, before the target standby machine is deployed, its status is still marked as non-production. For example, the Cloudbot platform triggers the installation platform to install and configure the target standby machine through a new device deployment workflow, and then the Cloudbot platform marks the target standby machine's status in the CMDB as non-production through the new device deployment workflow again. It should be noted that the status of the faulty device or the target standby machine in this embodiment can be referred to as the application status.
[0076] S105. After removing the faulty equipment from the user's computer room and installing the target standby equipment in the user's computer room, the status of the target standby equipment is changed to production status through the new equipment deployment workflow to end the new equipment deployment workflow.
[0077] Maintenance of edge computing nodes includes removing faulty equipment from the user's data center and installing new equipment in the user's data center. After the target standby machine is installed, its status is changed from non-production to production to complete the deployment of the new equipment.
[0078] S106. After repairing the faulty equipment, test the repaired faulty equipment through the faulty equipment offline workflow, and add the faulty equipment that passes the test to the standby pool to end the faulty equipment offline workflow.
[0079] After a faulty device is removed from the user's data center, it needs to be repaired. After the repair, the cloudbot platform triggers the installation platform to test the repaired faulty device through the faulty device offline workflow. If the test is passed, it is added to the standby pool so that the device can be used as a standby device in the future. At this point, the faulty device offline process is completed.
[0080] This application provides an edge computing node maintenance method, which offers a fully automated maintenance solution for edge computing nodes experiencing hardware failures in the production environment. It enables the removal of faulty devices from the network and the replacement of faulty devices with new ones, thus ensuring the service capabilities of edge computing nodes.
[0081] The workflow in the above embodiments will be further described below. Figure 2 This is a flowchart illustrating another edge computing node maintenance method provided in an embodiment of this application. Figure 2 As shown, the method includes:
[0082] S201. In response to edge cloud computing node failure and maintenance events, generate a workflow for taking the faulty device offline and a workflow for deploying the new device.
[0083] This step is similar to S101 in the above embodiment, and will not be described again here.
[0084] Optionally, edge computing node failure maintenance events are triggered after a user migrates or releases services on the faulty device. Users can migrate services on the faulty device to a backup device or other edge computing nodes with available inventory. If migration is not possible, services can be released to facilitate the subsequent decommissioning of the faulty device.
[0085] S202. The target standby machine in the standby pool is identified as the new equipment to be deployed through the new equipment deployment workflow.
[0086] This step is similar to S102 in the above embodiment, and will not be described again here.
[0087] S203. Modify the status of the faulty equipment to non-production status through the faulty equipment offline workflow, and disable monitoring and alarms for the faulty equipment.
[0088] The cloudbot platform changes the status of faulty devices in the CMDB to non-production status through the faulty device offline workflow, and the cloudbot platform no longer monitors or issues alarms for the faulty devices. At this point, the user services and cloud services on the faulty devices have been completely cleaned up.
[0089] S204. Generate a faulty equipment logistics order through the faulty equipment offline workflow. The faulty equipment logistics order is used to indicate the transportation of the faulty equipment from the user's computer room to the repair location.
[0090] To improve efficiency and reduce labor costs, the logistics order for faulty equipment can be linked to the logistics order for new equipment. The logistics of faulty equipment can begin after the new equipment arrives, and the removal of faulty equipment and the installation of new equipment can be completed in one trip to the user's computer room.
[0091] S205. Modify the logical configuration information of the target standby machine to the logical configuration information of the faulty machine through the new equipment deployment workflow; allocate switch ports and configure production routes for the target standby machine; install the operating system and service components for the target standby machine; configure the network for the target standby machine.
[0092] Logical configuration information can include network resource information, such as IP addresses and network segments, as well as information such as geographic location and data center. After modifying the logical configuration information of the target standby machine, the target standby machine will have the logical configuration information of the faulty device, but its physical location will not yet be in the user's data center. Network configuration for the target standby machine includes deploying a security gateway for it and modifying the backend service addresses that service components depend on.
[0093] S206. Through the faulty equipment offline workflow, using the same computing service component version as the cloud computing node cluster, add the target standby machine to the cloud computing node cluster in an expansion manner; lock the target standby machine and mark the target standby machine's status as non-production.
[0094] Optionally, the failure device offline workflow triggers the cluster management platform, which then adds the target standby machine to the cloud computing node cluster in an expansion manner, using the same computing service component version as the cloud computing node cluster.
[0095] S207. Add a preset tag to the target standby machine through the faulty equipment offline workflow. The preset tag is used to indicate that the target standby machine is the user's edge cloud computing node. Scan the target standby machine to determine whether the preset information exists. If it exists, delete the preset information. Turn off the monitoring and alarms for the target standby machine and shut down the target standby machine.
[0096] A pre-defined tag is added to the target standby machine to distinguish it from other types of cloud computing nodes, thereby preventing any impact on the edge cloud computing node when maintenance personnel operate or maintain other types of cloud computing nodes. Optionally, the pre-defined information can be sensitive information that does not meet security compliance requirements. For example, sensitive information may include keys or other information that may exist during the installation process described above.
[0097] S208. Generate a new equipment logistics order through the new equipment deployment workflow. The new equipment logistics order is used to indicate the delivery of the target backup machine to the user's data center. The departure date of the faulty equipment logistics order is the arrival date of the new equipment logistics order.
[0098] S209. After removing the faulty equipment from the user's computer room and installing the target standby equipment in the user's computer room, change the network mode of the target standby equipment to the field mode through the new equipment deployment workflow; perform full lifecycle management, network and operation and maintenance collaborative verification on the target standby equipment and the central cloud; enable monitoring and alarms for the target standby equipment, unlock the target standby equipment, and change the status of the target standby equipment to the production status to end the new equipment deployment workflow.
[0099] After the target backup machine arrives at the user's data center, the faulty device is first taken off the rack, and then the target backup machine is put back on the rack. After that, the network of the target backup machine is modified to field mode to redirect the uplink and downlink traffic, so that the data packets of the target backup machine proxies the security gateway client to enter the central cloud control area through the tunnel.
[0100] S210. After repairing the faulty equipment, test the repaired faulty equipment through the faulty equipment offline workflow, and add the faulty equipment that passes the test to the standby pool to end the faulty equipment offline workflow.
[0101] The method in this application provides an end-to-end maintenance method for edge computing nodes, which solves problems such as hardware installation, software deployment, network setup, effect verification, security compliance, logistics relocation, and hardware repair in the entire maintenance process of edge computing nodes, and provides an automated maintenance means for edge computing nodes that have hardware failures in the production environment.
[0102] Combination Figure 3 The method described in the embodiments of this application will be explained.
[0103] 1. When the fault monitoring platform determines that the cloud box (i.e., the edge cloud computing node) has a hardware fault, it generates a hardware fault repair work order and notifies the cloud box operation and maintenance team of the hardware fault details via telephone alarm.
[0104] 2. The fault monitoring platform triggers the ECS event system to send a cloud box fault repair event notification to the user terminal. At the same time, it provides the user with a selectable time window to schedule the date for the cloud box maintenance team to come to the site to replace the faulty cloud box (i.e. the faulty device). Subsequently, the ECS event system enters a state of waiting and periodically polling for fault repair events.
[0105] 3. The cloud box maintenance team communicates with users about the cloud box malfunction and repair incident, its background, and details, assisting users in resolving business issues on the faulty cloud box, for example:
[0106] a. If the user has a backup cloud box, then migrate the services on the faulty cloud box to the backup cloud box.
[0107] b. If the user does not have a backup cloud box, but has inventory on other health cloud boxes, then the services on the faulty cloud box will be migrated to other health cloud boxes.
[0108] c. If the customer does not have a backup cloud box and there is no inventory on other healthy cloud boxes, the services on the faulty cloud box will be released.
[0109] 4. After completing the migration or release of services on the faulty cloud box, the user should confirm the cloud box fault repair event in the ECS event system console, and select a suitable date within the appointment time window provided by the ECS event system for the cloud box operation and maintenance team to replace the faulty cloud box on-site.
[0110] 5. Once the ECS event system detects that a user has confirmed a cloud box failure repair event and selected a date for the cloud box maintenance team to replace the faulty cloud box, it triggers the cloudbot platform to generate a new device deployment workflow and a faulty device offline workflow.
[0111] 6. The new device deployment workflow of the cloudbot platform initiates standby scheduling, and selects a target standby machine from the cloud box standby machine pool that is the same or similar to the faulty cloud box as the new device to be deployed.
[0112] 7. The Cloudbot platform's faulty device offline workflow changes the application status of the faulty cloud box in the CMDB from production status (working_online) to non-production status (working_offline), and disables the Cloudbot platform's monitoring and alarms for the faulty cloud box. At this point, user services and cloud infrastructure services on the faulty cloud box have been cleaned up, leaving only a faulty physical server, which can be called a faulty machine.
[0113] 8. The Cloudbot platform initiates a faulty equipment offline workflow to move the faulty machine from the user's data center to the cloud factory (i.e., the repair location).
[0114] 9. The new device deployment workflow on the cloudbot platform initiates the logical migration of the new device to modify the logical configuration information of the target standby machine, such as network resources, region, and data center, in the CMDB to the logical configuration information of the faulty cloud box. At this time, the physical location of the target standby machine is still in the cloud factory.
[0115] 10. The new device deployment workflow on the cloudbot platform assigns IP addresses and switch ports to the target standby machine and connects the target standby machine to CEN.
[0116] 11. The new device deployment workflow of the cloudbot platform installs the operating system on the target standby machine through the installation platform, and performs a series of health checks on the installed target standby machine, such as system uptime and disk mount.
[0117] 12. The new device deployment workflow of the cloudbot platform installs basic service components such as VPC, load balancing, Domain Name System (DNS) resolution, Shell front-end package manager (Yum), log collection, monitoring, Tianji, staragent, and OOB on the target standby machine through the installation platform.
[0118] 13. The new device deployment workflow on the cloudbot platform deploys a security gateway for the target standby machine.
[0119] 14. The new device deployment workflow on the cloudbot platform changes the backend service address on which the basic service components installed in step 12 depend from the cloud factory intranet virtual IP (VIP) to the user-side public VIP.
[0120] 15. The new device deployment workflow of the cloudbot platform triggers the cluster management platform. The cluster management platform uses the same firmware version of the computing service components as the original cloud box cluster where the faulty cloud box was located to expand the capacity and deploy the target standby machine to the cloud box cluster. After the deployment is completed, the target standby machine is locked and its application status is marked as non-production status.
[0121] 16. The new device deployment workflow on the cloudbot platform assigns a pre-defined tag to the target standby machine in the CMDB, such as the cloud_box.customer_idc tag specifically for cloud boxes.
[0122] 17. The new device deployment workflow of the cloudbot platform scans for various sensitive information such as access keys (AK), secret access keys (SK), and keys on the target standby machine in accordance with security compliance requirements, and cleans up the sensitive information detected.
[0123] 19. The new device deployment workflow of the cloudbot platform gracefully shuts down the deployed target standby machine, that is, it turns off the monitoring and alarms of the target standby machine and then performs the shutdown.
[0124] 20. Disconnect the target backup machine from the network and power supply, and disconnect the cable.
[0125] 21. The new device deployment workflow on the cloudbot platform initiates a new device logistics order, specifying that the target standby machine will be moved from the cloud factory to the user's data center.
[0126] 22. The logistics provider transports the target backup machine from the cloud factory to the user's data center.
[0127] 23. The CloudBox maintenance team will travel to the user's data center on the date selected by the user for on-site replacement of the faulty equipment. Upon arrival, they will first remove the faulty equipment, then install the new equipment, and connect the new equipment to the power and network. It should be noted that if the user-selected date for on-site replacement of the faulty equipment coincides with the arrival date of the new equipment's logistics order, the arrival date of the new equipment's logistics order is the departure date of the faulty equipment's logistics order. If the user-selected date for on-site replacement of the faulty equipment falls on a date after the arrival date of the new equipment's logistics order, the departure date of the faulty equipment's logistics order is the user-selected date for on-site replacement of the faulty equipment.
[0128] 24. The new device deployment workflow of the cloudbot platform redirects the uplink and downlink of the target standby machine. That is, it changes the network of the target standby machine from the factory mode (cloud factory) to the cloud box mode (user site). The cloud box mode can be called the site mode, so that the security gateway client proxies the cloud box data packets to enter the central cloud control area through the tunnel.
[0129] 25. The new device deployment workflow of the cloudbot platform triggers the fault monitoring platform to initiate cloud-edge collaborative verification of the cloud box in operation state, including full lifecycle management, network and operation and maintenance collaborative verification, in order to accept the target standby machine.
[0130] 26. The new device deployment workflow on the cloudbot platform enables monitoring and alarms for the target standby machine, unlocks the target standby machine, and modifies its application status to production status in the CMDB. At this point, the new device deployment workflow on the cloudbot platform is completed, realizing the replacement of the faulty device by the target standby machine.
[0131] 27. Once the ECS event system detects that the new device deployment workflow on the cloudbot platform has been completed, it sends a cloud box fault repair event completion notification to the user.
[0132] 28. The logistics provider transports the faulty machine from the user's data center to the cloud factory.
[0133] 29. Cloudbot Platform Fault Equipment Offline Workflow: After the faulty machine arrives at the cloud factory, confirm the cloud box hardware fault repair work order in step 1.
[0134] 30. The cloudbot platform's faulty equipment offline workflow is triggered to repair faulty machines according to the hardware fault repair process of the central cloud computing node. After the faulty machines are repaired, they become machines waiting to be put into the warehouse.
[0135] 31. The cloudbot platform's faulty device offline workflow triggers the installation platform to perform necessary stress tests and hardware checks on the machines to be put into storage. After the checks are passed, the machines to be put into storage are transferred to the cloud box standby pool. At this point, the cloudbot platform's faulty device offline workflow is completed.
[0136] The edge cloud computing node maintenance method of this application embodiment covers the entire process from the fault monitoring platform discovering the hardware failure of the edge cloud computing node to completing the user-side hardware replacement. It includes notifying the user, processing the service migration on the faulty cloud box, installing and deploying the new equipment, network configuration, scanning and cleaning sensitive information, logistics relocation, hardware replacement, and effect verification. At the same time, it includes the process of taking the faulty machine replaced from the user's data center offline, logistics relocation, hardware repair, stress testing and hardware inspection, and putting it into the backup machine pool, providing an end-to-end edge cloud computing node maintenance solution.
[0137] Figure 4 This is a schematic block diagram of the operation and maintenance platform for edge cloud computing nodes provided in this application embodiment. Figure 4 As shown, the operation and maintenance platform 400 may include at least one processor 401 for implementing the edge cloud computing node maintenance method provided in the embodiments of this application.
[0138] Optionally, the operation and maintenance platform 400 further includes at least one memory 402 for storing program instructions and / or data. The memory 402 is coupled to the processor 401. The coupling in this embodiment is an indirect coupling or communication connection between devices, units, or modules, which can be electrical, mechanical, or other forms, for information exchange between devices, units, or modules. The processor 401 may operate in conjunction with the memory 402. The processor 401 may execute program instructions stored in the memory 402. At least one of the at least one memory may be included in the processor.
[0139] Optionally, the operation and maintenance platform 400 further includes a communication interface 403 for communicating with other devices via a transmission medium, thereby enabling the operation and maintenance platform 400 to communicate with other devices. The communication interface 403 may be, for example, a transceiver, interface, bus, circuit, or a device capable of transmitting and receiving functions. The processor 401 can utilize the communication interface 403 to transmit and receive data and / or information, and to implement the methods provided in the embodiments of this application. For details, please refer to the detailed descriptions in the preceding embodiments; further elaboration is not repeated here.
[0140] This application embodiment does not limit the specific connection medium between the processor 401, memory 402, and communication interface 403. This application embodiment... Figure 4 The processor 401, memory 402, and communication interface 403 are connected via bus 404. Bus 404 is... Figure 4 The connections between other components are shown in thick lines only and are not intended to be limiting. This bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, Figure 4 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0141] It should be understood that the processor in the embodiments of this application can be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method embodiments can be completed by the integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.
[0142] It should also be understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0143] This application also provides a cloud computing system, including an edge cloud computing node and an operation and maintenance platform for the edge cloud computing node as described in the foregoing embodiments.
[0144] This application also provides a computer-readable storage medium storing a computer program (also referred to as code or instructions). When the computer program is run, it causes the computer to perform the methods as described in any of the foregoing embodiments.
[0145] The terms “unit”, “module”, etc., used in this specification may be used to refer to computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution.
[0146] Those skilled in the art will recognize that the various illustrative logical blocks and steps described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application. In the several embodiments provided in this application, it should be understood that the disclosed apparatus, devices, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for example, the division of units is merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the shown or discussed mutual couplings or direct couplings or communication connections may be through some interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.
[0147] The unit described as a separate component may or may not be physically separate. The component shown as a unit may or may not be a physical unit; that is, it may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0148] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0149] In the above embodiments, the functions of each functional unit can be implemented entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. This computer program product includes one or more computer instructions (programs). When the computer program instructions (programs) are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be magnetic media (e.g., floppy disk, hard disk, magnetic tape), optical media (e.g., digital video disc (DVD)), or semiconductor media (e.g., solid-state disk (SSD)).
[0150] If this function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0151] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for maintaining an edge computing node, characterized in that, include: In response to edge cloud computing node failure and maintenance events, generate workflows for taking faulty equipment offline and deploying new equipment. The new equipment deployment workflow identifies target standby machines in the standby pool as new equipment to be deployed. The status of the faulty equipment is changed to a non-production status through the faulty equipment offline workflow. The target standby machine is installed using the new equipment deployment workflow, and the target standby machine is added to the cloud computing node cluster where the faulty equipment is located, and the status of the target standby machine is marked as non-production status. After removing the faulty device from the user's computer room and installing the target backup device in the user's computer room, the status of the target backup device is changed to production status through the new device deployment workflow to end the new device deployment workflow; After the faulty equipment is repaired, the repaired faulty equipment is tested through the faulty equipment offline workflow, and the faulty equipment that passes the test is added to the standby pool to end the faulty equipment offline workflow.
2. The method according to claim 1, characterized in that, After changing the status of the faulty equipment to a non-production status, the method further includes: The faulty equipment offline workflow generates a faulty equipment logistics order, which is used to instruct the faulty equipment to be transported from the user's computer room to the repair location. After marking the target standby machine as non-production, the method further includes: A new equipment logistics order is generated through the new equipment deployment workflow. The new equipment logistics order is used to instruct the delivery of the target backup equipment to the user's data center. The departure date of the faulty equipment logistics order is the arrival date of the new equipment logistics order.
3. The method according to claim 2, characterized in that, After changing the status of the faulty equipment to a non-production status, the method further includes: Disable monitoring and alarms for the faulty device.
4. The method according to claim 1, characterized in that, The step of installing the target standby machine using the new equipment deployment workflow and adding the target standby machine to the cloud computing node cluster where the faulty device resides includes: The logical configuration information of the target standby device is modified to the logical configuration information of the faulty device through the new device deployment workflow; Perform switch port allocation and production route configuration on the target standby machine; The target standby machine is then installed with an operating system and service components. Configure the network for the target standby machine; The target standby machine is added to the cloud computing node cluster in an expansion manner, using the same computing service component version as the cloud computing node cluster.
5. The method according to claim 4, characterized in that, After marking the target standby machine as non-production, the method further includes: Add a preset tag to the target standby machine, the preset tag being used to indicate that the target standby machine is the user's edge cloud computing node.
6. The method according to claim 4, characterized in that, After marking the target standby machine as non-production, the method further includes: Scan the target backup device to determine if preset information exists; if it does, delete the preset information.
7. The method according to claim 4, characterized in that, Before marking the target standby machine as non-production, the method further includes: Lock the target backup device; After marking the target standby machine as non-production, the method further includes: The monitoring and alarms for the target backup machine are turned off, and the target backup machine is powered off.
8. The method according to claim 1, characterized in that, After the target standby machine is racked in the user's computer room, the method further includes: Change the network mode of the target standby machine to field mode.
9. The method according to claim 8, characterized in that, After modifying the network mode of the target standby machine to the field mode, the method further includes: The target standby machine and central cloud are subject to full lifecycle management, network and operation and maintenance collaborative verification.
10. The method according to claim 7, characterized in that, Before modifying the status of the target standby machine to production status through the new equipment deployment workflow, the method further includes: Enable monitoring and alarms for the target backup machine, and unlock the target backup machine.
11. The method according to any one of claims 1-10, characterized in that, The edge computing node failure repair event is confirmed and triggered after the user migrates or releases the services on the faulty device.
12. An operation and maintenance platform for edge cloud computing nodes, characterized in that, include: Memory and processor; The memory is used to store computer programs; The processor is configured to execute a computer program stored in the memory, wherein when the computer program is executed, the processor performs the method described in any one of claims 1-11.
13. A cloud computing system, characterized in that, This includes edge computing nodes and the operation and maintenance platform for edge computing nodes as described in claim 12.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to perform the method as described in any one of claims 1-11.
Citation Information
Patent Citations
Equipment full-life-cycle management system
CN113011845A
Model deployment method and device, computer equipment and storage medium
CN113805546A