Method and apparatus for deployment of applications on edge cluster
By setting redundant services and considering constraints such as cost, resources and security when the application is deployed to an edge cluster, the problem of inadequate application deployment and failure to fully consider constraints in the existing technology is solved, and the optimized deployment and reliability of the application are achieved.
Patent Information
- Application Number
- CN202311623164.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-30
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art fails to fully consider the reliability and various constraints of the application, such as cost, resource requirements and security, when deploying applications on edge clusters.
By setting up redundant service copies for the services in the application, a deployment plan that maximizes application reliability is planned, and constraints such as equipment cost, resources and security are fully considered.
It realizes optimized deployment of applications on edge clusters, improves application reliability, and reduces deployment costs and resource consumption while meeting various constraints.
Smart Images

Figure CN120066757A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of computers, and more particularly to methods and apparatuses for the deployment of applications on edge clusters. Background Art
[0002] At present, edge computing technology has been widely applied. Edge computing can place computing resources near the user or terminal device, so compared with cloud computing, it can reduce the latency and bandwidth consumption of data communication and provide real-time processing close to the data source. A group of edge computing devices can form an edge cluster, and applications can be deployed on the edge cluster, enabling the edge computing devices in the edge cluster to run the application to achieve specific functions. Thus, how to deploy applications on edge clusters has become an issue that needs to be focused on.
[0003] The current deployment mechanisms for applications on edge clusters have various problems. One problem that needs to be solved urgently is that these deployment mechanisms do not fully consider whether the deployed applications meet the reliability requirements. In addition, the deployment of applications on edge clusters may also be restricted by some constraint conditions (such as cost issues, resource demand issues, security issues). However, the current deployment mechanisms do not comprehensively consider these constraint conditions either.
[0004] Therefore, the present disclosure proposes a method for the deployment of applications on edge clusters, which aims to maximize the reliability of applications and fully and comprehensively considers various constraint conditions that may affect the deployment, so as to achieve the optimized deployment of applications on edge clusters. Summary of the Invention
[0005] It is desired to provide a method for the deployment of applications on edge clusters, which can set corresponding redundant service replicas for the services included in the application, so as to plan a deployment plan that maximizes the reliability of the application, and then can implement the deployment of the application on the edge cluster based on the planned deployment plan. The proposed method also fully and comprehensively considers a series of constraint conditions that may affect the deployment of the application, including device cost constraints, resource constraints, security constraints, and so on. Thus, through this improved method proposed by the present disclosure, it is beneficial to achieve the optimized deployment of applications on edge clusters.
[0006] According to one aspect of the present disclosure, there is provided a method for planning the deployment of an application on an edge cluster, including: determining a plurality of main services included in the application; determining a plurality of candidate device types of edge computing devices for constituting the edge cluster; and determining a deployment planning scheme based on the plurality of main services and the plurality of candidate device types; wherein the deployment planning scheme indicates the number of redundant services to be set for each of the plurality of main services and the number of edge computing devices to be set for each of the plurality of candidate device types.
[0007] According to one aspect of the present disclosure, there is provided a method for deploying an application on an edge cluster, including: obtaining a deployment planning scheme determined according to the above deployment planning method; and controlling the deployment of the application on the edge cluster according to the deployment planning scheme.
[0008] According to yet another aspect of the present disclosure, there is provided a device for planning the deployment of an application on an edge cluster, including: a memory and a processor. The processor is coupled to the memory and is configured to execute the method according to any one of the various embodiments of the present disclosure for deployment planning.
[0009] According to yet another aspect of the present disclosure, there is provided a device for deploying an application on an edge cluster, including: a memory and a processor. The processor is coupled to the memory and is configured to execute the method according to any one of the various embodiments of the present disclosure for deploying an application.
[0010] According to still another aspect of the present disclosure, there is provided a computer-readable medium storing a computer program including instructions, which when executed by a processor cause the processor to be configured to execute the method according to any one of the various embodiments of the present disclosure for deployment planning or for deploying an application. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Various embodiments of the claimed subject matter will now be described by way of example with reference to the drawings. In the different drawings, the same reference numerals are used to denote the same or similar components.
[0012] Figure 1 A schematic diagram showing the overall architecture of edge computing according to an example embodiment of the present disclosure is shown.
[0013] Figure 2 A schematic diagram showing the architecture of an edge cluster according to an example embodiment of the present disclosure is shown.
[0014] Figure 3A schematic diagram of a deployment applied to an edge cluster without setting up redundant services according to an exemplary embodiment of the present disclosure is shown.
[0015] Figure 4 A schematic diagram of an optimized deployment applied to an edge cluster with redundant services set up according to an exemplary embodiment of the present disclosure is shown.
[0016] Figure 5 A flowchart of a method for planning a deployment applied to an edge cluster according to an exemplary embodiment of the present disclosure is shown.
[0017] Figure 6 A flowchart of a method for deploying an application on an edge cluster according to an exemplary embodiment of the present disclosure is shown.
[0018] Figure 7 A block diagram of a device according to an exemplary embodiment of the present disclosure is shown. Detailed implementation manners
[0019] In the following description, many specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure. However, those skilled in the relevant art will recognize that the present disclosure may be practiced without one or more of the specific details, or may be practiced using alternative methods, components, etc. In some instances, well-known structures and operations are not shown or described in detail so as not to unnecessarily obscure the present disclosure.
[0020] As discussed above, edge computing can place computing resources closer to the user or the terminal device, so compared with cloud computing, it can reduce the latency and bandwidth consumption of data communication and provide real-time processing closer to the data source. The following will be combined with Figure 1 Briefly discuss an example overall architecture of edge computing.
[0021] Figure 1 A schematic diagram of the overall architecture of edge computing according to an exemplary embodiment of the present disclosure is shown. Figure 1 The shown example overall architecture of edge computing may include terminal devices 102 (including terminal devices 102-1 to 102-M), edge computing devices 104 (including edge computing devices 104-1 to 104-N), and cloud devices 108 arranged hierarchically. Among them, the edge computing devices 104 may be located at an intermediate level between the terminal devices 102 and the cloud devices 108. Each edge cluster may be composed of a group of edge computing devices and be deployed to run a specific application. According to Figure 1In the example architecture shown, multiple edge computing devices 104-1 to 104-N can be divided into different edge clusters 106-1 to 106-G. For example, edge cluster 106-1 can include edge computing devices 104-1, 104-2... 104-J, and so on. Edge cluster 106-G can include edge computing devices 104-K to 104-N. Such an arrangement is particularly beneficial for compute-intensive applications because compute-intensive applications (such as the roadside perception applications discussed below) often require high computing capabilities that cannot be provided by a single edge computing device. In this case, by combining multiple edge computing devices into an edge cluster and through scheduling or management among the multiple edge computing devices in the edge cluster, the requirements of the application for high computing capabilities can be met. Edge cluster 106 can interact with terminal device 102 (through unidirectional or bidirectional data communication). For example, it can receive sensor information from terminal device 102, and so on. At the same time, edge cluster 106 can also interact with cloud device 108 to implement functions such as cloud computing or cloud storage when necessary.
[0022] For ease of understanding, the following will further discuss the content related to Figure 1 the overall architecture of edge computing shown in conjunction with an example application scenario. In an example aspect, edge computing technology can be applied in a vehicle-to-everything (V2X) scenario. In this example scenario, the edge cluster can be a roadside edge cluster, and the edge cluster (for example, edge cluster 106-1) can be deployed with a roadside perception application to enable perception of the environment around the vehicle and transmit the perception information to the in-vehicle processing unit to make corresponding control decisions (such as making the vehicle decelerate, brake, steer, etc.). In this example scenario, terminal device 102 can include: a roadside image sensor (such as a traffic camera installed at a traffic signal) 102-1, a roadside radar 102-2, and an in-vehicle communication module 102-3, and so on. Edge cluster 106-1 can receive information collected by the sensors from roadside image sensor 102-1 and roadside radar 102-2 that is associated with objects detected on the road (such as vehicles, pedestrians, etc.), and generate prediction data indicating the trajectory of the target object based at least on such information. Then, edge cluster 106-1 can further send the generated prediction data to in-vehicle communication module 102-3 so that the in-vehicle processing unit can make corresponding control decisions based on the prediction data. Edge cluster 106-1 can also transmit the generated prediction data to cloud device 108 for cloud storage or for further processing.
[0023] In the present disclosure, an application deployed on an edge cluster may adopt a service-based architecture. With this service-based architecture, an application can be split / decomposed into a series of services (or similarly called microservices), or a collection of such a series of services can constitute the application. Multiple services can run in a pipeline form such that the output of a previous service can be used as the input of a subsequent service. Alternatively, multiple services can also run in parallel such that some services can perform different operations based on the same input. In addition, there is generally low correlation between individual services, enabling each service to be deployed, managed, and tested independently. Thus, this service-based architecture can provide higher flexibility for the deployment of applications (especially some complex applications). For example, some services can be conveniently reorganized and replaced to enable the application to achieve different functions. In addition, splitting an application into a series of services also helps to improve the reliability of the application by setting redundancy for some services, as will be discussed in detail below.
[0024] Splitting an application into a series of services can be achieved during the design and development stage of the application. During the design and development stage, several program functions can be packaged together to form a container, and the formed containers can respectively correspond to services. For the example application of roadside perception discussed above, the roadside perception application can be divided into: a detection service that can detect target objects from image data from a roadside image sensor; a projection service that can project the two-dimensional coordinates of the target object extracted from the image data into three-dimensional coordinates; a tracking service that can predict the future movement trajectory of the target object based on the identification information, movement information, and historical movement trajectory of the target object; a fusion service that can achieve sensor fusion between image data from a roadside image sensor and radar data from a roadside radar; a recording service that can achieve real-time recording and storage of perception data; and a visualization service that can convert the perception data into an intuitive visual representation; and so on. It should be noted that the above-listed services are only examples, and which services are included in an application deployed on an edge cluster can be specifically set according to the functional requirements for the application.
[0025] The applicant has noticed that services are not absolutely reliable during operation but may encounter various failures. Since each service included in an application is an essential part for the normal operation of the application, a failure of any one service may cause the entire application to fail, thus affecting the reliability of the application. Especially in fields such as the Internet of Vehicles, how to improve the reliability of applications is an important challenge.
[0026] To improve the reliability of service-based applications, backups or replicas can be set up for one or more of the services into which the application is split. For clarity, in this disclosure, the services into which the application is split are referred to as primary services, while the backups or replicas set up for the primary services are referred to as redundant services. Redundant services are created by replicating the corresponding primary services. Redundant services can perform the same operations as the primary services and provide the same functions as the primary services. In the event of a failure of the primary service, the redundant services can be promptly activated to take over from the primary service to complete its functions.
[0027] Figure 2 FIG. shows a schematic diagram of the architecture of an edge cluster according to an exemplary embodiment of the present disclosure.
[0028] Figure 2 illustrates the architecture of the edge cluster using the Kubernetes architecture as an example. The edge cluster can be any of the edge clusters 106 discussed in conjunction with Figure 1 The edge cluster can include a primary edge computing device 202 and worker edge computing devices 206. The primary edge computing device and the worker edge computing devices are divided according to the roles of the edge computing devices in the edge cluster.
[0029] The primary edge computing device 202 can also be referred to as the Master Node. As the control plane of the entire edge cluster, it is responsible for scheduling and managing the worker edge computing devices 206. The primary edge computing device 202 can include a controller 204. The controller 204 is the core control component of the primary edge computing device 202 and can be used to implement the control of the deployment of services on the worker edge computing devices 206 as discussed in this disclosure. In one example, the controller 204 can be implemented as a Replication Controller and / or a Deployment controller in the Kubernetes architecture, etc. The primary edge computing device 202 can also include Figure 2 other components not shown in, for example, a repository (such as etcd) for storing the configuration data and status information of the edge cluster, an interface component (such as an API server) for maintaining the status information of the edge cluster, a scheduler for scheduling the running of services on the worker edge computing nodes 206, and so on.
[0030] In one example, the method for deploying an application on an edge cluster discussed in this disclosure can be executed by Figure 2 the primary edge computing device 202 shown in. More specifically, the method for deploying an application on an edge cluster discussed in this disclosure can be executed by, for example, the controller 204 in the primary edge computing device 202.
[0031] The working edge computing device 206 can also be referred to as a slave node or a worker node, which is used to run services under the management and scheduling of the master edge computing device 202. Each working edge computing device 206 may include one or more Pods 208. A Pod 208 is the smallest deployable unit in the Kubernetes architecture. Each Pod 208 may include one or more containers 210, and each container can be used to run a program image (e.g., a Docker Image). The working edge computing device 206 may also include Figure 2 other components not shown in the figure, such as a resource management component (e.g., Kubelet) for managing the running of Pods, a container runtime component (e.g., ContainerRuntime) for managing the running of containers, a service management component (e.g., Kube-proxy) for managing the access entry of services, and so on.
[0032] Deploying an application on an edge cluster as discussed in this disclosure may include deploying the services included in the application on each working edge computing device under the control of the master edge computing device. For example, through the management and scheduling of the master edge computing device, the corresponding working edge computing device can deploy the service on the working edge computing device by starting a container using a Docker Image.
[0033] Although Figure 2 the architecture of the edge cluster is shown by taking Kubernetes as an example, however, the edge cluster can adopt other architectures, such as the Mesos architecture, the Swarm architecture, etc., without departing from the scope of this disclosure.
[0034] In addition, Figure 2 the number of each component shown in the figure is only exemplary and does not limit the scope of this disclosure.
[0035] Next, Figure 3 it is discussed how to deploy an application based on services on an edge cluster without setting up redundant services.
[0036] Figure 3 shows a schematic diagram of the deployment of an application on an edge cluster without setting up redundant services according to an exemplary embodiment of this disclosure. As Figure 3As shown, the application 302 may include a plurality of services 304-1, 304-2, 304-3 arranged in series. The application 302 is to be deployed on the edge cluster 306, and the edge cluster 306 includes two edge computing devices 308-1 and 308-2. Without utilizing the optimized deployment of the present disclosure, deploying the application 302 on the edge cluster 306 may include randomly deploying the plurality of services 304-1, 304-2, 304-3 on the plurality of edge computing devices 308-1 and 308-2 of the edge cluster 306. For example, services 304-1 and 304-2 are made to run on the edge computing device 308-1, while service 304-3 runs on the edge computing device 308-2. Thus, service 304-1 may receive input data 310 and generate first service data, service 304-2 may receive the first service data and generate second service data, and service 304-3 may receive the second service data and generate output data 312, thereby implementing the predetermined function of the application 302.
[0037] In this case, as discussed above, since no redundant services are set, the failure of any one of the services 304-1, 304-2, 304-3 may cause the application 302 to fail to implement its normal function (for example, the failure of service 304-2 results in the inability to generate second service data, and thus, service 304-3 cannot generate output data 312 based on the second service data), resulting in the application 302 being unable to provide the required reliability. At the same time, the deployment of these services 304-1, 304-2, 304-3 on the edge computing devices 308-1 and 308-2 does not comprehensively consider the limitations of some possible constraint conditions (such as cost issues, resource requirement issues, security issues), and therefore, its deployment method may not be optimized.
[0038] The following combines Figure 4 to discuss the optimized deployment of a service-based application on an edge cluster in the case of setting redundant services according to the method of the present disclosure.
[0039] Figure 4 A schematic diagram of the optimized deployment of an application on an edge cluster in the case of setting redundant services according to an exemplary embodiment of the present disclosure is shown. In Figure 4 a service-based application 402 is shown, which can be deployed on an edge cluster 404. The application 402 can be divided into a series of services, including primary services S1, S2, S3... SN, and these primary services cooperate with each other to enable the application 302 to implement its predetermined function. To improve the reliability of the application 402, redundant services can be set for one or more of the primary services S1, S2, S3... SN. As Figure 4As shown, two redundant services R1-1 and R1-2 can be set for the primary service S1, no redundant service can be set for the primary service S2, one redundant service R3-1 can be set for the primary service S3, and so on, and one redundant service RN-1 can be set for the primary service SN.
[0040] The edge cluster 404 can include multiple edge computing devices 406-1, 406-2... 406-M. These edge computing devices can be of different device types. This type can be determined based on the model or configuration of the edge computing device. For example, an edge computing device with a first model or configuration can be called a first type of edge computing device, an edge computing device with a second model or configuration can be called a second type of edge computing device, and so on. As Figure 3 shown, the first type of edge computing device can include edge computing devices 406-1(A) and 406-1(B), the second type of edge computing device can include edge computing devices 406-2(A), 406-2(B) and 406-2(C), and so on, and the Mth type of edge computing device can include edge computing device 406-M. After the application 402 is deployed on the edge cluster 404 by using the optimization deployment mechanism of the present disclosure, any one of the primary service S1 and the redundant services R1-1 and R1-2 can receive the input data 408 and generate first service data, the primary service S2 can receive the first service data and generate second service data, any one of the primary service S3 and the redundant service R3-1 can receive the second service data and generate third service data, and so on, and any one of the primary service SN and the redundant service RN-1 can receive the service data generated by the previous type of service and generate the output data 410, so that the application 402 can achieve its predetermined function.
[0041] Through this redundancy setting of the service-based application 402, when the primary service (or one of its corresponding redundant services) fails, one of the corresponding redundant services (or another redundant service) of the primary service can be enabled to continue to complete the normal function of this type of service. Thus, the probability of the entire application failing due to a single service failure is reduced, and the reliability of the application is significantly improved.
[0042] In one example, the deployment of the services discussed in the present disclosure on the edge computing devices is not limited to deploying one service on one edge computing device, but also includes deploying one service on multiple edge computing devices, or deploying multiple services on one edge computing device, and so on. The specific deployment method can be based on, for example, the architecture of the edge cluster Figure 2 discussed, without limiting the scope of the present disclosure.
[0043] The optimization deployment mechanism of the application of the present disclosure on the edge cluster will be further discussed in detail below, including planning the deployment of the application on the edge cluster, and then deploying the application to the edge cluster according to the deployment plan determined through the planning.
[0044] First, as discussed above, it is possible to determine the main services included in the application to be deployed on the edge cluster. For example, as discussed above, it is possible to determine that the application is divided into N main services. After determining the N main services, information related to the attributes of each main service can be obtained, as discussed below, and deployment planning can be performed based on these service attributes to determine the deployment plan.
[0045] Then, it is possible to determine multiple candidate device types of the edge computing devices that make up the edge cluster. This can be understood as defining a pool of available device types, such that it is possible to select which types of edge computing devices to use from this pool of device types to form the edge cluster. After determining the candidate device types of the edge computing devices, information related to the attributes of the edge computing devices of each candidate device type can be obtained, as discussed below, and deployment planning can be performed based on these device attributes to determine the deployment plan.
[0046] In one example, a constraint optimization algorithm can be used to determine the deployment plan. The constraint optimization algorithm takes maximizing the reliability of the application as the optimization goal, while considering one or more constraints. The determined deployment plan can indicate the number of redundant services to be set for each main service and the number of edge computing devices to be set for each candidate device type.
[0047] The constraint optimization algorithm discussed in the present disclosure can take maximizing the reliability of the application as the optimization goal. The reliability of the application can be calculated based on the failure rate of each main service (since the redundant service is generated by replicating the main service, the redundant service has the same failure rate as the main service), and the number of redundant services set for each main service. Specifically, the optimization goal can be expressed as:
[0048]
[0049] where K is the number of main services included in the application, s i is the number of the i-th main service and its corresponding redundant services (or, the number of redundant services set for the i-th main service plus 1), F i represents the failure rate of the i-th main service and can be determined based on the service attributes mentioned above, and represents the s i power of F i times.
[0050] In addition, one or more constraint conditions of the constrained optimization algorithm can be set. Such constraint conditions contribute to the rapid convergence of the constrained optimization algorithm.
[0051] The first constraint condition can include a device cost constraint. It can be understood that when deploying an edge cluster, the budget for purchasing edge computing devices is usually limited. Thus, a predetermined device cost threshold can be set based on this budget, so that the deployment planning scheme given by the constrained optimization algorithm does not cause the total cost of the edge computing devices to exceed the predetermined device cost threshold. This cost constraint can be expressed as:
[0052]
[0053] where Budget is the set predetermined device cost threshold, J is the number of candidate device types of edge computing devices, n j is the number of the j-th type of edge computing devices, and DeviceP ′ cost ′ [j] is the cost of each j-th type of edge computing device and can be determined based on the device attributes mentioned above. Thus, through the above formula (2), the total cost of the edge computing devices constituting the edge cluster can be made not to exceed the predetermined device cost threshold.
[0054] The second constraint condition can include a resource constraint. The applicant notes that applications may require different computing resources. Some computationally intensive applications usually require the computing devices running the applications to have higher computing capabilities. For example, more central processing unit (CPU) cores, higher random access memory (RAM) capacity, and so on. If such computationally intensive applications are deployed on edge computing devices that do not have the corresponding computing capabilities, the applications may not run properly. Thus, taking the resource constraint as a constraint condition of the constrained optimization algorithm helps to achieve optimized deployment by deploying applications on edge computing devices with compatible computing capabilities.
[0055] In one exemplary aspect, the resource constraint can be embodied as the requirement of the application for the number of CPU cores of the edge computing devices in the edge cluster, as follows:
[0056]
[0057] where K is the number of main services included in the application, s i is the number of the i-th main service and its corresponding redundant services (or, the number of redundant services set for the i-th main service plus 1), J is the number of candidate device types of edge computing devices, n j is the number of the j-th type of edge computing devices, and ServiceP ′CPU′][i] is the number of CPU cores required for the i-th primary service (since redundant services are generated by replicating primary services, redundant services require the same number of CPU cores as primary services) and can be determined based on the service attributes mentioned above, and DeviceP ′ CPU′][j] is the number of CPU cores that each j-th type of edge computing device can provide and can be determined based on the device attributes mentioned above. Thus, through such a constraint condition, the sum of the number of CPU cores required for all primary services and redundant services of the application does not exceed the sum of the number of CPU cores that all edge computing devices can provide.
[0058] In another exemplary aspect, the resource constraint can be reflected in the application's requirement for the capacity of the video memory of the edge computing devices in the edge cluster. As those skilled in the art can understand, video memory refers to the memory inside a computing device for storing graphic data associated with various graphic processing processes of a graphics processing unit (GPU) (for example, rendering data to be displayed by a display, etc.). The resource constraint regarding the capacity of the video memory can be expressed as:
[0059]
[0060] where K is the number of primary services included in the application, s i is the number of the i-th primary service and its corresponding redundant services (or, the number of redundant services set for the i-th primary service plus 1), J is the number of candidate device types of edge computing devices, n j is the number of the j-th type of edge computing devices, ServiceP ′ GPU memory′][i] is the capacity of the video memory required for the i-th primary service (since redundant services are generated by replicating primary services, redundant services require the same capacity of the video memory as primary services) and can be determined based on the service attributes mentioned above, and DeviceP ′ GPUmemory′][j] is the capacity of the video memory that each j-th type of edge computing device can provide and can be determined based on the device attributes mentioned above. Thus, through such a constraint condition, the sum of the capacities of the video memory required for all primary services and redundant services of the application does not exceed the sum of the capacities of the video memory that all edge computing devices can provide.
[0061] In another example aspect, the resource constraint can be embodied as the requirement of the application for the capacity of the random access memory (RAM) of the edge computing devices in the edge cluster. As those skilled in the art can understand, RAM refers to the volatile memory inside the computing device that is used to temporarily store various instructions and data involved in the process of the CPU running programs, so as to achieve direct interaction with the CPU. The resource constraint regarding the capacity of the RAM can be expressed as:
[0062]
[0063] where K is the number of main services included in the application, s i is the number of the i-th main service and its corresponding redundant services (or, the number of redundant services set for the i-th main service plus 1), J is the number of candidate device types of the edge computing devices, n j is the number of the j-th type of edge computing devices, ServiceP ′ RAM′][i] is the capacity of the RAM required by the i-th main service (since the redundant service is generated by replicating the main service, the redundant service requires the same RAM capacity as the main service) and can be determined based on the service attributes mentioned above, and DeviceP ′ RAM′][j] is the capacity of the RAM that each j-th type of edge computing device can provide and can be determined based on the device attributes mentioned above. Thus, through such a constraint condition, the total capacity of the RAM required by all the main services and redundant services of the application can be made not to exceed the total capacity of the RAM that all the edge computing devices can provide.
[0064] In another example aspect, the resource constraint can be embodied as the requirement of the application for the capacity of the storage device of the edge computing devices in the edge cluster, where the storage device can refer to the non-volatile memory of the edge computing device, which can be embodied as a solid-state drive, a disk, an optical disc, and any other form of optical storage medium or magnetic storage medium, etc. The resource constraint regarding the capacity of the storage device can be expressed as:
[0065]
[0066] where K is the number of main services included in the application, s i is the number of the i-th main service and its corresponding redundant services (or, the number of redundant services set for the i-th main service plus 1), J is the number of candidate device types of the edge computing devices, n j is the number of the j-th type of edge computing devices, ServiceP ′Storage′][i] is the capacity of the storage device required for the i-th primary service (since redundant services are generated by replicating primary services, redundant services require the same storage device capacity as primary services) and can be determined based on the service attributes mentioned above, and DeviceP ′ Storage′][j] is the capacity of the storage device that each type-j edge computing device can provide and can be determined based on the device attributes mentioned above. Thus, through such constraint conditions, the sum of the capacities of the storage devices required for all primary services and redundant services of the application can be made not to exceed the sum of the capacities of the storage devices that all edge computing devices can provide.
[0067] It should be noted that the four types of resource constraints discussed in detail above are only examples, and those skilled in the art can conceive of taking the requirements for other types of resources as resource constraint conditions, such as power, network transmission rate, and so on, and such variations will not depart from the protection scope of the present disclosure.
[0068] In addition, the third constraint condition may include a security constraint. Generally speaking, different types of edge computing devices can provide different security levels, and devices with lower failure rates are usually considered to be more secure devices. Correspondingly, various types of services usually have different security requirements. As discussed in the above example scenario of roadside applications, the detection service generally requires a higher security level compared to the recording service, because the detection service usually results in more serious consequences in the event of a failure. Therefore, the detection service can be deployed on more secure edge computing devices to meet the security requirements.
[0069] Combined with the example scenario of roadside applications discussed above, the security level can be determined based on the Automotive Safety Integrity Level (ASIL). As stipulated in the standard ISO 26262 (“Road vehicles - Functional safety ASIL”) formulated by the International Organization for Standardization (ISO), the ASIL of components (including software and hardware) can be divided into five levels - QM, A, B, C, D, where, in the order from QM to D, the automotive safety integrity level gradually increases, that is, the requirements for functional safety gradually increase.
[0070] To help take security as a constraint condition in the constraint optimization algorithm, security can be quantified. For example, the five ASIL levels can be quantified respectively as shown in Table 1 below.
[0071] Table 1
[0072] ASIL Level Quantified Safety Level QM 0 A 1 B 2 C 3 D 4
[0073] In an example aspect, the security constraint can be expressed as:
[0074]
[0075] where K is the number of primary services included in the application, s i is the number of the i-th primary service and its corresponding redundant services (or, the number of redundant services set for the i-th primary service plus 1), J is the number of candidate device types of the edge computing device, n j is the number of the j-th type of edge computing device, ServiceP ′ ASIL′][i] is the quantified security level required by the i-th primary service (since the redundant service is generated by replicating the primary service, the redundant service has the same quantified security level as the primary service. For example, the security level can take values in [0, 4], as defined in Table 1 above) and can be determined based on the service attributes mentioned above, and DeviceP ′ ASIL′][j] is the quantified security level that each j-th type of edge computing device can provide (for example, it can take values in [0, 4], as defined in Table 1 above) and can be determined based on the device attributes mentioned above. Thus, through such constraint conditions, the sum of the quantified security levels required by all primary services and redundant services of the application does not exceed the sum of the quantified security levels that all edge computing devices can provide.
[0076] In another example aspect, the security constraint can also be expressed as:
[0077]
[0078] The meanings of the symbols in formula (8) are the same as those of the same symbols in formula (7) discussed above. Thus, through such constraint conditions, the maximum value of the quantified security levels required by all primary services and redundant services does not exceed the sum of the quantified security levels that all edge computing devices can provide.
[0079] Generally speaking, the security constraint associated with the sum of security levels discussed in conjunction with formula (7) imposes more stringent security requirements but may lead to waste of computing resources. On the contrary, the security constraint associated with the maximum value of security levels discussed in conjunction with formula (8) imposes more relaxed security requirements but is more cost-effective in terms of hardware. Thus, a choice can be made between these two types of security constraints based on specific requirements.
[0080] It should be noted that although the content regarding safety constraints has been discussed above by taking ASIL as an example, it is conceivable to perform such safety constraints based on other safety metrics (such as safety integrity level SIL, etc.) without departing from the scope of the present disclosure.
[0081] In an exemplary aspect, to determine the above-mentioned service attributes and device attributes, property tables for the edge computing device and the service respectively can be obtained to maintain values related to various attributes of various types of edge computing devices and various types of services involved in the above-mentioned constraint optimization algorithm. For example, the property table can be established and provided by the provider of the corresponding edge computing device or service.
[0082] As shown in the example of Table 2 below, the property table of the edge computing device can maintain values associated with attributes such as resources that various types of edge computing devices can provide (including the number of CPU cores, the capacity of video memory, the capacity of RAM, the capacity of the storage device), safety level, cost, etc. Among them, the resource attribute, the safety level attribute, and the cost attribute can be used in association with the above-mentioned resource constraint, safety constraint, and device cost constraint respectively. More specifically, the cost attribute can be used to determine the value of DeviceP ′ cost ′ , the resource attribute can be used to determine the values of DeviceP ′ CPU′], DeviceP ′ GPUmemory′], DeviceP ′ RAM′], DeviceP ′ Storage′], and the safety level attribute can be used to determine the value of DeviceP ′ ASIL′].
[0083] Table 2
[0084] Device Type CPU Core Video Memory RAM Storage Device Safety Cost Type 1 6 2GB 8GB 16GB 1 399 USD Type 2 8 8GB 32GB 32GB 1 699 USD Type 3 12 16GB 64GB 64GB 3 1999 USD
[0085] The number of CPU cores, the capacity of video memory, the capacity of RAM, the capacity of the storage device, and the cost that various types of edge computing devices can provide can be obtained from their specifications, etc., and the safety level can be determined through safety level assessment. For example, a professional third-party testing agency can perform the safety level assessment for the edge computing device, and determine the above-mentioned safety level according to the results of the safety level assessment.
[0086] As shown in Table 3 below, the attribute table of the service can maintain values associated with attributes such as the resources required for various types of services (including the number of CPU cores, the capacity of video memory, the capacity of RAM, and the capacity of the storage device), the security level, and its failure rate. Among them, the resource attribute and the security level attribute can be used in association with the above-mentioned resource constraint and security constraint respectively, and the failure rate can be used in association with the above-mentioned reliability optimization objective. More specifically, the failure rate attribute can be used to determine the value of F i value, and the resource attribute can be used to determine ServiceP ′ CPU′], ServiceP ′ GPU memory′], Service ′ RAM′], Service ′ Storage′] value, and the security level attribute can be used to determine Service ′ ASIL′] value.
[0087] Table 3
[0088] Service CPU Core Video Memory RAM Storage Device Safety Failure Rate 1 2 4GB 6GB 1GB 3 0.1 2 4 8GB 8GB 1GB 2 0.2 3 2 4GB 2GB 10GB 1 0.05
[0089] The number of CPU cores, the capacity of video memory, the capacity of RAM, the capacity of the storage device, and the security level required for each service can be determined by technicians during the development and design stage of the application, and its failure rate can be determined through failure testing.
[0090] Thus, one or more optimal solutions can be obtained using the constraint optimization algorithm discussed in detail above. In one example aspect, relevant software can be used to solve the above-mentioned constraint optimization problem. The obtained optimal solution corresponds to the planned deployment plan, and this deployment plan can indicate how many redundant services should be set for each main service, and how many edge computing devices should be set for each candidate device type among the above-mentioned candidate device types.
[0091] After obtaining the deployment plan, based on this deployment plan, it can be determined how to deploy the application on the edge cluster. In one example, after determining through the deployment plan how many redundant services each main service has and how many edge computing devices each candidate device type includes, the main edge computing device in the edge cluster (for example, the main edge computing device 202 discussed above in combination with Figure 2 discussion) can control the deployment of the application on the edge cluster based on such information. In one example, the controller in the main edge computing device can control which or which working edge computing devices the main service and redundant services should be deployed on based on the deployment plan, and the corresponding working edge computing devices can complete the deployment of the service under the control of the main edge computing device.
[0092] In one example, when deploying services on edge computing devices based on a deployment planning scheme, further restrictions can be imposed. The service sorting of multiple primary services can be determined based on the resource requirements of each primary service (the redundant services of the primary service have the same resource requirements). The device sorting of multiple candidate device types can be determined based on the resources that can be provided by the edge computing devices of each candidate device type. Among them, the resource requirements of the services and the resources that can be provided by the edge computing devices can be for various resource items discussed above. For example, the number of CPU cores, the capacity of video memory, the capacity of RAM, the capacity of the storage device, and any other available resource items. In one example, the above service sorting can be a rough sorting, where services with the same or similar resource requirements can have equal rankings in the determined service sorting. Similarly, the above device sorting can be a rough sorting, where edge computing devices that can provide the same or similar resources can have equal rankings in the determined sorting. Since the sum of the resource requirements of all services has been ensured not to exceed the sum of the resources that all edge computing devices can provide through the constraint optimization algorithm discussed above, this rough sorting will not affect the execution of the deployment process.
[0093] Then, based on the determined service sorting and device sorting, each primary service and its redundant service can be deployed on one or more edge computing devices. For example, the primary service with the first ranking in the determined service sorting and its redundant service can be deployed first, and they can be deployed on the edge computing devices belonging to the device type with the first ranking in the determined device sorting, and so on to complete the deployment of other services.
[0094] Figure 5 The flowchart of a method for planning the deployment of applications on an edge cluster according to an example embodiment of the present disclosure is shown. In one example, Figure 5 the operations shown can be performed by the primary edge computing device 202 in the edge cluster as discussed in conjunction with Figure 2 In another example, Figure 5 the operations shown can be performed on any computing device capable of performing deployment planning. In this example, the establishment of the edge cluster may not be completed yet, and the determined deployment planning scheme can be used to guide which edge computing devices to use to form the edge cluster. After determining the deployment planning scheme through the operations shown in Figure 5 the primary edge computing device (e.g., its controller) can control how to deploy services on the corresponding working edge computing devices based on this deployment planning scheme, as discussed below in conjunction with Figure 6
[0095] In S502, multiple primary services included in an application to be deployed on an edge cluster can be determined. As discussed above, during the development and design phase of the application, it can be divided into a series of primary services. In one example, redundant services of the primary service can be generated by replicating the primary service.
[0096] In S504, multiple candidate device types of edge computing devices for forming the edge cluster can be determined. As discussed above, this can be understood as defining a pool of available device types, such that which types of edge computing devices to use for forming the edge cluster can be selected from this pool of device types.
[0097] In S506, based on the multiple primary services and the multiple candidate device types, a deployment planning scheme can be determined. In one example, with maximizing the reliability of the application as the optimization goal, a constraint optimization algorithm can be used to determine the deployment planning scheme, where the deployment planning scheme can indicate the number of redundant services to be set for each primary service and the number of edge computing devices to be set for each candidate device type. In one example aspect, one or more of device cost constraints, resource constraints, and security constraints can be used as the constraint conditions of the constraint optimization algorithm. Thus, by using the method of the present disclosure, a deployment planning scheme that maximizes the reliability of the application while satisfying the above constraint conditions can be obtained, thereby realizing the optimized deployment of the application on the edge cluster.
[0098] Figure 6 The flowchart of a method for deploying an application on an edge cluster according to an example embodiment of the present disclosure is shown. In one example, Figure 6 The operations shown can be performed by a primary edge computing device 202 (e.g., its controller 204) in the edge cluster as discussed in conjunction with Figure 2 discussion.
[0099] In S602, a deployment planning scheme can be obtained, and the deployment planning scheme can be determined by using the method discussed above in conjunction with Figure 5 discussion.
[0100] In S604, according to the deployment planning scheme, the deployment of the application on the edge cluster can be controlled. For example, based on the resource requirements of each primary service, the service sorting of the multiple primary services can be determined, and based on the resources that can be provided by the edge computing devices of each candidate device type, the device sorting of the multiple candidate device types can be determined. Then, according to the service sorting and the device sorting, the deployment of each primary service and the redundant services of this primary service on the corresponding one or more edge computing devices can be controlled.
[0101] Figure 7A block diagram of a device according to an example embodiment of the present disclosure is shown, and the device can implement the method for optimized deployment applied to an edge cluster discussed above.
[0102] The exemplary device 700 includes a processor 704 connected to an internal communication bus 702. The processor 704 is configured to execute instructions in a memory 706 to implement the method for optimized deployment applied to an edge cluster described in detail above. Examples of the processor 704 may include a central processing unit (CPU), a microcontroller, and so on. The memory 706 suitable for tangibly embodying computer program instructions and data includes various forms of memory, such as EPROM, EEPROM, and flash memory devices, and so on. The device 700 may also include an input interface 708 and an output interface 710. The input interface 708 is used to receive input signals and data. The output interface 710 is used to send output signals and data.
[0103] The computer program may include instructions executable by a computer, and the instructions are used to cause the processor 704 of the device 700 to execute the method for optimized deployment of the present disclosure applied to an edge cluster. The program may be recorded on any data storage medium including a memory. For example, the program may be implemented in digital electronic circuits, or in computer hardware, firmware, software, or a combination thereof. The process / method steps described in the present disclosure may be executed by a programmable processor executing program instructions to execute the method, steps, and operations by operating on input data and generating outputs.
[0104] In addition to what is described herein, various modifications may be made to the disclosed embodiments and implementations of the present invention without departing from the scope thereof. Therefore, the descriptions and examples herein should be construed as illustrative rather than restrictive. The scope of the present invention should be measured only by reference to the claims.
Claims
1. A method for planning the deployment of an application on an edge cluster, comprising: determining a plurality of primary services included in the application; determining a plurality of candidate device types of edge computing devices for constituting the edge cluster; and based on the plurality of primary services and the plurality of candidate device types, determining a deployment planning scheme; wherein the deployment planning scheme indicates the number of redundant services to be set for each primary service among the plurality of primary services and the number of edge computing devices to be set for each candidate device type among the plurality of candidate device types.
2. The method according to claim 1, wherein, the determining the deployment planning scheme includes: using a constraint optimization algorithm to determine the deployment planning scheme, and the constraint optimization algorithm takes maximizing the reliability of the application as the optimization goal.
3. The method according to claim 2, wherein, the reliability of the application is calculated based on the failure rate of each primary service and the number of redundant services to be set for each primary service.
4. The method according to claim 3, wherein, the using a constraint optimization algorithm to determine the deployment planning scheme includes using one or more of the following as constraint conditions of the constraint optimization algorithm: device cost constraint; resource constraint; and security constraint.
5. The method according to claim 4, wherein, the device cost constraint is used to make the total device cost of all edge computing devices for constituting the edge cluster not exceed a predetermined device cost threshold.
6. The method according to claim 4, wherein: the resource constraint is used to make, for each resource item, the total resource requirements of all primary services and redundant services not exceed the total resources that all edge computing devices for constituting the edge cluster can provide; and the resource item includes one or more of the following: the number of central processing unit (CPU) cores; the capacity of video memory; the capacity of random access memory (RAM); and the capacity of the storage device.
7. The method according to claim 4, wherein, the security constraint includes: the total security requirements of all primary services and redundant services do not exceed the total security that all edge computing devices for constituting the edge cluster can provide; or the maximum value of the security requirements of all primary services and redundant services does not exceed the total security that all edge computing devices can provide.
8. The method according to claim 7, wherein, the security requirements of each primary service and redundant service and the security that each edge computing device can provide are characterized by the quantified automotive safety integrity level (ASIL).
9. The method according to claim 1, wherein: the edge cluster includes a roadside edge cluster; and the application includes a roadside perception application.
10. A method for deploying an application on an edge cluster, comprising: obtaining a deployment planning scheme determined by the method according to any one of claims 1-9; and controlling the deployment of the application on the edge cluster according to the deployment planning scheme.
11. The method according to claim 10, wherein, Controlling the deployment of the application on the edge cluster according to the deployment planning scheme includes: Determining the service sorting of the multiple primary services based on the resource requirements of each primary service; Determining the device sorting of the multiple candidate device types based on the resources that can be provided by the edge computing devices of each candidate device type; and Controlling the deployment of each primary service and the redundant service of the primary service on one or more corresponding edge computing devices according to the service sorting and the device sorting.
12. An apparatus for planning the deployment of an application on an edge cluster, including: A memory; A processor coupled to the memory, the processor being configured to execute the method according to any one of claims 1-9.
13. An apparatus for deploying an application on an edge cluster, including: A memory; A processor coupled to the memory, the processor being configured to execute the method according to any one of claims 10-11.
14. The apparatus according to claim 13, wherein, The apparatus includes a primary edge computing device in the edge cluster, and the primary edge computing device is used to manage the edge cluster.
15. A computer-readable medium storing a computer program including instructions that, when executed by a processor, cause the processor to be configured to execute the method according to any one of claims 1-9 or according to any one of claims 10-11.