Method and device for deploying and scheduling application instances
Through the global management platform, selecting appropriate site deployment application examples based on the user's service quality needs, the problem that the existing technology cannot provide edge cloud services with low latency and large traffic is solved, and the edge cloud services with high service quality are achieved.
Patent Information
- Application Number
- CN202010086421.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-02-11
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2040-02-11
AI Technical Summary
现有技术难以提供低时延与大流量的边缘云服务,无法满足用户的高服务质量需求。
Receive user service quality requirements information through the global management platform, select site deployment application instances that meet the delay requirements, and optimize the deployment location of application instances based on the number of connections and delay requirements.
It realizes providing users with high service quality edge cloud services, meeting the needs of low latency and large traffic.
Smart Images

Figure CN113259260B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of edge computing, and in particular to a method and device for deploying and scheduling application instances. Background Art
[0002] Currently, in the process of deploying edge application instances, users first specify resources, regions, and locations based on their own coverage requirements. Then, the global management platform filters available edge sites at the user-specified location and deploys the application instance. After that, users can view the region and geographic location where the application instance is located. However, creating edge application instances based on user-specified regions and locations is only a compromise under the current global management platform capabilities, and cannot truly provide low-latency and high-traffic edge cloud services. Summary of the invention
[0003] The present application provides a method for deploying and scheduling application instances, which can provide users with edge cloud services with high service quality.
[0004] In a first aspect, a method for deploying an application instance is provided, the method comprising: a global management platform receiving quality of service (QoS) requirement information from a first client, the QoS requirement information including a first delay requirement, a second delay requirement, and a first number of connections, the second delay requirement being superior to the first delay requirement, the QoS requirement information being input by a first user to the first client; the global management platform selecting a first available site that meets the first delay requirement from managed sites; the global management platform deploying one or more first application instances on the first available site, the number of connectable connections of the first application instance being less than or equal to the first number of connections.
[0005] Based on the above technical solution, the global management platform selects to deploy the application instance on a site that meets the first latency requirement input by the first user to the first client. Therefore, the application instance deployed by the global management platform can provide edge cloud services with high service quality for the first user's customers.
[0006] In combination with the first aspect, in certain implementations of the first aspect, the global management platform selects a first available site that meets the first latency requirement from the managed sites, including: the global management platform selects the first available site that meets the first latency requirement from the managed sites according to a first global QoS information table, and the first global QoS information table includes QoS information of the sites managed by the global management platform.
[0007] In combination with the first aspect, in certain implementations of the first aspect, the method also includes: the global management platform selects a second application instance that does not meet the second latency requirement from one or more of the first application instances; the global management platform selects a second available site that meets the second latency requirement from the managed sites; and the global management platform deploys the second application instance on the second available site.
[0008] Based on the above technical solution, the global management platform optimizes and updates the deployment location of application instances that do not meet the second latency requirement based on the second latency requirement, so that the optimized application instances can provide higher quality edge cloud services to the customers of the first user.
[0009] In combination with the first aspect, in certain implementations of the first aspect, the method also includes: the global management platform sets a resource reservation threshold based on the number of connectable third application instances deployed on the first site managed; when the number of remaining connections at the first site is less than the resource reservation threshold, the global management platform deploys a fourth application instance on the first site, and the QoS information of the fourth application instance is the same or equivalent to the QoS information of the third application instance.
[0010] In combination with the first aspect, in certain implementations of the first aspect, the method also includes: the global management platform predicts an increase in the number of connections in a first network segment based on historical access data, and the first network segment is the network segment that carries the third application instance; the global management platform calculates the number of the fourth application instances based on the increase in the number of connections in the first network segment.
[0011] In combination with the first aspect, in certain implementations of the first aspect, the method also includes: the global management platform receives a first request message from a second client or a regional management platform, the first request message is used to request scheduling of an application instance for the second client, and the first request message includes identification information of the second client; the global management platform filters out available application instances from one or more of the first application instances based on the second request message and the second global QoS information table, and the second global QoS information table includes QoS information of the first application instance.
[0012] In combination with the first aspect, in some implementations of the first aspect, the first request message also includes a third delay requirement input by the second user to the second client.
[0013] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: the global management platform scheduling an application instance for the second user according to the QoS level of the second user.
[0014] In a second aspect, a method for deploying an application instance is provided, the method comprising: a first client receives QoS requirement information input by a first user through a first interface, the QoS requirement information comprising a first delay requirement, a second delay requirement, and a first number of connections, the second delay requirement being superior to the first delay requirement; the first client sends the QoS requirement information to a global management platform.
[0015] Based on the above technical solution, the global management platform selects to deploy the application instance on a site that meets the first latency requirement input by the first user to the first client. Therefore, the application instance deployed by the global management platform can provide edge cloud services with high service quality for the first user's customers.
[0016] In combination with the second aspect, in some implementations of the second aspect, the first interface includes an application programming interface.
[0017] According to a third aspect, a method for scheduling application instances is provided, the method comprising: a regional management platform receives a second request message from a second client, the second request message is used to request scheduling of an application instance for the second client, the second request message includes identification information of the second client; the regional management platform filters out available application instances from one or more fifth application instances based on the second request message and a service QoS information table, the fifth application instance being an application instance deployed by a global management platform on a site managed by the regional management platform based on QoS requirement information from a first client, the service QoS information table including QoS information of the fifth application instance.
[0018] Based on the above technical solution, the regional management platform schedules an application instance for the second client based on the request message from the second client and the service QoS information table, and can provide users with high-quality edge cloud services.
[0019] In combination with the third aspect, in certain implementations of the third aspect, the second request message also includes a third delay requirement input by the second user to the second client.
[0020] In combination with the third aspect, in certain implementations of the third aspect, the method further includes: the regional management platform scheduling an application instance for the second user according to the QoS level of the second user.
[0021] In combination with the third aspect, in certain implementations of the third aspect, the method further includes: the regional management platform sends a first request message to the global management platform based on the second request message, and the first request message is used to request scheduling of an application instance for the second client.
[0022] Based on the above technical solution, when the regional management platform is unable to schedule an available application instance for the second client, the regional management platform may request the global management platform to schedule an application instance for the second client.
[0023] In a fourth aspect, a global management platform is provided, including a transceiver unit and a processing unit: the transceiver unit is used to receive QoS requirement information from a first client, the QoS requirement information including a first delay requirement, a second delay requirement and a first number of connections, the second delay requirement is better than the first delay requirement, and the QoS requirement information is input by a first user to the first client; the processing unit is used to select a first available site that meets the first delay requirement from the managed sites; the processing unit is also used to deploy one or more first application instances on the first available site, and the number of connectable instances of the first application instance is less than or equal to the first number of connections.
[0024] In combination with the fourth aspect, in certain implementations of the fourth aspect, the processing unit is specifically used to select a first available site that meets the first latency requirement from the managed sites based on a first global QoS information table, and the first global QoS information table includes QoS information of the sites managed by the global management platform.
[0025] In combination with the fourth aspect, in certain implementations of the fourth aspect, the processing unit is also used to select a second application instance that does not meet the second latency requirement from one or more of the first application instances; the processing unit is also used to select a second available site that meets the second latency requirement from the managed sites; the processing unit is also used to deploy the second application instance on the second available site.
[0026] In combination with the fourth aspect, in certain implementations of the fourth aspect, the processing unit is also used to set a resource reservation threshold based on the number of connectable third application instances deployed on the first managed site; when the number of remaining connections at the first site is less than the resource reservation threshold, the processing unit is also used to deploy a fourth application instance on the first site, and the QoS information of the fourth application instance is the same or equivalent to the QoS information of the third application instance.
[0027] In combination with the fourth aspect, in certain implementations of the fourth aspect, the processing unit is also used to predict an increase in the number of connections in the first network segment based on historical access data, where the first network segment is the network segment carried by the third application instance; the processing unit is also used to calculate the number of the fourth application instances based on the increase in the number of connections in the first network segment.
[0028] In combination with the fourth aspect, in certain implementations of the fourth aspect, the transceiver unit is also used to receive a first request message from a second client or a regional management platform, the first request message being used to request scheduling of an application instance for the second client, the first request message including identification information of the second client; the processing unit is also used to filter out available application instances from one or more of the first application instances based on the second request message and the second global QoS information table, the second global QoS information table including QoS information of the first application instance.
[0029] In combination with the fourth aspect, in certain implementations of the fourth aspect, the first request message also includes a third delay requirement input by the second user to the second client.
[0030] In combination with the fourth aspect, in certain implementations of the fourth aspect, the processing unit is further used to schedule an application instance for the second user according to the QoS level of the second user.
[0031] In the fifth aspect, a client is provided, including a receiving unit and a sending unit: the receiving unit is used to receive QoS requirement information input by a first user through a first interface, the QoS requirement information including a first delay requirement, a second delay requirement and a first number of connections, and the second delay requirement is better than the first delay requirement; the transceiver unit is used to send the QoS requirement information to the global management platform.
[0032] In combination with the fifth aspect, in some implementations of the fifth aspect, the first interface includes an application programming interface.
[0033] In a sixth aspect, a regional management platform is provided, comprising a transceiver unit and a processing unit: the transceiver unit is used to receive a second request message from a second client, the second request message is used to request scheduling of an application instance for the second client, and the second request message includes identification information of the second client; the processing unit is used to filter out available application instances from one or more fifth application instances based on the second request message and a service QoS information table, the fifth application instance being an application instance deployed by the global management platform on the site managed by the regional management platform based on QoS requirement information from the first client, and the service QoS information table includes QoS information of the fifth application instance.
[0034] In combination with the sixth aspect, in certain implementations of the sixth aspect, the second request message also includes a third delay requirement input by the second user to the second client.
[0035] In combination with the sixth aspect, in certain implementations of the sixth aspect, the processing unit is further used to schedule an application instance for the second user according to the QoS level of the second user.
[0036] In combination with the sixth aspect, in certain implementations of the sixth aspect, the transceiver unit is further used to send a first request message to the global management platform based on the second request message, and the first request message is used to request scheduling of an application instance for the second client.
[0037] In a seventh aspect, a global management platform is provided, comprising a processor, wherein the processor is coupled to a memory and can be used to execute instructions in the memory to implement the method in the first aspect or any possible implementation of the first aspect.
[0038] In an eighth aspect, a client is provided, comprising a processor, wherein the processor is coupled to a memory and can be used to execute instructions in the memory to implement the method in the second aspect or any possible implementation of the second aspect.
[0039] In a ninth aspect, a regional management platform is provided, comprising a processor, wherein the processor is coupled to a memory and can be used to execute instructions in the memory to implement the method in the third aspect or any possible implementation of the third aspect.
[0040] In a tenth aspect, a processor is provided, comprising: an input circuit, an output circuit, and a processing circuit. The processing circuit is used to receive a signal through the input circuit and transmit a signal through the output circuit, so that the processor executes the method in the first to third aspects or any possible implementation of the first to third aspects.
[0041] In the specific implementation process, the processor can be a chip, the input circuit can be an input pin, the output circuit can be an output pin, and the processing circuit can be a transistor, a gate circuit, a trigger, and various logic circuits. The input signal received by the input circuit can be, for example, but not limited to, received and input by a receiver, and the signal output by the output circuit can be, for example, but not limited to, output to a transmitter and transmitted by the transmitter, and the input circuit and the output circuit can be the same circuit, which is used as an input circuit and an output circuit at different times. The embodiments of the present application do not limit the specific implementation methods of the processor and various circuits.
[0042] In an eleventh aspect, a processing device is provided, comprising a processor and a memory. The processor is used to read instructions stored in the memory, and can receive signals through a receiver and transmit signals through a transmitter to execute the method in the first to third aspects or any possible implementation of the first to third aspects.
[0043] Optionally, the number of the processors is one or more, and the number of the memories is one or more.
[0044] Optionally, the memory may be integrated with the processor, or the memory may be provided separately from the processor.
[0045] In the specific implementation process, the memory can be a non-transitory memory, such as a read-only memory (ROM), which can be integrated with the processor on the same chip or can be set on different chips respectively. The embodiments of the present application do not limit the type of memory and the setting method of the memory and the processor.
[0046] It should be understood that the relevant data interaction process, such as sending indication information, can be a process of outputting indication information from a processor, and receiving capability information can be a process of receiving input capability information from a processor. Specifically, the processed output data can be output to a transmitter, and the input data received by the processor can come from a receiver. Among them, the transmitter and the receiver can be collectively referred to as a transceiver.
[0047] The processing device in the above-mentioned eleventh aspect can be a chip. The processor can be implemented by hardware or by software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor can be a general-purpose processor, which is implemented by reading the software code stored in the memory. The memory can be integrated in the processor or can be located outside the processor and exist independently.
[0048] In the twelfth aspect, a computer program product is provided, which includes: a computer program (also referred to as code, or instruction), which, when executed, enables a computer to execute the method in the above-mentioned first to third aspects or any possible implementation of the first to third aspects.
[0049] In the thirteenth aspect, a computer-readable medium is provided, which stores a computer program (also referred to as code, or instruction) which, when executed on a computer, enables the computer to execute the method in the above-mentioned first to third aspects or any possible implementation of the first to third aspects.
[0050] In the fourteenth aspect, an edge cloud scheduling system is provided, including the aforementioned global management platform, a first client and a regional management platform. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a schematic diagram of an application scenario provided by an embodiment of the present application.
[0052] Figures 2 to 4 It is a schematic flowchart of a method for deploying an application instance provided in an embodiment of the present application.
[0053] Figure 5It is a schematic flowchart of a method for scheduling application instances provided in an embodiment of the present application.
[0054] Figure 6 to Figure 7 It is a schematic block diagram of the global management platform provided in an embodiment of the present application.
[0055] Figures 8 to 9 It is a schematic block diagram of the first client provided in an embodiment of the present application.
[0056] Figure 10 to Figure 11 It is a schematic block diagram of the regional management platform provided in an embodiment of the present application. DETAILED DESCRIPTION
[0057] Edge cloud computing, or edge cloud for short, is a cloud computing platform built on edge infrastructure based on the core of cloud computing technology and edge computing capabilities. It forms an elastic cloud platform with comprehensive computing, networking, storage, security and other capabilities at the edge, and forms an end-to-end technical architecture of "cloud-edge-end collaboration" with the central cloud and IoT terminals. By placing network forwarding, storage, computing, intelligent data analysis and other tasks at the edge, it reduces response latency, reduces cloud pressure, reduces bandwidth costs, and provides cloud services such as full-network scheduling and computing power distribution.
[0058] Edge cloud is a type of public cloud, based on widely covered small sites, generally content delivery network (CDN), Internet access point (POP), mobile edge computing (MEC), each node in a small cluster form to provide public cloud services to the outside world. It has the characteristics of low latency, generally 5ms latency in the area, large bandwidth (40Gb-600Gb+), and the number will reach more than one thousand, or even tens of thousands. It is also generally called fog computing, edge computing, edge cloud, micro cloud computing (cloudlet), etc. in the industry.
[0059] Region and available zone (AZ) are concepts involved in cloud computing. Region refers to the area where the data center is located, which can be a region (South China) or a city (Shenzhen, Dongguan); available zone refers to the physical area where the computer room and site are located, which has the characteristics of independent energy consumption and network. A region usually contains multiple low-latency interconnected availability zones, which are used for scenarios and services such as disaster recovery and load balancing in the same region. Most cloud computing services do not support cross-regional operations.
[0060] Edge cloud users pursue the ultimate experience, that is, what edge cloud users ultimately need is actually the guarantee of latency and traffic. Currently, in the process of creating edge application instances, users first specify resources, regions, and locations based on their own coverage requirements. Then, the global management platform filters available sites and deploys application instances at the user-specified location. After that, users can view the region and geographic location where the application instance is located. However, creating edge application instances based on user-specified regions and locations is only a compromise under the current capabilities of the global management platform, and cannot truly provide low-latency and high-traffic edge cloud services. Application instances can be virtual machines, containers, or software modules.
[0061] Therefore, an embodiment of the present application provides a method for deploying application instances and scheduling application instances to provide users with edge cloud services with low latency and high traffic.
[0062] It should be noted that the first user mentioned in the embodiment of the present application is a user using client #1, and the second user is a user using client #2. The first application instance is an initially deployed application instance, the second application instance is an application instance in the first application instance that does not meet the second latency requirement, the third application instance is an application instance that is finally deployed, and the fourth application instance is a newly added application instance when the third application instance is insufficient.
[0063] Figure 1 A schematic diagram showing an application scenario applicable to the method provided in the embodiment of the present application is shown. Figure 1 As shown, the application scenario of the present application may be in the field of edge computing.
[0064] The global management platform manages all sites in multiple edge clouds. Figure 1 In the example, the global management platform manages sites in edge cloud #1 and sites in edge cloud #2.
[0065] The edge cloud can be composed of edge clusters, which can include multiple processing devices. Each processing device in the edge cloud can connect to other clients that are not in the edge cloud, collect information from other clients, transmit it to the regional management platform in the edge cloud, and then transmit it to the global management platform or process it at the regional management platform. Figure 1 In the example, edge cloud #1 consists of site #1, site #2, and regional management platform #1, and regional management platform #1 can manage site #1 and site #2. Edge cloud #2 consists of site #3, site #4, and regional management platform #2, and regional management platform #2 can manage site #3 and site #4.
[0066] The physical form of a site can be a single processing device, and the global management platform can deploy application instances on the site.
[0067] An application instance refers to a specific application of the same application service deployed on different sites, that is, one application service can correspond to multiple application instances. Figure 1 As shown, multiple application instances (application instance #1 to application instance #5) corresponding to the same application service can be deployed on the same or different sites. For example, application instance #1 and application instance #2 are deployed on site #1, application instance #3 is deployed on site #2, and application instance #4 and application instance #5 are deployed on site #3.
[0068] It should be understood that the client connected to the processing device can be an access terminal, a user unit, a user station, a mobile station, a mobile station, a remote station, a remote terminal, a mobile device, a user terminal, a terminal, a wireless communication device, a user agent or a user device. The physical form of the site can also be a cellular phone, a cordless phone, a session initiation protocol (SIP) phone, a wireless local loop (WLL) station, a personal digital assistant (PDA), a handheld device with wireless communication function, a computing device or other processing device connected to a wireless modem, a vehicle-mounted device, a wearable device, etc., and the embodiments of the present application are not limited to this.
[0069] Figure 2 is a method for deploying an application instance provided in an embodiment of the present application, and the method 200 shows the process of interaction between the global management platform and the client. The global management platform and the client can be, for example, Figure 1 The global management platform and client in Figure 2 As shown, the method 200 includes S201 to S203, and each step is described in detail below.
[0070] S201, client #1 (an example of the first client) sends QoS requirement information to the global management platform. Correspondingly, the global management platform receives the QoS requirement information from client #1.
[0071] The QoS requirement information includes a first delay requirement, a second delay requirement, and a first number of connections, wherein the second delay requirement is superior to the first delay requirement. The QoS requirement information is input by the first user to the client #1. For example, the first delay requirement may be the maximum delay that the user can accept; the second delay requirement may be the delay that the user expects; and the first number of connections may be the maximum number of connections that the application instance can have.
[0072] The first user may input QoS requirement information through an application program interface (API).
[0073] Taking edge cloud service as an example, the existing edge cloud service API interface does not have relevant settings for business QoS requirement information. Therefore, the embodiment of the present application adds detailed descriptions related to business QoS requirement information on the basis of the existing edge cloud service API interface, forming an edge cloud service API interface based on business QoS requirement information. An example of the edge cloud service API interface provided by the embodiment of the present application is as follows:
[0074] "coverage": {
[0075] "coverageLevel": "region",
[0076] "coveragePolicy": "discrete",
[0077] "coverageSites": [{
[0078] "site": "Xibei, Huabei",
[0079] "demand":["CMCC:10"]}],
[0080] "domainname": www.myservice.com
[0081] "QoS": {
[0082] "Delay": {
[0083] "max": "30ms",
[0084] "expection":"5ms"},
[0085] "Access capability": {
[0086] "maxlinksPerInstance": 2000
[0087] "minlinksNum": 3W
[0088] "maxlinksNum": 1000W
[0089] }},
[0090] As described above, in the edge cloud service API interface provided by the embodiment of the present application, the assignment method of the parameter "demand" in the region description is removed, that is, the description of the number of instances is removed. According to the method of the embodiment of the present application, the number of instances can be determined based on the parameters in the access capability. For example, it can be calculated according to formula (1):
[0091] InstanceNum=(minlinksNum / maxlinksPerInstance)*120% Formula (1)
[0092] The embodiment of the present application adds the parameter "QoS" to the resource description of the existing solution, and adds the description of the QoS requirement information. Among them, the parameter "Delay" is a custom data structure, which represents the description of the delay QoS. The parameter "Accesscapability" is also a custom data structure, which represents the description of the access number QoS. The "Delay" structure contains two parameters: the parameter "max" represents the maximum delay; the parameter "expection" represents the expected delay. The "Accesscapability" structure contains three parameters: the parameter "maxlinksPerInstance" represents the maximum number of connections for each application instance; the parameter "maxlinksNum" represents the maximum number of connections for each site; the parameter "minlinksNum" represents the minimum number of connections for each site.
[0093] It should be understood that the API interface shown above only adds the delay and the maximum number of connections for each application instance as an example for illustration, but this should not limit the embodiments of the present application. In the API interface provided in the embodiments of the present application, descriptions of parameters such as transmission bandwidth and data packet loss rate can also be added.
[0094] It is understood that, through the edge service API interface described above, the first user can input the first latency requirement, the second latency requirement, and the first connection quantity configuration-related parameters to client #1. The assignment of the parameters described above is only an example and should not limit the embodiments of the present application. It should be understood that the first user can assign the above parameters according to his own needs.
[0095] Optionally, client #1 may also predict QoS requirement information based on the delay requirement information historically input by the first user, and then send the QoS requirement information to the global management platform.
[0096] S202: The global management platform selects a first available site that meets a first latency requirement from the managed sites.
[0097] Optionally, the global management platform selects a first available site that meets the first latency requirement from all or part of the managed sites.
[0098] The global management platform can filter out the first available site that meets the first delay requirement information according to the QoS information of the sites stored locally.
[0099] Optionally, the global management platform may establish a first global QoS information table, wherein the first global QoS information table includes the QoS information of the site. Then, the global management platform filters out the first available site that meets the first QoS requirement information according to the first global QoS information table.
[0100] The method for the global management platform to establish the first global QoS information table may be:
[0101] The global management platform establishes a first global QoS information table according to the managed sites and the designated operator network segments.
[0102] The designated operator may be, for example, China Mobile, China Telecom, China Unicom, or other operators, which is not limited in the embodiments of the present application. It is understood that different operators correspond to different network segments, and different regions correspond to different network segments. Therefore, in the case of designated operators and designated regions, the corresponding network segments are also fixed.
[0103] Table 1 shows an example of a first QoS information table.
[0104] Table 1
[0105] Site Identifier network segment Latency area Site #1 10.1.1.1-10.2.2.2 5ms Longitude-latitude Site #1 10.3.3.3-10.4.4.4 15ms Longitude-latitude Site #1 10.6.6.6-10.7.7.7 25ms Longitude-latitude Site #2 10.1.1.1-10.2.2.2 25ms Longitude-latitude Site #2 10.3.3.3-10.4.4.4 15ms Longitude-latitude Site #2 10.6.6.6-10.7.7.7 5ms Longitude-latitude Site #3 10.1.1.1-10.2.2.2 25ms Longitude-latitude Site #3 10.3.3.3-10.4.4.4 5ms Longitude-latitude Site #3 10.6.6.6-10.7.7.7 25ms Longitude-latitude Site #4 10.1.1.1-10.2.2.2 25ms Longitude-latitude Site #4 10.3.3.3-10.4.4.4 25ms Longitude-latitude Site #4 10.6.6.6-10.7.7.7 25ms Longitude-latitude
[0106] As shown in Table 1, taking the global management platform managing sites #1 to #4 as an example, and after specifying the operator and region, the corresponding network segments are: 10.1.1.1-10.2.2.2, 10.3.3.3-10.4.4.4, 10.6.6.6-10.7.7.7 as an example, the contents of the first QoS information table are explained.
[0107] The site identifier is used to identify the site, and the site identifier may be a different number assigned by the global management platform to each site.
[0108] The corresponding relationship between sites, network segments and delays is: the communication delay of a client with an Internet Protocol (IP) address belonging to the network segment to access the site. As shown in Table 1, the communication delay of a client with an IP address of 10.1.1.1-10.2.2.2 to access site #1 is 5ms; the communication delay of a client with an IP address of 10.3.3.3-10.4.4.4 to access site #1 is 15ms; the communication delay of a client with an IP address of 10.6.6.6-10.7.7.7 to access site #1 is 25ms. Similarly, Table 1 also shows the communication delays of clients with IP addresses belonging to different network segments to access sites #2 to #4 respectively.
[0109] The corresponding relationship between a site and an area is: the site is set in the area. As shown in Table 1, edge #1 to site #4 are set in the same area.
[0110] Taking the first latency requirement information as an example, the maximum communication latency is 20ms. According to Table 1, the global management platform knows that the communication latency of all clients accessing edge point #4 is 25ms, so site #4 is unavailable. Therefore, the global management platform can determine sites #1 to #3 as the first available sites.
[0111] S203: The global management platform deploys one or more first application instances on the first available site.
[0112] The connectable number of the one or more first application instances is less than or equal to the first connection number.
[0113] For example, the number of instances calculated according to formula (1) is 5. The global management platform can randomly deploy application instances #1 to #5 at sites #1 to #3.
[0114] For example, Figure 1 As shown, the global management platform randomly selects sites #1 to #4 in the area specified by the first user, and randomly deploys application instance #1 and application instance #2 on site #1, deploys application instance #3 on site #2, and deploys application instance #4 and application instance #5 on site #3 respectively.
[0115] Then, the global management platform allocates the network segment that each application instance is responsible for and calculates the communication delay for the client to access each application instance.
[0116] For example, application instance #1 is responsible for 10.1.1.1-10.2.2.2, application instance #2 is responsible for 10.3.3.3-10.4.4.4, application instance #3 is responsible for 10.6.6.6-10.7.7.7, application instance #4 is responsible for 10.1.1.1-10.2.2.2, and application instance #5 is responsible for 10.3.3.3-10.4.4.4.
[0117] The communication delay of the client accessing the application instance can also be understood as the communication delay of the client accessing the site where the application instance is located. According to Table 1, the communication delay of the client with IP address 10.1.1.1-10.2.2.2 accessing site #1 is 5ms. Therefore, it can be seen that the communication delay of the client with IP address 10.1.1.1-10.2.2.2 accessing application instance #1 is also 5ms. Similarly, according to Table 1, the communication delay for a client with an IP address of 10.3.3.3-10.4.4.4 to access application instance #2 is 15ms; the communication delay for a client with an IP address of 10.6.6.6-10.7.7.7 to access application instance #3 is 5ms; the communication delay for a client with an IP address of 10.1.1.1-10.2.2.2 to access application instance #4 is 25ms; and the communication delay for a client with an IP address of 10.3.3.3-10.4.4.4 to access application instance #5 is 5ms.
[0118] It should be understood that if there is only one network segment of the designated operator in the designated area, the global management platform does not perform the step of allocating the network segment responsible for each application instance. In this case, each application instance is responsible for the only network segment in this area.
[0119] In some implementations, the global management platform may also adjust the deployment location of the first application instance according to the second latency requirement. Figure 3 As shown, the method 200 may further include S204 to S206.
[0120] S204: The global management platform selects a second application instance that does not meet the second latency requirement from the one or more deployed first application instances.
[0121] Taking the second delay requirement information that the expected delay is 5ms as an example, and taking the first deployed application instance as the application instance #1 to application instance #5 described above as an example, as described above, the communication delay for the client whose IP address belongs to 10.1.1.1-10.2.2.2 to access application instance #1 is 5ms, and the communication delay for the client whose IP address belongs to 10.3.3.3-10.4.4.4 to access application instance #2 is 15ms; the communication delay for the client whose IP address belongs to 10.6.6.6-10.7.7.7 to access application instance #3 is 5ms; the communication delay for the client whose IP address belongs to 10.1.1.1-10.2.2.2 to access application instance #4 is 25ms; the communication delay for the client whose IP address belongs to 10.3.3.3-10.4.4.4 to access application instance #5 is 5ms.
[0122] Therefore, application instance #2 and application instance #4 may be determined as second application instances that do not meet the second latency requirement.
[0123] S205: The global management platform selects a second available site that meets the second latency requirement from the managed sites.
[0124] Optionally, the global management platform selects a second available site that meets the second latency requirement from all or part of the managed sites.
[0125] The global management platform can filter out the second available site that meets the second delay requirement information according to the QoS information of the site stored locally.
[0126] Optionally, the global management platform may filter out a second available site that meets the second delay requirement according to the first global QoS information table.
[0127] Taking application instance #2 and application instance #4 as the second application instance as an example, according to Table 1, the communication delay for clients with IP addresses belonging to 10.3.3.3-10.4.4.4 to access site #3 is 5ms, which meets the second delay requirement. The communication delay for clients with IP addresses belonging to 10.1.1.1-10.2.2.2 to access site #1 is 5ms, which meets the second delay requirement. Therefore, the global management platform can determine site #1 and site #3 as the second available sites.
[0128] S206: The global management platform deploys a second application instance on a second available site.
[0129] After the global management platform selects the second application instance that does not meet the second latency requirement, it can query the optimal deployment location of the network segment that the second application instance is responsible for according to the first QoS information table, and move the second application instance. Then, after the second application instance is moved, the total communication delay of the client accessing all application instances is calculated. If it is better than the original total communication delay, the deployment location of the second application instance is changed.
[0130] Taking application instance #2 and application instance #4 as the second application instance, according to Table 1, the communication delay of the client with IP address 10.3.3.3-10.4.4.4 accessing site #3 is 5ms. Therefore, the global management platform can move application instance #2 from site #1 to site #3. After moving application instance #2, the total communication delay of the client accessing application instance #1 to application instance #5 is 45ms, which is less than the original 55ms. Therefore, the deployment location of application instance #2 is changed to site #3.
[0131] According to Table 1, the communication delay for clients with IP addresses 10.1.1.1-10.2.2.2 to access site #1 is the smallest. Therefore, application instance #4 is moved from site #3 to site #1. After moving application instance #4, the total communication delay for clients to access all application instances is 25ms, which is less than the original 45ms. Therefore, the deployment location of application instance #4 is changed to site #1.
[0132] It can be understood that in S205 and S206, the scheduling target of the global management platform for moving the deployment location of the first application instance is all network segments provided by the designated operator within the coverage area of the first application instance.
[0133] It should be noted that the termination conditions of the above deployment location update process are: (1) the network segments covered by all application instances meet the user's expected latency; or (2) the scheduling time is greater than 1 minute. As long as any of the termination conditions is met, the above update process is terminated.
[0134] For example, in the process of updating the deployment locations of application instance #2 and application instance #4, if the scheduling time reaches one minute after the deployment location of application instance #2 is updated, the update process will be terminated even if the step of updating application instance #4 has not been completed. For another example, within one minute, the update process of application instance #2 and application instance #4 is completed. At this time, the coverage network segments of all application instances meet the user's expected latency, and the update process is terminated.
[0135] It is understandable that after a period of time, the number of connections for some application instances may reach the maximum. If a new client requests an application instance, an available application instance cannot be scheduled for it. In this case, Figure 4As shown, the method 200 may further include S207 and S208.
[0136] S207: The global management platform sets a resource reservation value according to the number of connectable third application instances deployed on the first site.
[0137] The connectable number may be the maximum number of connections of the deployed third application instance. The resource reservation threshold may be, for example, 20% of the maximum number of connections of the third application instance. When the remaining number of connections of the third application instance is less than the resource reservation threshold, the deployment of a new application instance is increased in advance to meet the access needs of more clients.
[0138] Optionally, the global management platform may set different resource reservation thresholds for sites of different QoS levels.
[0139] As an example, for sites with different QoS levels, the global management platform can set a fixed resource reservation threshold. For example, for a site with a QoS level of 99, the global management platform can set 10% of the maximum number of connections of the site as the resource reservation threshold; for a site with a QoS level of 9999, the global management platform can set 40% of the maximum number of connections of the site as the resource reservation threshold; for the site with the highest priority, the global management platform can also set 100% of the maximum number of connections of the site as the resource reservation threshold.
[0140] As another example, for each QoS level site, the global management platform can dynamically adjust the resource reservation threshold based on historical access data over a period of time (e.g., one year). If the number of client #2s accessing a site of a certain QoS level is large over a period of time, the resource reservation threshold is increased; if the number of client #2s accessing a site of a certain QoS level is small over a period of time, the resource reservation threshold is reduced. For example, in the past year, the number of client #2s accessing sites with a QoS level of 9999 was small, so the global management platform can reduce the reserved resources from 40% to 20%.
[0141] S208: The global management platform deploys a fourth application instance on the first site.
[0142] The QoS information of the fourth application instance is the same or equivalent to the QoS information of the third application instance. For example, the third application instance is deployed on the first site, and the network segment that the third application instance is responsible for is 10.1.1.1-10.2.2.2, and the maximum number of connections is 2000. Then the fourth application instance is also deployed on the first site, and the fourth application instance is also responsible for 10.1.1.1-10.2.2.2, and the maximum number of connections is 2000.
[0143] The relevant information for deploying the fourth application parameter can be determined based on the relevant information for deploying the third application instance. In addition, the global management platform can also determine the number of the fourth application instance.
[0144] The method for the global management platform to determine the number of fourth application instances may include the following steps:
[0145] Step 1: The global management platform predicts the increase in the number of connections in the first network segment based on historical access data. The first network segment is the network segment carried by the third application instance.
[0146] For example, within a certain period of time, the global management platform counts the network segments where the IP addresses of each client accessing the same service domain name are located; then the global management platform counts the number of clients whose IP addresses belong to the first network segment; then the global management platform predicts the increase in the number of connections in the first network segment based on the number of clients. For example, the increase in the number of connections in the first network segment may be 100 times the number of clients corresponding to the first network segment.
[0147] Step 2: The global management platform calculates the number of fourth application instances that need to be added based on the predicted increase in the number of connections in the first network segment.
[0148] The number of the fourth application instances that need to be added is equal to the increase in the number of connections in the first network segment divided by the maximum number of connections in the fourth application instance. The maximum number of connections in the fourth application instance is equal to the maximum number of connections in the third application instance.
[0149] Step 3: The global management platform deploys the fourth application instance on the first site.
[0150] In the embodiment of the present application, the global management platform can deploy application instances according to the latency requirement information that the user can accept, so that the global management platform can provide users with services required by the user. Furthermore, the global management platform can also optimize and update the deployment location of the application instance according to the latency requirement information expected by the user, so that the global management platform can provide users with better services. For example, it can provide users with latency and high traffic services.
[0151] After the global management platform deploys the application instance, the global management platform may record the application instance identifier, the network segment that the application instance is responsible for, and the QoS information of the application instance, and establish a second QoS information table.
[0152] Table 2 shows an example of the second QoS information table.
[0153] Table 2
[0154] Service domain name Application instance identifier Site Identifier Bearer network segment Latency Number of connections 1111 Application Example #1 Site #1 10.1.1.1-10.2.2.2 5ms 2000 1111 Application Example #2 Site #3 10.3.3.3-10.4.4.4 5ms 1500 1111 Application Example #3 Site #2 10.6.6.6-10.7.7.7 5ms 1000 1111 Application Example #4 Site #1 10.1.1.1-10.2.2.2 5ms 1000 1111 Application Example #5 Site #3 10.3.3.3-10.4.4.4 5ms 800
[0155] The application instance identifier is used to identify the application instance. The application instance identifier may be a number assigned by the global management platform to each application instance. The number of connections is the number of users currently connected to the application instance.
[0156] Above, combined Figures 2 to 4 The method for deploying an application instance provided in the embodiment of the present application is described. Figure 5 The method for scheduling application examples provided in an embodiment of the present application is described.
[0157] Figure 5 FIG. 1 is a schematic flow chart of a method for scheduling an application example provided in an embodiment of the present application. Figure 5 As shown, the method 300 includes S301 to S304, and each step is described in detail below.
[0158] S301, the regional management platform receives a request message #1 (an example of the second request message) from a client #2 (an example of the second client), where the request message #1 is used to request the regional management platform to allocate an application instance to the client #2.
[0159] The request message #1 carries the address and IP information of the client #2.
[0160] Optionally, the request message #1 may also carry a third delay requirement input by the second user into the client #2. The third delay requirement may be the expected delay of the second user.
[0161] It can be understood that the request message #1 can be sent by the second user using the client #2 to the regional management platform through the client #2. For example, when the second user accesses a service domain name through the client #2, the client #2 sends the request message #1 to the regional management platform closest to the client #2.
[0162] like Figure 1 As shown, the service domain name accessed by the second user through client #2 is 1111, and the edge cloud closest to client #2 is edge cloud #1. Therefore, client #2 sends request message #1 to regional management platform #1 in edge cloud #1 to request regional management platform #1 to allocate an application instance for it. The address information carried in the request message #1 is the address where client #2 is located, and the IP information of client #2 carried in the request message #1 is the IP address of the client. Figure 1 As shown, the IP address of client #2 is 10.1.1.3.
[0163] S302: The regional management platform filters available application instances from one or more fifth application instances according to the service QoS information table, and redirects to the IP address of the optimal application instance.
[0164] The service QoS information table is established by the regional management platform for the sites included in the edge cloud to which it belongs and the application instances deployed on the sites. The fifth application instance is an application instance deployed by the global management platform on the sites managed by the regional management platform according to the QoS requirement information from client #1.
[0165] like Figure 1 As shown in Table 3, regional management platform #1 manages site #1 and site #2, and regional management platform #2 manages site #3 and site #4. Therefore, the service QoS information table established by regional management platform #1 is shown in Table 3.
[0166] Table 3
[0167] Service domain name Application instance identifier Site Identifier Bearer network segment Latency Number of connections 1111 Application Example #1 Site #1 10.1.1.1-10.2.2.2 5ms 2000 1111 Application Example #3 Site #2 10.6.6.6-10.7.7.7 5ms 1000 1111 Application Example #4 Site #1 10.1.1.1-10.2.2.2 5ms 1000
[0168] The service QoS information table established by regional management platform #2 is shown in Table 4.
[0169] Table 4
[0170] Service domain name Application instance identifier Site Identifier Bearer network segment Latency Number of connections 1111 Application Example #2 Site #3 10.3.3.3-10.4.4.4 5ms 1500 1111 Application Example #5 Site #3 10.3.3.3-10.4.4.4 5ms 800
[0171] After receiving the request message #1 from client #2, the regional management platform can first allocate an application instance to the client #2 according to the address information in the request message #1. Then, the regional management platform determines the network segment where the IP address is located according to the IP address carried in the request message #1; then, the regional management platform finds the application instance responsible for the network segment according to the service QoS information table; finally, it determines whether the application instance is available according to the QoS information corresponding to the application instance.
[0172] For example, it is determined whether the number of connections of the site where the application instance is located has reached a maximum value. If so, the application instance is unavailable; if not, the application instance is available.
[0173] For another example, if the request message #1 sent by the client to the regional management platform also carries a third delay requirement, the regional management platform can also determine whether the communication delay corresponding to the application instance meets the third delay requirement. If it meets the third delay requirement, the application instance is available; if it does not meet the third delay requirement, the application instance is unavailable.
[0174] like Figure 1As shown, the IP address of client #2 is 10.1.1.3. After receiving the request message #1 from client #2, regional management platform #1 allocates application instance #1 to client #2. Then, according to the IP address carried in the request message #1, it determines that the network segment where the IP address is located is 10.1.1.1. Then, according to Table 3, regional management platform #1 finds that the application instances responsible for 10.1.1.1 are application instance #1 and application instance #4.
[0175] Taking the process of the first user creating an edge application instance and assigning a value of 2000 to the parameter "maxlinksPerInstance" as an example, the regional management platform can know from Table 3 that the current number of connections of application instance #1 has reached the maximum number of connections, so application instance #1 is unavailable; and the current number of connections of application instance #4 is 1000, which has not reached the maximum number of connections, so application instance #4 is available.
[0176] According to Table 3, the communication delay corresponding to application instance #4 is 5ms. If the request message #1 sent by the client to the regional management platform #1 carries the third delay requirement, and the third delay requirement is less than 5ms, then application instance #4 is also unavailable; if the third delay requirement is greater than or equal to 5ms, then application instance #4 is available.
[0177] If the regional management platform determines that all application instances are unavailable based on the request message #1 from the client and the service QoS information table, the method 300 may further execute S304-S305.
[0178] S303, the regional management platform sends a request message #2 (an example of the first request message) to the global management platform, where the request message #2 is used to request allocation of an application instance to the client #2.
[0179] The content of request message #2 corresponds to that of request message #1. For example, if request message #1 carries the address and IP information of client #2, then request message #2 also carries the address and IP information of client #2; for another example, if request message #1 carries the address, IP information and the third latency requirement of client #2, then request message #2 also carries the address, IP information and the third latency requirement of client #2.
[0180] S304: The global management platform filters available application instances from one or more first application instances according to the second QoS information table, and directs them to the IP address of the optimal application instance.
[0181] After the global management platform receives the request message #2 from the regional management platform, it first allocates an application instance to the client #2 based on the address information carried in the request message #2; then, it determines the network segment where the IP address is located based on the IP address; then, the global management platform finds the application instance responsible for the network segment based on the service QoS information; finally, it determines whether the application instance is available based on the QoS information corresponding to the application instance.
[0182] For example, it is determined whether the number of connections of the site where the application instance is located has reached a maximum value. If so, the application instance is unavailable; if not, the application instance is available.
[0183] For another example, if the request message #2 sent by client #2 to the global management platform also includes a third delay requirement, the global management platform can also determine whether the communication delay corresponding to the application instance meets the third delay requirement. If it meets the third delay requirement, the application instance is available; if it does not meet the third delay requirement, the application instance is unavailable.
[0184] like Figure 1 As shown, the IP address of client #2 is 10.3.3.5. After the global management platform receives the request message #2 from the regional management platform #1, it allocates the application instance #1 to the client #2 nearby, and then determines that the network segment where the IP address is located is 10.3.3.3 based on the IP address carried in the request message #2.
[0185] Then, the global management platform finds that the application instances responsible for 10.3.3.3 are application instance #2 and application instance #5 according to Table 2. Taking the case where the user creates an edge application instance and assigns a value of 2000 to the parameter "maxlinksPerInstance", the global management platform can see from Table 2 that the current number of connections of application instance #2 and the current number of connections of application instance #5 have not reached the maximum number of connections, so application instance #2 and application instance #5 are available.
[0186] According to Table 2, the communication delays corresponding to application instance #2 and application instance #5 are both 5ms. If the request message #2 sent by the regional management platform #1 to the global management platform carries a third delay requirement, and the third delay requirement is less than 5ms, then application instance #2 and application instance #5 are also unavailable; if the third delay requirement is greater than or equal to 5ms, then application instance #2 and application instance #5 are available.
[0187] If the global management platform determines that there is an available application instance according to request message #2 and the second QoS information table, the global management platform sends a response message to the client, which carries the identifier of the available application instance. In addition, the response message is also used to redirect the client to the edge cloud that manages the site where the available application instance is located.
[0188] If the global management platform determines that there is no available application instance based on request message #2 and the second QoS information table, the global management platform starts resource scheduling and adds a new application instance to meet the user's service request.
[0189] Optionally, the regional management platform or the global management platform may schedule application instances for client #2 according to the QoS level of client #2.
[0190] For example, the regional management platform or the global management platform preferentially schedules application instances for client #2 with a high QoS level.
[0191] The embodiment of the present application does not limit how to determine the QoS level of client #2.
[0192] For example, the QoS level management device sets the QoS level of client #2 according to the ratio of the total number of requests satisfied by client #2 to the total number of requests of client #2. The statistics of this ratio can be in years. If the ratio of the total number of requests satisfied by client #2 to the total number of requests of client #2 is 99%, the QoS level management device sets the QoS level of client #2 to 99; if the ratio of the total number of requests satisfied by client #2 to the total number of requests of client #2 is 99.99%, the QoS level management device sets the QoS level of client #2 to 9999.
[0193] It can be understood that when client #2 initially accesses the regional management platform or the global management platform, it can be considered that the regional management platform or the global management platform can satisfy the request of client #2.
[0194] Optionally, in extreme cases (for example, the reserved resources have been used up and no new application instances have been added), client #2 with a high QoS level can preempt the instance resources of client #2 with a low QoS level. For example, client #2 with a QoS level of 9999 can preempt the instance resources of client #2 with a QoS level of 99.
[0195] Combination of the above Figure 3The process of the terminal client requesting the regional management platform to allocate an application instance is shown. It should be understood that the client #1 can also directly send a request message #1 to the global management platform to request the global management platform cloud to allocate an application instance to the client. The method of client #1 directly requesting the global management platform to allocate an application instance to it can refer to the description of S303 to S304 above. For the sake of brevity, this embodiment of the application will not be repeated.
[0196] Above, combined Figures 2 to 5 The method for deploying application instances and scheduling application instances provided by the embodiments of the present application is described in detail. Figures 6 to 11 The device provided in the embodiments of the present application is described in detail.
[0197] Figure 6 1 is a schematic block diagram of a global management platform 500 provided in an embodiment of the present application. As shown in the figure, the global management platform 500 may include: a transceiver unit 510 and a processing unit 520.
[0198] Specifically, the global management platform 500 may include a Figures 2 to 4 Method 200 and Figure 5 The units of the method performed by the global management platform in the method 300 in the above are respectively for implementing Figures 2 to 4 Method 200, Figure 5 It should be understood that the specific process of each unit executing the above corresponding steps has been described in detail in the above method embodiment, and for the sake of brevity, it will not be repeated here.
[0199] It should be understood that the transceiver unit in the global management platform 500 may correspond to Figure 7 The communication interface 620 in the global management platform 600 shown in FIG. 5 may correspond to the processing unit 520 in the global management platform 500. Figure 7 The processor 610 in the global management platform 600 is shown in FIG.
[0200] Figure 7 6 is a schematic block diagram of a global management platform 600 provided in an embodiment of the present application. As shown in the figure, the global management platform 600 may include: a communication interface 620, a processor 610 and a memory 630.
[0201] Optionally, the global management platform 600 may further include a bus 640. The communication interface 620, the processor 610, and the memory 630 may be interconnected via the bus 640; the bus 640 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus 640 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0202] The memory 630 can be used to store program codes and data executed by the computer system. Therefore, the memory 630 can be a storage unit inside the processor 610, or an external storage unit independent of the processor 610, or a component including a storage unit inside the processor 610 and an external storage unit independent of the processor 610.
[0203] The processor 610 may be composed of one or more general-purpose processors, such as a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of the present application. The processor may also be a combination that implements a computing function, such as a combination of multiple microprocessors, a combination of a DSP and a microprocessor, and the like. The processor 610 may be used to run a program that processes a function in a related program code. That is, the processor 610 may implement the functions of determining a module and creating a module by executing the program code. For details about the functions of determining a module and creating a module, please refer to the relevant description in the aforementioned embodiment.
[0204] In a possible implementation, the processor 610 is used to run relevant program codes to implement the above-mentioned Figures 2 to 4 The method described in S202 to S208 shown in the figure, or to implement the above-mentioned Figure 5The method described in S304 shown, and / or other steps for implementing the technology described herein, etc., are not described in detail or limited in this application.
[0205] The communication interface 620 may be a wired interface (eg, an Ethernet interface) or a wireless interface (eg, a cellular network interface or a wireless local area network interface) for communicating with other modules / devices.
[0206] The memory 630 may include a volatile memory, such as a random access memory (RAM); the memory may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD); the memory 630 may also include a combination of the above-mentioned types of memory. The memory 630 may be used to store a set of program codes so that the processor 610 calls the program codes stored in the memory 630 to implement the functions of the communication module and / or the processing module involved in the embodiments of the present invention.
[0207] When the program code in the memory 630 is executed by the processor 610, the global management platform 600 can execute the method in the above-mentioned method embodiment 200 or 300.
[0208] Figure 8 700 is a schematic block diagram of a first client 700 provided in an embodiment of the present application. As shown in the figure, the first client 700 may include: a receiving unit 710 and a sending unit 720.
[0209] Specifically, the first client 700 may include a processor for executing Figure 2 The units of the method performed by the first client in the method 200 in the embodiment. Furthermore, the units in the first client 700 and the above-mentioned other operations and / or functions are respectively for implementing Figures 2 to 4 It should be understood that the specific process of each unit executing the above corresponding steps has been described in detail in the above method embodiment, and for the sake of brevity, it will not be repeated here.
[0210] It should be understood that the receiving unit and the sending unit in the first client 700 may correspond to Fig. 9 The transceiver 820 in the first client 800 is shown in FIG.
[0211] Fig. 9800 is a schematic block diagram of a first client 800 provided in an embodiment of the present application. As shown in the figure, the first client 800 includes: a processor 810 and a transceiver 820. The processor 810 is coupled to a memory and is used to execute instructions stored in the memory to control the transceiver 820 to send signals and / or receive signals. Optionally, the first client 800 also includes a memory 830 for storing instructions.
[0212] It should be understood that the processor 810 and the memory 830 can be combined into one processing device, and the processor 810 is used to execute the program code stored in the memory 830 to implement the above functions. In specific implementation, the memory 830 can also be integrated into the processor 810, or independent of the processor 810.
[0213] It should also be understood that the transceiver 820 may include a receiver (or receiver) and a transmitter (or transmitter). The transceiver may further include an antenna, and the number of antennas may be one or more.
[0214] Specifically, the first client 800 may include a processor for executing Figure 2 The units of the method performed by the first client in the method 200 in the embodiment. Furthermore, the units in the first client 800 and the above-mentioned other operations and / or functions are respectively for implementing Figures 2 to 4 It should be understood that the specific process of each unit executing the above corresponding steps has been described in detail in the above method embodiment, and for the sake of brevity, it will not be repeated here.
[0215] Fig.10 1 is a schematic block diagram of a regional management platform 1000 provided in an embodiment of the present application. As shown in the figure, the regional management platform 1000 may include: a transceiver unit 1010 and a processing unit 1020.
[0216] Specifically, the regional management platform 1000 may include a Figure 5 The units of the method performed by the regional management platform in the method 300 in the above are respectively for implementing Figure 5 It should be understood that the specific process of each unit executing the above corresponding steps has been described in detail in the above method embodiment, and for the sake of brevity, it will not be repeated here.
[0217] It should be understood that the transceiver unit in the regional management platform 1000 may correspond to Fig.11 The communication interface 1120 in the regional management platform 1100 shown in FIG. 1100 may correspond to the processing unit 1020 in the regional management platform 1000. Fig.11 The processor 1110 in the regional management platform 1100 is shown in FIG.
[0218] Fig.11 11 is a schematic block diagram of a regional management platform 1100 provided in an embodiment of the present application. As shown in the figure, the regional management platform 1100 may include: a communication interface 1120, a processor 1110 and a memory 1130.
[0219] Optionally, the regional management platform 1100 may further include a bus 1140. The communication interface 1120, the processor 1110 and the memory 1130 may be interconnected via the bus 1140; the bus 1140 may be a PCI bus or an EISA bus, etc. The bus 1140 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.11 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0220] The memory 1130 may be used to store program codes and data executed by the computer system. Therefore, the memory 1130 may be a storage unit inside the processor 1110, or an external storage unit independent of the processor 1110, or a component including a storage unit inside the processor 1110 and an external storage unit independent of the processor 1110.
[0221] The processor 1110 may be composed of one or more general-purpose processors, such as a CPU, a general-purpose processor, a DSP, an ASIC, an FPGA or other programmable logic device, a transistor logic device, a hardware component or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of the present application. The processor may also be a combination that implements a computing function, such as a combination of multiple microprocessors, a combination of a DSP and a microprocessor, and the like. The processor 1110 may be used to run a program that processes functions in a related program code. In other words, the processor 1110 executes the program code to implement the functions of determining a module and creating a module. For details about the functions of determining a module and creating a module, please refer to the relevant description in the aforementioned embodiment.
[0222] In a possible implementation, the processor 1110 is used to run relevant program codes to implement the above-mentioned Figure 5 The method described in S302 shown, and / or other steps for implementing the technology described herein, etc., are not described in detail or limited in this application.
[0223] The communication interface 1120 may be a wired interface (eg, an Ethernet interface) or a wireless interface (eg, a cellular network interface or a wireless local area network interface) for communicating with other modules / devices.
[0224] The memory 1130 may include a volatile memory, such as a RAM; the memory may also include a non-volatile memory, such as a ROM, a flash memory, a HDD or a SSD; the memory 1130 may also include a combination of the above-mentioned types of memories. The memory 1130 may be used to store a set of program codes so that the processor 1110 calls the program codes stored in the memory 1130 to implement the functions of the communication module and / or the processing module involved in the embodiments of the present invention.
[0225] When the program code in the memory 1130 is executed by the processor 1110 , the area management platform 1100 can execute the method in the above method embodiment 300 .
[0226] According to the method provided in the embodiment of the present application, the present application also provides a computer program product, which includes: a computer program code, when the computer program code is run on a computer, the computer executes Figures 2 to 5 A method according to any one of the embodiments shown.
[0227] According to the method provided in the embodiment of the present application, the present application also provides a computer-readable medium, which stores a program code, and when the program code is run on a computer, the computer executes Figures 2 to 5 A method according to any one of the embodiments shown.
[0228] According to the method provided in the embodiment of the present application, the present application also provides a system, which includes the aforementioned global management platform, the first client and the regional management platform.
[0229] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions may be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (digital subscriber line, DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a high-density digital video disc (DVD)), or a semiconductor medium (eg, a solid state disk (SSD)).
[0230] Each network element in each of the above-mentioned device embodiments may completely correspond to each network element in the method embodiment, and the corresponding unit or unit performs the corresponding steps. For example, the transceiver unit (transceiver) performs the steps of receiving or sending in the method embodiment, and other steps except sending and receiving may be performed by the processing unit (processor). The functions of the specific units may refer to the corresponding method embodiments. Among them, there may be one or more processors.
[0231] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.
[0232] The terms "unit", "system", etc. used in this specification are used to represent computer-related entities, hardware, firmware, a combination of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process, a processor, an object, an executable file, an execution thread, a program and / or a computer running on a processor. By way of illustration, both the application and the computing device running on a computing device can be components. One or more components can reside in a process and / or an execution thread, and a component can be located on a computer and / or distributed between two or more computers. In addition, these components can be executed from various computer-readable media having various data structures stored thereon. Components can, for example, communicate through local and / or remote processes according to signals having one or more data packets (e.g., data from two components interacting with another component between a local system, a distributed system and / or a network, such as the Internet interacting with other systems through signals).
[0233] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0234] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0235] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0236] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0237] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0238] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage media include: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks or optical disks.
[0239] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A method for deploying an application instance, It is characterized in that include: The global management platform receives quality of service (QoS) requirement information from a first client, wherein the QoS requirement information includes a first delay requirement, a second delay requirement, and a first number of connections, wherein the second delay requirement is superior to the first delay requirement, the QoS requirement information is input by a first user to the first client, and the first number of connections indicates a maximum number of connections that can be made for an application instance; The global management platform selects a first available site that meets the first latency requirement from the managed sites; The global management platform deploys one or more first application instances on the first available site, the number of connectable first application instances is less than or equal to the first number of connections, and the number of connectable first application instances is the maximum number of connections of the first application instance.
2. The method according to claim 1, It is characterized in that The global management platform selects a first available site that meets the first latency requirement from the managed sites, including: The global management platform selects a first available site that meets the first latency requirement from the managed sites according to a first global QoS information table, where the first global QoS information table includes QoS information of the sites managed by the global management platform.
3. The method according to claim 1 or 2, It is characterized in that The method further comprises: The global management platform selects a second application instance that does not meet the second latency requirement from one or more of the first application instances; The global management platform selects a second available site that meets the second latency requirement from the managed sites; The global management platform deploys the second application instance on the second available site.
4. The method according to claim 1 or 2, It is characterized in that The method further comprises: The global management platform sets a resource reservation threshold according to the number of connectable third application instances deployed on the managed first site; When the number of remaining connections of the first site is less than the resource reservation threshold, the global management platform deploys a fourth application instance on the first site, and the QoS information of the fourth application instance is the same or equivalent to the QoS information of the third application instance.
5. The method according to claim 1 or 2, It is characterized in that The method further comprises: The global management platform receives a first request message from a second client or a regional management platform, where the first request message is used to request scheduling of an application instance for the second client, and the first request message includes identification information of the second client; The global management platform filters available application instances from one or more of the first application instances according to the first request message and a second global QoS information table, wherein the second global QoS information table includes QoS information of the first application instances.
6. The method according to claim 5, It is characterized in that The first request message also includes a third delay requirement input by the second user to the second client.
7. A method for deploying an application instance, It is characterized in that include: The first client receives, through the first interface, quality of service QoS requirement information input by the first user, the QoS requirement information including a first delay requirement, a second delay requirement, and a first number of connections, the second delay requirement being better than the first delay requirement, and the first number of connections indicating a maximum number of connections that can be made for the application instance; The first client sends the QoS requirement information to the global management platform.
8. The method according to claim 7, It is characterized in that The first interface comprises an application programming interface.
9. A method for scheduling an application instance, It is characterized in that include: The regional management platform receives a second request message from the second client, where the second request message is used to request scheduling of an application instance for the second client, and the second request message includes identification information of the second client; The regional management platform filters out available application instances from one or more fifth application instances based on the second request message and the business service quality QoS information table. The fifth application instance is an application instance deployed on the site managed by the regional management platform by the global management platform based on the QoS requirement information from the first client. The business QoS information table includes QoS information of the fifth application instance.
10. The method according to claim 9, It is characterized in that The second request message also includes a third delay requirement input by the second user to the second client.
11. The method according to claim 9 or 10, It is characterized in that The method further comprises: The regional management platform sends a first request message to the global management platform according to the second request message, where the first request message is used to request scheduling of an application instance for the second client.
12. A global management platform, It is characterized in that Including transceiver unit and processing unit: The transceiver unit is used to receive quality of service QoS requirement information from a first client, the QoS requirement information includes a first delay requirement, a second delay requirement and a first connection quantity, the second delay requirement is better than the first delay requirement, the QoS requirement information is input by a first user to the first client, and the first connection quantity indicates the maximum connectable quantity of the application instance; The processing unit is used to select a first available site that meets the first delay requirement from the managed sites; The processing unit is further used to deploy one or more first application instances on the first available site, the number of connectable first application instances is less than or equal to the first number of connections, and the number of connectable first application instances is the maximum number of connections of the first application instance.
13. The global management platform according to claim 12, It is characterized in that The processing unit is specifically configured to select a first available site that meets the first delay requirement from the managed sites according to a first global QoS information table, wherein the first global QoS information table includes QoS information of the sites managed by the global management platform.
14. The global management platform according to claim 12 or 13, It is characterized in that The processing unit is further configured to select a second application instance that does not meet the second latency requirement from one or more of the first application instances; The processing unit is further configured to select a second available site that meets the second delay requirement from the managed sites; The processing unit is further configured to deploy the second application instance on the second available site.
15. The global management platform according to claim 12 or 13, It is characterized in that The processing unit is further configured to set a resource reservation threshold according to the number of connectable third application instances deployed on the managed first site; When the number of remaining connections of the first site is less than the resource reservation threshold, the processing unit is also used to deploy a fourth application instance on the first site, and the QoS information of the fourth application instance is the same or equivalent to the QoS information of the third application instance.
16. The global management platform according to claim 12 or 13, It is characterized in that The transceiver unit is further used to receive a first request message from a second client or a regional management platform, the first request message is used to request scheduling of an application instance for the second client, and the first request message includes identification information of the second client; The processing unit is further configured to filter available application instances from one or more of the first application instances according to the first request message and a second global QoS information table, wherein the second global QoS information table includes QoS information of the first application instances.
17. The global management platform according to claim 16, It is characterized in that The first request message also includes a third delay requirement input by the second user to the second client.
18. A client, It is characterized in that Including receiving unit and sending unit: The receiving unit is used to receive quality of service QoS requirement information input by a first user through a first interface, where the QoS requirement information includes a first delay requirement, a second delay requirement, and a first number of connections, where the second delay requirement is better than the first delay requirement, and the first number of connections indicates a maximum number of connections that can be made for an application instance; The sending unit is used to send the QoS requirement information to the global management platform.
19. The client according to claim 18, It is characterized in that The first interface comprises an application programming interface.
20. A regional management platform, It is characterized in that Including transceiver unit and processing unit: The transceiver unit is used to receive a second request message from a second client, where the second request message is used to request scheduling of an application instance for the second client, and the second request message includes identification information of the second client; The processing unit is used to filter out available application instances from one or more fifth application instances based on the second request message and the business service quality QoS information table. The fifth application instance is an application instance deployed by the global management platform on the site managed by the regional management platform based on the QoS requirement information from the first client. The business QoS information table includes QoS information of the fifth application instance.
21. The regional management platform according to claim 20, It is characterized in that The second request message also includes a third delay requirement input by the second user to the second client.
22. The regional management platform according to claim 20 or 21, It is characterized in that The transceiver unit is further used to send a first request message to the global management platform according to the second request message, where the first request message is used to request scheduling of an application instance for the second client.
Citation Information
Patent Citations
Dynamic network device selection for containerized application deployment
US10470059B1
Method and system for efficient deployment of web applications in a multi-datacenter system
US20120136697A1