Task processing method and device based on cloud server, equipment, medium and product
By using the control group function in cloud server instances to constrain resource usage of batch job clients (Agents), the problem of overoccupancy caused by Agent is solved, and the performance and reliability of ECS instances are improved.
Patent Information
- Application Number
- CN202510025190.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-09
AI Technical Summary
A batch job client (Agent) in a cloud server instance may cause a large amount of CPU resources to be occupied or memory leaked, resulting in large-scale accidents in ECS instances.
By adding the batch job client installed in the cloud server instance to the system daemon, the control group function of the system kernel is used to constrain the resource use of the Agent to ensure that its resource use is within a certain range.
It effectively limits the resource use of Agent, prevents accidents caused by excessive resource utilization, and improves the performance and reliability of ECS instances.
Smart Images

Figure CN119960980A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method, device, equipment, medium, and product for task processing based on a cloud server. Background Art
[0002] The batch job client, also known as the cloud assistant (Agent), is a native automated command execution service built for the cloud server (Elastic Compute Service, ECS). Users can use this client to remotely perform batch operations on ECS instances and execute operation and maintenance commands (such as Shell, Python, Powershell, Bat, etc.) without passwords, logins, virtual private clouds (Virtual Private Cloud) and subnets, and without using a jump server.
[0003] In addition, because the Agent runs inside the ECS instance, when the stability and robustness of the Agent are not good, problems such as the central processing unit (CPU) resources being occupied or memory leaks may occur during the operation of the Agent, which may cause large-scale accidents in the ECS instance. Summary of the invention
[0004] In order to solve the above technical problems, the present application provides a task processing method, device, equipment, medium, and product based on a cloud server to ensure that the amount of resources used by the Agent is limited to a certain range, thereby facilitating better improvement of the performance of the ECS instance on which the Agent is installed.
[0005] In order to achieve the above objectives, the technical solutions provided by this application are as follows:
[0006] The present application provides a task processing method based on a cloud server, the method comprising: in response to starting a cloud server instance, adding a batch job client installed in the cloud server instance to a system daemon of the cloud server instance, the system daemon being used to constrain resource usage of the batch job client by using a control group function of a system kernel in the cloud server instance; in response to starting the batch job client, running the batch job client in the cloud server instance according to the constraints implemented by using the control group function.
[0007] In one possible implementation, the system kernel includes a first control group for constraining resource usage of the batch job client and a parent control group of the first control group, the resource usage constraint of the first control group is determined based on the resource usage constraint inherited from the parent control group and a resource isolation parameter configured for the batch job client, the resource isolation parameter is used to indicate an upper limit of resource usage of the batch job client; running the batch job client in the cloud server instance according to the constraint implemented by utilizing the control group function includes: running the batch job client in the cloud server instance according to the resource usage constraint of the first control group.
[0008] In one possible implementation, if the resource isolation parameter and the inherited resource usage constraint are both used to limit the upper limit of usage of the target resource, the upper limit of usage limited for the target resource by the resource isolation parameter shall not exceed the upper limit of usage limited for the target resource by the inherited resource usage constraint.
[0009] In one possible implementation, the system kernel includes a first control group for constraining resource usage of the batch job client; the method further includes: in response to starting the batch job client, creating a second control group in the system kernel for constraining resource usage of user tasks run through the batch job client, the second control group being different from the first control group, and there is no parent-child relationship between the second control group and the first control group; and running the user task in the cloud server instance according to the resource usage constraint of the second control group.
[0010] In one possible implementation, the system kernel also includes a parent control group of the first control group and a parent control group of the second control group, the parent control group of the second control group is different from the parent control group of the first control group, and there is no parent-child relationship between the parent control group of the second control group and the parent control group of the first control group; the resource usage constraints of the second control group are determined based on the resource usage constraints inherited from the parent control group of the second control group and the resource usage constraints configured by the user task.
[0011] In a possible implementation manner, the parent control group of the second control group and the parent control group of the first control group are both child control groups of a root control group, and the parent control group of the root control group does not exist in the system kernel.
[0012] In one possible implementation, running the user task in the cloud server instance according to the resource usage constraints of the second control group includes: in response to a run request triggered by the batch job client for the user task, creating a subprocess in the first control group, the subprocess being used to execute the user task; after moving the subprocess to the second control group, running the user task through the subprocess, the running process of the user task satisfying the resource usage constraints of the second control group.
[0013] In one possible implementation, the method further includes: in response to a run request triggered by the batch job client for the user task, determining whether a second control group has been created for constraining resource usage of the user task run by the batch job client; if the second control group has not been created, creating the second control group in the system kernel.
[0014] In one possible implementation, the method further includes: after moving the sub-process to the second control group, updating the configuration parameters of the system daemon process, wherein the updated configuration parameters are used to indicate that the sub-process will not be moved from the second control group back to the first control group when the configuration file of the system daemon process is reloaded.
[0015] In one possible implementation, after running the user task in the cloud server instance according to the resource usage constraint of the second control group, the method further includes: in response to the resource usage of the batch job client exceeding the resource upper limit limited by the resource usage constraint of the first control group, closing the batch job client and keeping the user task in a running state.
[0016] The present application provides a task processing device based on a cloud server, comprising: an adding unit, which is used to add a batch job client installed in the cloud server instance to the system daemon of the cloud server instance in response to starting a cloud server instance, and the system daemon is used to use the control group function of the system kernel in the cloud server instance to constrain the resource usage of the batch job client; and an operating unit, which is used to operate the batch job client in the cloud server instance according to the constraints implemented by using the control group function in response to starting the batch job client.
[0017] The present application provides an electronic device, which includes: a processor and a memory; the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs in the memory, so that the electronic device performs the cloud server-based task processing method provided by the present application.
[0018] The present application provides a computer-readable medium, in which instructions or computer programs are stored. When the instructions or computer programs are executed on a device, the device executes the cloud server-based task processing method provided by the present application.
[0019] The present application provides a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the cloud server-based task processing method provided by the present application.
[0020] Compared with the related art, this application has at least the following advantages:
[0021] In the resource isolation scheme implemented based on the control group (Cgroup) mechanism provided in the present application, for any ECS instance, when starting the ECS instance, the Agent installed in the ECS instance is added to the system daemon (systemd) of the ECS instance, so that the systemd can use the Cgroup function of the system kernel in the ECS instance to constrain the resource usage of the Agent, so that when it is detected that the Agent is started, the Agent is run in accordance with the constraint in the ECS instance to ensure that the maximum resources used by the Agent during operation do not exceed the resource usage upper limit limited by the constraint, so that not only the amount of resources used by the Agent can be limited within a certain range, but also the Agent can be effectively closed in time when the resource usage of the Agent exceeds the limit, so as to effectively overcome the defect caused by the inability to close the Agent in time when the resource usage of the Agent exceeds the limit, and thus it is conducive to better improving the performance of the ECS instance.
[0022] In addition, in some possible implementations, for the above system kernel, the system kernel includes a first control group (such as assist.client.system) for constraining the resource usage of the Agent and a parent control group (such as cgroup.system.slice) of the first control group, so that the resource usage constraint of the first control group is determined according to the resource usage constraint inherited from the parent control group and the resource isolation parameter configured for the Agent and used to indicate the upper limit of the resource usage of the Agent, so that when the Agent is detected to be started, it is run in the ECS instance according to the resource usage constraint of the first control group. The Agent is configured to ensure that the maximum resources used by the Agent during operation do not exceed the resource usage upper limit limited by the resource usage constraints of the first control group, such as the resource limits described by the resource isolation parameters configured for itself and the resource limits inherited from the parent control group. This not only limits the amount of resources used by the Agent within a certain range, but also effectively ensures that the Agent can be shut down in time when the resource usage of the Agent exceeds the limit, thereby effectively overcoming the defect caused by the inability to shut down the Agent in time when the resource usage of the Agent exceeds the limit, and thus helping to better improve the performance of the ECS instance. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0024] Figure 1 A flowchart of a cloud server-based task processing method provided in an embodiment of the present application;
[0025] Figure 2 A schematic diagram of a hierarchical structure provided in an embodiment of the present application;
[0026] Figure 3 A schematic diagram of an embodiment of the present application that does not isolate the running process of a user job from the running process of an Agent;
[0027] Figure 4 A schematic diagram of isolating the running process of a user job from the running process of an Agent provided in an embodiment of the present application;
[0028] Figure 5 A schematic diagram of the structure of a task processing device based on a cloud server provided in an embodiment of the present application;
[0029] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0030] In order to facilitate understanding of the technical solution of the present application, some technical terms are introduced below.
[0031] An ECS instance is a virtual computer created on a cloud server. It has its own operating system, processor, memory, storage space, network connection and other resources. Users can create multiple instances to meet different applications and workloads according to their needs. Each instance is independent and users can use them to deploy and manage applications just like traditional physical servers.
[0032] The batch job client, also known as the cloud assistant (Agent), refers to a client installed in the ECS instance for performing some operations on the ECS instance (such as batch operation and maintenance operations such as changing the instance password), so that the client can provide automated command execution services for the ECS instance. In addition, the client can not only execute tasks configured for the client itself (such as system tasks), but also pull up (such as execute) some user jobs through the client. It should be noted that this application does not limit the way to pull up user jobs. For example, when the interface of the user job (such as the mini program interface) is deployed in the client, the user can pull up the user job by clicking on the interface.
[0033] The kernel is the most basic part of the operating system. It is a piece of software that provides secure access to computer hardware for many applications. This access is limited, and the kernel determines when and for how long a program can operate on a certain piece of hardware. In addition, the kernel is the core of an operating system, the first layer of software expansion based on hardware, providing the most basic functions of the operating system and the basis for the operation of the operating system. It is responsible for managing the system's processes, memory, device drivers, files, and network systems, and determines the system's performance and stability.
[0034] systemd is a system and service manager for the Linux operating system. It is responsible for starting system components when the system starts and managing system processes during system operation. In addition, using the Cgroup function of the Linux kernel, systemd can limit and monitor the resource usage of some services, such as CPU, memory, and input / output (I / O). In addition, by default, systemd creates a new Cgroup under cgroup.system.slice for each service it monitors (such as the above-mentioned Agent).
[0035] Cgroup is a mechanism written into the Linux kernel that provides a task aggregation and division mechanism for the Linux kernel, so that the Linux kernel has the following functions: it is a hierarchical process group running in a system, and users can allocate resources to it (such as CPU time, system memory, network bandwidth or a combination of these resources), so that Cgroup can be used to limit, record and isolate the resource usage of a group of processes (such as CPU, memory, disk I / O, etc.). In addition, the principle of Cgroup is that when the Linux kernel finds that the process resources are used too much, it uses the kernel state mechanism to ensure that the process is closed in time. In addition, Cgroup is a mechanism for managing processes by group. From the user level, Cgroup technology organizes all processes in the system into independent trees. Each node of the tree is a process group, and each tree is associated with one or more subsystems. The role of the tree is to group processes, and the role of the subsystem is to operate on these groups. In addition, some concepts involved in Cgroup are as follows:
[0036] (1) A task in a Cgroup refers to a process or thread.
[0037] (2) A control group (Cgroup) is a collection of one or more processes whose resource usage can be restricted, recorded, and isolated. It can be seen that a control group is a group of processes subject to the same resource restrictions. Resource control in Cgroup is implemented in units of control groups. A process can join a control group and migrate from one process group to another. A process in a process group can use the resources allocated by Cgroup in units of control groups, and is subject to the restrictions set by Cgroup in units of control groups.
[0038] (3) Hierarchy refers to a tree structure consisting of a control group and its child control groups. Each child control group can inherit certain resource restrictions of the parent control group. In addition, since control groups exist in the form of directories, control groups can be organized into a hierarchical form, that is, a tree consisting of multiple control groups. The child node control group (referred to as child control group) on the control group tree is the child of the parent node control group (referred to as parent control group) and inherits the specific attributes of the parent control group.
[0039] (4) Subsystem is a kernel module that is responsible for actual resource management, such as CPU, memory, disk I / O, etc. In addition, a subsystem is a resource controller. For example, the CPU subsystem is a controller that controls CPU time allocation. A subsystem must be attached to a level to work. After a subsystem is attached to a level, all control groups on this level are controlled by this subsystem. In addition, Subsystem can be used to schedule or limit the resources of each process group, but this statement is not entirely accurate, because sometimes we group processes just to do some monitoring and observe their status.
[0040] Through research, it is found that the problem shown in the background technology section can be solved by means of resource isolation, wherein resource isolation is used to limit the amount of resources that an agent can use within a certain range.
[0041] The research also found that some resource isolation methods have some defects. For example, when using virtual machines (VMs) to achieve resource isolation, creating a VM in a cloud server instance consumes a lot of resources, so the resource isolation solution achieved with the help of VMs is not suitable for cloud computing scenarios. For another example, when using the application container engine (Docker) to achieve resource isolation, although the additional resources consumed by running a Docker in a cloud server instance can be ignored, the cloud server instance may not have a Docker environment, resulting in the resource isolation solution achieved with the help of Docker not being a universal solution. For another example, in some scenarios, the solution of implementing resource isolation based on self-developed program self-monitoring can be implemented. This solution is to check whether the program's own resource usage has reached a certain threshold, and if so, end its own operation process. However, this solution has at least the defects shown in the following ①-③.
[0042] ① When the CPU usage of the Agent is too large, the program self-monitoring process cannot be allocated to the CPU time slice, resulting in the inability to terminate the Agent's running process in time with the help of the process, which in turn affects the performance of the cloud server instance.
[0043] ② Because the program self-monitoring process runs periodically, when the CPU usage of the Agent is too large, the program self-monitoring process triggered in the current cycle cannot be allocated to the CPU time slice, causing the process to enter the queue state. This can easily cause the program self-monitoring process triggered in multiple consecutive cycles to enter the queue state, thereby bringing a large amount of additional performance overhead to the cloud server instance, which in turn seriously affects the performance of the cloud server instance.
[0044] ③ Because the program self-monitoring code is independently developed by cloud vendors based on their own needs, the program self-monitoring code has not been verified in a large-scale environment, which makes it easy for bugs to appear in the program self-monitoring code, leading to unexpected situations.
[0045] Based on the above research, in order to better overcome the above defects, the present application provides a resource isolation solution based on the Cgroup mechanism. In the solution, for the Agent installed in the ECS instance, by adding the Agent to the systemd of the ECS instance, all tasks (such as system tasks) configured by the Agent itself are set as a process group that needs to be managed by systemd, so that systemd can subsequently limit and monitor the resource usage of the process group with the help of the Cgroup mechanism to ensure that when it is detected that the resources used by the process group exceed the limit, the Agent is promptly closed through the kernel state mechanism, so as to achieve the purpose of resource isolation for the Agent through the Cgroup mechanism, so as to effectively ensure that the Agent cannot be closed in time when the resources exceed the limit due to human reasons or system resource shortage reasons, thereby effectively solving the problems of unavailability of cloud server instances or inability to execute user tasks caused by exhaustion of resources by the Agent for any reason.
[0046] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0047] In order to better understand the technical solution provided by this application, the following first describes the cloud server-based task processing method provided by this application in conjunction with some drawings. Figure 1 As shown, the cloud server-based task processing method provided in the embodiment of the present application includes the following S1-S2.
[0048] S1: In response to starting a cloud server instance, a batch job client installed in the cloud server instance is added to a system daemon of the cloud server instance, where the system daemon is used to constrain resource usage of the batch job client using a control group function of a system kernel in the cloud server instance.
[0049] The system kernel is the most basic part of the operating system in the cloud server instance. For example, when the operating system in the cloud server instance is a Linux operating system, the system kernel is the kernel of the Linux operating system, so that the system kernel can constrain the resource usage of some processes (such as Agent) with the help of the control group function.
[0050] In addition, for a cloud server instance, the system daemon process of the cloud server instance is used to constrain resource usage of some processes (such as Agent) by utilizing the control group function of the system kernel in the cloud server instance.
[0051] In addition, in order to better constrain the resource usage of some processes (such as Agent), the above-mentioned system kernel may include a first control group for constraining the resource usage of the batch job client and a parent control group of the first control group. The resource usage constraint of the first control group is determined based on the resource usage constraint inherited from the parent control group and the resource isolation parameter configured for the batch job client. The resource isolation parameter is used to indicate the upper limit of resource usage of the batch job client.
[0052] The first control group refers to a Cgroup (such as a batch job client, such as an Agent) created to constrain the resource usage of the batch job client. Figure 2 Cgroup4 as shown in the figure, so that the Cgroup is used to indicate the resource usage upper limit of all tasks (such as system tasks) configured by the Agent itself, so that the amount of resources that the Agent can use can be limited to a certain range based on the Cgroup. It can be seen that in a possible implementation, the first control group not only includes all the tasks configured by the Agent itself as a process group, but also includes resource usage constraints (such as resource usage constraints) for limiting the resource usage upper limit of each process in the process group. Figure 2 The resource usage constraint 2 inherited from Cgroup2 and the resource usage constraint 4 set for Cgroup4 are shown.
[0053] It should be noted that the present application is not limited to the first control group. For example, it can be implemented using the Cgroup of assist.client.system, so that the Cgroup is used to limit the resource usage of each task of the batch job client itself.
[0054] The parent control group of the first control group refers to the control group that exists in the Hierarchy and is the parent node of the first control group (such as Figure 2 Cgroup2 as shown in the figure) so that the first control group can inherit some resource restrictions of the parent control group, so that the resources that can be used by the Agent used to build the first control group are limited by the parent control group. In this way, it is ensured that the maximum resource usage limit of the Agent does not exceed the resource limit of itself or the parent control group.
[0055] In addition, the present application does not limit the implementation method of the parent control group of the first control group. For example, it can be implemented using the control group cgroup.system.slice, so that the processes restricted by the first control group (such as the processes used by the Agent itself) can be treated as system tasks for resource restriction and monitoring, so that the kernel state mechanism of the system can be used to promptly discover and close the processes with excessive resources. Among them, cgroup.system.slice is used to limit the upper limit of resource usage of system tasks; and cgroup.system.slice is a child control group of the root control group (Root Cgroup), so that the cgroup.system.slice can inherit some resource restrictions of the RootCgroup. It should be noted that the Root Cgroup refers to the root node of a hierarchical structure, so that the parent node of the root node does not exist in the hierarchical structure. For example, the Root Cgroup can refer to the folder indicated by the directory / sys / fs / cgroup.
[0056] It can be seen that in a possible implementation, for a hierarchy managed by systemd, the first layer of the hierarchy includes the Root Cgroup (such as Figure 2 The second level of the Hierarchy includes some child nodes of the Root Cgroup (such as Figure 2 Cgroup2 and Cgroup3 shown in the figure), and these child nodes include at least the control group cgroup.system.slice, so that these child nodes can inherit some resource restrictions of the Root Cgroup (such as Figure 2The third level of the Hierarchy includes the child nodes of each node in the second level (such as Figure 2 Cgroup4 and Cgroup5 shown in the figure) so that the third layer of the Hierarchy includes at least the child control groups of the cgroup.system.slice control group, such as the assist.client.system control group.
[0057] The resource usage constraint of the first control group is used to indicate the resource usage upper limit of each process included in the first control group; and the resource usage constraint of the first control group is determined based on the resource usage constraint inherited from the parent control group of the first control group and the resource isolation parameter configured for the batch job client. The resource isolation parameter refers to the resource usage constraint set for the batch job client and used to indicate the resource usage upper limit of the batch job client, such as Figure 2 The configuration information "CPU limit = 10; Memory limit = 512MB" is shown.
[0058] In addition, the present application does not limit the determination process of the resource usage constraint of the first control group. For example, it can be specifically as follows: first, the resource usage constraint inherited from the parent control group of the first control group is compared with the resource isolation parameter configured for the batch job client to obtain a comparison result, so that the comparison result can indicate whether there is an overlapping item between the two, so that the overlapping item can indicate a resource item (such as CPU) that is restricted in the constraint information of the two sources; if there is no overlapping item, it can be determined that the constraint information of the two sources is complementary, so the union of the constraint information of the two sources can be determined as the resource usage constraint of the first control group;
[0059] However, if there is an overlapping item, it may be determined whether the usage upper limit defined by the resource isolation parameter for the overlapping item exceeds the usage upper limit defined by the inherited resource usage constraint for the overlapping item;
[0060] If it does not exceed, it can be determined that the usage upper limit defined by the resource isolation parameter for the overlapping item is in a valid state in the first control group, but the usage upper limit defined by the inherited resource usage constraint for the overlapping item is in an invalid state in the first control group, so the usage upper limit defined for the overlapping item can be first deleted from the inherited resource usage constraint to obtain the deleted resource usage constraint; then, the resource usage constraint of the first control group is determined based on the deleted resource usage constraint and the resource isolation parameter;
[0061] If it exceeds, it can be determined that the usage upper limit defined by the resource isolation parameter for the overlapping item is in an invalid state in the first control group, but the usage upper limit defined by the inherited resource usage constraint for the overlapping item is in a valid state in the first control group. Therefore, the usage upper limit defined for the overlapping item can be deleted from the resource isolation parameter first to obtain the deleted resource isolation parameter; and then the resource usage constraint of the first control group is determined based on the deleted resource isolation parameter and the inherited resource usage constraint.
[0062] It can be seen that, under a possible implementation mode, in order to better improve efficiency, the above-mentioned resource isolation parameters can at least meet the following constraints: if the resource isolation parameter and the above-mentioned inherited resource usage constraint are both used to limit the usage upper limit of the target resource (such as the above-mentioned overlapping item), then the usage upper limit defined by the resource isolation parameter for the target resource does not exceed the usage upper limit defined by the inherited resource usage constraint for the target resource, thereby ensuring that the resource usage upper limits defined by the resource isolation parameter are all in a valid state, so that the resource usage constraint of the first control group finally determined can at least include all of the resource isolation parameters, so that the resource limit described by the resource usage constraint of the first control group comes from the resource isolation parameter, or from the inherited resource usage constraint, and further the resource usage upper limit of each process included in the first control group (such as the Agent's own process) does not exceed its own or the parent Cgroup's resource limit. In this way, it can be ensured that the resource upper limit of the Agent does not exceed its own or the parent Cgroup's resource limit.
[0063] Based on the relevant content of S1 above, it can be known that for any ECS instance, when starting the ECS instance, the Agent installed in the ECS instance is added to the systemd of the ECS instance, so that the systemd can use the Cgroup function of the system kernel in the ECS instance to constrain the resource usage of the Agent. Among them, because the resources that the Agent can use are limited by the control group cgroup.system.slice, so that the maximum resource limit of the Agent does not exceed the resource limit of itself or the parent Cgroup, it can effectively ensure that when the resource usage of the Agent exceeds the limit, the system kernel can timely discover and close the Agent with the help of some mechanisms, so as to effectively overcome the defects caused by the inability to close the Agent in time when the resource usage of the Agent exceeds the limit, which is conducive to better improving the performance of the ECS instance.
[0064] S2: In response to starting the batch job client, the batch job client is run in the cloud server instance according to the constraints implemented by using the control group function.
[0065] It should be noted that the present application does not limit the implementation method of the above-mentioned S2. For example, when the system kernel in the cloud server instance includes a first control group for constraining the resource usage of the batch job client, the S2 can specifically be: if it is detected that the Agent is started, the Agent is run in each ECS instance according to the resource usage constraint of the first control group, so that the running process of the Agent satisfies the resource restriction described by the resource usage constraint. This can effectively ensure that when the resource usage of the Agent exceeds the limit, the Agent can be discovered and closed in time, thereby effectively overcoming the defect caused by the inability to close the Agent in time when the resource usage of the Agent exceeds the limit, which is conducive to better improving the performance of the ECS instance.
[0066] Based on the relevant contents of S1 to S2 above, it can be known that in the resource isolation solution implemented based on the Cgroup mechanism provided in the present application, for any ECS instance, when starting the ECS instance, the Agent installed in the ECS instance is added to the system daemon (systemd) of the ECS instance, so that the systemd can use the Cgroup function of the system kernel in the ECS instance to constrain the resource usage of the Agent, so that when it is detected that the Agent is started, the Agent is run in accordance with the constraint in the ECS instance to ensure that the maximum resources used by the Agent during operation do not exceed the resource usage upper limit limited by the constraint. In this way, not only can the amount of resources used by the Agent be limited to a certain range, but also it can effectively ensure that the Agent can be closed in time when the resource usage of the Agent exceeds the limit, so as to effectively overcome the defect caused by the inability to close the Agent in time when the resource usage of the Agent exceeds the limit, and thus it is beneficial to better improve the performance of the ECS instance.
[0067] In addition, in some possible implementations, for the above system kernel, the system kernel includes a first control group (such as assist.client.system) for constraining the resource usage of the Agent and a parent control group (such as cgroup.system.slice) of the first control group, so that the resource usage constraint of the first control group is determined according to the resource usage constraint inherited from the parent control group and the resource isolation parameter configured for the Agent and used to indicate the upper limit of the resource usage of the Agent, so that when the Agent is detected to be started, it is run in the ECS instance according to the resource usage constraint of the first control group. The Agent is configured to ensure that the maximum resources used by the Agent during operation do not exceed the resource usage upper limit limited by the resource usage constraints of the first control group, such as the resource limits described by the resource isolation parameters configured for itself and the resource limits inherited from the parent control group. This not only limits the amount of resources used by the Agent within a certain range, but also effectively ensures that the Agent can be shut down in time when the resource usage of the Agent exceeds the limit, thereby effectively overcoming the defect caused by the inability to shut down the Agent in time when the resource usage of the Agent exceeds the limit, and thus helping to better improve the performance of the ECS instance.
[0068] After research, it was found that the resource isolation solution mentioned above, which is implemented by program self-monitoring, also has the following defects: during the process shutdown process, not only the Agent's own process will be shut down, but also the jobs run by the user through the Agent will be shut down, affecting the execution effect of the user's job.
[0069] After further research, it was found that the problem shown in the previous paragraph would also occur when only the Cgroup mechanism was used to implement Agent resource isolation. The reasons are as follows: the child process of Linux will inherit the Cgroup information of the parent process, and the job started by the user through the Agent (such as Figure 3 The User Task shown in FIG. 1 belongs to a child process of the Agent, so that the launched job will also be constrained by the Cgroup information of the Agent, so that the launched job will be directly closed when the Agent is closed.
[0070] It should be noted that for Figure 3For some of the content involved, TaskExecutor is a task runner, which has various resources to run tasks. The exec() series of functions are used to execute new programs in the current process, and they are an important part of implementing process control in the Linux system. In addition, Bash is used as the default command line interpreter in Linux. Bash is a powerful and flexible tool that can execute various commands and operating system tasks. User Task refers to a user job that needs to be started by Agent, such as a third-party application. In addition, fork can be used in Bash to create a child process and execute tasks in the child process. In Bash, the built-in wait() command can be used to wait for the background process to complete. The wait() command suspends the execution of the script until all background processes (that is, processes started with the & symbol) are completed.
[0071] After research, it was found that for the job started by the user through the Agent, in order to better ensure the execution effect of the job, the job should not be subject to the resource restrictions described by the Cgroup information of the Agent. Therefore, in order to better meet the requirements, this application provides a solution, which is to Figure 4 The job launching method shown in the figure is implemented so that the Agent process itself and the process launched by the Agent (such as user jobs) are in two different Cgroups with no parent-child relationship at any time. This ensures that the Agent's resources are limited while the jobs launched by the Agent are not limited.
[0072] Based on the above research, the present application also provides a possible implementation method of a cloud server-based task processing method. In this implementation method, the cloud server-based task processing method not only includes the above S1-S2, but may also include the following steps 11-12.
[0073] Step 11: In response to starting the batch job client, a second control group is created in the system kernel for constraining resource usage of user tasks run (such as pulled up) through the batch job client. The second control group is different from the first control group, and there is no parent-child relationship between the second control group and the first control group.
[0074] The second control group refers to a Cgroup (such as a batch job client) created for a user job launched by an agent and used to constrain the resource usage of the user job. Figure 2As shown in Cgroup5, the Cgroup can indicate that the user job is not subject to resource restrictions. It can be seen that in a possible implementation, the second control group not only includes the process group that uses the user job pulled by the Agent as a process, but also includes resource usage constraints (such as resource usage limits) for limiting the resource usage upper limit of each process (such as each user job) in the process group. Figure 2 The resource usage constraint 3 inherited from Cgroup3 and the resource usage constraint 5 specified for Cgroup5 are shown).
[0075] In addition, the second control group and the first control group above at least satisfy the following constraints: the second control group is different from the first control group, and there is no parent-child relationship between the second control group and the first control group, so that the second control group is independent of the first control group, so that the characteristics of resource restrictions presented by the processes included in the second control group (such as user jobs) can be completely different from the characteristics of resource restrictions presented by the processes included in the first control group (such as Agent processes), and thus, with the help of these two control groups, it is possible to ensure that the resources of the Agent are limited and the jobs launched by the Agent are not restricted.
[0076] In addition, in order to better improve the isolation effect, for any ECS instance with Agent installed, the system kernel in the ECS instance includes not only the first control group, the parent control group of the first control group, and the second control group, but also the parent control group of the second control group (such as Figure 2 Cgroup3 shown). In which, because the parent control group of the second control group is different from the parent control group of the first control group, and there is no parent-child relationship between the parent control group of the second control group and the parent control group of the first control group, the parent control group of the second control group is independent of the parent control group of the first control group, thereby better ensuring that the resource constraints required to be satisfied by the second control group are almost unlikely to be related to the resource constraints required to be satisfied by the first control group.
[0077] It should be noted that the present application does not limit the second control group. For example, it can be implemented using the Cgroup of assist.client.task, so that the Cgroup is used to limit the resource usage of user jobs launched by the batch job client.
[0078] The parent control group of the second control group refers to the control group that exists in the Hierarchy and is the parent node of the second control group (such as Figure 2 Cgroup3 as shown in the figure) so that the second control group can inherit some resource restrictions of the parent control group (such as Figure 2 The resource usage constraints shown in 3).
[0079] In addition, the present application does not limit the implementation method of the parent control group of the second control group. For example, it can be implemented using the control group cgroup.user.slice, so that the processes restricted by the second control group (such as user jobs launched by Agent) can be treated as non-system tasks (such as user tasks) for resource restriction and monitoring. Among them, cgroup.user.slice is used to limit the upper limit of resource usage of user jobs; and cgroup.user.slice is a child control group of the Root Cgroup, so that the cgroup.user.slice can inherit some resource restrictions of the Root Cgroup.
[0080] It can be seen that in one possible implementation, for a Hierarchy managed by systemd (such as Figure 2 For example, the first layer of the Hierarchy includes the Root Cgroup (such as Figure 2 The second level of the Hierarchy includes some child nodes of the Root Cgroup (such as Figure 2 Cgroup2 and Cgroup3 shown in the figure), and these child nodes include at least the cgroup.system.slice control group and the cgroup.user.slice control group, so that these child nodes can inherit some resource restrictions of the Root Cgroup (such as Figure 2 The third level of the Hierarchy includes the child nodes of each node in the second level (such as Figure 2 The third layer of the Hierarchy includes at least the sub-control groups of the cgroup.system.slice control group (such as the assist.client.system control group) and the sub-control groups of the cgroup.user.slice control group (such as the assist.client.task control group). In this way, the Agent's own processes and the tasks launched by the Agent are always in two different Cgroups with no parent-child relationship. This ensures that the Agent's resources are limited while the user jobs launched by the Agent are not limited.
[0081] The resource usage constraint of the second control group is used to indicate the upper limit of resource usage of each process included in the second control group; and the resource usage constraint of the second control group is determined based on the resource usage constraint inherited from the parent control group of the second control group and the resource usage constraint configured for the user task pulled up by the batch job client. It should be noted that the process of determining the resource usage constraint of the second control group is similar to the process of determining the resource usage constraint of the first control group above, and for the sake of brevity, it will not be repeated here.
[0082] In addition, the present application does not limit the implementation method of the resource usage constraint of the above-mentioned second control group. For example, it may specifically include: CPU limit = -1, memory limit = -1, ..., so that the resource usage constraint can indicate that there is no upper limit on resource usage for the task launched by the Agent, thereby enabling the resource usage constraint to indicate that the task is not restricted by resources.
[0083] Based on the relevant content of the above step 11, it can be known that if the Agent is detected to be started, the Task Cgroup subdirectory is initialized in each ECS instance, so that the subdirectory is used as a newly created second control group to restrict the resource usage of the user task pulled by the Agent, and enables the subdirectory to prepare an isolated environment for subsequent task pulling. Among them, the user task refers to a third-party service pulled by the Agent, such as a mini-program and other services.
[0084] Step 12: Run the user task in the cloud server instance according to the resource usage constraint of the second control group, so that the running process of the user task meets the resource limit described by the resource usage constraint.
[0085] Based on the relevant contents of steps 11 to 12 above, it can be known that in some scenarios, for any ECS instance, when starting the ECS instance, the Agent installed in the ECS instance is added to systemd, and the isolation parameters of the Agent (such as memory 512Mb, CPU1 / 5 cores, etc.) are set, so that the resource usage constraints of the Agent can be determined based on the isolation parameters and the resource usage constraints inherited from cgroup.system.slice. When the Agent is started, the Agent is run according to the resource usage constraints of the Agent, and the Task Cgroup subdirectory is initialized to prepare an isolated environment for subsequent task pulling. In this way, it is possible to ensure that the Agent's resources are limited and the user jobs pulled by the Agent are not restricted.
[0086] In addition, in order to better improve the effect, this application also provides a method for implementing user tasks pulled by Agent (such as Figure 4 The pull method shown in FIG. 1 may be as follows: in response to a run request triggered by a batch job client for the user task, a child process is first created in the first control group (eg, Figure 4 The Bash intermediate process shown in the figure) is used to execute the user task; after moving the subprocess to the second control group, the user task is run through the subprocess to ensure that the running process of the user task meets the resource usage constraints of the second control group.
[0087] It can be seen that for any task other than the Agent's own task, if it is detected that a run request is triggered for the task through the Agent, it can be determined that the user wants to start the task through the Agent, so a Bash process for executing the task can be created in the first control group first; then the Bash process can be moved from the first control group to the second control group in a certain way (such as echo bash>>new_path); and then the task can be started through the Bash process. Among them, after the Bash process is moved from the first control group to the second control group, the Bash process is not subject to the resource restrictions required by the first control group, but is subject to the resource restrictions required by the second control group, so that the tasks run by the Bash process are no longer subject to the resource restrictions of the Agent. This can effectively solve the defect that when resource isolation is achieved by means of Cgroup, the user job pulled up by the Agent is also subject to the resource restrictions of the Agent, thereby ensuring that the user job is completely isolated from the Agent at any time, and further ensuring that the running process of the user job will not be affected by the running state of the Agent, thus ensuring that the user job can still run normally when the Agent resource exceeds the limit and is closed.
[0088] It can be seen that in a possible implementation, for a user task launched through the Agent, after running the user task, in response to the resource usage of the batch job client exceeding the resource upper limit limited by the resource usage constraint of the first control group, the batch job client is closed, and the user task is kept in a running state, so that it can be ensured that the user job launched through the Agent is completely independent of the Agent, thereby ensuring that the job launched by the user through the Agent is not affected by the Agent itself.
[0089] In addition, in order to better improve the effect, the present application also provides an implementation method for user tasks pulled by Agent, which can be: in response to a run request triggered by a batch job client for the user task, first determine whether a second control group has been created for constraining the resource usage of the task pulled up by the batch job client; if the second control group has not been created, create a second control group in the above-mentioned system kernel, and create a subprocess in the first control group; if the second control group has been created, create a subprocess in the first control group, so that after the subprocess is moved to the second control group, the task is run through the subprocess, and the running process of the task meets the resource usage constraints of the second control group.
[0090] It can be seen that in a possible implementation, for a user job pulled by an Agent, the running process of the user job satisfies the following constraints: before the user job runs, check whether the current system has created a Cgroup dedicated to the user job (such as Figure 2 Cgroup5) as shown in the figure. If it does not exist, create it. Moreover, before the user job runs, you can move the Bash executed by the user job to the Cgroup through echo bash>>new_path, and then use the Bash to start the user job for running. This ensures that the user job is not subject to the resource restrictions of the Agent.
[0091] After research, it was found that for systemd, when systemd daemon-reloaded, the process may drift in the Cgroup, so that the above-mentioned Bash process may move from the second control group back to the first control group, affecting the running effect of the user's job.
[0092] Based on the above research, in order to overcome the problems shown in the previous paragraph, the cloud server-based task processing method provided in the present application may also include: after moving the above-mentioned child process (such as a Bash process) to the second control group, updating the configuration parameters of the above-mentioned system daemon process, and the updated configuration parameters are used to indicate that when the system daemon process configuration file is reloaded (systemd daemon-reload), the child process will not be moved from the second control group back to the first control group.
[0093] It can be seen that in a possible implementation mode, the present application prevents the process from drifting in the Cgroup during systemd daemon-reload by setting systemd parameters, so as to ensure that the Bash process does not move from the second control group back to the first control group, thereby ensuring that the operation of the user job at any time will not be affected by the Agent.
[0094] Based on the relevant contents of the above-mentioned cloud server-based task processing method, it can be known that the resource isolation solution based on Cgroup provided in this application has the advantages shown in (i) to (iv) below.
[0095] (I) The resource isolation solution provided by the present application ensures that the resources available to the Agent are limited to the Cgroup cgroup-system-slice in the ECS instance by adding the Agent process to systemd. In this way, some mechanisms of the system kernel in the ECS instance can effectively ensure that when the CPU, memory or other resources occupied by the Agent process itself exceed the limit, the Agent process is discovered and closed in time, and the situation where the Agent cannot be closed in time will not occur.
[0096] (ii) Since the resource isolation solution provided by the present application is implemented with the help of the Cgroup function of the Linux kernel itself, the resource usage monitoring process involved in the solution can be implemented with the help of the resource usage checking mechanism configured by the Linux kernel for this function, so that there is no need to use additional processes (such as the self-monitoring process of the program mentioned above), and thus the solution will not have additional performance overhead caused by periodically checking the resource usage of the Agent with the help of additional processes, which is conducive to better improving the performance of the ECS instance.
[0097] (III) The resource isolation solution provided by the present application is implemented with the help of some mechanisms of the Linux kernel itself (such as the Cgroup mechanism). Since these mechanisms have been verified in large-scale environments, the possibility of bugs in these mechanisms is much lower than the possibility of bugs in the program self-monitoring code independently developed by the cloud vendor, which is conducive to better improving the resource isolation effect.
[0098] (IV) Because the resource isolation solution provided by the present application can realize that the running process of the user job launched by the Agent is completely independent of the running process of the Agent, so that when the Agent is closed through the kernel state mechanism, the running of the user job will not be affected, which is conducive to improving the running effect of the user job.
[0099] Based on the cloud server-based task processing method provided in the embodiment of the present application, the embodiment of the present application also provides a cloud server-based task processing device. Figure 5 Explain and illustrate. Figure 5 This is a schematic diagram of a cloud server-based task processing device provided in an embodiment of the present application. It should be noted that for the technical details of the cloud server-based task processing device provided in an embodiment of the present application, please refer to the relevant content of the cloud server-based task processing method above.
[0100] like Figure 5 As shown, the cloud server-based task processing device 500 provided in the embodiment of the present application includes:
[0101] An adding unit 501 is used for adding a batch job client installed in the cloud server instance to a system daemon process of the cloud server instance in response to starting the cloud server instance, wherein the system daemon process is used for constraining resource usage of the batch job client by using a control group function of a system kernel in the cloud server instance;
[0102] The running unit 502 is configured to, in response to starting the batch job client, run the batch job client in the cloud server instance according to the constraints implemented by using the control group function.
[0103] In a possible implementation manner, the system kernel includes a first control group for constraining resource usage of the batch job client and a parent control group of the first control group, the resource usage constraint of the first control group is determined according to the resource usage constraint inherited from the parent control group and a resource isolation parameter configured for the batch job client, the resource isolation parameter is used to indicate an upper limit of resource usage of the batch job client;
[0104] The running unit 502 is specifically configured to run the batch job client in the cloud server instance according to the resource usage constraint of the first control group in response to starting the batch job client.
[0105] In one possible implementation, if the resource isolation parameter and the inherited resource usage constraint are both used to limit the upper limit of usage of the target resource, the upper limit of usage limited for the target resource by the resource isolation parameter shall not exceed the upper limit of usage limited for the target resource by the inherited resource usage constraint.
[0106] In a possible implementation manner, the system kernel includes a first control group for constraining resource usage of the batch job client; the cloud server-based task processing device 500 also includes:
[0107] a creating unit, configured to create, in response to starting the batch job client, in the system kernel a second control group for constraining resource usage of user tasks running through the batch job client, wherein the second control group is different from the first control group, and there is no parent-child relationship between the second control group and the first control group;
[0108] A processing unit is used to run the user task in the cloud server instance according to the resource usage constraint of the second control group.
[0109] In one possible implementation, the system kernel also includes a parent control group of the first control group and a parent control group of the second control group, the parent control group of the second control group is different from the parent control group of the first control group, and there is no parent-child relationship between the parent control group of the second control group and the parent control group of the first control group; the resource usage constraints of the second control group are determined based on the resource usage constraints inherited from the parent control group of the second control group and the resource usage constraints configured by the user task.
[0110] In a possible implementation manner, the parent control group of the second control group and the parent control group of the first control group are both child control groups of a root control group, and the parent control group of the root control group does not exist in the system kernel.
[0111] In one possible implementation, the processing unit is specifically used to: in response to a run request triggered by the batch job client for the user task, create a subprocess in the first control group, wherein the subprocess is used to execute the user task; after moving the subprocess to the second control group, run the user task through the subprocess, and the running process of the user task satisfies the resource usage constraints of the second control group.
[0112] In one possible implementation, the processing unit is further used to: in response to a run request triggered by the batch job client for the user task, determine whether a second control group has been created for constraining resource usage of the user task run by the batch job client; if the second control group has not been created, create the second control group in the system kernel.
[0113] In one possible implementation, the processing unit is also used to: after moving the sub-process to the second control group, update the configuration parameters of the system daemon process, and the updated configuration parameters are used to indicate that the sub-process will not be moved from the second control group back to the first control group when the configuration file of the system daemon process is reloaded.
[0114] In a possible implementation manner, the cloud server-based task processing device 500 further includes:
[0115] A closing unit is used to close the batch job client and keep the user task in a running state in response to the resource usage of the batch job client exceeding the resource upper limit limited by the resource usage constraint of the first control group after running the user task in the cloud server instance according to the resource usage constraint of the second control group.
[0116] Based on the above-mentioned content of the task processing device 500 based on the cloud server, it can be known that the working principle of the device 500 includes: for any ECS instance, when starting the ECS instance, the Agent installed in the ECS instance is added to the systemd of the ECS instance, so that the systemd can use the Cgroup function of the system kernel in the ECS instance to constrain the resource usage of the Agent, and make the system kernel include a first control group (such as assist.client.system) for constraining the resource usage of the Agent and a parent control group (such as cgroup.system.slice) of the first control group, so that the resource usage constraint of the first control group is based on the resource usage constraint inherited from the parent control group and the resource usage constraint configured for the Agent and used to instruct the Agent The resource isolation parameter of the resource usage upper limit is determined, so that when it is detected that the Agent is started, the Agent is run in the ECS instance according to the resource usage constraint of the first control group to ensure that the maximum resources used by the Agent during operation do not exceed the resource usage upper limit limited by the resource usage constraint of the first control group, such as the resource constraints described by the resource isolation parameters configured for itself and the resource constraints inherited from the parent control group. In this way, not only can the amount of resources used by the Agent be limited within a certain range, but it can also effectively ensure that the Agent can be closed in time when the resource usage of the Agent exceeds the limit, thereby effectively overcoming the defect caused by the inability to close the Agent in time when the resource usage of the Agent exceeds the limit, and thus helping to better improve the performance of the ECS instance.
[0117] In addition, an embodiment of the present application also provides an electronic device, which includes a processor and a memory: the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs in the memory, so that the electronic device performs any implementation of the cloud server-based task processing method provided in the embodiment of the present application.
[0118] See also Figure 6 , which shows a schematic diagram of the structure of an electronic device 600 suitable for implementing the embodiment of the present disclosure. The terminal device in the embodiment of the present disclosure may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0119] like Figure 6 As shown, the electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0120] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 6 The electronic device 600 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0121] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.
[0122] The electronic device provided by the embodiment of the present disclosure and the method provided by the above embodiment belong to the same inventive concept. The technical details not fully described in this embodiment can be referred to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0123] An embodiment of the present application also provides a computer-readable medium, in which instructions or computer programs are stored. When the instructions or computer programs are executed on a device, the device executes any implementation of the cloud server-based task processing method provided in the embodiment of the present application.
[0124] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0125] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (Hyper Text Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0126] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0127] The computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device can execute the method.
[0128] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or a combination thereof, including, but not limited to, object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0129] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0130] The units involved in the embodiments described in the present disclosure may be implemented by software or hardware, wherein the name of a unit / module does not, in some cases, constitute a limitation on the unit itself.
[0131] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0132] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0133] It should be noted that the various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the various embodiments can be referred to each other. For the system or device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description.
[0134] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0135] It should also be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0136] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0137] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A task processing method based on a cloud server, characterized in that: The method comprises: In response to starting a cloud server instance, adding a batch job client installed in the cloud server instance to a system daemon of the cloud server instance, the system daemon being used to constrain resource usage of the batch job client by using a control group function of a system kernel in the cloud server instance; In response to initiating the batch job client, the batch job client is run in the cloud server instance according to the constraints implemented using the control group function.
2. The method according to claim 1, characterized in that: The system kernel includes a first control group for constraining resource usage of the batch job client and a parent control group of the first control group, wherein the resource usage constraint of the first control group is determined according to the resource usage constraint inherited from the parent control group and a resource isolation parameter configured for the batch job client, wherein the resource isolation parameter is used to indicate an upper limit of resource usage of the batch job client; The running of the batch job client in the cloud server instance according to the constraints implemented by using the control group function includes: The batch job client is run in the cloud server instance according to the resource usage constraint of the first control group.
3. The method according to claim 1, characterized in that The system kernel includes a first control group for constraining resource usage of the batch job client; The method further comprises: In response to starting the batch job client, creating in the system kernel a second control group for constraining resource usage of user tasks running through the batch job client, wherein no parent-child relationship exists between the second control group and the first control group; The user task is run in the cloud server instance according to the resource usage constraint of the second control group.
4. The method according to claim 3, characterized in that The system kernel further includes a parent control group of the first control group and a parent control group of the second control group, and there is no parent-child relationship between the parent control group of the second control group and the parent control group of the first control group; The resource usage constraint of the second control group is determined according to the resource usage constraint inherited from the parent control group of the second control group and the resource usage constraint configured by the user task.
5. The method according to claim 4, characterized in that The parent control group of the second control group and the parent control group of the first control group are both child control groups of the root control group, and the parent control group of the root control group does not exist in the system kernel.
6. The method according to any one of claims 3 to 5, characterized in that: The running the user task in the cloud server instance according to the resource usage constraint of the second control group includes: In response to a running request triggered by the batch job client for the user task, creating a subprocess in the first control group, the subprocess being used to execute the user task; After the sub-process is moved to the second control group, the user task is run through the sub-process, and the running process of the user task satisfies the resource usage constraint of the second control group.
7. The method according to claim 6, characterized in that The method further comprises: In response to a run request triggered by the batch job client for the user task, determining whether a second control group for constraining resource usage of the user task run by the batch job client has been created; If the second control group is not created, the second control group is created in the system kernel.
8. The method according to claim 6, characterized in that The method further comprises: After moving the subprocess to the second control group, the configuration parameters of the system daemon are updated, and the updated configuration parameters are used to indicate that the subprocess will not be moved from the second control group back to the first control group when the configuration file of the system daemon is reloaded.
9. The method according to claim 3, characterized in that: After running the user task in the cloud server instance according to the resource usage constraint of the second control group, the method further includes: In response to the resource usage of the batch job client exceeding the resource upper limit limited by the resource usage constraint of the first control group, the batch job client is closed, and the user task is kept in a running state.
10. A task processing device based on a cloud server, characterized in that: include: an adding unit, configured to, in response to starting a cloud server instance, add a batch job client installed in the cloud server instance to a system daemon process of the cloud server instance, wherein the system daemon process is configured to constrain resource usage of the batch job client by using a control group function of a system kernel in the cloud server instance; The running unit is configured to, in response to starting the batch job client, run the batch job client in the cloud server instance according to the constraints implemented by using the control group function.
11. An electronic device, characterized in that: The device comprises: a processor and a memory; The memory is used to store instructions or computer programs; The processor is used to execute the instructions or computer programs in the memory so that the electronic device executes the method according to any one of claims 1 to 9.
12. A computer readable medium, characterized in that The computer-readable medium stores instructions or computer programs, and when the instructions or computer programs are executed on a device, the device executes the method according to any one of claims 1 to 9.
13. A computer program product, characterized in that It comprises a computer program carried on a non-transitory computer-readable medium, the computer program comprising a program code for executing the method according to any one of claims 1 to 9.