Branch Prediction Method and Device for Serverless Computing Based on Process Parasitism
By adding parasitic processes to the container image of serverless computing, pre-executing template functions and training branch predictors, the problem of low branch prediction accuracy in serverless computing is solved, and the function execution performance is improved.
Patent Information
- Application Number
- CN202111560316.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-18
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-12-18
AI Technical Summary
In serverless computing, the accuracy of branch predictors is low, resulting in a degradation of function execution performance. Existing solutions usually require redesigning the hardware branch predictors, affecting universality.
By adding parasitic processes to the serverless compute container image, pre-execute template functions and train branch predictors with system calls to improve branch prediction accuracy.
Improves branch prediction accuracy and improves the execution performance of functions in serverless computing. It is suitable for various architectures, including ARM and RISC-V.
Smart Images

Figure CN116266242B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of serverless computing, and particularly relates to a branch prediction method and device for serverless computing based on process parasitism, an electronic device, and a readable storage medium. Background Art
[0002] Serverless Computing refers to building and running applications without managing infrastructure such as servers. It describes a finer-grained deployment model in which an application is broken down into one or more fine-grained functions that are uploaded to a platform and then executed, scaled, and billed according to current requirements.
[0003] Serverless Computing does not mean that servers are no longer used to host and run code, nor does it mean that operations engineers are no longer needed. Instead, it means that consumers of serverless computing no longer need to perform server configuration, maintenance, updates, scaling, and capacity planning. These tasks and functions are handled by the serverless platform and are completely abstracted from developers and IT / operations teams. Therefore, developers can focus on writing the business logic of the application, and operations engineers can elevate their focus to more critical business tasks.
[0004] In computer architecture, a branch predictor is a digital circuit that guesses which branch will be executed before the execution of a branch instruction ends, in order to improve the performance of the instruction pipeline of the processor. The purpose of using a branch predictor is to improve the process of instruction pipelining.
[0005] A branch predictor requires a certain amount of training to achieve a relatively stable high prediction accuracy. Therefore, when a function in serverless computing is scheduled to a server, the branch prediction accuracy in the initial stage is usually very low. And the running time of functions in serverless computing is usually in milliseconds. High branch prediction errors usually result in more performance overhead, thus reducing the execution performance of functions in serverless computing.
[0006] Current solutions usually involve redesigning the branch predictor and the branch predictor algorithm. By expanding the perception range of the branch predictor and making full use of the principle of temporal locality, the overall branch prediction accuracy is improved. However, since the branch predictor is a hardware device, redesigning the branch predictor requires modification at the hardware level, which will reduce the generality of branch prediction. Summary of the Invention
[0007] The object of the present invention is to provide a branch prediction method and device, an electronic device and a readable storage medium for serverless computing based on process parasitism, which can improve the branch prediction accuracy rate and the execution performance of functions in serverless computing without changing the branch predictor hardware.
[0008] The present invention is implemented as follows:
[0009] To achieve the above object, the present invention provides a branch prediction method for serverless computing based on process parasitism, including the following steps:
[0010] Receive a call request from a user for a target function;
[0011] In the case of needing to expand capacity, schedule the container executing the target function to a new server that has not executed the target function in a short time; wherein a parasitic process is pre-added to the base image of the container.
[0012] Trigger the parasitic process when initializing the container on the new server, and the parasitic process is used to initiate a system call to trigger the system kernel to select a target template function according to the type of the target function and copy it N times.
[0013] Use the execution data of the N copied target template functions as training data to train the branch predictor on the new server.
[0014] Further, after receiving the call request from the user for the target function, it further includes:
[0015] Judge whether there is an instance that has not executed the function task running in the current computing environment;
[0016] If so, schedule the target function to the instance that has not executed the function task in the current computing environment, and the instance executes the computing task of the target function.
[0017] Further, the branch method for serverless computing based on process parasitism further includes:
[0018] If there is no instance that has not executed the function task running in the current computing environment, judge whether the current computing environment needs to expand capacity;
[0019] If it does not need to expand capacity, generate an instance in the current computing environment, and the instance executes the computing task of the target function.
[0020] Further, the judgment of whether the current computing environment needs to expand capacity includes:
[0021] Determine whether the CPU usage of all instances in the current computing environment exceeds a preset value. If so, determine that the current computing environment needs to be expanded.
[0022] Further, the branch method of serverless computing based on process parasitism further includes:
[0023] After initializing the container on the new server, generate an instance that executes the computing task of the target function.
[0024] Further, the type of the target function is inferred using a python deep learning algorithm.
[0025] Further, the target template function takes programming language, if-else logical structure, for loop position feature, and function feature as the core.
[0026] To achieve the above object, the present invention also provides a branch prediction device for serverless computing based on process parasitism, including:
[0027] A receiving module, configured to receive a call request from a user for a target function;
[0028] A scheduling module, configured to, in the case of needing to expand, schedule the container executing the target function to a new server that has not executed the target function within a short period of time; wherein a parasitic process is pre-added to the base image of the container;
[0029] A calling module, configured to trigger the parasitic process when initializing the container on the new server, and the parasitic process is used to initiate a system call to trigger the system kernel to select a target template function according to the type of the target function and copy it N times;
[0030] A training module, configured to use the execution data of the N copied target template functions as training data to train the branch predictor on the new server.
[0031] To achieve the above object, the present invention also provides an electronic device, including a processor and a memory, where a computer program is stored on the memory, and when the computer program is executed by the processor, the steps of the branch prediction method for serverless computing based on process parasitism described in any one of the above are implemented.
[0032] To achieve the above object, the present invention also provides a readable storage medium, where a computer program is stored in the readable storage medium, and when the computer program is executed by a processor, the steps of the branch prediction method for serverless computing based on process parasitism described in any one of the above are implemented.
[0033] Compared with the prior art, the present invention has the following beneficial effects:
[0034] 1. Compared with redesigning the branch predictor, the present invention has universality. By means of pre-executing template functions, the present invention can improve the branch prediction accuracy of all types of servers, enhance the execution performance of functions in serverless computing, and is applicable to all architectures (including ARM, RISC-V, etc.).
[0035] 2. Compared with the temporal locality of the branch prediction algorithm, the present invention pre-executes template functions, making full use of the temporal locality of the branch predictor.
[0036] Other features and advantages of the present invention will be described in the following specification, and in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained by the structures specifically pointed out in the written specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] To more clearly illustrate the technical solutions of the present invention, the drawings required for description will be briefly introduced below. Obviously, the drawings in the following description are an embodiment of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts:
[0038] Figure 1 is the overall design architecture diagram of the branch prediction method for serverless computing based on process parasitism provided by an embodiment of the present invention;
[0039] Figure 2 is the flowchart of the branch prediction method for serverless computing based on process parasitism provided by an embodiment of the present invention;
[0040] Figure 3 is the flowchart of the branch prediction method for serverless computing based on process parasitism in a specific example of the present invention;
[0041] Figure 4 is the structural diagram of the branch prediction device for serverless computing based on process parasitism provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The exemplary embodiments of the present invention will be described in more detail below with reference to the drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.
[0043] To solve the problems existing in the prior art, the present invention provides a branch prediction method, apparatus, electronic device and readable storage medium for serverless computing based on process parasitism.
[0044] Some concept explanations in serverless computing are as follows:
[0045] ① Serverless computing: Serverless computing is a method of providing backend services on demand. Serverless providers allow users to write and deploy code without worrying about the underlying infrastructure. Users who obtain backend services from serverless providers will pay according to the amount of computing and resource usage. Since this service is automatically scaled, there is no need to reserve and pay for a fixed amount of bandwidth or servers.
[0046] ② Container: A container contains an application and all the elements required for the application to run properly, including system libraries, system settings, and other dependencies. Any type of application can run in a container, and no matter where the containerized application is hosted, it will run in the same way. Similarly, a container can also carry serverless computing applications (i.e., functions) and then run on any server in the cloud platform.
[0047] ③ Instance: An instance refers to the runtime environment in which an application is running. For example, a container A running a certain service can be considered as an instance of this service. In principle, the functions of the serverless computing platform can be reduced to 0. Due to the automatic scaling of serverless computing, a large number of serverless function instances can be launched within a short time.
[0048] The basic idea of the present invention is as follows:
[0049] (1) Construct a serverless computing template function
[0050] The present invention designs a template function with programming languages, if-else logical structures, for loop position features, and function features as the core by researching mainstream serverless function workloads. The code volume of the template function is usually 20-30% of that of a normal function, without generating any network requests or disk operations, and the execution time is usually 5-10 ms. For example, if multiple functions use Python to perform deep learning inference and are of the same type of class functions, then there is one template function corresponding to multiple functions because their execution processes are basically the same: loading libraries, loading algorithm models, reading parameters, performing inference, and returning results.
[0051] (2) Design a pre-running process for parasitic containers
[0052] The present invention redesigns the basic container image and adds a pre-executed process in the basic image. The pre-executed process starts to execute at the beginning of the container startup, and calls the system call in advance to trigger the process of the kernel copying the template function.
[0053] (3) System call development of Fork template function
[0054] The present invention realizes the rapid copying of the specified template function by adding a system call in the system kernel. The system call passes which template function needs to be copied by means of parameters. For example, in Python deep learning, template functions include webtemplate, bigdatatemplate, MLtemplate, and Streamtemplate. The overall design architecture of the above process is as follows Figure 1 shown.
[0055] Please refer to Figure 2 The present invention provides a branch prediction method for serverless computing based on process parasitism, comprising the following steps:
[0056] Step S100, receiving a user's call request for a target function;
[0057] Step S200, when capacity expansion is required, scheduling the container executing the target function to a new server that has not executed the target function in a short period of time; wherein a parasitic process is pre-added to the base image of the container;
[0058] Step S300, triggering the parasitic process when initializing the container on the new server, the parasitic process is used to initiate a system call, triggering the system kernel to select a target template function according to the type of the target function and copy it N times;
[0059] Step S400: Using the copied N execution data of the target template functions as training data, the branch predictor on the new server is trained.
[0060] The following combination Figure 3 The above steps of the present invention are described in detail.
[0061] In step S100, the user initiates a call request for the target function through a client, and the client can make the request call in the form of a Web interface, a command line tool, a RESTful API, etc.
[0062] Before performing step S200, first determine whether there is an instance that has not executed a function task running in the current computing environment; if so, schedule the target function to the instance that has not executed the function task in the current computing environment, and this instance executes the computing task of the target function. It can be understood that if there is a function instance running in the environment, it means that the function is in the warm-up state at this time. Therefore, scheduling the target function task to these machines can improve the correctness of branch prediction. If not, then consider how to use the present invention to improve performance.
[0063] If there is no instance that has not executed a function task running in the current computing environment, then determine whether the current computing environment needs to be expanded; if it does not need to be expanded, then generate an instance in the current computing environment, and this instance executes the computing task of the target function. Specifically, determine whether expansion is needed according to whether the CPU usage of all instances in the current computing environment exceeds a preset value. For example, when the CPU usage of all instances exceeds, it is considered that the load is relatively large, so expansion is needed. If expansion is not needed, an instance can be directly generated in the current computing environment, and this instance is used to execute the computing task of the target function.
[0064] If expansion is needed, then perform step S200, and schedule the container that executes the target function to a new server that has not executed the target function in a short period of time (that is, schedule the container to the new server).
[0065] In step S300, since a parasitic process is pre-added to the base image of the container, thus, when initializing the container on the new server, the process buried in the container image (that is, the parasitic process) will be executed first, and the parasitic process will initiate a system call, triggering the system kernel to select a target template function according to the type of the target function and copy it N times. The type of the target function is inferred using a python deep learning algorithm. Since one function type corresponds to one template function, after determining the type of the target function, the corresponding target template function can be selected.
[0066] In step S400, the N copied target template functions are automatically executed, and the execution data can be used as training data to train the branch predictor on the new server.
[0067] It can be understood that when the container is scheduled to a new server, due to the branch predictor (a hardware design) being unfamiliar with this type of function, there will be more incorrect predictions. Therefore, in the present invention, the template function is executed in advance to let the branch predictor become familiar with this function and achieve a warm-up effect. Branch prediction generally occurs only in cases where there is code logic routing such as if-else. Therefore, as long as the template function also has this design result, the branch predictor can be made familiar with this logical structure in advance. After the same type of function is executed multiple times, the branch predictor will automatically become familiar with this function model and thus make accurate predictions. The specific training process of the branch predictor belongs to the category of branch predictor algorithm design and will not be elaborated here.
[0068] Furthermore, after the container is successfully initialized on the new server described in step S300, an instance is generated, and this instance executes the computing task of the target function. Since triggering the parasitic process, initiating the system call, and copying out N template processes are carried out during the container initialization process, and generating an instance to execute the computing task of the target function is carried out after the container is successfully initialized. Before the computing task of the target function is executed, the branch predictor has been trained with the execution data of N target template functions. That is, when the computing task of the target function is executed, the branch predictor already has a warm-up effect on the target function. Therefore, the accuracy of branch prediction can be improved, and further the execution performance of the function in serverless computing can be improved.
[0069] In summary, the present invention designs a template function based on function characteristics. When the container is initialized, a parasitic process is used to call the system call, and the system call quickly forks the template process, thereby improving the branch prediction accuracy through the template process and enhancing the execution performance of the function in serverless computing. The present invention has conducted sufficient experiments, and the results show that the present invention has improved the branch prediction accuracy by 49% and the overall throughput by 38%, indicating that the design scheme of the present invention is feasible.
[0070] Based on the same inventive concept, the present invention also provides a branch prediction device for serverless computing based on process parasitism, as Figure 4 shown, including:
[0071] A receiving module 100, configured to receive a user's call request for a target function;
[0072] A scheduling module 200, configured to, in the case of needing to expand capacity, schedule the container that executes the target function to a new server that has not executed the target function in a short period of time; wherein a parasitic process is pre-added to the base image of the container;
[0073] A calling module 300 is used to trigger the parasitic process when initializing the container on the new server. The parasitic process is used to initiate a system call, triggering the system kernel to select a target template function according to the type of the target function and copy it N times.
[0074] A training module 400 is used to use the execution data of the N copied target template functions as training data to train the branch predictor on the new server.
[0075] Optionally, the branch prediction device for serverless computing based on process parasitism further includes:
[0076] A first judgment module is used to judge whether there is an instance that has not executed the function task running in the current computing environment after the receiving module 100 receives a call request from the user for the target function; if so, trigger the first execution module.
[0077] The first execution module is used to schedule the target function to the instance that has not executed the function task in the current computing environment, and this instance executes the computing task of the target function.
[0078] Optionally, the branch prediction for serverless computing based on process parasitism further includes:
[0079] A second judgment module is used to judge whether the current computing environment needs to be expanded if there is no instance that has not executed the function task running in the current computing environment; if not, trigger the second execution module.
[0080] The second execution module is further used to generate an instance in the current computing environment, and this instance executes the computing task of the target function.
[0081] Optionally, when the second judgment module judges whether the current computing environment needs to be expanded, specifically:
[0082] Judge whether the CPU usage of all instances in the current computing environment exceeds a preset value. If so, determine that the current computing environment needs to be expanded.
[0083] Optionally, the branch prediction device for serverless computing based on process parasitism further includes:
[0084] A third execution module is used to generate an instance after initializing the container on the new server, and this instance executes the computing task of the target function.
[0085] Optionally, the type of the target function is inferred using a python deep learning algorithm.
[0086] Optionally, the target template function takes programming language, if-else logical structure, for loop position feature, and function feature as the core.
[0087] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, please refer to the partial description of the method embodiment.
[0088] Based on the same inventive concept, the present invention also provides an electronic device, including a processor and a memory. A computer program is stored in the memory, and when the processor executes the computer program, the steps of the branch prediction method for serverless computing based on process parasitism as described above are implemented.
[0089] In some embodiments, the processor may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor is generally used to control the overall operation of the electronic device. In this embodiment, the processor is used to run the program code stored in the memory or process data, such as running the program code of the branch prediction method for serverless computing based on process parasitism.
[0090] The memory includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory may be an internal storage unit of the electronic device, such as the hard disk or memory of the electronic device. In other embodiments, the memory may also be an external storage device of the electronic device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. Of course, the memory may also include both the internal storage unit and the external storage device of the electronic device. In this embodiment, the memory is generally used to store the operation methods and various application software installed in the electronic device, such as the program code of the branch prediction method for serverless computing based on process parasitism. In addition, the memory may also be used to temporarily store various data that have been output or will be output.
[0091] Based on the same inventive concept, the present invention also provides a readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of the branch prediction method for serverless computing based on process parasitism as described above are implemented.
[0092] In summary, a branch prediction method, device, electronic device and readable storage medium for serverless computing based on process parasitism provided by the present invention have the following advantages and positive effects:
[0093] 1. Compared with redesigning the branch predictor, the present invention has universality. By the method of pre-executing the template function, the present invention can improve the branch prediction accuracy of all types of servers, improve the execution performance of functions in serverless computing, and is applicable to all architectures (including ARM, RISC-V, etc.).
[0094] 2. Compared with the temporal locality of the branch prediction algorithm, the present invention pre-executes the template function and makes full use of the temporal locality of the branch predictor.
[0095] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.
[0096] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0097] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1The functions specified in one or more boxes.
[0098] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide for implementing in the process Figure 1 One process or more processes and / or boxes Figure 1 The steps of the functions specified in one or more boxes.
[0099] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.
Claims
1. A branch prediction method for serverless computing based on process parasitism, characterized in that, It includes the following steps: Receive a user's call request for a target function; In the case of needing to expand capacity, schedule the container executing the target function to a new server that has not executed the target function in a short time; wherein a parasitic process is pre-added to the base image of the container; Trigger the parasitic process when initializing the container on the new server, and the parasitic process is used to initiate a system call to trigger the system kernel to select a target template function according to the type of the target function and copy it N times; Use the execution data of the N copied target template functions as training data to train the branch predictor on the new server.
2. The branch prediction method for serverless computing based on process parasitism according to claim 1, wherein After receiving the user's call request for the target function, it further includes: Judge whether there is an instance running a function task that has not been executed in the current computing environment; If so, schedule the target function to the instance in the current computing environment that has not executed the function task, and this instance executes the computing task of the target function.
3. The branch prediction method for serverless computing based on process parasitism according to claim 2, wherein The method further includes: If there is no instance running a function task that has not been executed in the current computing environment, judge whether the current computing environment needs to expand capacity; If it does not need to expand capacity, generate an instance in the current computing environment, and this instance executes the computing task of the target function.
4. The branch prediction method for serverless computing based on process parasitism according to claim 3, wherein The judgment of whether the current computing environment needs to expand capacity includes: Judge whether the CPU usage of all instances in the current computing environment exceeds a preset value. If so, determine that the current computing environment needs to expand capacity.
5. The branch prediction method for serverless computing based on process parasitism according to claim 1, characterized in that, The method further includes: After initializing the container on the new server, generate an instance, and this instance executes the computing task of the target function.
6. The branch prediction method for serverless computing based on process parasitism according to claim 1, wherein The type of the target function is inferred using a python deep learning algorithm.
7. The branch prediction method for serverless computing based on process parasitism according to claim 1, wherein The target template function takes programming language, if-else logical structure, for loop position features, and function features as the core.
8. A branch prediction device for serverless computing based on process parasitism, characterized in that, It includes: A receiving module, used to receive a user's call request for a target function; A scheduling module, used to schedule the container executing the target function to a new server that has not executed the target function in a short time in the case of needing to expand capacity; wherein a parasitic process is pre-added to the base image of the container; A calling module, used to trigger the parasitic process when initializing the container on the new server, and the parasitic process is used to initiate a system call to trigger the system kernel to select a target template function according to the type of the target function and copy it N times; A training module, used to use the execution data of the N copied target template functions as training data to train the branch predictor on the new server.
9. An electronic device, characterized in that, It includes a processor and a memory, and a computer program is stored on the memory. When the computer program is executed by the processor, it implements the steps of the branch prediction method for serverless computing based on process parasitism according to any one of claims 1 to 7.
10. A readable storage medium, characterized in that, A computer program is stored in the readable storage medium. When the computer program is executed by a processor, it implements the steps of the branch prediction method for serverless computing based on process parasitism according to any one of claims 1 to 7.
Citation Information
Patent Citations
Request processing method and device
CN112860450A
Apparatus and method for efficient branch prediction using machine learning
WO2020247829A1