Cache management method and device, electronic equipment, storage medium and program product
By dividing the cache pool into independent cache partitions and dynamically adjusting the capacity in the Serverless computing platform, the cache contention problem is solved, and the request processing efficiency and the number of function instance calls are improved.
Patent Information
- Application Number
- CN202510906280.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-10-17
AI Technical Summary
In the Serverless computing platform, the existing cache pool uses a single cache pool, which leads to cache contention and reduces request processing efficiency.
The cache pool is divided into multiple independent cache partitions. Each cache partition stores the corresponding hot function instance. The cache partition capacity is dynamically adjusted by monitoring the cold start ratio to avoid cache contention between hot functions.
The number of function instance calls for hot functions is increased, cold start delay is reduced, and request processing efficiency is improved.
Smart Images

Figure CN120803713A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of cloud computing, and particularly relates to a cache management method and device, electronic equipment, storage medium and program product. BACKGROUND
[0002] In a Serverless computing architecture, an application is composed of multiple stateless functions, and the function instances are only generated and initialized when receiving a call request. When a cold start occurs, it takes hundreds of milliseconds to several seconds to generate and initialize the function instance, and the high latency of the cold start reduces the service experience of the interactive application.
[0003] In order to reduce the overhead of the cold start, the cache of the function instance is a method widely used in the existing Serverless computing platform, which aims to call the function instance from the hot start state. By keeping the function instance alive after the user request is completed, the subsequent request can reuse the cached function instance, achieving almost zero cold start delay.
[0004] However, the existing Serverless computing platform mainly adopts a monolithic cache pool, in which all functions within the process of each worker node share cache resources, which can cause cache contention and reduce request processing efficiency.
[0005] It should be noted that the information disclosed in the above background section is only used to strengthen the understanding of the background of the present disclosure, and therefore can include information that does not constitute prior art known to those of ordinary skill in the art. SUMMARY
[0006] The present disclosure provides a cache management method and device, electronic equipment, storage medium and program product, which at least partially overcome the problem of cache contention in the related art.
[0007] Other characteristics and advantages of the present disclosure will become apparent from the following detailed description, or will be learned by practice of the present disclosure.
[0008] According to one aspect of the present disclosure, a cache management method is provided, comprising: obtaining actual values of cold start ratios respectively corresponding to at least two cache partitions in a cache pool, wherein each cache partition is used to store function instances of a corresponding hot function, and the actual value of the cold start ratio is determined by the cache miss rate of the corresponding cache partition; and adjusting the cache partition capacity based on the actual values of the cold start ratios respectively corresponding to the at least two cache partitions and a target value of the cold start ratio.
[0009] In some possible embodiments, the adjusting the cache partition capacity based on the cold start ratio actual value and the cold start ratio target value corresponding to each of the at least two cache partitions comprises: calculating, for the cache partition, a cache miss margin value based on the cold start ratio actual value and the cold start ratio target value, wherein the cache miss margin value is positively correlated with the difference between the cold start ratio target value and the cold start ratio actual value; and adjusting the cache partition capacity based on the cache miss margin value of the cache partition.
[0010] In some possible embodiments, the adjusting the cache partition capacity based on the cache miss margin value of the cache partition comprises: increasing the cache partition capacity when the cache miss margin value of the cache partition is less than or equal to a first threshold.
[0011] In some possible embodiments, the adjusting the cache partition capacity based on the cache miss margin value of the cache partition further comprises: decreasing the cache partition capacity when the cache miss margin value of the cache partition is greater than a second threshold, wherein the second threshold is greater than the first threshold.
[0012] In some possible embodiments, the method further comprises: determining a maximum cache miss margin value and a minimum cache miss margin value from the cache miss margin values of the at least two cache partitions; and the decreasing the cache partition capacity when the cache miss margin value of the cache partition is greater than the second threshold comprises: if the minimum cache miss margin value is greater than the first threshold and the maximum cache miss margin value is greater than the second threshold, decreasing the cache partition capacity corresponding to the maximum cache miss margin value.
[0013] In some possible embodiments, the method further comprises: starting a timer when the cache miss margin value of any one of the cache partitions is less than or equal to the first threshold; resetting the timer when the cache miss margin values of a preset number of cache partitions are all greater than the first threshold after the adjusting the cache partition capacity; and increasing the capacity of the cache pool and / or migrating the function instance of the hot function if the timer exceeds a third threshold.
[0014] In some possible embodiments, the method further comprises: constructing the cache pool; and dividing the cache pool into a plurality of independent cache partitions, wherein the initial capacities of each of the cache partitions are the same.
[0015] In some possible embodiments, the method further comprises: receiving a user request; starting a new function instance if the cache pool does not have a function instance corresponding to the user request; invoking the new function instance to process the user request; and storing the new function instance in a corresponding cache partition if the function corresponding to the user request is a hot function after the invocation of the new function instance ends.
[0016] According to another aspect of the present disclosure, there is also provided a cache management apparatus, comprising: an obtaining module configured to obtain actual cold start ratio values corresponding to at least two cache partitions in a cache pool respectively, wherein each cache partition is configured to store function instances of a corresponding hotspot function, and the actual cold start ratio value is determined by a cache miss ratio of the corresponding cache partition; and an adjusting module configured to adjust cache partition capacities based on the actual cold start ratio values and target cold start ratio values corresponding to the at least two cache partitions respectively.
[0017] According to another aspect of the present disclosure, there is also provided an electronic device, comprising: a processor; and a memory configured to store executable instructions of the processor; wherein the processor is configured to perform the cache management method of any one of the above via execution of the executable instructions.
[0018] According to another aspect of the present disclosure, there is also provided a computer readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the cache management method of any one of the above.
[0019] According to another aspect of the present disclosure, there is also provided a computer program product, comprising: a computer program or instructions, the computer program or instructions being executed by a processor to implement the cache management method of any one of the above.
[0020] The cache management method provided in the embodiments of the present disclosure first obtains actual cold start ratio values corresponding to at least two cache partitions in a cache pool respectively, wherein each cache partition is configured to store function instances of a corresponding hotspot function, and the actual cold start ratio value is determined by a cache miss ratio of the corresponding cache partition; and then adjusts cache partition capacities based on the actual cold start ratio values and target cold start ratio values corresponding to the at least two cache partitions respectively. In the embodiments, function instances of different hotspot functions are respectively cached in different cache partitions to avoid cache contention among the hotspot functions, and the capacities of the cache partitions are dynamically adjusted by the cold start ratio of the cache partitions to further reduce the cache contention among the hotspot functions, improve the number of function instance invocations of the hotspot functions, and improve the request processing efficiency.
[0021] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0022] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments consistent with the present disclosure and serve to explain the principles of the present disclosure. It is apparent that the accompanying drawings described below are only some embodiments of the present disclosure, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0023] Figure 1 A flow chart illustrating a cache management method in an embodiment of the present disclosure is shown;
[0024] Figure 2 A framework diagram illustrating a cache management based on a serverless computing platform in an embodiment of the present disclosure is shown;
[0025] Figure 3 A flow chart illustrating another cache management method in an embodiment of the present disclosure is shown;
[0026] Figure 4 A flow chart illustrating a cache partition capacity adjustment method in an embodiment of the present disclosure is shown;
[0027] Figure 5 A flow chart illustrating another cache partition capacity adjustment method in an embodiment of the present disclosure is shown;
[0028] Figure 6 A flow chart illustrating a global cache capacity management method in an embodiment of the present disclosure is shown;
[0029] Figure 7 A comparison chart of overall performance under Trace A workload in an embodiment of the present disclosure is shown;
[0030] Figure 8 A comparison chart of overall performance under Trace B workload in an embodiment of the present disclosure is shown;
[0031] Figure 9 A comparison chart of cold start variance in an embodiment of the present disclosure is shown;
[0032] Figure 10 A schematic diagram of a cache management apparatus in an embodiment of the present disclosure is shown;
[0033] Figure 11 A structural block diagram of an electronic device in an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0034] Example implementations will now be described more fully with reference to the accompanying drawings. Example implementations may, however, be implemented in many different forms and should not be construed as limited to the examples set forth herein; rather, these implementations are provided so that this disclosure will be thorough and complete, and will fully convey the scope of example implementations to those skilled in the art. The described features, structures, or characteristics can be combined in one or more implementations.
[0035] In addition, the accompanying drawings are only schematic and are non-limiting illustrative of the present disclosure. Like references denote like elements throughout the accompanying drawings. It will be evident to those skilled in the art that the present disclosure can be practiced without parts which are not specifically shown in the drawings. Particularity, some of the blocks in the drawings can be functional blocks that represent functions implemented by software, hardware, or a combination of software and hardware. The functional blocks can be implemented in software, hardware, or a combination of software and hardware.
[0036] For the purpose of promoting the understanding and facilitating appreciation of the present disclosure, the following terms will first be explained:
[0037] Serverless computing: also known as serverless, is a new cloud computing application development paradigm based on stateless functional programming, which has a series of advantages such as user-friendly programming, free resource operation and maintenance, high elasticity, and on-demand billing. For example, the common FaaS (Function as a Service) is a typical Serverless computing scenario.
[0038] Cold start: the Serverless computing platform takes the function instance as the minimum business unit, which can process the corresponding user request and return the result. The function instance usually runs in a container and only serves one user request at the same time. If there is no available function instance when the user request arrives, a new function instance needs to be created in the Serverless computing platform. It takes hundreds of milliseconds to several seconds to create a new function instance and call it, which brings a long cold start delay problem to the business.
[0039] Function cache: in the Serverless computing platform, when the function instance completes the last call, it is saved in the memory for a period of time instead of being released immediately, so as to reduce the occurrence of subsequent request cold start calls.
[0040] The specific implementation of the embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0041] Figure 1 A flow chart of a cache management method in the embodiments of the present disclosure is shown in FIG. 1. Figure 1 As shown in FIG. 1, the cache management method provided in the embodiments of the present disclosure includes the following steps S102-S104.
[0042] The cache management method provided in the embodiments of the present disclosure is applied to the Serverless computing platform, and a system architecture of the Serverless computing platform is provided. As shown in FIG. 2, the system architecture of the Serverless computing platform includes a function cache module 202, a function instance module 204, a function request module 206, and a function result module 208. Figure 2As shown, the Serverless computing platform 200 includes a function scheduler 210 and a plurality of worker nodes 220, each of which includes a request router 221, a cache pool 222, and a cache pool manager 223, and each of the cache pools 222 can include a plurality of cache partitions 2221.
[0043] The function scheduler 210 is a traffic scheduling center of the Serverless computing platform, which is configured to determine the worker node corresponding to a user request and assign the user request to the corresponding worker node to avoid overloading of a single worker node through load balancing; in combination with a hotspot awareness method, the function scheduler 210 is configured to identify the user request corresponding to a hotspot function and preferentially schedule the user request corresponding to the hotspot function to a worker node with sufficient resources to reduce the cold start impact.
[0044] The worker node 220 is a runtime carrier of a function in the Serverless computing platform, which is configured to carry an execution environment of a function instance, receive a scheduling request of the function scheduler 210, run code, and process a user request.
[0045] The request router 221 is a traffic distributor in the worker node 220, which is configured to distribute a user request received by the worker node 220 to a cache partition in which a function instance corresponding to the user request is located, to ensure accurate traffic hitting and realize accurate matching between the user request and the function instance.
[0046] The cache pool manager 223 is configured to manage the capacity of the cache partition, dynamically adjust the capacity of the cache partition according to monitoring data, directly solve the cache contention problem of the function instance, make the cache pool adapt to the actual load, and improve the resource utilization rate.
[0047] The cache pool 222 is a storage area set in the worker node 220 for temporarily storing the function instance of the hotspot function and the like, which is configured to store the function instance of the hotspot function in the cache pool, when a new user request arrives, preferentially call the stored function instance from the cache pool, without creating a new function instance, i.e., skip the code loading, dependency initialization and other processes, directly execute the logic, and reduce the cold start probability.
[0048] In a possible implementation, the cache management method provided by the embodiment can be applied to a Serverless computing service scenario on the side of a cloud computing vendor. The cache management method can be deployed as an independent control plane component in each worker node of a Serverless cluster, for example, the cache management method can be executed by the cache pool manager 223 in the worker node 220. Since the components in each worker node run independently of each other, the scope of action of the components is limited to function cache control and function instance lifecycle management in the worker node where the components are located. In an implementation, the cache management method can be deployed as a module in a function scheduler in a Serverless computing platform.
[0049] In the embodiment, the cache management method is deployed as an independent control plane component in each worker node of a Serverless cluster.
[0050] S102, obtain actual values of cold start ratios of at least two cache partitions in a cache pool, wherein each cache partition is used to store function instances of a corresponding hot function, and the actual value of the cold start ratio is determined by a cache miss rate of the corresponding cache partition.
[0051] The cache pool can be understood as a storage area in a worker node of a Serverless computing platform, used to temporarily store high-frequency access data, function instances, and the like. By caching function instances of hot functions, the time consumption of function cold start is reduced.
[0052] The cache partition definition can be understood as dividing the cache pool into multiple independent storage units, and each cache partition is used to store function instances of a corresponding hot function. One cache partition stores one or more function instances of one hot function. For example, getProductDetail() is a function that receives a product ID and returns product details. A separate cache partition 1 is created for getProductDetail(), and the cache partition 1 stores multiple function instances of the getProductDetail() function, for example, 10 function instances. sendNotification() is a function that pushes new message notifications to users. A separate cache partition 2 is created for sendNotification(), and the cache partition 2 stores multiple function instances of the sendNotification() function, for example, 20 function instances.
[0053] Through the cache partition, the cache data of different hot functions is realized without interference, and the cache efficiency is improved.
[0054] In a serverless computing platform, a hotspot function can be understood as a function with a much higher invocation frequency than other functions due to a surge in request volume, critical business logic, or concentrated data access. Hotspot functions are characterized by a large number of concurrent requests within a specific time period. For example, product detail query functions during promotional activities, order payment functions, popular content push functions during social traffic peaks, and like functions.
[0055] In a serverless computing platform, a function instance is the smallest resource unit that carries out function code execution. Each function instance is an independent running environment, i.e., each function instance includes independent CPU (Central Processing Unit), memory, network stack, and other resources. Life cycle management is automatically created, scheduled, and destroyed by the serverless computing platform, and users do not need to manually manage servers. Function instances have stateless characteristics, i.e., data is not shared between function instances, and the state of a function instance is not retained after each user request processing is completed.
[0056] Cache miss rate can be understood as the proportion of function instances required by user requests that are not found in their corresponding cache partitions. In other words, cache miss rate is the ratio of the number of times user requests miss cache instances to the total number of user requests. Cache miss rate is positively correlated with cold start rate, i.e., the higher the cache miss rate, the higher the cold start rate, and the lower the cache miss rate, the lower the cold start rate. In one possible implementation, the cache miss rate is directly used as the actual value of the cold start rate.
[0057] The actual value of the cold start rate can be understood as the real data of the cache partition cold start rate collected by monitoring the system, reflecting the current working node running state, derived from the cache miss rate. The actual value of the cold start rate can reflect the optimization effect of the cache mechanism on the function start efficiency. The lower the cold start rate, the fewer the function cold starts, and the higher the function processing efficiency.
[0058] In one possible implementation, for each cache partition corresponding to a hotspot function, the total number of user requests received by the cache partition and the number of cold starts of user requests are periodically counted. The ratio of the number of cold starts of user requests to the total number of user requests is used as the actual value of the cold start rate corresponding to the cache partition. The counting period can be set according to actual conditions, for example, the actual value of the cold start rate corresponding to each cache partition is counted every 5 minutes.
[0059] S104, based on the actual value of the cold start rate corresponding to at least two cache partitions and the target value of the cold start rate, adjusting the cache partition capacity.
[0060] The cold start ratio target value can be understood as an expected threshold of the cold start ratio set according to the business requirements, which is usually determined by the SLA (Service-Level Agreement). For example, an e-commerce platform requires that the cold start ratio of the payment function does not exceed 1%, i.e., the target value is 1%.
[0061] The cache partition capacity can be understood as the maximum resource amount that the cache partition can carry, including memory capacity, instance capacity, data storage amount, etc. The memory capacity can be understood as the amount of data that can be cached, for example: each cache partition is allocated 2GB of memory; the instance capacity can be understood as the number of requests that the cache partition can handle at the same time, for example: each cache partition supports a maximum of 100 function instance concurrent processing. In this embodiment, the cache partition capacity is taken as an example to illustrate the memory capacity.
[0062] Adjusting the cache partition capacity can be understood as increasing or decreasing the capacity of each cache partition, so that the actual value of the cold start ratio corresponding to the cache partition is less than the target value of the cold start ratio.
[0063] In one possible implementation, by comparing the difference between the actual value and the target value of the cold start ratio of each cache partition, different levels of capacity adjustment are triggered.
[0064] Specifically, for each cache partition, when the difference between the actual value and the target value of the cold start ratio of the current cache partition is less than or equal to 5% of the target value, increase the capacity of the current cache partition by 10%, for example: the memory capacity of the current cache partition is increased from 1GB to 1.1GB. When the difference between the actual value and the target value of the cold start ratio of the current cache partition is greater than 5% of the target value and less than 20% of the target value, the capacity of the current cache partition is kept unchanged. When the difference between the actual value and the target value of the cold start ratio of the current cache partition is greater than 20% of the target value, reduce the capacity of the current cache partition by 5%, for example: the memory capacity of the current cache partition is increased from 1GB to 0.95GB.
[0065] In one possible implementation, the capacity adjustment model is trained through historical data to predict the future trend of the cold start ratio and to plan the capacity in advance.
[0066] The multi-dimensional data such as the cold start rate, the total amount of user requests, and a time stamp in the past period of time is collected, and an LSTM (Long Short-Term Memory) model is used to capture the time sequence characteristics of the cold start rate. The input data of the capacity adjustment model can include the cold start rate, the total amount of user requests, and a time stamp in a period of time, and the output data can include the cold start rate prediction value of the next 5 minutes. According to the difference between the cold start rate prediction value and the target value, different levels of capacity adjustment are triggered. Alternatively, the output data can also include the capacity adjustment of each cache partition, for example, the cache partition 1 is adjusted to 1.1 GB, the cache partition 2 remains unchanged, and the cache partition 3 is adjusted to 1.15 GB.
[0067] In the embodiment, the function instances of different hot functions are respectively cached in different cache partitions to avoid cache contention between the hot functions, and the capacity of each cache partition is dynamically adjusted through the cold start rate of the cache partition, further reducing the cache contention between the hot functions, improving the number of calls of the function instances of the hot functions, and improving the request processing efficiency.
[0068] On the basis of the above-mentioned embodiments, the cache management method is optimized in the embodiments of the present application, as shown in Figure 3 The optimized cache management method includes steps S302-S320.
[0069] S302, a cache pool is constructed.
[0070] The cache pool can be understood as a shared resource pool for storing reusable function instances in a Serverless computing platform, which is divided from the memory of a worker node to avoid reloading each time the function instance is called.
[0071] In one possible implementation, constructing the cache pool includes: creating a shared cache space as the cache pool through memory allocation, resource preloading, and the like.
[0072] S304, the cache pool is divided into a plurality of independent cache partitions, wherein the initial capacity of each cache partition is the same.
[0073] The cache partition can be understood as a unified cache pool divided into independent logical units, and each independent logical unit can be independently managed in terms of resource allocation, life cycle and cache strategy. For example, each hot function corresponds to a cache partition, for example: getProductDetail() is a function that receives a product ID and returns product details, and a separate cache partition 1 is created for getProductDetail(), and the cache partition 1 stores multiple function instances of the getProductDetail() function, for example: 10 function instances. sendNotification() is a function that pushes a new message notification to the user, and a separate cache partition 2 is created for sendNotification(), and the cache partition 2 stores multiple function instances of the sendNotification() function, for example: 20 function instances.
[0074] The same initial capacity of each cache partition can be understood as allocating the same basic initial capacity to each cache partition when the cache partition is created.
[0075] In one possible implementation, a mapping relationship between hot functions and cache partitions is established, that is, each hot function corresponds to a cache partition, and in the initialization stage, the same memory space is allocated to each cache partition, for example: the initial capacity of each cache partition is 512MB.
[0076] In one possible implementation, the container (Docker) or lightweight virtualization technology is used to implement the physically isolated cache partition. Specifically, each cache partition runs in an independent container, and the memory and CPU usage are limited by resource quota. When creating the container, the number of CPU cores, memory size and storage volume space are uniformly allocated.
[0077] In this embodiment, in the cache partition initialization stage, the same initial capacity is allocated to each cache partition to achieve fair resource allocation of each cache partition in the initialization stage and avoid load imbalance caused by initial configuration differences.
[0078] S306, receiving a user request.
[0079] The user request can be understood as an operation instruction sent by the user to the Serverless computing platform, and the user request includes specific business requirements, such as querying data, submitting a form, etc. In the Serverless scenario, the user request is delivered to the Serverless computing platform through the API (Application Programming Interface, application programming interface) gateway, message queue and other entry components, triggering the Serverless computing platform function to perform corresponding tasks.
[0080] In a possible implementation, when the cache management method is deployed in the function scheduler, the function scheduler receives a user request sent by a user end through an API gateway, a message queue or the like.
[0081] In a possible implementation, when the cache management method is deployed in the worker node, after the function scheduler receives a user request sent by a user end through an API gateway, a message queue or the like, the function scheduler determines a worker node for processing the user request, and sends the user request to a request router of the corresponding worker node. In other words, the request router of the worker node receives a user request sent by the function scheduler, and the user request is sent by a user end to the function scheduler.
[0082] S308, determining whether a function instance corresponding to the user request exists in the cache pool, if yes, performing S310, and if not, performing S312.
[0083] The function instance corresponding to the user request can be understood as a function instance capable of processing the user request.
[0084] In a possible implementation, the hot function is determined and the hot function list is constructed through historical load data, and the hot function list includes a plurality of hot functions. The hot function can be determined through a function call frequency. If the number of calls of a function in a time window in the historical load data accounts for more than a preset threshold in the total number of calls, the function is marked as a hot function. The functions in the hot function list can be dynamically adjusted in real time according to actual conditions. In this embodiment, only the determination method of the hot function is exemplarily described but not limited.
[0085] After receiving the user request, the target function name is extracted from the user request, the target function name is queried in the hot function list, it is determined whether the target function corresponding to the user request is a hot function, if the target function name does not exist in the hot function list, it is indicated that the cache pool does not store the function instance of the target function corresponding to the user request, that is, the cache pool does not exist the function instance corresponding to the user request, S312 is executed, and a new function instance is created.
[0086] If the target function name exists in the hotspot function list, the user request corresponding to the target function is a hotspot function, indicating that the cache pool stores the function instance of the hotspot function corresponding to the user request. Based on the mapping relationship of the target function name in the hotspot function and the cache partition, the target cache partition corresponding to the hotspot function is queried. It is queried in the target cache partition whether there is an available function instance. If there is, it indicates that there is a function instance corresponding to the user request in the cache pool, and S310 is executed to call the function instance corresponding to the user request in the cache partition of the cache pool. If not, it indicates that there is no function instance corresponding to the user request in the cache pool, and S312 is executed to create a new function instance.
[0087] Querying whether there is an available function instance in the target cache partition includes: querying the function instance list of the cache partition to find the function instance in the "idle" state. If there is a function instance in the "idle" state in the function instance list, there is an available function instance in the target cache partition, and the function instance in the "idle" state is taken as the target function instance corresponding to the user request. Further, the function instance list is sorted by an LRU (Least Recently Used) strategy, and the most recently active function instance is preferentially reused as the target function instance corresponding to the user request.
[0088] S310, calling the function instance corresponding to the user request in the cache pool.
[0089] The function instance corresponding to the user request in the cache pool is called to process the user request, the parameters of the user request are passed to the corresponding function instance, the initialized resources of the function instance are reused, and the function logic is executed. For example: the initialized resources include database connections, dependent libraries, etc.
[0090] After the function instance processes the user request, it is marked as an "idle" state and re-joins the function instance list to wait for the next call.
[0091] S312, if there is no function instance corresponding to the user request in the cache pool, a new function instance is started.
[0092] The absence of the corresponding function instance in the cache pool can be understood as that no function instance in the "idle" state is found in the corresponding cache partition according to the target function name of the user request. This includes: the target function is a non-hotspot function, and no corresponding cache partition is allocated; all instances in the function partition corresponding to the hotspot function are in a non-idle state.
[0093] When the cache pool does not have a function instance corresponding to the user request, the worker node creates a new function instance, including but not limited to: loading code, initializing environment, establishing external connection and other steps, that is, triggering the cold start process.
[0094] S314, invoke a new function instance to process the user request.
[0095] The invoking of the new function instance to process the user request can include: after the new function instance is created, receiving parameters of the user request and executing function logic. The specific processing logic is the same as that of invoking the function instance stored in the cache pool that can be reused, and the description in the above embodiments can be referred to.
[0096] After the new function instance processes the user request, if the function corresponding to the user request is a non-hot function, the new function instance is directly destroyed to release the memory resources of the cache pool.
[0097] S316, after the invocation of the new function instance ends, if the function corresponding to the user request is a hot function, the new function instance is stored in the corresponding cache partition.
[0098] In one possible implementation, after the invocation of the new function instance ends, it is queried whether the target function name corresponding to the user request is in the hot function list; if it is a hot function, the cache partition corresponding to the hot function is queried according to the mapping relationship between the hot function and the cache partition, the new function instance is added to the function instance list of the corresponding cache partition, and the LRU strategy is sorted.
[0099] In one possible implementation, if it is determined in the processing process of S308 that the function corresponding to the user request is a hot function, and since there is no available function instance in the cache partition corresponding to the hot function, a new function instance is directly created in the cache partition, and after the new function instance processes the user request, it is directly saved and not destroyed. In this way, the secondary query of the hot function and the cache partition can be avoided.
[0100] In this embodiment, after the invocation of the new function instance ends, if it belongs to a hot function, it is directly stored in the corresponding cache partition, and when subsequent similar requests arrive, the function instance can be directly reused from the cache partition, skipping the code loading, environment initialization and other steps of cold start, reducing the number of cold starts and improving the processing efficiency.
[0101] S318, obtaining actual values of cold start ratios respectively corresponding to at least two cache partitions in the cache pool, wherein each cache partition is used to store function instances of a hot function corresponding thereto, and the actual value of the cold start ratio is determined by the cache miss ratio of the corresponding cache partition.
[0102] S320, for the cache partition, calculating a cache miss margin value based on the actual value of the cold start ratio and the target value of the cold start ratio, wherein the cache miss margin value and the difference between the target value of the cold start ratio and the actual value of the cold start ratio are positively correlated.
[0103] The cache miss slack value can be understood as a quantitative index for measuring the difference between the actual value of the cold start ratio and the target value, and is positively correlated with the difference between the two. The greater the difference between the target value and the actual value of the cold start ratio, the greater the cache miss slack value, indicating that the cold start ratio of the cache partition is greater.
[0104] In one possible implementation, the cache miss slack value is the ratio of the difference between the target value of the cold start ratio and the actual value of the cold start ratio to the target value of the cold start ratio. That is, the cache miss slack value is the ratio of the first difference to the target value of the cold start ratio, and the first difference is the difference between the target value of the cold start ratio and the actual value of the cold start ratio.
[0105] In one possible implementation, the cache miss slack value is calculated by formula (1).
[0106]
[0107] Wherein, slack represents the cache miss slack value, slo_target represents the target value of the cold start ratio, and cache_miss represents the actual value of the cold start ratio.
[0108] S322, adjusting the cache partition capacity based on the cache miss slack value of the cache partition.
[0109] In one possible implementation, for each cache partition, when the cache miss slack value of the current cache partition is less than or equal to a first threshold, the capacity of the current cache partition is increased by 10%, for example: the memory capacity of the current cache partition is increased from 1GB to 1.1GB. When the cache miss slack value of the current cache partition is greater than a second threshold, the capacity of the current cache partition is reduced by 5%, for example: the memory capacity of the current cache partition is increased from 1GB to 0.95GB.
[0110] In this embodiment, by calculating the cache miss slack value, the quantification of the cold start ratio deviation is realized, and according to the real-time cold start data adjustment, the waste of resources caused by blind expansion or the increase of the cold start ratio caused by blind contraction is avoided, and the balance between performance (reduction of cold start ratio) and cost (reduction of memory consumption) is realized.
[0111] In one possible implementation, a specific application example is provided, as shown in Figure 2 The flow of function scheduling and cache management in the Serverless computing platform can be divided into five links: user request access, scheduling decision, request routing, performance monitoring, and partition capacity optimization, which are explained in detail as follows.
[0112] 1. User request access
[0113] The user terminal sends a user request to the function scheduler 210, and the user request is used to request a function call.
[0114] 2. Dispatching strategy execution
[0115] The function dispatcher 210 generates a function dispatching request and delivers it to the worker node 220 by using a Hash load balancing or hotspot-aware dispatching strategy.
[0116] The Hash load balancing can be understood as performing Hash calculation according to a user request parameter, and uniformly distributing to different worker nodes to avoid single node overload.
[0117] The hotspot-aware dispatching can be understood as identifying high-frequency requests, and preferentially dispatching to a worker node with sufficient cache to reduce cold start.
[0118] 3. Request routing and instance calling
[0119] After receiving the function dispatching request, the worker node 220 is allocated to a specific cache partition by the request router 221, and a function instance in the cache partition is called to process the request.
[0120] Each cache partition corresponds to a specific hotspot function, and the request router 221 accurately matches the demand and reuses the cache instance to avoid cold start.
[0121] 4. Performance monitoring feedback
[0122] The performance monitor 224 collects processing data of user requests in real time, such as the number of cold starts and the total number of user requests, and reports them to the cache pool manager 223.
[0123] 5. Dynamic optimization of partition capacity
[0124] The cache pool manager 223 calculates the cache miss margin value based on the performance monitoring data, and adjusts the capacity of the cache partition.
[0125] If the cold start is frequent, the partition capacity is increased to improve the reuse rate. If the cold start is rare, the partition capacity is reduced to release resources to other hotspot functions.
[0126] The entire process forms a closed loop of "request access, dispatching, routing, monitoring, capacity optimization, and re-dispatching", which reduces the cold start rate through hotspot function preferential dispatching and cache partition reuse, and avoids resource waste through dynamic adjustment of partition capacity, thereby achieving higher resource utilization.
[0127] On the basis of the above embodiment, the cache partition capacity adjustment method is optimized, as shown in Figure 4 The optimized cache partition capacity adjustment method includes steps S402-S414.
[0128] S402, acquire actual values of cold start ratios respectively corresponding to at least two cache partitions in a cache pool, wherein each cache partition is used to store function instances of a corresponding hot function, and the actual value of the cold start ratio is determined by the cache miss ratio of the corresponding cache partition.
[0129] S404, for each cache partition, calculate a cache miss margin value based on the actual value of the cold start ratio and a target value of the cold start ratio, wherein the cache miss margin value and the difference between the target value of the cold start ratio and the actual value of the cold start ratio are positively correlated.
[0130] S406, determine whether the cache miss margin value is less than or equal to a first threshold value, if yes, execute S408, if no, execute S410.
[0131] The first threshold value can be understood as a pre-set capacity expansion judgment threshold value. When the cache miss margin value is less than or equal to the first threshold value, it indicates that the actual value of the cold start ratio is large, i.e. the number of cold starts is large, and optimization is performed by increasing the cache partition capacity. For example, the first threshold value is 0.05.
[0132] S408, increase the cache partition capacity when the cache miss margin value of the cache partition is less than or equal to the first threshold value.
[0133] Increasing the cache partition capacity can include increasing the cache partition capacity by a set percentage or increasing the partition capacity by a set capacity. For example, the capacity of the cache partition is increased to 10% of the current capacity value, or the current capacity value of the cache partition is increased by 0.1 GB.
[0134] The calculated cache miss margin value is compared with the pre-set first threshold value to determine whether the expansion condition is met. When the cache miss margin value is less than or equal to the first threshold value, the expansion operation is performed, i.e. the capacity of the cache partition is increased.
[0135] It should be noted that if the cache miss margin values of multiple cache partitions are all less than or equal to the first threshold value, the capacity of each cache partition is increased.
[0136] In this embodiment, the available capacity of the cache partition is expanded to increase the number of function instances that can be stored in the cache partition, improve the cache hit ratio, and reduce the number of cold starts.
[0137] S410, determine whether the cache miss margin value is greater than a second threshold value, if yes, execute S412, if no, execute S414.
[0138] The second threshold value can be understood as a pre-set capacity reduction judgment threshold value, and the second threshold value is greater than the first threshold value. When the cache miss margin value is greater than the second threshold value, it indicates that the resource utilization in the cache partition is low, and the partition capacity can be reduced. For example, the second threshold value is equal to 0.2.
[0139] S412, reducing the cache partition capacity. The second threshold value is greater than the first threshold value.
[0140] Reducing the cache partition capacity can include reducing the cache partition capacity by a set proportion or reducing the cache partition capacity by a set capacity. For example, the capacity of the cache partition is reduced to 5% of the current capacity value, or the current capacity value of the cache partition is reduced by 0.05 GB.
[0141] After the cache miss margin value is greater than the first threshold value, indicating that the expansion condition is not triggered, the cache miss margin value is further compared with the second threshold value to determine whether the reduction condition is met. When the cache miss margin value is greater than the second threshold value, the capacity of the current cache partition is reduced, the waste of resources such as memory is reduced, and the resources can be allocated to the cache partition that needs more capacity.
[0142] It should be noted that if the cache miss margin values of multiple cache partitions are all greater than the second threshold value, the capacity of each cache partition is reduced.
[0143] In this embodiment, when the cache miss margin value is greater than the second threshold value, the idle cache memory is released, the memory cost is reduced, and the released memory can be reallocated to other cache partitions that need more capacity, thereby improving the overall resource utilization.
[0144] S414, maintaining the cache partition division scheme.
[0145] The cache partition division scheme can be understood as a rule and configuration for dividing the cache pool into independent cache partitions, including the cache partition corresponding to the hot function, the memory size of the cache partition, etc.
[0146] Maintaining the cache partition division scheme can be understood as maintaining the current cache partition corresponding to the hot function, the memory size of the cache partition, etc.
[0147] When the cache miss margin value is between the first threshold value and the second threshold value, the number, corresponding function, capacity, etc. of the current cache partition are maintained, so as to avoid frequent adjustment affecting the system stability.
[0148] On the basis of the above embodiment, the application optimizes the cache partition capacity adjustment method, as shown in Figure 5 The optimized cache partition capacity adjustment method includes steps S502-S512.
[0149] S502, acquire actual values of cold start ratios corresponding to at least two cache partitions in a cache pool respectively, wherein each cache partition is used to store function instances of a corresponding hot function, and the actual value of the cold start ratio is determined by the cache miss rate of the corresponding cache partition;
[0150] S504, for each cache partition, calculate a cache miss margin value based on the actual value of the cold start ratio and a target value of the cold start ratio, wherein the cache miss margin value and the difference between the target value of the cold start ratio and the actual value of the cold start ratio are positively correlated.
[0151] S506, determine a maximum cache miss margin value and a minimum cache miss margin value from the cache miss margin values of the at least two cache partitions.
[0152] The maximum cache miss margin value can be understood as the maximum value among the cache miss margin values of the plurality of cache partitions, indicating that the target value of the cold start ratio is much larger than the actual value of the cold start ratio, i.e., the actual value of the cold start ratio is much smaller than the target value, and the cache resource is idle and needs to be released.
[0153] The minimum cache miss margin value can be understood as the minimum value among the cache miss margin values of the plurality of cache partitions, indicating that the target value of the cold start ratio is only a little larger than the actual value of the cold start ratio, or the actual value of the cold start ratio is larger than the target value of the cold start ratio, i.e., the minimum cache miss margin value is negative, and the capacity of the cache partition needs to be increased.
[0154] The cache miss margin values calculated for the plurality of cache partitions are compared and filtered to select the maximum cache miss margin value and the minimum cache miss margin value.
[0155] S508, determine whether the minimum cache miss margin value is less than or equal to a first threshold value, if yes, execute S510, if no, execute S512.
[0156] S510, increase the capacity of the cache partition when the minimum cache miss margin value is less than or equal to the first threshold value.
[0157] The smaller the minimum cache miss margin value, the closer the actual value of the cold start ratio of at least one cache partition to the target value, or the actual value of the cold start ratio is larger than the target value, and the capacity of the cache partition needs to be expanded.
[0158] Determine the cache partition corresponding to the minimum cache miss margin value and increase the capacity thereof, for example, determine cache partition A corresponding to the minimum cache miss margin value, and increase the capacity of cache partition A from 1GB to 1.1GB.
[0159] In a possible implementation, starting from the cache partition corresponding to the minimum cache miss margin, the cache partitions with a cache miss margin less than the first threshold are respectively increased by 10% of the capacity.
[0160] S512, determining whether the maximum cache miss margin is greater than a second threshold, if yes, performing S514, if no, performing S516.
[0161] S514, reducing the capacity of the cache partition, wherein the second threshold is greater than the first threshold.
[0162] The greater the maximum cache miss margin, the more the cold start rate of at least one cache partition is far below the target, and the more the cache resource is idle, and the more the capacity needs to be reduced to release the resource.
[0163] When the minimum cache miss margin is greater than the first threshold, it is indicated that the expansion condition is not triggered, and then it is determined whether the maximum cache miss margin is greater than the second threshold. If the maximum cache miss margin is greater than the second threshold, it is determined that the cache partition corresponding to the maximum cache miss margin is reduced in capacity, for example, the capacity of the cache partition B corresponding to the minimum cache miss margin is increased from 1GB to 0.95GB.
[0164] S516, maintaining the cache partition division scheme.
[0165] When the minimum cache miss margin is greater than the first threshold and the maximum cache miss margin is less than or equal to the second threshold, it is indicated that the difference between the actual value and the target value of the cold start rate of all cache partitions is moderate, and the resource allocation is reasonable, and the capacity does not need to be adjusted. The capacity and division scheme of the current cache partition are maintained unchanged, so as to avoid system fluctuation caused by frequent adjustment.
[0166] In the embodiment, the cache partition corresponding to the minimum cache miss margin has an actual value of the cold start rate closest to the target value, or the actual value is less than the target value, and at this time, the cold start rate of the cache partition will exceed the standard or has already exceeded the standard, and the capacity needs to be expanded preferentially. The cache partition corresponding to the maximum cache miss margin has an actual value of the cold start rate far below the target value, and at this time, the resource utilization rate of the cache partition is low, and the capacity needs to be reduced to release the resource. Through the maximum value and the minimum value scheme, the cache partition which needs to be adjusted in capacity preferentially is accurately positioned, and the optimization efficiency is improved.
[0167] On the basis of the above embodiment, the cache management method is optimized, and a global cache management method is provided, as shown in Figure 6 The optimized cache management method includes steps S602-S606.
[0168] S602, when the cache miss margin value of any one cache partition is less than or equal to the first threshold, starting a timer.
[0169] The timer is used to record the duration that the cache partitions are at risk of cold start. When the cache miss margin value of any one of the cache partitions is less than or equal to the first threshold value, the timer starts. When the cache miss margin value of all the cache partitions is greater than the first threshold value, the timer is reset.
[0170] When the cache miss margin value of at least one of the cache partitions reaches or is lower than the first threshold value, the condition for starting the timer is triggered, indicating that the cold start rate of the cache partition is high. The time recording is started, and the duration from the first time when the cache miss margin value of the cache partition is less than or equal to the first threshold value to the current time is calculated.
[0171] S604, after adjusting the cache partition capacity, if the cache miss margin value of a set number of cache partitions is greater than the first threshold value, the timer is reset.
[0172] The set number can be set according to actual conditions. For example, the set number can be the total number of cache partitions in the cache pool.
[0173] In one possible implementation, after the cache partition capacity is adjusted, when the cache miss margin value of all the cache partitions exceeds the first threshold value, it indicates that the cold start risk has been alleviated. The timer counting duration is cleared, and the current time recording is stopped.
[0174] S606, if the timer counting duration exceeds a third threshold value, the capacity of the cache pool is increased, and / or the function instances of the hot function are migrated.
[0175] The third threshold value can be understood as a critical value of the pre-set maximum timer counting duration. When the counting duration exceeds the third threshold value, it indicates that the serverless computing platform has a long-term high cold start rate, and the operation of increasing the cache pool capacity and / or migrating the function instances needs to be performed. For example, the third threshold value can be set according to actual conditions. The third threshold value is 30 minutes.
[0176] When the timer continuously records the time exceeding the pre-set third threshold value, it indicates that there is a long-term high cold start rate risk, and the cache partition capacity adjustment fails to solve the problem of high cold start rate, triggering the global optimization measures, such as increasing the cache pool capacity or migrating the function instances.
[0177] Increasing the capacity of the cache pool can be understood as expanding the total resource scale of the entire cache pool to provide more available resources for all cache partitions and improve the overall cache capacity.
[0178] The function instance of the migration hotspot function can be understood as migrating the function instance of the hotspot function from the current worker node to other worker nodes with lower load, balancing the resource distribution in the entire serverless computing platform. Through cross-node resource scheduling, the cold start rate of a single worker node due to long-term high load is avoided.
[0179] In one possible implementation, the cache miss margin values of the cache partitions are monitored in real time, and when the cache miss margin value of any one cache partition is less than or equal to a first threshold value, a timer is started to record time. A 10% expansion operation is performed on the cache partition with a cache miss margin value less than or equal to the first threshold value to try to solve the local risk. Moreover, during the timing period, the expansion operation can be performed periodically multiple times, and after multiple expansion adjustments, the cache miss margin values of all cache partitions are checked. If the cache miss margin values of all cache partitions are greater than the first threshold value, the timer is reset and the current monitoring is ended. If the timer timing duration exceeds a third threshold value, it indicates that the expansion adjustment of the cache partition is invalid, and a global expansion or migration of the hotspot function instance to other worker nodes is performed. The global expansion can include increasing the cache pool capacity by 10%.
[0180] In this embodiment, when the cache miss margin value of any one cache partition is less than or equal to the first threshold value, a timer is started, and when the timer timing duration exceeds the third threshold value, a global optimization measure such as increasing the cache pool capacity or migrating the function instance is triggered to reduce the server load and prevent long-term server performance decline.
[0181] In this embodiment, an application example of cache management is provided. A test platform is configured with 256 GB of memory and a CPU with 48 cores and 2.50 GHz, and runs a set operating system. Real-world business applications are used for testing, including four workloads of serverless computing services, each containing 384 different function execution times, memory sizes, and call timestamps within 24 hours.
[0182] The cache management method provided in this embodiment is compared with an existing cold start scheme, which uses a monolithic cache pool design, and all functions share the same cache pool.
[0183] The serverless computing platform runs in standalone mode in one worker node, the size of the cache pool is initialized to 30% of the memory capacity of the worker node, and since the number of hot functions is small, the number of partitions in the cache pool is set to 20, and the cold start rate target value of the hot function is defined as 1%. The monitoring period is set to 5 minutes. Replay Trace A to generate the workload, which includes 11,363,701 call records from 384 functions in a day, an average of 132 access concurrency per second, and more than 90% of user requests come from 4 hot functions. 384 functions are deployed in the local serverless cluster, and the load generator is deployed separately in an additional worker node. The request calls in Trace B are sequentially issued to simulate the real arrival of user requests.
[0184] When a user request is issued, it first reaches the function scheduler of the serverless computing platform, which is forwarded to the request router, which queries whether there is an available function instance in the worker node. If so, the user request is forwarded to the cache partition where the available function instance is located for processing; if not, a cold start operation is triggered, and the function scheduler creates a function instance in a worker node according to a specific strategy, and forwards the user request to the worker node. Inside the worker node, the user request first reaches the request router, which forwards the user request to the corresponding function instance, and when the function instance executes the request, the cache pool manager is responsible for the life cycle management of the function instance. The cache pool manager performs the following operations, first, it collects historical data of the function to determine whether it is a hot function, if it is a hot function, it allocates a cache partition for it, and caches the function instance in the partition. If the function is not a hot function, the function instance is not cached.
[0185] As the system runs, the cache pool manager dynamically and periodically adjusts the capacity of the cache partition where each hot function is located through heuristic algorithms. For example, the cache miss margin values of hot function 1, hot function 2, and hot function 3 are 0.02, 0.06, and 0.08 respectively. The minimum cache miss margin value is 0.02, at this time the cache partition capacity of hot function 1 will be updated from 5% to 5%*1.1=5.5%. Hot function 2 and hot function 3 do not operate. For another example, if the cache miss margin values of hot function 1, hot function 2, and hot function 3 are 0.05, 0.06, and 0.3 respectively. The maximum cache miss margin value is 0.3, at this time the cache partition capacity of hot function 3 will be updated from 5% to 5%*0.95=0.475%. Hot function 1 and hot function 2 do not operate. As the serverless computing platform runs, the above operations are repeated to complete the allocation and management of function instance caching and cache partition resources.
[0186] Keeping the running environment configuration unchanged in the above examples, the workload is modified to Trace B, which contains 3,111,827 call records from 384 functions in a day, with an average of 36 accesses per second and a concurrency of 36, and more than 90% of the requests come from 14 hot functions. The technical test process is repeated to obtain new comparison results.
[0187] Figure 7 The overall performance comparison under the Trace A workload is shown, and for ease of comparison, the experimental result data is normalized to 1. From the comparison results, it can be seen that the performance of the serverless computing platform using the cache management method provided in the embodiment is higher than that of the cold start optimization method of the serverless computing platform in the related art, and the cold start call occurrence rate of the overall function is reduced by about 10%, that is, the number of hot calls is more, and the delay growth caused by the cold start is less. At the same time, it can be seen that the use amount of cache resources is reduced by 5%, which also shows that the cache management method provided in the embodiment can achieve better results with less cloud resource cost. Figure 7
[0188] From the comparison results, it can be seen that the performance of the serverless computing platform using the cache management method provided in the embodiment is higher than that of the cold start optimization method of the serverless computing platform in the related art, and the cold start call occurrence rate of the overall function is reduced by about 10%, that is, the number of hot calls is more, and the delay growth caused by the cold start is less. At the same time, it can be seen that the use amount of cache resources is reduced by 5%, which also shows that the cache management method provided in the embodiment can achieve better results with less cloud resource cost. Figure 8 In addition,
[0189] Further, the cold start rate variance of 384 functions is shown, and it can be seen that the cache management method provided in the embodiment can effectively reduce the performance difference between different functions (in a multi-tenant scenario), and can reduce the performance difference by nearly 60% compared with the existing scheme. Figure 9 It should be noted that the acquisition, storage, use, processing, etc. of data in the technical solutions of the present disclosure comply with the relevant provisions of relevant laws and regulations. The personal identity data, operation data, behavior data, etc. of various types of data obtained in the embodiments of the present disclosure have obtained the consent of the user.
[0190] Based on the same inventive concept, the present disclosure also provides a cache management device, as described in the following embodiments. Since the principles of the device embodiments solve problems similar to the above-mentioned method embodiments, the implementation of the device embodiments can be referred to the implementation of the above-mentioned method embodiments, and the repeated parts will not be described again.
[0191]
[0192] Figure 10 Fig. 1 shows a schematic diagram of a cache management apparatus according to an embodiment of the present disclosure. Figure 10 As shown, the apparatus comprises an obtaining module 1010 and an adjusting module 1020.
[0193] The obtaining module 1010 is configured to obtain actual values of cold start ratios respectively corresponding to at least two cache partitions in a cache pool, wherein each cache partition is configured to store function instances of a corresponding hotspot function, and the actual value of the cold start ratio is determined by a cache miss ratio of the corresponding cache partition.
[0194] In some possible embodiments, the adjusting module 1020 comprises a missing margin value calculation unit configured to calculate, for a cache partition, a cache missing margin value based on the actual value of the cold start ratio and the target value of the cold start ratio, wherein the cache missing margin value is positively correlated with the difference between the target value of the cold start ratio and the actual value of the cold start ratio; and a capacity adjusting unit configured to adjust the capacity of the cache partition based on the cache missing margin value of the cache partition.
[0195] In some possible embodiments, the capacity adjusting unit is specifically configured to increase the capacity of the cache partition when the cache missing margin value of the cache partition is less than or equal to a first threshold value.
[0196] In some possible embodiments, the capacity adjusting unit is specifically configured to decrease the capacity of the cache partition when the cache missing margin value of the cache partition is greater than a second threshold value, wherein the second threshold value is greater than the first threshold value.
[0197] In some possible embodiments, the adjusting module 1020 further comprises a maximum and minimum value selection unit configured to determine a maximum cache missing margin value and a minimum cache missing margin value from the cache missing margin values of the at least two cache partitions; and the capacity adjusting unit is specifically configured to decrease the capacity of the cache partition corresponding to the maximum cache missing margin value if the minimum cache missing margin value is greater than the first threshold value and the maximum cache missing margin value is greater than the second threshold value.
[0198] In some possible embodiments, the apparatus further comprises a global optimization module configured to start a timer when the cache missing margin value of any one of the cache partitions is less than or equal to the first threshold value; reset the timer if the cache missing margin values of a set number of the cache partitions are all greater than the first threshold value after the capacities of the cache partitions are adjusted; and increase the capacity of the cache pool and / or migrate the function instances of the hotspot functions if a time length of the timer exceeds a third threshold value.
[0199] In some possible embodiments, the method further includes: constructing a cache pool; and dividing the cache pool into a plurality of independent cache partitions, wherein an initial capacity of each cache partition is the same.
[0200] In some possible embodiments, the method further includes: receiving a user request; starting a new function instance if a function instance corresponding to the user request does not exist in the cache pool; invoking the new function instance to process the user request; and storing the new function instance into a corresponding cache partition if the function corresponding to the user request is a hot function after the invocation of the new function instance ends.
[0201] It should be noted that each module in the apparatus embodiments and the corresponding steps in the method embodiments have the same examples and application scenarios, but are not limited to the content disclosed in the above method embodiments. It should be noted that the above modules as part of the apparatus can be executed in a computer system such as a group of computer executable instructions.
[0202] Those skilled in the art can understand that each aspect of the present disclosure can be implemented in the form of a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software aspects, which can be collectively referred to as "circuitry", "module" or "system".
[0203] Based on the same inventive concept, the electronic device provided in the embodiments of the present disclosure includes: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the cache management method of any one of the above by executing the executable instructions. Since the principle of solving problems of the electronic device embodiments is similar to that of the above method embodiments, the implementation of the electronic device embodiments can refer to the implementation of the above method embodiments, and the repeated parts will not be described here.
[0204] The electronic device 1100 according to this implementation of the present disclosure will be described below with reference to Figure 11 . Figure 11 The electronic device 1100 shown is merely an example and should not impose any limitation on the function and scope of use of the embodiments of the present disclosure.
[0205] As shown in Figure 11 , the electronic device 1100 is in the form of a general computing device. The components of the electronic device 1100 can include, but are not limited to, the above at least one processing unit 1110, the above at least one storage unit 1120, and a bus 1130 connecting different system components, including the storage unit 1120 and the processing unit 1110.
[0206] The storage unit stores program codes which can be executed by the processing unit 1110, so that the processing unit 1110 performs the steps described in the above “Exemplary Method” section according to various exemplary embodiments of the present disclosure. For example, the processing unit 1110 can perform the following steps of the above method embodiments: constructing a first noise library; obtaining a noise pattern through a noise pattern adding interface, and adding the noise pattern to the first noise library to obtain a second noise library; obtaining a first text; cleaning the first text based on the second noise library to obtain a second text; and outputting the second text.
[0207] The storage unit 1120 can include a readable medium in the form of volatile storage such as a random access memory (RAM) 11201 and / or cache 11202, and can further include a read-only memory (ROM) 11203.
[0208] The storage unit 1120 can further include program / utility 11204 having a set of programs / modules 11205, including operating systems, one or more application programs, other programs, and programmatic modules, each of which can implement aspects of the networks environment, and each or a combination thereof.
[0209] The bus 1130 can represent one or more of several types of bus structures, including a storage bus or bus controller, a peripheral bus, a graphics bus, a processor or local bus using any of a variety of bus architectures.
[0210] The electronic device 1100 can also communicate with one or more external devices 1140 such as a keyboard or pointing device, a Bluetooth device, etc.; other devices such as printers, scanners, etc.; and / or various types of networks including a local area network (LAN) or a wide area network (WAN), such as the Internet. As illustrated, the electronic device 1100 can communicate with one or more networks 1160 via a network adapter 1160. It will be appreciated that the network connections shown are exemplary and other communications devices, such as a modem, a wireless link or other devices, can also be used. The input / output (I / O) interface 1150 can include, among other things, a keyboard; a mouse; a scanner; a microphone; a camera; a display; a speaker; etc. The I / O interface 1150 can also include an interface to one or more external devices 1140. As shown, the I / O interface 1150 can communicate with the other components of the electronic device 1100 via the bus 1130. However, the I / O interface 1150 can communicate with the other components using any appropriate communication media, including via a wireless link.
[0211] Those skilled in the art can clearly understand, through the description of the above embodiments, that the example embodiments described herein can be implemented by software, or by software in combination with necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a U disk, a mobile hard disk, etc.) or a network, and includes a plurality of instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to perform the method according to the embodiments of the present disclosure.
[0212] Based on the same inventive concept, the present disclosure also provides a computer readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the cache management method of any one of the above. Since the principle of solving problems of the computer readable storage medium embodiment is similar to that of the above method embodiment, the implementation of the computer readable storage medium embodiment can be referred to the implementation of the above method embodiment, and the repeated parts will not be described here.
[0213] More specific examples of the computer readable storage medium in the present disclosure can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0214] In the present disclosure, the computer readable storage medium can include a data signal carried in a baseband or as a part of a carrier wave, in which readable program codes are borne. Such a propagated data signal can take various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The readable signal medium can also be any readable medium other than the readable storage medium, which can send, propagate or transmit programs for use by or in connection with an instruction execution system, apparatus or device.
[0215] Optionally, the program code contained on the computer readable storage medium can be transmitted by any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.
[0216] In particular embodiments, the program code can execute entirely on the user's computing device, partly on the user's computing device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device or entirely on the remote computing device or server. In the latter scenario, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computing device such as through the Internet using an Internet Service Provider. The user's computing device can also be connected to the remote computing device through a wireless network.
[0217] Based on the same inventive concept, the embodiments of the present disclosure further provide a computer program product, comprising a computer program product, comprising: a computer program or instructions, which, when executed by a processor, implement the cache management method of any one of the above method embodiments. Since the principle of solving problems of this computer program product embodiment is similar to the above method embodiments, the implementation of this computer program product embodiment can be referred to the implementation of the above method embodiments, and the repeated parts will not be described here.
[0218] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, such division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units for embodiment.
[0219] In addition, although the steps of the method in the present disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in this specific order, or that all the shown steps must be performed to achieve the desired result. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step, one step can be divided into multiple steps, etc.
[0220] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by software in combination with necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash disk, a mobile hard disk, or the like) or a network, and includes a number of instructions to make a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) execute the methods according to the embodiments of the present disclosure.
[0221] Other embodiments of the present disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the features disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure following the general principles thereof and including such departures from the present disclosure that come within known use or custom in the art to which the present disclosure pertains. The specification and examples are to be regarded as illustrative only, and the true scope and spirit of the present disclosure are indicated by the appended claims.
Claims
1. A cache management method, characterized in that: include: Obtaining actual cold start ratio values corresponding to at least two cache partitions in the cache pool, wherein each cache partition is used to store a function instance of its corresponding hot function, and the actual cold start ratio value is determined by a cache miss rate of the corresponding cache partition; The cache partition capacity is adjusted based on actual cold start ratio values and target cold start ratio values respectively corresponding to at least two of the cache partitions.
2. The cache management method according to claim 1, wherein: The adjusting the cache partition capacity based on the actual cold start ratio value and the target cold start ratio value respectively corresponding to at least two of the cache partitions includes: a cache miss margin value calculated for each of the cache partitions based on the actual cold start ratio value and the target cold start ratio value, wherein the cache miss margin value is positively correlated with a difference between the target cold start ratio value and the actual cold start ratio value; The cache partition capacity is adjusted based on the cache miss margin value of the cache partition.
3. The cache management method according to claim 2, wherein: The adjusting the cache partition capacity based on the cache miss margin value of the cache partition includes: When the cache miss margin value of the cache partition is less than or equal to a first threshold, the cache partition capacity is increased.
4. The cache management method according to claim 3, wherein: The adjusting the cache partition capacity based on the cache miss margin value of the cache partition further includes: When the cache miss margin value of the cache partition is greater than a second threshold, the cache partition capacity is reduced, wherein the second threshold is greater than the first threshold.
5. The cache management method according to claim 4, wherein: Also includes: Determining a maximum cache miss margin and a minimum cache miss margin from the cache miss margin values of at least two of the cache partitions; When the cache miss margin value of the cache partition is greater than a second threshold, reducing the cache partition capacity includes: If the minimum cache miss margin is greater than the first threshold, and the maximum cache miss margin is greater than the second threshold, the cache partition capacity corresponding to the maximum cache miss margin is reduced.
6. The cache management method according to claim 2, wherein: Also includes: When the cache miss margin value of any cache partition is less than or equal to the first threshold, starting a timer; After adjusting the cache partition capacity, if the cache miss margin values of a set number of the cache partitions are all greater than the first threshold, resetting the timer; If the timing duration of the timer exceeds a third threshold, the capacity of the cache pool is increased, and / or the function instance of the hot function is migrated.
7. The cache management method according to claim 1, wherein: Also includes: Build a cache pool; The cache pool is divided into a plurality of independent cache partitions, wherein the initial capacity of each cache partition is the same.
8. The cache management method according to claim 7, wherein: Also includes: Receive user requests; If the function instance corresponding to the user request does not exist in the cache pool, starting a new function instance; Calling the new function instance to process the user request; After the new function instance call is completed, if the function corresponding to the user request is the hot function, the new function instance is stored in the corresponding cache partition.
9. A cache management device, characterized in that: include: an acquisition module, configured to obtain actual cold start ratio values corresponding to at least two cache partitions in a cache pool, each of which is configured to store a function instance of a corresponding hotspot function, and wherein the actual cold start ratio value is determined by a cache miss rate of the corresponding cache partition; The adjustment module is configured to adjust the cache partition capacity based on actual cold start ratio values and target cold start ratio values corresponding to at least two of the cache partitions.
10. An electronic device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; The processor is configured to execute the cache management method according to any one of claims 1 to 8 by executing the executable instructions.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the cache management method according to any one of claims 1 to 8 is implemented.
12. A computer program product comprising: A computer program or instruction, characterized in that when the computer program or instruction is executed by a processor, it implements the cache management method described in any one of claims 1 to 8.