Distributed function scheduling optimization method and device under serverless architecture

By adopting dynamic priority scheduling and distributed cache optimization methods in the cloud computing platform, the problem of low function scheduling efficiency in the serverless architecture is solved, efficient function task processing and resource utilization are achieved, and the system stability and robustness are improved.

CN120315853BActive Publication Date: 2025-09-26NORTHEASTERN UNIV CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510184281.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-09-26
Estimated Expiration
2045-02-19

AI Technical Summary

Technical Problem

In serverless cloud computing platforms, existing function scheduling strategies are unable to dynamically adjust priorities, resulting in delayed execution of critical tasks. This is especially true in high-concurrency scenarios, where they cannot meet the needs of latency-sensitive applications. Furthermore, static priorities cannot adapt to real-time load changes, leading to uneven resource allocation and affecting overall efficiency.

Method used

A dynamic priority scheduling strategy is adopted to calculate the comprehensive priority by obtaining priority rules and real-time monitoring data, dynamically adjust the priority of function tasks, and combine distributed caching and monitoring feedback mechanisms to optimize function scheduling.

Benefits of technology

It improves system response speed and resource utilization, reduces cold start delay, supports priority scheduling in high-concurrency scenarios, and improves user experience and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120315853B_ABST
    Figure CN120315853B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a distributed function scheduling optimization method and device under a serverless architecture. The method adopts a dynamic priority scheduling strategy to dynamically adjust the priority of function tasks according to predefined priority rules and real-time monitoring data to ensure that critical tasks are given priority and improve the system response speed and resource utilization efficiency. A distributed cache is used to store function tasks, execution environment snapshots and dependent libraries, and a preheating mechanism is supported to reduce cold start delays and improve the execution efficiency of function tasks. Through the monitoring feedback mechanism, indicator data related to system operation is collected in real time, and the priority and cache strategy are dynamically adjusted according to the indicator data to achieve adaptive optimization and improve system stability and robustness. Through dynamic priority scheduling and cache optimization, the system throughput and resource utilization are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of cloud computing technology, and in particular to a distributed function scheduling optimization method and device under a serverless architecture. Background Art

[0002] Cloud computing is a service model that provides computing resources over the internet. Cloud computing can provide computing resources such as servers, storage, databases, networks, and software to clients with diverse needs. Cloud computing models include server-based and serverless architectures. In a serverless architecture, users do not need to manage servers. Instead, they upload code snippets to the cloud platform, which automatically allocates resources and executes the code to meet their specific needs.

[0003] In a serverless cloud computing model, the computing platform needs to schedule and optimize functions based on user input requests. Since cloud computing platforms are widely used in various scenarios, they involve numerous functions. Therefore, the computing platform needs to schedule functions according to a pre-defined scheduling policy. For example, the computing platform can schedule functions based on a first-in-first-out (FIFO) queue. This approach fails to prioritize functions, potentially delaying the execution of critical tasks. This approach, especially in high-concurrency scenarios, fails to meet the needs of latency-sensitive applications.

[0004] Some cloud computing platforms also support setting static function priorities. This means that by pre-configuring the priority of each function, the computing platform schedules functions according to the configured priority. However, function scheduling based on static priorities cannot dynamically adjust to real-time load and business needs. When load patterns change, static priorities can lead to uneven resource allocation, affecting overall efficiency. Summary of the Invention

[0005] In view of this, the embodiments of the present application provide a distributed function scheduling optimization method and device under a serverless architecture, which can improve the efficiency and performance of the serverless computing platform, reduce cold start delay, support priority scheduling in high concurrency scenarios, improve user experience and resource utilization, and solve the problem of low function scheduling efficiency of the cloud computing platform.

[0006] According to one aspect of the present application, a distributed function scheduling optimization method in a serverless architecture is provided, the method comprising:

[0007] Obtain priority rules and real-time monitoring data;

[0008] Calculating a comprehensive priority based on the priority rule and the real-time monitoring data, the comprehensive priority being a weighted sum of scores of multiple monitoring items in the real-time monitoring data;

[0009] Loading a target function into a multi-level task queue according to the comprehensive priority, wherein the target function is a function determined according to the comprehensive priority and a queue length of the multi-level task queue;

[0010] Extracting a function task from the multi-level task queue, and executing startup function scheduling according to the cache status of the function task, the startup function includes a first function and a second function, the first function is a function task reused from the cache; the second function is a newly created function task;

[0011] Record the indicator data after the execution of the startup function, and adjust the function scheduling strategy according to the indicator data and the real-time monitoring data, the indicator data including execution performance, resource usage data and historical records; the function scheduling strategy includes the priority rules, the calculation weight parameters of the comprehensive priority and the queue management strategy of the multi-level task queue.

[0012] Optionally, the method further includes:

[0013] Get historical function call data;

[0014] Extracting function call parameters within a preset period from the historical function call data, the function call parameters including function name, call time, and call frequency;

[0015] Predicting a preheating function based on the function call parameters, and loading the preheating function into a distributed cache;

[0016] Function-related information corresponding to the preheating function is loaded into a distributed memory database, where the function-related information includes at least one of a function code compilation result, an execution environment snapshot, a function input parameter mode, and a function resource requirement.

[0017] Optionally, predicting a warm-up function based on the function call parameters includes:

[0018] Obtaining user behavior data, including user access paths, search records, and execution history;

[0019] Building a user behavior model based on the user behavior data, wherein the user behavior model is used to calculate the probability of the function task being called;

[0020] Obtaining system context information, and determining a preheating strategy based on the system context information, wherein the preheating strategy is dynamically adjusted based on the system context information and a deep Q learning algorithm;

[0021] The preheating function list is generated according to the preheating strategy, and the preheating function list includes a preset number of the preheating functions screened based on the preheating strategy; the preheating function is a function task whose probability of being called is greater than or equal to a probability threshold and satisfies the preheating strategy; the preset number is set according to the queue length of the multi-level task queue.

[0022] Optionally, the method further includes:

[0023] Collect update information, which is information generated after the function task is executed; the update information includes updated code compilation results and the latest status information of the execution environment;

[0024] Obtaining a baseline parameter of a cache update strategy, the baseline parameter including at least one of a version number of a function task, an execution timestamp, and a degree of change impact;

[0025] The function task in the distributed cache is updated based on the benchmark parameter and the update information.

[0026] Optionally, the method further includes:

[0027] Recording access information of the function task during cache usage, wherein the access information includes access time and access frequency;

[0028] Setting a hierarchical elimination strategy for the multi-level cache according to the accessed information;

[0029] When the cache space of the current level of the multi-level cache is full, searching for an object to be replaced according to the hierarchical elimination strategy of the current level;

[0030] The object to be replaced is removed from the multi-level cache, and cache information corresponding to the new function resource is loaded into the multi-level cache.

[0031] Optionally, starting function scheduling is performed according to the cache status of the function task, including:

[0032] Periodically selecting functions to be executed from the multi-level task queue in order of priority;

[0033] Searching for related resources of the function to be executed in the distributed cache;

[0034] If relevant resources of the to-be-executed function are found in the distributed cache, the to-be-executed function is started using the execution environment in the cache to start the first function;

[0035] If relevant resources of the function to be executed are not found in the distributed cache, resources are allocated for the function to be executed, and an execution environment is constructed to start the second function.

[0036] Optionally, periodically selecting a function to be executed from the multi-level task queue in order of priority includes:

[0037] Acquiring task requirement information, wherein the task requirement information includes at least one of a business characteristic of the system, a function call frequency, and an expected response timeliness;

[0038] Setting scheduling cycle parameters according to the task requirement information;

[0039] Based on the scheduling period parameters, functions to be executed are periodically selected from the multi-level task queue in the order of the comprehensive priority.

[0040] Optionally, searching for related resources of the function to be executed in the distributed cache includes:

[0041] Sending a query request to the distributed cache;

[0042] Recording the waiting time after sending the query request;

[0043] If the waiting time is less than or equal to the time threshold, mark that the relevant resources of the to-be-executed function are found in the distributed cache;

[0044] If the waiting time is greater than a time threshold, it is marked that no relevant resources of the to-be-executed function are found in the distributed cache.

[0045] Optionally, recording indicator data after the startup function is executed, and adjusting the scheduling strategy according to the indicator data and the real-time monitoring data, includes:

[0046] Periodically acquire multi-dimensional indicator data based on the collection interval parameter, including system resource utilization, function execution performance, and queue status;

[0047] Performing preprocessing on the indicator data according to preset processing items, wherein the preset processing items include one or more combinations of removing outliers, filling missing values, and format conversion;

[0048] The priority rules, the calculation weight parameters of the comprehensive priority and the queue management strategy of the multi-level task queue are adjusted based on the preprocessed indicator data.

[0049] According to another aspect of the present application, a distributed function scheduling optimization device in a serverless architecture is provided, the device comprising:

[0050] Acquisition module, used to obtain priority rules and real-time monitoring data;

[0051] A dynamic priority management module, configured to calculate a comprehensive priority based on the priority rules and the real-time monitoring data, wherein the comprehensive priority is a weighted sum of scores of multiple monitoring items in the real-time monitoring data;

[0052] A distributed cache module, configured to load a target function into a multi-level task queue according to the comprehensive priority, wherein the target function is a function determined according to the comprehensive priority and a queue length of the multi-level task queue;

[0053] A scheduling module is used to extract function tasks from the multi-level task queue and execute startup function scheduling according to the cache status of the function tasks, wherein the startup function includes a first function and a second function, the first function is a function task reused from the cache; the second function is a newly created function task;

[0054] A monitoring feedback module is used to record the indicator data after the execution of the startup function, and to adjust the function scheduling strategy based on the indicator data and the real-time monitoring data. The indicator data includes execution performance, resource usage data and historical records; the function scheduling strategy includes the priority rules, the calculation weight parameters of the comprehensive priority and the queue management strategy of the multi-level task queue.

[0055] According to another aspect of the present application, a computer device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein when the processor executes the program, the distributed function scheduling optimization method under the above-mentioned serverless architecture is implemented.

[0056] According to another aspect of the present application, a storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the distributed function scheduling optimization method under the above-mentioned serverless architecture is implemented.

[0057] By means of the above technical solution, an embodiment of the present application provides a distributed function scheduling optimization method and device under a serverless architecture, wherein the method adopts a dynamic priority scheduling strategy, and dynamically adjusts the priority of function tasks according to predefined priority rules and real-time monitoring data, ensuring that critical tasks are given priority and improving the system response speed and resource utilization efficiency. A distributed cache is used to store function tasks, execution environment snapshots and dependent libraries, and a preheating mechanism is supported to reduce cold start delays and improve the execution efficiency of function tasks. Through a monitoring feedback mechanism, indicator data related to system operation is collected in real time, and the priority and cache strategy are dynamically adjusted according to the indicator data to achieve adaptive optimization and improve system stability and robustness. Through dynamic priority scheduling and cache optimization, the system throughput and resource utilization are improved.

[0058] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0060] Figure 1 A schematic diagram of the cloud computing platform structure provided in an embodiment of the present application;

[0061] Figure 2 A flow chart of a distributed function scheduling optimization method under a serverless architecture provided in an embodiment of the present application;

[0062] Figure 3 A schematic diagram of the distributed function scheduling optimization interaction process provided in an embodiment of the present application;

[0063] Figure 4 A schematic diagram of the function scheduling process provided in an embodiment of the present application;

[0064] Figure 5 A schematic diagram of the preheating function scheduling process provided in an embodiment of the present application;

[0065] Figure 6 A schematic diagram of the process of updating information in a distributed cache provided in an embodiment of the present application;

[0066] Figure 7 A schematic diagram of the distributed cache object replacement process provided in an embodiment of the present application;

[0067] Figure 8A schematic diagram of the structure of a distributed function scheduling optimization device under a serverless architecture provided in an embodiment of the present application;

[0068] Figure 9 A schematic diagram of the hardware structure of the function scheduling optimization device provided in an embodiment of the present application;

[0069] Figure 10 This is a timing diagram of the function scheduling optimization provided in an embodiment of the present application. DETAILED DESCRIPTION

[0070] The present application will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that, unless there is a conflict, the embodiments and features in the embodiments of the present application can be combined with each other.

[0071] In the embodiments of the present application, the function does not refer to a function in the mathematical sense, but refers to a computing task or a collection of computing tasks pre-built according to a specific cloud computing method. It is a task application deployed in a distributed service system that can be called by a client or the like with data processing and communication functions. Therefore, the function is also called a function task.

[0072] Cloud computing is a service model that provides computing resources over the internet. Cloud computing can provide computing resources such as servers, storage, databases, networks, and software to clients with varying needs. For example, cloud computing can be used to build e-commerce platforms to meet the interactive needs of users and operators.

[0073] In some embodiments, cloud computing models include server-based and serverless architectures. In a serverless architecture, users do not need to manage servers. Instead, they upload code snippets to a cloud platform, which automatically allocates resources and executes the code to meet their specific needs. For example, e-commerce operators can use serverless computing platforms to handle user requests, such as product browsing, order processing, and payment.

[0074] In some application scenarios, within a serverless cloud computing model, the computing platform must schedule and optimize functions based on user-input requests. For example, during promotional events, order processing requests surge, and payment requests are prioritized to ensure transaction completion. Because cloud computing platforms are widely used in various scenarios, they involve numerous functions, requiring them to execute function scheduling according to predefined scheduling policies.

[0075] In some embodiments, the computing platform can perform function scheduling based on a First In First Out (FIFO) queue. When scheduling functions based on a FIFO queue, the function tasks can be processed in the order in which they enter the queue, i.e., enqueue, and a function task is added to the tail of the queue. Correspondingly, the element that enters first will be taken out first, i.e., dequeue, and a function task will be removed from the head of the queue. This method cannot distinguish the priority of the function, resulting in the delayed execution of critical tasks, especially in high concurrency situations, and cannot meet the requirements of applications that are sensitive to delays.

[0076] In some embodiments, cloud computing platforms also support setting static function priorities. This means that by pre-configuring the priority of each function, the computing platform can schedule functions according to the configured priority. However, function scheduling based on static priorities cannot dynamically adjust to real-time load and business needs. When load patterns change, static priorities can lead to uneven resource allocation, affecting overall efficiency.

[0077] like Figure 1 As shown, the cloud computing platform can schedule function tasks associated with user requests in response to user requests. The scheduling process of function tasks involves the storage of function tasks, which can include persistent storage and runtime cache. The storage location can be a two-tier architecture of distributed cache and persistent storage. Caching can reduce direct access to persistent storage and reduce latency. Among them, persistent storage can be implemented based on storage media such as code repositories and image repositories. Code repositories can store function source code, configuration files, and dependency lists. Image repositories can save execution environment snapshots.

[0078] Runtime caching can cache function code compilation results, execution environment snapshots, and dependent libraries through storage media such as distributed in-memory databases (Hazelcast), accelerating subsequent calls. To execute function tasks, cloud computing platforms can use specific caching strategies to reduce access pressure on backend storage, improve data access speed, and optimize cloud computing platform performance.

[0079] In some embodiments, cloud computing platforms can use basic caching mechanisms to store function code or execution environments. For example, they can store function code or execution environments in a storage space shared by multiple application instances. However, this caching approach lacks fine-grained control and optimization. For example, the cache may only store function code without dependent libraries or pre-warmed execution environments, which can still result in cold start delays. Furthermore, cache replacement strategies may not be efficient enough to maximize cache resources.

[0080] After executing a function task, the cloud computing platform can generate execution result information based on the execution results of the function task. The generated execution result information can be sent to the user client through the communication connection relationship between the user client and the platform to realize specific interactive functions.

[0081] Since the cloud computing platform completes the execution of the corresponding cloud computing task after sending the execution result information to the user client, the cloud computing platform lacks a monitoring and feedback mechanism. For example, the cloud computing system lacks real-time monitoring and feedback of function execution performance, resource utilization and queue status, making it difficult for the system to dynamically adjust scheduling and caching strategies according to actual conditions, and unable to achieve adaptive optimization.

[0082] In order to solve the problems of low function scheduling efficiency, cold start delay, and inability to achieve adaptive optimization on cloud computing platforms, some embodiments of the present application provide a distributed function scheduling optimization method under a serverless architecture. The method can be run on a cloud computing platform, and specifically can be executed by multiple functional modules deployed in the cloud computing platform. For example, the execution subject of the function is a computing node of the cloud computing platform, such as Figure 2 As shown, the method includes:

[0083] S101. Obtain priority rules and real-time monitoring data.

[0084] like Figure 3 As shown, to perform function scheduling, the cloud computing platform can be configured with an acquisition module for acquiring priority rules and implementing monitoring data. Priority rules are used to calculate the execution priority of function tasks. Priority rules can be calculated based on specific evaluation metrics. Priority rules include, but are not limited to, business criticality, service level agreement (SLA) requirements, deadlines, call frequency, error rates, and the like.

[0085] In some embodiments, priority rules can be expanded across multiple dimensions based on specific application requirements. For example, in addition to basic metrics such as business criticality, service level agreement (SLA) requirements, deadlines, call frequency, average execution time, resource consumption, and error rate, metrics such as historical stability (HS), relative competitive intensity (RCI), and burstiness incidence (BI) can also be introduced.

[0086] Among them, historical stability is used to characterize the failure rate or abnormal fluctuations in past function executions. Resource contention intensity is used to characterize the current level of competition between system resources and other high-priority tasks. The burstiness factor is used to characterize the short-term surge in function call frequency, which is used to quickly respond to sudden demands.

[0087] While acquiring priority rules, the acquisition module can also monitor the entire cloud computing platform in real time to obtain real-time monitoring data. The real-time monitoring data can include function call patterns, call frequency, average execution time, resource consumption, error rate, burst traffic, etc.

[0088] S102: Calculate a comprehensive priority based on the priority rule and the real-time monitoring data.

[0089] After obtaining the priority rules and real-time monitoring data, the cloud computing platform can calculate a comprehensive priority based on the priority rules and the real-time monitoring data, wherein the comprehensive priority is a weighted sum of scores of multiple monitoring items in the real-time monitoring data.

[0090] To calculate the overall priority, a dynamic priority queue management module can be deployed in the cloud computing platform. This module is responsible for dynamically calculating and adjusting the priority of function tasks based on predefined priority rules and real-time monitoring data. This module can utilize an enhanced multi-level feedback queue structure, combined with an adaptive weight adjustment mechanism and a prediction-based priority boosting strategy, to achieve more intelligent and efficient function scheduling.

[0091] The dynamic priority queue management module calculates the overall priority based on a weighted adaptive priority algorithm. This algorithm dynamically adjusts weight parameters based on fixed weights, combining real-time system status and historical execution data. For example, when system resources are limited, the resource consumption weight (RCW) is dynamically increased; in bursty high-concurrency scenarios, the call frequency weight (CFW) and burstiness incidence weight (BIW) are increased.

[0092] In some embodiments, the comprehensive priority score (FPS) can be calculated according to the following formula:

[0093] FPS = (BCW*business criticality score) + (SLAW*SLA requirement score) + (DW*deadline score) + (CFW*call frequency score) + (AETW*average execution time score) + (RCW*resource consumption score) + (ERW*error rate score) + (HSW*historical stability score) + (RCIW*resource contention intensity score) + (BIW*burst factor score);

[0094] Among them, FPS represents the comprehensive priority score; BCW represents the business criticality weight; SLAW represents the service level agreement (SLA) requirement weight; DW represents the deadline weight; CFW represents the call frequency weight; AETW represents the average execution time weight; RCW represents the resource consumption weight; ERW represents the error rate weight; HSW represents the historical stability weight; RCIW represents the resource competition intensity weight; BIW represents the suddenness factor weight.

[0095] Since the comprehensive priority is the weighted sum of the scores of multiple monitoring items in the real-time monitoring data, and the weights corresponding to the multiple monitoring items can be obtained based on historical data prediction, in some embodiments, to calculate the comprehensive priority, the cloud computing platform can also use a regression model trained based on historical data to dynamically predict weight parameters for optimization.

[0096] As can be seen, by calculating comprehensive priorities, the cloud computing platform can dynamically adjust function priorities based on predefined rules and real-time monitoring data, ensuring that critical tasks are prioritized, improving system responsiveness and resource utilization efficiency. For example, during promotional events, order processing requests surge, and payment requests, which have a higher priority, require prioritization to ensure transaction completion. By calculating comprehensive priorities, the payment function can be assigned a high priority, ensuring that payment requests are prioritized in high-concurrency scenarios and avoiding transaction failures due to delays. This improves the user experience, increases order conversion rates, and reduces operating costs.

[0097] S103: Load the target function into a multi-level task queue according to the comprehensive priority.

[0098] After calculating the overall priority, the cloud computing platform can load the target function into a multi-level task queue based on the overall priority. Specifically, the dynamic priority queue management module continuously monitors function call patterns and system resource status, dynamically adjusts the priority rule weights, calculates an overall priority score (FPS), and then dynamically allocates functions to the multi-level queue based on the priority score.

[0099] In some embodiments, the objective function is a function determined according to the comprehensive priority and the queue length of the multi-stage task queue. For example, if the length of the multi-stage task queue includes N tasks, then the first N function tasks can be determined as the objective function in descending order of the comprehensive priority.

[0100] In some embodiments, the cloud computing platform can also implement dynamic multi-level queue adjustment and cross-queue migration mechanisms. By monitoring queue length, wait time, system resource load and other indicators in real time, the scheduling time slice and capacity threshold of each priority queue can be dynamically adjusted to achieve adaptive threshold adjustment.

[0101] In some embodiments, the cloud computing platform can also automatically merge low-priority queues or split high-priority queues according to system load conditions. By dynamically splitting and merging multi-level task queues, the flexibility of queue management can be improved.

[0102] In some embodiments, the cloud computing platform can also introduce a cooling-down time window through cross-queue migration and cooling-down strategies to limit functions from frequently crossing queues in a short period of time, thereby alleviating unnecessary overhead caused by frequent migration. When performing cross-queue migration, the migration conditions can also be optimized, such as adding multi-dimensional trigger conditions, including: the cumulative waiting time of the function exceeds a certain threshold; tasks in the high-priority queue require additional resources due to sudden increases in traffic; the priority score of the function significantly increases or decreases in multiple consecutive calculations, etc.

[0103] To ensure scheduling fairness and critical business assurance, cloud computing platforms can implement a critical business assurance mode. In this mode, a Critical Business Queue (CBQ) can be added to allocate independent resource pools to specific critical business functions, ensuring fast scheduling at all times.

[0104] Cloud computing platforms can also provide a dynamic critical business switching mechanism, dynamically adjusting which functions are marked as critical based on real-time business traffic analysis. By introducing a priority-weighted fair queuing algorithm (WFQ), high-priority tasks can be quickly responded to while ensuring scheduling fairness and preventing long-term starvation of low-priority tasks.

[0105] In some embodiments, the cloud computing platform can also implement real-time algorithm calibration and anomaly detection. Based on the priority calculation result verification mechanism, the cloud computing platform can use rule verification and machine learning models to verify the priority calculation results in parallel, avoiding abnormal scheduling behavior caused by improper weight settings. To improve abnormal scheduling behavior, the cloud computing platform can identify abnormal patterns. By analyzing function call data in real time, it can identify potential abnormal patterns, such as a sudden drop in call frequency, an abnormal increase in resource consumption, etc., and trigger reassessment.

[0106] S104 , extracting a function task from the multi-level task queue, and executing startup function scheduling according to the cache status of the function task.

[0107] After loading the target function into the multi-level task queue, the cloud computing platform can also extract the function task from the multi-level task queue and execute the startup function scheduling based on the cache status of the function task. To implement the startup function scheduling, the cloud computing platform can deploy a scheduling module that can preferentially select functions from the high-priority queue, check the cache status, and execute the scheduling.

[0108] The startup function includes a first function and a second function. The first function is a function task reused from the cache; the second function is a newly created function task. To this end, when the scheduling module selects a function to be executed from the priority queue, it can synchronously query the distributed cache module. The distributed cache module can utilize an optimized scheduling algorithm, intelligent cache search strategy, and asynchronous cache update mechanism to maximize cache hit rate and scheduling efficiency.

[0109] In some embodiments, the distributed cache module is responsible for storing function code compilation results, execution environment snapshots, function input parameter patterns, resource requirements, and other metadata information, using a high-performance distributed in-memory data grid such as Hazelcast and Redis. This accelerates function calls and execution through efficient in-memory storage capabilities and optimized data structures.

[0110] A distributed in-memory database (Hazelcast) supports cross-node memory sharing and low-latency access. For example, cached content includes function code compilation results to avoid repeated compilation. The distributed in-memory database can store execution environment snapshots—the runtime state of container images—to achieve instant startup. The distributed in-memory database is also used to store metadata, such as function input parameter patterns and resource requirements.

[0111] In some embodiments, third-party libraries required by functions can be preloaded to form dependency libraries. Furthermore, when caching, cache locations can include distributed in-memory databases, such as Hazelcast, and local caches, meaning that some high-frequency functions retain cache copies locally on the compute node.

[0112] The distributed cache module can use multi-level caching and tiered eviction strategies to maximize the use of cache resources. These strategies monitor cache usage and record access information for each cache item (function task), including access time and frequency.

[0113] The distributed cache module can be based on a multi-level cache, dividing the cache into multiple levels. For example, the distributed cache module can include cache levels such as L1 cache, L2 cache, and L3 cache. Among them, the L1 cache adopts the LRU strategy, which has the characteristics of high priority and small capacity. The L2 cache adopts the LFU strategy, which has the characteristics of medium priority and medium capacity. The L3 cache adopts a hybrid strategy based on frequency and time decay, Time-aware LFU (TLFU), which has the characteristics of low priority and large capacity.

[0114] In some embodiments, in order to implement the scheduling of function tasks, the scheduling module can perform cache query through the distributed cache module. Figure 4 As shown, executing the startup function scheduling according to the cache status of the function task also includes periodically selecting the to-be-executed function from the multi-level task queue in order of priority, and searching for the relevant resources of the to-be-executed function in the distributed cache. If the relevant resources of the to-be-executed function are found in the distributed cache, the to-be-executed function is started using the execution environment in the cache to start the first function; if the relevant resources of the to-be-executed function are not found in the distributed cache, resources are allocated for the to-be-executed function, and an execution environment is constructed to start the second function.

[0115] The scheduling module can periodically select functions to be executed from the priority queue in order of priority and query the distributed cache. When executing the scheduling of the startup function, the scheduling module can first interact with the distributed cache module to query whether there are relevant resources for the function in the cache. If the relevant resources of the function to be executed are found in the cache, that is, the cache hits, the function is directly started using the execution environment in the cache. By directly reusing the execution environment in the cache, the function can be started quickly. If the relevant resources of the function to be executed are not found in the cache, that is, the cache misses, the function creation process is started, including allocating resources, building the execution environment, etc. By creating a new execution environment, the relevant information is cached after the function execution is completed.

[0116] In some embodiments, in order to periodically select functions to be executed from the multi-level task queue in order of priority, the scheduling module may obtain task requirement information, wherein the task requirement information includes at least one of the business characteristics of the system, the function call frequency, and the expected timeliness of response. A scheduling period parameter is then set based on the task requirement information, and based on the scheduling period parameter, functions to be executed are periodically selected from the multi-level task queue in order of the comprehensive priority.

[0117] The cloud computing platform can use function selection and cache query algorithms to enable the scheduling module to start the scheduling process according to a set time period, such as every 10 milliseconds. The set time period is a configurable parameter that can be set according to the specific application scenario.

[0118] After the scheduling cycle is set, a timer can be used to trigger the scheduling. The scheduling module then locates the highest-priority queue first. Next, the function tasks to be executed are retrieved sequentially, starting from the head of the high-priority queue. Using a first-in, first-out principle, functions in the high-priority queue are prioritized. For each retrieved function, a query request is sent to the distributed cache module to check the cache for the corresponding resources, as well as other metadata such as the function's input parameter pattern and resource requirements. The corresponding resources for the function to be executed can be snapshots of the execution environment, including container images and dependent libraries.

[0119] If a cache hit occurs, the corresponding function execution environment and other resources are found in the distributed cache module. The scheduling module can directly obtain these resources from the cache, quickly start the function, and record relevant information about the function execution, such as execution time and resource usage, for subsequent monitoring and feedback modules.

[0120] If the cache misses, the resource allocation module is called to create a new execution environment according to the function task creation process, including allocating computing resources and building the execution environment based on the function code. After the function executes, the updated relevant information, such as the execution environment status and code compilation results, is cached in the distributed cache module according to the cache update strategy for subsequent reuse.

[0121] The scheduling period parameter determines how often the scheduling module selects and schedules functions, such as every 10 milliseconds. This period can be adjusted based on factors such as the business characteristics of the cloud computing platform, the frequency of function calls, and the expected timeliness of response. If the business requires extremely high real-time performance, the period can be shortened; if system resources are limited and function calls are relatively infrequent, the period can be extended. The range of values ​​is generally from a few milliseconds to tens of milliseconds or even seconds, depending on the specific scenario.

[0122] Similarly, queue priority parameters define the priority values ​​or levels of functions within a queue based on factors such as business criticality, service-level agreement (SLA) requirements, and deadlines. For example, functions with high business criticality, strict SLA requirements, and tight deadlines are assigned the highest priority (e.g., level 1), followed by lower priority levels (e.g., level 2, level 3, and so on). These priority levels determine the position of functions within the corresponding multi-level queues and the order in which they are scheduled.

[0123] In some embodiments, in order to search for the relevant resources of the function to be executed in the distributed cache, the cloud computing platform can set a cache query timeout to improve query efficiency. That is, the scheduling module can send a query request to the distributed cache and record the waiting time after sending the query request. Then compare the waiting time with the pre-set time threshold. If the waiting time is less than or equal to the time threshold, it is marked that the relevant resources of the function to be executed are found in the distributed cache, that is, the cache hit; if the waiting time is greater than the time threshold, it is marked that the relevant resources of the function to be executed are not found in the distributed cache, that is, the cache miss.

[0124] By setting cache-related parameters, such as the cache query timeout, the scheduling module can set a maximum wait time (i.e., a time threshold) after sending a query request to the distributed cache. This can alleviate long scheduling process blockages caused by cache module failures or network issues. The time threshold can be appropriately set based on the network environment and cache module performance. For example, if it is set to 100 milliseconds, if no cache response is received after this time, the query is considered a failure and treated as a cache miss.

[0125] For example, the scheduling module can obtain configuration information such as scheduling time cycle parameters, queue priority parameters, and cache-related parameters by reading configuration files or system initialization parameters, and complete the initialization settings of the scheduling module itself and related modules based on these parameters, such as multi-level queue structure initialization, distributed cache module connection configuration, etc.

[0126] Each scheduling cycle is then initiated at intervals specified by the scheduling time period parameter. In each cycle, the priority queues are traversed, starting with the highest priority queue and retrieving pending function tasks in order. For each retrieved function task, a query request is sent to the distributed cache module, passing along the function's associated identifiers, including the function name, version number, and other unique identifiers, for accurate cache querying.

[0127] The priority queue is checked at a set interval, and function tasks are retrieved starting with the highest priority queue. A query request is then sent to the distributed cache module to check for a cache hit. If a cache hit is found, the execution environment is retrieved from the cache and the function is started. Execution-related information, such as the function's execution time, is recorded for subsequent monitoring and feedback. If a cache hit is found, the resource allocation module is called to create a new execution environment, start the function, and cache the relevant information after the function completes.

[0128] In some embodiments, intelligent snapshot lifecycle management can be used to dynamically adjust snapshot lifecycles and retirement policies based on function call frequency, resource consumption, and version updates.

[0129] It should be noted that the computing nodes of the cloud computing platform can be triggered for execution by the scheduling module, that is, the scheduling module selects the function instance to be executed based on the dynamic priority queue and cache status. If the cache hits, that is, the function code and execution environment exist in the distributed cache, the scheduling module will directly call the idle computing node and quickly start the function by reusing the execution environment snapshot (container image, dependency library) in the cache. If the cache misses, the computing node needs to pull the function code from the persistent storage (such as the code repository), create a new execution environment (such as starting a container), execute the function after initialization, and cache the execution environment snapshot to the distributed storage. The resources (CPU, memory) of the computing node are dynamically allocated by the scheduling module, giving priority to ensuring the resource requirements of high-priority functions.

[0130] S105 , recording the indicator data after the startup function is executed, and adjusting the function scheduling strategy according to the indicator data and the real-time monitoring data.

[0131] After executing the startup function scheduling, the cloud computing platform can also record the indicator data after the startup function execution and adjust the function scheduling strategy based on the indicator data and real-time monitoring data. The indicator data includes execution performance, resource usage data, and historical records; the function scheduling strategy includes the priority rules, the calculation weight parameters of the comprehensive priority, and the queue management strategy of the multi-level task queue.

[0132] To this end, the cloud computing platform can also deploy a monitoring and feedback module. This module is responsible for real-time monitoring of key indicators such as resource utilization, function execution performance, and queue status across the entire cloud computing platform. It then feeds the processed monitoring data back to the dynamic priority queue management module and distributed cache system to achieve adaptive scheduling and cache optimization.

[0133] The monitoring feedback module can adopt multi-dimensional monitoring, intelligent anomaly detection and predictive feedback mechanisms to improve the accuracy of monitoring and the timeliness of feedback. In some embodiments, the monitoring feedback module can monitor system resource utilization, function execution performance, queue status and other indicators in real time. Probe technology or built-in monitoring components of the system can be used to collect multi-dimensional data such as system resource utilization, function execution performance, queue status, etc. in real time. Among them, system resource utilization can include parameters such as CPU, memory, network bandwidth, etc.; function execution performance can include parameters such as execution time, response time, error rate, etc.; queue status can include parameters such as queue length and queue waiting time. After collecting and obtaining the indicator data, the indicator data can be transmitted to the dynamic priority queue management module so that the dynamic priority queue management module can adjust the function priority and caching strategy according to data changes.

[0134] In some embodiments, the monitoring feedback module can also periodically obtain multi-dimensional indicator data according to the collection time interval parameter when recording the indicator data after the execution of the startup function and adjusting the scheduling strategy according to the indicator data and the real-time monitoring data. The indicator data includes system resource utilization, function execution performance and queue status. The indicator data is then preprocessed according to the preset processing items, and the preset processing items include one or more combinations of removing outliers, filling missing values ​​and format conversion. Based on the preprocessed indicator data, the priority rules, the calculation weight parameters of the comprehensive priority and the queue management strategy of the multi-level task queue are adjusted.

[0135] After collecting and obtaining the indicator data, the monitoring feedback module can remove outliers and erroneous data from the collected data based on the data preprocessing algorithm. For example, if the CPU usage rate collected is obviously beyond the reasonable range due to probe failure or other reasons, such data will be eliminated. For function execution performance data with extremely short or extremely long abnormal execution time due to special network fluctuations and other reasons, it can be filtered by setting a threshold range. As for missing values, if some data collection points fail to obtain the corresponding data due to temporary failures or other reasons, they can be filled with the valid data collected last time, or filled according to the average value, median, etc. of the same type of indicators to ensure data integrity.

[0136] After feeding indicator data and real-time monitoring data back to the dynamic priority queue management module, the dynamic priority queue management module can dynamically adjust function priorities and caching strategies to achieve adaptive scheduling optimization. To this end, after receiving monitoring data, the dynamic priority queue analyzes it based on built-in algorithms and rules. For example, if it finds that the average execution time of a function continues to increase and its corresponding business criticality is high, it will use a predefined priority adjustment algorithm to comprehensively consider factors such as the magnitude of the change in execution time and the weight of business importance, and increase the priority of the function. If it monitors that the call frequency of a certain type of function has dropped significantly, its priority will be lowered accordingly.

[0137] For cache policy adjustments, if the cache hit rate is found to be too low based on monitoring data, you may decide to adjust the cache preheating mechanism, such as increasing the preheating function range, increasing the preheating frequency, etc., or change the cache replacement policy, such as switching from LRU to LFU, etc. This can be determined based on cache resource utilization and changes in function access patterns.

[0138] For example, the monitoring and feedback module in a cloud computing platform can set data collection interval parameters to collect system metrics at regular intervals (e.g., every second). The collected data can then be preprocessed, such as by cleaning and format conversion. The target format to which the raw data should be converted, such as JSON or XML, is specified, along with specific format specifications, including data field naming conventions and data structure organization, to ensure consistent understanding and processing of the data across all modules.

[0139] The processed data is then transmitted to the dynamic priority queue management module. The dynamic priority queue management module can adjust the priority weight parameters. Based on the adjustment information fed back by the dynamic priority queue management module, such as when a function's priority is increased and cache preheating is required, the module notifies related modules such as the distributed cache module to take corresponding actions.

[0140] By applying the technical solutions of the above-mentioned embodiments, the distributed function scheduling optimization method under a serverless architecture provided by the embodiments of the present application can effectively alleviate the problems of function cold start delay and low scheduling efficiency under a serverless architecture through multi-dimensional dynamic priority scheduling, deep caching of environment snapshots and predictive adaptive optimization. Among them, multi-dimensional dynamic priority scheduling is different from the priority allocation of static rules. By introducing multi-dimensional dynamic priority calculation, real-time monitoring data such as call frequency, execution time, resource consumption, context information, etc. are combined with predefined rules such as SLA, deadline, and business criticality to construct a dynamic priority vector, which more comprehensively and accurately reflects the actual importance of the function at a specific moment.

[0141] By building a priority vector model, the overall priority is no longer a single value, but a vector containing multiple dimensions, such as business priority, call frequency, resource consumption ratio, execution time variance, and context relevance. Each dimension can be assigned a different weight and dynamically adjusted based on monitoring data.

[0142] Adaptive weight adjustment introduces a feedback loop to adjust the weights of various dimensions based on system performance indicators such as average response time, queue length, and resource utilization, ensuring that priority calculations are more closely aligned with actual operating conditions. Furthermore, a hierarchical queue structure is implemented based on multi-level priority queue groups, with each group corresponding to a different scheduling strategy, such as latency-sensitive queues, throughput-optimized queues, and burst-response queues.

[0143] It's also possible to combine rules and machine learning algorithms to dynamically adjust the position of functions in each queue, achieving dynamic priority adjustment. For example, the weighted moving average method can smooth monitoring data, and decision trees / neural networks can predict a function's future resource requirements and call patterns.

[0144] By setting up a priority decay and promotion mechanism, we can avoid starvation of low-priority functions by introducing priority decay, and avoid long waiting times for high-priority functions by introducing priority promotion.

[0145] This method also supports deep caching of environment snapshots. By building a layered caching system, employing a multi-level caching architecture and caching strategies such as local caching, distributed memory caching, and persistent storage, the distributed cache module caches function code and a complete snapshot of the execution environment, including pre-configured container images, pre-loaded dependency libraries, and pre-warmed runtime context such as established database connections and cached data. This deep caching significantly reduces function startup time.

[0146] Furthermore, as a refinement and extension of the specific implementation of the above embodiment, in order to fully illustrate the specific implementation process of this embodiment, some embodiments of this application also provide a distributed function scheduling optimization method under a serverless architecture, such as Figure 5 As shown, based on the above embodiment, the method can also introduce a cache preheating mechanism, that is, the method includes:

[0147] S201, obtaining historical function call data;

[0148] S202, extracting function call parameters within a preset period from the historical function call data, the function call parameters including function name, call time, and call frequency;

[0149] S203: predicting a preheating function based on the function call parameters, and loading the preheating function into a distributed cache;

[0150] S204. Load the function-related information corresponding to the preheating function into a distributed memory database, where the function-related information includes at least one of a function code compilation result, an execution environment snapshot, a function input parameter mode, and a function resource requirement.

[0151] Cloud computing platforms can preload the code and execution environment of commonly used functions into the cache based on cache preheating algorithms, predicted function call patterns, historical data, user behavior, and contextual information, thereby reducing cold start delays.

[0152] When executing the warm-up algorithm, the cloud computing platform can first collect and analyze data, that is, collect historical function call data through the acquisition module, and obtain detailed information about the function calls within a certain period of time from the system records, including function name, call time, call frequency and other information.

[0153] We then use data analysis methods such as time series analysis to analyze call patterns and process the collected data to identify functions that are likely to be frequently called in the recent past. For example, by analyzing the trend of function call frequency over time, if a function shows a clear upward trend in call frequency within a specific time period, and this trend continues until approximately the current time, we can determine that the function is likely to be frequently called.

[0154] After acquiring data and obtaining analysis results, we can filter out a list of preheating functions that require preheating based on the analysis results. These preheating functions are then loaded into the cache. For each of these preheating functions, the corresponding function code compilation results, execution environment snapshots, function input parameter patterns, resource requirements, and other information are loaded into a distributed in-memory database such as Hazelcast. This storage is accomplished using the distributed in-memory database's efficient in-memory storage capabilities, paving the way for subsequent rapid function calls.

[0155] In some embodiments, during data collection and analysis, user behavior data, such as user access paths, search history, and purchase history, can be combined to build a user behavior model. This model can then be used to predict the function calls that a user may trigger next. Furthermore, system context information, such as the current time, holidays, and promotional events, can be combined to predict changes in function call patterns. Q-learning can be used to dynamically adjust the preheating strategy based on metrics such as the preheating hit rate and performance improvement.

[0156] Therefore, when predicting the preheating function based on the function call parameters, the cloud computing platform can first obtain user behavior data, and the user behavior data includes user access paths, search records, and execution history. Then, a user behavior model is constructed based on the user behavior data, and the user behavior model is used to calculate the probability of the function task being called. Then, the system context information is obtained, and a preheating strategy is determined based on the system context information, and the preheating strategy is dynamically adjusted based on the system context information and the deep Q learning algorithm. And the preheating function list is generated according to the preheating strategy, and the preheating function list includes a preset number of the preheating functions screened based on the preheating strategy; the preheating function is a function task whose probability of being called is greater than or equal to the probability threshold and meets the preheating strategy; the preset number is set according to the queue length of the multi-level task queue.

[0157] By building a user behavior model and determining a warm-up strategy based on system context information, predictive adaptive optimization can be achieved. That is, by introducing a machine learning-based prediction model, historical function call patterns, resource consumption, and business trends are analyzed to predict function call conditions and resource requirements in the future, thereby achieving proactive performance optimization.

[0158] Cloud computing platforms can predict resource demands based on call patterns and demand, employing time series analysis, machine learning algorithms, and statistical models. Predictions-based cache preheating and resource preallocation can pre-load snapshots into the cache and pre-allocate computing resources based on the predicted results. This facilitates dynamic resource adjustment and elastic scaling. This allows for dynamic resource allocation adjustments to compute nodes based on predicted resource demands, increasing the number of compute nodes or resource quotas in advance during periods of predicted high load.

[0159] In some embodiments, in order to obtain more accurate model prediction results, the cloud computing platform can also perform model training and iterative optimization, that is, by continuously collecting system operation data, training and optimizing the prediction model, and improving prediction accuracy.

[0160] By applying the technical solutions of the above embodiments, the method can set up a cache preheating mechanism to implement predictive analysis of function call patterns, pre-loading function resources that may be frequently called into the cache, further reducing the possibility of cold starts. Through proactive cache management strategies, the system's response speed and throughput can be significantly improved in high-concurrency scenarios. Furthermore, predictive adaptive optimization can be combined for more intelligent preheating.

[0161] Furthermore, as a refinement and extension of the specific implementation of the above embodiment, in order to fully illustrate the specific implementation process of this embodiment, some embodiments of this application also provide a distributed function scheduling optimization method under a serverless architecture, such as Figure 6As shown, based on the above embodiment, the method can also introduce a differential snapshot update strategy, that is, the method includes:

[0162] S301, collecting update information, wherein the update information is information generated after the function task is executed; the update information includes updated code compilation results and the latest status information of the execution environment;

[0163] S302: Obtaining benchmark parameters of a cache update strategy, where the benchmark parameters include at least one of a version number of a function task, an execution timestamp, and a degree of change impact;

[0164] S303: Update the function task in the distributed cache based on the benchmark parameter and the update information.

[0165] After the cloud computing platform executes the scheduling of the startup function, the function tasks can be executed in sequence according to the scheduling process. Information can be obtained after the function is executed. That is, when a function is completed, the updated related information generated by it is collected, such as the updated code compilation results, the latest status information of the execution environment, etc.

[0166] Then, the update decision is made according to the established cache update strategy. For example, based on version number comparison, if the version number of the newly executed function is higher than the corresponding version number stored in the cache, the cache needs to be updated. Alternatively, based on timestamp update, if the new execution timestamp is newer and meets the preset update conditions, the cache is determined to be updated.

[0167] In some embodiments, the impact of code changes on function execution results and performance can also be assessed based on change impact assessment. When the change is likely to result in significant performance improvement or functional changes, a cache update is performed.

[0168] Then, based on the determined update decision, the cache update operation is performed, and the updated function-related information obtained is updated to the distributed cache according to the corresponding rules, ensuring that the function information in the cache is always up-to-date and available, so that it can be accurately reused when called again in the future.

[0169] In some embodiments, for large functions or images, an incremental update strategy can be adopted, that is, only the changed parts are updated.

[0170] As can be seen, by introducing a differential snapshot update strategy, incremental updates and differential storage can be implemented, updating only the changed parts of the snapshot, reducing overhead and network transmission volume. For example, using layered mirroring or differential backup.

[0171] In some embodiments, in order to better utilize the distributed cache space, the cloud computing platform can also make the most of the limited cache resources by replacing the memory resources of the function tasks. Figure 7 As shown, the method further includes:

[0172] S401, recording access information of the function task during cache usage, wherein the access information includes access time and access frequency;

[0173] S402, setting a hierarchical elimination strategy for a multi-level cache according to the accessed information;

[0174] S403: When the cache space of the current level of the multi-level cache is full, searching for an object to be replaced according to the hierarchical elimination strategy of the current level;

[0175] S404: Remove the object to be replaced from the multi-level cache, and load cache information corresponding to the new function resource into the multi-level cache.

[0176] Through replacement judgment, when the cache space at a certain level is full and new function resources need to be loaded into the cache, the objects to be replaced can be found according to the elimination strategy of the cache at that level. For some caches that have little impact on performance, a probabilistic elimination strategy is adopted. Among them, the cache replacement strategy-related parameters may include the time accuracy parameters for recording the access information of each cache item, as well as reinforcement learning-related parameters such as learning rate, discount factor, exploration rate, etc. The cache update strategy-related parameters may include specific rules for version number comparison, timestamp accuracy and time interval thresholds, as well as change impact assessment thresholds, etc.

[0177] Then, a replacement operation is performed on the found object to be replaced. That is, the found cache item to be replaced can be removed from the cache, and the cache information corresponding to the new function resource can be loaded into the cache to complete the cache replacement operation, so as to ensure that the cache always stores the most commonly used and most valuable function resources, and maximize the use of limited cache resources.

[0178] In some embodiments, the distributed cache module can use historical data windows of different sizes based on different analysis purposes to obtain different amounts of historical data. The distributed cache module can also set cache space parameters, including the total cache space and capacity quotas based on different cache tiers.

[0179] For example, after the cloud computing platform is started, the cache warm-up process can be started based on historical data or pre-configured information, and function resources that may be frequently called can be loaded into the distributed memory database to initialize the reinforcement learning model.

[0180] After a function completes execution, its associated updated code compilation results and execution environment information are updated to the distributed cache according to a specific cache update strategy. No longer needed cache items are then eliminated based on the cache elimination strategies and probabilistic elimination mechanisms at each level. A reinforcement learning model continuously adjusts the warm-up strategy based on system performance indicators.

[0181] Furthermore, as a specific implementation of the distributed function scheduling optimization method under the serverless architecture described in the above embodiment, some embodiments of this application also provide a distributed function scheduling optimization device under the serverless architecture, such as Figure 8 As shown, the device includes:

[0182] Acquisition module, used to obtain priority rules and real-time monitoring data;

[0183] A dynamic priority management module, configured to calculate a comprehensive priority based on the priority rules and the real-time monitoring data, wherein the comprehensive priority is a weighted sum of scores of multiple monitoring items in the real-time monitoring data;

[0184] A distributed cache module, configured to load a target function into a multi-level task queue according to the comprehensive priority, wherein the target function is a function determined according to the comprehensive priority and a queue length of the multi-level task queue;

[0185] A scheduling module is used to extract function tasks from the multi-level task queue and execute startup function scheduling according to the cache status of the function tasks, wherein the startup function includes a first function and a second function, the first function is a function task reused from the cache; the second function is a newly created function task;

[0186] A monitoring feedback module is used to record the indicator data after the execution of the startup function, and to adjust the function scheduling strategy based on the indicator data and the real-time monitoring data. The indicator data includes execution performance, resource usage data and historical records; the function scheduling strategy includes the priority rules, the calculation weight parameters of the comprehensive priority and the queue management strategy of the multi-level task queue.

[0187] like Figure 9 As shown, the distributed function scheduling and optimization device in the serverless architecture is a distributed system consisting of multiple modules, including a processor, memory, input / output interfaces, and communication interfaces. The processor is responsible for executing the logic of each module, including priority calculation, cache query, execution environment management, resource allocation, and monitoring data analysis. It can be a general-purpose CPU, GPU, or other specialized processor.

[0188] Memory is used to store various data. This memory can include main memory, distributed cache, and persistent storage. Main memory is used to store running program code, function code, execution environment snapshots, metadata, and more. Distributed cache is used to build distributed in-memory databases (such as Hazelcast and RedisCluster) or storage systems compatible with memory semantics. It can store function code compilation results, execution environment snapshots, dependency libraries, input parameter patterns, and more. Persistent storage is used to store priority rule configurations, monitoring data history, and more.

[0189] The input / output interface is used to interact with external devices or systems. It can include a function call request receiving interface, a monitoring data collection interface, a management interface, and a log output interface. The function call request receiving interface receives function call requests from clients. The monitoring data collection interface receives monitoring data from the monitoring system. The management interface allows administrators to configure priority rules, caching policies, and other features. The log output interface outputs system logs and monitoring data.

[0190] Communication interfaces are used for communication between modules and with external systems. They include inter-module communication interfaces, interfaces for communication with distributed caches, and interfaces for communication with monitoring systems. Inter-module communication interfaces can use methods such as RPC (remote procedure calls) or message queues for inter-module communication. Communication interfaces with distributed caches can use the client API provided by the distributed cache. Communication interfaces with monitoring systems can use APIs or protocols provided by the monitoring system.

[0191] The distributed function scheduling optimization device within the entire serverless architecture can address latency issues caused by cold start functions in serverless architectures, as well as scheduling efficiency issues in high-concurrency scenarios. By dynamically adjusting function priorities and leveraging distributed caching, it prioritizes high-priority functions with cache hits, thereby improving system throughput and resource utilization.

[0192] like Figure 10 As shown in the figure, in specific applications, users can submit tasks through the client, that is, the client initiates a function call request. The scheduling module can receive the submitted tasks and query the distributed cache module based on the received tasks. The distributed cache module then performs cache query and processing. If the required function and its execution environment hit the cache (including preheating the resource pool), the scheduler obtains it directly from the cache and hands it to the function execution engine for execution, thereby shortening the startup time. If the cache does not hit, the scheduler triggers the resource preparation module, creates a new execution environment according to the creation process, and then hands it to the function execution engine for execution. After execution is completed, the relevant information will be cached in the distributed cache for subsequent call reuse.

[0193] The device then uses the function execution engine to execute the function, that is, the function execution engine actually executes the function code. After the function execution is completed, the result is returned to the client.

[0194] The monitoring feedback module in the device can monitor the system operation status in real time, such as resource utilization, function execution performance, etc., and feed back the monitoring data to the dynamic priority queue management module for adjusting function priority and caching strategy.

[0195] The device further includes a dynamic priority queue management module which can dynamically adjust the priority of the function according to monitoring data and predefined priority rules, and provide the priority information to the scheduler.

[0196] For example, the device can calculate and adjust function priorities based on predefined rules and real-time monitoring data, placing functions into corresponding multi-level queues. Function priorities are dynamically calculated and adjusted based on pre-set rules and real-time monitoring data obtained from the monitoring and feedback module. Functions are then placed into different queues based on their priorities. This allows function tasks to be pushed based on queue priority. Specifically, the dynamic priority queue management module pushes pending function tasks to the scheduling module based on queue priority.

[0197] After receiving the function task, the scheduling module first queries the distributed cache and determines whether the required function code and execution environment have been cached. Then the query result (hit or miss) is returned to the scheduling module. The scheduling module is made to perform branch processing, that is, the process is divided into two branches according to the cache hit situation. When the cache hits, the execution environment in the cache is reused to start the function. The scheduling module directly obtains the cached execution environment from the distributed cache, quickly starts the function execution, and avoids the delay of cold start. When the cache misses, a new execution environment is created according to the creation process, including downloading the function code, configuring the running environment, loading dependent libraries, etc.

[0198] After being scheduled by the scheduling module, the function can be executed in the function execution engine. After the function is executed, the function execution engine caches the function code, execution environment snapshot, input parameter mode and other related information in the distributed cache for subsequent call reuse.

[0199] The device also uses a monitoring and feedback module to monitor resource utilization and function execution performance in real time, including CPU, memory, and network utilization, as well as performance indicators such as function execution time and error rate. The monitoring and feedback module feeds the collected monitoring data back to the dynamic priority queue management module, which dynamically adjusts function priorities and the distributed cache's caching strategy, forming a closed-loop optimization process.

[0200] By deploying an acquisition module, a dynamic priority management module, a distributed cache module, a scheduling module, and a monitoring feedback module within the device, the device achieves extremely fast cold starts. By pre-warming and caching function code, environments, and dependencies in a distributed in-memory database, high-priority functions with cache hits are prioritized during scheduling. This eliminates repeated loading and initialization during cold starts, enabling almost instant function startup and reducing cold start latency by over 80%. For example, in e-commerce promotion scenarios, product recommendation functions can respond to user requests instantly.

[0201] Furthermore, the device can achieve higher throughput. A dynamic priority queue management module rationally arranges execution order based on factors such as function importance and resource requirements, prioritizing high-priority functions and preventing resources from being occupied by lower-priority functions for extended periods. Combined with distributed caching, this reduces resource waiting time during function startup and execution, increasing system throughput by 50%-200%. For example, social platforms can more quickly process massive amounts of messages and user status updates.

[0202] This device also provides improved resource utilization by monitoring function call patterns and resource requirements in real time, dynamically adjusting function priorities and precisely allocating resources to the functions that need them most. The cache system reuses function resources, reducing duplicate allocation and initialization, improving CPU and memory utilization by 30%-60%, and reducing cloud computing costs. For example, data-intensive applications can use resources more efficiently, avoiding waste.

[0203] The device also ensures rapid response times for critical services. A dynamic prioritization mechanism ensures that critical business functions (such as financial transactions and payment confirmations) receive high priority and are prioritized for execution by retrieving resources from the cache. In high-concurrency scenarios, response times for critical business functions can be controlled to milliseconds, ensuring normal business operation and service quality. For example, online payment systems can quickly complete payment confirmations even during peak hours.

[0204] By applying the technical solutions of the above-mentioned embodiments, the embodiments of the present application provide a distributed function scheduling optimization device under a serverless architecture, which can adopt a dynamic priority scheduling strategy to dynamically adjust the priority of function tasks according to predefined priority rules and real-time monitoring data, ensuring that critical tasks are given priority and improving the system response speed and resource utilization efficiency. A distributed cache is used to store function tasks, execution environment snapshots and dependent libraries, and a preheating mechanism is supported to reduce cold start delays and improve the execution efficiency of function tasks. Through the monitoring feedback mechanism, indicator data related to system operation is collected in real time, and the priority and cache strategy are dynamically adjusted according to the indicator data to achieve adaptive optimization and improve system stability and robustness. Through dynamic priority scheduling and cache optimization, the system throughput and resource utilization are improved.

[0205] It should be noted that for other corresponding descriptions of the various functional units involved in the distributed function scheduling optimization device under a server-less architecture provided in an embodiment of the present application, reference can be made to the corresponding descriptions in the distributed function scheduling optimization method under a server-less architecture provided in the above embodiment, and no further details will be given here.

[0206] The embodiment of the present application also provides a computer device, which can be specifically a personal computer, a server, a network device, etc. The computer device includes a bus, a processor, a memory and a communication interface, and may also include an input and output interface and a display device. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store location information. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the steps in each method embodiment are implemented.

[0207] Those skilled in the art will understand that the structure of the above-mentioned computer device is only a partial structure related to the solution of the present application and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components, or combine certain components, or have a different component arrangement.

[0208] In one embodiment, a computer-readable storage medium is further provided. The computer-readable storage medium may be non-volatile or volatile, and stores a computer program thereon. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0209] In one embodiment, a computer program product is further provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0210] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0211] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing related hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods.

[0212] Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc.

[0213] Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0214] The database involved in each embodiment provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchain. The processor involved in each embodiment provided herein may be, but is not limited to, a general-purpose processor, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, and the like.

[0215] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0216] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A distributed function scheduling optimization method under a serverless architecture, characterized in that: The method comprises: Obtain priority rules and real-time monitoring data; Calculating a comprehensive priority based on the priority rule and the real-time monitoring data, the comprehensive priority being a weighted sum of scores of multiple monitoring items in the real-time monitoring data; Loading a target function into a multi-level task queue according to the comprehensive priority, wherein the target function is a function determined according to the comprehensive priority and a queue length of the multi-level task queue; Extracting a function task from the multi-level task queue, and executing startup function scheduling according to the cache status of the function task, the startup function includes a first function and a second function, the first function is a function task reused from the cache; the second function is a newly created function task; Record the indicator data after the execution of the startup function, and adjust the function scheduling strategy according to the indicator data and the real-time monitoring data, the indicator data including execution performance, resource usage data and historical records; the function scheduling strategy including the priority rules, the calculation weight parameters of the comprehensive priority and the queue management strategy of the multi-level task queue; the function scheduling strategy is used to introduce a feedback loop through adaptive weight adjustment, and adjust the weights of each dimension according to the system performance indicators, so that the priority calculation fits the actual operating conditions; the system performance indicators include average response time, queue length, and resource utilization; the function scheduling strategy is also used to form a grouped hierarchical queue structure based on a multi-level priority queue group, each group of queues corresponds to a different function scheduling strategy, and the grouped hierarchical queue structure includes a delay-sensitive queue group, a throughput-optimized queue group, and a burst traffic response queue group.

2. The method according to claim 1, characterized in that The method further comprises: Get historical function call data; Extracting function call parameters within a preset period from the historical function call data, the function call parameters including function name, call time, and call frequency; Predicting a preheating function based on the function call parameters, and loading the preheating function into a distributed cache; Function-related information corresponding to the preheating function is loaded into a distributed memory database, where the function-related information includes at least one of a function code compilation result, an execution environment snapshot, a function input parameter mode, and a function resource requirement.

3. The method according to claim 2, characterized in that Predicting a warm-up function based on parameters of the function being called includes: Obtaining user behavior data, including user access paths, search records, and execution history; Building a user behavior model based on the user behavior data, wherein the user behavior model is used to calculate the probability of the function task being called; Obtaining system context information, and determining a preheating strategy based on the system context information, wherein the preheating strategy is dynamically adjusted based on the system context information and a deep Q learning algorithm; The preheating function list is generated according to the preheating strategy, and the preheating function list includes a preset number of the preheating functions screened based on the preheating strategy; the preheating function is a function task whose probability of being called is greater than or equal to a probability threshold and satisfies the preheating strategy; the preset number is set according to the queue length of the multi-level task queue.

4. The method according to claim 2, characterized in that The method further comprises: Collect update information, which is information generated after the function task is executed; the update information includes updated code compilation results and the latest status information of the execution environment; Obtaining a baseline parameter of a cache update strategy, the baseline parameter including at least one of a version number of a function task, an execution timestamp, and a degree of change impact; The function task in the distributed cache is updated based on the benchmark parameter and the update information.

5. The method according to claim 1, wherein The method further comprises: Recording access information of the function task during cache usage, wherein the access information includes access time and access frequency; Setting a hierarchical elimination strategy for the multi-level cache according to the accessed information; When the cache space of the current level of the multi-level cache is full, searching for an object to be replaced according to the hierarchical elimination strategy of the current level; The object to be replaced is removed from the multi-level cache, and cache information corresponding to the new function resource is loaded into the multi-level cache.

6. The method according to claim 1, wherein The startup function scheduling is performed according to the cache status of the function task, including: Periodically selecting functions to be executed from the multi-level task queue in order of priority; Searching for related resources of the function to be executed in the distributed cache; If relevant resources of the to-be-executed function are found in the distributed cache, the to-be-executed function is started using the execution environment in the cache to start the first function; If relevant resources of the function to be executed are not found in the distributed cache, resources are allocated for the function to be executed, and an execution environment is constructed to start the second function.

7. The method according to claim 6, characterized in that Periodically selecting functions to be executed from the multi-level task queue in order of priority, including: Acquiring task requirement information, wherein the task requirement information includes at least one of a business characteristic of the system, a function call frequency, and an expected response timeliness; Setting scheduling cycle parameters according to the task requirement information; Based on the scheduling period parameters, functions to be executed are periodically selected from the multi-level task queue in the order of the comprehensive priority.

8. The method according to claim 6, characterized in that Searching for resources related to the function to be executed in the distributed cache includes: Sending a query request to the distributed cache; Recording the waiting time after sending the query request; If the waiting time is less than or equal to the time threshold, mark that the relevant resources of the to-be-executed function are found in the distributed cache; If the waiting time is greater than a time threshold, it is marked that no relevant resources of the to-be-executed function are found in the distributed cache.

9. The method according to claim 1, characterized in that Recording indicator data after the startup function is executed, and adjusting the scheduling strategy according to the indicator data and the real-time monitoring data, including: Periodically acquire multi-dimensional indicator data based on the collection interval parameter, including system resource utilization, function execution performance, and queue status; Performing preprocessing on the indicator data according to preset processing items, wherein the preset processing items include one or more combinations of removing outliers, filling missing values, and format conversion; The priority rules, the calculation weight parameters of the comprehensive priority and the queue management strategy of the multi-level task queue are adjusted based on the preprocessed indicator data.

10. A distributed function scheduling optimization device under a serverless architecture, characterized in that: The device comprises: Acquisition module, used to obtain priority rules and real-time monitoring data; A dynamic priority management module, configured to calculate a comprehensive priority based on the priority rules and the real-time monitoring data, wherein the comprehensive priority is a weighted sum of scores of multiple monitoring items in the real-time monitoring data; A distributed cache module, configured to load a target function into a multi-level task queue according to the comprehensive priority, wherein the target function is a function determined according to the comprehensive priority and a queue length of the multi-level task queue; A scheduling module is used to extract function tasks from the multi-level task queue and execute startup function scheduling according to the cache status of the function tasks, wherein the startup function includes a first function and a second function, the first function is a function task reused from the cache; the second function is a newly created function task; A monitoring feedback module is used to record the indicator data after the execution of the startup function, and to adjust the function scheduling strategy according to the indicator data and the real-time monitoring data, wherein the indicator data includes execution performance, resource usage data and historical records; the function scheduling strategy includes the priority rules, the calculation weight parameters of the comprehensive priority and the queue management strategy of the multi-level task queue; the function scheduling strategy is used to introduce a feedback loop through adaptive weight adjustment, and adjust the weights of each dimension according to system performance indicators, so that the priority calculation fits the actual operating conditions; the system performance indicators include average response time, queue length, and resource utilization; the function scheduling strategy is also used to form a grouped hierarchical queue structure based on a multi-level priority queue group, each group of queues corresponds to a different function scheduling strategy, and the grouped hierarchical queue structure includes a delay-sensitive queue group, a throughput-optimized queue group, and a burst traffic response queue group.

Citation Information

Patent Citations

  • Cluster system comprehensive scheduling energy saving method and device

    CN105528054A

  • Resource processing method and device, server and storage medium

    CN117149392A