Server non-awareness machine learning inference system and method based on WebAssembly and container mixed deployment

By adopting a hybrid deployment method of WebAssembly and containers in server-free perception computing, combining clustering deep reinforcement learning decision module and running instance recycling module, the frequent cold start problems caused by low resource allocation efficiency and unpredictability in the existing technology are solved, efficient resource configuration and instance management are achieved, and management costs are reduced.

CN120123097APending Publication Date: 2025-06-10SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510287552.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In the existing servers' unaware computing, machine learning inference has problems such as low resource configuration efficiency, unpredictability of requests, frequent cold starts, secure isolation and high resource management costs.

Method used

The server-free machine learning inference system based on WebAssembly and container hybrid deployment is adopted, including WebAssembly-assisted hybrid runtime module, clustering deep reinforcement learning decision module and running instance recycling module. Through WebAssembly's fast cold start and container resource management characteristics, efficient resource configuration and instance management are achieved.

Benefits of technology

It improves the resource configuration efficiency and computing performance of server-free computing, reduces the cost of instance management and resource configuration, and achieves efficient and accurate response to highly dynamic time-sensitive user computing needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123097A_ABST
    Figure CN120123097A_ABST
Patent Text Reader

Abstract

The invention discloses a server non-perceptual machine learning inference system and method based on WebAssembly and container mixed deployment, and the system comprises a WebAssembly-assisted mixed runtime module which is used for constructing WebAssembly and containerization examples according to user task demands and task calculation characteristics, and constructing a scheduling management table of data between the WebAssembly and a container according to the linear storage characteristics of the WebAssembly; comprising a clustering deep reinforcement learning decision-making module which completes efficient and accurate configuration of system computing resources and task demand resources, and optimizes deep reinforcement learning decision-making steps through a clustering thought, thereby reducing decision-making cost and improving decision-making efficiency; and the operation instance recovery module is used for quickly judging an optimal processing mode of the operation instance by utilizing historical request characteristics and equipment affinity difference during operation, so that the management cost during operation is reduced. According to the method, efficient and accurate response of the high-dynamic time-sensitive user computing requirements is realized, and reference and help are provided for resource allocation decision making of the high-dynamic time-sensitive computing requirements by a server non-perception computing platform.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer information computing, and particularly relates to the resource allocation of serverless machine learning inference, and mainly relates to a serverless machine learning inference system and method based on hybrid deployment of WebAssembly and containers. Background Art

[0002] Serverless computing is known for its characteristics of management-free deployment and pay-per-use, and has become a transformative force in building and deploying machine learning (ML) models. Serverless ML inference, which means deploying ML inference tasks in isolated containers (i.e., serverless functions) and instantiating them as needed, is widely adopted in large companies and startups. Typical applications include network video coding, online super-resolution, etc. Many well-known Internet companies and open source communities also support various serverless inference platforms, such as AWS SageMaker, Azure ML, and Knative.

[0003] Serverless inference workloads exhibit the following two characteristics: (1) ML inference functions are usually integrated into large-scale online services with strict real-time service level objectives (SLOs). (2) Serverless workloads are highly unpredictable and may surge by an order of magnitude within a minute, forcing the platform to immediately spawn new instances (i.e., cold starts). Specifically, existing research has been dedicated to either real-time or resource efficiency, or their optimal balance. For example, various efforts have been advocated to meet application service standards. One approach is to pre-warm running instances, and keeping active serverless instances ensures timely inferences when incoming traffic surges. This sacrifices resource efficiency, especially when requests are highly unpredictable. Another approach is to reuse existing containers to execute multiple instances, which may break the security isolation between instances and requires additional resources for fine-grained management of instances. On the other hand, batch processing mechanisms are widely adopted to improve the resource efficiency of ML inference. However, the queuing delay for forming a suitable combination queue conflicts with strict service standards and low latency requirements. Fortunately, the emergence and increasing adoption of WebAssembly (Wasm) bring hope for accelerating the performance of serverless ML inference.

[0004] In a serverless computing environment, the startup speed of Wasm is 100 times faster than that of containers, making it very suitable for latency-sensitive bursty serverless inference. Although Wasm offers many significant advantages, it doesn't mean it can completely replace containers. Wasm is mainly designed to provide a secure and efficient execution environment for functions, but it cannot provide system-level features such as direct access to the host's computing resources, while containers are very suitable for more complex instances that may require additional system resources and dependencies. Therefore, Wasm can only be used to deploy specific services or in specific scenarios. Summary of the Invention

[0005] The present invention precisely addresses the problems existing in machine language inference in existing serverless computing, and proposes a serverless machine learning inference system and method based on hybrid deployment of WebAssembly and containers, including a WebAssembly-assisted hybrid runtime module, a clustering deep reinforcement learning decision module, and a running instance recycling module; the WebAssembly-assisted hybrid runtime module constructs WebAssembly and containerized instances according to user task requirements and task computing characteristics, and constructs a scheduling management table for data between WebAssembly and containers according to the linear storage characteristics of WebAssembly; the clustering deep reinforcement learning decision module efficiently and accurately configures the system computing resources and task requirement resources, optimizes the deep reinforcement learning decision steps through clustering ideas, thereby reducing the decision cost and improving the decision efficiency; the running instance recycling module uses the characteristics of historical requests and the differences in device affinity during runtime to quickly determine the optimal processing method for running instances, reducing the runtime management cost. The method of the present invention realizes an efficient and accurate response to the high-dynamic time-sensitive user computing requirements, and provides reference and help for the resource configuration decision of the serverless computing platform for high-dynamic time-sensitive computing requirements.

[0006] To achieve the above object, the technical solution adopted by the present invention is: a serverless machine learning inference system based on hybrid deployment of WebAssembly and containers, at least including a WebAssembly-assisted hybrid runtime module, a clustering deep reinforcement learning decision module, and a running instance recycling module;

[0007] The WebAssembly-assisted hybrid runtime module: constructs WebAssembly and containerized instances according to user task requirements and task computing characteristics; constructs a scheduling management table for data between WebAssembly and containers according to the linear storage characteristics of WebAssembly;

[0008] The clustering deep reinforcement learning decision-making module: Based on the current environmental devices and instance task requirements, it predicts the computing costs under different loads and device states, constructs a resource allocation decision-making agent using historical requests, abstracts the resource allocation problem of instance computing into a Markov decision problem, maps resource scheduling to the model inference of deep reinforcement learning, and optimizes the decision-making steps and decision-making space of the agent using the clustering idea;

[0009] The running instance recycling module: According to the historical request information and the device affinity characteristics during runtime, it discriminates the optimal processing methods for reusing and recycling running instances to achieve running instance recycling.

[0010] As an improvement of the present invention, in the WebAssembly-assisted hybrid runtime module, the running instances include E-WASM, P-WASM, and traditional containers. The E-WASM is the startup instance, and the P-WASM is responsible for serving complete computing tasks and, as a new adjustable knob, assists traditional containers to implement inference services; and based on the import / export and dynamic linking mechanisms of WebAssembly, data sharing between multiple running instances based on dynamic linking is realized.

[0011] As another improvement of the present invention, the method for data sharing between multiple running instances based on dynamic linking is specifically as follows:

[0012] For WebAssembly instances, the dynamic linking technology is used to load WebAssembly instances together, and data sharing is transformed into data transfer between variables;

[0013] For the WebAssembly-container sharing surface, the memory export mechanism of WebAssembly is used to selectively share the linear memory containing instance outputs to the next-hop container to implement a hybrid shared execution schedule and separate workflow information and scheduling decisions from the instances.

[0014] As another improvement of the present invention, in the clustering deep reinforcement learning decision-making module, the clustering deep reinforcement learning model consists of two layers. The upper layer rearranges and combines the entire request queue into a new cluster request queue KQ, and the lower layer uses the clustering deep reinforcement learning model to judge the execution result of KQ. The process of generating KQ is specifically as follows:

[0015] Qsort = Sort(RQ),

[0016] KQ = Kmeans(Qsort, K, Y),

[0017]

[0018] Among them, Sort(RQ) represents sorting the request queue RQ in ascending order according to the element values, and Kmeans(Qsort, K, Y) represents clustering each element m in Qsort i with its position i in the queue. The number of clusters in the clustering result is K, and the maximum number of elements in each cluster is Y. represents rounding up the result, and Tol represents the expected degree of accuracy of the decision result.

[0019] As another improvement of the present invention, the reward calculation method of the clustering depth reinforcement learning model is specifically as follows:

[0020]

[0021] Among them, num(b i ) represents the number of requests in the batch processing queue b i , sr ji represents the reward at the j-th step in b i , Sl(b i ) represents the penalty generated when the resources required to execute b j exceed the remaining resources of the system, and runt(b i ) refers to the time required to execute b i . m and k i respectively represent the number of all request queues and the number of requests in the request queue.

[0022] As a further improvement of the present invention, in the running instance recycling module, the optimal value of the number of requests n in the final execution queue is determined by using the computational model-level affinity and the computational time benefit relationship. The specific calculation method is as follows:

[0023]

[0024] SR cpu -a cpu (n)>m 1 ,

[0025] SR gpu -a gpu (n)>M 2 ,

[0026] SR mem -a mem (n)>M 3 ,

[0027]

[0028] Among them, t now represents the minimum remaining time of the requests in the current batch processing queue, SR cpu , SRgpu and SR mem represent the remaining CPU, GPU, and memory resources of the current system, and t a (n), a cpu (n), a gpu (n), and a mem (n) represent the predicted computing time, CPU, GPU, and resource cost consumption, and M 1 , M 2 , and M 3 represent the minimum remaining amounts of CPU, GPU, and memory in the system used to prevent system overload.

[0029] To achieve the above object, the technical solution adopted by the present invention is also: a server-unaware machine learning inference method based on hybrid deployment of WebAssembly and containers, including the following steps:

[0030] S1, user request computing cost evaluation: According to the current computing device resources and the characteristics of the user computing request, predict the resource consumption and execution time of different numbers of requests using different hardware devices in different carriers;

[0031] S2, resource configuration decision-making modeling: Utilize the current system resource status and the evaluation result of the user request cost in step S1 to implement resource configuration decision-making modeling for all requests in the current request queue;

[0032] S3, clustering deep reinforcement learning decision-making: Use a deep reinforcement learning model to solve the modeling problem in step S2, and utilize the clustering idea to optimize the decision space and cost to improve the decision-making efficiency;

[0033] S4, instance recycling management: According to the historical request information and the device affinity characteristics at runtime, discriminate the optimal processing methods for reusing and recycling running instances to achieve running instance recycling.

[0034] Compared with the prior art, the beneficial effects of the present invention are as follows: In the machine language inference method of serverless computing of the present invention, the advantages of WebAssembly and container technologies are fully utilized in a mixed and complementary manner. A serverless machine learning inference system and method based on the hybrid deployment of WebAssembly and container are disclosed, making full use of the computing characteristics and advantages of different runtimes to improve the computing efficiency of the system and reduce the computing cost. Utilize the efficient cold start advantage of WebAssembly to build a preprocessing and simple working instance to achieve a rapid start of computing tasks. Considering the computing characteristics of different tasks, introduce the execution instances of subsequent work: the running instance based on WebAssembly and the running instance based on the container to achieve the optimal and stable computing of computing tasks and avoid the problem of insufficient generality existing in a single execution instance. Considering the data interaction problem of different runtime instances, utilize the linear storage characteristic of WebAssembly to build a data jump table to achieve accurate and efficient data sharing. Abstract the resource configuration problem of serverless computer machine language inference into a Markov decision problem, use a deep reinforcement learning model to complete the decision-making, and use the clustering idea to optimize the decision space cost to improve the decision-making efficiency. The present invention realizes the optimal resource configuration decision, improves the overall computing efficiency of the system, and reduces the computing cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 FIG. is a schematic structural diagram of a serverless machine learning inference system based on the hybrid deployment of WebAssembly and container according to the present invention;

[0036] Figure 2 FIG. is a schematic structural diagram of a WebAssembly-assisted hybrid runtime module in the system of the present invention;

[0037] Figure 3 FIG. is a schematic structural diagram of a clustering deep reinforcement learning decision module in the system of the present invention;

[0038] Figure 4 FIG. is a benefit evaluation result diagram of the instance recycling and reuse in Case 2 by the instance recycling module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] The present invention will be further clarified below in conjunction with the drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.

[0040] Embodiment 1

[0041] A serverless machine learning inference system based on the hybrid deployment of WebAssembly and container, as Figure 1As shown in the figure, it at least includes a WebAssembly-assisted hybrid runtime module, a clustering deep reinforcement learning decision module, and a running instance recycling module, and uses WebAssembly to assist in the rapid startup and operation of complete container instances, aiming to solve the problems of efficient inference and resource allocation of machine learning for serverless computing.

[0042] To reduce the startup speed and computing cost of running instances, WebAssembly technology is introduced to assist in building a hybrid deployment runtime for containers. The running instances include E-WASM, P-WASM, and traditional containers. E-WASM is designed as a startup instance by leveraging the fast cold start feature of WebAssembly to complete the preprocessing of computing tasks and simple tasks. Different from E-WASM, P-WASM is responsible for serving complete computing tasks, that is, acting as a new adjustable knob to assist traditional containers in achieving efficient ML inference services.

[0043] The WebAssembly-assisted hybrid runtime module, as Figure 2 shown in the figure, when a user sends a request, according to the user request requirements and computing characteristics, WebAssembly and containerized instances are constructed, where the WebAssembly instance will be decoupled into P-WASM and E-WASM. The WebAssembly instance will be decoupled into P-WASM and E-WASM. To make full use of the fast cold start feature of the WebAssembly instance, first, the E-WASM instance uses its efficient cold start feature to quickly complete the preprocessing of the user's computing request, and is designed as a preprocessing instance for providing simple and preparatory work. The subsequent work will be completed through the P-WASM instance or the container instance according to the computing characteristics. P-WASM is responsible for serving complete tasks and acting as a new adjustable knob for inference services.

[0044] Considering the problem of efficient and low-cost data interaction among P-WASM, E-WASM, and traditional containers after decoupling, we make full use of the import / export and dynamic linking mechanisms of WebAssembly to achieve data sharing among multiple running instances based on dynamic linking. The present invention designs two different data sharing methods for different runtimes. Among WebAssembly instances, the dynamic linking technology is used to load WebAssembly instances together, where data sharing is transformed into data transfer between variables. For the WebAssembly-container sharing surface, the memory export mechanism of WebAssembly is used to selectively share the linear memory containing instance outputs to the next-hop container. Thus, without breaking the security isolation between instances, the data sharing and interaction overhead are minimized; finally, by designing a hybrid sharing execution schedule, the workflow information (i.e., the next-hop) and scheduling decisions are separated from the instances, thereby avoiding the reconstruction of instances, further improving the reusability of WebAssembly instances, and minimizing the storage overhead at runtime.

[0045] To achieve efficient instance computing resource configuration, a deep reinforcement learning resource configuration decision mechanism based on the clustering idea is introduced. The problem of instance computing resource configuration is abstracted as a Markov decision problem, and then the resource scheduling is mapped into the model inference of deep reinforcement learning, and the decision-making steps and decision space are optimized through the clustering idea. The clustering deep reinforcement learning decision module, as Figure 3 shown, first, based on the current environment and instance task requirements, the computing costs under different loads and device states are predicted according to the current computing request queue, so as to provide an accurate basis for resource configuration decisions. Then, the resource configuration decision agent is constructed using historical requests, and the decision-making cost and efficiency of the agent are optimized using the clustering idea to ensure the efficiency of resource configuration. Secondly, the resource scheduling is mapped into the model inference of deep reinforcement learning, and the resource configuration is completed through the clustering deep reinforcement learning model (CDQN) constructed by combining the clustering idea.

[0046] Considering that the computing efficiency of deep reinforcement learning is linearly related to the number of requests in the request queue RQ, and the requests in RQ are not evenly distributed over time. Therefore, the clustering idea is combined to construct the clustering deep reinforcement learning model (CDQN). CDQN consists of two parts: (1) rearranging and combining the entire request queue into a new cluster request queue KQ, and (2) using DQN to judge the execution results of KQ, reducing the number of DQN judgments and inference time, and improving the computing efficiency of DQN. The process of generating KQ is as follows:

[0047] Qsort = Sort(RQ),

[0048] KQ = Kmeans(Qsort, K, Y),

[0049]

[0050] where Sory(RQ) represents sorting RQ in ascending order by element value, and Kmeans(Qsort, K, Y) represents clustering each element m in Qsort i with its position i in the queue. The number of clusters in the clustering result is K, and the maximum number of elements in each cluster is Y. represents rounding up the result, and Tol represents the expected degree of accuracy of the decision result.

[0051] Construct a CDQN model according to the computing requirements. Among them, the policy network and the value network are also updated according to the reward function to obtain a better action strategy. The inference configuration can be obtained through online inference. The trained CDQN model can generate the best actions based on the policy and value networks and output the best combination of batch size, hardware target, and selected server oblivious runtime. The CDQN model consists of four main parts: the agent, the state, the action, and the reward. We set the decision space to six actions: (1) Merge the current request into the current batch; (2) Skip; (3 / 4) Execute; there is a container on the CPU / GPU; (5 / 6) Execute. Use WebAssembly on the CPU / GPU. All action selections are mutually exclusive and there is no overlap. The model decision goal is to maximize the total reward. The reward signal is given by the environment based on the agent's actions, guiding the agent's behavior. In a fixed environment, a faster execution time can complete more requests, thus effectively fulfilling the user's computing requirements. Set the following formula to calculate the reward,

[0052]

[0053] where num(b i ) represents the number of requests in the batch queue b i , sr ji represents the reward at the j-th step in b i , Sl(b i ) represents the penalty incurred when the resources required to execute b j exceed the remaining resources of the system, and runt(b i ) refers to the time required to execute b i . m and k i represent the number of all request queues and the number of requests in the request queue respectively.

[0054] To achieve better utilization of system resources and reduce the costs of instance computing and management, a discriminant mechanism for reusing and recycling running instances is introduced. By leveraging historical request information and the device affinity characteristics during runtime, the optimal processing methods for reusing and recycling running instances are quickly determined, reducing the management costs of running instances. In the running instance recycling module, when a running instance completes the computing request of the current user, it is necessary to effectively manage the instance, reset its lifecycle, accurately evaluate the reuse value of the container through the constructed recycling evaluation model, ensure the high efficiency and low cost of processing user computing requirements, and fully consider the performance advantages compared with other solutions (i.e., other execution strategies involved in CDQN). By using the affinity at the computational model level and the relationship between computational time benefits, the optimal value of the number of requests n in the final execution queue is determined on the premise of ensuring resource availability. The specific calculation method is as follows,

[0055]

[0056] SR cpu -a cpu (n)>M 1 ,

[0057] SR gpu -a gpu (n)>M 2 ,

[0058] SR mem -a mem (n)>M 3 ,

[0059]

[0060] where t now represents the minimum remaining time of the requests in the current batch processing queue, SR cpu , SR gpu and SR mem represent the remaining CPU, GPU, and memory resources in the current system, t a (n), a cpu (n), a gpu (n) and a mem (n) represent the predicted computing time, CPU, GPU, and resource cost consumption, M 1 , M 2 and M 3 represent the minimum remaining amounts of CPU, GPU, and memory in the system to prevent system overload.

[0061] Thus, the real-time resource configuration and instance management process for seamless machine learning inference on the entire server are completed, ensuring the system operation efficiency and user service quality.

[0062] Example 2

[0063] A server - unaware machine - learning inference method based on the hybrid deployment of WebAssembly and containers includes the following steps:

[0064] S1, User request computing cost assessment: According to the current computing device resources and the characteristics of user computing requests, predict the resource consumption and execution time of different numbers of requests when using different hardware devices in different carriers. In this embodiment, 2 types of hardware devices (CPU, GPU) and 2 types of running carriers (WebAssembly, Docker) are adopted. Different numbers of requests (5, 10, 50) are selected respectively, and an evaluation model M is constructed through multiple samplings.

[0065] S2, Resource configuration decision - making modeling: Utilize the current system resource status and the evaluation results of S1 on user request costs to implement resource configuration decision - making modeling for all requests in the current request queue. The modeling is CPU, GPU, memory = M(carrier, device, number of requests). In this embodiment, CPU utilization ratio 1%, GPU requirement 8000MB, memory requirement 10000MB = M(WebAssembly, GPU, 100).

[0066] S3, Clustering - based deep reinforcement learning decision - making: Use a deep reinforcement learning model to solve the modeling problem in S2, and utilize the clustering idea to optimize the decision - making space and cost to improve the decision - making efficiency.

[0067] In this embodiment, in the current system state, the remaining available ratio of CPU is 75%, the remaining amount of GPU is 7000MB, and the remaining amount of memory is 12000MB. Y is set to 5, and Tol is set to 90%. Then, the final configuration decision result is (carrier: Docker, device: GPU, number of requests to be executed 80, number of remaining requests to be executed later: 20).

[0068] S4, Instance recycling management: Considering the advantages of the efficient cold start and low computing overhead of WebAssembly, achieve the optimal operation instance recycling and management by evaluating the optimal benefit between the recycling management cost and computing efficiency. While ensuring computing efficiency and service quality, reduce the demand computing and instance management costs. Achieve an efficient and accurate response to the high - dynamic time - sensitive user computing demands, and ensure the resource configuration decision - making of the server - unaware computing platform for high - dynamic time - sensitive computing demands.

[0069] An evaluation model M2 is obtained through a large number of samplings. Its essence is that M1 removes the resource cost consumed by cold start. The performance results of different configurations in [1, 90] are as Figure 4As shown in the figure. In this embodiment, the carrier of the currently executed instance is Docker, and the default device configuration is GPU 1000MB. Then the number of directly executable instances is [1, 15]. The current remaining number of requests to be executed is 20, and the shortest deadline is 2.5s. The minimum remaining amounts of CPU, GPU, and memory are set to 10%, 1000MB, and 1500MB respectively. Then, according to M2, the costs and latencies of different carriers (reused Docker, cold-started Webassembly), different computing devices (CPU, GPU), and different request numbers [1, 20] are evaluated. The evaluation results are as Figure 4 shown in (the part where batch processing is [1, 20]). Then the final decision result is: (carrier: Webassembly, device: CPU, number of executed requests: 15, remaining number of requests to be executed subsequently: 0). The previous Docker instance is recycled and waiting for subsequent utilization.

[0070] In summary, in the present invention, the WebAssembly-assisted hybrid runtime module is used to efficiently complete part of the pre / post work of data for Docker or other running carriers. Through the fast cold-start characteristic of WebAssembly and the robustness of the container execution environment, real-time and accurate response to task calculation is achieved; the clustering deep reinforcement learning decision module efficiently and accurately configures the system computing resources and task demand resources, optimizes the deep reinforcement learning decision steps through the clustering idea, thereby reducing the decision cost and improving the decision efficiency; the running instance recycling module quickly discriminates the optimal processing method of the running instance by using the historical request characteristics and the device affinity difference during runtime, reducing the runtime management cost. The system and method of the present invention reduce the demand calculation and instance management costs while ensuring the calculation efficiency and service quality, achieve efficient and accurate response to the high-dynamic and time-sensitive user computing requirements, and ensure the resource configuration decision of the server's transparent computing platform for high-dynamic and time-sensitive computing requirements.

[0071] It should be noted that the above content only illustrates the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. For those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements all fall within the protection scope of the claims of the present invention.

Claims

1. A server-aware machine learning inference system based on hybrid deployment of WebAssembly and containers, characterized by: At least includes a WebAssembly-assisted hybrid runtime module, a clustering deep reinforcement learning decision module, and a running instance recycling module; The WebAssembly-assisted hybrid runtime module: constructs WebAssembly and containerized instances according to user task requirements and task computing characteristics; constructs a scheduling management table for data between WebAssembly and containers according to the linear storage characteristics of WebAssembly; The clustering deep reinforcement learning decision module predicts the computing costs under different loads and equipment states based on the current environment equipment and instance task requirements, builds a resource configuration decision agent using historical requests, abstracts the resource configuration problem of instance computing into a Markov decision problem, maps resource scheduling to the model reasoning of deep reinforcement learning, and uses clustering ideas to optimize the decision steps and decision space of the agent; The running instance recycling module determines the optimal processing method for reusing and recycling the running instance according to the historical request information and the device affinity characteristics during the running time, and realizes the recycling of the running instance.

2. The server-unaware machine learning reasoning system based on hybrid deployment of WebAssembly and containers as claimed in claim 1, characterized in that: In the WebAssembly-assisted hybrid runtime module, the running instances include E-WASM, P-WASM and traditional containers. The E-WASM is the startup instance, and the P-WASM is responsible for serving the complete computing task, serving as a new adjustable knob to assist the traditional container in implementing reasoning services; and based on the import / export and dynamic linking mechanism of WebAssembly, data sharing between multiple running instances based on dynamic links is realized.

3. The server-unaware machine learning reasoning system based on hybrid deployment of WebAssembly and containers as described in claim 2, characterized in that: The data sharing method between multiple running instances based on dynamic link is specifically as follows: For WebAssembly instances, dynamic linking technology is used to load WebAssembly instances together, and data sharing is converted into data transfer between variables; For the WebAssembly-container shared surface, we leverage WebAssembly’s memory export mechanism to selectively share the linear memory containing instance outputs to the next-hop container, thus achieving a hybrid shared execution schedule that separates workflow information and scheduling decisions from instances.

4. The server-unaware machine learning reasoning system based on hybrid deployment of WebAssembly and containers as claimed in claim 1, characterized in that: In the clustering deep reinforcement learning decision module, the clustering deep reinforcement learning model consists of two layers. The upper layer rearranges the entire request queue and combines it into a new cluster request queue KQ. The lower layer uses the clustering deep reinforcement learning model to determine the execution result of KQ. The process of generating KQ is specifically as follows: Qsort=Sort(RQ), KQ=Kmeans(Qsort,K,Y), Where Sort(RQ) means sorting the request queue RQ in ascending order by element value, Kmeans(Qsort,K,t) means clustering each element mi in Qsort with its position i in the queue, the number of clusters in the clustering result is K, and the number of elements in each cluster is at most Y. Indicates that the result is rounded up, and Tol indicates the expected accuracy of the decision result.

5. The server-unaware machine learning reasoning system based on hybrid deployment of WebAssembly and containers as described in claim 4, characterized in that: The reward calculation method of the clustering deep reinforcement learning model is specifically as follows: Where num(b i ) represents batch queue b i The number of requests in sr ji Indicates b i The reward for step j in i ) means that when b is executed j The penalty incurred when the required resources exceed the remaining resources in the system, runt(b i ) means the execution of b i The time required. m and k i Respectively represent the number of all request queues and the number of requests in the request queues.

6. The server-unaware machine learning reasoning system based on hybrid deployment of WebAssembly and containers as claimed in claim 1, characterized in that: In the running instance recovery module, the optimal value of the number n of requests in the final execution queue is determined by calculating the model-level affinity and the computing time-benefit relationship. The specific calculation method is: SR cpu -a cpu (n)>m1, <h2 style=";text-align:left;direction:ltr">SR<h2 style=";text-align:left;direction:ltr"> gpu <h2 style=";text-align:left;direction:ltr"> -a<h2 style=";text-align:left;direction:ltr"> gpu <h2 style=";text-align:left;direction:ltr"> (n)>m2, <h2 style=";text-align:left;direction:ltr">SR<h2 style=";text-align:left;direction:ltr"> mem <h2 style=";text-align:left;direction:ltr"> -a<h2 style=";text-align:left;direction:ltr"> mem <h2 style=";text-align:left;direction:ltr"> (n)>M3, Among them, t now Indicates the minimum remaining time of the requests in the current batch queue, SR cpu , SR gpu and SR mem Indicates the remaining CPU, GPU and memory resources of the current system, t a (n), a cpu (n), a gpu (n) and a mem (n) represents the predicted computing time, CPU, GPU, and resource cost consumption, and M1, M2, and M3 represent the minimum remaining amount of CPU, GPU, and memory in the system to prevent system overload.

7. A server-unaware machine learning reasoning method based on hybrid deployment of WebAssembly and containers using the system as claimed in claim 1, characterized in that: The steps include: S1, user request computing cost evaluation: based on the current computing device resources and user computing request characteristics, predict the resource consumption and execution time of different numbers of requests using different hardware devices in different carriers; S2, resource allocation decision modeling: using the current system resource status and the evaluation result of the user request cost in step S1, the resource allocation decision modeling of all requests in the current request queue is realized; S3, clustering deep reinforcement learning decision: use the deep reinforcement learning model to solve the modeling problem in step S2, and use clustering ideas to optimize the decision space and cost to improve decision efficiency; S4, instance recycling management: According to the historical request information and the device affinity characteristics during runtime, the optimal processing method for running instance reuse and recycling is determined to achieve running instance recycling.