Lightweight web container-oriented end-side AI inference optimization method and device

CN122816901APending Publication Date: 2026-09-25CHINA UNICOM ONLINE INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611130281.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-28
Publication Date
2026-09-25

AI Technical Summary

Benefits of technology

本发明通过将每个算子类型的计算实现分别编译为独立的WebAssembly模块,仅加载模型实际使用的算子,有效降低了推理模块的包体积,使其满足轻量级Web容器的严格限制;通过为同一算子生成一种或多种不同优化策略的Kernel版本,并在运行时根据终端设备的实际计算能力选择匹配版本,实现了跨高低端设备的自适应优化;通过在推理运行过程中监控设备状态并动态替换Kernel版本,当设备资源充裕时切换至速度优先版本以提升推理速度,当设备资源紧张时切换至体积优先版本以降低内存占用,显著提升了推理性能与运行稳定性;通过按调用频率将Kernel文件划分为主加载包和按需加载包,优先加载高频算子以缩短冷启动时间;通过算子级独立模块化设计,支持单个算子的独立更新与替换,降低了模型更新成本。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122816901A_ABST
    Figure CN122816901A_ABST
Patent Text Reader

Abstract

The application discloses a kind of end-side AI inference optimization methods and devices for lightweight web container, comprising: parsing target AI model calculation graph, extracting operator type and compiling into independent WebAssembly module, generating multiple versions of Kernel for the same type of operator;According to the call frequency, it is divided into main loading package and on-demand loading package, the main package is loaded when starting, and the on-demand package is loaded when the low-frequency operator is called for the first time during inference;Device capacity is detected to select the matching version for the operator and build a mapping table;When the device state changes, load the target version and update the mapping table function pointer, keep the calculation graph and tensor data unchanged to complete dynamic replacement.The application can reduce the volume of inference package, realize device adaptive optimization, support runtime dynamic switching of operator to improve inference performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and computer software technology, and in particular to an edge AI inference optimization method and apparatus for lightweight web containers. Background Technology

[0002] With the rapid development of artificial intelligence technology, more and more AI models are being deployed on terminal devices to perform inference tasks in order to achieve goals such as low latency, privacy protection, and offline availability. Currently, mainstream edge inference solutions include: native edge inference engines such as TensorFlow Lite and ONNX Runtime Mobile. These solutions run directly at the device's operating system layer and can fully utilize hardware acceleration capabilities, but they need to be compiled into platform-specific binaries and cannot run directly in cross-platform lightweight web containers; cloud inference solutions that deploy AI models on cloud servers are not limited by terminal computing power, but they suffer from problems such as high network latency, the need to upload private data, and offline unavailability.

[0003] In recent years, WebAssembly, as a cross-platform binary instruction format, can be executed in browsers and various lightweight web containers. Some solutions attempt to compile the entire inference engine into a single WebAssembly module to achieve cross-platform deployment. However, lightweight web containers typically have strict constraints: the main package size is limited (e.g., the main package usually does not exceed 2MB); the running terminals range from high-end flagship mobile phones to low-end entry-level devices, with computing power differences of up to 10 times; and they do not support dynamic code execution.

[0004] AI inference models consist of a series of operators organized according to a computation graph structure. Typical operators include convolution operators, fully connected operators, and activation operators. The core function of an inference engine is to execute the computation of each operator sequentially according to the computation graph. The specific computational implementation of each operator is called a kernel. In traditional inference engines, the kernel implementations of all operators are statically compiled into the engine to form a unified executable program.

[0005] Based on the aforementioned technical background, existing WebAssembly-based inference solutions suffer from the following main drawbacks: First, the overall compilation results in excessively large module sizes. When compiling the complete inference engine into a single WebAssembly module, even if the model uses only a small number of operators, the compiled output still contains implementations of all operators supported by the engine, causing the file size to typically exceed the package size limit of lightweight containers. Second, static compilation cannot adapt to device differences. Existing solutions determine the operator implementation strategy during the compilation phase, making it impossible to dynamically adjust based on the actual capabilities of the terminal device at runtime. High-end devices cannot fully utilize computing power, while low-end devices may experience inference lag or memory overflow due to excessive resource consumption by operator implementations. Third, runtime dynamic optimization is not possible. In traditional solutions, operator implementations are fixed after compilation. When the device state changes during runtime, it is impossible to switch to a more suitable operator implementation version for the current state, lacking runtime adaptability. Fourth, long cold start times. Loading a single, large WebAssembly module results in a long compilation and instantiation time during container startup, leading to high latency in the first inference response. Fifth, model updates are costly. When it is necessary to optimize the implementation of a certain operator, the entire inference engine module must be recompiled and redeployed, and it is not possible to update individual operators independently.

[0006] Therefore, existing technologies are insufficient to meet the requirements of lightweight web containers for edge AI inference in terms of package size, device adaptability, and runtime dynamic optimization. Summary of the Invention

[0007] The main objective of this invention is to provide an edge AI inference optimization method for lightweight web containers.

[0008] Another objective of this invention is to provide an edge AI inference optimization device for lightweight web containers.

[0009] To achieve the above objectives, a first aspect of the present invention proposes an edge AI inference optimization method for lightweight web containers, comprising: The computation graph of the target AI model is analyzed, the set of operator types actually used by the model is extracted, and the computation implementation of each operator type is compiled into an independent WebAssembly module, wherein at least two different Kernel versions with different optimization strategies are generated for the same operator type. Based on the calling frequency of each operator in the computation graph, the kernel files corresponding to the operators are divided into a main loading package and an on-demand loading package. The kernel in the main loading package is loaded when the container starts, and the corresponding on-demand loading package is loaded when a low-frequency operator is called for the first time during inference. The computing power of the terminal device is detected, and a matching kernel version is selected for each operator based on the detection results. A mapping table between operator identifiers and kernel function entry points is constructed. During inference, the device status is monitored. When the device status changes to meet the preset conditions, the target kernel version is loaded and the function pointers of the corresponding operators in the mapping table are updated to complete the dynamic replacement of the kernel without reloading the model or rebuilding the computation graph.

[0010] In one embodiment of the present invention, the construction of the mapping table between operator identifiers and kernel function entry points includes: During the inference initialization phase, mapping table entries are constructed, each entry containing an operator type identifier, an operator instance identifier, a pointer to the current kernel function, and an identifier for the current kernel version; Based on the equipment capability test results, select an initial kernel version for each operator, fill in the operator's type identifier and instance identifier into the corresponding entries, and fill in the selected kernel version identifier and the corresponding function pointer into the same entry.

[0011] In one embodiment of the present invention, loading the corresponding on-demand loading package when the low-frequency operator is first invoked during inference includes: When a low-frequency operator is invoked for the first time, the corresponding kernel file is loaded from the on-demand loading package; The loaded kernel is registered in the mapping table and cached locally for direct use in subsequent calls.

[0012] In one embodiment of the present invention, all kernel modules follow a unified interface specification, which defines input parameters including input tensor memory address and operator parameter address, output parameters including output tensor memory address, and return value including execution status code.

[0013] In one embodiment of the present invention, updating the function pointer of the corresponding operator in the mapping table includes: Determine the target operator and target kernel version that need to be replaced, load the WebAssembly module of the target kernel into memory and instantiate it; Obtain the function entry address of the target kernel, update the function pointer of the corresponding operator entry in the mapping table to point to the new kernel using atomic operations, and release the memory occupied by the old kernel.

[0014] In one embodiment of the present invention, the step of monitoring the device status during inference operation and loading the target Kernel version when the device status change meets a preset condition includes: Dynamic replacement is performed between a speed-priority version of the kernel and a size-priority version of the kernel; wherein, the size-priority version adopts a code size optimization compilation strategy, and the speed-priority version adopts an execution speed optimization compilation strategy; When device resources are plentiful, switch to the speed-priority version to improve inference speed; when device resources are scarce, switch to the size-priority version to reduce memory usage. The preset conditions include the current available memory being lower than a preset lower threshold or the current CPU load exceeding a preset upper threshold.

[0015] In one embodiment of the present invention, it further includes: During inference, the operator loading hierarchy strategy is dynamically adjusted based on the actual call frequency of the operators. The kernels corresponding to operators with increased call frequency are moved to the main loading package, while the kernels corresponding to operators with decreased call frequency are moved to the on-demand loading package.

[0016] In one embodiment of the present invention, it further includes: Multiple adjacent operators in the computation graph are merged and compiled into a single WebAssembly Kernel module to reduce the number of kernel calls and the interaction overhead between the host environment and the WebAssembly module.

[0017] In one embodiment of the present invention, it further includes: Establish a unified Tensor memory pool to reuse Tensor memory space during inference, thereby avoiding redundant allocation and reducing memory fragmentation.

[0018] To achieve the above objectives, a second aspect of the present invention provides an edge AI inference optimization device for lightweight web containers, comprising: The model parsing module is used to parse the computation graph of the target AI model and extract the set of operator types; The Kernel compilation module is used to compile each operator into an independent WebAssembly module and generate multiple versions of the Kernel. The hierarchical loading module is used to divide the main loading package and the on-demand loading package according to the operator call frequency and control the loading timing. The Kernel Manager is used to maintain the Kernel registry, version management, and operator-Kernel mapping tables. The device capability detection module is used to detect the computing capabilities of the terminal device and select a matching kernel version for the operator; The hot-swap control module is used to monitor the device status and trigger dynamic kernel replacement when preset conditions are met.

[0019] The embodiments of the present invention have the following beneficial effects: This invention effectively reduces the package size of the inference module by compiling the computation implementation of each operator type into an independent WebAssembly module, loading only the operators actually used by the model, thus meeting the strict limitations of lightweight web containers. It achieves adaptive optimization across high-end and low-end devices by generating one or more kernel versions with different optimization strategies for the same operator and selecting the matching version at runtime based on the actual computing power of the terminal device. By monitoring the device status and dynamically replacing the kernel version during inference, switching to a speed-priority version when device resources are sufficient to improve inference speed and switching to a size-priority version when device resources are limited to reduce memory usage, it significantly improves inference performance and operational stability. By dividing the kernel files into a main loading package and an on-demand loading package based on call frequency, it prioritizes loading high-frequency operators to shorten cold start time. Through an independent modular design at the operator level, it supports independent updates and replacements of individual operators, reducing model update costs. Attached Figure Description

[0020] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 A flowchart illustrating an edge AI inference optimization method for lightweight web containers, provided as an embodiment of the present invention; Figure 2 This invention provides an overall architecture diagram of an edge AI inference optimization system for lightweight web containers. Figure 3 This is a flowchart of operator-independent compilation provided in an embodiment of the present invention; Figure 4 A hierarchical loading timing diagram provided for embodiments of the present invention; Figure 5 A flowchart of the inference execution process provided for embodiments of the present invention; Figure 6 This is a flowchart of the Kernel hot-swap process provided in an embodiment of the present invention; Figure 7 This is a diagram illustrating the internal structure of the Kernel Manager provided in an embodiment of the present invention. Figure 8 This is a structural diagram of an edge AI inference optimization device for lightweight web containers provided in an embodiment of the present invention. Detailed Implementation

[0021] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0022] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0023] The following description, with reference to the accompanying drawings, describes an edge AI inference optimization method and apparatus for lightweight Web containers according to an embodiment of the present invention.

[0024] Example 1 This embodiment provides an edge AI inference optimization method for lightweight web containers, such as... Figure 1 As shown, the method includes the following steps: S1: Analyze the computation graph of the target AI model, extract the set of operator types actually used by the model, and compile the computation implementation of each operator type into an independent WebAssembly module, where at least two different Kernel versions with different optimization strategies are generated for the same operator type.

[0025] S2, based on the calling frequency of each operator in the computation graph, divides the Kernel file corresponding to the operator into a main loading package and an on-demand loading package. The Kernel in the main loading package is loaded when the container starts, and the corresponding on-demand loading package is loaded when a low-frequency operator is called for the first time during inference.

[0026] S3 detects the computing power of the terminal device, selects a matching kernel version for each operator based on the detection results, and constructs a mapping table between operator identifiers and kernel function entry points.

[0027] S4 monitors the device status during inference. When the device status changes to meet preset conditions, it loads the target kernel version and updates the function pointers of the corresponding operators in the mapping table, thus completing the dynamic replacement of the kernel without reloading the model or rebuilding the computation graph.

[0028] like Figure 2 As shown, the system architecture of this embodiment is divided into a model and compilation layer, a kernel storage layer, a runtime layer, and a container layer from top to bottom.

[0029] The model and compilation layer includes AI model files and a kernel compilation module. The output of the AI ​​model files is connected to the kernel compilation module, and the compilation process of the kernel compilation module includes operator extraction, independent compilation, and multi-version generation in sequence.

[0030] The kernel storage layer includes a main loading package and an on-demand loading package. The main loading package and the on-demand loading package are distinguished by a hierarchical strategy: the main loading package contains kernel files with high-frequency operators, while the on-demand loading package contains kernel files with low-frequency operators.

[0031] The runtime layer includes an inference execution framework, a computation graph scheduler, a hot update control module, a device capability detection module, a kernel manager, and an operator-kernel mapping table. The inference execution framework is connected to the computation graph scheduler, the hot update control module, and the device capability detection module, respectively; the outputs of the computation graph scheduler, the hot update control module, and the device capability detection module are connected to the kernel manager, respectively; the output of the kernel manager is connected to the operator-kernel mapping table.

[0032] The container layer includes a lightweight web container runtime and a WebAssembly virtual machine, wherein the lightweight web container runtime includes the WebAssembly virtual machine. The container layer and the runtime layer are connected via a bidirectional channel.

[0033] Specifically, step S1 includes: Obtain the target AI model file (such as a .tflite format file in TensorFlow Lite), parse the model file, and extract the computation graph structure information. The computation graph includes the definitions of input and output nodes, the types and parameter configurations of each computation node, and the data dependencies between nodes. Traverse all computation nodes in the computation graph and extract the set of operator types actually used by the model. For example, for an image classification model, the extracted operator set might be {Conv2D, DepthwiseConv, ReLU, FullyConnected, Softmax}. Only operators in this set are subsequently compiled; unused operators are not included in the compilation, thus achieving operator pruning.

[0034] like Figure 3 As shown, the kernel implementation code for each extracted operator is compiled into an independent WebAssembly module. During compilation, one or more versions of the kernel are generated for each operator. The size-first version uses the -Os optimization level and trims debugging information to minimize code segments, resulting in a small file size, making it suitable for low-memory devices or initial loads. The speed-first version uses the -O3 optimization level and enables optimizations such as loop unrolling, resulting in fast execution speed and making it suitable for high-performance devices.

[0035] The following uses the Conv2D operator as an example to illustrate the core computation logic of the kernel. Let the input tensor be X, the convolution kernel weights be W, the bias be b, and the output tensor be Y. For the output value of the output channel oc at spatial location (oh, ow), the core calculation is as follows: Y[oc, oh, ow]= sum_{ic, kh, kw}(X[ic, oh stride_h+kh-pad_h, ow stride_w+kw-pad_w] W[oc, ic, kh, kw]) + b[oc] Where ic represents the input channel index, kh / kw represents the height / width direction index of the convolution kernel, stride_h / stride_w represents the stride, and pad_h / pad_w represents the padding amount.

[0036] The kernel execution steps include: calculating the starting position of the corresponding input window based on the output coordinates; performing element-wise multiply-accumulate (MAC) accumulation on the input window and the convolution kernel; adding the bias term b[oc] after accumulation; if the operator in the computation graph has an activation function, performing the activation operation (e.g., ReLU: Y=max(0,Y)); and writing the result back to the corresponding position of the output tensor. In the quantization implementation (INT8 Kernel), the above multiply-accumulate is first completed in the INT32 accumulation domain, then dequantized or rescaled according to the quantization parameters, and finally written back to the INT8 or FP32 output. The volume-first version and the speed-first version are mathematically equivalent, differing only in loop unrolling, memory access layout, and instruction scheduling strategy.

[0037] Each compiled artifact comes with a metadata description file that records the operator type identifier, kernel version identifier, input and output tensor specifications, compiled artifact size, and applicable device capabilities.

[0038] All kernel modules adhere to a unified interface specification to ensure substitutability and interface consistency between different operators and different kernel versions. Parameter passing adopts a structure of "fixed header + Tensor description area + operator parameter area," the complete definition of which is as follows: KernelCallParams { / / Fixed Header uint32_t abi_version; / / Interface version number, for backward compatibility uint32_t op_type; / / Operator type identifier (e.g., Conv2D=1, Relu=2) uint32_t input_count; / / Number of input Tensors uint32_t output_count; / / Number of output Tensors uint32_t attr_bytes; / / Length of operator parameter area in bytes uint32_t reserved; / / Reserved field, defaults to 0 / / Tensor description area (inputs + outputs) TensorDesc tensors[input_count + output_count]; / / Operator parameter area (interpreted by op_type) uint8_t attrs[attr_bytes]; } TensorDesc { uint64_t data_offset; / / Tensor data offset relative to Wasm linear memory base address uint32_t n; uint32_t c; uint32_t h; uint32_t w; uint32_t dtype; / / Data type (INT8 / INT32 / FP32, etc.) int32_t zero_point; / / Quantization zero point; can be set to 0 in non-quantization scenarios. float scale; / / Quantization scaling factor; can be set to 1.0 in non-quantization scenarios. } The attrs are encoded using TLV (Type-Length-Value) or according to operator conventions. Taking Conv2D as an example, attrs include at least the fields stride_h, stride_w, pad_h, pad_w, dilation_h, dilation_w, and activation_type. When the inference execution framework calls the Kernel, it first verifies abi_version and op_type, then locates the input / output data address using TensorDesc, and finally parses the attrs to perform the corresponding computation.

[0039] The unified entry point can be defined as: int32_t kernel_entry(const KernelCallParams params); The return code convention is as follows: 0 indicates success, and non-zero indicates failure (such as invalid parameters, memory access violation, or data type mismatch). This unified parameter passing format ensures the substitutability and interface consistency between different operators and different kernel versions. A unified interface enables kernel interchangeability, laying the foundation for subsequent hot-swap functionality.

[0040] Furthermore, for multiple adjacent operators in the computation graph with fixed calling relationships, the system supports merging and compiling them into a single WebAssembly Kernel module. For example, adjacent convolution and activation operators can be merged into a combined kernel, and the activation operation can be performed directly within the same module after the convolution calculation. This combines the operation that originally required two kernel calls and two data transfers into one, thereby reducing the number of kernel calls and the interaction overhead between the host environment and the WebAssembly module, and further improving inference efficiency.

[0041] Specifically, step S2 includes: Static analysis is performed on the computational graph of the target model, and the frequency of each operator appearing in the computational graph is counted as the call frequency. Operators are divided into two levels according to a preset threshold: operators with a call frequency greater than or equal to the threshold are high-frequency operators, such as Conv2D and ReLU, which appear repeatedly in the computational graph; operators with a call frequency less than the threshold are low-frequency operators, such as Softmax and Reshape, which appear only once at the end of the model.

[0042] The kernel files corresponding to high-frequency operators are assigned to the main loading package, while the kernel files corresponding to low-frequency operators are assigned to the on-demand loading package. For example... Figure 4 As shown, when the container starts, all high-frequency kernels in the main loading package are loaded, and the inference execution framework completes initialization. During inference, when a low-frequency operator is called for the first time, the inference execution framework finds that the target kernel is not loaded in the mapping table, and sends a loading request to the on-demand loading package. After the low-frequency kernel is loaded and registered in the mapping table, inference continues. Subsequent calls to the same low-frequency operator are directly retrieved from the cache, without needing to load it again.

[0043] Furthermore, during inference execution, the system continuously monitors the actual call frequency of each operator. When it detects that the call frequency of a low-frequency operator originally belonging to the on-demand loading package has significantly increased in recent inference tasks, exceeding a preset frequency threshold, the system automatically moves the corresponding kernel file of that operator from the on-demand loading package to the main loading package. This allows subsequent inference tasks to directly load the operator when the container starts, avoiding frequent triggering of on-demand loading during runtime. Conversely, when it detects that the call frequency of a high-frequency operator originally belonging to the main loading package has continuously decreased in recent inference tasks, falling below a preset frequency threshold, the system automatically moves the corresponding kernel file of that operator from the main loading package to the on-demand loading package to reduce the size of the main loading package. Through this dynamic adjustment mechanism, the operator loading strategy continuously adapts to changes in the actual inference load, achieving adaptive optimization of the main loading package and the on-demand loading package.

[0044] Specifically, step S3 includes: During the inference initialization phase, a mapping table between operator identifiers and kernel functions is constructed. Each entry in the mapping table contains an operator type identifier, an operator instance identifier, a pointer to the current kernel function, and an identifier for the current kernel version.

[0045] During the inference startup phase, the capabilities of the terminal device are detected through the system capability interface provided by the lightweight web container. The detection items include device performance level and memory size. Based on the detection results, the device capability level is determined: if the performance level and memory size meet the high-performance criteria, it is classified as a high-performance device, and a speed-priority version of the kernel is selected for each operator; otherwise, it is classified as a resource-constrained device, and a size-priority version of the kernel is selected for each operator. For example, if benchmarkLevel is greater than or equal to 30 and memorySize is greater than or equal to 2048MB, it is classified as a high-performance device; otherwise, it is classified as a resource-constrained device. Based on the device capability detection results, an initial kernel version is selected for each operator, and the operator type identifier, instance identifier, selected kernel version identifier, and corresponding function pointer are filled into the mapping table.

[0046] like Figure 5 As shown, the inference execution framework executes each operator sequentially according to the computation graph topology. When an operator not loaded in the mapping table is encountered, the corresponding kernel file is loaded on demand and registered in the mapping table before execution continues.

[0047] Specifically, step S4 includes: During inference operation, device status metrics are continuously monitored, including current available memory and current CPU load. When available memory is detected to be below a preset lower threshold (e.g., below 512MB) or CPU load is detected to be above a preset upper threshold (e.g., above 80%), the preset conditions are determined to be met and a hot-swap decision is triggered.

[0048] like Figure 6 As shown, the hot replacement execution process is as follows: Determine the target operator and target kernel version to be replaced, for example, when memory is insufficient, replace the speed-priority version with the size-priority version; Load the WebAssembly module of the target kernel into memory and instantiate it; Obtain the function entry address of the target kernel; Update the function pointer of the corresponding operator entry in the mapping table to point to the new kernel; Release the memory occupied by the WebAssembly module of the old kernel.

[0049] The key technical features of hot replacement are: only function pointers are replaced, without modifying the computation graph structure or reallocating tensor memory; only the function addresses in the mapping table are updated; the memory addresses and data of the input and output tensors remain unchanged before and after the replacement, ensuring the continuity of inference data; the model file and computation graph remain unchanged in memory, and the replacement process does not involve model parsing; the pointer update of the mapping table is an atomic operation, ensuring that no intermediate state occurs during the inference execution process.

[0050] like Figure 7 As shown, to achieve the above functions, this embodiment also includes a Kernel Manager as a core component. The Kernel Manager includes: The Kernel registry stores the metadata (operator type, version, entry address) of all registered kernels. Version manager manages multiple versions of the same operator kernel and provides a version switching interface; The load / unload controller is responsible for the instantiation and memory release of WebAssembly modules; The mapping table maintains the mapping relationship between operator instance identifiers and kernel function pointers.

[0051] The core functions provided by the Kernel Manager include registering a Kernel (registering the metadata and function entry points of a newly loaded Kernel to the registry), querying a Kernel (querying available Kernels based on operator type and version identifier), switching versions (receiving instructions from the hot-swap control module to update the mapping table), and releasing a Kernel (unloading modules that are no longer in use and reclaiming memory).

[0052] The following uses an image classification scenario as an example to illustrate the complete workflow of this embodiment.

[0053] A MobileNet V2 image classification model is run in a lightweight web container. The model includes five operators: Conv2D, DepthwiseConv, ReLU, FullyConnected, and Softmax.

[0054] The first stage is offline compilation: parsing the MobileNet V2 model file and extracting the operator set {Conv2D, DepthwiseConv, Relu, FullyConnected, Softmax}; compiling each operator into a size-priority version and a speed-priority version, with only one version generated for Relu and Softmax due to their small size; through frequency analysis, Conv2D appears 53 times, DepthwiseConv appears 17 times, and Relu appears 70 times, classifying them as high-frequency operators, while FullyConnected appears once and Softmax appears once, classifying them as low-frequency operators; the speed-priority versions of the high-frequency operators are allocated to the main loading package (e.g., Conv2D_fast.wasm 85KB, DepthwiseConv_fast.wasm 43KB, Relu.wasm 8KB, totaling 136KB), and the low-frequency operators are allocated to the on-demand loading package (e.g., FC_fast.wasm 35KB, Softmax.wasm 12KB, totaling 47KB).

[0055] The second phase is runtime execution: The container starts, loads the main loader, and initializes the inference execution framework; device detection shows 4 CPU cores and 3GB of available memory, indicating a high-performance device, and a speed-priority version is selected; a mapping table is built, and the entry addresses of Conv2D, DepthwiseConv, and ReLU are filled in; the first inference request is received, and each operator is executed according to the computation graph; when FullyConnected is executed, it is found that it is not loaded, triggering on-demand loading and registration; similarly, loading is triggered when Softmax is executed.

[0056] The third stage is runtime hot-swapping: After inference has been running for a period of time, the monitoring module detects that available memory has dropped to 400MB (below the 512MB threshold); the hot-swapping control module decides to replace the larger Conv2D_fast.wasm with Conv2D_small.wasm; it loads and instantiates Conv2D_small.wasm, obtains the function entry address; updates the function pointers of the Conv2D entries in the mapping table; and releases the memory occupied by Conv2D_fast.wasm. Subsequent inference automatically calls Conv2D_small.wasm, memory usage decreases, and inference continues to run normally. Throughout the entire hot-swapping process, the computation graph structure remains unchanged, the intermediate tensor data remains unchanged, and the inference process is not interrupted.

[0057] Furthermore, this embodiment establishes a unified Tensor memory pool. During the inference initialization phase, the system pre-allocates a contiguous memory space as the Tensor memory pool, from which memory for all input and output tensors is allocated. During inference, when an intermediate tensor is no longer needed for subsequent computations, its occupied memory space is not immediately released but marked as reusable and returned to the memory pool for direct use by subsequent operators. Through this memory reuse mechanism, frequent memory allocation and release operations are effectively avoided, memory fragmentation is reduced, and memory management overhead is lowered, further improving memory utilization efficiency during inference runtime.

[0058] In summary, this embodiment effectively reduces the package size of the inference module by independently compiling and generating multiple versions at the operator level, loading only the operators actually used by the model; it shortens the cold start time by prioritizing high-frequency operators through a hierarchical loading strategy; it achieves adaptive optimization across high-end and low-end devices through runtime device capability detection and kernel version selection; and it significantly improves inference performance and operational stability by dynamically switching operators when device resources change through device status monitoring and a kernel hot-swap mechanism.

[0059] Example 2 This invention also provides an edge AI inference optimization device for lightweight web containers, such as... Figure 8 As shown, the device 10 includes: The model parsing module 100 is used to parse the computation graph of the target AI model and extract the set of operator types; Kernel compilation module 200 is used to compile each operator into an independent WebAssembly module and generate multiple versions of the Kernel; The hierarchical loading module 300 is used to divide the main loading package and the on-demand loading package according to the operator call frequency and control the loading timing. Kernel Manager 400 is used to maintain the Kernel registry, version management, and operator-Kernel mapping tables; The device capability detection module 500 is used to detect the computing capabilities of the terminal device and select a matching kernel version for the operator. The hot-swap control module 600 is used to monitor the device status and trigger dynamic kernel replacement when preset conditions are met.

[0060] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

[0061] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0062] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

Claims

1. A method for optimizing edge AI inference for lightweight web containers, characterized in that, Includes the following steps: The computation graph of the target AI model is analyzed, the set of operator types actually used by the model is extracted, and the computation implementation of each operator type is compiled into an independent WebAssembly module, wherein at least two different Kernel versions with different optimization strategies are generated for the same operator type. Based on the calling frequency of each operator in the computation graph, the kernel files corresponding to the operators are divided into a main loading package and an on-demand loading package. The kernel in the main loading package is loaded when the container starts, and the corresponding on-demand loading package is loaded when a low-frequency operator is called for the first time during inference. The computing power of the terminal device is detected, and a matching kernel version is selected for each operator based on the detection results. A mapping table between operator identifiers and kernel function entry points is constructed. During inference, the device status is monitored. When the device status changes to meet the preset conditions, the target kernel version is loaded and the function pointers of the corresponding operators in the mapping table are updated to complete the dynamic replacement of the kernel without reloading the model or rebuilding the computation graph.

2. The method according to claim 1, characterized in that, The mapping table between the operator identifier and the kernel function entry point includes: During the inference initialization phase, mapping table entries are constructed, each entry containing an operator type identifier, an operator instance identifier, a pointer to the current kernel function, and an identifier for the current kernel version; Based on the equipment capability test results, select an initial kernel version for each operator, fill in the operator's type identifier and instance identifier into the corresponding entries, and fill in the selected kernel version identifier and the corresponding function pointer into the same entry.

3. The method according to claim 1, characterized in that, The loading of the corresponding on-demand package when the low-frequency operator is first invoked during the inference process includes: When a low-frequency operator is invoked for the first time, the corresponding kernel file is loaded from the on-demand loading package; The loaded kernel is registered in the mapping table and cached locally for direct use in subsequent calls.

4. The method according to claim 1, characterized in that, All kernel modules follow a unified interface specification, which defines input parameters including the memory address of the input tensor and the address of the operator parameter, output parameters including the memory address of the output tensor, and return values ​​including the execution status code.

5. The method according to claim 1, characterized in that, The function pointers for the corresponding operators in the update mapping table include: Determine the target operator and target kernel version that need to be replaced, load the WebAssembly module of the target kernel into memory and instantiate it; Obtain the function entry address of the target kernel, update the function pointer of the corresponding operator entry in the mapping table to point to the new kernel using atomic operations, and release the memory occupied by the old kernel.

6. The method according to claim 1, characterized in that, The process of monitoring device status during inference operation and loading the target Kernel version when device status changes meet preset conditions includes: Dynamic replacement is performed between a speed-priority version of the kernel and a size-priority version of the kernel; wherein, the size-priority version adopts a code size optimization compilation strategy, and the speed-priority version adopts an execution speed optimization compilation strategy; When device resources are plentiful, switch to the speed-priority version to improve inference speed; when device resources are scarce, switch to the size-priority version to reduce memory usage. The preset conditions include the current available memory being lower than a preset lower threshold or the current CPU load exceeding a preset upper threshold.

7. The method according to claim 1, characterized in that, Also includes: During inference, the operator loading hierarchy strategy is dynamically adjusted based on the actual call frequency of the operators. The kernels corresponding to operators with increased call frequency are moved to the main loading package, while the kernels corresponding to operators with decreased call frequency are moved to the on-demand loading package.

8. The method according to claim 1, characterized in that, Also includes: Multiple adjacent operators in the computation graph are merged and compiled into a single WebAssembly Kernel module to reduce the number of kernel calls and the interaction overhead between the host environment and the WebAssembly module.

9. The method according to claim 1, characterized in that, Also includes: Establish a unified Tensor memory pool to reuse Tensor memory space during inference, thereby avoiding redundant allocation and reducing memory fragmentation.

10. A device for optimizing edge AI inference for lightweight web containers, used to implement the method described in any one of claims 1-9, characterized in that, include: The model parsing module is used to parse the computation graph of the target AI model and extract the set of operator types; The Kernel compilation module is used to compile each operator into an independent WebAssembly module and generate multiple versions of the Kernel. The hierarchical loading module is used to divide the main loading package and the on-demand loading package according to the operator call frequency and control the loading timing. The Kernel Manager is used to maintain the Kernel registry, version management, and operator-Kernel mapping tables. The device capability detection module is used to detect the computing capabilities of the terminal device and select a matching kernel version for the operator; The hot-swap control module is used to monitor the device status and trigger dynamic kernel replacement when preset conditions are met.