End-side large model parameter protection method and system based on trusted execution environment

Through pipeline scheduling and control-data plane separation NPU driver design, the memory utilization and NPU resource sharing issues of large language models on the terminal side are solved, and efficient and secure model parameter protection and fast inference are achieved in the TEE environment.

CN120744908APending Publication Date: 2025-10-03SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510892299.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

When deploying large language models on the client side, existing technologies face the contradiction between memory utilization and fast inference startup, as well as the difficulty of efficiently and securely sharing NPU resources between TEE and REE. This makes model parameters easy to steal and causes excessive inference delays.

Method used

A pipeline scheduling mechanism is used to execute model parameter loading, decryption, and memory allocation operations in parallel. Combined with topology-aware memory allocation and release order, an NPU driver with control-data plane separation is designed to achieve lightweight resource switching between TEE and REE, dynamically expand secure memory, and retain some parameter cache.

Benefits of technology

Significantly reduces inference startup latency, improves memory utilization, ensures security and performance, and supports efficient inference of large models on resource-constrained devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744908A_ABST
    Figure CN120744908A_ABST
Patent Text Reader

Abstract

The invention provides an end-side large model parameter protection method and system based on a trusted execution environment, and the method comprises the steps: enabling a large model client application program to receive the input of a user, and transmitting the input of the user to a trusted application, namely, a large model security application; the REE OS kernel forwards the request of the user mode to the trusted application, proxy I / O and NPU of the trusted application and continuous memory allocation requests are achieved, and trusted application thread scheduling is achieved; the security monitor is used for forwarding a request of the REE OS kernel to the TEE OS kernel; the TEE OS kernel is responsible for carrying out security configuration related to the TrustZone with high privilege and realizing address space isolation between security applications; and the large model security application realizes pipeline recovery accelerated reasoning, dynamic memory capacity expansion and NPU device calling by an NPU driver to perform calculation acceleration. According to the method, the hardware acceleration capability of the reasoning task is guaranteed, meanwhile, the scale of a TEE trusted computing base (TCB) is controlled, and the overall safety and the performance expandability of the system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a method and system for protecting terminal-side large model parameters based on a trusted execution environment. Background Art

[0002] In recent years, the rapid development of large language models (LLMs) in the field of artificial intelligence has had an increasingly significant impact on global production and life. Their power lies in their ability to generate natural language text that resembles human speech, making them of significant practical value in many fields. However, these models also face challenges such as bias, data privacy, and the risk of generating inappropriate content. The development of large models also brings with it challenges such as increased resource consumption and data privacy issues.

[0003] Currently, there are two main deployment methods for large language models: cloud-side and device-side. The advantages of deploying large language models on the cloud include leveraging powerful computing power and storage resources, enabling the running of models with larger parameters, thereby improving model performance. However, the disadvantage is its reliance on network connectivity and cloud service platforms, which can lead to availability issues. Furthermore, cloud-side processing involves data transmission, which carries a high risk of privacy leakage. The advantage of device-side deployment is that it can run on local devices without relying on network connectivity or cloud service platforms, resulting in higher availability. Device-side deployment also helps protect user privacy because data does not need to be transmitted to the cloud. However, device-side deployment is limited by device computing power and storage space, making it incapable of supporting large models.

[0004] The inherent advantages of on-device deployment in terms of privacy and usability have made on-device deployment of large language models a research hotspot in recent years. The sheer number of parameters in large language models leads to high training costs, meaning that these parameters should be considered model assets of the model provider. Because on-device inference requires storing all or part of these parameters on the user device, they are vulnerable to theft by attackers, such as jailbroken devices, memory dump attacks, or malware penetration. Theft of these parameters not only results in the loss of the model provider's substantial initial investment in technical assets and intellectual property, but also could allow competitors to illegally obtain key technical secrets, severely weakening their market competitiveness. Although some manufacturers have attempted to encrypt model files stored on-device devices, model parameters are still stored in plaintext in device memory during inference, making it difficult to completely eliminate the risk of leakage. Therefore, efficiently and securely protecting the parameters and intermediate computation results of large language models in an on-device environment has become a pressing technical need.

[0005] A Trusted Execution Environment (TEE) is a technology used to enhance the security of computing devices. It provides an isolated execution environment, allowing sensitive data and code to run within a protected area, thereby preventing external attacks and unauthorized access. The core function of a TEE is to ensure that code and data are not tampered with during execution. However, existing TEE technologies generally face the limitation of requiring pre-partitioning of continuous physical memory and being unable to dynamically expand at runtime. This limitation is particularly prominent for large-model inference tasks that require dynamic expansion of large memory resources, limiting their full security and flexibility.

[0006] Although deploying large-scale model inference on-device within a TEE environment can significantly improve its security, current technical implementation still faces many difficulties, mainly reflected in the following two aspects: Challenge 1: The conflict between efficient memory utilization and fast inference startup Traditional TEE environments often statically reserve a fixed amount of secure memory space at device startup. While this static allocation method is simple to implement, it is extremely inflexible, especially on resource-constrained mobile devices. When LLM inference tasks require several GB or more of memory, if the statically reserved secure memory is too large, it will seriously squeeze the resources required for the normal operation of the REE. Conversely, if the reserved secure memory is too small, the secure memory must be dynamically expanded before each inference, resulting in additional memory allocation, data migration, file loading, and decryption, and other time-consuming operations, significantly increasing the LLM inference startup time (Time-to-First-Token (TTFT)).

[0007] For example, an 8-bit quantized Llama-3-8B model typically requires several GB of contiguous physical memory. Dynamic memory expansion of this scale can incur additional overhead of over ten seconds, far exceeding the user's acceptable latency threshold. Therefore, a new mechanism is urgently needed that can dynamically expand secure memory size while reducing the additional latency to an acceptable level for users.

[0008] Challenge 2: Efficient and secure sharing of NPU resources between TEE and REE Large models on-device typically rely on neural processing units (NPUs) for acceleration to achieve efficient inference. Currently, the NPUs in most mobile devices are typically configured in a REE environment to support a wide range of AI tasks, such as vision and audio. When running LLM inference in a TEE, NPU resources are also urgently needed to accelerate the inference process. However, there is currently a lack of a mechanism that can securely protect NPU task data in the TEE environment from leakage while efficiently enabling fast switching of shared NPU resources between the TEE and REE.

[0009] Simply redeploying the complete REE NPU driver within the TEE would not only result in additional reinitialization delays when switching the NPU, but would also introduce potential security risks due to the dramatic increase in code complexity and size within the TEE. On the other hand, completely prohibiting REE from using NPU resources would severely limit the normal user experience of the device. Therefore, designing a lightweight, secure, and efficient NPU resource sharing mechanism that both avoids excessive complexity of the driver within the TEE and efficiently enables fast resource switching between the TEE and REE has become a major technical challenge that needs to be addressed.

[0010] To protect the parameters and intermediate states of large language models running on edge devices, research has attempted to introduce memory isolation, virtualization mechanisms, or hardware-assisted access control technologies. A typical approach is a memory access isolation mechanism based on the Stage-2 Page Table (S2PT). This mechanism uses virtualization technology to add an additional address mapping layer when accessing physical pages, thereby controlling access to sensitive memory pages. This approach offers certain advantages in protecting memory security, especially on servers or chip platforms with virtualization support. However, S2PT significantly increases memory access latency in practice, especially when using a 4KB page granularity. Each TLB miss requires a double page table walk, increasing inference latency and reducing overall performance by nearly half. Although using large pages (such as 2MB or 1GB) can alleviate this problem to some extent, the high fragmentation of mobile device memory makes it difficult to consistently use large page mappings, making S2PT solutions difficult to implement on edge devices.

[0011] At the kernel level, Linux has introduced a contiguous memory allocator mechanism that supports on-demand allocation of large blocks of contiguous physical memory at runtime, thereby meeting the contiguous memory requirements of hardware, including neural network accelerators. The contiguous memory allocator mechanism can efficiently complete memory allocation under low memory pressure. However, in scenarios where resources are tight or there are a large number of locked pages, the contiguous memory allocator triggers frequent page migration operations, resulting in a significant increase in memory allocation latency, which in severe cases can affect the real-time performance of other system tasks. Furthermore, the contiguous memory allocator often fails to closely coordinate with the model loading and decryption processes, preventing the processing flow from being fully parallelized and affecting the overall inference response time.

[0012] In addition to memory management technologies, some work has attempted to reduce the risk of model leakage by encrypting model files, decrypting them during inference, and immediately clearing plaintext model data after inference. While this approach can prevent direct extraction of model files to a certain extent, since model parameters still need to reside in memory in plaintext during inference, it is still difficult to defend against attacks such as DMA attacks, memory scraping, or malicious code injection. Furthermore, this encryption-decryption process must be re-executed before each inference begins, and the frequent I / O and computational overhead become key bottlenecks affecting the performance of multiple rounds of inference.

[0013] Trusted Execution Environments (TEEs) were once used to protect small neural network models by caching a small amount of model parameters within the TEE or migrating some model layers to the TEE to ensure that core model logic is not leaked. However, as the parameter size of large language models has increased significantly, this static partitioning protection method has become difficult to adapt to the memory and computing resource requirements of LLMs. Furthermore, some systems, in order to run the entire model in limited secure memory, frequently load and remove parameter data during inference. This strategy not only results in frequent memory migration and context switching, but can also lead to severe performance degradation due to resource competition, and even affect the accuracy and robustness of the model.

[0014] In summary, existing technologies still have significant shortcomings in protecting large models on the edge. Virtualization isolation mechanisms have high performance overhead and cannot meet the needs of real-time inference. Dynamic memory allocation strategies are unstable under high loads and lack parallel optimization capabilities. Encryption and clearing strategies are only effective for static model files and lack protection for running data. Traditional TEE technology struggles to support LLM-level resource requirements without sacrificing inference performance. These limitations indicate that securing large models on the edge requires a completely new system architecture design that balances dynamic resource scheduling, high-performance inference, and full model confidentiality. Summary of the Invention

[0015] In view of the defects in the prior art, the purpose of the present invention is to provide a terminal-side large model parameter protection method and system based on a trusted execution environment.

[0016] According to the present invention, a terminal-side large model parameter protection system based on a trusted execution environment is provided, comprising: a large model client application, a large model security application, a REE OS kernel, a TEE OS kernel, and a security monitor; The large model client application receives user input and passes the user input to the trusted application, i.e., the large model security application; The REE OS kernel forwards user-mode requests to trusted applications, proxies I / O, NPU, and continuous memory allocation requests for trusted applications, and schedules trusted application threads. The security monitor is used to forward the request of the REE OS kernel to the TEE OS kernel; The TEE OS kernel is responsible for high-privilege TrustZone-related security configuration and address space isolation between secure applications; Large model security applications implement pipeline recovery to accelerate inference, dynamic memory expansion, and NPU driver calls NPU devices for computing acceleration.

[0017] Preferably, the large model client application and the REE OS kernel are in a rich execution environment, and the large model security application and the TEE OS kernel are in a trusted execution environment.

[0018] Preferably, it also includes a secure memory expansion and recycling interface. During the inference process, large model security applications can dynamically request TEE OS to expand secure memory. The underlying system collaborates with REE to allocate continuous physical pages from the continuous memory allocator area and mark them as areas protected by TZASC.

[0019] Preferably, the memory release operation is performed in reverse order according to the topological order of the model layers.

[0020] Preferably, only a lightweight data plane NPU driver is retained in the TEE, which is responsible for building the model execution context and initiating tasks; while the control plane is still managed by the complete NPU driver in the REE, which is responsible for general control tasks including scheduling and energy consumption regulation.

[0021] Preferably, each NPU task on the TEE side corresponds to a shadow placeholder task on the REE side; The REE control plane driver only acts as a scheduler, responsible for selecting the execution order. Once a security task is scheduled, its actual execution is triggered and verified by the TEE side driver.

[0022] According to a method for protecting large model parameters on an end-side based on a trusted execution environment provided by the present invention, the system for protecting large model parameters on an end-side based on a trusted execution environment is adopted. The protection method includes: Step S1: The large model client application writes the prompt input by the user into the shared memory with the large model security application and then calls the ioctl interface provided by the TEE driver in the REE OS to initiate inference. The TEE driver in the REE OS enters the security monitor through the security monitor call. The security monitor forwards the inference request to TEEOS through an exception return. The TEE OS wakes up the inference thread in the large model trusted application to start inference. Step S2: After starting inference, a thread in the large model security application performs secure memory expansion. The thread invokes a system call of the TEE OS to initiate a secure content expansion request. The TEE OS enters the secure monitor through a secure monitor call, and the secure monitor returns to the TEE driver in the REE OS through an exception. Step S3: After the TEE driver calls the Linux contiguous memory allocator to allocate contiguous memory, the calling driver enters the security monitor through the security monitor call. The security monitor forwards the allocated contiguous memory to TEEOS through an exception return. The TEE OS checks whether the current memory is contiguous with the previously allocated memory, and finally maps the contiguous memory to the large model security application. Step S4: After the memory is expanded, the large model security application writes the parameter loading information into the shared memory and then calls the TEE OS system call to initiate a parameter loading request. The TEE OS enters the security monitor through the security monitor call. The security monitor returns to the TEE driver in the REE OS through an exception. The TEE driver returns the current ioctl to enter the thread of the large model client application. Step S5: The thread of the large model client application obtains the parameter loading information in the shared memory, calls the I / O interface provided by Linux to read the parameters, and again calls the ioctl interface provided by the TEE driver to resume reasoning. After the interface enters the TEE driver of the REE OS, it calls the security monitor through the security monitor. The security monitor forwards the resume reasoning request to the TEE OS through an exception return. The TEE OS wakes up the reasoning thread in the large model trusted application to resume reasoning. Step S6: The large model security application initiates an NPU computing request during inference calculation and calls the TEE OS system call. The TEE OS enters the security monitor through the security monitor call. The security monitor forwards the NPU computing request to the TEE driver in the REE OS through an exception return. Step S7: The TEE driver hands over the NPU computing task to the NPU driver in the REE OS. The NPU driver will proactively initiate a secure NPU task to transfer NPU ownership and enter the secure monitor through a secure monitor call. The secure monitor forwards the secure NPU task to the TEE OS through an exception return. After the TEE OS configures the NPU to a secure state, it returns to the secure NPU driver in user mode. Step S8: After the large model security application completes the inference, it writes the generated string into the shared memory and initiates the inference through the TEEOS system call. The TEE OS enters the security monitor through the security monitor call. The security monitor forwards the inference completion to the TEE driver in the REE OS through an exception return. The TEE driver completes the current ioctl and returns it to the large model client application. The large model client application obtains the inference result from the shared memory.

[0023] Preferably, the steps S2 and S3 adopt a two-step interface, first completing memory allocation and then performing permission protection.

[0024] Preferably, it also includes a flow scheduling mechanism; The pipeline scheduling mechanism is based on the model's inference graph (DAG). It treats each operator as a node in the graph and identifies the parameter set that each calculation depends on in advance based on the graph's topological order. It includes the following sub-steps: Step S101: First, perform static analysis on the inference graph (DAG) of the large model and construct a dependency graph to form a topological sorting sequence; Step S102: The parameter preparation process for each operator is further refined into three heterogeneous subtasks, including continuous memory allocation, parameter loading (I / O reading), and parameter decryption. The heterogeneous subtasks are further broken down into micro-operation units, which constitute the basic scheduling units in the recovery pipeline. Step S103: assigning priorities to tasks and maintaining priority scheduling queues; Step S104: After the inference is started, the tasks in the high-priority scheduling queue are scheduled first until the inference is completed.

[0025] Preferably, some commonly used parameters are temporarily retained in the trusted memory after the inference is completed; During the pipeline scheduling process, model parameters are loaded layer by layer according to the topological order in the graph. After the inference is completed, the system releases the memory layer by layer in the reverse order, forming a first-in-last-out usage mode.

[0026] Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention schedules model parameter loading, decryption, and memory allocation operations through pipeline scheduling, which are executed in parallel with the model calculation process, effectively hiding the high overhead of cold start and significantly reducing the inference startup delay (TTFT).

[0027] 2. The present invention avoids memory fragmentation problems by dynamically expanding and reclaiming continuous secure memory areas, and combines topology-aware allocation and release order, thereby achieving on-demand, secure, and stable support for large-model reasoning on resource-constrained devices, improving overall memory utilization, and achieving efficient memory utilization while ensuring secure isolation.

[0028] 3. This invention adopts an NPU driver design with separated control and data planes, deploying only minimal data execution logic in the TEE. This avoids the migration of complex drivers and their dependencies into the TEE, balancing acceleration performance and security boundary control. This enables the NPU to switch between the TEE and the REE with low latency and security, supporting NPU acceleration while maintaining a minimum trusted computing base (TCB). BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments with reference to the following drawings: Figure 1 This is a system architecture diagram of the present invention; Figure 2 This is a schematic diagram of the large model reasoning process of the present invention; Figure 3 This is a schematic diagram of the large model inference pipeline scheduling of the present invention. DETAILED DESCRIPTION

[0030] The present invention will be described in detail below with reference to specific embodiments. The following examples will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those skilled in the art, several changes and improvements can be made without departing from the scope of the present invention. These all fall within the scope of protection of the present invention.

[0031] This invention addresses the risk of proprietary model parameter theft faced by large models in edge-side scenarios, the dilemma faced by existing trusted execution environment (TEE) technology protection solutions between memory efficiency and fast inference, and the lack of TEE support for NPUs. This design accelerates inference cold start latency by introducing pipeline recovery and enables lightweight switching between the NPU's non-secure and secure states through control-plane and data-plane separation, thereby enabling the NPU within the TEE.

[0032] This invention leverages the deterministic computational graph characteristics of large-model inference processes to schedule model parameter memory allocation, loading, and decryption operations in parallel with the computational process, building a heterogeneous operation pipeline that significantly reduces inference startup latency (TTFT) while ensuring data security. A three-stage secure memory interface ("expand-protect-reclaim") is constructed, combined with a topology-order-aware parameter caching and release strategy to achieve on-demand allocation and efficient reclamation of continuous secure memory within the TEE, balancing model scale adaptability with physical memory utilization efficiency. A collaborative NPU driver design with control-data separation utilizes a streamlined data-plane NPU driver deployed within the TEE, while the control plane remains running within the REE. NPU task scheduling and switching are collaboratively implemented through secure interfaces, ensuring hardware acceleration for inference tasks while controlling the size of the TEE's trusted computing base (TCB) and improving overall system security and performance scalability.

[0033] According to the present invention, a terminal-side large model parameter protection system based on a trusted execution environment is provided. Figure 1 As shown, the system consists of a secure world and a non-secure world. Each world includes a corresponding operating system, user-mode environment, and a security monitor in the highest-level state. The secure world refers to the trusted execution environment (TEE), while the non-secure world refers to the rich execution environment (REE). Specifically, it includes a large-model client application, a large-model security application, the REE OS kernel, the TEE OS kernel, and a security monitor. The large-model client application accepts user input and passes it to the trusted application, namely the large-model security application. The REE OS (Linux) kernel forwards user-mode requests to the trusted application, proxies I / O, NPU, and contiguous memory allocation requests for the trusted application, and schedules the trusted application threads. The security monitor forwards requests from the REE OS kernel to the TEE OS kernel. The TEE OS kernel primarily performs high-privilege TrustZone-related security configuration and implements address space isolation between secure applications. The large-model security application implements pipeline recovery for accelerated inference, dynamic memory expansion, and the NPU driver invokes the NPU device for computational acceleration.

[0034] To address the dual security and performance challenges of large-scale language model inference on-device, this paper proposes a holistic system architecture that combines flexible and secure memory management with a collaborative NPU-driven design. Based on Arm TrustZone technology, this system systematically restructures model loading, inference scheduling, and accelerator usage, enabling large models on-device to achieve strong security protection within a TEE environment while achieving response speeds approaching those of conventional inference environments.

[0035] In terms of memory management, this invention employs a "pipeline recovery" mechanism, leveraging the deterministic computational graph of large-model inference to parallelize operations such as model parameter loading, decryption, and memory allocation with the inference process. Specifically, model parameters are restored layer by layer in topological order, and the corresponding inference operator is immediately launched once the required parameters for each layer are available. This approach overlaps I / O with computation, effectively hiding the high latency during model initialization. To further reduce "pipeline bubbles" (i.e., computations waiting for unfinished parameter recovery), the system introduces a priority-based scheduling mechanism and a micro-operation-level preemptive execution strategy. This, in conjunction with a "partial parameter caching" strategy, retains some model parameters in secure memory after inference completes, thereby optimizing the startup time for the next round of inference.

[0036] To enable flexible management of large model parameters, this paper designs a new secure memory expansion and deallocation interface. During inference, trusted applications can dynamically request secure memory expansion from the TEE OS. The underlying system, in collaboration with the REE, allocates contiguous physical pages from the contiguous memory allocator area and marks them as protected by TZASC. Memory deallocation operations are then executed in reverse order of the model layer's topological order, ensuring a consistent memory layout and preventing fragmentation. This "first-in, last-out" allocation and deallocation strategy not only ensures security but also greatly simplifies memory management logic, avoiding complex defragmentation processes.

[0037] To address the issue of NPU sharing between TEE and REE, this paper proposes a collaborative driver design with "control-data separation." Specifically, the system retains only a lightweight data-plane NPU driver in the TEE, responsible for building the model execution context and initiating tasks; the control plane is still managed by the complete NPU driver in the REE, responsible for general control tasks such as scheduling and energy consumption regulation. The two collaborate to enable rapid switching of the NPU between the two worlds. The system dynamically adjusts the NPU's access rights and interrupt routing through the TrustZone hardware mechanism, ensuring that the NPU is not accessed or interfered with by the REE when executing TEE tasks, while quickly restoring REE control of the NPU after the inference task is completed.

[0038] The entire system architecture ensures the security of model parameters, KV cache, and intermediate results while maximizing the performance optimization capabilities of the inference framework outside the TEE. Modifications to the TEE operating system itself are minimal, resulting in excellent practicality and portability. This design not only addresses key obstacles to running large models within the TEE but also provides a practical path for secure AI inference on mobile devices.

[0039] This invention utilizes a collaborative driver design with separate control and data planes, separating the NPU control logic from the execution logic. Only the streamlined data plane driver is integrated into the TEE, responsible for initializing model task contexts, triggering acceleration tasks, and receiving completion status. The control plane driver remains in the REE, handling general control functions such as scheduling and device management. Communication between the two planes occurs through secure monitor calls, achieving functional decoupling and privilege isolation, thus avoiding the TCB bloat caused by moving the entire NPU driver into the TEE.

[0040] At runtime, when model inference in the TEE requires the use of the NPU, the REE driver first switches NPU control to the TEE. The TEE driver then configures the NPU's required execution context, including register instructions, I / O page tables, and input / output buffers. To ensure security, this invention leverages TrustZone's hardware mechanisms to dynamically reconfigure NPU access permissions. Specifically, TZPC is configured to prohibit REEs from accessing the NPU's MMIO interface. TZASC is used to restrict the NPU to specific protected memory areas, preventing arbitrary DMA access. Interrupt signals are also routed to the TEE via the GIC, ensuring a complete control chain for closed-loop execution within a secure environment.

[0041] To prevent malicious replay or tampering attacks, this invention also introduces a shadow task mechanism into the NPU task scheduling process. Each TEE-side NPU task corresponds to a shadow placeholder task on the REE. The REE control plane driver acts solely as a scheduler, responsible for selecting the execution order. Once a secure task is scheduled, its actual execution is triggered and verified by the TEE-side driver, ensuring that the execution context has not been modified or reused, and guaranteeing task integrity and execution order.

[0042] According to a method for protecting large model parameters on the terminal side based on a trusted execution environment provided by the present invention, the large model parameter protection system on the terminal side based on the trusted execution environment is adopted. Figure 2 As shown, the protection methods include: Step S1: The large model client application writes the user input prompt to shared memory with the large model security application and then calls the ioctl interface provided by the TEE driver in the REE OS to initiate inference. The TEE driver in the REE OS calls the secure monitor, which then forwards the inference request to the TEE OS via an exception return (eret). The TEE OS then wakes up the inference thread in the large model trusted application to begin inference.

[0043] Step S2: After starting inference, the thread in the large model security application will need to expand the secure memory. The thread calls the TEE OS system call to initiate a secure content expansion request. The TEE OS enters the security monitor through the security monitor call, and the security monitor returns to the TEE driver in the REE OS through an exception.

[0044] Step S3: After the TEE driver calls the Linux continuous memory allocator to allocate continuous memory, the calling driver enters the security monitor through the security monitor call. The security monitor forwards the allocated continuous memory to TEEOS through an exception return. The TEE OS checks whether the current memory is continuous with the previously allocated memory, and finally maps the continuous memory to the large model security application.

[0045] The dynamically scalable contiguous secure memory allocation mechanism leverages the contiguous memory allocator in the Linux kernel to dynamically migrate and reorganize physical pages. During model inference, trusted applications can request memory expansion from the TEE operating system as needed. The TEE OS collaborates with the REE-side TEE driver to allocate contiguous physical memory from the contiguous memory allocator area and configures this memory as "secure memory" through TZASC, prohibiting access from the non-secure world. This process utilizes a two-step interface design, first completing memory allocation and then performing permission protection. This effectively avoids the use of additional transfer buffers, reducing I / O latency and memory copy overhead.

[0046] Step S4: After the memory is expanded, the large model security application writes the parameter loading information into the shared memory and then calls the TEE OS system call to initiate a parameter loading request. The TEE OS enters the security monitor through the security monitor call. The security monitor returns to the TEE driver in the REE OS through an exception. The TEE driver returns the current ioctl to enter the thread of the large model client application.

[0047] Step S5: The thread of the large model client application obtains the parameter loading information in the shared memory, calls the I / O interface provided by Linux to read the parameters, and calls the ioctl interface provided by the TEE driver again to resume reasoning. After the interface enters the TEE driver of the REE OS, it calls the security monitor and enters the security monitor. The security monitor forwards the resume reasoning request to the TEE OS through an exception return. The TEE OS wakes up the reasoning thread in the large model trusted application to resume reasoning.

[0048] Step S6: The large model security application initiates an NPU computing request during inference calculation and calls the TEE OS system call. The TEE OS enters the security monitor through the security monitor call. The security monitor forwards the NPU computing request to the TEE driver in the REE OS through an exception return.

[0049] Step S7: The TEE driver transfers the NPU computing task to the NPU driver in the REE OS. The NPU driver proactively initiates a secure NPU task to transfer NPU ownership, calls the secure monitor, and then enters the secure monitor. The secure monitor then forwards the secure NPU task to the TEE OS via an exception return. After the TEE OS configures the NPU to a secure state, it returns to the secure NPU driver in user mode. The subsequent process of restoring the NPU to a non-secure state after the secure NPU driver completes the secure NPU task is similar and is omitted in the figure.

[0050] Step S8: After the large model security application completes the inference, it writes the generated string into the shared memory and initiates the inference through the TEEOS system call. The TEE OS enters the security monitor through the security monitor call. The security monitor forwards the inference completion to the TEE driver in the REE OS through an exception return. The TEE driver completes the current ioctl and returns it to the large model client application. The large model client application obtains the inference result from the shared memory, completing the entire inference process.

[0051] like Figure 2 The figure shows a flowchart of serial execution. In order to effectively reduce the cold start delay in the inference process of large language models, the present invention proposes a pipeline scheduling mechanism to parallelize the model parameter recovery process and the inference calculation process. The traditional inference initialization process usually relies on the completion of operations such as model parameter loading, memory allocation and decryption before starting the model calculation. This serial method will cause significant time delays when processing large models, especially in the context of limited end-side device resources. The scale of model parameters often reaches several GB, and its loading and preparation process can easily become a performance bottleneck.

[0052] like Figure 3As shown, the pipeline scheduling mechanism is based on the inference calculation graph (DAG) of the model, regards each calculation operator as a node in the graph, and identifies in advance the set of parameters that each calculation depends on according to the topological order of the graph. During the execution of the model, the system splits the parameter "recovery" process into three schedulable independent operations: continuous memory allocation, parameter loading (I / O), and decryption operations, and builds a unified scheduling queue together with the corresponding inference calculation operations. The scheduler uniformly schedules these heterogeneous operations according to the readiness status and priority of each operation, realizing asynchronous interleaved execution of parameter preparation and calculation, so that the model calculation can be started before all model parameters are ready, effectively hiding the delay of the initialization stage. The pipeline scheduling mechanism includes the following sub-steps: Step S101: First, perform static parsing on the inference graph (DAG) of the large model and construct a dependency graph to form a topological sorting sequence.

[0053] Step S102: The parameter preparation process for each operator is further broken down into three heterogeneous subtasks: continuous memory allocation, parameter loading (I / O reading), and parameter decryption. These heterogeneous subtasks are further broken down into micro-operation units, which form the basic scheduling units in the recovery pipeline.

[0054] Step S103: Assign priorities to tasks and maintain priority scheduling queues.

[0055] Step S104: After the inference is started, the tasks in the high-priority scheduling queue are scheduled first until the inference is completed.

[0056] To further improve the efficiency of pipeline scheduling, the present invention introduces a priority strategy based on critical path analysis, enabling the scheduler to prioritize resource allocation to critical recovery operations that could block subsequent computational paths. Furthermore, to reduce resource usage and blockage caused by lengthy operations (such as memory migration or large-block data decryption), the scheduler supports a preemptive execution mechanism. This mechanism breaks down certain recovery tasks into interruptible "micro-operation units," which can be prioritized for execution when critical computational tasks are ready, improving system responsiveness. In other words, since recovery tasks are broken down into "micro-operation units," they can be preempted by newly prepared, higher-priority tasks between "micro-operation units," enabling "preemptive switching" to higher-priority tasks.

[0057] Furthermore, to reduce pipeline "startup bubbles" (i.e., the waiting time for the first batch of computations due to unready parameters), this paper proposes a partial parameter caching mechanism. This mechanism temporarily retains some commonly used parameters in trusted memory after inference completes, avoiding repeated loading and speeding up the next round of inference. This cache uses a "first-in, last-out" allocation and release order to ensure physical continuity of memory areas and avoid fragmentation.

[0058] To ensure the controlled release of contiguous physical memory, this paper incorporates a topology-aware memory reclamation mechanism based on the model's inference graph structure. Specifically, during pipeline scheduling, model parameters are loaded layer by layer according to the topological order in the graph. After inference is complete, the system releases memory layer by layer in the reverse order, creating a "First-In-Last-Out" (FILO) usage model. This order aligns with best practices for contiguous memory management, avoiding fragmentation within the TEE and eliminating the need for complex memory cleanup operations.

[0059] Furthermore, the present invention supports the independent expansion and protection of multiple trusted memory areas. For example, the model parameter area and the KV cache area can be separated into different TZASC areas, allowing for independent scaling and management, thereby improving memory efficiency and simplifying mapping logic. The system also provides an interface for the TEE OS to proactively notify the TA to reclaim some cache areas when memory is low, ensuring the coordination of dynamic memory scheduling.

[0060] This invention aims to dynamically, securely, and efficiently expand trusted memory on resource-constrained devices to support large model inference while minimizing inference startup latency. It also aims to efficiently and securely share neural network accelerators (such as NPUs) between TEEs and REEs, avoiding excessive switching overhead and Trusted Compute Base (TCB) expansion. It aims to protect large model parameters through Trusted Execution Environment technology and reduce the inference performance overhead caused by this protection.

[0061] During the model inference process, the present invention refines the parameter loading, decryption and memory allocation operations into schedulable recovery tasks, and builds a unified execution pipeline with the inference calculation operations, making full use of the topological structure of the large model inference graph to achieve parallel parameter recovery and calculation, thereby significantly reducing cold start delays. A dynamic secure memory expansion and recovery interface is designed, combined with the topological determinism of the parameter loading order and release order to ensure that the secure memory allocated in the TEE always remains physically continuous and avoid fragmentation problems. By deploying a minimized data plane driver in the TEE, secure execution control of the NPU task context is achieved, and the control plane is still managed and scheduled by the REE. This architecture ensures that NPU resources can be switched safely and efficiently between the TEE and the REE, and limits the trusted computing base within the TEE to a minimum range, thereby improving the security of the system.

[0062] Those skilled in the art will appreciate that, in addition to implementing the system and its various devices, modules, and units provided by the present invention in purely computer-readable program code, it is entirely possible to implement the same functions of the system and its various devices, modules, and units provided by the present invention in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered a hardware component, and the devices, modules, and units included therein for implementing various functions can also be considered as structures within the hardware component; the devices, modules, and units for implementing various functions can also be considered as both software modules implementing the method and structures within the hardware component.

[0063] The above describes specific embodiments of the present invention. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art may make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. The embodiments of this application and the features in the embodiments may be combined with each other in any manner unless there is a conflict.

Claims

1. A system for protecting large model parameters on the terminal side based on a trusted execution environment, characterized in that: include: Large model client application, large model security application, REE OS kernel, TEE OS kernel and security monitor; The large model client application receives user input and passes the user input to the trusted application, i.e., the large model security application; The REE OS kernel forwards user-mode requests to trusted applications, proxies I / O, NPU, and continuous memory allocation requests for trusted applications, and schedules trusted application threads. The security monitor is used to forward the request of the REE OS kernel to the TEE OS kernel; The TEE OS kernel is responsible for high-privilege TrustZone-related security configuration and address space isolation between secure applications; Large model security applications implement pipeline recovery to accelerate inference, dynamic memory expansion, and NPU driver calls NPU devices for computing acceleration.

2. The terminal-side large model parameter protection system based on a trusted execution environment according to claim 1 is characterized in that: The large model client application and the REE OS kernel are in a rich execution environment, and the large model security application and the TEE OS kernel are in a trusted execution environment.

3. The terminal-side large model parameter protection system based on a trusted execution environment according to claim 1 is characterized in that: It also includes a secure memory expansion and recycling interface. During the inference process, large model security applications can dynamically request the TEE OS to expand secure memory. The underlying system collaborates with the REE to allocate continuous physical pages from the continuous memory allocator area and mark them as areas protected by TZASC.

4. The terminal-side large model parameter protection system based on a trusted execution environment according to claim 1, characterized in that: The memory release operation is performed in reverse order according to the topological order of the model layer.

5. The terminal-side large model parameter protection system based on a trusted execution environment according to claim 1 is characterized in that: Only a lightweight data plane NPU driver is retained in the TEE, which is responsible for building the model execution context and initiating tasks; the control plane is still managed by the complete NPU driver in the REE, which is responsible for general control tasks including scheduling and energy consumption regulation.

6. The terminal-side large model parameter protection system based on a trusted execution environment according to claim 5, characterized in that: Each NPU task on the TEE side corresponds to a shadow placeholder task on the REE side; The REE control plane driver only acts as a scheduler, responsible for selecting the execution order. Once a security task is scheduled, its actual execution is triggered and verified by the TEE side driver.

7. A method for protecting large model parameters on the terminal side based on a trusted execution environment, characterized in that: Using the terminal-side large model parameter protection system based on the trusted execution environment, the protection method includes: Step S1: The large model client application writes the prompt input by the user into the shared memory with the large model security application and then calls the ioctl interface provided by the TEE driver in the REE OS to initiate inference. The TEE driver in the REE OS enters the security monitor through the security monitor call. The security monitor forwards the inference request to the TEE OS through an exception return. The TEE OS wakes up the inference thread in the large model trusted application to start inference. Step S2: After starting inference, a thread in the large model security application performs secure memory expansion. The thread invokes a system call of the TEE OS to initiate a secure content expansion request. The TEE OS enters the secure monitor through a secure monitor call, and the secure monitor returns to the TEE driver in the REE OS through an exception. Step S3: After the TEE driver calls the Linux contiguous memory allocator to allocate contiguous memory, the calling driver enters the security monitor through the security monitor. The security monitor forwards the allocated contiguous memory to the TEE OS through an exception return. The TEE OS checks whether the current memory is contiguous with the previously allocated memory and finally maps the contiguous memory to the large model security application. Step S4: After the memory is expanded, the large model security application writes the parameter loading information into the shared memory and then calls the TEEOS system call to initiate a parameter loading request. The TEE OS enters the security monitor through the security monitor call. The security monitor returns to the TEE driver in the REE OS through an exception. The TEE driver returns the current ioctl to enter the thread of the large model client application. Step S5: The thread of the large model client application obtains the parameter loading information in the shared memory, calls the I / O interface provided by Linux to read the parameters, and again calls the ioctl interface provided by the TEE driver to resume reasoning. After the interface enters the TEE driver of the REE OS, it calls the security monitor through the security monitor. The security monitor forwards the resume reasoning request to the TEE OS through an exception return. The TEE OS wakes up the reasoning thread in the large model trusted application to resume reasoning. Step S6: The large model security application initiates an NPU computing request during inference calculation and calls the TEE OS system call. The TEE OS enters the security monitor through the security monitor call. The security monitor forwards the NPU computing request to the TEE driver in the REE OS through an exception return. Step S7: The TEE driver hands over the NPU computing task to the NPU driver in the REE OS. The NPU driver will proactively initiate a secure NPU task to transfer NPU ownership and enter the secure monitor through a secure monitor call. The secure monitor forwards the secure NPU task to the TEE OS through an exception return. After the TEE OS configures the NPU to a secure state, it returns to the secure NPU driver in user mode. Step S8: After the large model security application completes the inference, it writes the generated string into the shared memory and initiates the inference through the TEE OS system call. The TEE OS enters the security monitor through the security monitor call. The security monitor forwards the inference completion to the TEE driver in the REE OS through an exception return. The TEE driver completes the current ioctl and returns it to the large model client application. The large model client application obtains the inference result from the shared memory.

8. The method for protecting large model parameters on the terminal side based on a trusted execution environment according to claim 7, characterized in that: The steps S2 and S3 adopt a two-step interface, first completing memory allocation and then performing permission protection.

9. The method for protecting large model parameters on the terminal side based on a trusted execution environment according to claim 7, characterized in that: It also includes a flow scheduling mechanism; The pipeline scheduling mechanism is based on the model's inference graph (DAG). It treats each operator as a node in the graph and identifies the parameter set that each calculation depends on in advance based on the graph's topological order. It includes the following sub-steps: Step S101: First, perform static analysis on the inference graph (DAG) of the large model and construct a dependency graph to form a topological sorting sequence; Step S102: The parameter preparation process for each operator is further refined into three heterogeneous subtasks, including continuous memory allocation, parameter loading (I / O reading), and parameter decryption. The heterogeneous subtasks are further broken down into micro-operation units, which constitute the basic scheduling units in the recovery pipeline. Step S103: assigning priorities to tasks and maintaining priority scheduling queues; Step S104: After the inference is started, the tasks in the high-priority scheduling queue are scheduled first until the inference is completed.

10. The method for protecting large model parameters on the terminal side based on a trusted execution environment according to claim 7, characterized in that: After inference is completed, some common parameters are temporarily retained in the trusted memory; During the pipeline scheduling process, model parameters are loaded layer by layer according to the topological order in the graph. After the inference is completed, the system releases the memory layer by layer in the reverse order, forming a first-in-last-out usage mode.

Citation Information

Cited By

  • End-side large model weight protection method and system based on private memory of decryption reasoning card

    CN122153981A