Method for NPU firmware to support multi-model fast switching
By introducing priority scheduling and model caching mechanisms in NPU firmware, dynamically manage the loading and switching of multiple models, solving the problems of multi-model switching delay and resource waste in traditional methods, achieving efficient multi-tasking and real-time improvement.
Patent Information
- Application Number
- CN202510274919.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
Smart Images

Figure CN120216175A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of neural processing units and deep learning hardware, and specifically relates to a method for an NPU firmware to support fast multi-model switching. Background Art
[0002] During the deep learning inference process, many application scenarios require simultaneous processing of multiple models, such as object detection, path planning, and semantic segmentation in autonomous driving, or multi-task running on edge computing devices. These scenarios require the models to be loaded and switched quickly, but traditional methods often have problems of high latency and resource waste.
[0003] Existing technologies usually adopt static allocation or preloading methods to handle multi-model tasks, which are difficult to adapt to the dynamic changes of tasks, and cannot efficiently utilize storage and computing capabilities under limited hardware resources. Therefore, it is of great practical significance to develop a firmware framework that supports fast multi-model switching and optimizes the switching latency by combining model caching and priority mechanisms. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for an NPU firmware to support fast multi-model switching in order to solve the above-mentioned problems.
[0005] The technical solution adopted by the present invention is as follows: A method for an NPU firmware to support fast multi-model switching, the method comprising the following steps:
[0006] S1: Receive a multi-task request, analyze the task type through a priority scheduling module, dynamically allocate an initial priority by combining task real-time performance and importance, and generate a task execution queue;
[0007] S2: According to the task frequency and priority, the model caching module preloads high-frequency / high-priority models into the cache, stores low-frequency models in the main storage, and adopts a lazy loading mechanism to reduce resource occupancy;
[0008] S3: Monitor the system load and task status in real time, and dynamically increase the priority of urgent tasks through the priority scheduling module to support high-priority tasks to preempt the model resources of low-priority tasks;
[0009] S4: When switching models, the fast switching mechanism only loads the differentiated weights or parameters, avoids full-model reloading, and combines DMA to accelerate data transmission, shortening the switching time to the millisecond level;
[0010] S5: Utilize the hardware co-optimization module to quickly save and restore the computing context at the firmware level to ensure seamless model switching;
[0011] S6: Support parallel processing of high-priority model inference and low-priority model preloading through the parallel computing units and memory channels built into the NPU, maximizing the utilization of hardware resources;
[0012] S7: Collect task latency, cache hit rate, and hardware resource occupancy metrics, and feedback them to the scheduling module to dynamically optimize the cache policy and priority weights;
[0013] S8: After the task is completed, release the computing resources occupied by the model, and retain high-frequency models or unload low-frequency models according to the cache policy to free up storage space for subsequent tasks;
[0014] S9: Periodically update the hierarchical cache rules and priority algorithm parameters based on historical task data to adapt to long-term task mode changes and improve the system's adaptive ability
[0015] In a preferred embodiment, in step S1, after the system receives a multi-model task request from the application layer, the priority scheduling module first identifies the task type, combines preset real-time metrics and task criticality tags, and dynamically generates a priority scoring table; this module statistically analyzes the historical task execution frequency based on the sliding window algorithm, assigns higher initial priority weights to high-frequency tasks, and simultaneously constructs a two-way task queue with time constraints to ensure that high-real-time tasks are preferentially scheduled at the front of the queue and reserve a dynamic insertion window for burst tasks.
[0016] In a preferred embodiment, in step S2, the model cache module initiates three-level storage partitioning according to the priority scoring table: fully load the model with the highest score into the SRAM cache tightly coupled to the NPU; compress the second-priority model through pruning and store it in the shared LLC cache pool; only retain the metadata descriptors for low-frequency models, and store the actual parameters in the external DDR; the cache manager uses a heat-weight hybrid prediction model to track the call times, most recent timestamp of use, and task correlation data of each model in real time, and dynamically adjusts the storage level; when the SRAM space is insufficient, trigger a priority-based eviction policy to demote the low-heat model for storage and retain its metadata fingerprint.
[0017] In a preferred embodiment, in step S3, during the task execution, the hardware event monitoring unit continuously detects the task execution status; once the preemption condition is triggered, the scheduling module immediately freezes the DMA transfer of the current low-priority model, clears its prefetch instruction queue, and mounts the physical address space of the high-priority model to the NPU memory access path through memory remapping technology; at the same time, update the mutex status of the task queue to ensure that the preempted task can reload the context from the checkpoint when it resumes.
[0018] In a preferred embodiment, in step S4, when it is necessary to switch to the target model, the parameter difference analysis engine compares the weight matrix of the currently running model with the binary image of the target model, and uses a bitmap to mark the different blocks; the loading engine generates a minimized transfer sequence according to the difference bitmap, and realizes the jump transfer of the discontinuous address space through the chained DMA descriptor; for the convolutional layer parameters, the block incremental coding technology is adopted to only transfer the parameter blocks whose changes exceed the preset threshold, and the CRC check code is combined to ensure the transmission integrity, shortening the typical model switching time from 50 ms in the traditional scheme to within 3 ms.
[0019] In a preferred embodiment, in step S5, at the moment of model switching, the firmware layer triggers a dedicated context management coprocessor, encrypts the register state of the current model and writes it into the non-blocking buffer area; the pre-stored context of the target model is quickly restored to the register group through the hardware-accelerated decompression engine, and at the same time the memory management unit synchronously switches the page table base address register to realize the seamless switching of the virtual address space; this process uses the NPU internal bus bandwidth isolation technology to ensure that the context transmission does not affect the data stream of the high-priority tasks being executed.
[0020] In a preferred embodiment, in step S6, the NPU computing resources are divided into multiple virtualized partitions. The high-priority model exclusively occupies the main computing array to execute the inference task, while the coprocessor unit parallelly executes the parameter prefetching and weight initialization decoding of the low-priority model; the memory controller adopts an interleaved access mode to disperse the model weight loading requests to different DRAM Bank groups, and combines the bus clock phase offset technology to eliminate the memory access conflict; for the quantized model, the fixed-point format conversion hardware is started in the preloading stage to reduce the computing delay during operation.
[0021] In a preferred embodiment, in step S7, the embedded performance probe continuously collects the time series data of the L1 cache hit rate, DMA transfer error rate, and task switching time, and generates a resource utilization heat map through the online learning module; the feedback engine dynamically adjusts the cache partition ratio according to the heat map features: when it is detected that the cache hit rate of the semantic segmentation model continuously drops below 40%, its storage level is automatically increased to SRAM, and the cache priority of the speech recognition model is correspondingly reduced; at the same time, the priority weight coefficient is adjusted by attenuation or gain according to the task timeout history record.
[0022] In a preferred embodiment, in step S8, the task termination instruction triggers a multi-level resource recovery pipeline: First, the binding relationship between the model and the computing unit is released, and the corresponding TLB entries and cache tags are cleared; then, according to the reference count table, the model reuse possibility is judged. High-frequency models retain metadata and are marked as standby, while low-frequency models trigger a storage recovery interrupt; finally, physical memory fragmentation is performed, the released storage blocks are merged and added to the free pool, and the storage alignment efficiency of subsequent models is optimized through address remapping.
[0023] In a preferred embodiment, in step S9, the system starts an offline optimization engine during the idle period to perform spatio-temporal pattern mining on historical task logs; uses a reinforcement learning framework to train a cache prediction model to generate cache policy templates for different scenarios; at the same time, analyzes the success rate and error patterns of incremental loading, and optimizes the parameters of the differential bitmap generation algorithm; after the updated policy file passes the digital signature verification, it is written into the policy storage area of the firmware through the secure boot process to achieve the adaptive evolution of system-level parameters.
[0024] In summary, due to the adoption of the above technical solutions, the beneficial effects of the present invention are as follows:
[0025] 1. In the present invention, the system real-time performance and resource utilization rate in multi-task scenarios are significantly improved. Through the dynamic hierarchical cache policy and priority scheduling mechanism, the system can intelligently allocate cache resources, giving priority to ensuring the rapid loading of high-frequency and high-urgency models. For example, in the autonomous driving scenario, the switching delay of the path planning model can be reduced to the millisecond level, ensuring seamless switching of critical tasks. The incremental parameter loading technology is deeply integrated with the firmware-level context management, only transmitting model difference data and using hardware acceleration to restore the computing state, greatly reducing the bandwidth waste caused by traditional full-scale loading. When edge computing devices perform multi-task parallelism such as speech recognition and image processing, the model switching efficiency is increased by more than 10 times, while maintaining the computing accuracy without loss.
[0026] 2. In the present invention, long-term operation stability and scenario adaptability are achieved through hardware co-optimization and self-feedback mechanisms. The priority preemption strategy combined with real-time performance monitoring can dynamically adjust resource allocation when a sudden task is triggered, avoiding low-priority tasks from blocking critical processes. The periodic global policy optimization continuously improves the cache rules and scheduling algorithms based on historical task patterns, enabling the system to autonomously adapt the best switching strategy in diverse scenarios such as smart homes and data centers. In addition, the hierarchical storage management and parallel resource release mechanisms effectively reduce the memory fragmentation problem, supporting the device to run stably for a long time with limited hardware resources, providing a highly reliable and low-latency general solution for multi-model real-time inference scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 It is a flowchart of the method of the present invention. Detailed Implementation Modes
[0028] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0029] Embodiment:
[0030] Referring to Figure 1 , a method for an NPU firmware to support multi-model fast switching, the method includes the following steps:
[0031] S1: Receive multi-task requests, analyze the task types through a priority scheduling module, dynamically allocate initial priorities in combination with task real-time performance and importance, and generate a task execution queue.
[0032] S2: According to the task frequency and priority, the model cache module pre-loads high-frequency / high-priority models into the cache, stores low-frequency models in the main memory, and adopts a lazy loading mechanism to reduce resource occupancy.
[0033] S3: Monitor the system load and task status in real time, dynamically increase the priority of urgent tasks through the priority scheduling module, and support high-priority tasks to preempt the model resources of low-priority tasks.
[0034] S4: When switching models, the fast switching mechanism only loads the differentiated weights or parameters, avoids full-model reloading, combines DMA to accelerate data transmission, and shortens the switching time to the millisecond level.
[0035] S5: Utilize the hardware co-optimization module to quickly save and restore the computing context at the firmware level to ensure seamless model switching.
[0036] S6: Through the parallel computing unit and memory channels built into the NPU, support parallel processing of high-priority model inference and low-priority model preloading to maximize the utilization of hardware resources.
[0037] S7: Collect metrics such as task latency, cache hit rate, and hardware resource occupancy, and feedback them to the scheduling module to dynamically optimize the cache policy and priority weights.
[0038] S8: After the task is completed, release the computing resources occupied by the model, and retain high-frequency models or unload low-frequency models according to the cache policy to free up storage space for subsequent tasks.
[0039] S9: Based on historical task data, periodically update the hierarchical cache rules and priority algorithm parameters to adapt to long-term task mode changes and improve the system's adaptability.
[0040] In step S1, after the system receives a multi-model task request from the application layer, the priority scheduling module first identifies the task type (such as object detection, semantic segmentation, or path planning in the autonomous driving scenario), combines the preset real-time metrics (such as the response time threshold) and the task criticality labels (such as safety-related task markers), and dynamically generates a priority scoring table. This module statistically calculates the historical task execution frequency based on the sliding window algorithm, assigns a higher initial priority weight to high-frequency tasks, and constructs a two-way task queue with time constraints to ensure that high-real-time tasks are preferentially scheduled at the front of the queue and reserves a dynamic insertion window for burst tasks.
[0041] In step S2, the model caching module starts three-level storage partitioning according to the priority scoring table: fully loads the model with the highest score into the SRAM cache tightly coupled to the NPU; the model with the second-highest priority is pruned and compressed and then stored in the shared LLC cache pool; the low-frequency model only retains the metadata descriptor, and the actual parameters are stored in the external DDR. The cache manager uses a heat-weight hybrid prediction model to track the call count, the most recent usage timestamp, and the task correlation data of each model in real time and dynamically adjusts the storage level. When the SRAM space is insufficient, a priority-based eviction policy is triggered to demote the low-heat model for storage and retain its metadata fingerprint.
[0042] In step S3, during the task execution, the hardware event monitoring unit continuously detects the task execution status (such as the emergency signal generated by the path planning task due to environmental mutations). Once the preemption condition is triggered, the scheduling module immediately freezes the DMA transfer of the current low-priority model, clears its prefetch instruction queue, and mounts the physical address space of the high-priority model to the NPU memory access path through the memory remapping technology. At the same time, it updates the mutex status of the task queue to ensure that the preempted task can reload the context from the checkpoint when it resumes.
[0043] In step S4, when it is necessary to switch to the target model, the parameter difference analysis engine compares the weight matrix of the currently running model with the binary image of the target model and marks the different blocks with a bitmap. The loading engine generates a minimized transfer sequence based on the difference bitmap and realizes the skip transfer of the discontinuous address space through the chained DMA descriptor. For the convolutional layer parameters, the block-by-block incremental coding technology is adopted to only transfer the parameter blocks whose changes exceed the preset threshold, and the CRC check code is combined to ensure the transfer integrity, reducing the typical model switching time from 50 ms in the traditional scheme to within 3 ms.
[0044] In step S5, at the moment of model switching, the firmware layer triggers a dedicated context management coprocessor to encrypt and write the register state of the current model (including accumulators, address pointers, and pipeline buffer data) into a non-blocking buffer. The pre-stored context of the target model is quickly restored to the register set through a hardware-accelerated decompression engine. Meanwhile, the memory management unit (MMU) synchronously switches the page table base address register to achieve a seamless switch of the virtual address space. This process utilizes the NPU internal bus bandwidth isolation technology to ensure that context transmission does not affect the data stream of high-priority tasks being executed.
[0045] In step S6, the NPU computing resources are divided into multiple virtualized partitions. The high-priority model exclusively occupies the main computing array to execute inference tasks, while the coprocessor unit concurrently executes parameter prefetching and weight initialization decoding of low-priority models. The memory controller adopts an interleaved access mode to disperse model weight loading requests to different DRAM Bank groups, combined with the bus clock phase offset technology to eliminate memory access conflicts. For quantized models, the fixed-point format conversion hardware is started during the preloading phase to reduce the runtime computing latency.
[0046] In step S7, the embedded performance probe continuously collects data such as the L1 cache hit rate, DMA transfer error rate, and task switching time series, and generates a resource utilization heat map through an online learning module. The feedback engine dynamically adjusts the cache partition ratio according to the heat map characteristics: when it detects that the cache hit rate of the semantic segmentation model continuously drops below 40%, it automatically elevates its storage level to SRAM and correspondingly reduces the cache priority of the speech recognition model. Meanwhile, the priority weight coefficient is adjusted for attenuation or gain based on the task timeout history record.
[0047] In step S8, the task termination instruction triggers a multi-level resource recovery pipeline: first, it releases the binding relationship between the model and the computing unit and clears the corresponding TLB entries and cache tags; then, it determines the model reuse possibility according to the reference count table. High-frequency models retain metadata and are marked as standby, while low-frequency models trigger a storage recovery interrupt; finally, it performs physical memory fragmentation sorting, merges the released storage blocks and adds them to the free pool, and optimizes the storage alignment efficiency of subsequent models through address remapping.
[0048] In step S9, the system starts an offline optimization engine during the idle period to mine the spatio-temporal patterns of historical task logs. It trains a cache prediction model using a reinforcement learning framework to generate cache policy templates for different scenarios. Meanwhile, it analyzes the success rate and error patterns of incremental loading and optimizes the parameters of the differential bitmap generation algorithm. After the updated policy file passes the digital signature verification, it is written into the policy storage area of the firmware through a secure boot process to achieve the adaptive evolution of system-level parameters.
[0049] It can be known from the above that:
[0050] In the present invention, the system real-time performance and resource utilization rate in multi-task scenarios are significantly improved. Through the dynamic hierarchical caching strategy and the priority scheduling mechanism, the system can intelligently allocate cache resources, giving priority to ensuring the rapid loading of high-frequency and high-urgency models. For example, in the autonomous driving scenario, the switching delay of the path planning model can be reduced to the millisecond level, ensuring seamless switching of critical tasks. The incremental parameter loading technology is deeply integrated with the firmware-level context management, only transmitting the model difference data and using hardware acceleration to restore the computing state, greatly reducing the bandwidth waste caused by traditional full-scale loading. When edge computing devices perform multi-tasks in parallel such as speech recognition and image processing, the model switching efficiency is increased by more than 10 times while maintaining the computing accuracy without loss.
[0051] In the present invention, long-term operation stability and scenario adaptability are achieved through hardware collaborative optimization and self-feedback mechanism. The priority preemption strategy combined with real-time performance monitoring can dynamically adjust resource allocation when sudden tasks are triggered, avoiding low-priority tasks from blocking critical processes. The periodic global policy optimization continuously improves the caching rules and scheduling algorithms based on historical task patterns, enabling the system to autonomously adapt the best switching strategy in diverse scenarios such as smart homes and data centers. In addition, the hierarchical storage management and parallel resource release mechanism effectively reduce the memory fragmentation problem, supporting the device to operate stably for a long time under limited hardware resources, providing a highly reliable and low-latency general solution for multi-model real-time inference scenarios.
[0052] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0053] The above embodiments are only used to illustrate the technical solutions of the present invention, not to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for NPU firmware to support fast switching of multiple models, characterized in that: The method comprises the following steps: S1: Receive multi-task requests, analyze task types through the priority scheduling module, dynamically assign initial priorities based on task real-time and importance, and generate task execution queues; S2: According to the task frequency and priority, the model cache module preloads high-frequency / high-priority models into the cache, stores low-frequency models into the main storage, and uses a delayed loading mechanism to reduce resource usage; S3: monitors system load and task status in real time, dynamically increases the priority of urgent tasks through the priority scheduling module, and supports high-priority tasks to preempt model resources of low-priority tasks; S4: When switching models, the fast switching mechanism only loads differentiated weights or parameters to avoid overloading the full model, and combines DMA to accelerate data transmission, shortening the switching time to milliseconds; S5: Use the hardware co-optimization module to quickly save and restore the computing context at the firmware level to ensure seamless model switching; S6: Through the built-in parallel computing unit and memory channel of NPU, it supports parallel processing of high-priority model inference and low-priority model preloading, maximizing hardware resource utilization; S7: Collect task delay, cache hit rate, and hardware resource usage indicators, and feed them back to the scheduling module to dynamically optimize the cache strategy and priority weight; S8: After the task is executed, the computing resources occupied by the model are released, and the high-frequency model is retained or the low-frequency model is unloaded according to the cache strategy to free up storage space for subsequent tasks; S9: Based on historical task data, periodically update the hierarchical cache rules and priority algorithm parameters to adapt to long-term task mode changes and improve the system's adaptability.
2. A method for NPU firmware to support fast switching of multiple models as claimed in claim 1, characterized in that: In step S1, after the system receives the multi-model task request from the application layer, the priority scheduling module first identifies the task type, and dynamically generates a priority score table based on the preset real-time indicators and task criticality tags; this module counts the historical task execution frequency based on the sliding window algorithm, assigns a higher initial priority weight to high-frequency tasks, and constructs a bidirectional task queue with time constraints to ensure that high real-time tasks are scheduled first at the front end of the queue, and reserves a dynamic insertion window for burst tasks.
3. A method for NPU firmware to support fast switching of multiple models as claimed in claim 1, characterized in that: In the step S2, the model cache module starts the three-level storage partition according to the priority score table: the highest-scoring model is completely loaded into the NPU tightly coupled SRAM cache; The second-priority model is pruned and compressed and then stored in the shared LLC cache pool; Only metadata descriptors are retained for low-frequency models, and actual parameters are stored in external DDR. The cache manager uses a heat-weight hybrid prediction model to track the number of calls, most recent use timestamps, and task-related data of each model in real time, and dynamically adjust the storage hierarchy. When SRAM space is insufficient, a priority-based eviction strategy is triggered to downgrade the storage of low-frequency models and retain their metadata fingerprints.
4. A method for NPU firmware to support fast switching of multiple models as claimed in claim 1, characterized in that: In step S3, during the task running, the hardware event monitoring unit continuously detects the task execution status; once the preemption condition is triggered, the scheduling module immediately freezes the DMA transmission of the current low-priority model, clears its prefetch instruction queue, and mounts the physical address space of the high-priority model to the NPU memory access path through memory remapping technology; at the same time, the mutex lock status of the task queue is updated to ensure that the context can be reloaded from the checkpoint when the preempted task is restored.
5. The method for NPU firmware to support fast switching of multiple models as claimed in claim 1, characterized in that: In step S4, when it is necessary to switch to the target model, the parameter difference analysis engine compares the weight matrix of the currently running model with the binary image of the target model, and uses a bitmap to mark the difference blocks; the loading engine generates a minimized transmission sequence based on the difference bitmap, and realizes the jump transmission of non-continuous address space through the chained DMA descriptor; for the convolution layer parameters, the block incremental encoding technology is adopted, and only the parameter blocks whose changes exceed the preset threshold are transmitted, and the CRC check code is combined to ensure the transmission integrity, thereby shortening the typical model switching time from 50ms of the traditional solution to less than 3ms.
6. A method for NPU firmware to support fast switching of multiple models as claimed in claim 1, characterized in that: In step S5, at the moment of model switching, the firmware layer triggers the dedicated context management coprocessor to encrypt the register state of the current model and write it into the non-blocking cache area; the pre-stored context of the target model is quickly restored to the register group through the hardware accelerated decompression engine, and the memory management unit synchronously switches the page table base address register to achieve seamless switching of the virtual address space; This process utilizes NPU internal bus bandwidth isolation technology to ensure that context transfer does not affect the data flow of high-priority tasks being executed.
7. A method for NPU firmware to support fast switching of multiple models as claimed in claim 1, characterized in that: In step S6, the NPU computing resources are divided into multiple virtualized partitions, the high-priority model exclusively occupies the main computing array to perform reasoning tasks, and the co-processing unit performs parameter pre-fetching and weight initialization decoding of the low-priority model in parallel; The memory controller uses an interleaved access mode to distribute model weight loading requests to different DRAM Bank groups, and combines bus clock phase offset technology to eliminate memory access conflicts. For quantized models, the fixed-point format conversion hardware is started during the preloading phase to reduce runtime computing delays.
8. The method for NPU firmware to support fast switching of multiple models as claimed in claim 1, characterized in that: In step S7, the embedded performance probe continuously collects L1 cache hit rate, DMA transfer error rate, and task switching time series data, and generates a resource utilization heat map through an online learning module; the feedback engine dynamically adjusts the cache partition ratio according to the heat map characteristics: when it is detected that the cache hit rate of the semantic segmentation model is continuously lower than 40%, its storage level is automatically upgraded to SRAM, and the cache priority of the speech recognition model is correspondingly reduced; at the same time, the priority weight coefficient is attenuated or gain adjusted according to the task timeout history record.
9. A method for NPU firmware to support fast switching of multiple models as claimed in claim 1, characterized in that: In step S8, the task termination instruction triggers a multi-level resource recovery pipeline: first, the binding relationship between the model and the computing unit is released, and the corresponding TLB entries and cache tags are cleared; then, the possibility of model reuse is determined according to the reference count table, and the high-frequency model retains metadata and is marked as standby, and the low-frequency model triggers a storage recovery interrupt; finally, physical memory defragmentation is performed, and the released storage blocks are merged and added to the free pool.
10. The method for NPU firmware to support fast switching of multiple models as claimed in claim 1, characterized in that: In step S9, the system starts the offline optimization engine during the idle period to perform spatiotemporal pattern mining on the historical task logs; The cache prediction model is trained using a reinforcement learning framework to generate cache policy templates for different scenarios. The success rate and error pattern of incremental loading are analyzed to optimize the parameters of the difference bitmap generation algorithm. The updated policy file is digitally signed and verified before being written into the firmware’s policy storage area through a secure boot process.
Citation Information
Cited By
Distributed computing power scheduling method and system based on AIGC
CN120508404A
Firmware loading method and device based on dual-chip collaboration
CN120704717A
End side intelligent model-based intelligent controller architecture with body and operation method
CN120742773A
Dynamic energy consumption optimization method and device for edge device speech recognition and medium
CN120877739A
Video target identification method based on multi-model hot switching
CN122116247A