The application discloses a
model inference acceleration method,
system, device and medium of a vehicle-mounted
edge device, the method comprising: by analyzing the
inference task request and dynamically hot loading the model weight, mapping the pre-filling stage and the decoding stage to different
CUDA streams for
parallel processing, and introducing a forced
preemption mechanism to ensure the real-time performance of high-priority tasks. Through
paging memory management and KV Cache hierarchical replacement strategy, the memory resources are optimized, and combined with
power consumption monitoring and bottom layer hardware
instruction set optimization, the technical problems of
low resource utilization, high-priority task
delay and excessive
power consumption of the existing vehicle-mounted
edge device in multi-task
concurrency are solved. The application can effectively improve the
inference throughput and
system stability of the model in the vehicle-mounted
edge device, solve the technical problems of
low resource utilization, high-priority task
delay and excessive
power consumption of the existing vehicle-mounted edge device in multi-task
concurrency, and is suitable for complex computing scenarios of intelligent vehicles.