End-side model inference dynamic optimization method and device, equipment and medium

By collecting runtime metrics in the edge processor and building an offline configuration library, and combining offline prior constraints and discrete neighborhood search, low-overhead adjustment of CPU frequency and thread count is performed, solving the real-time and thermal safety issues of the edge processor when running large language models, and achieving efficient dynamic optimization.

CN122347231APending Publication Date: 2026-07-07VOYAH AUTOMOBILE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
VOYAH AUTOMOBILE TECH CO LTD
Filing Date
2026-05-27
Publication Date
2026-07-07

AI Technical Summary

Technical Problem

When running large language models, edge processors cannot simultaneously balance real-time performance, thermal safety, and processing efficiency, resulting in insufficient optimization effects.

Method used

The system collects operational metrics within the current observation window, aligns the data based on a unified timeline, constructs an offline configuration library, and performs low-overhead joint adjustment of CPU frequency and thread count through discrete neighborhood search under offline prior constraints, combined with minimum dwell time, debouncing, finite memory, and controlled exit mechanisms. It also performs atomic configuration switching at the lexical generation boundary.

Benefits of technology

It achieves a dynamic balance between real-time performance and energy efficiency during edge model inference, improves thermal stability and robustness, avoids configuration oscillation and local optima problems, and significantly improves the performance of edge processors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122347231A_ABST
    Figure CN122347231A_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an end-side model inference dynamic optimization method, device, equipment and medium. The method surrounds the end-side single-flow CPU inference path, establishes a performance, energy consumption and thermal state measurement system covering the prompt processing stage and the self-recurrence generation stage, and adopts a fixed word window unified index statistics and configuration decision. The offline measurement is performed on the CPU frequency, thread number and affinity configuration space, the feasible boundary and preferred priori are refined, and a discrete neighborhood search dynamic optimization method under the offline priori constraint is proposed. Thus, taking a unified comprehensive cost function as an optimization target, the online search is limited in the feasible neighborhood obtained through offline screening under the first word delay, thermal safety and frequency accessibility constraint, and combined with the minimum residence, anti-shake, limited memory and controlled jump-out mechanism, the low-overhead joint adjustment of the CPU frequency and thread number is realized, while the adaptability to state changes is improved.
Need to check novelty before this filing date? Find Prior Art