End-side model inference dynamic optimization method and device, equipment and medium
By collecting runtime metrics in the edge processor and building an offline configuration library, and combining offline prior constraints and discrete neighborhood search, low-overhead adjustment of CPU frequency and thread count is performed, solving the real-time and thermal safety issues of the edge processor when running large language models, and achieving efficient dynamic optimization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- VOYAH AUTOMOBILE TECH CO LTD
- Filing Date
- 2026-05-27
- Publication Date
- 2026-07-07
AI Technical Summary
When running large language models, edge processors cannot simultaneously balance real-time performance, thermal safety, and processing efficiency, resulting in insufficient optimization effects.
The system collects operational metrics within the current observation window, aligns the data based on a unified timeline, constructs an offline configuration library, and performs low-overhead joint adjustment of CPU frequency and thread count through discrete neighborhood search under offline prior constraints, combined with minimum dwell time, debouncing, finite memory, and controlled exit mechanisms. It also performs atomic configuration switching at the lexical generation boundary.
It achieves a dynamic balance between real-time performance and energy efficiency during edge model inference, improves thermal stability and robustness, avoids configuration oscillation and local optima problems, and significantly improves the performance of edge processors.
Smart Images

Figure CN122347231A_ABST