面向高并发API服务的AI网关多级缓存同步方法和系统
By employing a multi-level caching synchronization method in the AI gateway, request routing and security authentication are performed in memory. Leveraging the high-speed read and write characteristics of memory, the problems of traffic management and load balancing for high-concurrency API services are solved, resulting in an AI gateway system with high concurrency and low latency response.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU PINGPONG INTELLIGENT TECH CO LTD
- Filing Date
- 2025-05-28
- Publication Date
- 2026-07-17
AI Technical Summary
AI gateways face challenges in traffic management and load balancing when handling high-concurrency API services. Furthermore, the high cost and frequent switching required when users use multiple large AI models lead to performance bottlenecks and increased latency.
A multi-level caching synchronization method is adopted to perform request routing, security authentication and flow control in memory. It leverages the high-speed read and write characteristics of memory to reduce dependence on backend cache and database. Data consistency is maintained through a publish-subscribe mechanism, supporting high-concurrency request processing.
It achieves ultra-high throughput and low latency response, reduces reliance on backend caching and databases, avoids performance bottlenecks, supports tens of thousands of concurrent requests per second, and ensures data consistency and system stability.
Smart Images

Figure CN120528975B_ABST