Adaptive CPU-GPU AI Inference System Based on Model Family and Runtime Metrics
Patent Information
- Authority / Receiving Office
- TR · TR
- Patent Type
- Applications
- Current Assignee / Owner
- TURKCELL TEKNOLOJI ARASTIRMA & GELISTIRME AS
- Filing Date
- 2026-05-18
- Publication Date
- 2026-06-22
Smart Images

Figure 00000014_0000
Abstract
Description
- 1 - TARIFF ADAPTIVE BASED ON MODEL FAMILY AND WORKING TIME METRICS CPU-GPU Artificial Intelligence Inference System TECHNICAL FIELD The invention enables real-time AI inference and image processing on edge devices. in the areas of processing, hardware acceleration and CPU / GPU resource optimization, AI models that operate via camera or video stream, device 10 central processing unit (CPU) and graphics processing unit (GPU) resources adaptive based on model family and runtime performance metrics. a family of models that enables guidance in this way and runtime metrics The invention relates to an adaptive CPU-GPU artificial intelligence inference system. The system uses OpenCV-based image preprocessing, OpenVINO-based CPU inference, 15 ONNX / CUDA-based GPU inference and NeuralEngine wrapper layer using appropriate inference lines for different model families This invention provides artificial intelligence-powered image analytics, and advanced capabilities. IT, real-time object detection, face / pose / keypoint analysis, smart camera. systems, industrial Internet of Things (IoT), smart store solutions, 20 security systems, smart city applications, and low-latency edge device data It can be used in the field of processing systems. PREVIOUS TECHNIQUE In current approaches, AI inference on edge devices is mostly done via GPUs. This is done based on either acceleration or fixed backend selection. In such configurations, CPU resources are often used only for preprocessing or control. It is used for tasks, while the GPU is central for all heavy inference tasks. It is positioned as an accelerator. This means that the GPU has an excessive 30% load. load, temperature increase, power consumption increase, decreasing frame rate, and reality This leads to technical problems such as timely processing delays. - 2 - Document CN116432754A describes AI inference from video streaming. the process involves different processing units such as CPU, GPU, NPU and DPU working together It describes speeding up through the use of cameras. In the solution in question, the camera... The CPU splits the data into squares, distributing tasks to different units. While there is an approach to combining the results, there are 5 model family-based approaches. OpenVINO / CPU and ONNX / CUDA distinction, NeuralEngine wrapper layer, mathematical cost function and backend selection, hysteresis-based access control, and The feedback model profile update mechanisms are explicitly included. It does not receive. Document number CN110333946A describes CPU, GPU, and FPGA processing units as 10 By using them together, general-purpose high-performance computing can be achieved. It defines the CPU as the main control node in the solution in question. heterogeneous computing architecture that distributes tasks and combines results However, based on a family of models specific to real-time image streaming. Adaptive backend routing, OpenCV DNN-based NeuralEngine wrapper, 15 Cost function based on runtime metrics and hysteresis decision the mechanism is missing and the CPU is used as the active inference pipeline It is not specified. OpenVINO AUTO Device Selection and related commercial and technical documentation, as well as academic information. Studies explore automatic selection or heterogeneous planning between CPU / GPU 20 It presents approaches and the integrated NeuralEngine layer of the invention, the model profiled mathematical cost function, hysteresis-based access control and feedback It does not include a feed-based profile update mechanism. Current technical documentation offers either a general task allocation or a static selection process. and based on model family and runtime metrics (CPU / GPU load, temperature, power 25 Adaptive CPU / OpenVINO based on consumption, latency, dropped frame rate, confidence score. and GPU / CUDA routing, common pattern management with NeuralEngine wrapper. and a stable real-time inference architecture with a hysteresis cost function It does not present the GPU in an integrated way. Therefore, current solutions avoid overloading, keep CPU resources passive, energy efficiency 30 It cannot solve these problems and ensure long-term edge device stability within a single architecture. - 3 - A BRIEF DESCRIPTION OF THE INVENTION Real-time AI inference and image processing on edge devices, In the field of hardware acceleration and CPU / GPU resource optimization, for cameras or Model family and study 5 of artificial intelligence models that work with video streaming. invention for adaptively managing time based on performance metrics. The subject is adaptive CPU / GPU based on model family and runtime metrics. An artificial intelligence inference system has been developed. The developed system includes real-time data acquisition, preprocessing, and model profiling. management, OpenVINO CPU inference, ONNX / CUDA GPU inference, NeuralEngine 10 Wrapper management, runtime performance monitoring, mathematical backend through selection, results comparison, and feedback profile update steps working to create more efficient, low-latency, and energy-efficient artificial intelligence on edge devices. It enables inference. The developed system supports different model families (MediaPipe Palm Detection, FaceMesh, Pose Estimation, MobileNet-SSD, YOLOv5 etc.) 15 Model profile and instantaneous OpenVINO / CPU and ONNX / CUDA / GPU lines. by guiding according to its metrics, cost function, hysteresis access control and It offers integrated management with the NeuralEngine wrapper and GPU overclocking. It reduces the load, energy consumption, and the falling square rate. This Integrated architecture, sustainable in heterogeneous edge devices including CPU and GPU 20 It improves real-time image analytics performance. DESCRIPTION OF THE FIGURES Figure 1. Adaptive 25 Based on Model Family and Runtime Metrics. CPU-GPU Artificial Intelligence Inference System System Architecture The corresponding part numbers shown in the figures are given below. 1. Real-Time Data Acquisition and Frame Generation Module 30 2. OpenCV Preprocessing and Common Data Interface Module 3. Model Registration and Model Working Profile Module - 4 - 4. OpenVINO Model Conversion and CPU Inference Module 5. ONNX Model Conversion and CUDA / GPU Inference Module 6. NeuralEngine Base Class / Wrapper Module 7. Runtime Performance Monitoring Module 8. Mathematical Backend Selection and Load Balancing Module 5 9. Model Family Based Inference Path Selection Module 10. Results Comparison and Final Output Selection Module 11. Feedback Profile Update Module 12. Adaptive Operating Mode Determination Module 13. Optimized Output and Recording Module 10 DETAILED DESCRIPTION OF THE INVENTION Real-time AI inference and image processing on edge devices, In the area of hardware acceleration and CPU / GPU resource optimization, the camera or 15 Model family and study of artificial intelligence models that work with video streaming. at least for time to be directed adaptively according to performance metrics. a central processing unit (CPU), graphics processing unit (GPU), memory units and data on a heterogeneous edge device with CPU and GPU containing storage units The family of model subjects developed for operation and the runtime of the invention are 20 Adaptive CPU / GPU AI inference system based on metrics; real-time Data acquisition and frame generation module (1), OpenCV preprocessing and common data interface module (2), model registration and model working profile module (3), OpenVINO model conversion and CPU inference module (4), ONNX model conversion and CUDA / GPU inference module (5), NeuralEngine base class / wrapper module 25 (6), runtime performance monitoring module (7), mathematical backend selection and load balancing module (8), model family based inference path selection module (9), Result comparison and final output selection module (10), feedback profile update module (11), adaptive operating mode determination module (12) and optimize It includes the generated output and recording module (13). 30 The invention is based on real-time data from an artificial intelligence inference system. Acquisition and frame generation module (1); USB camera, IP camera, RTSP video stream or - 5 - Real-time streaming from a GStreamer-compatible video source via OpenCV It receives the image stream. Real-time data acquisition and frame production module (1), It splits the incoming video stream into timestamped frames, and the frame generation rate is decreasing. By measuring the number of frames and the time difference between frames, Fdrop(t) = DroppedFrames(t) / The formula TotalFrames(t) is used to calculate the rate of fallen squares, and this value is set to 5. The backend uses it in selection decisions. Real-time data retrieval and frame rate. production module (1), on the camera interfaces and memory resources of the end device It works by providing standard square input for subsequent layers. The invention concerns OpenCV preprocessing within an artificial intelligence inference system. and common data interface module (2); raw image conforms to model requirements 10 resizing, color space transformation, pixel normalization, Channel sorting, trimming, scaling, and batch format conversion. It performs the following operations: OpenCV preprocessing and common data interface module. (2) provides a common data interface to OpenVINO / CPU and CUDA / GPU lines and Inlet size, channel structure, and normalization method for each model: Model 15 It produces consistent data representation by retrieving data from the Work Profile. OpenCV preprocessing and Common data interface module (2), efficient front on CPU and memory resources It processes inputs to produce suitable inputs for different model families. The invention concerns model registration and the artificial intelligence inference system. Model working profile module (3); profile 20 for each model Mᵢ = {Tᵢ, Rᵢ, Sᵢ, Cᵢ, Pᵢ, Aᵢ} It keeps the vector on record. Model registration and model working profile module. (3), models Class-1 (low / medium weighted, CPU prioritized), Class-2 (dynamic) and Classifying it as Class-3 (intensive, GPU-prioritized), the initial default background Defining backend assignments and runtime performance It updates these assignments according to its measurements. Model registration and model operation 25 profile module (3) keeps profile data on the memory units of the end device It forms the basis for working time decisions. The invention is based on the OpenVINO model, an artificial intelligence inference system. conversion and CPU inference module (4); models to OpenVINO IR format It converts. OpenVINO model conversion and CPU inference module (4), 30 Models like MediaPipe Palm Detection, FaceMesh, and Pose Estimation support the CPU. It performs active inference, switching the CPU from a passive control unit to an active one. - 6 - converting it to an inference pipeline and in each cycle, the CPU inference time, load rate, It generates temperature, power consumption, and safety score. OpenVINO model Transformation and CPU extraction module (4), the device's central processing unit resources It reduces GPU load by using it as an active inference pipeline. The invention concerns the ONNX model 5, an artificial intelligence inference system. conversion and CUDA / GPU extraction module (5); models to ONNX format It converts. ONNX model conversion and CUDA / GPU inference module (5), Extraction of intensive patterns like YOLOv5 on a CUDA-enabled GPU It performs, GPU load rate, temperature, power consumption, memory usage and The model measures the confidence score and the device's graphics processing unit resources. It generates GPU metrics. ONNX model conversion and CUDA / GPU inference. module (5) enables efficient operation of models requiring high parallel processing. It provides. The invention is based on the NeuralEngine, which is part of the artificial intelligence inference system. class / wrapper module (6); each model 15 by wrapping OpenCV DNN modules for loadModel (model path, device type), preprocess (raw data, target format), selectBackend (hardware infrastructure, working engine), infer (prepared input, confidence threshold), postprocess (inference output, result labels), measureRuntime (total time, performance data) and updateProfile (user settings, system It standardizes the functions (update). NeuralEngine base class / 20 The wrapper module (6) allows all models to use a common software interface. loading, preprocessing, backend selection, inference, and result. Production and performance metric reporting under a single control logic. It combines the NeuralEngine base class / wrapper module (6), end device Integrated model management on processor and memory resources 25 This ensures the modular structure of the system. The invention concerns the runtime involved in the artificial intelligence inference system. Performance monitoring module (7); CPU / GPU load ratio (Lcpu(t), Lgpu(t)), temperature (Tcpu(t), Tgpu(t)), power consumption (Pcpu(t), Pgpu(t)), memory usage, inference It monitors the delay, the falling squares rate (Fdrop(t)), and the model confidence score. 30 Runtime performance monitoring module (7), device resource status vector D(t) = {Lcpu(t), Lgpu(t), Tcpu(t), Tgpu(t), Pcpu(t), Pgpu(t), Bmem(t), Fdrop(t)} - 7 - It generates and calculates moving averages for the last N squared. The study time performance monitoring module (7), end device hardware sensors and processor By providing real-time metrics through its resources, it enables adaptive decision-making. It forms the basis of the mechanism. The invention concerns the mathematical background involved in the artificial intelligence inference system. End selection and load balancing module (8); CostCPU(i,t) = α·Lcpu(t) for each model + β·Tcpu(t) + γ·Pcpu(t) + δ·LatCPU(i,t) − ε·AccCPU(i) and CostGPU(i,t) cost It calculates the functions. Mathematical backend selection and load balancing. module (8) makes the Backend(i,t) decision with hysteresis coefficient H, CostCPU(i,t) + In the case of H < CostGPU(i,t), select the CPU / OpenVINO line and 10 In the case where |CostCPU(i,t) − CostGPU(i,t)| ≤ H, the previous backend It protects. The mathematical backend selection and load balancing module (8) of the device By performing load balancing on CPU and GPU resources, the system It strengthens its resolve. The invention is based on a family of 15 models within an artificial intelligence inference system. Inference path selection module (9); initial backend according to model technical profile It is making assignments and transferring Class-1 models to the OpenVINO / CPU pipeline, Class-3 It directs its models towards the ONNX / CUDA pipeline. Model family-based inference. The path selection module (9) dynamically assigns with runtime metrics. updating and monitoring profiles and metrics on memory and processor resources 20 It provides guidance by evaluating the situation. The invention concerns the comparison of results within an artificial intelligence inference system. and final output selection module (10); results from different backends R(i,b,t) = λ₁·Conf(i,b,t) + λ₂·Stab(i,b,t) − λ₃·Lat(i,b,t) − λ₄·Power(b,t) with reliability score It evaluates. Result comparison and final output selection module (10), object 25 In its determination, IoU = Area(BoxA ∩ BoxB) / Area(BoxA ∪ BoxB) and the key point In their models, they perform consistency checks with the normalized Euclidean distance DKP and the most It selects the result with the highest reliability score as the final output. Result Comparison and final output selection module (10), end device processor resources It produces a stable output. 30 The invention concerns a feedback profile within an artificial intelligence inference system. Update module (11); model delay, power after each inference cycle - 8 - It records consumption, temperature, and accuracy / reliability score. It is feedback-based. The profile update module (11) updates the Model Working Profile, GPU When temperature or power consumption thresholds are exceeded, the models are switched to the CPU / OpenVINO line. When CPU latency increases, it shifts to the CUDA / GPU bus and memory units. It supports long-term adaptation. 5 The invention focuses on adaptive computing within an artificial intelligence inference system. mode selection module (12); performance, balanced, energy saving and thermal It switches between protection modes. Adaptive operating mode selection. module (12) evaluates metrics such as GPU temperature and power consumption to model Optimize operating frequency, input resolution, and backend preference together. and guarantee sustainable operation within the limitations of end-device hardware. is doing. The invention concerns an optimized artificial intelligence inference system. Output and recording module (13); the final result, selected backend information, extraction time, power consumption, temperature value, confidence score, model run 15 It records the profile and the falling frame rate. Optimized output and recording. module (13) coordinates the recording operations on the data storage units of the end device. This increases the system's accountability. The invention involves acquiring real-time data in an artificial intelligence inference system and Real-time frames from camera or video stream with frame generation module (1) 20 The data is being retrieved and the falling frame rate is being calculated, along with OpenCV preprocessing and shared data. With the interface module (2), images are prepared for model inputs, model registration and model profiles are loaded and classified with the model working profile module (3), OpenVINO model conversion and CPU inference module (4) with CPU line active Performing inference, ONNX model conversion and CUDA / GPU inference 25 With module (5), GPU line intensive models are activated, NeuralEngine All models via a common interface with the base class / wrapper module (6) It is managed by the runtime performance monitoring module (7) with the D(t) vector and Moving averages are generated, mathematical backend selection and load balancing are performed. Backend decision using cost functions and hysteresis H with module (8) 30 is provided with the model family based inference path selection module (9) and model profile. Guidance is provided accordingly, and the results comparison and final output selection module is used. - 9 - (10) The final result is selected using the R(i,b,t) score, IoU and Dkp criteria, Profiles are dynamically updated with the feedback profile update module (11). It is being updated with the adaptive operating mode determination module (12) to the device conditions. Depending on the performance, the system switches to balanced, energy-saving, or thermal protection mode. and with the optimized output and recording module (13) all results, metrics and 5 Decisions are being recorded. Central processing unit (CPU), graphics processing unit (GPU), on a heterogeneous end device containing memory units and data storage units The AI inference system in operation; model family and runtime metrics. Adaptive routing based on NeuralEngine wrapper, mathematical cost The function includes hysteresis-based transition control and feedback profile updating. It differs from classic fixed back-end systems. 20 30
Claims
- 10 - SYSTEMS 1. Real-time AI inference and image processing on edge devices, In the field of hardware acceleration and CPU / GPU resource optimization, the camera or a family of AI models that work via video streaming and 5 adaptive guidance based on work time performance metrics for at least one central processing unit (CPU), graphics processing unit (GPU), memory Heterogeneous edge devices with CPU and GPU containing data storage units and components. The family of model designs and the work to be carried out on the subject of the invention It is an adaptive CPU / GPU AI inference system based on time metrics, 10 Features; real-time data acquisition and frame generation module (1), OpenCV front Processing and common data interface module (2), model registration and model working profile module (3), OpenVINO model conversion and CPU inference module (4), ONNX Model conversion and CUDA / GPU inference module (5), NeuralEngine basic class / wrapper module (6), runtime performance monitoring module (7), 15 mathematical backend selection and load balancing module (8), model family based inference path selection module (9), result comparison and final output selection module (10), feedback profile update module (11), adaptive operating mode Detection module (12) and optimized output and recording module (13) It is characterized by its inclusion. 20 2. It is an artificial intelligence inference system according to claim 1, and its features include; USB camera, IP from camera, RTSP video stream or GStreamer compatible video source Real-time data acquisition and frame production module receiving image stream (1) It is characterized by its inclusion.
3. It is an artificial intelligence inference system according to claim 1, and its feature is; it analyzes the incoming video stream 25 Dividing into timestamped frames and Fdrop(t) = DroppedFrames(t) / Real-time data that calculates the rate of falling squares using the TotalFrames(t) formula. It is characterized by containing the acquisition and square generation module (1).
4. It is an artificial intelligence inference system according to claim 1, and its feature is; it analyzes the raw image. Resizing, color space transformation, pixel normalization, and batch 30 OpenCV preprocessing and common processes perform the conversion operations to the specified format. It is characterized by containing a data interface module (2). - 11 - 5. It is an artificial intelligence inference system according to claim 1, and its features include: OpenVINO / CPU and Providing a common data interface to CUDA / GPU lines and Model Work OpenCV front-ends generate consistent data representation by retrieving input parameters from the profile. It is characterized by containing the processing and common data interface module (2).
6. It is an artificial intelligence inference system according to claim 1, and its characteristic is; Mᵢ = 5 for each model. Model record and model working profile holding the profile vector {Tᵢ, Rᵢ, Sᵢ, Cᵢ, Pᵢ, Aᵢ} It is characterized by containing module (3).
7. It is an artificial intelligence inference system according to Claim 1, and its characteristic is that its models are Class-1. Classifying initial backend assignments as Class-2 and Class-3. Model record 10 defines and updates assignments with runtime metrics. and is characterized by including the model working profile module (3).
8. It is an artificial intelligence inference system according to claim 1, and its characteristic feature is that its models are OpenVINO. Performing active inference on the CPU by converting to IR format. OpenVINO includes a model conversion and CPU inference module (4) It is characterized by: 15 9. It is an artificial intelligence inference system according to Claim 1, and its feature is; MediaPipe Palm Models like Detection, FaceMesh, and Pose Estimation are on the CPU line. OpenVINO enables this operation and transforms the CPU into an active inference pipeline. It is characterized by containing a model transformation and CPU inference module (4).
10. It is an artificial intelligence inference system according to Claim 1, and its feature is that its models are ONNX 20. ONNX performs inferences on CUDA-enabled GPUs by converting the data to a specific format. with model conversion and CUDA / GPU inference module (5) It is characteristic.
11. It is an artificial intelligence inference system according to claim 1, and its feature is that it uses dense elements like YOLOv5. 25 that enable the models to run on the GPU pipeline and generate GPU metrics ONNX includes model conversion and CUDA / GPU inference module (5) It is characteristic.
12. According to claim 1, it is an artificial intelligence inference system, and its characteristic is; loadModel (model path, device type), preprocess (raw data, target format), selectBackend (hardware infrastructure, working engine), infer (prepared input, confidence threshold), postprocess 30 (inference output, result labels), measureRuntime (total time, performance) the functions updateProfile (user settings, system update) and updateData (user settings, system update). - 12 - by including the NeuralEngine base class / wrapper module (6) which standardizes It is characteristic.
13. It is an artificial intelligence inference system according to claim 1, and its characteristic is that all models NeuralEngine is the basis for managing software through a common software interface. It is characterized by containing the class / wrapper module (6). 5 14. It is an artificial intelligence inference system according to claim 1, and its characteristic is; CPU / GPU load rate, temperature, power consumption, memory usage, latency, frame rate drop, and Runtime performance monitoring module that monitors the confidence score (7) It is characterized by its inclusion.
15. It is an artificial intelligence inference system according to claim 1, and its feature is; it generates a vector D(t). and runtime performance monitoring that calculates moving averages It is characterized by containing module (7).
16. It is an artificial intelligence inference system according to claim 1, and its feature is; CostCPU(i,t) and CostGPU(i,t) is a mathematical backend selection that calculates cost functions. and is characterized by having a load balancing module (8). 15 17. According to claim 1, it is an artificial intelligence inference system, and its characteristic is the hysteresis coefficient. H is a mathematical backend selection and load balancing system that makes backend decisions. It is characterized by containing module (8). According to claim 18, it is an artificial intelligence inference system, and its characteristic is; according to the model family. 20 model family-based systems that perform initial assignments and update them at runtime. It is characterized by containing an inference path selection module (9).
19. It is an artificial intelligence inference system according to claim 1, and its characteristic is; R(i,b,t) reliability. The results are evaluated using the score, and consistency checks are performed using IoU and Dkp. It is characterized by containing a comparison and final output selection module (10).
20. It is an artificial intelligence inference system according to claim 1, and its characteristic is that each inference is 25 Updates model profiles after the cycle and backends when thresholds are exceeded. (backend) scrolling feedback profile update module (11) It is characterized by its inclusion.
21. It is an artificial intelligence inference system according to Claim 1, and its characteristic is performance, Adaptive 30 that switches between balanced, energy-saving and thermal protection modes. It is characterized by containing a working mode determination module (12). - 13 - According to Claim 1, it is an artificial intelligence inference system whose characteristic is that the final result is... Optimized output that records the selected backend and metrics. It is characterized by containing a registration module (13). 10 20 30