Distributed GPU (Graphics Processing Unit) scheduling, data encryption management and dynamic charging system oriented to privatized AI (Artificial Intelligence) deployment

By using technical means such as point-to-point encryption transmission, data sandbox and permission control, and data audit mechanism in the GPU scheduling system, the problems of data security and isolation defects and rigid billing model are solved, and data privacy protection, efficient resource utilization and user experience improvement are achieved.

CN120075225APending Publication Date: 2025-05-30徐绍华
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510290621.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing GPU scheduling system has data security and isolation defects, and the risk of sensitive data leakage is high. At the same time, the billing model is rigid and cannot support small and high-frequency use scenarios for individual users, resulting in waste of resources and a decline in user experience.

Method used

Point-to-point encryption transmission, data sandbox and permission control, data audit mechanism, node weight calculation formula, elastic scaling strategy and micro-billing model are adopted to ensure data security and computing power elasticity.

Benefits of technology

End-to-end encryption is realized to ensure data privacy; data security and isolation are ensured through containerized isolation and permission control; micro-billing is supported, resource waste is reduced, and user experience is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
Patent Text Reader

Abstract

The invention discloses a distributed GPU scheduling, data encryption management and dynamic charging system oriented to privatized AI deployment, and belongs to the technical field of distributed computing and artificial intelligence. The system comprises a decentralized scheduling module, which distributes tasks to distributed GPU nodes through a dynamic weight scoring model (weight factors comprise GPU utilization rate, network delay and task priority); the central control load balancing unit is used for monitoring node states in real time and dynamically adjusting a scheduling strategy; the data encryption module is used for establishing a point-to-point encryption channel between a user terminal and a GPU node based on a WebRTC technology, and data is encrypted by a national cryptographic SM4 algorithm and is decrypted only at a target node; and the dynamic billing unit is used for realizing 0.01-element micro billing according to the GPU computing power consumption and automatically distributing earnings through an intelligent contract. Experiments show that the task completion time of the system in a 100-node cluster is shortened by 30%, the resource utilization rate is improved by 45%, the data leakage risk is reduced by 98%, and the system is suitable for enterprise sensitive data protection and personal user elastic computing power requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of distributed computing and artificial intelligence, and particularly relates to a distributed GPU scheduling, data encryption management and dynamic charging system for private AI deployment, which is especially suitable for individual users and enterprise scenarios with high requirements for data privacy and computing power elasticity. Background Art

[0002] Defects in data security and isolation: Existing GPU scheduling systems (such as the Kubernetes default scheduler) adopt a centralized architecture, and user data needs to be uploaded to a third-party cloud server for processing, which poses a risk of sensitive data leakage (refer to the "2023 Data Breach Investigation Report": 73% of the leakage incidents in enterprise-level AI deployments originate from third-party service providers).

[0003] Rigid charging model: Traditional charging models rely on whole machine leasing or hourly charging (such as AWS EC2 instances), and cannot support the scenarios of small and frequent usage by individual users (such as a charge of less than 0.1 yuan for a single API call), resulting in resource waste and a decline in user experience. Summary of the Invention

[0004] Point-to-point encrypted transmission: An encrypted DTLS-SRTP channel is established between the user terminal and the GPU node through WebRTC, supporting end-to-end encryption (E2EE). The key is generated locally by the user and stored in a hardware security module (HSM); the data is encrypted by the national cipher SM4 algorithm before uploading, and only the target GPU node holds the decryption key.

[0005] Data sandbox and permission control: Containerized isolation technology (such as Docker / Kata Containers) is adopted, and AI model inference can only access the desensitized data set; API-level permission control is implemented based on RBAC (Role-Based Access Control) to reject unauthorized calls (such as when user A has no right to access the face recognition API, the system returns an HTTP 403 status code).

[0006] Data audit mechanism: The uploaded data is subjected to format verification (such as file type, size) and content compliance detection (based on regular expressions / machine learning models) by a preprocessor; the illegal data is automatically marked and an alarm is triggered (such as when an undensitized ID number is detected, the transmission is blocked and the administrator is notified).

[0007] Node weight calculation formula: W = α*(1 - CPU_util) + β*(1 - GPU_util) + γ*(1 - Network_latency), where: α = 0.4, β = 0.5, γ = 0.1 (weights can be adjusted dynamically), and the central control server collects the status of GPU nodes (CPU / GPU utilization, memory, network latency) every 60 seconds; failed API Determination rule: APIs that time out in three consecutive heartbeat detections or have an error rate > 5% will be automatically taken offline.

[0008] Elastic scaling policy: When the API queue length exceeds the threshold (Q > 100), GPU node expansion is automatically triggered (based on Kubernetes HPA); new nodes are automatically deployed and registered to the scheduling cluster through Ansible scripts.

[0009] Experimental data: Performance comparison test: Test environment: 100-node cluster (NVIDIA A100 GPUs), and the task types include image training (ResNet-50) and inference (BERT-Large); Results: Indicator Traditional centralized scheduling The system of the present invention Improvement amplitude Average task completion time 320s 224s 30% Peak GPU utilization rate 65% 94% 45% API call error rate 8% 0.5% 93.75%

[0010] Encryption performance: The encryption throughput of the SM4 algorithm reaches 15 Gbps (Intel Xeon Platinum 8380), and the latency < 1 ms; Billing accuracy: Supports micro-billing at the 0.01 yuan level (test case: 1 million API calls, and the billing error < 0.001%). Description of the drawings System architecture diagram (marking the interaction relationships of the scheduling module, encryption channel, and billing unit).

Claims

1. System architecture: A system for private AI deployment, including: Decentralized GPU scheduling module: dynamically allocates tasks to distributed GPU nodes based on load balancing algorithm; And new nodes can be added or damaged nodes can be deleted dynamically; Data management module: supports encrypted collection (WebRTC technology) and isolated storage of user data within the local area network; Private deployment module: provides a localized deployment interface for personal / enterprise private data and AI models; Dynamic billing unit: Automatically generate bills and distribute revenue based on the GPU usage time and computing power consumption of any terminal.

2. Enhanced data security: Encrypted transmission uses the national secret algorithm (SM4) and is bound to the hardware security module (HSM) to ensure that data is only decrypted in the user's authorized environment. And according to user roles, permissions and groups, it can ensure that user data is only used within the home or enterprise. Scheduling optimization: Introduce a task allocation strategy based on topology awareness to optimize the communication efficiency between multiple machines and multiple cards (RDMA technology).

Citation Information

Cited By

  • Resource regulation and control method, device and system, electronic equipment and storage medium

    CN121691826A