Distributed resource dynamic scheduling method based on scene awareness
Through the scenario-aware distributed resource dynamic scheduling method, the resource allocation of domestic applications is identified and prioritized in real time, solving the static allocation limitations and insufficient domestic adaptation problems in the resource scheduling of the Xinchuang Cloud Computer, achieving efficient resource scheduling and the smoothness of domestic applications, and adapting to domestic operating systems and hardware platforms.
Patent Information
- Application Number
- CN202510831973.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-10-14
AI Technical Summary
The existing Xinchuang cloud computer resource scheduling solution has static allocation limitations, insufficient domestic adaptation and priority conflicts. It cannot dynamically adapt to the real-time computing power requirements of scenarios such as video rendering and document processing, and has not optimized resource allocation for domestic operating systems such as Kylin and Tongxin.
A distributed resource dynamic scheduling method based on scenario perception is adopted. By collecting user operation behavior and system resource status data in real time, a lightweight convolutional neural network is used to identify scenario types, generate dynamic resource allocation strategies, give priority to CPU/GPU/storage resources for domestic applications, combine resource allocation with domestic operating system kernel interfaces, and achieve elastic scaling of resources by adjusting strategies through feedback data.
Significantly reduces the response delay of domestic applications, improves the stability of video rendering frame rate, supports domestic operating systems such as Kylin and Tongxin, reduces resource contention and overall power consumption, and is compatible with domestic chip hardware platforms such as Feiteng and Kunpeng.
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the technical field of Xinyuan localization software, and particularly relates to a distributed resource dynamic scheduling method based on scene perception. BACKGROUND
[0002] Cloud computer resource scheduling centrally manages computing resources through virtualization technology, dynamically allocates hardware such as CPU, memory and GPU to user terminals, uses intelligent algorithms (such as load balancing and priority queue) to monitor demand in real time, elastically expands and shrinks capacity on demand, ensures performance stability in high-concurrency scenarios, supports resource pooling across physical servers, realizes rapid deployment and migration in combination with containerization technology, reduces delay and improves utilization, and typical applications include game streaming, remote office and other low-delay high-quality scenarios.
[0003] However, the existing Xinyuan cloud computer resource scheduling scheme has the following problems:
[0004] ① Static allocation limitation: traditional resource scheduling relies on fixed rules and cannot dynamically adapt to real-time computing power requirements of video rendering, document processing and other scenarios;
[0005] ② Insufficient localization adaptation: resource allocation is not optimized for the kernel characteristics (such as process scheduling mechanism and file system structure) of domestic operating systems such as Kirin and United;
[0006] ③ Priority conflict: when domestic applications and general applications compete for resources, there is a lack of dynamic priority adjustment mechanism. SUMMARY
[0007] The purpose of the present application is to provide a distributed resource dynamic scheduling method based on scene perception to solve the problems of static allocation limitation, insufficient localization adaptation and priority conflict in the background art.
[0008] To achieve the above purpose, the present application provides the following technical scheme: a distributed resource dynamic scheduling method based on scene perception, comprising:
[0009] The following steps:
[0010] Step S1: Real-time collection of user operation behavior and system resource state data;
[0011] Step S2: Identify the current operation scene type through a scene classification model;
[0012] Step S3: Generate a dynamic resource allocation strategy according to the scene type, and prioritize the CPU / GPU / storage resource quota of domestic applications;
[0013] Step S4: Issue resource allocation instructions through the kernel interface of the domestic operating system and monitor the execution effect;
[0014] Step S5: Adjust the strategy based on the feedback data to realize resource elasticity expansion.
[0015] Further, the scene classification model in step S2 is a lightweight convolutional neural network, and the input features include application window title, GPU memory occupancy rate, and CPU instruction set usage mode, and the output is a set of predefined scene labels.
[0016] Further, the resource allocation strategy in step S3 includes:
[0017] ① Reserve at least 50% GPU computing power and large core CPU threads for video rendering type domestic applications;
[0018] ② Assign low-latency memory channels and NVMe storage cache areas to document processing type applications;
[0019] ③ Dynamically limit the CPU occupancy rate of non-domestic background processes.
[0020] Further, step S4 dynamically adjusts the process cgroup weight and I / O priority by interacting with the kernel scheduler of Kirin / United System operating system, and modifies the file system log submission strategy to reduce the delay of domestic applications.
[0021] Further, the feedback mechanism of step S5 includes:
[0022] ① Monitor the frame rate and response delay indicators of domestic applications, and trigger GPU computing power elasticity expansion when the threshold is exceeded;
[0023] ② When idle resources are detected, automatically enter low-power sleep state.
[0024] Further, the "domestic application" in step S2 needs to be verified by digital signature whitelist, and the verification process includes:
[0025] ① Extract the hash value of the application developer certificate;
[0026] ② Compare with the pre-set Xinyuan certification certificate library;
[0027] ③ The application that passes the verification is marked as domestic priority.
[0028] Further, the resource allocation weight calculation formula in step S3 is:
[0029] W_{national}=0.6S_{scene}+0.3R_{user}+0.1P_{OS}Wnational=0.6Sscene+0.3Ruser+0.1POS
[0030] Wherein S scene S scene is a scene urgency score, R user R user is a user setting coefficient, and P OS POS is an operating system adaptation coefficient.
[0031] Compared with the prior art, the application has the beneficial effects that:
[0032] (1) The scheduling method can significantly reduce the average response delay of domestic applications, improve the stability of the video rendering frame rate, automatically hibernate idle resources, reduce overall power consumption, support mainstream domestic operating systems such as Kirin V10 and UOS, and be suitable for hardware platforms such as Feiteng and Kunpeng chips.
[0033] (2) The scheduling method can dynamically allocate CPU / GPU / storage resources in real time according to the user operation scene, preferentially guarantee the smoothness of the operation of domestic applications, and optimize the scheduling strategy in combination with the underlying characteristics of the Kirin / UOS operating system to reduce the resource contention rate. DETAILED DESCRIPTION
[0034] The technical solutions in the embodiments of the application are described clearly and completely, and obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the application.
[0035] An embodiment provided by the application is a distributed resource dynamic scheduling method based on scene perception.
[0036] The method comprises the following steps:
[0037] Step S1: Real-time collection of user operation behavior and system resource state data, GPU memory occupancy rate is a key indicator for measuring the use efficiency of GPU memory, representing the percentage of the occupied GPU memory capacity in the total GPU memory capacity, and its core value lies in revealing the data storage pressure: high occupancy rate is often found in large 3A games, AI model training and other scenes, which belongs to normal resource consumption; but if the light application suddenly occupies the memory, or the system is abnormally high when idle, it may be caused by malicious processes or driver faults, which needs to be checked in time.
[0038] Step S2: Identify the current operation scene type through a scene classification model, and the domestic operating system is a key infrastructure for breaking through the core technical bottleneck and realizing the independent controllability of the information industry in China, which is mainly developed based on the Linux kernel and covers multiple scenes such as desktop and server.
[0039] Step S3: Generate dynamic resource allocation strategy according to scene type, prioritize guaranteeing CPU / GPU / storage resource quota for localization application, GPU memory occupancy rate is a key indicator to measure the efficiency of video card memory usage, representing the percentage of current task memory capacity occupied in total memory capacity, its core value lies in revealing data storage pressure: high occupancy rate is common in large 3A games, AI model training and other scenarios, which belongs to normal resource consumption; but if the light application suddenly occupies the memory, or the system is idle with abnormally high running, it may be caused by malicious processes, driver failure, which needs to be checked in time;
[0040] Step S4: Issue resource allocation instructions through domestic operating system kernel interface and monitor execution effect, domestic operating system is a key infrastructure for China to break through core technical bottlenecks and achieve independent controllability in information industry, mainly developed based on Linux kernel, covering desktop, server and other multiple scenarios;
[0041] Step S5: Adjust the strategy based on feedback data to realize resource elastic scaling, resource elastic scaling is the core capability of cloud computing, which automatically adjusts the scale of computing resources through real-time monitoring of business load, realizes rapid expansion in peak period to ensure service stability, and automatically shrinks in low period to reduce cost, its key technologies include dynamic perception trigger mechanism, multi-resource type scheduling and intelligent prediction algorithm.
[0042] Further, the scene classification model in step S2 is a lightweight convolutional neural network, the input features include application window title, GPU memory occupancy rate, CPU instruction set usage mode, and the output is a set of predefined scene labels, GPU memory occupancy rate is a key indicator to measure the efficiency of video card memory usage, representing the percentage of current task memory capacity occupied in total memory capacity, its core value lies in revealing data storage pressure: high occupancy rate is common in large 3A games, AI model training and other scenarios, which belongs to normal resource consumption; but if the light application suddenly occupies the memory, or the system is idle with abnormally high running, it may be caused by malicious processes, driver failure, which needs to be checked in time.
[0043] Further, the resource allocation strategy in step S3 includes:
[0044] ① Reserve at least 50% GPU computing power and large core CPU threads for video rendering type domestic application, video rendering is the process of converting raw materials into continuous pictures through computer algorithm, its core is to combine three-dimensional scene data or edit timeline frame by frame to synthesize final video output, key technology points include: rendering pipeline layered processing, hardware acceleration, algorithm optimization;
[0045] ②Assign low-latency memory channels and NVMe storage cache areas to document processing applications. NVMe storage cache areas are special areas in NVMe solid-state drives for temporarily storing data, which can improve overall performance by accelerating data read and write operations. The core includes DRAM cache, SLC Cache, and HMB technology, which significantly improves 4K random read and write speed and sequential transmission efficiency, and reduces latency. Cache strategy directly affects the stable performance of SSD;
[0046] ③Dynamically limit the CPU occupancy rate of non-domestic background processes. CPU occupancy rate is a core indicator of processor workload, indicating the proportion of CPU processing time, reflecting real-time system performance. Low occupancy means idle resources, and continuous high occupancy may cause lag or overheating. Note that for multi-core CPUs, the overall / single-core data should be considered. Short-term peak is normal, but long-term full load needs to be checked for background programs or viruses.
[0047] Further, the step S4 dynamically adjusts the process cgroup weight and I / O priority by interacting with the kernel scheduler of Kirin / United States operating system, and modifies the file system log submission strategy to reduce the delay of domestic applications. With safety and reliability as the core, it has formed an operating system product line covering desktop, server, Internet of Things, and industry, etc. The latest version of Galaxy Kirin desktop operating system V10 SP1 deeply integrates AI capabilities, becoming the first domestic operating system supporting end-side intelligence, and adapting to Feiteng, Kunpeng, and Longxin domestic CPU platforms through innovative architecture.
[0048] Further, the feedback mechanism of step S5 includes:
[0049] ①Monitor the frame rate and response delay of domestic applications. When the threshold is exceeded, trigger GPU computing power elastic expansion. GPU computing power elastic expansion is a technology that dynamically adjusts GPU resource size to respond to computing demand fluctuations. Its core lies in the use of virtualization, containerization, and intelligent scheduling system to realize flexible allocation of resources. This technology can automatically expand or shrink the size of GPU clusters according to real-time load, supporting training, inference, rendering, and other scenarios, significantly improving resource utilization and reducing costs;
[0050] ②When idle resources are detected, automatically enter low-power sleep state to reduce overall power consumption and save costs. GPU memory occupancy rate is a key indicator of GPU memory usage efficiency, representing the percentage of total GPU memory capacity occupied by the current task. Its core value lies in revealing data storage pressure. High occupancy rate is common in large 3A games, AI model training, and other scenarios, which is normal resource consumption. However, if a lightweight application suddenly occupies all GPU memory or the system runs abnormally high when idle, it may be caused by malicious processes or driver faults, which need to be checked in time.
[0051] Further, the "domestic application" in step S2 needs to be verified by digital signature white list, and the verification process includes:
[0052] ① Extract the hash value of the application developer certificate. Hash value is the unique digital fingerprint of fixed length converted from arbitrary length input data through hash algorithm, which has irreversibility, determinacy and collision resistance;
[0053] ② Compare with the preset Xingcheng authentication certificate library. White list is an authorization mechanism that predefines legal objects. Through the principle of "allowing only explicitly permitted entities to access or operate", it realizes safety control or resource optimization allocation, which is commonly used in financial supervision, network security, education certification and other fields;
[0054] ③ The application that passes the verification is marked as domestic priority. GPU memory occupancy rate is a key indicator to measure the efficiency of video card memory usage, which represents the percentage of the current task's memory capacity to the total memory capacity. Its core value lies in revealing the data storage pressure: high occupancy rate is common in large 3A games, AI model training and other scenarios, which is normal resource consumption; but if a lightweight application suddenly occupies the memory, or the system is idle with abnormally high running, it may be caused by malicious processes, driver failure, which needs to be checked in time.
[0055] Further, the resource allocation weight calculation formula in step S3 is:
[0056] W_{national}=0.6S_{scene}+0.3R_{user}+0.1P_{OS}Wnational=0.6Sscene+0.3Ruser+0.1POS
[0057] Where S_{scene}Sscene is the scene urgency score, R_{user}Ruser is the user setting coefficient, and P_{OS}POS is the operating system adaptation coefficient. Resource allocation weight is determined by quantitative indicators to determine the relative proportion of different subjects in resource allocation.
[0058] 1. Core architecture:
[0059] 1.1, Scene perception module:
[0060] Monitor user behavior data (application type, window focus, input device event);
[0061] Classify scenes (such as video rendering, document editing, multi-task parallel) through lightweight CNN model;
[0062] 1.2, Resource dynamic allocation engine:
[0063] CPU / GPU scheduling:
[0064] Video rendering scenario: reserve 70% GPU computing power, bind large core CPU threads;
[0065] Document processing scenario: allocate low-latency memory channel, enable CPU energy-efficient core;
[0066] 1.3, storage optimization:
[0067] Allocate NVMe cache area for domestic applications, and downgrade cold data to QLC storage pool;
[0068] 1.4, domestic OS coordination layer:
[0069] Interact with Kirin / United Kernel, dynamically adjust process priority (such as increase domestic application cgroup weight);
[0070] Optimize I / O scheduling queue depth based on domestic file system features (such as UKFS).
[0071] 1.5, data flow:
[0072] 1) Scene-aware module collects user behavior and system state in real time;
[0073] 2) Resource allocation engine generates dynamic strategy and issues through OS interface;
[0074] 3) Monitor feedback mechanism adjusts strategy (such as triggering elastic expansion when load exceeds threshold).
[0075] 2, mathematical description of scene classification model:
[0076] Input feature vector: X = [x_1, x_2, …, x_{10}] X = [x1, x2, …, x10], where:
[0077] x_1x1 ~ x_3x3: Window title keyword matching degree (TF-IDF value);
[0078] x_4x4 ~ x_6x6: GPU memory occupancy, CPU AVX512 instruction set usage frequency, disk IOPS;
[0079] x_7x7 ~ x_{10}x10: Mouse movement speed, keyboard event frequency, external device signal strength, network delay;
[0080] Model loss function:
[0081] L = -\frac{1}{N}\sum_{i=1}^{N}\sum_{c=1}^{C}y_{ic}\log(p_{ic})+\lambda\|W\|_2^2L =_N1∑i=1N∑c=1C yic log(pic)+λ∥W∥22
[0082] where CC is the number of scene categories, \lambda = 0.001 \lambda = 0.001 is the L2 regularization coefficient.
[0083] Resource allocation weight calculation (formalized strategy):
[0084] Priority weight of domestic applications:
[0085] W_{national}=\alpha\cdot S_{scene}+\beta\cdot R_{user}+\gamma\cdotP_{OS}Wnational=α·Sscene+β·Ruser+γ·POS
[0086] where:
[0087] S_{scene}Sscene: Scene urgency score (video rendering = 1.0, document editing = 0.7);
[0088] R_{user}Ruser: User manually set priority coefficient (default = 1.0);
[0089] P_{OS}POS: Operating system adaptation coefficient (Kylin = 1.2, United = 1.1);
[0090] \alpha=0.6,\beta=0.3,\gamma=0.1α=0.6,β=0.3,γ=0.1 are adjustable parameters.
[0091] Experimental data comparison
[0092] Index Traditional scheduling scheme Invention scheme Lifting range Domestic application response delay (ms) 220 132 40% Video rendering frame rate stability (%) 72.5 92.1 27% Resource contention rate (times / min) 18.7 6.2 66.8% System overall power consumption (W) 45.3 33.9 25.2%
[0093] Working principle:
[0094] 1. Implementation of scene perception module
[0095] 1.1. Input data:
[0096] Window title keywords (such as "WPS text" "Premiere Pro");
[0097] GPU memory usage, CPU instruction set features (such as AVX512 usage frequency);
[0098] External device events (such as digital pen pressure signals);
[0099] Classification model:
[0100] Using MobileNetV3 lightweight model, input is a 10-dimensional feature vector, and output is a scene label; the model is deployed in user mode, and the inference delay is <1ms.
[0101] 1.2, Resource allocation strategy example
[0102] Scenario 1: Multi-task of domestic office software
[0103] Detecting WPS and Foxit PDF running simultaneously:
[0104] Assign CPU big core to WPS main process, reserve 2 GPU CUDA cores for document rendering; set I / O priority upper limit for UCOS DDE desktop process;
[0105] Scenario 2: 4K video editing
[0106] Detecting that the professional version of Kuaipian calls OpenGL API:
[0107] Dynamically expand GPU virtual memory to 8GB and enable domestic decoder hardware decoding channel;
[0108] Limit the CPU occupancy rate of the background update process to ≤5%.
[0109] 1.3, Cooperative optimization of domestic OS
[0110] Kylin system adaptation:
[0111] Accelerate national secret algorithm through KAE (Kylin Accelerator Engine) interface;
[0112] Modify the kernel CFS scheduler to allocate vruntime compensation coefficient (+15%) for domestic applications;
[0113] UCOS adaptation:
[0114] Hijack resource request events of DDE desktop manager, preload high-frequency application resources;
[0115] Optimize the log submission strategy of UKFS file system to reduce small file write delay.
[0116] 1.4, Domestic application whitelist verification
[0117] Digital signature verification process:
[0118] Extract the SHA-256 hash value of the developer's certificate when the application starts;
[0119] Compare with the pre-installed MIIT digital creation whitelist certificate library;
[0120] The application that passes the verification is granted a "domestic priority" label, triggering a resource reservation policy. The above is the whole working principle of the present application.
[0121] It will be obvious to a person skilled in the art that the application is not limited to the details of the foregoing exemplary embodiments and can be implemented in other concrete forms without departing from the spirit or essential characteristics of the application. The embodiments are therefore to be considered in all respects as illustrative and not restrictive, the scope of the application being indicated by the appended claims rather than by the foregoing description, and all changes which come within the meaning and range of equivalency of the claims are therefore intended to be embraced therein.
Claims
1. A distributed resource dynamic scheduling method based on scenario perception, characterized in that: include: Follow these steps: Step S1: real-time collection of user operation behavior and system resource status data; Step S2: Identify the current operation scene type through the scene classification model; Step S3: Generate a dynamic resource allocation strategy based on the scenario type, giving priority to ensuring the CPU / GPU / storage resource quotas for domestic applications; Step S4: issuing resource allocation instructions through the domestic operating system kernel interface and monitoring the execution effect; Step S5: Adjust the strategy based on the feedback data to achieve elastic scaling of resources.
2. The method for dynamic scheduling of distributed resources based on scenario awareness according to claim 1, characterized in that: The scene classification model in step S2 is a lightweight convolutional neural network, the input features include application window title, GPU memory occupancy, CPU instruction set usage mode, and the output is a predefined scene label set.
3. The method for dynamic scheduling of distributed resources based on scenario awareness according to claim 1, characterized in that: The resource allocation strategy in step S3 includes: ① Reserve at least 50% of GPU computing power and large-core CPU threads for domestic applications such as video rendering; ② Allocate low-latency memory channels and NVMe storage cache areas for document processing applications; ③ Dynamically limit the CPU usage of non-domestic background processes.
4. The method for dynamic scheduling of distributed resources based on scenario awareness according to claim 1, characterized in that: The step S4 dynamically adjusts the process cgroup weight and I / O priority by interacting with the kernel scheduler of the Kylin / Tongxin operating system, and modifies the file system log submission policy to reduce the delay of domestic applications.
5. The method for dynamic scheduling of distributed resources based on scenario awareness according to claim 2, characterized in that: The feedback mechanism of step S5 includes: ① Monitor the frame rate and response delay indicators of domestic applications, and trigger elastic expansion of GPU computing power when the threshold is exceeded; ②When idle resources are detected, it automatically enters a low-power sleep state.
6. The method for dynamic scheduling of distributed resources based on scenario awareness according to claim 2, characterized in that: In step S2, the "domestic application" must pass the digital signature whitelist verification. The verification process includes: ① Extract the hash value of the application developer certificate; ②Compare with the preset Xinchuang authentication certificate library; ③Applications that pass the verification are marked as domestic priority.
7. The method for dynamic scheduling of distributed resources based on scenario awareness according to claim 2, characterized in that: The resource allocation weight calculation formula in step S3 is: W_{national}=0.6S_{scene}+0.3R_{user}+ 0.1P_{OS}Wnational=0.6Sscene+0.3Ruser+0.1POS Where S_{scene}Sscene is the scene urgency score, R_{user}Ruser is the user setting coefficient, and P_{OS}POS is the operating system adaptation coefficient.