An intelligent process scheduling method based on an AI large model in a desktop cloud scenario

CN122614477BActive Publication Date: 2026-09-29CHANGSHA LINGWEI INNOVATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611105410.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-24
Publication Date
2026-09-29
Estimated Expiration
2046-07-24

AI Technical Summary

Technical Problem

现有方法未设计离线降级机制,系统可用性在异常场景下难以保障

Benefits of technology

[0031]本发明实施例提供的技术方案带来的有益效果至少包括:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122614477B_ABST
    Figure CN122614477B_ABST
Patent Text Reader

Abstract

The application discloses a kind of intelligent process scheduling methods based on AI big model in desktop cloud scene, belong to desktop cloud computing technology field, this method includes: in desktop cloud virtual machine customer operating system inside deployment collection agent, according to the feature data of each running process is collected in collection cycle;Characteristic data is constructed as structured prompt word and is submitted to AI big model, obtains the semantic label corresponding to each process;Scheduling execution continues to collect scheduling effect index, injects AI big model context update semantic label judgment description.The application also designs offline degradation mechanism, when AI big model inference service is unavailable, switch to local rule engine mode.The application realizes the process differentiation scheduling based on AI big model semantic understanding, host and customer machine double-path collaborative scheduling and scheduling effect closed-loop feedback optimization, improves desktop cloud foreground business application response performance and user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of desktop cloud computing technology, and in particular to an intelligent process scheduling method based on a large AI model in a desktop cloud scenario. Background Technology

[0002] Desktop cloud systems deploy user desktop environments as virtual machines on data center servers, allowing users to access these virtual desktops remotely via thin clients. In a desktop cloud environment, a single physical server hosts multiple virtual machines, each running multiple processes. Competition for CPU resources increases latency in foreground applications, degrading the user experience. Existing desktop cloud process scheduling typically employs the operating system's default fair scheduling policy, failing to differentiate between the resource requirements of foreground processes and background auxiliary processes. Critical interactions in foreground applications may experience lag due to background processes consuming CPU resources.

[0003] Existing desktop cloud process scheduling methods typically prioritize processes based on a single metric such as CPU utilization, automatically demoting high-utilization processes. However, some high-utilization processes may be critical background support processes (such as database services or build processes), and blindly demoting them could affect the integrity of business functions; conversely, some low-utilization processes may be non-critical idle processes, and maintaining their default priority would waste resources. Existing methods do not incorporate the business semantic attributes of processes for differentiated scheduling, resulting in scheduling decisions lacking understanding of the business context.

[0004] Existing desktop cloud process scheduling methods typically only perform scheduling within the virtual machine guest operating system or at the hypervisor level, without establishing a collaborative scheduling mechanism between the host and guest layers. When the guest operating system kernel promotes the priority of foreground processes, the hypervisor is unaware of this semantic information and still allocates physical CPU time slices according to the default weights. Conversely, when the hypervisor adjusts the virtual CPU scheduling weights, the guest operating system does not synchronously adjust process priorities. This creates information gaps in the transmission of scheduling instructions between the two layers, making it difficult to achieve coordinated optimization of scheduling performance.

[0005] Existing desktop cloud process scheduling methods typically employ static rules or lightweight machine learning models for classification, failing to leverage the semantic understanding capabilities of large language models to perform deep semantic analysis on textual features such as process names and startup paths. When encountering novel application processes or processes with non-standard naming conventions, static rule matching fails, lightweight models lack generalization ability, and process classification accuracy is limited.

[0006] Existing desktop cloud process scheduling methods typically lack a closed-loop feedback mechanism after the scheduling policy takes effect. When the response latency of the foreground application does not improve or even worsens after the scheduling weight is assigned, existing methods do not automatically analyze scheduling performance metrics and feed them back to the classification model for policy correction. A closed loop for continuous optimization is not established between the scheduling policy and business performance.

[0007] Existing desktop cloud process scheduling methods rely on cloud-based AI inference services for process semantic classification. When the network is interrupted or the inference service is unavailable, the scheduling function completely fails. Existing methods do not include an offline degradation mechanism, making it difficult to guarantee system availability in abnormal scenarios. Summary of the Invention

[0008] In view of this, embodiments of the present invention provide an intelligent process scheduling method based on a large AI model in a desktop cloud scenario.

[0009] On the one hand, an intelligent process scheduling method based on an AI large model is provided for desktop cloud scenarios, including the following steps:

[0010] Step 1: Process Feature Collection: Deploy a collection agent within the guest operating system of the desktop cloud virtual machine. The collection agent collects feature data of each running process within the guest operating system according to the collection period T. The feature data includes process identifier, parent process identifier, process name, CPU utilization, memory usage, and startup path.

[0011] Step 2, Semantic Classification: Construct structured prompt words from the feature data of each process and submit them to the AI ​​big model. Obtain the semantic labels corresponding to each process output by the AI ​​big model. The semantic labels are one of four categories: key front-end business processes, key back-end support processes, non-critical high-occupancy processes, and non-critical low-occupancy processes.

[0012] The classification results of semantic tags are cached for a period of M. If the cache expires, the classification is re-executed.

[0013] Step 3: Construction of the foreground business process set: Monitor the foreground activation window of the client operating system in real time. Take the process to which the foreground activation window belongs as the root node. Identify the associated processes that are related to the root node through three dimensions: process tree tracing, inter-process communication detection, and component dependency analysis. Merge the root node and each associated process into the foreground business process set.

[0014] Step 4: Calculate the scheduling weight W for each process using the following formula:

[0015]

[0016] In the formula, B is the semantic baseline score, calculated as follows:

[0017]

[0018] In the formula, X is the criticality coefficient, which is determined by the semantic tags output in step two. X=1 when the semantic tags contain key attributes and X=0 when they contain non-key attributes.

[0019] Y is the front-end attribute coefficient. Y=1 when the semantic tag contains front-end business attributes, otherwise Y=0.

[0020] Z represents the high occupancy coefficient. When a semantic tag contains a high occupancy attribute, Z=1; when it contains a low occupancy attribute, Z=0.

[0021] F is the front-end membership coefficient. F=1 when the process belongs to the set of front-end business processes, and F=0 otherwise.

[0022] G is the front-end gain, calculated using the following formula:

[0023]

[0024] In the formula, The upper limit of the scheduling weight is set to 100;

[0025] This represents the maximum value of the semantic benchmark score;

[0026] The lower limit of the scheduling weight is set to 10;

[0027] Step 5, Dual-path distribution: The scheduling weight W of each process is distributed in parallel via the host path and the guest path. The host path is to adjust the scheduling weight of the virtual machine's virtual central processing unit through the hypervisor interface, and the guest path is to adjust the process priority through the guest operating system kernel interface.

[0028] Step 6, Closed-loop feedback: After scheduling is executed, scheduling effect indicators are continuously collected. These indicators include changes in CPU utilization of each process, foreground application response latency, system average load, and virtual CPU ready time.

[0029] The scheduling performance metrics, along with the business scenario tags, are injected into the context of the AI ​​big model. The AI ​​big model then outputs an updated semantic tag judgment description, replacing the currently cached semantic tag judgment description.

[0030] Beneficial effects

[0031] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:

[0032] This invention deploys a data collection agent within the client operating system of a desktop cloud virtual machine. It collects characteristic data of each running process according to a collection cycle, including process identifier, parent process identifier, process name, CPU utilization, memory usage, and startup path. This characteristic data is then used to construct structured prompts and submitted to an AI model. The AI ​​model outputs semantic tags for each process, which are categorized into four types: critical front-end business processes, critical back-end support processes, non-critical high-utilization processes, and non-critical low-utilization processes. This invention solves the technical problem in existing technologies where priority adjustments are based on single indicators such as CPU utilization, and where differentiated scheduling is not combined with process business semantic attributes, leading to the blind degradation of critical back-end support processes. It achieves the technical effect of classifying process business semantics through deep semantic analysis of the AI ​​model and improving the business context understanding capability of scheduling decisions.

[0033] This invention monitors the active foreground window of the client operating system in real time, takes the process to which the active foreground window belongs as the root node, and identifies related processes associated with the root node through three dimensions: process tree tracing, inter-process communication detection, and component dependency analysis. The root node and all related processes are then incorporated into the foreground business process set. This solves the technical problems of existing technologies, such as the lack of a complete identification mechanism for foreground business processes and their associated supporting processes, and the incomplete construction of the foreground business set. It achieves the technical effect of completely identifying the foreground business process set through multi-dimensional association analysis and ensuring the integrity of foreground business functions.

[0034] This invention calculates the scheduling weight of each process using the following formula. The semantic baseline score is determined by the criticality coefficient, foreground coefficient, high occupancy coefficient, and foreground membership coefficient. The foreground membership coefficient is 1 if the process belongs to the set of foreground business processes, and 0 otherwise. A foreground gain is also set. This invention solves the technical problems in the prior art where the scheduling weight calculation does not distinguish the priority difference between foreground business processes and background processes, and the foreground business response guarantee is insufficient. It achieves the technical effect of improving the scheduling weight of foreground business processes and ensuring the response performance of foreground applications by using the foreground membership coefficient and the foreground gain.

[0035] This invention addresses the technical problem in existing technologies where scheduling is performed only within the virtual machine's guest operating system or at the virtual machine monitor level, without establishing a collaborative scheduling mechanism between the host and guest layers. By distributing the scheduling weights of each process in parallel via the host path and the guest path, the scheduling weights of the virtual machine's virtual central processing unit are adjusted through the virtual machine monitor interface, while the guest path adjusts process priorities through the guest operating system kernel interface. This achieves the technical effect of ensuring consistency between the host and guest layer scheduling strategies and improving scheduling efficiency through parallel distribution via dual paths. Attached Figure Description

[0036] Figure 1 A flowchart provided for an embodiment of this application. Detailed Implementation

[0037] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0038] like Figure 1 As shown, an intelligent process scheduling method based on an AI large model in a desktop cloud scenario includes the following steps:

[0039] Step 1: Process Feature Collection: Deploy a collection agent within the guest operating system of the desktop cloud virtual machine. The collection agent collects feature data of each running process within the guest operating system according to the collection period T. The feature data includes process identifier, parent process identifier, process name, CPU utilization, memory usage, and startup path.

[0040] Step 2, Semantic Classification: Construct structured prompt words from the feature data of each process and submit them to the AI ​​big model. Obtain the semantic labels corresponding to each process output by the AI ​​big model. The semantic labels are one of four categories: key front-end business processes, key back-end support processes, non-critical high-occupancy processes, and non-critical low-occupancy processes.

[0041] The classification results of semantic tags are cached for a period of M. If the cache expires, the classification is re-executed.

[0042] Step 3: Construction of the foreground business process set: Monitor the foreground activation window of the client operating system in real time. Take the process to which the foreground activation window belongs as the root node. Identify the associated processes that are related to the root node through three dimensions: process tree tracing, inter-process communication detection, and component dependency analysis. Merge the root node and each associated process into the foreground business process set.

[0043] Step 4: Calculate the scheduling weight W for each process using the following formula:

[0044]

[0045] In the formula, B is the semantic baseline score, calculated as follows:

[0046]

[0047] In the formula, X is the criticality coefficient, which is determined by the semantic tags output in step two. X=1 when the semantic tags contain key attributes and X=0 when they contain non-key attributes.

[0048] Y is the front-end attribute coefficient. Y=1 when the semantic tag contains front-end business attributes, otherwise Y=0.

[0049] Z represents the high occupancy coefficient. When a semantic tag contains a high occupancy attribute, Z=1; when it contains a low occupancy attribute, Z=0.

[0050] F is the front-end membership coefficient. F=1 when the process belongs to the set of front-end business processes, and F=0 otherwise.

[0051] G is the front-end gain, calculated using the following formula:

[0052]

[0053] In the formula, The upper limit of the scheduling weight is set to 100;

[0054] This represents the maximum value of the semantic benchmark score;

[0055] The lower limit of the scheduling weight is set to 10;

[0056] Step 5, Dual-path distribution: The scheduling weight W of each process is distributed in parallel via the host path and the guest path. The host path is to adjust the scheduling weight of the virtual machine's virtual central processing unit through the hypervisor interface, and the guest path is to adjust the process priority through the guest operating system kernel interface.

[0057] Step 6, Closed-loop feedback: After scheduling is executed, scheduling effect indicators are continuously collected. These indicators include changes in CPU utilization of each process, foreground application response latency, system average load, and virtual CPU ready time.

[0058] The scheduling performance metrics, along with the business scenario tags, are injected into the context of the AI ​​big model. The AI ​​big model then outputs an updated semantic tag judgment description, replacing the currently cached semantic tag judgment description.

[0059] It should be noted that the determination process of each coefficient in step four is as follows: the three coefficients X, Y, and Z are directly taken from the semantic label attributes output by the AI ​​large model;

[0060] The process of determining the baseline value 3, the grade difference 1, 2, -2 and the multiplier 10 is as follows: the scheduling weight is normalized according to the percentage system, and the non-critical low-occupancy process is used as the baseline grade, with X=0, Y=0, Z=0, and substituting them into the equation, we get B=30.

[0061] The grade difference is determined according to the direction of the attribute's influence on scheduling demand. The criticality coefficient increases by 1 grade for each grade, the front-end coefficient increases by 2 grades for each grade, and the high occupancy coefficient decreases by 2 grades for each grade.

[0062] Each gear has a score of 10. After the host path mapping, the difference in cpu.weight parameter between adjacent gears is 9999×10 / 90=1111. The scheduler can clearly distinguish between adjacent gears.

[0063] Substituting the values ​​X=1, Y=1, and Z=0 for the key front-end business processes, we get B=60, which means... =60, G=100-60=40;

[0064] For non-critical high-occupancy processes, X=0, Y=0, Z=1, substituting these values ​​gives B=10. Take the lowest value of the semantic benchmark score, which is 10;

[0065] =100 is the percentage-normalized upper bound of the scheduling weight;

[0066] Substitute the values ​​according to the calculation formula in step four: Key front-end business processes belonging to the front-end business process set W=60+40×1=100, key back-end support processes belonging to the front-end business process set W=40+40×1=80, key front-end business processes not belonging to the front-end business process set W=60+40×0=60, non-critical low-occupancy processes W=30, and non-critical high-occupancy processes W=10.

[0067] As an optional embodiment, the feature data in step one also includes virtual central processing unit (CPU) readiness time, which is read from the statistics interface exposed by the hypervisor.

[0068] The acquisition period T is determined by the following formula:

[0069]

[0070] In the formula, D is the average dwell time of the active foreground window, which is obtained by averaging the time intervals between two adjacent foreground window switches in the historical window switching log.

[0071] n is the number of samples during a single stay and n≥2;

[0072] When there is no historical window to switch logs, the collection period T is 5 seconds.

[0073] The cache validity period M mentioned in step two is determined by the following formula:

[0074]

[0075] In the formula, k is the cache effectiveness coefficient, which is determined by the following formula:

[0076]

[0077] In the formula, V is the average time interval between two consecutive fluctuations in the virtual central processing unit's ready time. The fluctuation refers to fluctuations whose amplitude exceeds its standard deviation. V is obtained by statistically analyzing the ready time records exposed by the virtual machine monitoring program.

[0078] When the statistical sample is insufficient, V is set to 30 seconds.

[0079] It should be noted that the virtual CPU readiness time reflects the degree of contention for the physical CPU by other virtual machines in a desktop cloud time-sharing scenario;

[0080] The value of n≥2 is based on the fact that at least two samplings must be completed during a single dwell period in order to identify changes in the process set during the dwell period;

[0081] floor is the floor function;

[0082] When V=30 seconds and T=5 seconds, k=floor(30 / 5)=6, M=k·T=30 seconds, and the cache validity period covers one complete fluctuation cycle of the system load;

[0083] The default value of 5 seconds for the collection cycle is based on the fact that feature collection, semantic classification, weight calculation and dual-path distribution are completed sequentially within a single collection cycle, and 5 seconds is the lower limit of the sum of the execution time of each step.

[0084] The default value of V is 30 seconds, which corresponds to 6 default collection cycles. This ensures that the cache validity period in the default state covers a complete load fluctuation cycle. After the system is running, V is replaced by the measured statistical value.

[0085] The classification results are reused within the cache validity period M, and fluctuations in a single collection do not change the semantic labels.

[0086] As an optional embodiment, the structured prompt words in step two include process feature fields, definitions of four types of semantic tags, output format constraints, and valid classification results within the previous cache validity period. The valid classification results are injected into the structured prompt words as reference examples.

[0087] The output of the AI ​​large model is first validated in terms of format. When the output content does not belong to the four types of semantic tags or the parsing fails, it is classified according to the fallback rule: processes with a central processing unit utilization rate not lower than the utilization threshold C are classified as non-critical high-utilization processes, and the remaining processes are classified as non-critical low-utilization processes. The AI ​​large model is then resubmitted for classification in the next collection cycle.

[0088] The occupancy threshold C is determined by the following formula:

[0089]

[0090] In the formula, μ is the mean of the CPU utilization rate of each process in the virtual machine, and σ is the standard deviation of the sample.

[0091] It should be noted that by injecting valid classification results from the previous cache validity period into the structured prompt words, the AI ​​large model uses the same classification criteria for the same process during adjacent collection weeks.

[0092] The calculation example of the occupancy threshold C is as follows: Suppose that the mean μ of the CPU occupancy rate of each process in the current virtual machine is 42 percent and the standard deviation σ is 18 percent, then C is 60 percent. Processes with a CPU occupancy rate of not less than 60 percent are classified as non-critical high-occupancy processes under the fallback rule.

[0093] As an optional embodiment, the specific process of process tree tracing in step three is as follows: starting from the process identifier of the root node, traverse upwards along the parent process identifier until the session root process, and at the same time traverse downwards along the child processes until the leaf processes.

[0094] All processes along the traversal path are merged into the candidate associated process set;

[0095] For Windows systems, process snapshots are used to enumerate each process and its parent-child relationships.

[0096] For Linux systems, the process and its parent-child relationship can be obtained by reading the status files of each process in the / proc directory.

[0097] It should be noted that the fields enumerated in the process snapshot of the Windows system are the process identifier field and the parent process identifier field, while the field read from the status file of the Linux system is the parent process identifier field.

[0098] The objects covered during upward traversal are the parent service processes that started the foreground application, and the objects covered during downward traversal are the child processes derived from the foreground application.

[0099] The traversal terminates when the session root process is reached or when a leaf process with no child processes is reached.

[0100] The processes along the traversal path have a direct derivation relationship with the active window in the foreground.

[0101] As an optional embodiment, the specific process of inter-process communication detection in step three is as follows: enumerate the named pipes, shared memory and socket handles held by each process in the candidate associated process set;

[0102] If a candidate process has a handle sharing or communication connection with the root node, the candidate process is determined to be an associated process.

[0103] The specific process of the component dependency analysis is as follows: enumerate the component object model service registration information and remote procedure call endpoint mapping;

[0104] If a candidate process provides a component object model component or remote procedure call service to the root node, then the candidate process is determined to be an associated process.

[0105] In step three, the window title of the foreground activation window and the process name of the root node are also submitted to the AI ​​big model to obtain the business scenario tags output by the AI ​​big model.

[0106] The number of related processes identified from three dimensions—process tree tracing, inter-process communication detection, and component dependency analysis—was counted respectively. , , Calculate the proportion of inter-process communication associations using the following formula:

[0107]

[0108] When the business scenario label output by the AI ​​model is video conferencing and If the number of cases is no more than 1 / 3, resubmit the window title, process name, and statistical counts for the three dimensions to the AI ​​large model for review of the business scenario tags.

[0109] It should be noted that the value of 1 / 3 is based on the expected proportion of each dimension when the three dimensions are evenly distributed.

[0110] Inter-process communication is active in video conferencing scenarios. Above this benchmark, When the value is not greater than this benchmark, there is a possibility of misjudgment in the business scenario label;

[0111] The calculation example is as follows: Let =2、 =5、 =1, then =5 / 8=0.625, which is greater than 1 / 3, consistent with the business scenario label of video conferencing, and does not require verification.

[0112] As an optional embodiment, the hypervisor interface in step five is one of the following: KVM's cgroups v2 interface, VMware ESXi's vSphere SDK interface, Hyper-V's resource group interface, and XEN's xl interface.

[0113] When using the KVM cgroups v2 interface, the scheduling weight W is linearly mapped to the cpu.weight parameter value P using the following formula:

[0114]

[0115] In step five, the guest operating system kernel interface for the Linux system uses the nice interface, mapping the scheduling weight W to the nice value Q using the following formula:

[0116]

[0117] The guest operating system kernel interface uses a process priority category interface for the Windows system.

[0118] After the dual-path scheduling is assigned, the actual effective cpu.weight parameter value and nice value of each process are read back. For processes whose read-back values ​​are inconsistent with the target values, the scheduling weight of the corresponding path is reassigned.

[0119] It should be noted that `round` is the rounding function;

[0120] The process of determining the mapping interval width of 9999 is as follows: The cgroups v2 interface specifies that the value range of the cpu.weight parameter is 1 to 10000, and the interval width is the difference between 10000 and 1, that is, 9999;

[0121] The process of determining the mapping interval width of 39 is as follows: the nice value ranges from -20 to 19, and the interval width is the difference between 19 and -20, which is 39;

[0122] In the P calculation formula, 1 represents the lower limit of the cpu.weight parameter, and in the Q calculation formula, 19 represents the upper limit of the nice value.

[0123] The readback verification is for cases where the virtual machine monitoring program interface loses adjustment requests during asynchronous execution in high-load desktop cloud scenarios. The readback value of the process that lost the adjustment request is still the value before the adjustment, which is inconsistent with the target value, and a resend is triggered accordingly.

[0124] As an optional embodiment, step five incorporates a debouncing mechanism: the scheduling weight change is only implemented when the scheduling weights calculated consecutively for K times are consistent, and the number of times K is consistent is determined by the following formula:

[0125]

[0126] At the same time, a hysteresis mechanism is set: when the absolute value of the difference between the recalculated scheduling weight and the currently effective weight is less than the minimum difference ΔW between gears, the currently effective weight remains unchanged.

[0127] ΔW is set to 10.

[0128] It should be noted that floor is the floor function;

[0129] The value of 2T in the denominator of the K calculation formula is based on the following: the anti-shake time window does not exceed half of the cache validity period, and at least one change judgment is completed within the validity period of the same classification result;

[0130] When M=30 seconds and T=5 seconds, K=floor(30 / 10)=3, the stabilization time window K·T is 15 seconds, which is less than the buffer validity period M=30 seconds;

[0131] The process of determining ΔW is as follows: the score corresponding to each level in the semantic benchmark score calculation formula is 10, the level difference between each attribute is 1 to 2 levels, and the minimum level difference between adjacent levels is 1 level. Therefore, ΔW = 10 × 1 = 10.

[0132] When the weight change is less than one minimum level difference, the relative level relationship of the process remains unchanged, and the original weight is maintained.

[0133] As an optional embodiment, the foreground application response delay L in step six is ​​the time difference between the foreground application receiving the input event and completing the interface refresh;

[0134] A sliding statistical window is constructed using response delay samples from the most recent 2K scheduling cycles;

[0135] When the response delay L within the sliding statistical window exceeds the target threshold for K consecutive scheduling cycles... At that time, the aggregated scheduling performance metrics and business scenario tags are injected into the context of the AI ​​big model, and the AI ​​big model outputs the updated semantic tag judgment description to replace the currently cached semantic tag judgment description.

[0136] The target threshold Determine by the following formula:

[0137]

[0138] In the formula, a is the mean of the front-end application response latency samples under the same business scenario label within the sliding statistical window;

[0139] r is the standard deviation of the sample.

[0140] It should be noted that the width of the sliding statistics window is 2K scheduling cycles. When K=3, it is 6 scheduling cycles, corresponding to a duration of 6T=30 seconds, which is consistent with the cache validity period M. The statistical samples and the currently effective classification results are in the same load range.

[0141] The value of coefficient 2 is determined as follows: when the response delay sample is approximately normally distributed, the mean plus 2 times the standard deviation is used as the upper bound of the sample distribution. Delays exceeding this upper bound are judged as abnormal delays and adjustments are triggered accordingly. Occasional jitters do not trigger adjustments.

[0142] As an optional embodiment, an offline degradation step is also included: sending probe requests to the AI ​​large model inference service according to the collection cycle;

[0143] When the response timeout reaches 2T, the inference service is deemed unavailable, and the system switches to local rule engine mode.

[0144] The local rule engine matches the process name and startup path in the most recent valid classification results, and the matched processes retain the corresponding semantic tags.

[0145] Processes that fail to hit are classified according to the fallback rule based on the occupancy threshold C, and scheduling weights are calculated and dual-path delivery is performed based on the classification results;

[0146] Once the inference service is available again, step two is re-executed to classify each process.

[0147] It should be noted that the timeout value of 2T is based on the following: if a probe request is sent in the current collection cycle and no response is received by the end of the next collection cycle, it is judged as a timeout, so 2 collection cycles are used.

[0148] The local rules engine classifies missing processes based on the same fallback rules as the online mode, and the classification criteria are the same before and after the downgrade.

[0149] As an optional implementation, when the available computing power of the node where the virtual machine is located is lower than the computing power threshold, the large AI model is replaced with a lightweight model with no more than 2 billion parameters and inference is performed locally on the virtual machine.

[0150] The structured prompts are truncated in descending order of process CPU utilization, retaining the feature data of the first m processes. The number of processes m is determined by the following formula:

[0151]

[0152] Decision delay in local inference Through actual measurement, when When the number of processes is greater than T / 2, the number of processes is truncated twice using the following formula:

[0153]

[0154] In the formula, H is the context window token capacity of the lightweight model, which is determined by the specification parameters of the lightweight model;

[0155] u represents the average number of tokens in a single process feature data, which is obtained by statistically analyzing the structured prompts in historical submissions.

[0156] m′ represents the number of processes after the second truncation.

[0157] It should be noted that the value of 2 billion parameters is based on the upper bound of the number of model parameters that actually satisfies the constraint that the decision delay is less than T / 2 at the edge nodes.

[0158] The value of T / 2 is based on the following: within a collection cycle, inference is completed in the first half of the cycle, and the second half of the cycle is used for scheduling weight calculation and dual-path distribution.

[0159] The calculation example for the second truncation is as follows: Let T = 5 seconds. =4 seconds, m=100, then m′=floor(100×5 / 8)=62, retain the 62 processes with the highest CPU utilization to submit local inference;

[0160] The computing power threshold is the minimum computing power required for the lightweight model to run locally on the node, and is determined by the product of the number of model parameters and the computing power overhead for inference per unit parameter.

[0161] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the scope and intent of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application is also intended to include such modifications and variations.

Claims

1. A method for intelligent process scheduling based on a large AI model in a desktop cloud scenario, characterized in that, Includes the following steps: Step 1: Process Feature Collection: Deploy a collection agent within the guest operating system of the desktop cloud virtual machine. The collection agent collects feature data of each running process within the guest operating system according to the collection period T. The feature data includes process identifier, parent process identifier, process name, CPU utilization, memory usage, and startup path. Step 2, Semantic Classification: Construct structured prompt words from the feature data of each process and submit them to the AI ​​big model. Obtain the semantic labels corresponding to each process output by the AI ​​big model. The semantic labels are one of four categories: key front-end business processes, key back-end support processes, non-critical high-occupancy processes, and non-critical low-occupancy processes. The classification results of semantic tags are cached for a period of M. If the cache expires, the classification is re-executed. Step 3: Construction of the foreground business process set: Monitor the foreground activation window of the client operating system in real time. Take the process to which the foreground activation window belongs as the root node. Identify the associated processes that are related to the root node through three dimensions: process tree tracing, inter-process communication detection, and component dependency analysis. Merge the root node and each associated process into the foreground business process set. Step 4: Calculate the scheduling weight W for each process using the following formula: In the formula, B is the semantic baseline score, calculated as follows: In the formula, X is the criticality coefficient, which is determined by the semantic tags output in step two. X=1 when the semantic tags contain key attributes and X=0 when they contain non-key attributes. Y is the front-end attribute coefficient. Y=1 when the semantic tag contains front-end business attributes, otherwise Y=0. Z represents the high occupancy coefficient. When a semantic tag contains a high occupancy attribute, Z=1; when it contains a low occupancy attribute, Z=0. F is the front-end membership coefficient. F=1 when the process belongs to the set of front-end business processes, and F=0 otherwise. G is the front-end gain, calculated using the following formula: In the formula, The upper limit of the scheduling weight is set to 100; This represents the maximum value of the semantic benchmark score; The lower limit of the scheduling weight is set to 10; Step 5, Dual-path distribution: The scheduling weight W of each process is distributed in parallel via the host path and the guest path. The host path is to adjust the scheduling weight of the virtual machine's virtual central processing unit through the hypervisor interface, and the guest path is to adjust the process priority through the guest operating system kernel interface. Step 6, Closed-loop feedback: After scheduling is executed, scheduling effect indicators are continuously collected. These indicators include changes in CPU utilization of each process, foreground application response latency, system average load, and virtual CPU ready time. The scheduling performance metrics, along with the business scenario tags, are injected into the context of the AI ​​big model. The AI ​​big model then outputs an updated semantic tag judgment description, replacing the currently cached semantic tag judgment description.

2. The intelligent process scheduling method based on a large AI model in a desktop cloud scenario according to claim 1, characterized in that, The feature data mentioned in step one also includes virtual CPU readiness time, which is read from the statistics interface exposed by the hypervisor. The acquisition period T is determined by the following formula: In the formula, D is the average dwell time of the active foreground window, which is obtained by averaging the time intervals between two adjacent foreground window switches in the historical window switching log. n is the number of samples during a single stay and n≥2; When there is no historical window to switch logs, the collection period T is 5 seconds. The cache validity period M mentioned in step two is determined by the following formula: In the formula, k is the cache effectiveness coefficient, which is determined by the following formula: In the formula, V is the average time interval between two consecutive fluctuations in the virtual central processing unit's ready time. The fluctuation refers to fluctuations whose amplitude exceeds its standard deviation. V is obtained by statistically analyzing the ready time records exposed by the virtual machine monitoring program. When the statistical sample is insufficient, V is set to 30 seconds.

3. The intelligent process scheduling method based on a large AI model in a desktop cloud scenario according to claim 2, characterized in that, The structured prompt words mentioned in step two include process feature fields, definitions of four types of semantic tags, output format constraints, and valid classification results within the previous cache validity period. The valid classification results are injected into the structured prompt words as reference examples. The output of the AI ​​large model is first validated in terms of format. When the output content does not belong to the four types of semantic tags or the parsing fails, it is classified according to the fallback rule: processes with a central processing unit utilization rate not lower than the utilization threshold C are classified as non-critical high-utilization processes, and the remaining processes are classified as non-critical low-utilization processes. The AI ​​large model is then resubmitted for classification in the next collection cycle. The occupancy threshold C is determined by the following formula: In the formula, μ is the mean of the CPU utilization rate of each process in the virtual machine, and σ is the standard deviation of the sample.

4. The intelligent process scheduling method based on a large AI model in a desktop cloud scenario according to claim 1, characterized in that, The specific process of process tree tracing in step three is as follows: starting from the process identifier of the root node, traverse upwards along the parent process identifier until the session root process, and at the same time traverse downwards along the child processes until the leaf processes. All processes along the traversal path are merged into the candidate associated process set; For Windows systems, process snapshots are used to enumerate each process and its parent-child relationships. For Linux systems, the process and its parent-child relationship can be obtained by reading the status files of each process in the / proc directory.

5. The intelligent process scheduling method based on a large AI model in a desktop cloud scenario according to claim 4, characterized in that, The specific process of inter-process communication detection in step three is as follows: enumerate the named pipes, shared memory and socket handles held by each process in the candidate associated process set; If a candidate process has a handle sharing or communication connection with the root node, the candidate process is determined to be an associated process. The specific process of the component dependency analysis is as follows: enumerate the component object model service registration information and remote procedure call endpoint mapping; If a candidate process provides a component object model component or remote procedure call service to the root node, then the candidate process is determined to be an associated process. In step three, the window title of the foreground activation window and the process name of the root node are also submitted to the AI ​​big model to obtain the business scenario tags output by the AI ​​big model. The number of related processes identified from three dimensions—process tree tracing, inter-process communication detection, and component dependency analysis—was counted respectively. , , Calculate the proportion of inter-process communication associations using the following formula: When the business scenario label output by the AI ​​model is video conferencing and If the number of cases is no more than 1 / 3, resubmit the window title, process name, and statistical counts for the three dimensions to the AI ​​large model for review of the business scenario tags.

6. The intelligent process scheduling method based on a large AI model in a desktop cloud scenario according to claim 1, characterized in that, The hypervisor interface mentioned in step five is one of the following: KVM's cgroups v2 interface, VMware ESXi's vSphere SDK interface, Hyper-V's resource group interface, and XEN's xl interface. When using the KVM cgroups v2 interface, the scheduling weight W is linearly mapped to the cpu.weight parameter value P using the following formula: In step five, the guest operating system kernel interface for the Linux system uses the nice interface, mapping the scheduling weight W to the nice value Q using the following formula: The guest operating system kernel interface uses a process priority category interface for the Windows system. After the dual-path scheduling is assigned, the actual effective cpu.weight parameter value and nice value of each process are read back. For processes whose read-back values ​​are inconsistent with the target values, the scheduling weight of the corresponding path is reassigned.

7. The intelligent process scheduling method based on an AI large model in a desktop cloud scenario according to claim 6, characterized in that, Step five incorporates a debouncing mechanism: the scheduling weight change is only implemented when the calculated scheduling weights are consistent for K consecutive times. The number of times the weights are consistent, K, is determined by the following formula: At the same time, a hysteresis mechanism is set: when the absolute value of the difference between the recalculated scheduling weight and the currently effective weight is less than the minimum difference ΔW between gears, the currently effective weight remains unchanged. ΔW is set to 10.

8. The intelligent process scheduling method based on a large AI model in a desktop cloud scenario according to claim 7, characterized in that, The foreground application response delay L mentioned in step six is ​​the time difference between the foreground application receiving the input event and completing the interface refresh. A sliding statistical window is constructed using response delay samples from the most recent 2K scheduling cycles; When the response delay L within the sliding statistical window exceeds the target threshold for K consecutive scheduling cycles... At that time, the aggregated scheduling performance metrics and business scenario tags are injected into the context of the AI ​​big model, and the AI ​​big model outputs the updated semantic tag judgment description to replace the currently cached semantic tag judgment description. The target threshold Determine by the following formula: In the formula, a is the mean of the front-end application response latency samples under the same business scenario label within the sliding statistical window; r is the standard deviation of the sample.

9. The intelligent process scheduling method based on a large AI model in a desktop cloud scenario according to claim 3, characterized in that, It also includes an offline degradation step: sending probe requests to the AI ​​large model inference service according to the collection cycle; When the response timeout reaches 2T, the inference service is deemed unavailable, and the system switches to local rule engine mode. The local rule engine matches the process name and startup path in the most recent valid classification results, and the matched processes retain the corresponding semantic tags. Processes that fail to hit are classified according to the fallback rule based on the occupancy threshold C, and scheduling weights are calculated and dual-path delivery is performed based on the classification results; Once the inference service is restored, step two is re-executed to classify each process.

10. The intelligent process scheduling method based on an AI large model in a desktop cloud scenario according to claim 3, characterized in that, When the available computing power of the node where the virtual machine is located is lower than the computing power threshold, the large AI model is replaced with a lightweight model with no more than 2 billion parameters and inference is performed locally on the virtual machine. The structured prompts are truncated in descending order of process CPU utilization, retaining the feature data of the first m processes. The number of processes m is determined by the following formula: Decision delay in local inference Through actual measurement, when When the number of processes is greater than T / 2, the number of processes is truncated twice using the following formula: In the formula, H is the context window token capacity of the lightweight model, which is determined by the specification parameters of the lightweight model; u represents the average number of tokens in a single process feature data, which is obtained by statistically analyzing the structured prompts in historical submissions. m′ represents the number of processes after the second truncation.

Citation Information

Patent Citations

  • AI model automatic deployment platform based on containerization technology

    CN120743426A

  • Systems, methods, kits, and apparatuses for using artificial intelligence for automation in value chain networks

    US20240144141A1