Self-adaptive resource scheduling and work order priority ranking method and system
By combining reinforcement learning models and large language models, the allocation and priority adjustment of work orders are dynamically optimized, solving the problem of uneven resource allocation in traditional work order processing and achieving more efficient IT environment problem solving and more stable resource utilization.
Patent Information
- Application Number
- CN202511168728.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-11-14
AI Technical Summary
Traditional IT environment issue ticket processing workflows rely on human experience and fixed rules, making it difficult to adapt to dynamic changes in the environment and team. They also fail to deeply understand the nature and urgency of the problem, leading to uneven resource allocation and impacting overall efficiency and employee satisfaction.
An adaptive resource scheduling and work order priority ranking method is adopted. By combining a reinforcement learning model with a large language model, the system dynamically perceives the environmental state and constructs a total reward function based on work order information, environmental information, and team personnel information to optimize work order allocation and priority adjustment, thereby achieving intelligent allocation and dynamic decision-making.
It improved work order processing efficiency, optimized resource utilization, ensured that critical issues were prioritized, achieved load balancing, reduced manual intervention costs, and enhanced environmental stability and employee satisfaction.
Smart Images

Figure CN120952628A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of work order system technology, and in particular to a method and system for adaptive resource scheduling and work order priority ranking. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] As enterprises grow in size and IT environments become increasingly complex, the number of environmental issues (including infrastructure, middleware and applications, permissions, etc.) is also growing exponentially. Enterprises typically use work order systems to record and track environmental issues.
[0004] Traditional IT environment issue ticket processing workflows typically rely on manual experience, fixed rules, or simple polling mechanisms for allocation and prioritization. This approach has the following drawbacks: rules struggle to adapt to dynamic changes in the environment and team; priorities, once set, are often difficult to adjust based on real-time conditions; relying solely on the limited fields of the ticket form makes it difficult to deeply understand the nature of the problem, its urgency, and its potential impact; it may lead to some teams or personnel being overloaded while others have idle resources, affecting overall efficiency and employee satisfaction; and the system cannot automatically learn and optimize allocation strategies from historical experience. The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0005] As enterprises grow in size and IT environments become increasingly complex, the number of environmental issues (including infrastructure, middleware and applications, permissions, etc.) is also growing exponentially. Enterprises typically use work order systems to record and track environmental issues.
[0006] Traditional IT issue ticket processing workflows typically rely on human experience, fixed rules, or simple polling mechanisms for allocation and prioritization. This approach has the following drawbacks: rules struggle to adapt to dynamic changes in the environment and team; priorities, once set, are often difficult to adjust based on real-time conditions; relying solely on the limited fields of the ticket form makes it difficult to deeply understand the nature of the problem, its urgency, and its potential impact; it may lead to some teams or personnel being overloaded while others have idle resources, affecting overall efficiency and employee satisfaction; and the system cannot automatically learn and optimize allocation strategies from historical experience. Summary of the Invention
[0007] To address the aforementioned issues, this invention proposes an adaptive resource scheduling and work order priority ranking method and system. This system dynamically senses the environmental state, understands the essence of work orders, and learns the optimal work order allocation and priority adjustment strategies, thereby improving problem-solving efficiency, optimizing resource utilization, and ensuring environmental stability.
[0008] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a method for adaptive resource scheduling and work order priority ranking, comprising: Using work order information, environmental information, and team personnel information as state observation variables, work order priority and allocation strategy as decision actions, and constructing a total reward function with SLA achievement reward, total duration reward, load balancing reward, allocation quality reward, and priority processing reward, a reinforcement learning model is trained. Receive pending work orders, determine the set of decision actions based on current environment information and team member information, and select the best priority for pending work orders and assign them to the best team or individual based on the total reward function.
[0009] As an alternative implementation method, the work order information includes work order features and original urgency extracted from the large language model, environmental information including service health and dependencies, and team / personnel information including skills, workload, and historical processing efficiency.
[0010] As an alternative implementation, the total reward function is defined as the weighted sum of the individual sub-reward functions at time step t, plus the cost of each action.
[0011] As an alternative implementation method, the SLA achievement reward is: ; in: It's a work order. ; It's a work order. The initial priority; It is the basic positive reward value for achieving SLA; This is the base negative reward value for failing to achieve the SLA; It's a work order. The actual resolution time; It's a work order. The SLA requires a certain duration; It is the additional penalty factor per unit time when the SLA is exceeded.
[0012] As an alternative implementation method, the total time reward is: ; in: It's a work order. The initial priority; This is the penalty coefficient for each unit in the market; It's a work order. The actual resolution time; or, ; in: It's the reward value for resolving the issue ahead of schedule; It is the penalty coefficient for delayed resolution; It's a work order. The expected resolution time.
[0013] As an alternative implementation method, the load balancing reward is: ; in: It is the standard deviation of the queue length; It is the penalty coefficient for load imbalance; or, ; in: It is the reward coefficient allocated to idle teams; The team before assignment queue length, It is to assign work orders to the team. Post-team The new queue length.
[0014] As an alternative implementation method, the quality bonus is allocated as follows:
[0015] in: It's a work order. With the team Skill matching degree; It is the reward value for skill matching; Penalty value for skill mismatch; If a reallocation occurs, then: ; in: It is the penalty value for redistribution.
[0016] As an alternative implementation method, the priority processing reward is as follows: ; in: It is a reward for the correct priority change; It is a penalty for changing the error priority.
[0017] Secondly, the present invention provides a system for adaptive resource scheduling and work order priority ranking, comprising: The training module is configured to use work order information, environmental information, and team personnel information as state observation variables, work order priority and allocation strategy as decision actions, and construct a total reward function with SLA achievement reward, total duration reward, load balancing reward, allocation quality reward and priority processing reward to train the reinforcement learning model. The processing module is configured to receive pending work orders, determine a set of decision actions based on current environment information and team personnel information, and select the best priority for pending work orders and assign them to the optimal team or individual based on the total reward function.
[0018] Thirdly, the present invention provides an electronic device including a memory and a processor, and computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the method described in the first aspect.
[0019] Fourthly, the present invention provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the method described in the first aspect.
[0020] Fifthly, the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the method described in the first aspect.
[0021] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention provides a reinforcement learning-based intelligent scheduling and prioritization method for IT environment work orders. By combining the semantic understanding capabilities of large language models with the dynamic decision-making optimization capabilities of reinforcement learning, it offers a new solution to the pain points of traditional work order processing. Through intelligent allocation, work orders are quickly directed to the most suitable handlers / teams, shortening the average resolution time. Dynamic priority adjustment ensures that critical issues are addressed first. Allocation is based on the real-time workload and skill matching of teams / personnel, achieving load balancing, avoiding resource bottlenecks and waste, improving the overall utilization of test environment resources, enhancing environmental stability, and reducing manual intervention costs. The system can continuously learn and optimize strategies from experience, adapting to environmental changes and new problem patterns without frequent manual rule adjustments. Faster problem response and resolution, and a more stable test environment, directly improve the satisfaction of developers and testers.
[0022] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0024] Figure 1 The flowchart is for the adaptive resource scheduling and work order priority sorting method provided in Embodiment 1 of the present invention. Detailed Implementation
[0025] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0026] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0027] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form as well. Furthermore, it should be understood that the terms “comprising” and “including”, and any variations thereof, are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0028] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0029] Example 1 This embodiment provides a method for adaptive resource scheduling and work order priority ranking, such as... Figure 1 As shown, it includes: Using work order information, environmental information, and team personnel information as state observation variables, work order priority and allocation strategy as decision actions, and constructing a total reward function with SLA achievement reward, total duration reward, load balancing reward, allocation quality reward, and priority processing reward, a reinforcement learning model is trained. Receive pending work orders, determine the set of decision actions based on current environment information and team member information, and select the best priority for pending work orders and assign them to the best team or individual based on the total reward function.
[0030] In this embodiment, a large model is used to perform in-depth analysis of the incoming work orders and extract key information. Then, a reinforcement learning agent (RL agent) takes this information, along with the current environmental state and team member information, as input. Through the learned strategies, it determines the optimal priority of the work orders and assigns them to the best team or individual. After executing the decision, the RL agent is rewarded or penalized based on the actual processing results (such as resolution time and user satisfaction), enabling it to continuously optimize its decision-making strategies. A human correction and feedback mechanism provides supervision and guidance to the system, accelerating the learning process and ensuring the rationality of the strategies.
[0031] The method of this embodiment will be described in detail below.
[0032] S1: Data Access and Preprocessing: Obtain raw work order information, environmental status information, and team personnel information from sources such as work order systems, monitoring systems, and IT operations and maintenance management systems (CMDB), and then clean, format, and integrate them.
[0033] S2: Large Model Semantic Understanding Work Orders: The cleaned work order text is input into the large model LLM module for semantic understanding, entity recognition (such as the services and components involved), intent recognition, sentiment analysis, and preliminary assessment of the severity and scope of the problem. The output is structured features and semantic vectors.
[0034] The applications of LLM and RL algorithms for large models include: Work order feature engineering: LLM is used to extract high-quality semantic features from unstructured text as part of the RL state.
[0035] Impact and urgency assessment: Leveraging the reasoning capabilities of LLM and combining knowledge from CMDB, we can more accurately assess the potential business impact and true urgency of work orders, assisting the RL Agent in making judgments.
[0036] Integration with RL Agent: The output of the LLM can be used as part of the state input of the RL Agent, or the LLM itself can be used as part of the RL Agent policy network (e.g., using the LLM to directly output the probability distribution of actions).
[0037] RL Algorithm Selection and Tuning: Select a suitable RL algorithm based on the characteristics of the problem (such as Deep Q-Networks for discrete action space).
[0038] S3: State Space Definition: The RL Agent combines the work order features output by the LLM with real-time information provided by the environmental state awareness module (such as the current load of each team, the skill matching degree of members, and the environmental health) to construct the current state, which specifically includes work order information, environmental information, and team personnel information. The work order information includes features extracted by LLM and the original urgency level; the environmental information includes service health and dependencies; and the team / personnel information includes skills, workload, and historical processing efficiency.
[0039] S4: Action Space Definition: Based on its internal policy network, the RL Agent selects an action for the current state, such as assigning it to one of N teams or adjusting it to one of M priorities. For complex scenarios, hierarchical decision-making or parameterized action space may be required. Specific decision actions include: assigning the work order to a specific team / individual, adjusting the work order priority, and suggesting triggering an automated repair process.
[0040] S5: Instruction Execution: The decision engine transforms the actions selected by the RL Agent into specific scheduling instructions, which are then sent to the corresponding processing teams / individuals or automation systems through the work order execution and notification module.
[0041] S6: Results Feedback and Reward Calculation: After the work order is completed (or a certain time point is reached), collect the processing results, such as resolution time, whether it is completed within the SLA, user feedback, etc.; design a reward function to calculate the reward value based on the processing results and preset optimization goals (such as minimizing the average resolution time, maximizing the SLA achievement rate, and balancing the team load).
[0042] The design of the reward function needs to accurately reflect the expected system behavior and business goals. For example: quickly resolving high-priority work orders: positive reward; work order timeout: negative reward; team load balancing: reward / penalty based on load standard deviation; successfully predicting and assigning to highly matched personnel: additional reward. A careful balance needs to be struck between short-term rewards and long-term goals to avoid unexpected negative behaviors.
[0043] Specifically, it includes: For work order scheduling and priority ranking systems, the total reward will be... Defined as the reward obtained at time step t (e.g., each time an allocation decision is made or every fixed time interval). Total reward At time step Defined as the weighted sum of all reward components, plus the optional cost per action, i.e.:
[0044] in: It is the weight of each sub-reward. It is a reward component At time step The generated reward value, It is the cost of each step of the action (e.g.) ).
[0045] Each sub-reward component is described in detail below.
[0046] (1) SLA (Service Level Agreement) achievement reward ( ), usually in work orders Calculated when closed: ; in: It's a work order. ; It's a work order. Business importance factor / initial priority; It is the basic positive reward value for achieving SLA; This is the base negative reward value for failing to achieve the SLA; It's a work order. The actual resolution time; It's a work order. The SLA requires a certain duration; It is the additional penalty factor per unit time when the SLA is exceeded.
[0047] (2) Solve the problem of time-based rewards / penalties There are two calculation methods, and in practical applications, either one can be chosen based on business objectives and scenario characteristics; specifically: Firstly, when a work order is closed, a penalty is applied based on the total duration: ; in: This is the penalty coefficient for each unit in the market.
[0048] In this approach, the longer the total time spent resolving a work order, the more points are deducted. The aim is to indiscriminately push the system to close all work orders as quickly as possible. This solution is simple and effective when the business objective is to minimize the average resolution time of all work orders, regardless of their complexity or type.
[0049] Secondly, when a work order is closed, the expected duration is considered: ; in: It's the reward value for resolving the issue ahead of schedule; It is the penalty coefficient for delayed resolution; It's a work order. The expected resolution time.
[0050] This approach pre-determines an "expected resolution time" for each work order. Completing the process faster than expected results in a reward, while delays incur penalties. This approach is more suitable when different types of work orders have clearly defined and distinct resolution time standards (similar to a more granular SLA). It encourages the system not only to be fast but also to complete tasks "within expectations," enabling it to handle tasks of varying complexity more flexibly.
[0051] During the process (more intensive rewards - open ticket penalties): for each time step All open work orders Total penalties:
[0052] in: It is the penalty coefficient for each open work order at each step. It is a fixed penalty coefficient, which can be understood as the points deducted for each work order of unit importance for each time step back. This represents the total penalty incurred at the current time step due to the backlog of work orders. It involves calculating all work orders that belong to the open work order set and then summing them up. It is a single open work order The penalty incurred at a time step. It's a work order. Importance / priority. This means that a backlog of an important work order will result in a greater ongoing penalty than a regular work order.
[0053] Specifically: This formula means that at every decision point in the system (time step t), a penalty is calculated. The penalty is equal to the sum of the penalties for all currently open (or closed) work orders. In simpler terms, this formula is like a "time-based billing meter"—points are continuously deducted as long as work orders remain unresolved. Furthermore, a backlog of important work orders will result in even greater point deductions.
[0054] This approach complements, but does not replace, the previous two schemes. The first two schemes offer one-time rewards or penalties. Only when a work order is closed will the system calculate the final, substantial reward or penalty based on the total time taken. This is called sparse reward because the reward signal is not continuous. The formula (open work order penalty) is a persistent or procedural penalty. Regardless of whether the work order is closed, as long as it exists in the to-do list, a small penalty is incurred at each time step. This is called dense reward. The ultimate goal of persistent penalty is to optimize the final result measured by the first two schemes. By continuously imposing small penalties during the process, the agent is guided not to procrastinate, and it is incentivized to take action to close the work order as soon as possible, thus achieving a better final score (i.e., the value calculated by scheme one or scheme two) when the work order is closed.
[0055] Together with time-based rewards / penalties, these measures constitute a measure of processing efficiency or time cost, but incentivize from different dimensions: 1. End-Result Oriented (the first two approaches): These define "what constitutes a good / bad resolution time," representing the ultimate goal. For example, shorter total time or faster than expected time is considered good. 2. Process Behavior Guidance (Open Work Order Penalty Formula): This provides continuous motivation to achieve the above goals. Without this process penalty, the AI might think, "As long as I can resolve the work order quickly in the end, a slight delay in the middle doesn't matter," but such delays themselves consume system resources and increase risk. The existence of this formula is to penalize delaying behavior, ensuring that the AI always feels a sense of urgency to reduce the backlog of work orders.
[0056] The open work order penalty formula is an important component of the "reward / penalty for resolution time" system. By imposing continuous, fine-grained penalties on the backlog of work orders, it serves as a behavioral guidance mechanism, supplementing the final reward and penalty that is only settled when the work order is closed. This more effectively incentivizes the AI system to process tasks quickly and continuously.
[0057] (3) Load balancing reward Each time a work order is assigned to the team Calculate later. Assume... The team before assignment Queue length / workload It is to assign work orders (workload) ) to the team Post-team New queue length ( , other ).
[0058] There are two calculation methods, which are different calculation methods designed to achieve the same goal (load balancing). In practical applications, one of them is usually chosen based on specific optimization preferences.
[0059] Specifically: Firstly, based on the standard deviation of the queue: ; in: It is the standard deviation of the queue length; It is the penalty coefficient for load imbalance.
[0060] This scheme aims for "absolute fairness" across the entire system. It measures the dispersion of work queue lengths across all teams. A smaller standard deviation indicates more similar workloads, a more balanced system, and lower penalties. Conversely, if any team's workload is significantly higher or lower than others, the standard deviation increases, leading to increased penalties. This incentivizes AI to make decisions that bring the entire system towards its optimal equilibrium. It serves as a more macro-level, holistic evaluation metric.
[0061] Secondly, rewards are allocated to teams with more free time: ; in: It is the reward coefficient allocated to teams with less free time.
[0062] This scheme encourages a "Robin Hood" style of single allocation. It doesn't concern itself with the absolute equilibrium of the entire system, but rather focuses on evaluating whether "this allocation was wise." Higher rewards are given when a new work order is assigned to a relatively idle team (i.e., the team's queue length is much shorter than the longest queue). The incentive is to make the most intuitive decision that directly reduces the system's maximum load in each allocation. It's a more specific, action-oriented evaluation metric.
[0063] If the desired final state of the system is that all teams maintain a strictly similar workload, then Option 1 is the better choice, focusing more on optimizing the final state. If the goal is to incentivize AI to learn a simple and effective allocation rule: "Always assign new tasks to the least busy person," then Option 2 is the better choice, focusing more on rewarding the current action.
[0064] (4) Allocating quality rewards Assigning work orders For the team Time calculation;
[0065] in: It's a work order. With the team (or the skill matching degree of the person handling the task) (0 to 1); It is the reward value for skill matching; Penalty value for skill mismatch.
[0066] If a reallocation occurs: ; in: It is the penalty value for redistribution.
[0067] (5) Prioritize reward processing When the agent's action is to adjust the work order When it comes to priority.
[0068] ; in: It is a reward for the correct priority change; It is a penalty for changing the error priority.
[0069] S7: Policy Learning and Update: After receiving the reward signal, the RL Agent updates its policy network according to the reinforcement learning algorithm it uses (such as Q-Learning, SARSA, PPO, A2C, etc.) so that it can make better decisions when it encounters similar states in the future.
[0070] S8: Data closed loop and continuous learning.
[0071] Cold start and exploration: In the early stages of the system, guidance based on rules or human experience is required, while the RL Agent conducts thorough exploration to collect experience.
[0072] Simulation environment: Construct a simulation environment to pre-train and test the RL agent, accelerating the learning process and reducing the risk of trial and error in the real environment.
[0073] Online / Offline Learning: Combining the advantages of online learning (real-time strategy updates) and offline learning (regular batch training using accumulated data).
[0074] S9: Manual Correction and Intervention: Test environment administrators review and adjust the system's allocation and priority ranking results through the manual correction and feedback platform. This correction information can directly correct current errors, or it can be used as high-quality training data or to adjust the reward function, helping the model converge to a better policy more quickly.
[0075] Among them, the manual feedback and correction mechanism has the following advantages: Human-machine collaboration: Human experts provide feedback, correct the model's erroneous decisions, provide high-quality demonstration data (imitation learning), or adjust the reward function; Explainability: Improve the explainability of the RL decision-making process, making it easier for administrators to understand and trust the system, and to intervene effectively.
[0076] This embodiment provides a reinforcement learning-based intelligent scheduling and prioritization method for IT environment work orders. By combining the semantic understanding capabilities of large language models with the dynamic decision-making optimization capabilities of reinforcement learning, it offers a new solution to address the pain points of traditional work order processing. Through intelligent allocation, work orders are quickly directed to the most suitable handlers / teams, shortening the average resolution time. Dynamically adjusting priorities ensures that critical issues are addressed first. Allocation is based on the real-time workload and skill matching of teams / personnel, achieving load balancing, avoiding resource bottlenecks and waste, improving the overall utilization of test environment resources, enhancing environmental stability, and reducing manual intervention costs. The system can continuously learn and optimize strategies from experience, adapting to environmental changes and new problem patterns without frequent manual rule adjustments. Faster problem response and resolution, and a more stable test environment, directly improve the satisfaction of developers and testers.
[0077] It should be noted that all data acquisition is conducted in accordance with laws and regulations and with user consent, and the data is used legally.
[0078] Example 2 This embodiment provides a system for adaptive resource scheduling and work order priority ranking, including: The training module is configured to use work order information, environmental information, and team personnel information as state observation variables, work order priority and allocation strategy as decision actions, and construct a total reward function with SLA achievement reward, total duration reward, load balancing reward, allocation quality reward and priority processing reward to train the reinforcement learning model. The processing module is configured to receive pending work orders, determine a set of decision actions based on current environment information and team personnel information, and select the best priority for pending work orders and assign them to the optimal team or individual based on the total reward function.
[0079] It should be noted that the above modules correspond to the steps described in Embodiment 1, and the examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in Embodiment 1. It should also be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.
[0080] In further embodiments, the following is also provided: An electronic device includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the method described in Embodiment 1. For brevity, further details are omitted here.
[0081] It should be understood that in this embodiment, the processor can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0082] Memory may include read-only memory and random access memory, and provides instructions and data to the processor. A portion of memory may also include non-volatile random access memory. For example, memory may also store information about the device type.
[0083] A computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the method described in Embodiment 1.
[0084] The method in Example 1 can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor. The software modules can reside in readily available storage media in the field, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, a detailed description is not provided here.
[0085] A computer program product includes a computer program that, when executed by a processor, implements the method described in Embodiment 1.
[0086] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which execute in a device on a target real or virtual processor to perform the processes / methods described above. Typically, program modules include routines, programs, libraries, objects, classes, components, data structures, etc., that perform specific tasks or implement specific abstract data types. In various embodiments, the functionality of program modules can be combined or divided among program modules as needed. The machine-executable instructions for the program modules can execute within a local or distributed device. In a distributed device, the program modules can reside in both local and remote storage media.
[0087] The computer program code used to implement the methods of the present invention may be written in one or more programming languages. This computer program code may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the computer or other programmable data processing device, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a computer, partially on a computer, as a stand-alone software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server.
[0088] In the context of this invention, computer program code or related data may be carried by any suitable carrier to enable a device, apparatus, or processor to perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, and the like. Examples of signals may include electrical, optical, radio, sound, or other forms of propagation signals, such as carrier waves, infrared signals, etc.
[0089] Those skilled in the art will recognize that the units and algorithm steps described in connection with the various examples of this embodiment can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this invention.
[0090] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention. To solve the above problems, the present invention proposes a method and system for adaptive resource scheduling and work order priority ranking, which dynamically senses the environmental state, understands the essence of work orders, and learns the optimal work order allocation and priority adjustment strategy, thereby improving problem-solving efficiency, optimizing resource utilization, and ensuring environmental stability.
[0091] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a method for adaptive resource scheduling and work order priority ranking, comprising: Using work order information, environmental information, and team personnel information as state observation variables, work order priority and allocation strategy as decision actions, and constructing a total reward function with SLA achievement reward, total duration reward, load balancing reward, allocation quality reward, and priority processing reward, a reinforcement learning model is trained. Receive pending work orders, determine the set of decision actions based on current environment information and team member information, and select the best priority for pending work orders and assign them to the best team or individual based on the total reward function.
[0092] As an alternative implementation method, the work order information includes work order features and original urgency extracted from the large language model, environmental information including service health and dependencies, and team / personnel information including skills, workload, and historical processing efficiency.
[0093] As an alternative implementation, the total reward function is defined as the weighted sum of the individual sub-reward functions at time step t, plus the cost of each action.
[0094] As an alternative implementation method, the SLA achievement reward is: ; in: It's a work order. ; It's a work order. The initial priority; It is the basic positive reward value for achieving SLA; This is the base negative reward value for failing to achieve the SLA; It's a work order. The actual resolution time; It's a work order. The SLA requires a certain duration; It is the additional penalty factor per unit time when the SLA is exceeded.
[0095] As an alternative implementation method, the total time reward is: ; in: It's a work order. The initial priority; This is the penalty coefficient for each unit in the market; It's a work order. The actual resolution time; or, ; in: It's the reward value for resolving the issue ahead of schedule; It is the penalty coefficient for delayed resolution; It's a work order. The expected resolution time.
[0096] As an alternative implementation method, the load balancing reward is: ; in: It is the standard deviation of the queue length; It is the penalty coefficient for load imbalance; or, ; in: It is the reward coefficient allocated to idle teams; The team before assignment queue length, It is to assign work orders to the team. Post-team The new queue length.
[0097] As an alternative implementation method, the quality bonus is allocated as follows:
[0098] in: It's a work order. With the team Skill matching degree; It is the reward value for skill matching; Penalty value for skill mismatch; If a reallocation occurs, then: ; in: It is the penalty value for redistribution.
[0099] As an alternative implementation method, the priority processing reward is as follows: ; in: It is a reward for the correct priority change; It is a penalty for changing the error priority.
[0100] Secondly, the present invention provides a system for adaptive resource scheduling and work order priority ranking, comprising: The training module is configured to use work order information, environmental information, and team personnel information as state observation variables, work order priority and allocation strategy as decision actions, and construct a total reward function with SLA achievement reward, total duration reward, load balancing reward, allocation quality reward and priority processing reward to train the reinforcement learning model. The processing module is configured to receive pending work orders, determine a set of decision actions based on current environment information and team personnel information, and select the best priority for pending work orders and assign them to the optimal team or individual based on the total reward function.
[0101] Thirdly, the present invention provides an electronic device including a memory and a processor, and computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the method described in the first aspect.
[0102] Fourthly, the present invention provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the method described in the first aspect.
[0103] Fifthly, the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the method described in the first aspect.
[0104] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention provides a reinforcement learning-based intelligent scheduling and prioritization method for IT environment work orders. By combining the semantic understanding capabilities of large language models with the dynamic decision-making optimization capabilities of reinforcement learning, it offers a new solution to the pain points of traditional work order processing. Through intelligent allocation, work orders are quickly directed to the most suitable handlers / teams, shortening the average resolution time. Dynamic priority adjustment ensures that critical issues are addressed first. Allocation is based on the real-time workload and skill matching of teams / personnel, achieving load balancing, avoiding resource bottlenecks and waste, improving the overall utilization of test environment resources, enhancing environmental stability, and reducing manual intervention costs. The system can continuously learn and optimize strategies from experience, adapting to environmental changes and new problem patterns without frequent manual rule adjustments. Faster problem response and resolution, and a more stable test environment, directly improve the satisfaction of developers and testers.
[0105] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention.
Claims
1. A method for adaptive resource scheduling and work order priority ranking, characterized in that, include: Using work order information, environmental information, and team personnel information as state observation variables, work order priority and allocation strategy as decision actions, and constructing a total reward function with SLA achievement reward, total duration reward, load balancing reward, allocation quality reward, and priority processing reward, a reinforcement learning model is trained. Receive pending work orders, determine the set of decision actions based on current environment information and team member information, and select the best priority for pending work orders and assign them to the best team or individual based on the total reward function.
2. The method for adaptive resource scheduling and work order priority ranking as described in claim 1, characterized in that, The work order information includes work order features and original urgency extracted from the large language model; the environmental information includes service health and dependencies; and the team / personnel information includes skills, workload, and historical processing efficiency. The total reward function is defined as the weighted sum of the sub-reward functions at time step t, plus the cost of each action.
3. The method for adaptive resource scheduling and work order priority ranking as described in claim 1, characterized in that, The reward for achieving SLA is: ; in: It's a work order. ; It's a work order. Initial priority; It is the basic positive reward value for achieving SLA; This is the base negative reward value for failing to achieve the SLA; It's a work order. The actual resolution time; It's a work order. The SLA requires a certain duration; It is the additional penalty factor per unit time when the SLA is exceeded.
4. The method for adaptive resource scheduling and work order priority ranking as described in claim 1, characterized in that, The total time reward is: ; in: It's a work order. Initial priority; This is the penalty coefficient for each unit in the market; It's a work order. The actual resolution time; or, ; in: It's the reward value for resolving the issue ahead of schedule; It is the penalty coefficient for delayed resolution; It's a work order. The expected resolution time.
5. The method for adaptive resource scheduling and work order priority ranking as described in claim 1, characterized in that, Load balancing rewards are: ; in: It is the standard deviation of the queue length; It is the penalty coefficient for load imbalance; or, ; in: It is the reward coefficient allocated to idle teams; The team before assignment queue length, It is to assign work orders to the team. Post-team The new queue length.
6. The method for adaptive resource scheduling and work order priority ranking as described in claim 1, characterized in that, The quality bonus is allocated as follows: in: It's a work order. With the team Skill matching degree; It is the reward value for skill matching; Penalty value for skill mismatch; If a reallocation occurs, then: ; in: It is the penalty value for redistribution. Priority processing rewards are: ; in: It is a reward for the correct priority change; It is a penalty for changing the error priority.
7. A system for adaptive resource scheduling and work order priority ranking, characterized in that, include: The training module is configured to use work order information, environmental information, and team personnel information as state observation variables, work order priority and allocation strategy as decision actions, and construct a total reward function with SLA achievement reward, total duration reward, load balancing reward, allocation quality reward and priority processing reward to train the reinforcement learning model. The processing module is configured to receive pending work orders, determine a set of decision actions based on current environment information and team personnel information, and select the best priority for pending work orders and assign them to the optimal team or individual based on the total reward function.
8. An electronic device, characterized in that, It includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, which, when executed by the processor, perform the method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, Used to store computer instructions, which, when executed by a processor, perform the method described in any one of claims 1-6.
10. A computer program product, characterized in that, Includes a computer program, which, when executed by a processor, implements the method described in any one of claims 1-6.