Resource allocation method and system based on product technical support indexes

By combining game theory models and hybrid optimization methods with reinforcement learning, the problem of reliance on human experience in resource allocation is solved, achieving dynamic optimization and fairness in resource allocation, and improving the objectivity and utility maximization of allocation results.

CN120875407APending Publication Date: 2025-10-31CHINESE PEOPLES LIBERATION ARMY UNIT 32181
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511017790.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing resource allocation methods rely excessively on human experience, lack objectivity and fairness, and result in allocation outcomes being heavily influenced by subjective preferences, while also lacking mathematical modeling capabilities.

Method used

A game theory model combined with a hybrid optimization method is adopted to generate an initial resource allocation scheme through task utility function and constraints, and then to optimize resource allocation by reinforcement learning and dynamic priority adjustment to generate the final scheme.

Benefits of technology

It achieves dynamic optimization and fair allocation of resources, with flexibility and adaptability, improving the objectivity of allocation results and maximizing utility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120875407A_ABST
    Figure CN120875407A_ABST
Patent Text Reader

Abstract

The invention discloses a resource allocation method and system based on a product technology guarantee index, and relates to the technical field of maintenance guarantee, and the method comprises the following steps: 1, based on a task demand set, an available guarantee resource set, a priority weight set, a task time factor and an initial strategy set, defining a task utility function, and obtaining a game model; 2, solving a Nash equilibrium point of resource allocation based on a game model through a hybrid optimization method, and generating an initial resource allocation scheme; step 3, optimizing the initial allocation scheme through a reinforcement learning feedback mechanism and a dynamic priority adjustment method to generate an optimized resource allocation scheme; and step 4, evaluating the optimized resource allocation scheme based on the task completion rate and the resource utilization rate to obtain a final resource allocation scheme. According to the invention, dynamic optimization and fair distribution of technical support indexes are realized based on resource competition and cooperative distribution of the game model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of maintenance and support technology, and more specifically to a resource allocation method and system based on product technical support indicators. Background Technology

[0002] Currently, all resource allocation work adopts a subjective qualitative approach, relying on hierarchical evaluation models (such as the Analytic Hierarchy Process) or expert experience scoring to achieve comprehensive resource allocation goals. However, this approach has the following significant drawbacks: (1) the allocation process relies excessively on human experience and judgment, leading to a significant influence of subjective preferences on the allocation results; (2) it lacks the mathematical modeling ability to ensure allocation fairness and maximize utility, resulting in a serious lack of objectivity and verifiability in the allocation results. Therefore, how to overcome these drawbacks has become an urgent problem for those skilled in the art. Summary of the Invention

[0003] In view of this, the present invention provides a resource allocation method and system based on product technical assurance indicators, which overcomes the above-mentioned defects.

[0004] To achieve the above objectives, the present invention adopts the following technical solution:

[0005] A resource allocation method based on product technical assurance indicators includes the following steps:

[0006] Step 1: Based on the task requirement set, available guaranteed resource set, priority weight set, task time factor, and initial strategy set, define the task utility function to obtain the game model; the game model is configured with constraints, namely resource upper limit constraint, minimum task requirement constraint, task time constraint, and resource sharing mechanism constraint.

[0007] Step 2: Solve for the Nash equilibrium point of resource allocation based on the game model using a hybrid optimization method to generate an initial resource allocation scheme;

[0008] Step 3: Optimize the initial allocation scheme by using a reinforcement learning feedback mechanism and a dynamic priority adjustment method to generate an optimized resource allocation scheme;

[0009] Step 4: Evaluate the optimized resource allocation scheme based on the task completion rate and resource utilization rate. If the preset threshold is met, the final resource allocation scheme is obtained; if the preset threshold is not met, proceed to step 2.

[0010] Optionally, the expression for the utility function is:

[0011]

[0012] In the formula, αij For resource r j For task d i Contribution weight; β i For task d i Resource consumption coefficient; x ij Assigned to task d i resources r j ;γ i For task d i Time sensitivity coefficient; T i This represents the task completion time factor.

[0013] Optionally, the expression for the resource sharing mechanism constraint is:

[0014]

[0015] In the formula, x ij Assigned to task d i resources r j ; For resource r j Maximum available quantity; To remove task d i In addition, the shared available quantity for the k-th task.

[0016] Optionally, the hybrid optimization method is a staged optimization algorithm. In the first stage, a genetic algorithm is used to screen the initial strategy based on fitness and output the screening scheme. In the second stage, reinforcement learning is used to dynamically optimize the screening scheme in the first stage and generate an initial resource allocation scheme.

[0017] Optionally, the optimization steps in the first stage are as follows:

[0018] Step 211: Initialize the population. Each individual in the population represents a resource allocation scheme.

[0019] Step 212: Calculate the fitness function, which is constructed based on the task utility function and priority weights, and is used to evaluate the resource allocation scheme;

[0020] Step 213: Perform crossover and mutation operations, using a single-point crossover and mutation operator to generate new individuals;

[0021] Step 214: Select K optimal individuals and jump to step 212 for iterative optimization until the iteration stops, and output the screening scheme.

[0022] Optionally, the optimization steps in the second stage are as follows:

[0023] Define the state, action space, and reward function;

[0024] The resource allocation scheme of the screening scheme is updated using the strategy gradient optimization method until it converges to the Nash equilibrium point, thus obtaining the initial resource allocation scheme.

[0025] Optionally, the step of obtaining the optimized resource allocation scheme is as follows:

[0026] Based on task time constraints, the task priority weights are dynamically adjusted using a reward function to obtain optimized task weights.

[0027] Check whether the resources in the resource allocation scheme meet the minimum requirement constraint of the task. If they do, output the optimized resource allocation scheme; otherwise, trigger the resource reallocation mechanism to allocate resources based on the optimized task weight.

[0028] A resource allocation system based on product technical assurance indicators, comprising:

[0029] The model building module is used to define a task utility function and obtain a game model based on a set of task requirements, a set of available guaranteed resources, a set of priority weights, a set of task time factors, and a set of initial strategies. The game model is configured with constraints, including resource ceiling constraints, minimum task requirement constraints, task time constraints, and resource sharing mechanism constraints.

[0030] The resource allocation module is used to solve the Nash equilibrium point of resource allocation based on a game model using a hybrid optimization method, and to generate an initial resource allocation scheme.

[0031] The scheme optimization module is used to optimize the initial allocation scheme through a reinforcement learning feedback mechanism and a dynamic priority adjustment method to generate an optimized resource allocation scheme;

[0032] The scheme evaluation module is used to evaluate the optimized resource allocation scheme based on the task completion rate and resource utilization rate. If the preset threshold is met, the final resource allocation scheme is obtained; if the preset threshold is not met, the initial resource allocation scheme is regenerated.

[0033] As can be seen from the above technical solutions, compared with the prior art, the present invention provides a resource allocation method and system based on product technical assurance indicators. The resource competition and cooperation allocation based on the game model realizes the dynamic optimization and fair allocation of technical assurance indicators, while having flexibility and adaptability. Attached Figure Description

[0034] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0035] Figure 1 This is a schematic diagram of the method flow provided by the present invention. Detailed Implementation

[0036] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0037] One embodiment of the present invention discloses a resource allocation method based on product technical assurance indicators, such as... Figure 1 As shown, the specific steps are as follows:

[0038] Step 1: Based on the task requirement set, available guaranteed resource set, priority weight set, task time factor, and initial strategy set, define the task utility function to obtain the game model; the game model is configured with constraints, namely resource upper limit constraint, minimum task requirement constraint, task time constraint, and resource sharing mechanism constraint.

[0039] Step 2: Solve for the Nash equilibrium point of resource allocation based on the game model using a hybrid optimization method to generate an initial resource allocation scheme;

[0040] Step 3: Optimize the initial allocation scheme by using a reinforcement learning feedback mechanism and a dynamic priority adjustment method to generate an optimized resource allocation scheme;

[0041] Step 4: Evaluate the optimized resource allocation scheme based on the task completion rate and resource utilization rate. If the preset threshold is met, the final resource allocation scheme is obtained; if the preset threshold is not met, proceed to step 2.

[0042] Furthermore, in step 1, the game theory model is defined as follows:

[0043] Its input information includes various technical support requirements (mainly the task requirement set D = {d1, d2, ..., d...}). i ,…,d n}), where d iThis represents the i-th task; the available resource set (resource set R = {r1, r2, ..., r...}) j ,…,r m}), where r j Represents the j-th resource; the priority weight set for each task (weight set W = {w1, w2, ..., w...}). i ,…,w n}), where w i It is task d i The initial priority weight will be dynamically adjusted based on the task completion progress and time factors; the task time factor (which is used to represent the timeliness of the task, defined as T={T1,T2,…,T…) i ,…,T n}), where T i It is task d i The estimated completion time influences the priority of resource allocation for the task. During execution, various technical support requirements will be considered. i (i.e., the task) is considered a participant in a game theory model; task d i The set of strategies represents its possible resource allocation schemes, which can be further represented as S. i ={x i1 ,x i2 ,…,x ij ,…,x im}, where x ij This indicates that the task d is assigned to it. i resources r j Based on the input information, define a target for task d. i Utility function:

[0044]

[0045] In the formula, α ij It is a resource r j For task d i Contribution weight, β i It is task d i The resource consumption coefficient, γ i It is task d i Time sensitivity coefficient, T i T is the task completion time factor. If the task is close to the deadline, then T... i Increase the resource allocation priority of the task.

[0046] To improve resource utilization, this model allows some tasks to share resources, i.e., task d i Can accept a shared resource r j Minimum allocation threshold

[0047]

[0048] If task d i Unable to obtain minimum allocation threshold Then its weight w will be adjusted first. i To increase the priority of the next round of allocation. Task weight w i Adjustments are made after each round of resource allocation so that unfinished critical tasks can receive higher resource allocation priority in the next round. The expression is as follows:

[0049]

[0050] In the formula, The adjusted task weights, The weights from the previous round are given, α is an adjustment factor, and T is the weight from the previous round. i This is the task completion time factor; This ensures that the shorter the remaining time of a task, the faster its weight increases, thus guaranteeing that urgent tasks receive resources first.

[0051] The constraints include:

[0052] Total resource limit (i.e., resource limit constraint): In the formula, n is the total number of tasks; This represents the maximum available resource; it indicates the resource r allocated to all tasks. j The total amount cannot exceed the maximum available amount of the resource.

[0053] Minimum task requirements constraints:

[0054]

[0055] In the formula, For task d i For resource r j Basic requirements; w k w represents the weight of the k-th task. i Let i be the weight of the i-th task; This represents the minimum acceptable allocation amount for the task when resources are shared. Thus, the task weight w... i Adjustments affecting resource requirements, urgent tasks (when w) i The minimum requirements remain unchanged (when the requirements are high), and lower priority tasks can accept fewer resources.

[0056] Task time constraints:

[0057]

[0058] In the formula, It is task d iThe allowable latest completion time is designed to prioritize resources for tasks with shorter completion times, thus preventing delays for high-priority tasks and allowing task time constraints to directly impact resource allocation strategies.

[0059] Resource sharing mechanism constraints:

[0060] In the formula, For resource r j The total amount; To remove task d i In addition, the amount of shared resources for the k-th task; representing task d i The allocated resources r j It cannot exceed the total resources minus the portion shared by other tasks.

[0061] In one embodiment, the hybrid optimization method is a staged optimization algorithm. In the first stage, a genetic algorithm is used to screen initial strategies based on fitness and output a screening scheme. In the second stage, reinforcement learning is used to dynamically optimize the screening scheme from the first stage to generate an initial resource allocation scheme.

[0062] Furthermore, this step is used to find the optimal equilibrium point for task allocation, that is, to ensure that, under given constraints, no task can improve its own utility by unilaterally adjusting its resource allocation strategy, thereby achieving a stable resource allocation scheme. The input information for this step includes: the utility function U of each task. i (x i ,x -i The set of strategies S for the task i ={x i1 ,x i2 ,……x im Constraints and dynamic adjustment mechanisms for task priorities.

[0063] To improve solution efficiency and optimization performance, a hybrid optimization method combining genetic algorithm (GA) and reinforcement learning (RL) is adopted. This method comprises two stages: GA initialization of the global solution and RL dynamic optimization.

[0064] In one embodiment, the optimization steps in the first stage are as follows:

[0065] Step 211: Initialize the population. Each individual in the population represents a resource allocation scheme.

[0066] Step 212: Calculate the fitness function, which is constructed based on the task utility function and priority weights, and is used to evaluate the resource allocation scheme;

[0067] Step 213: Perform crossover and mutation operations, using a single-point crossover and mutation operator to generate new individuals;

[0068] Step 214: Select K optimal individuals and jump to step 212 for iterative optimization until the iteration stops, and output the screening scheme.

[0069] Furthermore, the specific steps for initializing the global solution using a genetic algorithm (GA) are as follows:

[0070] First, initialize the population: generate multiple possible resource allocation schemes X = {x1, x2, ..., x}. n}, where x i ={x i1 ,x i2 ,……x im Meanwhile, an initial population size P is set, which sequentially covers the search space.

[0071] Secondly, the fitness function is calculated: task utility maximization is used as the fitness evaluation criterion, and the top K optimal solutions are selected by sorting. Its expression is:

[0072]

[0073] Secondly, crossover mutation: select individuals with higher probabilities for crossover to generate new candidate solutions x'. i Meanwhile, x is adjusted through the mutation operator. ij To increase search diversity.

[0074] Finally, select the best individuals: select the K best individuals to proceed to the next round of optimization.

[0075] In one embodiment, a state, an action space, and a reward function are defined;

[0076] The resource allocation scheme of the screening scheme is updated using the strategy gradient optimization method until it converges to the Nash equilibrium point, thus obtaining the initial resource allocation scheme.

[0077] Furthermore, reinforcement learning (RL) dynamic optimization has the following characteristics:

[0078] State representation: Define the state s in reinforcement learning t This includes the current resource allocation x for the task. i ={x i1 ,x i2 ,……x im Remaining time T for the task i and task weight w i .

[0079] Action space: Each task can choose to increase resources (i.e., increase x). ij ), reduce resources (i.e. reduce x) ij ), and maintain the current allocation (keep x) i constant).

[0080] Reward function: Maximizes total task utility and reduces resource waste; its expression is:

[0081] R t =U tatal (t)-U tatal (t-1);

[0082] After completing the above definition, the policy gradient optimization method is used to train the agent to continuously optimize the task allocation scheme, iterating until the reward value converges, and the final optimized solution is obtained.

[0083] In one embodiment, the step of obtaining the optimized resource allocation scheme is as follows:

[0084] Based on task time constraints, the task priority weights are dynamically adjusted using a reward function to obtain optimized task weights.

[0085] Check whether the resources in the resource allocation scheme meet the minimum requirement constraint of the task. If they do, output the optimized resource allocation scheme; otherwise, trigger the resource reallocation mechanism to allocate resources based on the optimized task weight.

[0086] Furthermore, after initial optimization of resource allocation, to further improve task execution performance, a reinforcement learning (RL) feedback mechanism and dynamic priority adjustment method are used to analyze and adjust the resource allocation scheme, thereby optimizing overall system performance. Input information includes: the optimized resource allocation scheme X. * ={x1 * x2 * ,…,x n *}, where x i * For task d i Resource allocation vector; reinforcement learning reward value Used to evaluate task d i Resource allocation effectiveness; task priority weights Dynamically adjusted by the reinforcement learning training process; task completion time constraint (i.e., task completion time factor) T i To ensure the priority execution of time-sensitive tasks; remaining resources Used to check for underutilized resources; including:

[0087] (1) Dynamic adjustment of task priority: The priority of tasks will be dynamically updated based on the historical reward values ​​of reinforcement learning. The expression of the update formula is:

[0088]

[0089] In the formula, For task d i The reinforcement learning reward in the t-th round of resource allocation; α is the adjustment coefficient; T i Set time constraints for task completion; This represents the task priority weight.

[0090] (2) Resource reallocation mechanism: If some tasks still do not meet the minimum resource requirements, the resource reallocation mechanism is executed, and its expression is:

[0091]

[0092] In the formula, For task d i For resource r j Basic needs; Weight the priority items to ensure that high-priority tasks receive more resources; For shared optimization items; Assigned to task d i Maximum resource quantity, n 共享 This is the number of tasks sharing the resource.

[0093] After completing the above adjustment process, the optimized task weights will be generated. It is used for resource allocation decisions in the next round, and also includes the adjusted resource allocation scheme X' = {x1', x2', ..., x n If the optimization request for a task is not met, or if some tasks still do not meet the minimum requirements, then proceed to the next round of optimization.

[0094] In one embodiment, the step requires calculating the task completion rate (Task_Success_Rate) and resource utilization rate (Resource_Utilization) based on the adjusted resource allocation scheme X', and then comparing them with predetermined indicators to complete the evaluation and feedback, specifically:

[0095] First, task completion rate analysis: This involves calculating the percentage of tasks that meet the minimum requirements, as this metric reflects the effectiveness of the resource allocation plan; among which... Indicates task d i If the minimum resource requirements are met, the value is 1; otherwise, it is 0. The calculation formula is as follows:

[0096]

[0097] Next, resource utilization analysis is performed: the optimized resource utilization is calculated. This indicator reflects the efficiency of resource use and ensures that the allocation plan does not lead to resource waste. The calculation formula is as follows:

[0098]

[0099] Finally, the final allocation scheme X is determined by comparing it with the predetermined task completion rate and resource utilization rate. final If the desired result is still not achieved, adjust the parameters and optimize again.

[0100] This embodiment further illustrates the use of aircraft maintenance tasks: An aviation maintenance department needs to allocate resources (maintenance personnel, equipment, and spare parts) for the maintenance of three aircraft to meet task priority and time requirements. The tasks include engine maintenance (d1) for aircraft A, navigation system maintenance (d2) for aircraft B, and fuselage inspection (d3) for aircraft C. The resource allocation method of this embodiment is used to allocate resources reasonably, maximizing task utility and ensuring fair resource allocation. The task requirement set is represented as: D = {d1, d2, d3}, n = 3; the available resource set can be further defined as: R = {r1, r2, r3} (maintenance personnel, equipment, spare parts), total amount; m = 3; Priority weight set: W = {w1 = 0.5, w2 = 0.3, w3 = 0.2}, α = 0.1; Task time factor: T = {T1 = 5, T2 = 8, T3 = 10}. The utility function parameters are defined as: B = {β1 = 0.2, β2 = 0.15, β3 = 0.1}, γ = (γ1 = 0.1, γ2 = 0.08, γ3 = 0.05), α ij The resulting contribution coefficient matrix A is as follows:

[0101]

[0102] The minimum requirement and sharing threshold are defined as follows:

[0103] Step 1, Define the game theory model: Based on the above information, define the following task utility function:

[0104]

[0105] The constraints are set as follows:

[0106] Total resource limit:

[0107] Minimum task requirements constraints:

[0108] Task time constraints:

[0109] Resource sharing mechanism constraints:

[0110] Step 2, solve for the Nash equilibrium, specifically:

[0111] Phase Optimization:

[0112] 1) Set the initial population size P = 10, and an example of one individual in the population is as follows:

[0113]

[0114] 2) Calculate the fitness function: Calculate the fitness of each individual in the population. The calculation process of the fitness for total utility is illustrated below using a subset of individuals as an example:

[0115]

[0116] U1=2log(1+3.5)+1.5log(1+1.5)+1log(1+2.5)-0.2(3.5+1.5+2.5)-0.1×5=2×1.544+1.5×0.916+1×1.299-1.5-0.5=3.088+1.374+1.299-2=3.761;

[0117] U2=1.8log(1+3.5)+1.2log(1+0.5)+0.8log(1+2.5)-0.15(3.5+0.5+2.5)-0.08× 8=1.8×1.544+1.2×0.405+0.8×1.299-1-0.64=2.779+0.486+1.039-1.64=2.664;

[0118] U3=1×log(1+2.5)+0.9log(1+1)+0.7log(1+2)-0.1(2.5+1+2)-0.05×10=1.299+0.624+0.485-1.05=1.358;

[0119] U total =0.5×3.761+0.3×2.664+0.2×1.358=2.811;

[0120] 3) Use a single-point crossover mutation operator to perform crossover mutation;

[0121] 4) Select the 5 best individuals from the population to proceed to the next round of optimization. After multiple rounds of genetic algorithm iterations, the population converges, and the scheme with the highest fitness is selected as the initial input for reinforcement learning.

[0122] Second phase optimization:

[0123] 1) Define the state s of reinforcement learning t ={x 11 =4,x 12 =1,x 13 =2,x 21=3.2,x 22 =1,x 23 =3,x 31 =2.8,x 32 =1,x 33 =2,w i =0.5,w i =0.3,w i =0.2};

[0124] Action space: Each task can choose to increase resources (i.e., increase x). ij ), reduce resources (i.e. reduce x) ij ), and maintain the current allocation (keep x) i constant).

[0125] The reward function takes the following form:

[0126] R t =U total (t)-U total (t-1);

[0127] 2) A Q-learning strategy was used for optimization, and the result still converged stably to x. 11 =4,x 12 =1,x 13 =2,x 21 =3.2,x 22 =1,x 23 =3,x 31 =2.8,x 32 =1,x 33 =2.

[0128] Step 3, optimize the generation of the resource allocation scheme, specifically as follows:

[0129] First, the task priority is dynamically adjusted by calculating the completion status of the time constraints for the three tasks. Based on the task time constraint formula, task d1 is not completed, while d2 and d3 are completed. Since task d1 is not completed, its priority needs to be increased. Therefore, the priority is updated according to the priority update formula, resulting in w1. (t+1) =0.6, w2 (t+1) =0.24, w3 (t+1) =0.16.

[0130] Secondly, the resource reallocation mechanism: based on the new weight update results, it checks whether the current resource allocation meets the minimum requirements (i.e., the task's minimum requirement constraint). After calculation, x in task d1... 11 =4≥1.2 (satisfied), x 12 =1≥0.6 (satisfied), x 13 =2≥1.8 (satisfied); x in task d221 =3.2≥0.5 (satisfied), x 22 =1≥0.5 (satisfied), x 23 =3≥0.5 (satisfied); x in task d3 31 =2.8≥0.5 (satisfied), x 32 =1≥0.3 (satisfied), x 33 =2≥0.5 (satisfied). Simultaneously, the total resource allocation also satisfies (as follows), therefore no resource reallocation is required.

[0131] r1 = 4 + 3.2 + 2.8 = 10 ≤ 10;

[0132] r2 = 1 + 1 + 1 = 3 ≤ 5;

[0133] r3 = 2 + 3 + 2 = 7 ≤ 20;

[0134] Step 4, evaluation and feedback: Based on the formula for task completion rate, the calculated task completion rate is not 100%, and the resource utilization rate is 57.14%, which meets the expected indicators (task completion rate not less than 67%, resource utilization rate not higher than 80%). Therefore, this allocation scheme is adopted.

[0135] This embodiment also discloses a resource allocation system based on product technical assurance indicators, including:

[0136] The model building module is used to define a task utility function and obtain a game model based on a set of task requirements, a set of available guaranteed resources, a set of priority weights, a set of task time factors, and a set of initial strategies. The game model is configured with constraints, including resource ceiling constraints, minimum task requirement constraints, task time constraints, and resource sharing mechanism constraints.

[0137] The resource allocation module is used to solve the Nash equilibrium point of resource allocation based on a game model using a hybrid optimization method, and to generate an initial resource allocation scheme.

[0138] The scheme optimization module is used to optimize the initial allocation scheme through a reinforcement learning feedback mechanism and a dynamic priority adjustment method to generate an optimized resource allocation scheme;

[0139] The scheme evaluation module is used to evaluate the optimized resource allocation scheme based on the task completion rate and resource utilization rate. If the preset threshold is met, the final resource allocation scheme is obtained; if the preset threshold is not met, the initial resource allocation scheme is regenerated.

[0140] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.

[0141] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A resource allocation method based on product technical assurance indicators, characterized in that, Includes the following steps: Step 1: Based on the task requirement set, available guaranteed resource set, priority weight set, task time factor, and initial strategy set, define the task utility function to obtain the game model; the game model is configured with constraints, namely resource upper limit constraint, minimum task requirement constraint, task time constraint, and resource sharing mechanism constraint. Step 2: Solve for the Nash equilibrium point of resource allocation based on the game model using a hybrid optimization method to generate an initial resource allocation scheme; Step 3: Optimize the initial allocation scheme by using a reinforcement learning feedback mechanism and a dynamic priority adjustment method to generate an optimized resource allocation scheme; Step 4: Evaluate the optimized resource allocation scheme based on the task completion rate and resource utilization rate. If the preset threshold is met, the final resource allocation scheme is obtained. If the preset threshold is not met, proceed to step 2.

2. The resource allocation method based on product technical assurance indicators according to claim 1, characterized in that, The expression for the utility function is: In the formula, α ij For resource r j For task d i Contribution weight; β i For task d i Resource consumption coefficient; x ij Assigned to task d i resources r j ;γ i For task d i Time sensitivity coefficient; T i This represents the task completion time factor.

3. The resource allocation method based on product technical assurance indicators according to claim 1, characterized in that, The expression for the constraint of the resource sharing mechanism is: In the formula, x ij Assigned to task d i resources r j ; For resource r j The total amount; To remove task d i In addition, the amount of shared resources for the k-th task.

4. The resource allocation method based on product technical assurance indicators according to claim 1, characterized in that, The hybrid optimization method is a phased optimization algorithm. In the first phase, a genetic algorithm is used to screen the initial strategy based on fitness and output the screening scheme. In the second phase, reinforcement learning is used to dynamically optimize the screening scheme in the first phase and generate the initial resource allocation scheme.

5. A resource allocation method based on product technical assurance indicators according to claim 4, characterized in that, The optimization steps in the first stage are as follows: Step 211: Initialize the population. Each individual in the population represents a resource allocation scheme. Step 212: Calculate the fitness function, which is constructed based on the task utility function and priority weights, and is used to evaluate the resource allocation scheme; Step 213: Perform crossover and mutation operations, using a single-point crossover and mutation operator to generate new individuals; Step 214: Select K optimal individuals and jump to step 212 for iterative optimization until the iteration stops, and output the screening scheme.

6. The resource allocation method based on product technical assurance indicators according to claim 5, characterized in that, The optimization steps in the second stage are as follows: Define the state, action space, and reward function; The resource allocation scheme of the screening scheme is updated using the strategy gradient optimization method until it converges to the Nash equilibrium point, thus obtaining the initial resource allocation scheme.

7. A resource allocation method based on product technical assurance indicators according to claim 6, characterized in that, The steps for obtaining the optimized resource allocation scheme are as follows: Based on task time constraints, the task priority weights are dynamically adjusted using a reward function to obtain optimized task weights. Check whether the resources in the resource allocation scheme meet the minimum requirement constraint of the task. If they do, output the optimized resource allocation scheme; otherwise, trigger the resource reallocation mechanism to allocate resources based on the optimized task weight.

8. A resource allocation system based on product technical assurance indicators, characterized in that, include: The model building module is used to define a task utility function and obtain a game model based on a set of task requirements, a set of available guaranteed resources, a set of priority weights, a set of task time factors, and a set of initial strategies. The game model is configured with constraints, including resource ceiling constraints, minimum task requirement constraints, task time constraints, and resource sharing mechanism constraints. The resource allocation module is used to solve the Nash equilibrium point of resource allocation based on a game model using a hybrid optimization method, and to generate an initial resource allocation scheme. The scheme optimization module is used to optimize the initial allocation scheme through a reinforcement learning feedback mechanism and a dynamic priority adjustment method to generate an optimized resource allocation scheme; The scheme evaluation module is used to evaluate the optimized resource allocation scheme based on the task completion rate and resource utilization rate. If the preset threshold is met, the final resource allocation scheme is obtained. If the preset threshold is not met, the initial resource allocation scheme will be regenerated.

Citation Information

Cited By

  • Intelligent equipment maintenance resource dynamic optimization distribution system

    CN121684873A