A method for optimizing operation and maintenance strategy of a mission critical system performing heterogeneous tasks
By optimizing the load and mode switching strategies of mission-critical systems through Bayesian updates and semi-Markov decision processes, the problem of optimizing operation and maintenance strategies under dynamic, stochastic, and heterogeneous mission backgrounds is solved, enabling efficient operation and flexible response of the system under unknown conditions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING INST OF TECH
- Filing Date
- 2026-02-27
- Publication Date
- 2026-07-03
Smart Images

Figure CN122332152A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of system operation and maintenance strategy optimization, and specifically discloses a method for optimizing the operation and maintenance strategy of a task-critical system that performs heterogeneous tasks. Background Technology
[0002] In critical sectors such as emergency medical services, manufacturing, and new energy development, systems often need to perform mission-critical tasks. The reliability and performance of these mission-critical systems directly affect public safety, transportation efficiency, infrastructure integrity, and economic benefits. Therefore, researching and optimizing operation and maintenance strategies for mission-critical systems is of great significance for reducing related costs and mitigating safety risks.
[0003] To mitigate the risk of unplanned system downtime, implementing efficient operation and management strategies is crucial, with maintenance strategies being the most widely used core approach in industry. However, in real-world engineering scenarios, executing maintenance plans requires significant time for tool allocation, spare parts preparation, and the deployment of skilled technicians, resulting in insufficient flexibility and difficulty in responding to sudden demands in a short period. Therefore, exploring effective pathways to mitigate system degradation without downtime is of great practical significance.
[0004] Adjusting load levels is one key technical means that combines feasibility and economy. System load level is a core indicator in the process of task execution, directly reflecting the efficiency of task completion within a specified time. This efficiency exhibits specific characteristics in different scenarios: in production systems, load level is positively correlated with productivity, with higher load meaning higher output rate; in transportation systems, it is related to the speed of cargo transfer or the weight carried; and for wind turbines, load level is specifically reflected in operating speed, directly affecting power generation efficiency.
[0005] The strong correlation between load level and system degradation makes dynamically adjusting the load to regulate the degradation process a highly feasible alternative risk control scheme. This method is applicable to a variety of engineering systems: for example, gearboxes and generators in wind turbines experience accelerated degradation under high-speed operation; conveyor belts are prone to failure at high speeds; trucks experience significantly shortened service life under overload conditions; large computer clusters experience a significant increase in failure rate during overload operation; and cutting tools experience accelerated wear rates during high-speed cutting.
[0006] Existing research on load level adjustment strategy optimization mainly falls into two categories: system age-based and system state-based. These studies have achieved gradual upgrades from single-component to multi-component, from static planning to dynamic control, and from single-objective to multi-objective optimization. However, existing research still has significant limitations: most focus on modeling and controlling the system's own degradation process, or design strategies only for fixed task scenarios with single tasks and long-term continuous execution. They fail to jointly optimize the task arrival process with load adjustment strategies, neglecting the critical impact of dynamic changes in task flow on system operation and maintenance decisions. However, in real-world engineering scenarios, task arrival is not entirely controllable, often exhibiting dynamic characteristics (dynamic task arrival), randomness (random task arrival time and workload), and heterogeneity (differences in workload and arrival rate among different tasks). This leads to dynamic fluctuations in the task waiting queue. Therefore, comprehensively considering the dynamic changes in the task arrival process and the task waiting queue when optimizing load adjustment strategies can maximize the overall system operational efficiency while ensuring system reliability.
[0007] Furthermore, mode switching is a widely used method to reduce the risk of system failure. Systems can operate in different modes, such as online mode, offline mode, and various standby modes, each with a different degradation rate. Making decisions about when and how to switch is crucial for ensuring optimal system performance, extending system lifespan, and minimizing the probability of unplanned downtime. Existing research largely focuses on passive mode switching in reserve systems (i.e., switching triggered by component failure), while research on dynamic switching between multiple system operating modes is relatively limited.
[0008] When optimizing the aforementioned operation and maintenance strategies, the limitations of traditional optimization and solution methods become increasingly apparent due to the need to comprehensively consider the dynamic, random, and heterogeneous nature of the task arrival phase, as well as the random degradation process of the system during task execution. This problem is particularly prominent in the context of heterogeneity. Existing research on task arrival processes focuses more on the dynamic and random characteristics of task arrival, while relatively insufficiently considering task heterogeneity. Task heterogeneity refers to the differences in core attributes such as task volume and arrival rate when multiple tasks originate from the same task group, and the decision-maker has limited prior information about this task group. With the advancement of modern information technology, by utilizing technologies such as the Internet of Things, real-time monitoring, and real-time data analysis, complete arrival and completion path information for each task can be utilized, including task arrival rate, task volume, and completion time, thereby relaxing assumptions about known task arrival processes. This leads to two key stochastic processes in the optimization of operation and maintenance strategies: first, the task arrival process affected by unknown arrival rates and task volume distributions; and second, the task execution process exhibiting random degradation dependent on load levels. Therefore, there is an urgent need to construct an integrated optimization method that combines update mechanisms and decision-making models to systematically solve the problem of optimizing operation and maintenance strategies under the above-mentioned multi-factor constraints. Summary of the Invention
[0009] To address the aforementioned problems in existing technologies, this invention proposes an optimization method integrating Bayesian update and semi-Markov decision process. Based on the system degradation level, system age, number of tasks arriving, and cumulative number of tasks arriving at the decision time, the method optimizes the load level adjustment strategy and mode switching strategy. Furthermore, based on the structural properties of the optimal strategy, the optimization solution algorithm is improved.
[0010] The technical solution of the present invention is as follows:
[0011] An optimization method for operation and maintenance strategies of task-critical systems performing heterogeneous tasks includes the following steps:
[0012] S1: Different tasks have different and unknown arrival rate and task quantity parameters. At the time of completion of each task, the posterior distribution of arrival rate and task quantity parameters is obtained based on the prior distribution of arrival rate and task quantity parameters of each task and the new information of the task. The hyperparameter update rule and sufficient statistics characterizing the randomness, dynamics and heterogeneity of the task are obtained.
[0013] S2: The degradation process of the system is described by a continuous-time Markov process, and the system mode switching strategy and load adjustment strategy are designed.
[0014] S3: Based on the sufficient statistics obtained in S1 and the degree of system degradation, construct optimization models for mode switching strategy and load adjustment strategy with the goal of minimizing the expected total cost based on the framework of semi-Markov decision process (SMDP).
[0015] S4: Use an improved algorithm to solve the optimization model described in S3 to obtain the optimal mode switching strategy and load adjustment strategy when each task is completed.
[0016] Furthermore, the specific process of step S1 is as follows:
[0017] S1.1: Define arrival rate parameter and task quantity parameters prior distribution
[0018] The arrival of the task follows a homogeneous Poisson process, and the arrival rate is determined by unknown parameters. This indicates that the workload of each task follows a parameter containing unknown parameters. The exponential distribution; the unknown parameters and Treat it as a random variable, and use and express;
[0019] Assumption The prior distribution is the shape parameter and scale parameters The Gamma distribution has a prior density distribution of . , The prior distribution is the shape parameter and scale parameters The Gamma distribution has a prior density distribution of . ;
[0020] S1.2: Obtaining hyperparameter update rules and sufficient statistics using Bayes' theorem
[0021] No. The task is up to the first The observation data for arriving at the task between the completion times of each task are: ,in, for The observed values, Indicates completion of the first The task is up to the first The random duration between tasks is determined by the task waiting time. and task completion time It consists of two parts; for The observed values, Indicates completion of the first The task is up to the first The number of random tasks arriving between tasks; for The observed values, of which Indicates the number of times reached The workload of each task; in addition, let Indicates completion of the first The task is up to the first The cumulative task volume during each task period, its observed value is... ;
[0022] Using the first When each task is completed and The joint posterior density function as the first The prior density function when each task is completed, i.e. Combined with the first The task is up to the first Observational data arriving at the task between the completion times of each task. Based on Bayes' theorem, we can obtain the first... When each task is completed and Joint posterior density function:
[0023]
[0024] According to formula (1), and The joint posterior density function is equal to two Gamma distributions. and The product of density functions leads to the following hyperparameter update rule: , , and ;
[0025] Therefore, given the initial hyperparameters and When the first The information about when a task is completed can be equivalent to a sufficient statistical measure. ,in , and They represent the first The number of tasks completed when each task is finished, and the age (from the initial time of system operation to the [number]th task). The cumulative time taken to complete each task and the cumulative number of tasks reached.
[0026] Furthermore, in step S2, the system operation process is characterized as follows: the system needs to execute... One task, complete The reward for each task is Using a continuous-time Markov process with non-negative, stationary, and independent increments. To describe the degradation process of the system, time interval Internal degradation increment The distribution function can be expressed as Let the first Maintenance will be performed upon completion of each task; if the system malfunctions, corrective maintenance will be carried out at a cost of [amount missing]. Otherwise, preventative maintenance will be carried out, at a cost of .
[0027] Furthermore, in step S2, the mode switching strategy is as follows: the system includes an online mode and an offline mode, and switches between the online and offline modes. When in online mode, the system continuously degrades and incurs operating costs, denoted as the operating cost per unit time. When the system has no pending tasks at the time of task completion, it switches to offline mode or remains in online mode. In offline mode, the system does not degrade, but switching back to online mode when a new task arrives incurs switching costs. When the system still has tasks to process at the time of task completion, the system continues to remain online and execute tasks; to balance operating costs and switching costs, a decision must be made at the decision point whether to switch modes.
[0028] The load adjustment strategy is as follows: the system has an operating load-degradation dependency, that is, when a higher operating load is used, the system is more likely to experience accelerated degradation; in order to balance benefits and maintenance costs, the operating load level to be used in the next task execution phase needs to be determined at the decision point.
[0029] Furthermore, in step S3, based on the framework of a semi-Markov decision process and using elements such as the state set, action set, and value function, an optimization model for the mode switching strategy and load adjustment strategy with the objective of minimizing the expected total cost is constructed, as follows:
[0030] Status: Completed When performing a task, the state of the semi-Markov decision process (SMDP) is used as... It means that, among them, Represents the completion of the first The degree of system degradation during each task; when the degree of degradation exceeds a fixed failure threshold. When this happens, the system transitions to a failure state. ;
[0031] Action set: each using and This indicates switching the system to online or offline mode; in the... When a task is completed, based on the state... Possible actions It can be represented as ,in ,and ,in Indicates the system load level. and These are the minimum load level and the maximum load level, respectively; specifically, if the first... There are remaining tasks when there are 1 task, that is... By default, the mode is switched to online mode, i.e. ;
[0032] Value function: Let and Respectively represent and and Related single-stage expected cost; Define value function For the stage Based on state The minimum expected total cost, i.e., the optimization model, can be expressed as:
[0033]
[0034] in, Representing a Markov process Regarding thresholds When the first arrival, The distribution function is ;
[0035] The boundary conditions for the optimization model are:
[0036] .
[0037] Furthermore, in step S4, an improved algorithm is used to solve the optimization model described in S3, as follows;
[0038] For all possible states in the Nth stage The corresponding value function is initialized, and the optimal mode switching strategy and load adjustment strategy for all states from the Nth stage to the 1st stage are solved using a backward iterative algorithm. In the solution process, the structural properties of the optimal strategy are used to solve the problem, namely: the degree of system degradation, the number of tasks reached, and the system age each have thresholds. For the optimal mode switching strategy of all states in the current stage that exceed any threshold, the optimal mode switching strategy is the opposite of the optimal mode switching strategy of the states in the current stage that do not reach the threshold.
[0039] For the optimal load level strategy for a certain state in the current stage, the number of tasks arriving is compared with the number of tasks arriving in the corresponding state of the known optimal load level strategy, or the cumulative number of tasks arriving is compared with the cumulative number of tasks arriving in the corresponding state of the known optimal load level strategy. Based on the comparison result and the known optimal load level strategy, the optimal load level for the current state is adjusted. Using the structural properties of the optimal strategy for solution can reduce the search space, accelerate convergence, and reduce computational complexity. For example, when... When, if the state is obtained The corresponding optimal mode switching strategy is For states with a larger number of tasks arriving, the optimal mode switching strategy must also be... If the state is obtained The corresponding optimal load is For states with a larger cumulative arrival volume of tasks, the optimal load will not be less than [a certain value]. .
[0040] Furthermore, the structural properties of the optimal strategy include the control limit properties of the optimal mode switching strategy and the threshold monotonicity of the optimal load level strategy;
[0041] The control limit property of the optimal mode switching strategy is: under a fixed load control strategy, there exists a threshold. ,when Switching to offline mode is optimal when... Switching to online mode is optimal at this time; a threshold exists. ,when Switching to online mode is optimal when... When this happens, switching to offline mode is optimal; [the text abruptly ends here, likely due to an incomplete sentence or missing information.] ,when Switching to offline mode is optimal when... Switching to online mode is optimal at this time;
[0042] The threshold monotonicity of the optimal load level strategy is: under a fixed mode switching strategy, for a fixed... status Optimal load level about It is non-incremental; for fixed... status Optimal load level about It is not subtractive.
[0043] The beneficial effects of this invention are as follows: This invention proposes an optimization method for the operation and maintenance strategy of a multimorphic system performing heterogeneous tasks. Heterogeneous tasks consist of multiple dynamically randomly arriving tasks with different and unknown arrival rates and task quantity parameters. During the random arrival process, this invention fully considers the cognitive uncertainty caused by unknown parameters in the task arrival process and task quantity distribution. During task execution, it fully considers the possibility of random degradation due to the dependency relationship between load and degradation, resulting in accidental uncertainty. To optimize mode switching and load control strategies under the above-mentioned mixed uncertainty, this invention integrates update and decision-making methods. The proposed framework combines Bayesian methods with semi-Markov decision processes to optimize situation-based dynamic mode switching and load control strategies, aiming to minimize the expected total cost by balancing task benefits with corrective maintenance, operation, and mode switching costs. Furthermore, this method proves and analyzes the monotonicity of the estimated random parameters and variables, establishes the structural characteristics of the optimal strategy, including monotonicity and control limit properties, and designs an algorithm to solve for the optimal strategy based on these characteristics. Attached Figure Description
[0044] Figure 1 This is a schematic diagram of the decision optimization based on Bayesian update in this invention.
[0045] Figure 2 This is a schematic diagram of the system operation process of the present invention; where (a) represents a successful task and (b) represents a system failure.
[0046] Figure 3 This is an optimal load level diagram for different task arrival numbers and degradation levels in Embodiment 2 of the present invention.
[0047] Figure 4 This is the optimal load level diagram for different numbers of arriving tasks and the cumulative number of arriving tasks in Embodiment 2 of the present invention.
[0048] Figure 5 This is an optimal load level diagram for different cumulative task amounts and degradation levels in Embodiment 2 of the present invention. Detailed Implementation
[0049] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0050] It should be noted that the terms "comprising" and "having," and any variations thereof, in the embodiments and accompanying drawings of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the steps or units listed, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.
[0051] Example 1: This embodiment of the invention provides a method for optimizing the operation and maintenance strategy of a task-critical system that performs heterogeneous tasks, including the following steps:
[0052] S1. Since each task comes from a heterogeneous task group, and considering that different tasks have different and unknown arrival rates and task volume parameters, it is necessary to update the parameters of each task in real time. Based on the prior information of the task arrival process and the new information of the arrival between two adjacent tasks, the updated posterior distribution of the task arrival process is obtained using Bayes' theorem. This result will become the prior probability distribution for the next update; the specific process is as follows:
[0053] Consider a system performing heterogeneous tasks, where the arrival of tasks follows a homogeneous Poisson process, and the arrival rate is determined by unknown parameters. This indicates that the workload of each task follows a parameter containing unknown parameters. The exponential distribution. Parameters for each task. and Real-time learning is required. Complete the first... The task is up to the first The time between tasks is denoted as The waiting time of the task and task completion time It consists of two parts. (After completing the first...) The task is up to the first During each task period, the arriving task random number and task quantity are respectively used as... and Indicate. Let. Indicates completion of the first The task is up to the first The cumulative number of tasks during a given task period.
[0054] Each task originates from a heterogeneous group of tasks, and each task has different arrival parameters. and This is something the decision-makers were unaware of. (Regarding the parameters...) and Treat it as a random variable, and use and Indicate. Assume. The prior distribution is the shape parameter and scale parameters The Gamma distribution has a prior density distribution of . , The prior distribution is the shape parameter and scale parameters The Gamma distribution has a prior density distribution of . .
[0055] The update method is as follows: Figure 1 As shown: There is no accumulated arrival data before the first task arrives. Therefore, it is assumed that... and It follows an independent prior density distribution. Let and Let represent the hyperparameters before the first task is reached. These hyperparameters can be obtained from historical data or test data. Then the joint prior density distribution can be used... Indicate. Using the first The task is up to the first Observational data arriving at the task between the completion times of each task. and the When each task is completed and The joint posterior density distribution, i.e. It can update the first When each task is completed and The joint posterior density distribution. and The joint posterior density distribution is equal to the two Gamma distributions. and The product of these factors leads to the hyperparameter update rule: , , ,and Therefore, given the initial hyperparameters... and , No. The information about when a task is completed can be equivalent to a sufficient statistical measure. ,in , and They represent the first The number of tasks completed, age, and cumulative number of tasks completed when a task is finished.
[0056] S2. Due to the load-degradation dependency of the system, a reliability model of the system based on dynamic mode switching and load control strategies is constructed. The schematic diagram of the system operation process is shown below. Figure 2 As shown;
[0057] The system requires an adjustable load. Finish One task, of which and These are the minimum and maximum load levels, respectively. The system's degradation process is represented by a continuous-time Markov process with non-negative, stationary, and independent increments. To describe, degradation increment In time interval The cumulative distribution function can be expressed as The system exhibits load-degradation dependence; in particular, it is more likely to experience accelerated degradation at higher load levels, making it more susceptible to damage. Consider two different load levels, denoted as... and ( ), No. The task is up to the first Each task generates a corresponding random degradation increment. and This means that at smaller load levels Above, degradation increment Greater than The probability is lower than at larger load levels. The increment below Greater than The possibility, that is , The system's degradation rate depends on the load level, using This indicates the load-degradation relationship.
[0058] Maintenance activities are assumed to take place at a predetermined time, such as the first... Upon completion of each task. If a system failure occurs, corrective maintenance will be performed, at a cost of [amount missing]. Otherwise, preventative maintenance will be carried out, the cost of which is Higher payload levels accelerate mission completion, thereby increasing the expected rewards associated with mission success. The reward for each task is However, it is important to note that higher load levels can also accelerate system degradation, thereby increasing the expected cost of system failure. The trade-off between task benefits and system degradation underscores the necessity of optimizing overall outcomes and minimizing potential costs and risks.
[0059] The system can switch between online and offline modes. During online operation, the system responds quickly to incoming tasks. However, it's worth noting that the system degrades even during idle periods. Furthermore, the system incurs operational costs in online mode; these costs are denoted as [insert cost here]. In contrast, in offline mode, the system is unaffected by degradation if there are no pending tasks; however, the system incurs additional startup costs when tasks arrive. This includes startup resources and potential overhead. To minimize expected startup costs, keeping the system online is crucial, as it allows for immediate startup when tasks arrive. However, it's important to recognize that maintaining online status without tasks can lead to unnecessary operational costs and increase the risk of system failure due to continuous operation. Therefore, finding a balance between reducing startup costs, managing operational costs, and effectively mitigating the risk of system failure is a critical consideration for maintaining optimal system performance.
[0060] S3. Based on the posterior distributions of the arrival rate parameters and task quantity parameters obtained in step S1, and the reliability model in step S2, and using elements such as the state set, action set, and value function, an optimization model for the mode switching strategy and load adjustment strategy with the objective of minimizing the expected total cost is constructed within the framework of a semi-Markov decision process; specifically as follows:
[0061] Consider a semi-Markov decision process in a finite-time domain, which is a generalization of the discrete-time Markov decision process, where dwell times are not equal but follow arbitrary distributions. Assume the decision-maker can only choose an action after completing the nth task. This embodiment focuses on the system's evolution and the decision-making time for task arrival. The optimization model is formally described by the following components:
[0062] Status: Completed When performing a task, the SMDP state is used. It means that, among them, Represents the completion of the first The degree of degradation during each task; when the degree of degradation exceeds a fixed failure threshold. When this happens, the system transitions to a failure state. .
[0063] Action set: each using and This indicates that the system is switched between online and offline modes. (In the...) When a task is completed, based on the state... Possible actions It can be represented as ,in ,and Specifically, if the first... There are remaining tasks when there are 1 task, that is... By default, the mode is switched to online mode, i.e. .
[0064] Value function: Let and Respectively represent and and The relevant single-stage expected cost. Definition. For the stage Based on state The minimum expected total cost can be expressed as:
[0065]
[0066] in, Representing a Markov process Regarding thresholds When it first arrived. The distribution function is .
[0067] The boundary conditions for the optimization model are:
[0068]
[0069] S4. Based on the update process in S1, obtain the posterior distribution of the arrival rate parameter and the task quantity parameter of the heterogeneous task, as well as the monotonicity of the information of new tasks arriving between two adjacent tasks with respect to the state of the semi-Markov decision process.
[0070] Inference 1. In the first When each task is completed, the arrival rate and workload The posterior distribution satisfies the following random order:
[0071] (a) For state , In the given Time about Randomly decreasing, in a given Time about Randomly increasing.
[0072] In the given Define two different durations and ( Based on the hyperparameter update rule in step S1, we can obtain... In different The corresponding random parameters The hyperparameters can be expressed as and .make and They represent based on and The random parameters. Using the likelihood ratio order, i.e., let... and They are random variables, and their density functions are respectively and If the ratio Follow Increasing, then Less than in likelihood ratio The likelihood ratio is stronger than that of a typical random order, if... ,but . We can obtain
[0073]
[0074] This formula is about Increasing. Therefore, we can obtain That is, for the state , In the given Time about Randomly decreasing. Similarly, we can obtain... That is, for the state , In the given Time about Randomly increasing.
[0075] (b) For the state , In the given Time about Randomly increasing, in a given Time about Randomly decreasing.
[0076] In the given Define the cumulative arrival count for two different tasks. Based on the hyperparameter update rule in step S1, we can obtain In different The corresponding random parameters The hyperparameters can be expressed as and .make and Representing different hyperparameters and The random parameters are as follows. We can obtain...
[0077]
[0078] This formula is about Increasing. Therefore, we can obtain That is, for the state , In the given Time about Randomly increasing. Similarly, for the state... , In the given Time about Randomly decreasing.
[0079] Inference 2. In the first When each task is completed, , , and Satisfying the following random order:
[0080] (a) For state , In the given Time about Randomly decreasing, in a given Time about Randomly increasing.
[0081] Based on Corollary 1, we can conclude that... and .make (corresponding hyperparameters) or ) and (corresponding hyperparameters) or )express Implementation values under different hyperparameters. Given , can be obtained
[0082]
[0083] This formula is about Increasing. Therefore, we can obtain It maintains random order under monotonic transformations, that is, for states , In the given Time about Randomly decreasing, in a given Time about Randomly increasing.
[0084] (b) For the state , In the given Time about Randomly decreasing, in a given Time about Randomly increasing.
[0085] Based on Corollary 1, we can conclude that... and .make (corresponding hyperparameters) or ) and (corresponding hyperparameters) or )express Implementation values under different hyperparameters. Given , can be obtained
[0086]
[0087] This formula is about Increasing. Therefore, we can obtain It maintains random order under monotonic transformations, that is, for states , In the given Time about Randomly decreasing, in a given Time about Randomly increasing.
[0088] (c) For the state , In the given Time about Randomly decreasing, in a given Time about Randomly increasing.
[0089] Based on inference 2(b), the task completion time during the working phase is random. and Positive correlation, therefore, for state , In the given Time about Randomly decreasing, in a given Time about Randomly increasing.
[0090] (d) For state , In the given Time about Randomly increasing, in a given Time about Randomly decreasing.
[0091] Based on inference 2(a), given , can be obtained
[0092]
[0093] This formula is about Increasing. Therefore, we can obtain It maintains random order under monotonic transformations, that is, for states , In the given Time about Randomly increasing, in a given Time about Randomly decreasing.
[0094] S5. Based on the optimization model described in S3 and the properties in S4, the structural properties of the optimal strategy are obtained, including the control limit properties or threshold monotonicity of the optimal mode switching strategy and the optimal load level strategy. An improved backward iterative algorithm based on the structural properties of the optimal strategy is then presented. Specifically:
[0095] Assume distribution function It is about A monotonically non-decreasing function, and is about The assumption captures two intuitively monotonic properties of the failure time distribution. First, for a given fixed load level, when the degradation level... As the threshold increases, the system becomes more prone to failure because the degree of degradation required to reach the failure threshold decreases. Therefore, the distribution function... It is about It is a monotonically non-decreasing function. Secondly, for a given degenerate state, the performance level is increased. This will lead to a higher rate of degradation, making about It is a monotonically non-decreasing function.
[0096] The steps based on the backward induction algorithm employ induction, letting...
[0097]
[0098] for , It can be represented as
[0099]
[0100] Corollary 3. Given the first The status when a task is completed The monotonicity of the optimal value function is established as follows:
[0101] (a) For fixed status , Follow Non-reduction.
[0102] Boundary conditions for valued functions, i.e., formulas about It is a non-decreasing function. (Based on the formula) ,when hour, and about monotony and about The monotonicity is consistent and non-decreasing. Let Represents the given in the first At each decision point In the The random degradation level at each decision point. Consider two different degradation levels. and ( Based on the independent incremental property, we can obtain
[0103]
[0104] Based on the definition of random order, we can obtain... In its first The value function corresponding to each decision point. and Assuming In the given In the case of, for ,about It is a non-decreasing function, so we can obtain Due to the following random order theorem: if and If it is any increasing (decreasing) function, then ( )and ( Therefore, we can obtain , and when hour, about It is a non-decreasing function, that is, for a given... status , Follow Non-reduction.
[0105] (b) For fixed status , Follow Non-incremental.
[0106] Boundary conditions for valued functions, i.e., formulas about It is a non-increasing function. (Based on the formula) ,when hour, and and Irrelevant. Represents the given in the first At each decision point In the The number of tasks randomly arriving at each decision point. Consider two different numbers of tasks. and ( ), can be obtained
[0107]
[0108] Based on the definition of random order, we can obtain... In its first The value function corresponding to each decision point. and Assuming In the given In the case of, for ,about It is a non-increasing function, so we can obtain Therefore, we can obtain , and when hour, about It is a non-increasing function, that is, for a given... status , Follow Non-incremental.
[0109] (c) For fixed status , Follow Non-reduction.
[0110] Boundary conditions for valued functions, i.e., formulas about It is a non-decreasing function. (Based on the formula) ,when hour, and and Irrelevant. Represents the given in the first At each decision point In the The age at each decision point. Consider two different ages. and ( ), can be obtained
[0111]
[0112] Based on the definition of random order, we can obtain... In its first The value function corresponding to each decision point. and Assuming In the given In the case of, for ,about It is a non-decreasing function, so we can obtain Therefore, we can obtain , and when hour, about It is a non-decreasing function, that is, for a given... status , Follow Non-reduction.
[0113] (d) For fixed status , Follow Non-reduction.
[0114] Boundary conditions for valued functions, i.e., formulas about It is a non-decreasing function. (Based on the formula) ,when hour, and and Irrelevant. Represents the given in the first At each decision point In the The cumulative amount of work done at each decision point. Consider two different amounts of work. and ( ), can be obtained
[0115]
[0116] Based on the definition of random order, we can obtain... In its first The value function corresponding to each decision point. and Assuming In the given In the case of, for ,about It is a non-decreasing function, so we can obtain Therefore, we can obtain , and when hour, about It is a non-decreasing function, that is, for a given... status , Follow Non-reduction.
[0117] Based on Corollary 3, control limit strategies for the optimal mode switching strategy and load control strategy are established in the following two theorems.
[0118] Theorem 1. Under a fixed load control strategy, the optimal mode switching strategy is established as follows:
[0119] (a) A threshold exists ,when Switching to offline mode is optimal when... Switching to online mode is optimal at this time.
[0120] Utilizing superadditivity, i.e., real-valued functions for and have ,but about It possesses superadditivity. If the inverse inequality above holds, then it is said to have superadditivity. It possesses subadditivity. Utilizing subadditivity and superadditivity, the control limits for optimal mode switching and load control strategies can be obtained. If , can be obtained
[0121]
[0122] according to ,in This is directly related to the workload. Based on the Bayesian update formula, the following posterior distribution can be obtained:
[0123]
[0124] Since the load level represents the task's execution rate, the duration of the work phase... The posterior distribution is
[0125]
[0126] also, The posterior distribution is
[0127]
[0128] therefore, The distribution can be represented by the following convolution.
[0129]
[0130] Therefore, for all achievable , about It is monotonically increasing, inequality It was established. Therefore... and about It has superadditivity.
[0131] According to the state transition probability, we can obtain
[0132]
[0133] Based on this, we can obtain It is about The non-decreasing function. Based on Corollary 3, It is about A non-decreasing function. Let and Let be a real-valued non-negative sequence, satisfying Assuming It is a real-valued sequence that satisfies the monotonicity condition, that is, for ,have , can be obtained Therefore, the following property can be obtained: assuming and Let be a random variable, and , Therefore, if Established, . We can obtain It exhibits superadditivity. An optimal mode-switching strategy exists, i.e., there exists... ,when Switching to offline mode is optimal when... Switching to online mode is optimal at this time.
[0134] (b) A threshold exists ,when Switching to online mode is optimal when... Switching to offline mode is the optimal choice.
[0135] To prove the control limit property, we utilize the additive nature of the expected cost function over a single period. The monotonicity. If You can get information about It is a non-increasing function.
[0136]
[0137] Based on inference 2, since The distribution of Decreasing in random order, and about Increasing, the expectation of this weighted integral is about It is decreasing. Therefore, about Monotonically decreasing. To prove... about The monotonicity of the difference is defined. The purpose is to prove that the difference is related to Strictly incremental.
[0138]
[0139] based on The distribution can be obtained
[0140]
[0141] This formula represents the derivative of a single Pareto density at a given point. According to... The posterior distribution can be obtained. .because The distribution is convolution Its derivative can be obtained as
[0142]
[0143] The derivative is less than Therefore, based on the previous proof, we can obtain the inequality. .
[0144] Using the different numbers of arrivals in Corollary 2 and of and The random order of states affects the transition probabilities of several states. We analyze the additivity of their transition probabilities separately:
[0145]
[0146]
[0147] The last one is about The decreasing function, in addition to Corollary 3(b), It is about Since the sum of subadditive functions is subadditive, we can obtain... This is an option that can be added. Therefore, the optimal action with the goal of minimizing costs can be determined. If needed, select online mode; otherwise, select offline mode.
[0148] (c) A threshold exists ,when Switching to offline mode is optimal when... Switching to online mode is optimal at this time.
[0149] Based on corollaries 2 and 3, we can obtain , and about The monotonicity of the policy. To further verify the optimality of the policy, it is necessary to prove the superadditivity of the expected cost function and the state transition probability within a single cycle. Based on this, it can be proven that there exists a... The threshold that guarantees the optimal strategy. If You can get information about Increasing function:
[0150]
[0151] Based on inference 2, since The distribution of Increasing in random order, and about Increasing, the expectation of this weighted integral is about It is increasing. Therefore, about Monotonically increasing. Because... The distribution can be represented as convolution. Its derivative can be obtained as
[0152]
[0153] Therefore, we can obtain We can obtain the inequality Referring to the proof in (b), the inequality can be obtained. Based on inference 3, about Non-subtraction yields It is superadditive. Therefore, it exists. ,when Switching to offline mode is optimal when... Switching to online mode is optimal at this time.
[0154] In the discussion of the monotonicity of the optimal load control strategy, the mode switching strategy is first set as a fixed decision, and the optimal load level is defined as... .
[0155] Theorem 2. Optimal Load Level The monotonicity is established as follows:
[0156] (a) For fixed status Optimal load level about It is not an increase.
[0157] Consider two different and ( By applying subadditivity, if , can be obtained
[0158]
[0159] Its derivative can be expressed as
[0160]
[0161] To prove that the derivative term is nonnegative, we need to prove the following:
[0162] (1) Due to about It is non-subtractive, for all , can be obtained . The larger the value, the earlier the failure time. . We can obtain Similarly, we can obtain .
[0163] (2) The random ordered property established from Corollaries 2 (c) and (d), and about It decreases in random order. Because , Also satisfies the requirements regarding It decreases in random order. Therefore, based on and Likelihood ratio and log-concavity, convolution Similarly about It exhibits logarithmic concavity. Furthermore, it can be obtained that for all... , It is true, that is If true, then we can obtain Established.
[0164] (3) Based on The posterior distribution can be obtained. At the same time, based on different , can be obtained .because , can be obtained
[0165]
[0166] therefore, Established, and about Non-subtractive, such that This holds true. Similarly, we can obtain the inequality... It was established. Therefore... and about It has subadditivity.
[0167] Load control strategies affect task completion time and system degradation. This can be achieved by applying Corollary 2, which considers different arrival times for each task. and Down The random order makes the state transition probabilities at different performance levels... about Monotonically decreasing. The transition probability of degradation level is non-decreasing for different numbers of completed tasks. about It is an increasing function. Therefore, the inequality... Established. Furthermore... about Since it is a non-increasing function, we can obtain This can be added next. Therefore, for fixed... status Optimal load level about It is not an increase.
[0168] (b) For fixed status Optimal load level about It is not subtractive.
[0169] if You can get a result about Non-decreasing functions:
[0170]
[0171] Its derivative can be obtained as follows:
[0172]
[0173] To prove that the derivative term is nonnegative, we need to prove the following:
[0174] (1) Similar to the proof in (a), we can obtain and .
[0175] (2) The random ordering property established from Corollary 2(c), about It is increasing in random order. Therefore, based on From the likelihood ratio order and log-concavity, we can obtain If true, then we can obtain Established.
[0176] (3) Based on The posterior distribution can be obtained. At the same time based on different , can be obtained Therefore, we can obtain ,and about It is non-subtractive. This makes the inequality... This holds true. Similarly, we can obtain the inequality... It was established. Therefore... and about It has subadditivity.
[0177] Inspired by the proof of Corollary 2(a), using , and The random order, while taking into account the cumulative workload of different arrival tasks. and , and thus .also, about Since it is a non-decreasing function, we can obtain It is super-additive. Therefore, for fixed... status Optimal load level about It is not subtractive.
[0178] A core challenge of applying backward induction to SMDP is the curse of dimensionality, which significantly increases computational complexity and limits its applicability. This is because many methods ignore the structural properties of the optimal policy, treating the entire policy space as the search domain. However, structural properties such as monotonicity and control bounds can be utilized to reduce the search space, accelerate convergence, and decrease computational complexity. This invention confirms the existence and monotonicity of the optimal threshold and incorporates these properties into the algorithm design, proposing an improved backward iterative algorithm for solving the optimal policy.
[0179]
[0180] Example 2: This example proposes an optimization method for the operation and maintenance strategy of a task-critical system executing heterogeneous tasks. The heterogeneous tasks consist of multiple dynamically arriving tasks with different and unknown arrival rates and task volume parameters. During the random arrival process, this example fully considers the cognitive uncertainty caused by unknown parameters in the task arrival process and task volume distribution. During task execution, this method fully considers the possibility of random degradation due to the dependency relationship between load and degradation, resulting in accidental uncertainty. To optimize mode switching and load control strategies under the above-mentioned mixed uncertainty, this example integrates update and decision-making methods.
[0181] This embodiment primarily utilizes a combination of Bayesian methods and semi-Markov decision processes to optimize situation-based dynamic mode switching and load control strategies. The aim is to minimize the expected total cost by balancing mission benefits with corrective maintenance, operational, and mode switching costs. The following section uses an unmanned aerial vehicle (UAV) system as an example to introduce a method for optimizing the operation and maintenance strategy of a mission-critical system performing heterogeneous tasks, as proposed in this invention. The method includes the following steps:
[0182] Step 1: Define the system's degradation process and the parameters designed for operation. Taking a UAV system performing a transportation mission as an example, with adjustable payload levels... The load-degradation relationship is considered to be... This is a linear function. Parameters and These represent the load levels as follows: and The degradation rate over time. In the model of this invention, the gamma process is used to capture the continuous energy degradation process of lithium-ion batteries, and its scale parameter is... The shape function is The shape parameter is equal to The scale parameter is determined as a function of the load level. ,in This represents the standard deviation degradation increment at the maximum load level.
[0183] Drones are highly adaptable platforms capable of seamlessly switching between online and offline modes to meet diverse operational needs. When online, drones actively perform transport tasks or remain idle. In idle mode, drones hover in fixed locations, demonstrating their flexibility and responsiveness, quickly initiating task execution upon receiving a request. However, even during idle periods, it must be acknowledged that drone systems experience energy degradation, leading to operational costs, including energy consumption, equipment depreciation, and various indirect expenses required to maintain system functionality and readiness. Conversely, when there are no pending tasks, the system can enter offline mode. In this mode, the drone system is unaffected by degradation. However, when a new task arises, the system requires a startup procedure, involving setup time, resource allocation, and potential latency. If the degradation level exceeds a fixed failure threshold... The system degradation level will transition to a failure state. If the system is in If it fails at the end of the cycle, corrective maintenance costs will be incurred. Otherwise, preventative maintenance costs will be incurred. and benefits Therefore, it is necessary to optimize the mode switching strategy and payload control strategy of UAVs.
[0184] Table 1. Values of indicators in the numerical examples
[0185]
[0186] Transportation tasks arrive randomly and follow a compound Poisson process. Each task has a random workload. , It is represented by the transportation distance it covers and follows an exponential distribution. Due to the lack of parameters regarding the number and workload of random tasks by decision-makers... and Therefore, it is necessary to utilize hyperparameters for this information. , and observed arrival data To learn the task arrival process. However, decision-makers lack true hyperparameters to control task arrival behavior. , Instead of accessing precise information about hyperparameters, they have only historical data that can be used for model calibration. In the absence of precise information about hyperparameters, decision-makers rely on the calibration process to estimate the model and align it with observed task arrival patterns.
[0187] Step 2: Based on the current parameters, analyze the optimal mode switching strategy with control limits under different states.
[0188] Table 2 illustrates how the optimal mode switching strategy changes with the number of arriving tasks and the level of system degradation. This means that there are no remaining unfinished tasks in the system, and the system needs to wait for new tasks to arrive. When the system is in a good state, the optimal system mode is to switch to online mode, allowing tasks to be executed immediately upon arrival. However, when the system state deteriorates, online mode may lead to more severe degradation during task waiting, manifested as an increased probability of system failure, thus reducing the reliability of task execution. In this case, by switching to offline mode, the system effectively reduces the risk of failure, reduces runtime, and thus lowers the probability of failure. Therefore, switching the system to offline mode becomes the optimal choice. This means there are still unexecuted tasks in the system, and the system mode will switch back to online mode.
[0189] Table 2 Different states during task execution , And the given Optimal mode switching strategy
[0190]
[0191] Table 3 illustrates how the optimal mode switching strategy changes with the number of arriving tasks and the system age. At this point, it means there are no remaining unfinished tasks in the system, and the system needs to wait for new tasks to arrive. According to Corollary 2, in the subsequent time period, the number of tasks arriving... With system age It is decreasing, while the task waiting time... Then it depends on the system age The rate of increase is incremental. This indicates that older systems are more likely to receive fewer tasks than newer systems, and older systems typically experience longer task wait times compared to newer systems. Therefore, when a system is younger, the probability of shorter task wait times is higher. Keeping the system online may therefore be advantageous, as it avoids the costs of switching and allows tasks to be executed immediately upon arrival. Conversely, as the system ages, the probability of longer task wait times increases. In this case, running the system online will lead to unnecessary operational costs and may exacerbate the risk of system failure. Therefore, when a system reaches a certain age, considering switching it to offline mode becomes more optimized, improving reliability and cost-effectiveness. When This means there are still unexecuted tasks in the system, and the system mode will switch back to online mode.
[0192] Table 3 Different states during task execution , And the given Optimal mode switching strategy
[0193]
[0194] Table 4 illustrates the variations in the optimal mode-switching strategy, exploring how this strategy responds to fluctuations in system degradation levels and age. It's important to note that this analysis was conducted under the condition that there are no remaining unfinished tasks in the system. On one hand, switching the system to offline mode is optimal when the degradation level reaches a certain threshold. As degradation intensifies, the risk of failure increases. To effectively mitigate this risk, switching the system to offline mode can effectively avoid potential degradation caused by the system being idle during task waiting periods, thereby extending the system's lifespan and improving overall cost-effectiveness. On the other hand, switching the system to offline mode is also considered optimal when the system age reaches a certain threshold. As the system ages, it tends to encounter longer task waiting times in subsequent periods. To maintain efficiency and reduce operating costs, switching the system to offline mode can effectively reduce the adverse financial impact associated with excessively long task waiting times, which not only optimizes resource allocation but also maintains operational efficiency. Essentially, switching to offline mode is a proactive strategy to address the challenges posed by degraded systems, ensuring continued operational efficiency while controlling operating costs. Through careful planning and implementation, decision-makers can effectively leverage the advantages of offline mode, simplify operations, maximize resource utilization, and enhance long-term cost sustainability.
[0195] Table 4 Different states during task execution , And the given Optimal mode switching strategy
[0196]
[0197] Step 3: Based on the current parameters, analyze the optimal load control strategy under different conditions. Figure 3 , 4 Figures 5 and 6 illustrate the monotonic relationships between the optimal load level and three states: the number of tasks arriving, the cumulative workload of arriving tasks, and the system degradation level. Based on the dependency between load and degradation, when the system degradation level is high, increasing the load level leads to a greater increase in degradation, thus exacerbating the overall system degradation and increasing the risk of system failure. Therefore, as system degradation intensifies, the optimal load level will decrease accordingly. Based on Corollary 2, the number of tasks arriving in the next stage... As the number of tasks arrives Decreasing, while the cumulative task volume increases as tasks are completed. Incremental. Therefore, as the number of arriving tasks increases, the optimal load level will decrease accordingly; while as the cumulative number of arriving tasks increases, the optimal load level will increase accordingly.
[0198] Finally, it should be noted that the above embodiments are intended to illustrate the technical solutions of the present invention and do not constitute any limitation on the present invention. Those skilled in the art should fully understand that modifications to the technical solutions described in the foregoing embodiments or equivalent substitutions for any part or all of the technical features are entirely feasible. Such modifications or substitutions, as long as they do not depart from the scope of protection defined by the claims of the present invention, should be considered reasonable extensions of the present invention.
Claims
1. A method for optimizing the operation and maintenance strategy of a task-critical system performing heterogeneous tasks, characterized in that, Includes the following steps: S1: Different tasks have different and unknown arrival rate and task quantity parameters. At the time of completion of each task, based on the prior distribution of arrival rate and task quantity parameters of each task and the new information of the task, the posterior distribution of arrival rate and task quantity parameters is obtained, and the hyperparameter update rule and sufficient statistics characterizing the randomness, dynamics and heterogeneity of the task are obtained. S2: The degradation process of the system is described by a continuous-time Markov process, and the system mode switching strategy and load adjustment strategy are designed. S3: Based on the sufficient statistics obtained in S1 and the degree of system degradation, construct optimization models for mode switching strategy and load adjustment strategy with the goal of minimizing the expected total cost based on the framework of a semi-Markov decision process. S4: Use an improved algorithm to solve the optimization model described in S3 to obtain the optimal mode switching strategy and load adjustment strategy when each task is completed.
2. The method for optimizing the operation and maintenance strategy of a task-critical system performing heterogeneous tasks according to claim 1, characterized in that, The specific process of step S1 is as follows: S1.1: Define arrival rate parameter and task quantity parameters prior distribution The arrival of the task follows a homogeneous Poisson process, and the arrival rate is determined by unknown parameters. This indicates that the workload of each task follows a parameter containing unknown parameters. The exponential distribution; the unknown parameters and Treat it as a random variable, and use and express; set up The prior distribution is the shape parameter and scale parameters The Gamma distribution has a prior density distribution of . , The prior distribution is the shape parameter and scale parameters The Gamma distribution has a prior density distribution of . ; S1.2: Obtaining hyperparameter update rules and sufficient statistics using Bayes' theorem No. The task is up to the first The observation data for arriving at the task between the completion times of each task are: ,in, for The observed values, Indicates completion of the first The task is up to the first The random duration between tasks is determined by the task's waiting time. and task completion time It consists of two parts; for The observed values, Indicates completion of the first The task is up to the first The number of random tasks arriving between tasks; for The observed values, of which Indicates the number of times reached The workload of each task; in addition, let Indicates completion of the first The task is up to the first The cumulative task volume during each task period, its observed value is... ; Using the first When each task is completed and The joint posterior density function as the first The prior density function when each task is completed, i.e. Combined with the first The task is up to the first Observational data arriving at the task between the completion times of each task. Based on Bayes' theorem, the first... When each task is completed and The joint posterior density function; and The joint posterior density function is equal to two Gamma distributions. and The product of density functions leads to the following hyperparameter update rule: , , and ; Therefore, given the initial hyperparameters and When the first The information about the arrival of a task upon completion is equivalent to a sufficient statistic. ,in , and They represent the first The number of tasks completed, age, and cumulative number of tasks completed when a task is finished.
3. The method for optimizing the operation and maintenance strategy of a task-critical system performing heterogeneous tasks according to claim 2, characterized in that, In step S2, the system operation process is characterized as follows: the system needs to execute... One task, complete The reward for each task is Using a continuous-time Markov process with non-negative, stationary, and independent increments. To describe the degradation process of the system, time interval Internal degradation increment The distribution function is expressed as Let the first Maintenance will be performed upon completion of each task; if the system malfunctions, corrective maintenance will be carried out at a cost of [amount missing]. Otherwise, preventative maintenance will be carried out, at a cost of .
4. The method for optimizing the operation and maintenance strategy of a task-critical system performing heterogeneous tasks according to claim 3, characterized in that, In step S2, the mode switching strategy is as follows: the system includes an online mode and an offline mode, and switches between the online and offline modes. When in online mode, the system continuously degrades and incurs operating costs, denoted as the operating cost per unit time. ; When the system has no pending tasks at the time of task completion, it switches to offline mode or remains in online mode. In offline mode, the system does not degrade, but switching back to online mode when a new task arrives incurs switching costs. When the system still has tasks to process at the time of task completion, the system continues to remain online and execute tasks; to balance operating costs and switching costs, a decision must be made at the decision point whether to switch modes. The load adjustment strategy is as follows: the system has an operating load-degradation dependency, that is, when a higher operating load is used, the system will experience accelerated degradation; in order to balance benefits and maintenance costs, the operating load level to be used in the next task execution phase needs to be determined at the decision point.
5. The method for optimizing the operation and maintenance strategy of a task-critical system performing heterogeneous tasks according to claim 4, characterized in that, In step S3, based on the framework of a semi-Markov decision process and incorporating elements such as the state set, action set, and value function, an optimization model for the mode switching strategy and load adjustment strategy is constructed with the objective of minimizing the expected total cost, as detailed below: Status: Completed When performing a task, the state of the semi-Markov decision process is used as... It means that, among them, Represents the completion of the first The degree of system degradation during each task; when the degree of degradation exceeds a fixed failure threshold. When this happens, the system transitions to a failure state. ; Action set: each using and This indicates switching the system to online or offline mode; in the... When a task is completed, based on the state... ,action Represented as ,in ,and ,in Indicates the system load level. and These are the minimum load level and the maximum load level, respectively; specifically, if the first... There are remaining tasks when there are 1 task, that is... By default, the mode is switched to online mode, i.e. ; Value function: Let and Respectively represent and and Related single-stage expected cost; Define value function For the stage Based on state The minimum expected total cost, i.e., the optimization model, is expressed as: in, Representing a Markov process Regarding thresholds When the first arrival, The distribution function is ; The boundary conditions for the optimization model are: 。 6. The method for optimizing the operation and maintenance strategy of a task-critical system performing heterogeneous tasks according to claim 5, characterized in that, In step S4, the improved algorithm is used to solve the optimization model described in S3, as follows; For all possible states in the Nth stage The corresponding value function is initialized, and the optimal mode switching strategy and load adjustment strategy for all states from the Nth stage to the 1st stage are solved using a backward iterative algorithm. In the solution process, the structural properties of the optimal strategy are used to solve the problem, namely: the degree of system degradation, the number of tasks reached, and the system age each have thresholds. For the optimal mode switching strategy of all states in the current stage that exceed any threshold, the optimal mode switching strategy is the opposite of the optimal mode switching strategy of the states in the current stage that do not reach the threshold. For the optimal load level strategy of a certain state in the current stage, the number of tasks arriving is compared with the number of tasks arriving in the state corresponding to the known optimal load level strategy, or the cumulative number of tasks arriving is compared with the cumulative number of tasks arriving in the state corresponding to the known optimal load level strategy. Based on the comparison result and the known optimal load level strategy, the optimal load level of a certain state in the current stage is adjusted.
7. The method for optimizing the operation and maintenance strategy of a task-critical system performing heterogeneous tasks according to claim 6, characterized in that, The structural properties of the optimal strategy include the control limit properties of the optimal mode switching strategy and the threshold monotonicity of the optimal load level strategy. The control limit property of the optimal mode switching strategy is: under a fixed load control strategy, there exists a threshold. ,when Switching to offline mode is optimal when... Switching to online mode is optimal at this time; a threshold exists. ,when Switching to online mode is optimal when... When this happens, switching to offline mode is optimal; [the text abruptly ends here, likely due to an incomplete sentence or missing information.] ,when Switching to offline mode is optimal when... Switching to online mode is optimal at this time; The threshold monotonicity of the optimal load level strategy is: under a fixed mode switching strategy, for a fixed... status Optimal load level about It is non-incremental; for fixed... status Optimal load level about It is not subtractive.