Termination strategy optimization method considering staged task system

By adopting Wiener process modeling and reverse dynamic programming algorithm in the phased task system, the task termination strategy is optimized, the termination problem of multiple continuous subtasks is solved, and costs are reduced and decision-making accuracy is improved.

CN120762864AActive Publication Date: 2025-10-10HEFEI UNIV OF TECH

Patent Information

Application Number
CN202511257102.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-10-10
Estimated Expiration
2045-09-04

AI Technical Summary

Technical Problem

The existing technology lacks a termination strategy optimization method for a staged task system consisting of multiple continuous subtasks, especially fails to effectively solve the termination problem of multiple stage tasks.

Method used

The Wiener process is used to model the degradation state of the system, and the task termination or continuation action of each decision point is optimized based on the decision interval and state transition probability through the reverse dynamic programming algorithm to generate a set of task termination strategies.

Benefits of technology

It effectively reduces task costs, improves the optimization efficiency and accuracy of task termination strategies, adapts to system state changes, and solves the gap in termination strategies for phased tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120762864A_ABST
    Figure CN120762864A_ABST
Patent Text Reader

Abstract

The invention provides a termination strategy optimization method considering a staged task system, and relates to the field of task termination strategies. Firstly, for each stage of a staged task, it is defined that the degradation state of a system obeys a Wiener process, and the number of decision points of each stage is obtained; secondly, converting the degradation state of each stage into a discrete state model to solve the state transition probability of the system; secondly, solving the task termination cost and the task continuation cost of each decision point in different system states, and selecting a smaller value as a corresponding expected cost; and finally, selecting task termination or task continuation action based on the current decision point and the expected cost in the system state by adopting a reverse dynamic programming algorithm to obtain a task termination strategy set. According to the optimization method provided by the invention, the blank that a staged task termination strategy is not considered in related technologies is filled up, and meanwhile, the optimization problem is expressed by adopting a random dynamic programming framework, so that the cost is effectively reduced compared with a fixed threshold strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of task termination strategies, and in particular to a termination strategy optimization method considering a phased task system. Background Art

[0002] Mission termination strategies play a crucial role in fields such as aerospace, submarines, and mechanical engineering. Specifically, when a safety-critical system is performing a mission within a specified timeframe, system survival takes precedence over mission completion for safety or economic reasons. When the risk of system failure becomes unacceptable, the mission can be aborted, and rescue procedures can be immediately activated to ensure the system remains intact.

[0003] In related technologies, research on task termination strategies can be roughly divided into two categories: One type is the termination strategy based on the number of shocks. That is, after the system is subjected to a certain number of external shocks, the system will suffer corresponding damage. The literature (Levitin G, Finkelstein M. Optimal mission abort policyfor systems operating in a random environment[J]. Risk Analysis, 2018, 38(4):795-803.) uses external shocks following a non-homogeneous Poisson process to simulate the impact of a random environment on the system, takes the number of shocks as a decision parameter, and evaluates the mission termination strategy in a random environment.

[0004] Another type is the termination strategy based on random processes, which uses Wiener, Gamma and other random processes and the obtained comprehensive information to guide the termination of the task. Literature (Zhao X, Sun J, Qiu Q, et al. Optimal inspection and mission abort policies for systems subject to degradation[J]. European Journal of Operational Research, 2021, 292(2): 610-621.) uses Gamma process to model the continuous degradation system, and analyzes the joint optimization problem of the optimal detection and dynamic task termination strategy. Literature (Yang L, Chen Y, Qiu Q, et al. Risk control of mission-critical systems: Abort decision-makings integrating health and age conditions[J]. IEEE Transactions on Industrial Informatics, 2022, 18(10): 6887-6894.) designs the task termination strategy based on the degradation degree and age information of the monitored system, and studies the optimal risk control strategy of the random degradation system through the Wiener process model.

[0005] It can be seen that the existing research on phased task termination strategy is less. In addition, although Chinese patent application CN119473779A uses "phased" Markov chain to depict the degradation process of the system and proposes a task termination strategy, the "phased" is for the "multi-state" of system degradation, not the phased of task, so the termination strategy is still for single-phase task. SUMMARY

[0006] (I) Technical problems solved In view of the shortcomings of the prior art, the present application provides a termination strategy optimization method considering a phased task system, which solves the termination problem of the overall task composed of multiple continuous sub-tasks (or phased tasks).

[0007] (II) Technical solutions In order to achieve the above purpose, the present application is realized by the following technical solutions: A termination strategy optimization method considering a phased task system, assuming that the system executes a phased task, the termination strategy optimization method comprising: for each stage of the staged task, define the degradation state of the system to be subject to a Wiener process; and based on a preset decision interval, obtain the number of decision points for each stage; convert the degradation state of each stage into a discrete state model, and solve the state transition probability of the system in the discrete state; based on the state transition probability, solve the task termination cost and task continuation cost of each decision point under different system states, and select the smaller value as the corresponding expected cost; wherein if the decision point is located at the boundary of adjacent stages, the expected cost at least includes the reward for completing the stage to which the decision point belongs; starting from the end point of the staged task, traverse each decision point using a backward dynamic programming algorithm, select a task termination or task continuation action based on the expected cost at the current decision point and system state, and save it to a task termination strategy set until the starting point of the staged task is reached.

[0008] Preferably, the staged task is composed of M stages, for the mth stage, the degradation state of the system is defined as subject to a Wiener process with drift parameter and diffusion coefficient , wherein represents being in the mth stage, represents the starting time of the mth stage, represents the end time of the mth stage.

[0009] Preferably, the decision interval is defined as , then the number of decision points for the mth stage is , wherein is an integer, represents the upward rounding symbol.

[0010] Preferably, the conversion of the degradation state of each stage into a discrete state model and the solving of the state transition probability of the system in the discrete state comprises: define a finite degradation state set , wherein 0 represents the initial health state of the system, and I represents the system failure state; state partition rule: set the discrete interval of the degradation state as d, when the degradation state of the system falls within the interval [id, (i+1)d], it is determined that the system is in state i; when the degradation state of the system is greater than or equal to Id, the system enters the failure state I; the midpoint value (i+0.5)d in the interval is used to represent each state i; during the mth stage, the system transfers from the current state i to other state k, based on the characteristics of the Wiener process, the transition probability is calculated, and classified as follows: 1) if i≠I and k≠I 2) if i≠I and k=I 3) if i=I where, Pm(i,k) denotes the probability of the system transitioning from state i to state k at the mth stage, after time t, is the cumulative distribution function of the standard normal distribution.

[0011] Preferably, the solving the task termination cost and the task continuation cost of each decision point in different system states based on the state transition probability, and selecting the smaller value as the corresponding expected cost, comprises: defining the state space of the system as ; wherein n represents the cumulative number of decision points; and defining the action space as {A, C}, wherein A represents terminating the task, and C represents continuing the task, and letting V(n, i) represent the expected cost at state (n, i): wherein min represents the minimization function, A(n, i) represents the cost of terminating the task at state (n, i), and C(n, i) represents the cost of continuing the task at state (n, i); The calculation of V(n, i) needs to consider the following cases: 1) if , it indicates that the phased task has been successfully completed, and the expected cost is: wherein, Rm represents the reward of completing the Mth stage task; 2) if , it indicates that the system state is at (n, i) within the mth stage of the executed task, and the cost C(n, i) of continuing the task is composed of the continuation failure cost CI and the continuation expected cost CE: wherein, Rm represents the phased task failure cost, AF represents the system failure cost; V(n+1, k) represents the future expected cost of the system state at (n+1, k); Pm(i,k) denotes the probability of the system transitioning from state i to state k at the mth stage, after time t, ; and The cost A(n, i) of terminating the task when the system state is at (n, i) is composed of the termination operation expected cost AF: wherein, denotes the probability of successfully rescuing the system after terminating the mission at the mth stage; 3) if denotes that the mission has been successfully completed at the mth stage, the cost C(n,i) of continuing the mission when the system state is at (n,i) is composed of the cost of continuing failure CI, the cost of continuing execution CE, and the stage reward PR: wherein, denotes the reward of completing the mth stage; the cost A(n,i) of terminating the mission when the system state is at (n,i) is composed of the expected cost of terminating operation AF and the stage reward PR: (13).

[0012] Preferably, the reverse dynamic programming algorithm is used to traverse each decision point from the end point of the staged mission, the mission termination or mission continuation action is selected based on the expected cost at the current decision point and the system state, and is saved to the mission termination strategy set until the starting point of the staged mission is reached, comprising: STEP0, initialize parameters including , and set denotes the mission termination strategy set when the system state is at (n,i), wherein denotes the continuation action, denotes the termination action; STEP1, calculate the total number of decision points ; STEP2, for all , calculate V(n,i) by formula (5); STEP3, let n=N-1; STEP4, if n<0, execute STEP12, otherwise execute STEP5; STEP5, set i=0; STEP6, if i>I-1, jump to STEP11, otherwise execute STEP7; STEP7, if there exists m∈{1,2,…,M-1} such that , calculate C(n,i) and A(n,i) by formulas (11) and (13) respectively, otherwise calculate C(n,i) and A(n,i) by formulas (6) and (9); STEP8, calculate , if V(n,i)=C(n,i), set , i = i + 1, execute STEP6; otherwise , execute STEP9; STEP9, for all k = i + 1 to I - 1, if there exists , such that , then calculate A(n, k) by formula (13), otherwise calculate A(n, k) by formula (9); STEP10, set V(n, k) = A(n, k), , turn to STEP11; STEP11, set n = n - 1, return to STEP4; STEP12, end the algorithm, output the task termination strategy set.

[0013] A termination strategy optimization system considering a staged task system, assuming that a system performs a staged task, the termination strategy optimization system comprising: a definition module, configured to define that a degradation state of the system is subject to a Wiener process for each stage of the staged task, and obtain a number of decision points of each stage based on a preset decision interval; a solution module, configured to convert the degradation state of each stage into a discrete state model, and solve a state transition probability of the system in the discrete state; and configured to solve a task termination cost and a task continuation cost of each decision point in different system states based on the state transition probability, and select a smaller value as a corresponding expected cost; wherein if the decision point is located at a boundary between adjacent stages, the expected cost at least includes a reward for completing the stage to which the decision point belongs; a decision module, configured to traverse each decision point from an end point of the staged task to a start point of the staged task by using a backward dynamic programming algorithm, select a task termination or task continuation action based on the expected cost in the current decision point and system state, and save to a task termination strategy set until the start point of the staged task is reached.

[0014] A storage medium storing a computer program for termination strategy optimization considering a staged task system, wherein the computer program causes a computer to execute the termination strategy optimization method as described above.

[0015] An electronic device comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the programs comprise a program for executing the termination strategy optimization method as described above.

[0016] (Three) beneficial effects The present invention provides a termination strategy optimization method for a phased task system. Compared with the existing technology, it has the following advantages: In the present invention, first, for each stage of a phased task, the degradation state of the system is defined as obeying a Wiener process, and the number of decision points in each stage is obtained. Secondly, the degradation state of each stage is converted into a discrete state model to solve the state transition probability of the system. Next, the task termination cost and task continuation cost of each decision point under different system states are solved, and the smaller value is selected as the corresponding expected cost. Finally, an inverse dynamic programming algorithm is used to select task termination or task continuation actions based on the current decision point and the expected cost under the system state, thereby obtaining a set of task termination strategies. The optimization method proposed in the present invention fills the gap in the related art that does not consider phased task termination strategies. At the same time, a stochastic dynamic programming framework is used to formulate the optimization problem, effectively reducing costs compared to fixed threshold strategies. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 A block diagram of a termination strategy optimization method considering a phased task system provided by an embodiment of the present invention; Figure 2 A flowchart of a reverse dynamic programming algorithm provided by an embodiment of the present invention; Figure 3 A schematic diagram of a termination decision process for a two-stage task provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention are clearly and completely described. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0020] The embodiment of the present application solves the termination problem of an overall task consisting of multiple consecutive subtasks (or stage tasks) by providing a termination strategy optimization method that takes into account a staged task system.

[0021] The technical solution in the embodiments of the present application is to solve the above technical problems, and the overall idea is as follows: Some related research, such as Chinese patent application CN119473779A, claims to have studied phased tasks. However, the concept of "phased" here essentially refers to a system degradation state, rather than multiple phased tasks (i.e., multiple consecutive subtasks). Consequently, it fails to address the termination of an overall task consisting of these multiple consecutive subtasks (or phased tasks). In response, embodiments of the present invention propose an optimization method for a phased task termination strategy. This method discretizes the task execution process into multiple time points. At each time point, a decision is made to continue or terminate the task based on the degradation state, thereby achieving an effective balance between task reliability and system survivability.

[0022] In addition, the embodiment of the present invention adopts a stochastic dynamic programming framework to express the optimization problem, and adopts a reverse induction method to find the optimal solution.

[0023] In summary, the embodiments of the present invention uniformly model and optimize the dual dynamic interaction between the system degradation state and the staged task process, and propose a termination strategy generation method that can be oriented to staged tasks and adaptive to the system state.

[0024] In order to better understand the above technical solution, the above technical solution will be described in detail below with reference to the accompanying drawings and specific implementation methods.

[0025] Example 1: like Figure 1 As shown, an embodiment of the present invention provides a termination strategy optimization method considering a phased task system. Assuming that the system executes a phased task, the termination strategy optimization method includes: S1. For each stage of the phased task, define the degradation state of the system to obey a Wiener process; and obtain the number of decision points in each stage based on a preset decision interval; S2. Convert the degradation state of each stage into a discrete state model and solve the state transition probability of the system in the discrete state; S3. Based on the state transition probabilities, solve the task termination cost and task continuation cost for each decision point under different system states, and select the smaller value as the corresponding expected cost; wherein, if the decision point is located at the boundary of adjacent stages, the expected cost at least includes the reward for completing the stage to which the decision point belongs; S4. Starting from the end point of the phased task, a reverse dynamic programming algorithm is used to traverse each decision point. Based on the current decision point and the expected cost under the system state, task termination or task continuation action is selected and saved to the task termination strategy set until the starting point of the phased task is reached.

[0026] The optimization method provided in the embodiment of the application fills the blank in the related art that the staged task termination strategy is not considered, and simultaneously uses a stochastic dynamic programming framework to express the optimization problem, thereby effectively reducing the cost compared with the fixed threshold strategy.

[0027] Next, each step of the above scheme will be described in detail: It should be noted that, in the embodiment of the application, it is assumed that the system performs a staged task, the task is composed of M stages, each stage has a corresponding execution time, and only when the current stage is completed, the next stage will be started.

[0028] First, the degradation state of the staged task system to which the application is directed is introduced, and the introduction of step S1 is referred to for details: In step S1, for each stage of the staged task, the degradation state of the system is defined to follow a Wiener process; and based on a preset decision interval, the number of decision points of each stage is obtained.

[0029] A Markov chain is a kind of stochastic process, which is widely used in stochastic process modeling and analysis in natural science, social science, engineering technology and other fields. If a process is a Markov process, and the change in unit time follows a normal distribution with an expectation of 0 and a variance of 1, the process is a Wiener process.

[0030] Considering that when the system performs the staged task, the damage degree of the system in each stage is different because different tasks are performed in each stage. Therefore, in this step, the degradation state of the system is defined to follow a Wiener process with a drift parameter and a diffusion coefficient , wherein represents being in the mth stage, represents the start time of the mth stage, represents the end time of the mth stage.

[0031] In addition, the decision interval is defined as , that is, every time unit, the decision maker needs to decide whether to continue the task, and the number of decision points of the mth stage is , wherein is an integer, represents the upward rounding symbol.

[0032] Secondly, the mathematical model of the task termination problem is introduced, and the introduction of steps S2 and S3 is referred to for details: In step S2, the degradation state of each stage is converted into a discrete state model, and the state transition probability of the system is solved in the discrete state.

[0033] To facilitate practical application and decision making, this step converts the degradation state of each stage into a discrete state model. Specifically, a finite set of degradation states is defined: , where 0 represents the initial healthy state of the system and I represents the fault state of the system.

[0034] State division rule: Set the discrete interval of degradation state to d. When the system degradation state falls within the interval [id, (i+1)d], the system is judged to be in state i; when the system degradation state is greater than or equal to Id, the system enters the fault state I; each state i is represented by the midpoint value of the interval (i+0.5)d.

[0035] Next, in the discrete state, solve the state transition probability of the system. The relevant steps are as follows: During the mth phase, i.e., the time interval In the system, the system transfers from the current state i to other states k. The transition probability is calculated based on the characteristics of the Wiener process and is classified as follows: 1) If i≠I and k≠I 2) If i≠I and k=I 3) If i=I in, represents the probability that the system transitions from state i to state k after time t at the mth stage, d represents the discrete interval, is the cumulative distribution function of the standard normal distribution.

[0036] In step S3, based on the state transition probability, the task termination cost and task continuation cost of each decision point under different system states are solved, and the smaller value is selected as the corresponding expected cost; if the decision point is located at the boundary of adjacent stages, then the expected cost at least includes the reward for completing the stage to which the decision point belongs.

[0037] Based on the state transition probability of the system solved in step S2, this step specifically includes: Define the state space of the system as ; where n represents the cumulative number of decision points; and define the action space as {A, C}, where A represents the termination task and C represents the continuation task, and let V(n,i) represent the expected cost at state (n,i): Where min represents the minimization function, A(n,i) represents the cost of terminating the task at state (n,i), and C(n,i) represents the cost of continuing the task at state (n,i). The following situations need to be considered when calculating V(n,i): 1) If , indicating that the phased task has been successfully completed, the expected cost is: in, Represents the reward for completing the Mth stage task; 2) If , which means that in the task of the mth stage of execution, when the system state is (n, i), the cost of continuing the task C(n, i) is composed of the failure cost CI and the expected cost CE of continuing execution: in, represents the failure cost of the phased task, represents the system failure cost; V(n+1,k) represents the future expected cost of the system state being (n+1,k); Indicates that at the mth stage, after time , the probability that the system transitions from state i to state k; The cost of terminating a task when the system state is (n,i) is composed of the expected cost AF of the termination operation: in, represents the probability of successfully rescuing the system after terminating the mission at the mth stage; 3) If , indicating that the task at stage m has been successfully completed and the cost of continuing the task when the system state is (n, i) is composed of the failure cost CI, the expected cost of continuing execution CE, and the stage reward PR: in, Represents the reward for completing the mth stage; The cost of terminating a task when the system state is (n,i) is composed of the expected cost AF of the termination operation and the stage reward PR: (13).

[0038] Finally, the reverse dynamic programming algorithm proposed in the embodiment of the present invention is introduced. For details, see the content of step S4: In step S4, starting from the end of the staged task, a backward dynamic programming algorithm is used to traverse each decision point, select the task termination or task continuation action based on the expected cost at the current decision point and system state, and save to the task termination strategy set until the beginning of the staged task is reached.

[0039] As shown in Figure 2 , this step specifically includes: STEP0, initialize parameters including , and set , which represents the task termination strategy set when the system state is (n, i), where represents the continuation action, represents the termination action; STEP1, calculate the total number of decision points ; STEP2, for all , calculate V(n, i) by formula (5); STEP3, set n = N-1; STEP4, if n < 0, execute STEP12, otherwise execute STEP5; STEP5, set i = 0; STEP6, if i > I-1, jump to STEP11, otherwise execute STEP7; STEP7, if there exists m ∈ {1, 2, …, M-1} such that , calculate C(n, i) and A(n, i) by formulas (11), (13) respectively, otherwise calculate C(n, i) and A(n, i) by formulas (6), (9); STEP8, calculate , if V(n, i) = C(n, i), set , i = i+1, execute STEP6; otherwise , execute STEP9; STEP9, for all k = i+1 to I-1, if there exists such that , calculate A(n, k) by formula (13), otherwise calculate A(n, k) by formula (9); STEP10, set V(n, k) = A(n, k), , go to STEP11; STEP11, set n = n-1, return to STEP4; STEP12, end the algorithm, output the task termination strategy set.

[0040] It can be understood that the algorithm fundamentally solves the optimization problem of the stage characteristics and the cost structure difference between different stages in the real complex task scene by introducing a staged task modeling framework.

[0041] It is particularly pointed out that the core innovation of the above algorithm is embodied in the design of STEP7 to STEP10, which skillfully distinguishes two key decision moments: intra-stage decision points and inter-stage transition points (i.e. the boundary points of adjacent stages); Specifically: When the decision point n is located at the end of a certain stage m and the start of the next stage m+1 (i.e. ), the algorithm uses special formulas (11), (13) to calculate C(n,i) and A(n,i), which fully considers the particularity of the stage transition moment, including the possible risk level change and the change of decision constraints between different stages; When the decision point is in the stage (i.e. ), the regular formulas (6), (9) are used for calculation. This double calculation mechanism enables the algorithm to accurately capture the dynamic complexity in the execution process of the staged task, avoiding the decision bias and suboptimal solution problem when facing tasks with obvious stage characteristics. In addition, STEP9 introduces an intelligent optimization mechanism based on state degradation monotonicity: once the termination strategy is determined to be the optimal choice in a certain degraded state i, the algorithm immediately uses the monotonicity feature of state degradation to directly set all states k (k=i+1 to I-1) worse than the current state as the termination strategy without repeating the calculation of the continuation cost of these states. This "one-time decision, cascading application" strategy significantly improves the calculation efficiency while ensuring the logical consistency of the decision. This design not only avoids a large amount of redundant calculation, but more importantly, it reflects the inherent logic of state-dependent decision-making in staged complex tasks, enabling the algorithm to generate an efficient staged task termination strategy while ensuring global optimality, with significant computational advantages and decision accuracy improvement.

[0042] To further illustrate the technical solutions provided by the embodiments of the present application, the following describes the implementation of the staged task termination strategy in combination with an actual scene implementation case.

[0043] This case takes the unmanned aerial vehicle emergency material distribution task in the urban environment as an example to clearly demonstrate the application of the staged task termination strategy. The task is clearly defined as containing two logically distinct, sequentially executed and target-independent stages: Stage 1: Delivery stage, the unmanned aerial vehicle takes off from the base and performs material delivery to the designated target area; The task execution time of this stage is fixed at 10 units of time, i.e. T1=10, the degradation of the unmanned aerial vehicle system (such as battery, motor, navigation system) is simulated using a Wiener process in this stage, the drift coefficient =0.6, diffusion coefficient =0.8, this stage has higher energy consumption (flight load), and the degradation rate is relatively fast (reflected in the higher drift coefficient and diffusion coefficient).

[0044] Stage two: return stage, after the unmanned aerial vehicle completes the material delivery, it flies back to the base. The task execution time of this stage is also fixed at 10 units of time, that is, T2=20, the degradation of the system in stage two is simulated by using a new Wiener process, but the initial state inherits the degradation state at the end of stage one, the drift coefficient of the process =0.4, diffusion coefficient =0.6, this stage has lighter load (the material has been delivered), and the expected degradation rate is relatively slow (reflected in the lower drift coefficient and diffusion coefficient).

[0045] In addition, the discrete interval d=0.5, the decision interval =1, the phased task failure cost and the system failure cost are 500 and 1000 respectively, the rewards for completing the first stage and the second stage are 300 and 150 respectively, and the rescue success probabilities after the first stage and the second stage terminate the task are 0.99 and 0.93 respectively.

[0046] As Figure 3 shown, Figure 3 the termination decision process of the two-stage task (the optimal task termination strategy) is shown, the dashed line divides the task into two stages, the upper solid line indicates terminating the task, and the lower indicates continuing the task.

[0047] Table 1 and Table 2 show the comparison between the proposed termination strategy and the fixed threshold termination strategy under different costs, and it can be seen that, regardless of the change of the task failure cost or the change of the system damage cost, the proposed strategy has certain advantages compared with the fixed threshold strategy, and can reduce the cost.

[0048] Table 1 Comparison of strategies under different phased task failure costs =1000) Table 2 Comparison of strategies under different system failure costs =500) Embodiment 2 The embodiment of the present application provides a termination strategy optimization system considering a phased task system, assuming that a system performs a phased task, the termination strategy optimization system comprises: a definition module configured to define, for each stage of the staged task, a degradation state of the system subject to a Wiener process, and obtain a number of decision points for each stage based on a preset decision interval; a solution module configured to convert the degradation state of each stage into a discrete state model, and solve a state transition probability of the system in the discrete state; and configured to solve a task termination cost and a task continuation cost of each decision point in different system states based on the state transition probability, and select a smaller value as a corresponding expected cost; wherein if the decision point is located at a boundary between adjacent stages, the expected cost at least includes a reward for completing the stage to which the decision point belongs. a decision module configured to traverse each decision point from an end point of the staged task to a start point of the staged task using a backward dynamic programming algorithm, select a task termination action or a task continuation action based on the expected cost at the current decision point and the system state, and save the selected action to a task termination strategy set.

[0049] Embodiment 3: The embodiment of the present application provides a storage medium storing a computer program for optimizing a termination strategy of a staged task system, wherein the computer program causes a computer to execute the termination strategy optimization method as described in Embodiment 1.

[0050] Embodiment 4: The embodiment of the present application provides an electronic device, comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the program comprises a program for executing the termination strategy optimization method as described in Embodiment 1.

[0051] It can be understood that the termination strategy optimization system, the storage medium and the electronic device provided by the embodiment of the present application correspond to the termination strategy optimization method provided by the embodiment of the present application, and the explanation, examples and beneficial effects of the related content can refer to the corresponding part in the termination strategy optimization method, which will not be repeated here.

[0052] In summary, compared with the prior art, the present application has the following beneficial effects: 1. The embodiment of the present application provides a staged task termination strategy optimization method, which fills the gap in the current task termination strategy which does not consider the staged task.

[0053] 2、The embodiment of the present application adopts a random dynamic programming framework to express the optimization problem, and adopts a reverse induction method to solve it. Compared with a traditional fixed threshold strategy, the above experimental results show that the termination strategy proposed in the embodiment of the present application has great advantages in cost.

[0054] It should be noted that, in this document, relational terms such as first and second and the like can be used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0055] The above embodiments are only used to illustrate the technical solutions of the present application, not limit the present application; even though the present application has been described in detail with the foregoing embodiments, it should be understood by those skilled in the art that: the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced equivalently; and the modification or replacement does not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A termination strategy optimization method considering a phased task system, characterized in that: Assuming that the system executes a phased task, the termination strategy optimization method includes: For each stage of the phased task, defining the degradation state of the system to obey a Wiener process; and obtaining the number of decision points in each stage based on a preset decision interval; Convert the degradation state of each stage into a discrete state model, and solve the state transition probability of the system under the discrete state; Based on the state transition probabilities, the task termination cost and task continuation cost for each decision point under different system states are calculated, and the smaller value is selected as the corresponding expected cost. If the decision point is located at the boundary of adjacent stages, the expected cost at least includes the reward for completing the stage to which the decision point belongs. Starting from the end point of the phased task, a reverse dynamic programming algorithm is used to traverse each decision point. Based on the current decision point and the expected cost under the system state, the task termination or task continuation action is selected and saved to the task termination strategy set until the starting point of the phased task is reached.

2. The termination strategy optimization method according to claim 1, characterized in that: The phased task consists of M phases. For the mth phase, the degradation state of the system is defined as The drift parameter is , the diffusion coefficient is The Wiener process, where Indicates that it is in the mth stage, represents the starting time of the mth stage, Indicates the end time of the mth stage.

3. The termination strategy optimization method according to claim 2, characterized in that: Define the decision interval as , then the number of decision points in the mth stage is ,in is an integer, Indicates the round-up symbol.

4. The termination strategy optimization method according to claim 3, wherein: The process of converting the degradation state of each stage into a discrete state model and solving the state transition probability of the system under the discrete state includes: Define a finite set of degenerate states , where 0 represents the initial healthy state of the system and I represents the faulty state of the system; State division rule: Set the discrete interval of degradation state to d. When the system degradation state falls within the interval [id, (i+1)d], the system is judged to be in state i. When the system degradation state is greater than or equal to Id, the system enters the fault state I. Each state i is represented by the midpoint value of the interval (i+0.5)d. During the mth phase, the system transitions from the current state i to another state k. The transition probability is calculated based on the characteristics of the Wiener process and is classified as follows: 1) If i≠I and k≠I 2) If i≠I and k=I 3) If i=I in, represents the probability that the system transitions from state i to state k after time t at the mth stage, d represents the discrete interval, is the cumulative distribution function of the standard normal distribution.

5. The termination strategy optimization method according to claim 4, characterized in that: The method of solving the task termination cost and task continuation cost of each decision point under different system states based on the state transition probability and selecting the smaller value as the corresponding expected cost includes: Define the state space of the system as ; where n represents the cumulative number of decision points; and define the action space as {A, C}, where A represents the termination task and C represents the continuation task, and let V(n,i) represent the expected cost at state (n,i): Where min represents the minimization function, A(n,i) represents the cost of terminating the task at state (n,i), and C(n,i) represents the cost of continuing the task at state (n,i). The following situations need to be considered when calculating V(n,i): 1) If , indicating that the phased task has been successfully completed, the expected cost is: in, Represents the reward for completing the Mth stage task; 2) If , which means that in the task of the mth stage of execution, when the system state is (n, i), the cost of continuing the task C(n, i) is composed of the failure cost CI and the expected cost CE of continuing execution: in, represents the failure cost of the phased task, represents the system failure cost; V(n+1,k) represents the future expected cost of the system state being (n+1,k); Indicates that at the mth stage, after time , the probability that the system transitions from state i to state k; The cost of terminating a task when the system state is (n,i) is composed of the expected cost AF of the termination operation: in, represents the probability of successfully rescuing the system after terminating the mission at the mth stage; 3) If , indicating that the task at stage m has been successfully completed and the cost of continuing the task when the system state is (n, i) is composed of the failure cost CI, the expected cost of continuing execution CE, and the stage reward PR: in, Represents the reward for completing the mth stage; The cost of terminating a task when the system state is (n,i) is composed of the expected cost AF of the termination operation and the stage reward PR: (13)。 6. The termination strategy optimization method according to claim 5, characterized in that: Starting from the end point of the phased task, a reverse dynamic programming algorithm is used to traverse each decision point, and based on the current decision point and the expected cost under the system state, a task termination or task continuation action is selected and saved to the task termination strategy set until the starting point of the phased task is reached, including: STEP0, initialization parameters include , and set represents the set of task termination strategies when the system state is (n,i), where Indicates continuing the action. Indicates the termination of an action; STEP 1. Calculate the total number of decision points ; STEP2, for all , calculate V(n,i) by formula (5); STEP 3, let n = N-1; STEP4. If n < 0, execute STEP12, otherwise execute STEP5. STEP 5. Set i=0; STEP6. If i>I-1, jump to STEP11, otherwise execute STEP7; STEP7. If there exists m∈{1,2,…,M-1}, such that , then calculate C(n,i) and A(n,i) respectively by formula (11) and (13), otherwise calculate C(n,i) and A(n,i) by formula (6) and (9); STEP 8. Calculation , if V(n,i)=C(n,i), then set , i=i+1, execute STEP6; otherwise , execute STEP9; STEP9, for all k=i+1 to I-1, if there exists , making , then A(n,k) is calculated by formula (13), otherwise A(n,k) is calculated by formula (9); STEP10, set V(n,k)=A(n,k), , go to STEP11; STEP11. Set n=n-1 and return to STEP4. STEP 12: End the algorithm and output the task termination strategy set.

7. A termination strategy optimization system considering a phased task system, characterized in that: Assuming that the system performs a phased task, the termination strategy optimization system includes: A definition module is configured to define, for each stage of the phased task, that the degradation state of the system obeys a Wiener process; and obtain the number of decision points in each stage based on a preset decision interval; The solution module is used to convert the degradation state of each stage into a discrete state model and solve the state transition probability of the system under the discrete state; and for solving, based on the state transition probabilities, the task termination cost and task continuation cost for each decision point under different system states, and selecting the smaller value as the corresponding expected cost; wherein if the decision point is located at the boundary of adjacent stages, the expected cost at least includes the reward for completing the stage to which the decision point belongs; The decision module starts from the end point of the phased task and uses the reverse dynamic programming algorithm to traverse each decision point. Based on the current decision point and the expected cost under the system state, it selects the task termination or task continuation action and saves it to the task termination strategy set until it reaches the starting point of the phased task.

8. A storage medium, characterized in that: It stores a computer program for optimizing a termination strategy considering a phased task system, wherein the computer program enables a computer to execute the termination strategy optimization method according to any one of claims 1 to 6.

9. An electronic device, characterized in that: include: one or more processors; Memory; and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, the programs including a method for executing the termination strategy optimization method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • ASRS task scheduling and goods allocation distribution method and system under classified storage

    CN115730789A

  • Evolution-based multi-objective reinforcement learning vehicle route planning method

    CN115907254A

  • Task termination strategy generation method and system based on dynamic programming algorithm

    CN119473779A

  • Lithium battery residual life prediction method based on multi-stage Wiener process

    CN119780729A

  • Method of managing a task

    US20070186214A1

Cited By

  • Task termination policy for polymorphic voting system equipped with protection device

    CN122022011A