Switching control method for a warm reserve system
Patent Information
- Application Number
- CN202611173797.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-04
- Publication Date
- 2026-09-04
AI Technical Summary
[0004]固定延迟时间切换策略的最优阈值是在全场景平均意义下优化得出的单一数值,反映的是所有可能工况下的“平均最优行为”;然而,实际单次运行中主部件的失效时间具有显著的随机性和个体差异,某一固定阈值可能在主部件寿命恰好等于分布均值时表现良好,但当主部件寿命远长于或远短于均值时,该阈值将严重偏离最优决策,也就是说,若主部件实际寿命较长,任务已接近完成,此时即使主部件失效,系统完全可依靠冗余期内的惯性维持至任务结束,而无需冒险激活备用部件;反之,若主部件实际寿命极短,任务尚处于初期阶段,延迟切换将大幅增加系统失效风险;因此,固定延迟时间切换策略是难以根据具备个体差异与动态变化的主部件的实际寿命以在冗余期自适应调整切换时机,从而在具有切换时间冗余的温储备系统的应用中面临较大的失效风险
本发明通过马尔可夫决策过程求解了基于主工作部件实际已运行寿命的最优切换阈值函数,该求解使得在整个历史任务周期的每一个冗余期决策时刻,都是根据当前系统状态重新计算“继续等待”与“立即切换”两种决策方案的期望成本,并选择其中成本更小的动作执行,其最优切换阈值随主部件寿命呈单调函数关系变化,即任务完成概率随寿命增加而增加时阈值单调非减,反之单调非增,则该最优切换阈值体现的是当前主部件寿命下的最优值,而并非是一个所有可能工况下的“平均最优行为”,因而整个过程是根据主工作部件在实际运行过程中所体现出的个体寿命差异,来确定在冗余期与该寿命相匹配的最优切换时机,从而此种动态延迟时间切换策略适配了主部件寿命的个体差异与动态变化,使得具有切换时间冗余的温储备系统在实际工作中能高效应用。
Smart Images

Figure CN122697643A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of temperature storage system control technology, and in particular to a switching control method, device, equipment and medium for a temperature storage system. Background Technology
[0002] A backup system, as a typical redundant architecture, consists of online operating components and standby components. When the primary component fails, the standby component can take over in a timely manner, thereby effectively improving the overall reliability and availability of the system.
[0003] Current backup systems can maintain operation for a short period after the failure of the primary component. Based on this characteristic, researchers have proposed a delayed switching strategy, which involves waiting for a preset time after the primary component fails before activating the backup component. Existing solutions use a fixed-delay switching strategy, the core idea of which is to derive the expected lifetime or expected operating profit of the system under different fixed delay times based on the law of total probability, and select a single fixed threshold that optimizes the performance index as the global switching standard through one-dimensional search or numerical optimization. For example, by analytically deriving the component lifetime distribution and redundancy time distribution, the optimal fixed waiting time that maximizes the average lifetime of the system can be solved.
[0004] The optimal threshold for a fixed-delay switching strategy is a single value optimized across all scenarios, reflecting the "average optimal behavior" under all possible operating conditions. However, in actual single-run operations, the failure time of the main component exhibits significant randomness and individual differences. A fixed threshold may perform well when the main component's lifespan is exactly equal to the distribution mean, but when the main component's lifespan is much longer or shorter than the mean, this threshold will deviate significantly from the optimal decision. In other words, if the actual lifespan of the main component is long and the task is nearing completion, even if the main component fails, the system can rely entirely on the inertia within the redundancy period to maintain operation until the task ends without risking the activation of the backup component. Conversely, if the actual lifespan of the main component is extremely short and the task is still in its early stages, delayed switching will significantly increase the risk of system failure. Therefore, the fixed-delay switching strategy struggles to adaptively adjust the switching timing during the redundancy period based on the actual lifespan of the main component, which exhibits individual differences and dynamic changes, thus facing a significant risk of failure in applications of temperature reserve systems with switching time redundancy. Summary of the Invention
[0005] This invention provides a switching control method for a temperature storage system, which can solve the problems existing in the prior art.
[0006] This invention provides a switching control method for a temperature reserve system, comprising the following steps: Acquire historical lifespan distribution data of the main working components and standby components within the thermal storage system; Based on historical lifetime distribution data, an optimal strategy table for the thermal storage system is constructed: with the optimization objective of minimizing the expected operating cost of the thermal storage system for any historical task cycle, a decision optimization model is established based on a Markov decision process and the optimal decision rule is obtained by solving it. The optimal decision rule is then converted into an optimal strategy table, which stores dynamic switching thresholds that correspond one-to-one with the lifetime values of each possible main working component. The state space of the Markov decision process is the cumulative waiting time of the standby component during the redundancy period, the action space includes continuing to wait and immediate switching, and the state transition probabilities include system task completion, system failure, and system continued survival. Based on the actual operational lifespan of the main working components in the temperature storage system during the current task cycle, the corresponding dynamic switching threshold is matched in the optimal strategy table. After the main working component fails and enters the redundancy period, the cumulative waiting time of the backup component during the redundancy period after the failure of the main working component is compared with the dynamic switching threshold, and the backup component is controlled to continue to remain in the waiting state or switch to the working state.
[0007] Preferably, the construction of the decision optimization model includes: After the temperature reserve system detects a failure in the main working component, the temperature reserve system immediately enters the redundancy period; The operation process of the temperature storage system during the redundancy period is abstracted into a Markov decision process. The cumulative waiting time of the standby component during the redundancy period is used as the system state variable space, and the standby component continuing to wait and immediately switching is used as the action space. The task completion state is defined as the temperature reserve system successfully completing all tasks during the waiting period; the failure termination state is defined as the temperature reserve system failing to switch over due to the failure of the backup component in the temperature reserve state, or failing to continue to maintain due to the accumulated waiting time exceeding the preset redundancy period threshold; the survival state is defined as the temperature reserve system neither completing the task nor failing, but smoothly transitioning to the next decision moment; state transition probabilities are established based on the task completion state, failure termination state, and survival state. A cost-benefit function is constructed based on the one-time switching activation cost, the economic loss cost caused by the overall failure of the temperature reserve system, and the task benefits obtained by successfully completing the task, and a decision optimization model is formed.
[0008] Preferably, the construction of the optimal strategy table includes: The decision optimization model is solved based on the Bellman optimality equation. For each system state of the temperature storage system in the current task cycle, the expected cost of the immediate switching action and the expected cost of the continued waiting action are calculated according to the Bellman optimality equation. The minimum expected cost between the two is taken as the optimal value function of the corresponding system state. Based on the optimal value function of each system state, select the action that minimizes the expected cost for each system state's optimal value function, and convert this optimal decision rule into a mapping table with the main component's lifetime as the index and the optimal switching threshold as the value, forming the optimal strategy table.
[0009] Preferably, determining the dynamic switching threshold in the optimal strategy table as a function of the actual service life of the main component includes: When the main working component fails after its service life exceeds the preset threshold, the workload of the temperature storage system decreases and the dependence on the backup component decreases. The optimal decision tends to extend the waiting time or not switch the backup component. Therefore, the dynamic switching threshold is set to monotonically non-decreasing as the actual service life of the main component increases. When the main working component fails after its service life is less than or equal to a preset threshold, the temperature storage system is in the early stage of operation and has a large workload. The temperature storage system needs to rely on the backup component to take over the operation in order to complete the task. The optimal decision tends to shorten the waiting time or directly switch the backup component. Therefore, the dynamic switching threshold is set to be monotonically non-increasing as the actual service life of the main component increases.
[0010] Preferably, obtaining the actual operational lifespan of the main working component in the current task cycle within the temperature storage system includes: Based on the real-time operating status of the main working components within the temperature storage system, the cumulative operating time of the main working components is continuously recorded from the start of the current task cycle until the failure of the main working components is detected. The cumulative operating time at that moment is then recorded as the actual operating life of the main working components in the current task cycle.
[0011] Preferably, the acquisition of the cumulative waiting time of the backup component during the redundancy period after the failure of the main working component includes: After the temperature reserve system detects a failure in the main working component, the temperature reserve system immediately enters the redundancy period; During the redundancy period, the temperature reserve system does not perform the switching action of the backup component. Instead, it continues to maintain the function of the temperature reserve system by utilizing the system inertia retained after the failure of the main working component. At the same time, it accumulates the total waiting time experienced by the backup component in the temperature reserve state from the start of the redundancy period to the current time, forming the cumulative waiting time of the backup component during the redundancy period after the failure of the main working component.
[0012] Preferably, the control backup component continues to remain in a standby state or switches to an operating state, including: Within the current task cycle, if the cumulative waiting time of the backup component after the failure of the main working component is less than the dynamic switching threshold, it is determined that the optimal switching time has not been reached. The temperature reserve system controls the backup component to continue to maintain the temperature reserve waiting state, does not perform the switching action, and continues to accumulate the waiting time, and enters the next decision cycle to compare and judge again. When the cumulative waiting time of the backup component after the failure of the main working component is greater than or equal to the dynamic switching threshold, the optimal switching time is determined to have arrived. The temperature reserve system immediately sends a switching command to the backup component, controlling the backup component to switch from the temperature reserve state to the working state, so as to take over the failed main working component and continue to perform the temperature reserve system tasks.
[0013] This invention also provides a switching control device for a temperature storage system, comprising: The data module is used to acquire historical lifespan distribution data of the main working components and standby components within the thermal storage system; The strategy construction module is used to construct an optimal strategy table for the temperature storage system based on historical lifetime distribution data. The optimization objective is to minimize the expected operating cost of the temperature storage system for any historical task cycle. A decision optimization model is established based on a Markov decision process, and the optimal decision rule is obtained by solving it. This optimal decision rule is then converted into an optimal strategy table, which stores dynamic switching thresholds corresponding one-to-one with the lifetime values of each possible main working component. The state space of the Markov decision process represents the cumulative waiting time of the backup component during the redundancy period. The action space includes continuing to wait and immediate switching. The state transition probabilities include system task completion, system failure, and system continued survival. The switching control module is used to match the corresponding dynamic switching threshold in the optimal strategy table based on the actual operating life of the main working component in the current task cycle of the temperature storage system. After the main working component fails and enters the redundancy period, the cumulative waiting time of the backup component during the redundancy period after the failure of the main working component is compared with the dynamic switching threshold, and the backup component is controlled to continue to remain in the waiting state or switch to the working state.
[0014] This invention also provides an electronic device, including a memory and a processor; The memory is used to store computer programs; When the processor executes the computer program stored in the memory, it implements the steps of the switching control method for a temperature storage system as described above.
[0015] This invention also provides a computer-readable storage medium for storing a computer program, which, when executed by a processor, implements the steps of a switching control method for a temperature storage system as described above.
[0016] This invention provides a switching control method for a temperature reserve system, which has the following advantages compared with the prior art: This invention solves for the optimal switching threshold function based on the actual operational lifespan of the main working component using a Markov decision process. This solution ensures that at each redundancy period decision point throughout the entire historical task cycle, the expected cost of the two decision options, "continue waiting" and "immediate switching," is recalculated based on the current system state, and the action with the lower cost is selected. The optimal switching threshold changes monotonically with the lifespan of the main component; that is, the threshold is monotonically non-decreasing when the task completion probability increases with lifespan, and monotonically non-increasing when the probability decreases. Therefore, the optimal switching threshold reflects the optimal value under the current lifespan of the main component, rather than an "average optimal behavior" under all possible operating conditions. Thus, the entire process determines the optimal switching time that matches the lifespan during the redundancy period based on the individual lifespan differences exhibited by the main working component during actual operation. This dynamic delay time switching strategy adapts to the individual differences and dynamic changes in the lifespan of the main component, enabling the temperature reserve system with switching time redundancy to be used efficiently in actual operation. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the overall process provided for an embodiment of the present invention; Figure 2 A schematic diagram of state transition probabilities provided for embodiments of the present invention; Figure 3 This is a schematic diagram illustrating the optimal switching decision under different strategies provided in embodiments of the present invention; Figure 4 This is a schematic diagram illustrating the optimal decision-making process for dynamic delay time switching provided in an embodiment of the present invention; Figure 5 A schematic diagram of the optimal value function provided in an embodiment of the present invention; Figure 6 Different virtual lifetime coefficients provided for embodiments of the present invention w A schematic diagram of the optimal decision-making process; Figure 7 Different system failure costs provided for embodiments of the present invention c f A schematic diagram of the optimal decision-making process; Figure 8 Different switching costs provided for embodiments of the present invention c s A schematic diagram of the optimal decision-making process; Figure 9 Different task benefits provided in the embodiments of the present invention r A schematic diagram of the optimal decision-making process; Figure 10This is a schematic diagram illustrating the optimal decision-making process under different maximum redundancy periods, as provided in an embodiment of the present invention. Detailed Implementation
[0018] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0019] Modern engineering systems are becoming increasingly complex and intelligent, and their operational reliability is directly related to public safety and economic benefits. As a typical redundant architecture, the backup system can quickly activate the standby component after the failure of the main component, thereby ensuring the continuous operation of the system. However, traditional switching strategies usually assume that switching occurs immediately after the failure of the main component, ignoring the switching time redundancy characteristic that is common in actual engineering. That is, the system can still maintain operation for a short period of time after the failure of the main component. Taking engineering systems such as mine ventilation duct systems, residential water supply systems, and commercial air conditioning and refrigeration systems as examples, if the standby component can complete the switching and start operation within a preset threshold time when the main working component fails, the system interruption caused by the failure of the main component can be ignored, and the system can still be regarded as being in normal working condition.
[0020] Currently, the main switching strategy is a fixed-delay system switchover strategy. This strategy waits for a pre-optimized fixed time after the primary component fails before activating the backup component. The optimal switching threshold for the fixed-delay strategy is obtained by maximizing the expected benefits, reflecting the average optimal behavior across all scenarios. However, since the failure time of the primary component in a single operation is random, this preset threshold often cannot guarantee that it will be optimal in every specific implementation. Considering the engineering scenario of UAVs performing surveying tasks, if the primary sensor has a long lifespan and the task is nearing completion, even if the primary component fails, the system can rely on the acquired information to maintain operation until the task ends because the remaining task load is minimal, without risking a switchover to the backup sensor. Conversely, if the primary component has a very short lifespan and the task is still in its early stages, the system still needs to rely on the sensor to complete a large amount of surveying and navigation work. In this case, delayed switching can significantly improve the overall lifespan of the system and increase the probability of completing the task. Therefore, the fixed-delay switching strategy is difficult to adapt to the individual differences and dynamic changes in the lifespan of the primary component.
[0021] To address the shortcomings of fixed-delay-time switching strategies, this invention proposes a dynamic-delay-time switching strategy. This strategy dynamically adjusts the switching time based on the actual lifespan of the primary working component, achieving an adaptive decision-making process. Specifically, after the primary working component fails, the system no longer executes a pre-set fixed waiting time, but flexibly determines the switching time of the backup component based on the actual lifespan of the component. This dynamic strategy can appropriately shorten the delay to reduce the risk of system failure when the lifespan of the primary component is short, and appropriately extend the delay to fully utilize the standby capability of the backup component when the lifespan of the primary component is long, thereby balancing system reliability and economy at a more refined level.
[0022] like Figure 1 As shown, the specific steps include: Step 1: Preliminary data collection and parameter calibration.
[0023] Historical operational data of the target system is collected, primarily including raw data on component lifespan and redundancy time. This data is then grouped and statistically analyzed to generate an empirical frequency distribution. Histograms are then fitted using specialized statistical software, and the distribution form that best reflects the true pattern is selected through a goodness-of-fit test, with its specific parameters determined. Furthermore, economic parameters such as switching costs, failure costs, task benefits, and downtime costs are obtained based on the enterprise's financial accounting standards.
[0024] Step 2: Offline model training and policy generation.
[0025] Using the precise parameters obtained in step one, the four core elements required for the Markov decision process model of this system are fully defined: all possible states of the system, all optional actions, state transition probabilities under each action, and the corresponding cost-benefit functions. This allows the construction of a Markov decision process mathematical model suitable for this system. Subsequently, this mathematical model is input into a computational program equipped with a value iteration algorithm, and offline iterative calculations are performed in the background until the algorithm converges. The program ultimately outputs a queryable optimal strategy table. The core logic of this table is that for any given duration of operation of the main component, there exists an optimal redundancy waiting limit.
[0026] Step 3: Monitor the status of main components online in real time.
[0027] After the system is put into actual operation, the operating status of the main working component is monitored in real time. Once a failure of the main component is detected, its total operating life up to the time of failure is recorded immediately, and this is used as an index to find the corresponding preset switching threshold in the optimal strategy table generated in step two.
[0028] Step 4: Dynamically execute adaptive switching decisions.
[0029] Once the redundancy period begins, the system starts timing. Based on the threshold obtained from the lookup table in step three, the following rules are executed: if the current cumulative redundancy period duration is less than the threshold, the system remains in a "wait" state and does not take any action. Once the cumulative redundancy period duration reaches or exceeds the threshold, and the backup component is still alive, the control center immediately issues a switching command to activate the backup component to take over operation.
[0030] Step 5: System feedback and rolling model updates.
[0031] Data such as the lifespan of the main components, the actual redundancy duration, whether the switchover was successful, and the final task results recorded in each actual operation are continuously transmitted back to the database. After accumulating enough new batch data, the process returns to step one to recalibrate parameters and perform offline training, continuously iterating and updating the optimal strategy table so that the decision model can adapt to system aging or changes in the external environment.
[0032] More specifically: I. Model Assumptions and Description.
[0033] This invention models the research problem of delayed switching strategies in thermal reserve systems. First, by clarifying the system structure, random variables, and relevant cost and benefit parameters, the problem is precisely defined, and on this basis, the expected cost function of the system during task execution is derived. It should be noted that although an analytical expression for the expected cost can be obtained, it is difficult to analytically characterize the structural features of the optimal solution, such as the threshold form of the optimal switching strategy or the monotonicity of state dependence. In contrast, the Markov Decision Process (MDP) modeling framework can not only derive the Bellman optimality equation, but also reveal the intrinsic structural properties of the optimal strategy, thereby generating decision rules with clear management significance.
[0034] Based on this, by constructing a state space, a set of feasible decisions, and state transition probabilities, a dynamic optimization problem with the goal of minimizing the expected cost is established, thus laying the foundation for the subsequent analysis of the structural properties of the optimal strategy.
[0035] 1. Expected cost.
[0036] Consider a two-component temperature reserve system for performing operations for a duration of... Z The task, among which Z Let be a random variable with integer values, and satisfy the following conditions: E ( Z The system contains two components: component 1, which is the working component, and component 2, which is the standby component. Let the lifespan of component 1 be an integer random variable. X The lifespan of component 2 during the temperature reserve phase is an integer random variable. YAssuming the lifespans of the two components are independent and the switching process is instantaneous and perfect; the backup component can only switch to the working state after component 1 fails, therefore the switching can be observed at the moment of switching. X The specific value that can be obtained is... X = x Even after component 1 fails, the system still has a switching time redundancy period, i.e., there is a time interval. U During this period, the system can still be considered to be operating normally; the spare component, i.e., component 2, only needs to be in operation during this period. U The switch can be completed within the specified time. U It is modeled as an integer random variable.
[0037] Suppose that the dwell time of the spare component during the redundancy period reaches a threshold. ℓ ( x The switch is triggered when ) ℓ ( x )≥0; After the switch is completed, component 2 enters the working state and continues to execute the remaining tasks. Its remaining lifetime is represented by a random variable. M express.
[0038] set up c s This represents the switching cost of switching a backup component to operational status. c f Represents the system failure cost, where c s and c f All are strictly positive numbers; given X = x Under the condition that, the expected cost of the system during the entire task execution period is denoted as . h x ( ℓ ( x Its expression is as follows: .
[0039] As shown in the above formula, a successful switch requires the reserve component to remain alive, and the redundancy period must not have ended at the time of the switch; after the switch is completed, if the remaining lifespan of component 2 is... M If the system cannot sustain the task until completion, it fails and incurs failure costs, as shown in the second line of the equation; conversely, if the task is successfully completed, a reward can be obtained. r As shown in the third line of the equation; during the redundancy period, if the spare component fails to switch to working status before the deadline, the system will fail, as shown in the fourth line of the equation; in addition, the task may also be completed with a certain probability during the redundancy period, as shown in the fifth line of the equation.
[0040] smaller ℓ (x A higher value ensures that backup components are reliably activated, but premature exposure to adverse operating environments accelerates their degradation, potentially leading to system failure before the task is completed—a phenomenon reflected in the second line of the equation above; conversely, a larger value... ℓ ( x The value will delay the switching, thereby increasing the risk of untimely switching and increasing the probability of system failure, as shown in the fourth line of the above equation; therefore, this invention aims to address each x Determine the optimal switching threshold to achieve the desired cost. E ( C The minimization of ) is expressed as follows: .
[0041] Although the optimization problem explicitly described by the above mathematical formula can be solved directly using numerical methods, this invention uses stochastic dynamic programming to remodel and analyze it. An important advantage of this method is that it can reveal the structural properties of the optimal strategy.
[0042] 2. Dynamic programming modeling.
[0043] This invention uses stochastic dynamic programming to model the problem of finding the minimum expected cost. At each time point during the redundancy period, the decision-maker must make a trade-off between two decisions: whether to switch the backup component to the operating state to prevent system failure, or to continue waiting to slow down the performance degradation of the backup component. To this end, this invention introduces a class of dynamic delay time switching decisions, aiming to achieve a balance between task completion and the risk of system failure.
[0044] by t This indicates the number of time points after the system enters the redundancy phase. t ∈ T ={0, 1, 2, ...}; If a switch is performed, the system will enter normal operation, denoted as ∆; if the switch fails in time, the system will fail and enter the absorption state Ψ; if the task is successfully completed during the redundancy period, it will enter the absorption state Θ; accordingly, the state space can be represented as { T ∪∆∪Ψ∪Θ};Let the waiting decision be denoted as W The switching action is recorded as S Therefore, in each state t The action set can be represented as follows: In action S Under these conditions, the system enters the absorption state ∆ with probability 1; if the system is in state ∆... t And take action W Then, the following three situations may occur: First, the task is completed during the waiting period, and the system completes the task with probability.p (Θ|( t , W The system transitions to the absorption state Θ; secondly, if component 2 fails or the accumulated waiting time exceeds the redundancy threshold, the system will... p (Ψ|( t , W The component fails and enters the absorption state Ψ; third, component 2 survives to the next time point, and the mission is in progress. t If the +1 step is not completed, the system will proceed with the probability step. p ( t +1|( t , W )) Transfer to state t +1; the corresponding state transition probability expression is: .
[0045] Among them: To indicate that the main component is running x The unit time, redundancy phase has continued t Under the premise of a unit time, the task is t The conditional probability of completion at time +1; This means that under the same conditions, the spare component is... t +1 time failure or redundancy period t The conditional probability of the cutoff time +1; the specific expression is as follows: by R x ( t ) indicates a given redundancy time t Online component runtime x At that time, take action S The expected cost incurred; specifically R x ( t ) is represented as: .
[0046] in: To indicate when the main component is running x When each unit of time is in the redundancy phase... t Take switching action S The conditional probability of the system failing afterward.
[0047] For calculation This invention employs a virtual age method, where the age of the temperature reserve component is neither zero nor equal to the time it has spent in the reserve state during switching; it is typically defined as... w ( u ),u Indicates the length of the reserve period, and w ( u )< u Therefore, if a spare component is in [0, u During this period, it operated in temperature reserve mode without failure, and at all times... u When switched to the active state, its remaining lifetime distribution is as follows: R ( w ( u )+ t ) / R ( w ( u )), here w ( u To satisfy w (0) = 0 is a non-decreasing function. This expression indicates that the component is in [0, u The degree of degradation experienced by the internal temperature storage mode is equivalent to that experienced by continuous operation under normal conditions [0, w ( u The degree of degradation of )]; from this, we can obtain The expression is as follows: For example Figure 2 The state transition probability diagram is shown; let's denote... C x ( t ) is in state x Take action below W Time, from moment t The expected cost incurred at the beginning, which measures the cost from the state x The expected cost of starting and executing the task; at the decision-making moment. t From the state t Minimum expected cost of departure V x ( t The solution is given by the following optimal value equation: According to the full expectation formula, the expected cost is... E ( C By measuring the lifespan of each main component X = x With task duration Z = z Expected Cost under Conditions V x (0) is obtained by weighting the sum according to the corresponding joint probabilities, therefore E ( C ) is represented as: Expected cost function E ( C The phased dynamic delay time switching decisions are integrated into a global optimization objective of the temperature reserve system within the switching time redundancy framework, minimizing the expected cost. E ( C That is equivalent to each x Determine the optimal switching threshold ℓ ∗ ( x This allows for a balance between mitigating the degradation of spare parts and the risk of control system failure.
[0048] II. Analysis of the properties of the optimal strategy.
[0049] This invention first characterizes the basic features of optimal switching decisions by performing monotonicity analysis on the optimal value function; based on the above analysis results, it further presents a method for determining the optimal switching strategy; the structural properties obtained here provide an important basis for understanding the underlying decision-making mechanism, and also help simplify the implementation process of delayed switching strategies. For ease of subsequent analysis, several assumptions are introduced:
[0050] Assumption 1: (a) When x Given, Follow t An increase rather than a decrease; when t Given, Follow x (b) when x Given, Follow t The increase rather than the increase; when t Given, Follow x An increase rather than a rise; or (c) It is a constant.
[0051] Assumption 2: (a) When x Given, Follow t An increase rather than a decrease; when t Given, Follow x (a) increases rather than decreases; (b) when x Given, Follow t The increase rather than the increase; when t Given, Follow x An increase rather than a rise; or (c) It is a constant.
[0052] Assumption 3: (a) When x Given, Followt The increase rather than the increase; when t Given, Follow x (a) an increase rather than an increase; or (b) when x Given, Follow t An increase rather than a decrease; when t Given, Follow x It increases rather than decreases.
[0053] Assumption 1(a) states that the task duration Z failure rate t or x It exhibits non-decreasing characteristics, a property commonly seen in actual production environments or project execution. Taking UAVs performing terrain reconnaissance missions as an example, as the operating time accumulates, the completeness of data collection is usually higher, thus increasing the probability of mission success. Hypothesis 1(b) indicates that... Z failure rate t or x It exhibits a non-increasing characteristic, which is common in mission scenarios with high timeliness requirements, such as search and rescue. After a disaster, the probability of survivors being found is usually higher in the initial stage, while the probability of mission success gradually decreases over time, reflecting a decreasing failure rate structure. Hypothesis 1(c) applies to mission duration. Z If the failure rate remains constant in the discrete-time frame, then Z If it follows a geometric distribution, this condition can be satisfied because the geometric distribution has a constant failure rate and is memoryless.
[0054] Assumption 2(a) holds if the spare parts have a lifespan during the storage period. Y With redundancy duration U Both have non-decreasing failure rates, or one has a non-decreasing failure rate and the other has a constant failure rate; conversely, the condition for assumption 2(b) to hold is that... Y and U failure rate t or x Both exhibit non-incremental characteristics, or one has a non-incremental failure rate while the other remains unchanged; if the redundancy duration U With spare life Y If both have a constant failure rate, then hypothesis 2(c) holds.
[0055] Under Hypothesis 3(a), the probability of failure after a spare component switches to operational status decreases with increasing lifetime. This is because a longer lifetime implies better inherent reliability. Conversely, Hypothesis 3(b) indicates that the failure probability increases with increasing lifetime. This is more common in practice and reflects the increasing failure rate of operational components. Both hypotheses can be applied to different component states and mission environments.
[0056] Proposition 1: (Monotonicity of the optimal value function), for a given... x : (a): Under the conditions that Assumption 1(a), Assumption 2(a) and Assumption 3(a) are true, if c f [ p (Ψ|( t +1, W ))- p (Ψ|( t , W ))]≤- r ,but V x ( t )about t Non-incremental.
[0057] (b): Under the conditions that Assumption 1(b), Assumption 2(b), and Assumption 3(b) hold, if c f [ p (Ψ|( t +1, W ))- p (Ψ|( t , W ))]≥ r Then V x ( t )about t Non-reduction.
[0058] (c): Under the conditions that Assumption 1 (c), Assumption 2 (c), and Assumption 3 (a) hold, V x ( t )about t Non-incremental.
[0059] (d): Under the conditions that Assumption 1 (c), Assumption 2 (c), and Assumption 3 (b) hold, V x ( t )about t Non-reduction.
[0060] Proposition 1(a) indicates that system cost decreases as the delay time of spare components during the reserve phase increases. Firstly, according to Hypothesis 1(a), a longer delay time increases the probability that a component will complete its task during the reserve period. Hypothesis 2(a) indicates that this delay also increases the probability of failure during the reserve period. According to Hypothesis 3(a), a longer delay time increases the probability of completing the task after switching. Therefore, it can be inferred that the positive effects of Hypothesis 1(a) and Hypothesis 3(a) outweigh the negative effect of Hypothesis 2(a), making the delayed switching strategy helpful in reducing system cost. In summary, delayed switching achieves effective cost savings by balancing the advantages of system inertia operation and extended component lifespan. For systems with switching time redundancy, when increasing component lifespan can improve the probability of task completion, switching should be delayed as much as possible to obtain greater value. Proposition 1(b) can be understood using the same logic; Propositions 1(c) and 1(d) show that when the component failure rate remains constant, there is a correlation between system cost and the failure rate during task duration; specifically, as shown in Proposition 1(c), a longer delay time increases the probability of task completion after switching, thereby reducing system cost; while Proposition 1(d) exhibits the opposite monotonicity characteristic; at the management level, Proposition 1 provides decision-makers with a basis for switching strategies within the redundancy time window; when delayed switching is beneficial to task completion, the delay time should be increased as much as possible, at which point the failure risk introduced by the delay is within a controllable range; conversely, when delayed switching is detrimental to task completion, the switching strategy should be implemented as early as possible.
[0061] Proposition 2: (Monotonicity of the optimal value function), for a given... t : (a) Under the condition that assumption 3(a) holds, V x ( t )about x Non-incremental.
[0062] (b) Under the condition that assumption 3(b) holds, V x ( t Regarding x Non-reduction.
[0063] Proposition 2(a) indicates that the total system cost decreases with the extension of the lifespan of the main components. This is because the longer lifespan of the main components, as set by Hypothesis 3(a), helps to increase the probability of task completion, which is common in engineering cases. Conversely, Proposition 2(b) indicates that under certain conditions, the extension of the lifespan of the main components may actually reduce the probability of task completion, leading to an increase in system cost. This phenomenon arises from the scenario depicted by Hypothesis 3(b), such as in search and rescue missions, where the probability of successful rescue decreases if survivors are not found in a short time. Therefore, in such situations, accelerating the task completion process is particularly crucial. For decision-makers, while strategy optimization can improve system economy, the inherent characteristics and operating environment of external equipment are equally important. In actual operation and maintenance, decision-makers should rationally configure the operating parameters of working components according to the task type and equipment attributes.
[0064] Theorem 1: (Existence of an optimal switching threshold) For a given... x If the following condition is met, then there exists a switching threshold. ℓ ∗ ( x ) makes when t < ℓ ∗ ( x The optimal decision at this time is to continue waiting. t ≥ ℓ ∗ ( x The optimal decision at this time is to perform a switchover: .
[0065] Theorem 1 reveals the optimal switching strategy for backup components during the redundancy phase, indicating the existence of an optimal switching threshold. ℓ ∗ ( x This means that when the redundancy time is less than a certain threshold, continuing to wait is the optimal decision; when the redundancy time reaches or exceeds the threshold, a switchover operation should be performed. This result reflects a trade-off mechanism: when the redundancy time is short, waiting helps slow down component aging, while when the redundancy time exceeds the threshold... ℓ ∗ ( x After that, the risk of missing the switching opportunity gradually increases and exceeds the benefits brought by the delay strategy. At this time, it is necessary to switch to the running state immediately to avoid irreversible failure of the system.
[0066] Theorem 2: (Monotonicity of the optimal switching threshold).
[0067] (a): Under the condition that assumption 3(a) holds, the optimal switching threshold is... ℓ ∗ (x Regarding x Non-reduction.
[0068] (b): Under the condition that assumption 3(b) holds, the optimal switching threshold is... ℓ ∗ ( x )about x Non-incremental.
[0069] Theorem 2(a) shows that, under the given conditions, the switching timing of the backup component is positively correlated with the lifespan of the primary component; specifically, the longer the lifespan of the primary component, the later the backup component should be switched. This conclusion is based on Assumption 3(a), which assumes that a longer primary component lifespan helps increase the probability of task completion, thus eliminating the need to activate the backup component prematurely, as premature activation would introduce unnecessary activation costs and failure risks. The mechanism revealed by Theorem 2(b) is consistent with Proposition 2(b), suggesting that in specific scenarios, early switching is more conducive to accelerating task completion. Theorem 2 summarizes the basic principle of dynamically adjusting the switching time of the backup component based on the actual lifespan of the primary component. This theorem provides decision-makers with clear management insights: in scenarios where the probability of task success increases with the working time of the primary component, when the primary component lifespan is long, switching should be appropriately delayed to fully utilize its remaining lifespan, thereby avoiding unnecessary costs and failure risks from premature activation of the backup component; while in situations with high timeliness requirements, when the primary component lifespan is long, the reserve component should be switched to working status as early as possible.
[0070] Lemma 1: (Sufficient condition for immediate switching) If the following condition holds, then the optimal switching threshold satisfies... ℓ ∗ ( x )=0, meaning the backup component is activated immediately after the main component fails. .
[0071] Lemma 1 provides the optimal switching threshold. ℓ ∗ ( x A sufficient condition for )=0 is that the backup component should be activated immediately after the primary component fails; the above condition requires that the marginal reduction in failure risk obtained by immediate switching compared to continuing to wait. It should be no less than the normalized switching cost. The normalized cost reflects the one-time activation cost. c s Consequences of system failure c f and rewards for successful tasks rThe condition describes the trade-off between risk and cost. Intuitively, this condition portrays the trade-off between risk and cost. When the risk reduction from immediately switching to a backup component is sufficient to cover the switching cost, immediate switching becomes the only optimal decision, because any delay will only expose the system to a higher risk of failure during the redundancy period and lead to an increase in expected costs. From a management perspective, this result provides clear operational guidelines for critical engineering systems. In high-risk scenarios such as power grid systems and medical life support systems, operators should prioritize the immediate activation of backup components. Therefore, the normalized cost threshold should be adjusted based on the specific failure cost and task benefits of the system to achieve a balance between cost and risk. At the same time, a proactive real-time monitoring mechanism should be established to trigger immediate activation once the failure risk gap crosses the threshold, thereby minimizing expected costs while ensuring the resilience and reliability of the system.
[0072] III. Numerical Examples
[0073] This invention verifies the effectiveness and robustness of the proposed dynamic delay time switching strategy for a temperature storage system that considers switching time redundancy through numerical simulation experiments.
[0074] 1. Strategy comparison.
[0075] To verify the effectiveness of the switching strategy of this invention in a real-world engineering context, we examine the UAV backup sensor system researched by Levitin et al. as an example. This UAV system can perform tasks such as environmental monitoring, battlefield reconnaissance, or smart agriculture. Assuming the UAV is equipped with two sensors, and the primary sensor is exposed to harsh environmental conditions with a relatively short expected lifespan, its lifespan... X It follows a discrete Weibull distribution with a shape parameter of 2 and a scale parameter of 2²; the spare sensor, i.e., component 2, has a relatively long lifespan because it is stored inside the UAV. Y The task duration required for the UAV to complete follows a discrete Weibull distribution with a shape parameter of 2 and a scale parameter of 28. Z It follows a discrete Weibull distribution with shape parameter 2 and scale parameter 30; task reward r Set to 210, system failure cost c f The switching cost is 160. c s The virtual lifetime coefficient is 80. w The value is set to 0.8; after the main sensor fails, the drone can continue to operate for a period of time based on the information already collected. U The duration follows a discrete uniform distribution of 5 to 10, and the specific variable values can be found in Table 1. Subsequently, two switching strategies are introduced for comparison with the dynamic delay time switching strategy. The specific definitions and expected cost calculations under the corresponding strategies are explained in detail below.
[0076] Table 1. Minimum Expected Cost Parameter Values Main component lifespan shape parameters and dimensional parameters (2,22) Reserve component lifespan shape parameters and dimensional parameters (2,28) Task duration, shape parameters, and scale parameters (2,30) Task Rewards r 210 System failure cost cf 160 Switching cost CS 80 Virtual lifetime coefficient w 0.8 Redundancy duration U upper and lower limits [5, 10] Failure-Triggered Switching (FTS): Under this strategy, if a running component fails and the backup component remains available, the backup component is immediately activated to replace the failed component. In other words, the switching threshold for this strategy is... ℓ ∗ ( x The strategy of setting the value to 0 is widely used in academic research and engineering practice because it can maintain the continuity of system operation. However, frequent switching may lead to increased switching costs and accelerated equipment aging, ultimately resulting in increased operating costs. The expected cost of this strategy is obtained from the following formula:
[0077] .
[0078] Fixed-Time Delayed Switching (FDS): This strategy stipulates that after a running component fails, if the backup component is still available when a pre-set threshold is reached, it will be switched to the running state. In other words, this strategy sets the switching threshold to... ℓ ∗ ( x )= δ , in δ To and x Irrelevant positive integers. This strategy is the fixed delay time switching strategy mentioned earlier. Its core idea is to reduce the overall system operating cost by optimizing the preset delay time; the expected cost under this strategy can be obtained by the following formula:
[0079] .
[0080] Dynamic Switching Strategy (DSS): This is the innovative switching strategy proposed in this invention. Under this strategy, the switching timing of the backup component changes dynamically according to the failure time of the running component, that is, the switching time is a function of the lifespan of the running component. By adaptively adjusting the switching decision based on the observed failure time, this strategy can effectively reduce operating costs while improving the probability of task completion.
[0081] Table 2 Expected Costs of Different Strategies FTS Strategy 46.4 50.9 49.2 51.6 FDS Strategy 39.2 41.7 42.6 39.5 DSS Strategy 36.2 39.8 39.3 36.9 2. Optimal switching strategy.
[0082] Table 2 presents the calculation results of the numerical examples. As shown in Table 2, under the fault-triggered switching strategy, the minimum expected cost of the system is 46.4. Figure 3 As shown, the optimal switching threshold for this strategy is 0; under the fixed-delay switching strategy, the optimal expected cost of the system is 39.2, and as... Figure 3 As shown, the optimal switching threshold is 8, meaning that regardless of when component 1 fails, component 2 will switch 8 time points after its failure. For the dynamic delay time switching strategy, the optimal cost is 36.2, which is the lowest among the three strategies, fully verifying the superiority of the strategy proposed in this invention. Table 2 also lists the optimal costs of the three strategies under different parameter configurations. It can be observed that the dynamic delay time switching strategy proposed in this invention can achieve the lowest operating cost in all test scenarios. Figure 3 Further evidence shows that ℓ ∗ ( x )about x The non-decreasing characteristic, consistent with the theoretical result of Theorem 2(a), indicates that the longer the service life of the primary component, the longer the switching time of the backup component can be delayed. This is because the extended service life of the primary component increases the probability of successful task completion, making premature switching unnecessary and thus avoiding additional switching costs and failure risks.
[0083] Depend on Figure 4 It can be seen that when the lifespan of component 1 is less than 10, the switching threshold for the spare component is 0, which means that an immediate switchover is performed. The corresponding switchover decision is marked with a red square in the figure. When the lifespan of component 1 is between 11 and 13, the switchover should be based on... ℓ ∗ ( x The optimal switching threshold exceeds the maximum redundancy duration of 10 for component 1, indicating that the spare component should always remain in its current reserve state and no switching is required. These observations are consistent with... Figure 3 The patterns observed are consistent, further validating the robustness of the proposed strategy.
[0084] Figure 5 Showing x =12 Time Value Function V x ( t ) t The trend of change, as shown in the graph, is that when... t When the value is less than 3, the expected cost of continuing to wait is low, so delaying the switch is the optimal choice at this time; when... t When the value is ≥3, the cost of switching is lower, therefore switching becomes the better decision; the above observations are consistent with... Figure 3 and Figure 4The patterns presented are consistent; furthermore, V x ( t )about t It exhibits non-decreasing characteristics, which means that the expected cost increases monotonically with the increase of the switching time redundancy. This monotonicity can be mainly attributed to the increasing risk of system failure as the waiting time increases.
[0085] 3. Sensitivity analysis.
[0086] To conduct sensitivity analysis, this invention employs a different parameter configuration than that used in the strategy comparison. Sensitivity analysis aims to reveal the impact of model parameter changes on system behavior and the optimal strategy. Under the parameter configuration used in the strategy comparison, system performance changes exhibit relatively small amplitudes or irregular trends, which may mask potential patterns of change. To ensure clearer and easier-to-interpret results, sensitivity analysis uses a different set of parameters: component 1 follows a discrete Weibull distribution with a shape parameter of 2 and a scale parameter of 100; the backup sensor component 2 follows a discrete Weibull distribution with a shape parameter of 2 and a scale parameter of 170. Z It follows a discrete Weibull distribution with a shape parameter of 1.8 and a scale parameter of 190. The task reward... r = 10000, System failure cost c f =100, Sensor activation cost c s =50, virtual lifetime factor set to 50. w =0.5, U It follows a discrete uniform distribution from 1 to 50.
[0087] Figure 6 Demonstrates virtual lifetime coefficient w Impact on strategy implementation, coefficient w Used to indicate the severity of the environment during the working phase compared to the preparation phase. w The closer the value is to 1, the more similar the standby environment is to the operating environment; as shown in the figure, when w When the value increases, it indicates that the environment in which the reserve components operate becomes more severe, and the optimal switching time for the reserve components needs to be brought forward accordingly. This is reflected in the graph as the area of the green waiting region decreases; more specifically, when... w = 0.3、 x When the value is 160, the optimal strategy is to wait for 8 time points after the main component fails, and if the reserve has not yet failed and the task has not been completed, then execute the switching decision; while when wWhen the value increases to 0.6, the environmental severity intensifies, and the optimal switching threshold drops to two time points. It is almost necessary to switch the standby component to the operating state immediately to reduce the risk of failure during the standby period. For decision-makers, the switching strategy should be adjusted according to the working environment and the reserve environment. When the reserve environment is relatively mild, the switching should be appropriately delayed, while when the reserve environment is more severe, the switching should be done earlier.
[0088] from Figure 7 The results show that the geometric shapes and area distributions of the waiting decision region and the switching decision region in the four sub-graphs are highly similar. This phenomenon indicates that the optimal switching decision has significant consistency under different system failure cost settings, meaning that the optimal switching strategy is insensitive to changes in this parameter. The main reason for this result is that regardless of whether the decision-maker chooses immediate switching or delayed switching, the system cannot completely avoid potential failure costs. Specifically, if immediate switching is chosen, the switched component will face a higher risk of operational failure in the working state, while if delayed switching is chosen, the component will stay in the reserve system for a longer period of time, thereby increasing the possibility of failure during the reserve phase. Both types of risks will ultimately be reflected in the form of system failure costs. Because the expected cost does not significantly favor one decision, the decisions show strong consistency.
[0089] Figure 8 This indicates that rising switching costs significantly delay the optimal switching time for backup components. For example, when the primary component's lifespan is 180 days, if the switching cost is 40, then it's advisable to wait 10 units of time after the primary component fails before switching the backup component to operational status. However, when the switching cost is 70, it's possible to wait 15 units of time after failure before switching, resulting in the lowest expected cost. This phenomenon arises because switching costs are expenses incurred only during the switching action; the waiting action itself does not incur these costs. Therefore, higher costs directly inhibit switching behavior during the decision-making process. When weighing the current switching cost against the risk of future failure, the system tends to postpone switching, ideally avoiding it altogether, to reduce the economic burden of high costs, thus demonstrating a clear delay trend in the optimal switching decision. Therefore, if switching requires significant manpower, material resources, and financial resources, it should be postponed. For instance, switching a power grid converter valve typically causes instantaneous voltage spikes, resulting in substantial component wear. In such scenarios, switching should be delayed as much as possible, ideally avoiding switching the component altogether.
[0090] Figure 9 The changing pattern of the switching threshold presented is consistent with Figure 7The two options share an inherent similarity because both choices—switching or maintaining the current operating state to utilize the redundancy period—can bring positive expected task benefits to the system. Specifically, if switching is chosen, the reserve component, in its working state, has a chance to successfully complete the task and generate benefits. If switching is not chosen, the system continues to operate relying on the redundancy period, also having the opportunity to obtain task benefits while avoiding switching costs. Since both decision paths imply non-zero expected task benefits, the benefit factor forms a certain degree of hedging effect in the decision function, thereby weakening its marginal influence on the switching threshold. Therefore, compared to more discriminative parameters such as system failure costs, the task benefit parameter has a relatively weaker impact on the optimal switching boundary. This observation indicates that in reserve systems with redundancy mechanisms, the decision is less sensitive to changes in task benefits, and decision-makers can focus more on dominant factors such as failure costs and switching costs.
[0091] Figure 10 The study reveals that the optimal delay switchover time increases with the length of the redundancy period, meaning that the longer the available redundancy time, the more likely the optimal switchover decision will be postponed. This conclusion aligns with intuitive understanding: a longer redundancy period provides greater margin for error in the event of system failure, allowing decision-makers to postpone the switchover without impacting system performance. From a management perspective, this finding suggests that when redundancy time is sufficient, such as when the map information collected by the UAV system is comprehensive or the terrain complexity is low, postponing the switchover helps to make fuller use of the system's remaining functionality and may reduce unnecessary switchover costs. Conversely, when redundancy time is limited, earlier switchover is more conducive to controlling the risk of system failure.
[0092] In this invention, during the redundancy period, the system needs to make a choice between "immediately switching or continuing to wait" at each decision point based on the current observed state. This choice not only affects the cost at the current moment but also influences the cost distribution of all future decision points through state transitions. This strategy transforms the decision variable from "a time length" to "a series of state-dependent decision rules." Furthermore, this invention adopts a closed-loop sequential decision structure, re-evaluating the optimal action under the current state at each decision point. This allows for full utilization of dynamically emerging real-time information during the redundancy period to continuously correct existing decision paths. Thus, this invention views decision-making as a dynamic process that continues as the system state evolves, rather than a static event completed at a specific point in time.
[0093] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.
Claims
1. A switching control method for a temperature reserve system, characterized in that, Includes the following steps: Acquire historical lifespan distribution data of the main working components and standby components within the thermal storage system; Based on historical lifetime distribution data, an optimal strategy table for the thermal storage system is constructed: with the optimization objective of minimizing the expected operating cost of the thermal storage system for any historical task cycle, a decision optimization model is established based on a Markov decision process and the optimal decision rule is obtained by solving it. The optimal decision rule is then converted into an optimal strategy table, which stores dynamic switching thresholds that correspond one-to-one with the lifetime values of each possible main working component. The state space of the Markov decision process is the cumulative waiting time of the standby component during the redundancy period, the action space includes continuing to wait and immediate switching, and the state transition probabilities include system task completion, system failure, and system continued survival. Based on the actual operational lifespan of the main working components in the temperature storage system during the current task cycle, the corresponding dynamic switching threshold is matched in the optimal strategy table. After the main working component fails and enters the redundancy period, the cumulative waiting time of the backup component during the redundancy period after the failure of the main working component is compared with the dynamic switching threshold, and the backup component is controlled to continue to remain in the waiting state or switch to the working state.
2. The switching control method for a temperature storage system according to claim 1, characterized in that, The construction of the decision optimization model includes: After the temperature reserve system detects a failure in the main working component, the temperature reserve system immediately enters the redundancy period; The operation process of the temperature storage system during the redundancy period is abstracted into a Markov decision process. The cumulative waiting time of the standby component during the redundancy period is used as the system state variable space, and the standby component continuing to wait and immediately switching is used as the action space. The task completion state is defined as the temperature reserve system successfully completing all tasks during the waiting period; the failure termination state is defined as the temperature reserve system failing to switch over due to the failure of the backup component in the temperature reserve state, or failing to continue to maintain due to the accumulated waiting time exceeding the preset redundancy period threshold; the survival state is defined as the temperature reserve system neither completing the task nor failing, but smoothly transitioning to the next decision moment; state transition probabilities are established based on the task completion state, failure termination state, and survival state. A cost-benefit function is constructed based on the one-time switching activation cost, the economic loss cost caused by the overall failure of the temperature reserve system, and the task benefits obtained by successfully completing the task, and a decision optimization model is formed.
3. The switching control method for a temperature storage system according to claim 2, characterized in that, The construction of the optimal strategy table includes: The decision optimization model is solved based on the Bellman optimality equation. For each system state of the temperature storage system in the current task cycle, the expected cost of the immediate switching action and the expected cost of the continued waiting action are calculated according to the Bellman optimality equation. The minimum expected cost between the two is taken as the optimal value function of the corresponding system state. Based on the optimal value function of each system state, select the action that minimizes the expected cost for each system state's optimal value function, and convert this optimal decision rule into a mapping table with the main component's lifetime as the index and the optimal switching threshold as the value, forming the optimal strategy table.
4. The switching control method for a temperature storage system according to claim 3, characterized in that, The determination of the dynamic switching threshold in the optimal strategy table as a function of the actual service life of the main component includes: When the main working component fails after its service life exceeds the preset threshold, the workload of the temperature storage system decreases and the dependence on the backup component decreases. The optimal decision tends to extend the waiting time or not switch the backup component. Therefore, the dynamic switching threshold is set to monotonically non-decreasing as the actual service life of the main component increases. When the main working component fails after its service life is less than or equal to a preset threshold, the temperature storage system is in the early stage of operation and has a large workload. The temperature storage system needs to rely on the backup component to take over the operation in order to complete the task. The optimal decision tends to shorten the waiting time or directly switch the backup component. Therefore, the dynamic switching threshold is set to be monotonically non-increasing as the actual service life of the main component increases.
5. The switching control method for a temperature storage system according to claim 1, characterized in that, The acquisition of the actual operational lifespan of the main working components in the temperature storage system during the current task cycle includes: Based on the real-time operating status of the main working components within the temperature storage system, the cumulative operating time of the main working components is continuously recorded from the start of the current task cycle until the failure of the main working components is detected. The cumulative operating time at that moment is then recorded as the actual operating life of the main working components in the current task cycle.
6. The switching control method for a temperature storage system according to claim 1, characterized in that, The acquisition of the cumulative waiting time of the backup component during the redundancy period after the failure of the main working component includes: After the temperature reserve system detects a failure in the main working component, the temperature reserve system immediately enters the redundancy period; During the redundancy period, the temperature reserve system does not perform the switching action of the backup component. Instead, it continues to maintain the function of the temperature reserve system by utilizing the system inertia retained after the failure of the main working component. At the same time, it accumulates the total waiting time experienced by the backup component in the temperature reserve state from the start of the redundancy period to the current time, forming the cumulative waiting time of the backup component during the redundancy period after the failure of the main working component.
7. The switching control method for a temperature storage system according to claim 6, characterized in that, The control backup component either remains in a standby state or switches to an operating state, including: Within the current task cycle, if the cumulative waiting time of the backup component after the failure of the main working component is less than the dynamic switching threshold, it is determined that the optimal switching time has not been reached. The temperature reserve system controls the backup component to continue to maintain the temperature reserve waiting state, does not perform the switching action, and continues to accumulate the waiting time, and enters the next decision cycle to compare and judge again. When the cumulative waiting time of the backup component after the failure of the main working component is greater than or equal to the dynamic switching threshold, the optimal switching time is determined to have arrived. The temperature reserve system immediately sends a switching command to the backup component, controlling the backup component to switch from the temperature reserve state to the working state, so as to take over the failed main working component and continue to perform the temperature reserve system tasks.
8. A switching control device for a temperature storage system, characterized in that, include: The data module is used to acquire historical lifespan distribution data of the main working components and standby components within the thermal storage system; The strategy construction module is used to construct an optimal strategy table for the temperature storage system based on historical lifetime distribution data. The optimization objective is to minimize the expected operating cost of the temperature storage system for any historical task cycle. A decision optimization model is established based on a Markov decision process, and the optimal decision rule is obtained by solving it. This optimal decision rule is then converted into an optimal strategy table, which stores dynamic switching thresholds corresponding one-to-one with the lifetime values of each possible main working component. The state space of the Markov decision process represents the cumulative waiting time of the backup component during the redundancy period. The action space includes continuing to wait and immediate switching. The state transition probabilities include system task completion, system failure, and system continued survival. The switching control module is used to match the corresponding dynamic switching threshold in the optimal strategy table based on the actual operating life of the main working component in the current task cycle of the temperature storage system. After the main working component fails and enters the redundancy period, the cumulative waiting time of the backup component during the redundancy period after the failure of the main working component is compared with the dynamic switching threshold, and the backup component is controlled to continue to remain in the waiting state or switch to the working state.
9. An electronic device, characterized in that, include: Memory and processor; The memory is used to store computer programs; When the processor executes the computer program stored in the memory, it implements the steps of the switching control method for a temperature storage system as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Used to store a computer program, which, when executed by a processor, implements the steps of a switching control method for a temperature storage system as described in any one of claims 1 to 7.