Early warning anomaly detection and risk avoidance industrial device control method and apparatus

By acquiring the controllable and statistical states of industrial equipment and using a control action decision model to generate real-time control actions, the problem of the inability to avoid risks in a timely manner in existing technologies is solved, thereby improving the safety and reliability of the equipment.

CN117032104BActive Publication Date: 2026-06-19TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL
Filing Date
2023-07-31
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

Existing technologies struggle to implement adaptive real-time output control strategies in industrial equipment control, making it difficult to mitigate risks in a timely manner during anomaly detection, which may lead to equipment damage and safety hazards.

Method used

By acquiring the controllable state and statistical state of preset parameters of the device to be controlled, and inputting them into the control action decision model, with the goal of maximizing benefits, the selection probability of each preset control action is determined by the objective function and constraints of the control action decision model, and real-time control actions are generated based on the selection probability.

Benefits of technology

It enables timely control actions when abnormal features are detected, reducing losses caused by equipment malfunctions and improving equipment safety and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117032104B_ABST
    Figure CN117032104B_ABST
Patent Text Reader

Abstract

This disclosure relates to an industrial equipment control method and apparatus for early warning anomaly detection and risk avoidance. The method includes: acquiring the controllable state and statistical state of preset parameters of the equipment to be controlled at the current moment; inputting the controllable state and statistical state into a control action decision model, and determining the selection probability of each preset control action by using the objective function and constraints of the control action decision model with the goal of maximizing benefits; determining the target control action based on the preset action selection strategy and the selection probability of each preset control action; and controlling the equipment to be controlled according to the target control action. Through the above technical solution, the controllable state and statistical state of preset parameters at the current moment can be input into the control action decision model to generate control actions in real time and control the equipment. This allows for timely implementation of corresponding control actions to handle anomalies when abnormal characteristics are detected, reducing losses caused by equipment anomalies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of industrial equipment technology, and more specifically, to an industrial equipment control method and apparatus for early warning anomaly detection and risk avoidance. Background Technology

[0002] In related technologies, to achieve control of industrial equipment, Markov Decision Process (MDP) models are often constructed based on the operating status of the equipment, and the optimal control problem of MDP model changes is handled from the perspective of Quickest Change Detection (QCD). However, although QCD is used to study the change problem of MDP, these technologies only focus on switching between optimal strategies for different MDPs and do not achieve adaptive real-time output control strategies, making them difficult to apply in industrial alarm environments. Summary of the Invention

[0003] The purpose of this disclosure is to provide an industrial equipment control and device for early warning anomaly detection and risk avoidance, in order to solve the above-mentioned technical problems.

[0004] To achieve the above objectives, the first aspect of this disclosure provides an industrial equipment control method for early warning anomaly detection and risk avoidance, the method comprising:

[0005] The controllable state and the statistical state of preset parameters of the device to be controlled at the current moment are obtained, and the statistical state is used to characterize the changing state of the device to be controlled.

[0006] The controllable state and the statistical state are input into the control action decision model. With the goal of maximizing the benefit, the selection probability of each preset control action is determined through the objective function and constraints of the control action decision model.

[0007] The target control action is determined based on the preset action selection strategy and the selection probability of each preset control action;

[0008] The device to be controlled is controlled according to the target control action.

[0009] Optionally, obtaining the statistical status of the device to be controlled at the current moment includes:

[0010] Based on a preset probability density function, the log-likelihood ratio of the preset parameter at each moment within a first time period is determined. The probability density function includes a first probability density function and a second probability density function. The first probability density function is used to characterize the probability distribution of the observed value of the preset parameter before the point of change, and the second probability density function is used to characterize the probability distribution of the observed value of the preset parameter after the point of change. The point of change is the time point at which the statistical average value of the observed value changes. The first time period is used to characterize the time period from the start time to the current time.

[0011] For each second time period, a cumulative log-likelihood ratio is determined, which is used to characterize the sum of the log-likelihood ratios of the preset parameters at each moment within the second time period. The second time period is used to characterize the time period from each moment within the first time period to the current moment.

[0012] The sum of the largest log-likelihood ratios is determined as the statistic for the preset parameter;

[0013] Based on the statistics and the preset correspondence between the statistical states and the statistics, the statistical state of the device to be controlled at the current moment is determined.

[0014] Optionally, the objective function is:

[0015]

[0016]

[0017]

[0018]

[0019] in, This represents a quadruple state consisting of a statistical state, a controllable state, a change-time state, and the current time state. S This represents the state space consisting of four-tuple states. a Indicates the preset control action. A This represents the action state space for preset control actions. Represents the target optimization variable. Indicates the state of the quadruple Execute preset control actions a The gains obtained Indicates taking the state of the quadruple. Current moment n controllable state , express n A state of controllability at any given moment. express n The target control action selected at any given time. express n Always under control Execute preset control actions The gains obtained Represents the steady-state probability of a state. This indicates the preset control action selection function, representing the state. Select preset control action The probability, This indicates the preset values ​​of parameters related to industrial equipment damage. This represents the preset minimum controllable state. This indicates the preset initial controllable state. It represents an arbitrary constant.

[0020] Optionally, the constraints include state constraints and safety constraints, wherein the state constraints characterize the conditions under which the system containing the controlled device reaches a steady state, and the safety constraints characterize the conditions under which the controlled device is kept in a normal operating state.

[0021] Optionally, the state constraint condition is:

[0022]

[0023]

[0024]

[0025]

[0026] in, This represents the state space consisting of statistical states, controllable states, time-varying states, and the current time state. S This represents the state space consisting of four-tuple states. a Indicates the preset control action. A This represents the action state space for preset control actions. Represents the steady-state probability of a state. This indicates the preset control action selection function, representing the state. Select preset control action a The probability, Represents the target optimization variable. Indicates from state Execute preset control actions a When the next state is obtained. The probability, Representing state steady-state probability, Representing state The steady-state probability multiplied by the state Select preset control action a The probability of.

[0027] Optionally, the security constraint is:

[0028]

[0029] in, A threshold representing the probability of equipment damage. This indicates that the states in the state space belong to the set of first preset states. The first preset states are used to characterize the state in which equipment malfunctions are handled safely. Indicates the first preset state. This indicates that the states in the state space belong to the set of second preset states. The second preset states are used to characterize the state where the device malfunction cannot be safely handled. This indicates the second preset state. This represents the optimization variable belonging to the second preset state. This represents the optimization variable belonging to the first preset state.

[0030] Optionally, determining the selection probability of each preset control action, with the goal of maximizing profit, through the objective function and constraints of the control action decision model, includes:

[0031] Determine the optimal solution for the target optimization variable, and determine the optimal steady-state probability based on the optimal solution;

[0032] Determine the magnitude of the preset steady-state probability and the optimal steady-state probability, and when the optimal steady-state probability is greater than the preset steady-state probability, determine the first selection probability of each preset control action based on the optimal solution;

[0033] The selection probability of each preset control action is determined based on the preset weighting coefficient and the first selection probability.

[0034] A second aspect of this disclosure provides an industrial equipment control device for early warning anomaly detection and risk avoidance, the device comprising:

[0035] The acquisition module is used to acquire the controllable state and the statistical state of preset parameters of the device to be controlled at the current moment. The statistical state is used to characterize the changing state of the device to be controlled.

[0036] The first determining module is used to input the controllable state and the statistical state into the control action decision model, with the goal of maximizing the benefit, and to determine the selection probability of each preset control action through the objective function and the constraints of the control action decision model.

[0037] The second determining module is used to determine the target control action based on the preset action selection strategy and the selection probability of each preset control action.

[0038] The control module is used to control the device to be controlled according to the target control action.

[0039] A third aspect of this disclosure provides a non-transitory computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the method described in any of the first aspects of this disclosure.

[0040] A fourth aspect of this disclosure provides an electronic device, comprising:

[0041] A memory on which computer programs are stored;

[0042] A processor for executing the computer program in the memory to implement the steps of the method according to any one of the first aspects of this disclosure.

[0043] The above technical solution allows the current controllable state and the statistical state of preset parameters to be input into the control action decision model to generate control actions in real time and control the equipment. This enables timely control actions to be taken when abnormal features are detected, thereby reducing losses caused by equipment malfunctions.

[0044] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description

[0045] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the following detailed description to explain the present disclosure, but do not constitute a limitation thereof. In the drawings:

[0046] Figure 1 This is a flowchart illustrating an industrial equipment control method for early warning anomaly detection and risk avoidance according to an exemplary embodiment;

[0047] Figure 2 This is a block diagram illustrating a sensor network according to an exemplary embodiment;

[0048] Figure 3 This is a block diagram illustrating an improved Markov system according to an exemplary embodiment;

[0049] Figure 4 This is a schematic diagram illustrating the results of average expected revenue and equipment damage rate under different strategies and geometric priors, according to an exemplary embodiment.

[0050] Figure 5 This is a schematic diagram illustrating the average normal operating duration and the degree of damage after abnormal changes, obtained according to an exemplary embodiment of the equipment control method.

[0051] Figure 6 This is a schematic diagram of a strategy table according to an exemplary embodiment;

[0052] Figure 7 This is a block diagram illustrating an industrial equipment control device for early warning anomaly detection and risk avoidance according to an exemplary embodiment;

[0053] Figure 8 This is a block diagram illustrating an electronic device according to an exemplary embodiment. Detailed Implementation

[0054] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0055] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0056] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0057] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0058] As mentioned in the background section, in order to control industrial equipment, MDP models are often constructed based on the operating status of the equipment, and the optimal control problem of MDP model changes is handled from the perspective of QCD. However, although QCD is used to study the changes in MDP, these related technologies only focus on the optimal strategy switching control for different MDPs and do not achieve adaptive real-time output control strategies, making them difficult to apply in industrial alarm environments.

[0059] Specifically, the CuSUM algorithm, the fastest detection algorithm proposed in related technologies, successfully balances detection latency and false alarm rate during anomaly detection. The CuSUM algorithm has been theoretically proven to perform anomaly detection at the fastest speed, such as detecting anomalous changes in statistical properties in random sequences. This technology has good applicability in non-stationary environments in industrial settings, such as detecting pressure, temperature, or arbitrary statistical values ​​related to hazard occurrence, and quickly identifying abnormal situations.

[0060] However, in scenarios where false alarms leading to device shutdowns could result in significant cost losses, it's necessary to increase the detection threshold of the CuSUM algorithm to reduce the false alarm rate. However, this approach increases detection latency. Therefore, in such cases, taking appropriate anomaly handling measures only after detecting the anomaly could lead to substantial losses, such as equipment damage or even explosions.

[0061] To address this issue, related technologies consider modeling the problem as a Markov decision process for analysis and optimal strategy derivation. Specifically, the optimal control problem of MDP model changes can be handled from the perspective of QCD. However, while these methods utilize QCD to study MDP changes, they only focus on the optimal strategy switching control between different MDPs and do not achieve adaptive real-time strategy output, making them difficult to apply in industrial alarm environments. Furthermore, these methods lack sufficient research on risk avoidance, especially when decision-makers can only move the controllable state to its adjacent state; simply using the fastest change detection and avoidance method may result in the controllable state failing to transition to a safe state in a timely manner.

[0062] In view of this, the present disclosure provides an industrial equipment control method and apparatus for early warning anomaly detection and risk avoidance. By inputting the current controllable state and the statistical state of preset parameters into the control action decision model, control actions are generated in real time and the equipment is controlled. This allows for timely implementation of corresponding control actions to handle anomalies when abnormal features are detected, thereby reducing losses caused by equipment anomalies.

[0063] Specifically, the problem to be solved can be modeled as a sensor network, such as Figure 2As shown, the sensor network can include four units: a sequence observation unit, a cumulative and statistical calculation and processing unit, a real-time control unit, and a controllable state unit for industrial equipment.

[0064] Among them, the sequence observation unit is used to detect time slots. k The observation sequence. Specifically, through the observation sequence. Figure 2 Analysis of the observed sequences monitored in the data shows that it is an independent Gaussian random variable. , in, Represents the statistical average of the observed sequence. Let represent the variance of the observed sequence, and and It is known that when the system's operating environment is in a normal state, in the time slot... k Observations The statistical average is The variance is At the point of parameter change The system's operating environment state jumps from a normal state to an abnormal state and remains in this state, causing the observed values ​​to... The statistical average from Become Note the variance at this point. Unchanged. Since changes cannot be visually identified through the observed sequence, anomalies can be identified using the QCD algorithm.

[0065] The statistical calculation and processing unit is used to detect whether the operating environment state of the system has changed and to obtain the statistics of preset parameters at the current time. Simultaneously, to reflect the degree of change in the controlled equipment and reduce subsequent data processing, multiple statistical states can be pre-set to characterize the changing state of the controlled equipment. Different statistical states correspond to different state values. Therefore, when the statistics of the preset parameters are obtained, the statistics can be converted into corresponding state values, and these state values ​​can be used to characterize the corresponding statistical state. Specifically, it can be used... This represents the statistics obtained through the QCD algorithm, which can then be adjusted according to the actual situation. Replace with the corresponding state value. Set time slot. n The statistical state is , ,in This represents the state value corresponding to each statistic state. h For a predefined threshold. In each time slot, the statistical state can be transmitted to the controller via a limited bandwidth channel.

[0066] The real-time control unit is used to generate real-time control actions based on the controllable state and the statistical state of preset parameters. Specifically, each time slot can be... n The controllable state is represented as , ,in, This represents the initial controllable state (maximum controllable state). This represents the minimum controllable state. Represents the adjacent states of the minimum controllable state. This represents the adjacent states to the initial controllable state. The controller operates within the time slot. n Simultaneously, the state is quantified using cumulative and statistical methods. and controllable state Output real-time control actions Real-time control actions can be represented as a , Where -1 represents transitioning to a smaller adjacent controllable state, 0 represents maintaining the current state, and 1 represents transitioning to a larger adjacent controllable state. If we use... b This represents the two-tuple (statistical state) observed by the controller. and controllable state If the state is ) Therefore, by determining the state b Select action a The probability of this can be used to determine how to select and implement the control action; that is, it can be determined by calculating... , Determine the observable state b Select action a The probability of.

[0067] The controllable state unit is used to transfer the controllable state of the device to an adjacent state or maintain the current state according to the real-time control action, so as to achieve higher benefits or lower risks in the system, and at the same time, it will also feed back the updated controllable state to the real-time control unit.

[0068] The embodiments of this disclosure will be further explained below with reference to the accompanying drawings.

[0069] Figure 1 This is a flowchart illustrating an industrial equipment control method for early warning anomaly detection and risk avoidance according to an exemplary embodiment of this disclosure, with reference to... Figure 1 This method can be executed by the controller and may include the following steps:

[0070] S101: Obtain the controllable state and the statistical state of preset parameters of the device to be controlled at the current moment. The statistical state is used to characterize the changing state of the device to be controlled.

[0071] It should be understood that the controllable state of equipment refers to the state in which the equipment can be adjusted and changed by manual or automatic control. These states can include the equipment's operating mode, parameter settings, and on / off status. For example, the controllable state of a light bulb can include its on and off states. For more complex equipment, such as robots or air conditioning systems, the controllable states may be more diverse. For instance, the controllable states of a robot might include its movement speed, direction of movement, and arm position. The controllable states of an air conditioning system might include its temperature setpoint, fan speed, and operating mode.

[0072] It should also be understood that the preset parameters can be selected according to the actual situation, and the embodiments disclosed herein do not impose any limitations on this. In possible implementations, the preset parameters may be motion speed or wind force, etc.

[0073] Furthermore, it should be understood that the changing state of the controlled device can be an abnormal changing state, a normal changing state, and / or a stable state. A stable state means that the state of the controlled device does not change. Abnormal changing state / normal changing state can be a general term for all abnormal changing states / normal changing states, or it can be classified according to the degree of change to obtain abnormal changing states / normal changing states of different degrees. For example, abnormal changing states can be divided into mild abnormal changing states, moderate abnormal changing states, and severe abnormal changing states, etc., and this disclosure does not impose any limitations on this.

[0074] Furthermore, it should be understood that unless the state of the controlled device changes significantly within a short period of time, it is difficult to detect the change immediately. Therefore, this embodiment determines whether the state of the controlled device has changed by monitoring the statistics of preset parameters. Simultaneously, to reflect the degree of change in the controlled device, a statistical status is also set, so that after obtaining the statistics, the statistical status can be determined based on the magnitude of the statistics.

[0075] The statistics of the controlled device can be used to detect abnormal changes, to detect changes in real time, or to detect changes quickly; this disclosure does not limit this. In a possible implementation, to facilitate rapid detection of changes in the state of the controlled device, a statistic for rapid change detection can be used as the statistic for the preset parameter; that is, the CuSUM statistic can be used as the statistic for the preset parameter.

[0076] Once the statistics of the device to be controlled are determined, the parameter values ​​of the preset parameters of the device to be controlled can be obtained, and the statistics of the device to be controlled can be obtained based on the parameter values.

[0077] The statistical quantities of the controlled device obtained based on the parameter values ​​can be obtained using statistical quantity acquisition methods in related technologies or using improved statistical quantity acquisition methods. This disclosure does not impose any restrictions on this method.

[0078] In a possible implementation, obtaining the statistical status of the device to be controlled at the current moment may include:

[0079] Based on a preset probability density function, the log-likelihood ratio of the preset parameter at each moment within a first time period is determined. The probability density function includes a first probability density function and a second probability density function. The first probability density function characterizes the probability distribution of the observed value of the preset parameter before the point of change, and the second probability density function characterizes the probability distribution of the observed value of the preset parameter after the point of change. The point of change is the time point at which the statistical average of the observed value changes. The first time period characterizes the time period from the start time to the current time. For each second time period, a cumulative log-likelihood ratio is determined. This cumulative log-likelihood ratio characterizes the sum of the log-likelihood ratios of the preset parameter at each moment within the second time period. The second time period characterizes the time period from each moment within the first time period to the current time. The largest cumulative log-likelihood ratio is determined as the statistic of the preset parameter. Based on the statistic and the preset correspondence between the statistic state and the statistic, the statistical state of the controlled device at the current moment is determined.

[0080] For example, it can be used Let represent the probability density function, where Denotes the first probability density function. Let represent the second probability density function. Then... k The log-likelihood ratio of the preset parameters at time t is:

[0081]

[0082] in, Indicates the preset parameters in k Log-likelihood ratio at time step.

[0083] If you need to obtain preset parameters k When calculating statistics at a given time, it is necessary to find a specific time. , This makes from time j At that time k The log-likelihood ratio is maximized by the cumulative sum. To obtain the time... j The time can be calculated. k Each previous moment to moment kThe log-likelihood ratio cumulative sum is calculated, and by comparing the magnitudes of each log-likelihood ratio cumulative sum, the largest log-likelihood ratio cumulative sum and its corresponding time are obtained. j In other words, time k The statistic can be determined by the following formula:

[0084]

[0085]

[0086] in, Indicates time k Statistical measure Indicates from time j At that time k The log-likelihood is greater than the cumulative sum. Indicates at time t The log-likelihood ratio.

[0087] Finally, based on the time... k The statistical quantities and the preset correspondence between statistical quantity states are used to determine the statistical quantity state of the device to be controlled at the current moment.

[0088] The preset correspondence between statistical states and statistical quantities can be set according to actual conditions, and this embodiment does not impose any restrictions on this. In possible implementations, the preset correspondence between statistical states and statistical quantities may be as shown in Table 1.

[0089] Table 1 Preset Correspondence

[0090]

[0091] As shown in Table 1, the statistical state can include a first state, a second state, and a third state, with state values ​​of 0, 1, and 2 corresponding to the first, second, and third states, respectively. When the statistic is greater than or equal to 0 and less than 0.5, it corresponds to the first state, and the state value 0 is used to replace the statistic. When the statistic is greater than or equal to 0.5 and less than 1.5, it corresponds to the second state, and the state value 1 is used to replace the statistic. When the statistic is greater than or equal to 1.5 and less than 2.5, it corresponds to the third state, and the state value 2 is used to replace the statistic.

[0092] It should be understood that the method of obtaining the preset parameter statistics here is merely illustrative and does not constitute a limitation on the scheme. In possible implementations, the preset parameter statistics can also be obtained based on the Shiryaev statistic, the Shiryaev-Roberts statistic (SR), the generalized likelihood ratio (GLR), and the natural log-likelihood ratio statistic in the Gaussian mixture model (GMM).

[0093] S102: Input the controllable state and the statistical state into the control action decision model. With the goal of maximizing the benefit, determine the selection probability of each preset control action through the objective function and constraints of the control action decision model.

[0094] The revenue can be set according to the actual situation. For example, it can be the average revenue, the expected revenue, or the maximum revenue. This embodiment of the disclosure does not impose any restrictions on this.

[0095] It should be understood that in industrial environments, optimizing performance and efficiency often involves operating equipment and systems closer to their safe operating boundaries. This may lead to increased potential benefits, but it may also increase the risk of infrastructure damage or failure. Properly balancing benefits and risks is crucial in ensuring the long-term reliability and safety of industrial equipment and processes. Therefore, to capture this balance, a benefit function can be established. This payoff function is used to analyze and understand the impact of different strategies or decisions.

[0096] In addition, to assess the risk, a damage parameter related to the equipment to be controlled can also be considered. d Furthermore, after the controlled equipment suddenly undergoes abnormal changes, the damage parameters... d With rate v The change is assumed here to be monotonically decreasing. Among them, the damage parameter... d The initial value can be set to Meanwhile, to prevent damage to the controlled equipment, it is necessary to ensure that the damage parameter remains within a controllable range at any given time. Therefore, in a possible implementation, the reward function can be expressed as:

[0097]

[0098] in, express n Always under control Execute preset control actions The resulting profit value This indicates the preset values ​​of parameters related to industrial equipment damage. This represents the preset minimum controllable state. This indicates the preset initial controllable state. Denotes any constant. express n A state of controllability at any given moment. express n The target control action selected at any given time.

[0099] Furthermore, it should be understood that in an industrial environment, the goal is to maintain optimal production efficiency while keeping the risk of infrastructure damage within a given threshold, thus achieving the highest possible level of profitability under control. Therefore, the return can be set as a long-term average expected return, and used... Indicating in strategy The long-term average expected return is given by Indicating in strategy The probability of damage to infrastructure, using Representation Strategy The expected total return. Taking these factors into account, the above problem can be transformed into:

[0100]

[0101] in, N This represents the total time calculated from the start time.

[0102] Furthermore, it should be understood that, given the controller's incomplete understanding of changes in the environmental state, a partially observable Markov decision model can be constructed, consisting of four element states. These four element states may include statistical states. Controllable state Change in time state and the current time status. Among them, the changing time state refers to the point in time when the controlled equipment may malfunction.

[0103] In possible implementations, it can be used S Representing the state space of a Markov decision model, using If we represent the states in the state space, then:

[0104]

[0105] in, A set representing the state of a statistical measure. Represents the set of controllable states. A set representing changing states over time. A set representing the current state at time.

[0106] In a possible implementation, the Markov decision model can consist of a Markov chain having an initial state set, an unreachable state set, a process state set, a safe process state set, and a damage-causing state set. The initial state set... This includes the initial statistical state, the initial controllable state, the change-time state, and the initial current time state, representing the starting point of the Markov chain. (Safe processing state set) This indicates the situation where, after an abnormal change, the controllable state promptly moves to a safe zone. In other words, it occurs when the statistic reaches a preset threshold. h The states that can be classified into three categories: the state of damage and the state of unreachable states. The first category includes situations where the controllable state reaches the minimum controllable state, or the current state reaches the maximum set state. Note that all three situations require the damage parameter value to be greater than the controllable state value simultaneously. The second category represents situations where the controlled equipment is damaged due to an anomaly that cannot be handled in time. The third category represents situations where the damage parameter value is less than the minimum controllable state value. The fourth category represents states that the equipment cannot reach. The fifth category represents a subset of the state space other than the four types of state sets mentioned above.

[0107] In a possible implementation, the Markov chain starts from any initial state in the initial state set and follows an initial distribution. , The system evolves through a set of processing states until it reaches either a safe processing state set or a damage-inducing state set. To perform long-term average expected return analysis, a transition can be added between the final absorbed state (damage-inducing state set, safe processing state set, and unreachable state set) and the initial state, thus creating an irreducible and aperiodic Markov chain. That is, the improved partially observable Markov system can be modeled as follows: Figure 3 As shown, it can be studied through steady-state probabilities to understand its long-term behavior. In this case, if a state in the Markov chain reaches a set of states that cause damage, a set of states that allow safe handling, or a set of unreachable states, it is transformed back to any initial state in the initial state set, and the probability of each initial state being chosen follows the initial distribution.

[0108] Since the modeled Markov chain is irreducible, each state has a steady-state probability. By utilizing these steady-state probabilities and decision function It can be calculated in the strategy and initial distribution Long-term average expected returns of industrial equipment This is used to effectively describe the long-term performance of a system under a given policy. That is, the long-run average expected return. It can be represented as:

[0109]

[0110] in, Indicates the state Select preset control action a The probability, Indicates the state Execute preset control actions a The benefits obtained.

[0111] Wherein, the steady-state probability of each state in the Markov chain One-step transition probability between quaternary states Confirmed, as shown in the following equation:

[0112]

[0113]

[0114]

[0115] in, Indicates from state to state The transition probability, Indicates from state Execute preset control actions a When the next state is obtained. The probability, Representing state The steady-state probability.

[0116] The transition probability of a statistical state can be calculated using the previous statistical state, the previous time-of-change state, and the previous current time state in the following equation:

[0117]

[0118] in, This indicates that the observed values ​​before the point of change follow a Gaussian distribution as their average. This indicates that the observed values ​​after the point of change follow a Gaussian distribution as their average. It represents the variance of the observed values ​​following a Gaussian distribution.

[0119] For the controller, there are two termination states: a safe handling state and a damage-causing state. Given a policy... and initial distribution We can calculate the probability that the device will be damaged due to the controller's failure to handle the anomaly. To capture the risks associated with the operation of industrial infrastructure under a given strategy.

[0120]

[0121] in, This indicates a state where the safety processing status is centralized. Represents a set of safe processing states. The steady-state probability represents the state of safe handling. This indicates a state in which damage is concentrated. This indicates a set of states that trigger damage. This represents the steady-state probability of triggering a damaging state.

[0122] In a possible implementation, this can be achieved by optimizing the probabilistic action selection function. With steady-state probability To maximize average expected return Therefore, the above problem can be transformed into an optimization problem, which can be concisely expressed as:

[0123]

[0124] However, the above problem presents a nonlinear programming challenge, due to the optimization variables in the objective function and constraints. and The coupling between these variables can be a significant challenge. However, this challenge can be addressed by using variable substitution, which transforms the problem into a more manageable equivalent. To simplify the problem and convert it into a linear programming (LP) problem, variable substitution can be used. Specifically, a new optimization variable can be defined. It is represented as follows:

[0125]

[0126] The optimization problem described above can then be equivalent to the following form:

[0127]

[0128] in, The optimization variable represents the state that belongs to the set of safe processing states. The state represents an optimization variable that belongs to the set of states that cause damage.

[0129] Therefore, by substituting the current controllable state and the statistical state of the preset parameters into the above formula, the selection probability of each preset control action can be obtained.

[0130] It should be understood that the controller's decisions depend on the observable binary states. Instead of a quaternary state However, applying policy equivalence constraints... Including this constraint introduces a nonlinear constraint, which can be difficult to handle. To address this issue, this embodiment also proposes a weighted average real-time control strategy to assist in computing the observable binary state strategy, thereby effectively simplifying the decision-making process.

[0131] Specifically, it can make This represents the optimal solution to the linear programming problem. In the optimal solution... Under these conditions, the steady-state probability can be determined by the following formula:

[0132]

[0133] If the steady-state probability Then we can directly start from the optimal solution. The optimal probabilistic strategy is derived from this. It can be determined according to the following formula:

[0134]

[0135] Let there be a quaternion state. With binary state Related, represented as The corresponding steady-state probability of the binary state can be determined according to the following formula:

[0136]

[0137] let The desired weighting coefficients can be determined using the following formula:

[0138]

[0139] The weighted average real-time control strategy It can be determined according to the following formula:

[0140]

[0141] Simplify the above formula and apply the weighted average real-time control strategy. The optimal solution can be obtained from the following formula. Sure:

[0142]

[0143] S103: Determine the target control action based on the preset action selection strategy and the selection probability of each preset control action.

[0144] It should be understood that the action selection strategy can be set according to the actual situation. For example, the action selection strategy can be to use the preset control action corresponding to the maximum selection probability as the target control action, or it can be to determine the target control action based on a probability sampling strategy. This disclosure does not impose any limitations on this. In order to make flexible choices in the face of uncertainty and change, in possible implementations, the target control action can be determined by a probability sampling strategy.

[0145] It should be understood that probability sampling strategies are random sampling strategies. Random sampling strategies can help make decisions under different circumstances and introduce a certain degree of randomness, thus allowing for flexible choices when facing uncertainty and change. At the same time, by using random sampling strategies each time a control action is selected, the potential effects of different control actions can be fully considered, which can, to some extent, avoid falling into fixed decision-making patterns and improve the flexibility of decision-making.

[0146] It should also be understood that probability sampling is a statistical method that makes random selections based on a given probability distribution; therefore, the selection result may not be optimal. Thus, in possible implementations, other factors, such as environment, task requirements, and prior experience, can be combined to select the optimal target control action.

[0147] S104: Control the device to be controlled according to the target control action.

[0148] In summary, based on the above technical solution, the current controllable state and the accumulated and statistical quantities of preset parameters can be input into the control action decision model to generate control actions in real time and control the equipment. This allows for timely control actions to handle abnormalities when abnormal features are detected, thereby reducing losses caused by equipment malfunctions.

[0149] To verify the feasibility of the industrial equipment control method for early warning anomaly detection and risk avoidance provided in the embodiments of this disclosure, in a possible implementation, the solution can be illustrated using a scenario of abnormal changes in energy storage batteries as an example.

[0150] Specifically, in this scenario, these batteries operate sequentially in normal and abnormal states. When a battery enters an abnormal state, its maximum withstand voltage decreases at a rate... vThe goal of this scenario is to maintain a high charging voltage to improve production efficiency, but it's crucial to ensure the battery's maximum withstand voltage does not drop below the charging voltage to prevent equipment damage and safety hazards. This method involves real-time monitoring of battery status and voltage changes, and taking appropriate measures based on the monitoring results to maintain battery operation within a safe range. In this scenario, maintaining a high charging voltage improves production efficiency, but if the maximum withstand voltage drops below the charging voltage, it could lead to equipment damage and safety hazards.

[0151] In this scenario, we examine the abnormal mean change mentioned in the technical solution. Let the changed mean be... The mean before the change was The variance is Given the extremely low false alarm rate (less than 0.1), the statistical threshold is set... h The value is set to 15, and the statistical state is divided into 11 levels, namely... The rate of decrease v of the damage-related parameter (maximum withstand voltage) is set to 1.5, and a controllable state with a state space size of 10 is defined, i.e. And set the size of the state space of the changing time state to be The size of the state space of the current time state is .

[0152] Based on the above definition, we initialize the maximum withstand voltage to 30 and set the adjustment range of the charging voltage to... In the definition of the payoff function, we will use a constant. Set to 100. And set the change time state set. and its initial distribution

[0153] ,in This indicates the point of change is at time 1, and typically also indicates the worst performance in the QCD algorithm. Meanwhile, This means maximizing profits even when no changes occur.

[0154] In the simulation, the time variation Follow different geometric distribution parameters ,For example By Setting different values ​​allows for varying degrees of risk avoidance strategies, thereby improving performance. Results under different strategies were obtained through over 4000 independent simulations. For comparison, a constant charging voltage can also be used to demonstrate non-interventional scenarios, such as... Figure 4 As shown.

[0155] Figure 4The results of a non-interventional strategy of related technologies were compared. In different environments, high production gains can be achieved without intervention by setting a high constant charging voltage. However, when an anomaly occurs, there may not be enough time to reduce the charging voltage to a safe range after detecting the anomaly alarm and then adjusting the controllable state, resulting in a damage rate of 97.2%. Even at the operating point of a charging voltage state with lower gains and risks, relying solely on the aforementioned sequential solution without extraction intervention still leads to a damage rate as high as 70%, which is the deficiency of related technologies in this scenario. In the proposed strategy, we observed that by reducing the voltage set in the equation... We can achieve a comparable risk-averse strategy without significantly reducing returns. Compared to a high-return, high-risk non-intervention strategy, the proposed strategy shows a slight decrease in returns while effectively mitigating the risk of equipment damage, demonstrating the practical effectiveness of our proposed strategy in real-world applications.

[0156] In possible implementations, experiments were also conducted to evaluate the performance of our proposed strategy in terms of the average duration of normal battery operation after abnormal changes and the degree of damage after abnormal changes, such as Figure 5 As shown, the results demonstrate that the proposed method outperforms the non-intervention strategy. The average duration of normal battery operation after anomaly indicates that the battery reaches the point of change... Afterwards, the voltage state is maintained within the safe range. The time interval is within 10 seconds. Figure 5 (a) shows how we lower the threshold The average duration of normal battery operation increased. Compared to the high-risk, high-reward non-interventional scenario (69.5s), the low-risk proposed strategy (113.5s) showed an improvement of 163.31%, providing more time for inspection and maintenance while reducing damage and hazards.

[0157] We further through Figure 5 (b) The experiment investigated the extent of damage after abnormal changes. The degree of battery damage depends on the voltage difference between the charging voltage and the maximum withstand voltage after the charging voltage is lower than the maximum withstand voltage. We will determine the alarm time. The degree of damage is defined as By lowering the threshold To mitigate risk, our proposed strategy effectively reduced the damage level from 12.8 to 4.45.

[0158] In a possible implementation, a strategy table can be designed by observing the changing patterns of the cumulative sum statistic and the controllable state. Based on this strategy table, the CuSUM statistic, and the controllable state, a corresponding fixed strategy can be determined. The strategy table is as follows: Figure 6 As shown.

[0159] Based on the same concept, this disclosure also provides an industrial equipment control device for early warning anomaly detection and risk avoidance, such as... Figure 7 As shown, the device 700 may include:

[0160] The acquisition module 701 is used to acquire the controllable state and the statistical state of preset parameters of the device to be controlled at the current moment, wherein the statistical state is used to characterize the changing state of the device to be controlled.

[0161] The first determining module 702 is used to input the controllable state and the statistical state into the control action decision model, with the goal of maximizing the benefit, and to determine the selection probability of each preset control action through the objective function and the constraints of the control action decision model.

[0162] The second determining module 703 is used to determine the target control action based on the preset action selection strategy and the selection probability of each preset control action;

[0163] The control module 704 is used to control the device to be controlled according to the target control action.

[0164] In a possible implementation, the acquisition module 701 may include:

[0165] The first determining unit is used to determine the log-likelihood ratio of the preset parameter at each moment within a first time period based on a preset probability density function. The probability density function includes a first probability density function and a second probability density function. The first probability density function is used to characterize the probability distribution of the observed value of the preset parameter before the point of change, and the second probability density function is used to characterize the probability distribution of the observed value of the preset parameter after the point of change. The point of change is the time point at which the statistical average value of the observed value changes. The first time period is used to characterize the time period from the start time to the current time.

[0166] The second determining unit is used to determine the cumulative sum of log-likelihood ratios for each second time period. The cumulative sum of log-likelihood ratios is used to characterize the sum of the log-likelihood ratios of the preset parameters at each moment in the second time period. The second time period is used to characterize the time period from each moment in the first time period to the current moment.

[0167] The third determining unit is used to determine the maximum cumulative sum of log-likelihood ratios as the statistic of the preset parameter;

[0168] The fourth determining unit is used to determine the statistical state of the device to be controlled at the current moment based on the statistical quantity and the preset correspondence between the statistical quantity state and the statistical quantity.

[0169] In a possible implementation, the first determining module 702 may include:

[0170] The fifth determining unit is used to determine the optimal solution of the target optimization variable and determine the optimal steady-state probability based on the optimal solution;

[0171] The sixth determining unit is used to determine the magnitude of the preset steady-state probability and the optimal steady-state probability, and when the optimal steady-state probability is greater than the preset steady-state probability, to determine a first selection probability for each preset control action based on the optimal solution;

[0172] The sixth determining unit is used to determine the selection probability of each preset control action based on the preset weighting coefficient and the first selection probability.

[0173] In a possible implementation, the objective function can be:

[0174]

[0175]

[0176]

[0177]

[0178] in, This represents a quadruple state consisting of a statistical state, a controllable state, a change-time state, and the current time state. S This represents the state space consisting of four-tuple states. a Indicates the preset control action. A This represents the action state space for preset control actions. Represents the target optimization variable. Indicates the state of the quadruple Execute preset control actions a The gains obtained Indicates taking the state of the quadruple. Current moment n controllable state , express n A state of controllability at any given moment. express n The target control action selected at any given time. express n Always under control Execute preset control actions The gains obtained Represents the steady-state probability of a state. This indicates the preset control action selection function, representing the state. Select preset control action a The probability, This indicates the preset values ​​of parameters related to industrial equipment damage. This represents the preset minimum controllable state. This indicates the preset initial controllable state. It represents an arbitrary constant.

[0179] In a possible implementation, the constraints include state constraints and safety constraints, wherein the state constraints characterize the conditions under which the system containing the controlled device reaches a steady state, and the safety constraints characterize the conditions under which the controlled device is maintained in a normal operating state.

[0180] In a possible implementation, the state constraint condition can be:

[0181]

[0182]

[0183]

[0184]

[0185] in, This represents the state space consisting of statistical states, controllable states, time-varying states, and the current time state. S This represents the state space consisting of four-tuple states. a Indicates the preset control action. A This represents the action state space for preset control actions. Represents the steady-state probability of a state. This indicates the preset control action selection function, representing the state. Select preset control action a The probability, Represents the target optimization variable. Indicates from state Execute preset control actions a When the next state is obtained. The probability, Representing state steady-state probability, Representing state The steady-state probability multiplied by the state Select preset control actiona The probability of.

[0186] In a possible implementation, the security constraint can be:

[0187]

[0188] in, A threshold representing the probability of equipment damage. This indicates that the states in the state space belong to the set of first preset states. The first preset states are used to characterize the state in which equipment malfunctions are handled safely. Indicates the first preset state. This indicates that the states in the state space belong to the set of second preset states. The second preset states are used to characterize the state where the device malfunction cannot be safely handled. This indicates the second preset state. This represents the optimization variable belonging to the second preset state. This represents the optimization variable belonging to the first preset state.

[0189] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0190] Based on the same concept, embodiments of this disclosure also provide a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the industrial equipment control method for early warning anomaly detection and risk avoidance provided in any of the first aspects of this disclosure.

[0191] Based on the same concept, embodiments of this disclosure also provide an electronic device, including:

[0192] A memory on which computer programs are stored;

[0193] A processor is configured to execute the computer program in the memory to implement the steps of the industrial equipment control method for early warning anomaly detection and risk avoidance provided in any of the first aspects of this disclosure.

[0194] Figure 8 This is a block diagram illustrating an electronic device 800 according to an exemplary embodiment. For example... Figure 8 As shown, the electronic device 800 may include a processor 801 and a memory 802. The electronic device 800 may also include one or more of a multimedia component 803, an input / output (I / O) interface 804, and a communication component 805.

[0195] The processor 801 controls the overall operation of the electronic device 800 to complete all or part of the steps in the device control method described above. The memory 802 stores various types of data to support the operation of the electronic device 800. This data may include, for example, instructions for any application or method operating on the electronic device 800, and application-related data such as contact data, sent and received messages, pictures, audio, video, etc. The memory 802 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The multimedia component 803 may include a screen and audio components. The screen may be, for example, a touchscreen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in memory 802 or transmitted via communication component 805. The audio component also includes at least one speaker for outputting audio signals. I / O interface 804 provides an interface between processor 801 and other interface modules, such as a keyboard, mouse, buttons, etc. These buttons may be virtual or physical buttons. Communication component 805 is used for wired or wireless communication between the electronic device 800 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G, NB-IoT, eMTC, or other 5G technologies, or combinations thereof, is not limited here. Therefore, the corresponding communication component 805 may include: a Wi-Fi module, a Bluetooth module, an NFC module, etc.

[0196] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the device control method described above.

[0197] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, which, when executed by a processor, implement the steps of the device control method described above. For example, the computer-readable storage medium may be the memory 802 including program instructions described above, which may be executed by the processor 801 of the electronic device 800 to complete the device control method described above.

[0198] The preferred embodiments of this disclosure have been described in detail above with reference to the accompanying drawings. However, this disclosure is not limited to the specific details of the above embodiments. Within the scope of the technical concept of this disclosure, various simple modifications can be made to the technical solutions of this disclosure, and these simple modifications all fall within the protection scope of this disclosure.

[0199] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, this disclosure will not describe the various possible combinations separately.

[0200] Furthermore, various different embodiments of this disclosure can be combined in any way, as long as they do not violate the spirit of this disclosure, they should also be regarded as the content disclosed in this disclosure.

Claims

1. A method for controlling industrial equipment with early warning anomaly detection and risk avoidance, characterized in that, The method includes: The process involves acquiring the controllable state of the device under control and the statistical state of preset parameters at the current moment. These statistical states characterize the changing state of the device under control. Acquiring the statistical state of the preset parameters at the current moment includes: determining the log-likelihood ratio of the preset parameters at each moment within a first time period based on a preset probability density function. The probability density function includes a first probability density function and a second probability density function. The first probability density function characterizes the probability distribution of observed values ​​of the preset parameters before the point of change, and the second probability density function characterizes the probability distribution of observed values ​​of the preset parameters after the point of change. The change point is the time point at which the statistical average of the observed values ​​changes. The first time period is used to characterize the time period from the start time to the current time. For each second time period, a cumulative log-likelihood ratio is determined. The cumulative log-likelihood ratio is used to characterize the sum of the log-likelihood ratios of the preset parameter at each time point within the second time period. The second time period is used to characterize the time period from each time point within the first time period to the current time. The largest cumulative log-likelihood ratio is determined as the statistic of the preset parameter. Based on the statistic and the preset correspondence between the statistic state and the statistic, the statistical state of the device to be controlled at the current time is determined. The controllable state and the statistical state are input into the control action decision model. With the goal of maximizing profit, the probability of selecting each preset control action is determined through the objective function and constraints of the control action decision model. The objective function is: in, This represents a quadruple state consisting of a statistical state, a controllable state, a change-time state, and the current time state. S This represents the state space consisting of four-tuple states. a Indicates the preset control action. A This represents the action state space for preset control actions. Represents the target optimization variable. Indicates the state of the quadruple Execute preset control actions a The gains obtained Indicates taking the state of the quadruple. Current moment n controllable state , express n A state of controllability at any given moment. express n The target control action selected at any given time. express n Always under control Execute preset control actions The gains obtained Represents the steady-state probability of a state. This indicates the preset control action selection function, representing the state. Select preset control action The probability, This indicates the preset values ​​of parameters related to industrial equipment damage. This represents the preset minimum controllable state. This indicates the preset initial controllable state. Let represent any constant; the step of determining the selection probability of each preset control action with the objective of maximizing profit, through the objective function and constraints of the control action decision model, includes: determining the optimal solution of the objective optimization variable, and determining the optimal steady-state probability based on the optimal solution; determining the magnitude of the preset steady-state probability and the optimal steady-state probability, and when the optimal steady-state probability is greater than the preset steady-state probability, determining the first selection probability of each preset control action based on the optimal solution; and determining the selection probability of each preset control action based on the preset weight coefficient and the first selection probability. The target control action is determined based on the preset action selection strategy and the selection probability of each preset control action; The device to be controlled is controlled according to the target control action.

2. The method according to claim 1, characterized in that, The constraints include state constraints and safety constraints. The state constraints characterize the conditions under which the system containing the controlled device reaches a steady state, and the safety constraints characterize the conditions under which the controlled device is maintained in a normal operating state.

3. The method according to claim 2, characterized in that, The state constraint condition is: in, This represents the state space, which consists of statistical states, controllable states, time-varying states, and the current time state. S This represents the state space consisting of four-tuple states. a Indicates the preset control action. A This represents the action state space for preset control actions. Represents the steady-state probability of a state. Select a function for the preset control action, indicating the state. Select preset control action a The probability, Represents the target optimization variable. Indicates from state Execute preset control actions a When the next state is obtained. The probability, Representing state The steady-state probability, Representing state The steady-state probability multiplied by the state Select preset control action a The probability of.

4. The method according to claim 2, characterized in that, The security constraints are as follows: in, A threshold representing the probability of equipment damage. This indicates that the states in the state space belong to the set of first preset states. The first preset states are used to characterize the state in which equipment malfunctions are handled safely. Indicates the first preset state. This indicates that the states in the state space belong to the set of second preset states. The second preset states are used to characterize the state where the device malfunction cannot be safely handled. This indicates the second preset state. This represents the optimization variable belonging to the second preset state. This represents the optimization variable belonging to the first preset state.

5. An industrial equipment control device for early warning anomaly detection and risk avoidance, characterized in that, The device includes: The acquisition module is used to acquire the controllable state and statistical status of preset parameters of the device to be controlled at the current moment. The statistical status is used to characterize the changing state of the device to be controlled. Acquiring the statistical status of the preset parameters of the device to be controlled at the current moment includes: determining the log-likelihood ratio of the preset parameters at each moment within a first time period based on a preset probability density function. The probability density function includes a first probability density function and a second probability density function. The first probability density function characterizes the probability distribution of the observed value of the preset parameter before the point of change, and the second probability density function characterizes the probability distribution of the observed value of the preset parameter after the point of change. The probability distribution is defined as follows: the point of change is the time point at which the statistical average of the observed values ​​changes; the first time period is used to characterize the time period from the start time to the current time; for each second time period, the cumulative log-likelihood ratio is determined, which is used to characterize the sum of the log-likelihood ratios of the preset parameter at each time point within the second time period; the second time period is used to characterize the time period from each time point within the first time period to the current time; the largest cumulative log-likelihood ratio is determined as the statistic of the preset parameter; based on the statistic and the preset correspondence between the statistic state and the statistic, the statistical state of the device to be controlled at the current time is determined; The first determining module is used to input the controllable state and the statistical state into the control action decision model, with the objective of maximizing the benefit, and to determine the selection probability of each preset control action through the objective function and constraints of the control action decision model; the objective function is: in, This represents a quadruple state consisting of a statistical state, a controllable state, a change-time state, and the current time state. S This represents the state space consisting of four-tuple states. a Indicates the preset control action. A This represents the action state space for preset control actions. Represents the target optimization variable. Indicates the state of the quadruple Execute preset control actions a The gains obtained Indicates taking the state of the quadruple. Current moment n controllable state , express n A state of controllability at any given moment. express n The target control action selected at any given time. express n Always under control Execute preset control actions The gains obtained Represents the steady-state probability of a state. This indicates the preset control action selection function, representing the state. Select preset control action The probability, This indicates the preset values ​​of parameters related to industrial equipment damage. This represents the preset minimum controllable state. This indicates the preset initial controllable state. Let represent any constant; the step of determining the selection probability of each preset control action with the objective of maximizing profit, through the objective function and constraints of the control action decision model, includes: determining the optimal solution of the objective optimization variable, and determining the optimal steady-state probability based on the optimal solution; determining the magnitude of the preset steady-state probability and the optimal steady-state probability, and when the optimal steady-state probability is greater than the preset steady-state probability, determining the first selection probability of each preset control action based on the optimal solution; and determining the selection probability of each preset control action based on the preset weight coefficient and the first selection probability. The second determining module is used to determine the target control action based on the preset action selection strategy and the selection probability of each preset control action; The control module is used to control the device to be controlled according to the target control action.

6. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method described in any one of claims 1-4.

7. An electronic device, characterized in that, include: A memory on which computer programs are stored; A processor for executing the computer program in the memory to implement the steps of the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Decision model training method, and strategy control method and device of target object

    CN115238891A