Real-time monitoring alarm method, device, electronic device and storage medium

By dynamically adjusting the window length and reward value calculation, the problems of false alarms and missed alarms in existing monitoring methods are solved, enabling more accurate alarm status judgment and improving the stability and safety of industrial processes.

CN119002323BActive Publication Date: 2025-10-28TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310559953.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-17
Publication Date
2025-10-28
Estimated Expiration
2043-05-17

Smart Images

  • Figure CN119002323B_ABST
    Figure CN119002323B_ABST
Patent Text Reader

Abstract

This disclosure relates to a real-time monitoring and alarm method, apparatus, electronic device, and storage medium. The method involves repeatedly acquiring monitored feature values ​​during the production process at preset time intervals as target feature values. An initial window length for the target feature value is determined, and monitoring feature values ​​acquired within the window length prior to the target feature value are identified as reference feature values. A reference action corresponding to the target feature value is determined based on the reward value corresponding to each reference feature value. The window length is adjusted based on the reference action corresponding to the target feature value to update the reference feature value. A cumulative alarm reward and a cumulative normal reward are determined based on the reward value corresponding to the updated reference feature value to determine the alarm state corresponding to the target feature value. This disclosure determines the alarm state for the next period based on feature values ​​whose alarm states have already been determined by setting reward values. This method considers multiple monitoring feature values ​​for a given period when determining the alarm state, avoiding false alarms caused by instantaneous feature value changes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of secure computing, and more particularly to a real-time monitoring and alarm method, apparatus, electronic device, and storage medium. Background Technology

[0002] With increasing production demands and technological advancements, process industries, such as steel, metallurgy, energy, and chemicals, are gradually becoming larger-scale and more complex. Ensuring the stable and normal operation of complex industrial processes is crucial for reducing production costs and improving efficiency and effectiveness. Acquiring process variables (indicators) reflecting production status through various sensing methods and analyzing this data to determine whether production is normal is a vital guarantee for stable industrial production. In actual production, traditional process monitoring methods mainly rely on manual threshold judgments for a few key indicators. However, as the complexity of production processes increases, the dimensions of monitoring indicators also increase, making manual monitoring insufficient. Currently, some alarm-based automatic monitoring methods have emerged, but these primarily compare real-time values ​​of monitoring statistics with control limits, triggering an alarm when the limits are exceeded. This approach is prone to false alarms and missed alarms. Summary of the Invention

[0003] In view of this, this disclosure proposes a real-time monitoring alarm method, device, electronic equipment, and storage medium, which are designed to accurately monitor the safety of the production process.

[0004] According to a first aspect of this disclosure, a real-time monitoring and alarm method is provided, the method comprising:

[0005] The monitoring feature values ​​monitored during the production process are acquired multiple times according to a preset time period, and the currently acquired monitoring feature value is determined to be the target feature value each time the monitoring feature value is acquired.

[0006] Determine the initial window length of the target feature value, and use the monitoring feature values ​​obtained within the window length before the target feature value as reference feature values;

[0007] A reference action corresponding to the target feature value is determined based on the reward value corresponding to each reference feature value. The reward value is determined based on the reference action corresponding to the reference feature value. The reference action includes alarm action and normal action.

[0008] The window length is adjusted according to the reference action corresponding to the target feature value, and the corresponding reference feature value is updated according to the adjusted window length;

[0009] The cumulative alarm reward and cumulative normal reward are determined based on the reward value corresponding to the updated reference feature value, and the alarm status corresponding to the target feature value is determined based on the cumulative alarm reward and cumulative normal reward, wherein the alarm status includes alarm and normal.

[0010] In one possible implementation, determining the initial window length of the target feature value includes:

[0011] The window length adjusted based on the reference action corresponding to the previous target feature value is determined as the initial window length for the current target feature value.

[0012] In one possible implementation, determining the reference action corresponding to the target feature value based on the reward value corresponding to each reference feature value includes:

[0013] The alarm confidence upper bound and the normal confidence upper bound corresponding to the target feature value are determined based on the reward value corresponding to each of the reference feature values;

[0014] In response to the alarm confidence upper bound being greater than the normal confidence upper bound, the reference action corresponding to the target feature value is determined to be an alarm action;

[0015] In response to the normal confidence upper bound being greater than the alarm confidence upper bound, the reference action corresponding to the target feature value is determined to be a normal action.

[0016] In one possible implementation, determining the alarm confidence upper bound and the normal confidence upper bound corresponding to the target feature value based on the reward value corresponding to each of the reference feature values ​​includes:

[0017] The alarm reward value and the normal reward value are included in the reward value corresponding to each of the aforementioned reference feature values;

[0018] The alarm confidence upper bound and the normal confidence upper bound corresponding to the target feature value are determined based on each of the alarm reward values ​​and each of the normal reward values.

[0019] In one possible implementation, determining the alarm confidence upper bound and the normal confidence upper bound corresponding to the target feature value based on each of the alarm reward values ​​and each of the normal reward values ​​includes:

[0020] The alarm estimate is determined based on the location of each alarm reward value and the corresponding reference feature value, as well as a preset forgetting factor;

[0021] The normal estimate is determined based on the position of each normal reward value and its corresponding reference feature value, as well as the forgetting factor;

[0022] An alarm metric is determined based on the forgetting factor, dynamic probability, and the number of alarm reward values ​​obtained statistically, wherein the dynamic probability is determined based on the window length corresponding to the target feature value.

[0023] The normal metric value is determined based on the forgetting factor, the dynamic probability, and the number of normal reward values ​​obtained statistically.

[0024] The product of the alarm metric and the preset control coefficient is calculated, and the sum of this product and the alarm estimate is used as the upper limit of the alarm confidence value.

[0025] The product of the normal metric and the preset control coefficient is calculated, and the sum of this product and the normal estimate is used as the upper bound of the normal confidence value.

[0026] In one possible implementation, determining the alarm estimate based on the position of each alarm reward value and its corresponding reference feature value, and a preset forgetting factor, includes:

[0027] Through formula Calculate the alarm estimate, where λ is the forgetting factor and N K (a1) represents the number of alarm reward values ​​obtained statistically, i represents the position of the reference feature value corresponding to the alarm estimate, and r represents the position of the reference feature value. a1 The alarm reward value is the reference feature value.

[0028] In one possible implementation, determining the normal estimate based on each normal reward value, the corresponding reference feature value position, and the forgetting factor includes:

[0029] Through formula Calculate the alarm estimate, where λ is the forgetting factor and N K (a2) represents the number of normal reward values ​​obtained statistically, where i is the position of the reference feature value corresponding to the normal estimated value, and r is the number of normal reward values ​​obtained statistically. a2 This is the normal reward value corresponding to the reference feature value.

[0030] In one possible implementation, determining the alarm metric based on the forgetting factor, dynamic probability, and the statistically obtained number of alarm reward values ​​includes:

[0031] Through formula Calculate the alarm metric value, where p is the dynamic probability value.

[0032] In one possible implementation, determining the normality metric based on the forgetting factor, the dynamic probability, and the statistically obtained number of normal reward values ​​includes:

[0033] Through formula Calculate the normal metric, where p is the dynamic probability value.

[0034] In one possible implementation, adjusting the window length based on the reference action corresponding to the target feature value includes:

[0035] In response to the fact that the reference action corresponding to the target feature value is the same as the reference action corresponding to the previous target feature value, it is determined that the window length will be extended by 1;

[0036] In response to the fact that the reference action corresponding to the target feature value is different from the reference action corresponding to the previous target feature value, the length of the window is scaled according to the preset scaling parameters.

[0037] In one possible implementation, scaling the window length according to preset scaling parameters includes:

[0038] Calculate the product of the window length and the scaling parameter, and round it to obtain the candidate length;

[0039] In response to the candidate length being greater than the preset minimum length, the candidate length is determined to be the adjusted window length;

[0040] In response to the candidate length being less than the shortest length, the shortest length is determined to be the adjusted window length.

[0041] In one possible implementation, the method further includes:

[0042] The corresponding reward value is calculated based on the target feature value and the reference action.

[0043] In one possible implementation, calculating the corresponding reward value based on the target feature value and the reference action includes:

[0044] In response to the reference action being an alarm action, the formula 1 / (1+e) is used. Stat-Stat_UCL Calculate the reward value corresponding to the target feature value, where Stat is the target feature value and Stat_UCL is a preset feature value threshold.

[0045] In one possible implementation, the step of calculating the corresponding reward value based on the target feature value and the reference action further includes:

[0046] In response to the reference action being a normal action, the formula 1-1 / (1+e) is used. Stat-Stat_UCL Calculate the reward value corresponding to the target feature value.

[0047] According to a second aspect of this disclosure, a real-time monitoring and alarm device is provided, the device comprising:

[0048] The target feature value determination module is used to acquire monitoring feature values ​​monitored during the production process multiple times according to a preset time period, and determine the currently acquired monitoring feature value as the target feature value each time the monitoring feature value is acquired;

[0049] The reference feature value determination module is used to determine the initial window length of the target feature value and to use the monitoring feature values ​​obtained within the window length before the target feature value as reference feature values.

[0050] An action determination module is used to determine a reference action corresponding to the target feature value based on a reward value corresponding to each reference feature value. The reward value is determined based on the reference action corresponding to the reference feature value. The reference action includes an alarm action and a normal action.

[0051] The window adjustment module is used to adjust the window length according to the reference action corresponding to the target feature value, and update the corresponding reference feature value according to the adjusted window length;

[0052] The monitoring and alarm module is used to determine the cumulative alarm reward and cumulative normal reward based on the reward value corresponding to the updated reference feature value, and to determine the alarm status corresponding to the target feature value based on the cumulative alarm reward and cumulative normal reward, wherein the alarm status includes alarm and normal.

[0053] In one possible implementation, the reference feature value determination module is further configured to:

[0054] The window length adjusted based on the reference action corresponding to the previous target feature value is determined as the initial window length for the current target feature value.

[0055] In one possible implementation, the action determination module is further configured to:

[0056] The alarm confidence upper bound and the normal confidence upper bound corresponding to the target feature value are determined based on the reward value corresponding to each of the reference feature values;

[0057] In response to the alarm confidence upper bound being greater than the normal confidence upper bound, the reference action corresponding to the target feature value is determined to be an alarm action;

[0058] In response to the normal confidence upper bound being greater than the alarm confidence upper bound, the reference action corresponding to the target feature value is determined to be a normal action.

[0059] In one possible implementation, the action determination module is further configured to:

[0060] The alarm reward value and the normal reward value are included in the reward value corresponding to each of the aforementioned reference feature values;

[0061] The alarm confidence upper bound and the normal confidence upper bound corresponding to the target feature value are determined based on each of the alarm reward values ​​and each of the normal reward values.

[0062] In one possible implementation, the action determination module is further configured to:

[0063] The alarm estimate is determined based on the location of each alarm reward value and the corresponding reference feature value, as well as a preset forgetting factor;

[0064] The normal estimate is determined based on the position of each normal reward value and its corresponding reference feature value, as well as the forgetting factor;

[0065] An alarm metric is determined based on the forgetting factor, dynamic probability, and the number of alarm reward values ​​obtained statistically, wherein the dynamic probability is determined based on the window length corresponding to the target feature value.

[0066] The normal metric value is determined based on the forgetting factor, the dynamic probability, and the number of normal reward values ​​obtained statistically.

[0067] The product of the alarm metric and the preset control coefficient is calculated, and the sum of this product and the alarm estimate is used as the upper limit of the alarm confidence value.

[0068] The product of the normal metric and the preset control coefficient is calculated, and the sum of this product and the normal estimate is used as the upper bound of the normal confidence value.

[0069] In one possible implementation, the action determination module is further configured to:

[0070] Through formula Calculate the alarm estimate, where λ is the forgetting factor and N K (a1) represents the number of alarm reward values ​​obtained statistically, i represents the position of the reference feature value corresponding to the alarm estimate, and r represents the position of the reference feature value. a1 The alarm reward value is the reference feature value.

[0071] In one possible implementation, the action determination module is further configured to:

[0072] Through formula Calculate the alarm estimate, where λ is the forgetting factor and N K (a2) represents the number of normal reward values ​​obtained statistically, where i is the position of the reference feature value corresponding to the normal estimated value, and r is the number of normal reward values ​​obtained statistically. a2 This is the normal reward value corresponding to the reference feature value.

[0073] In one possible implementation, the action determination module is further configured to:

[0074] Through formula Calculate the alarm metric value, where p is the dynamic probability value.

[0075] In one possible implementation, the action determination module is further configured to:

[0076] Through formula Calculate the normal metric, where p is the dynamic probability value.

[0077] In one possible implementation, the window adjustment module is further configured to:

[0078] In response to the fact that the reference action corresponding to the target feature value is the same as the reference action corresponding to the previous target feature value, it is determined that the window length will be extended by 1;

[0079] In response to the fact that the reference action corresponding to the target feature value is different from the reference action corresponding to the previous target feature value, the length of the window is scaled according to the preset scaling parameters.

[0080] In one possible implementation, the window adjustment module is further configured to:

[0081] Calculate the product of the window length and the scaling parameter, and round it to obtain the candidate length;

[0082] In response to the candidate length being greater than the preset minimum length, the candidate length is determined to be the adjusted window length;

[0083] In response to the candidate length being less than the shortest length, the shortest length is determined to be the adjusted window length.

[0084] In one possible implementation, the device further includes:

[0085] The reward value calculation module is used to calculate the corresponding reward value based on the target feature value and the reference action.

[0086] In one possible implementation, the reward value calculation module is further configured to:

[0087] In response to the reference action being an alarm action, the formula 1 / (1+e) is used. Stat-Stat_UCL Calculate the reward value corresponding to the target feature value, where Stat is the target feature value and Stat_UCL is a preset feature value threshold.

[0088] In one possible implementation, the reward value calculation module is further configured to:

[0089] In response to the reference action being a normal action, the formula 1-1 / (1+e) is used. Stat-Stat_UCL Calculate the reward value corresponding to the target feature value.

[0090] According to a third aspect of this disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to implement the above-described method when executing instructions stored in the memory.

[0091] According to a fourth aspect of this disclosure, a non-volatile computer-readable storage medium is provided that stores computer program instructions thereon, wherein the computer program instructions, when executed by a processor, implement the above-described method.

[0092] According to a fifth aspect of this disclosure, a computer program product is provided, including computer-readable code or a non-volatile computer-readable storage medium carrying the computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device performs the above-described method.

[0093] In this embodiment, monitoring feature values ​​observed during the production process are acquired multiple times at preset time intervals as target feature values. An initial window length for the target feature value is determined, and monitoring feature values ​​acquired within the window length preceding the target feature value are identified as reference feature values. A reference action corresponding to the target feature value is determined based on the reward value corresponding to each reference feature value. The window length is adjusted to update the reference feature value based on the reference action corresponding to the target feature value. The cumulative alarm reward and cumulative normal reward are determined based on the reward value corresponding to the updated reference feature value to determine the alarm status corresponding to the target feature value. This disclosure determines the alarm status for the next period based on feature values ​​whose alarm status has already been determined by setting reward values. This method considers multiple monitoring feature values ​​for a given period when determining the alarm status, avoiding false alarms caused by instantaneous feature value changes.

[0094] Further features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. Attached Figure Description

[0095] The accompanying drawings, which are included in and form part of this specification, illustrate exemplary embodiments, features, and aspects of this disclosure together with the specification and serve to explain the principles of this disclosure.

[0096] Figure 1 A flowchart of a real-time monitoring and alarm method according to an embodiment of the present disclosure is shown;

[0097] Figure 2 A schematic diagram showing a target feature value and a reference feature value according to an embodiment of the present disclosure is shown;

[0098] Figure 3 A schematic diagram of a real-time monitoring and alarm device according to an embodiment of the present disclosure;

[0099] Figure 4 A schematic diagram of an electronic device according to an embodiment of the present disclosure is shown;

[0100] Figure 5 A schematic diagram of another electronic device according to an embodiment of the present disclosure is shown. Detailed Implementation

[0101] Various exemplary embodiments, features, and aspects of this disclosure will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.

[0102] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.

[0103] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.

[0104] The real-time monitoring and alarm method of this disclosure can be executed by electronic devices such as terminal devices or servers. The terminal device can be any fixed or mobile terminal, such as user equipment (UE), mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, vehicle-mounted device, or wearable device. The server can be a single server or a server cluster consisting of multiple servers. Any electronic device can implement the real-time monitoring and alarm method of this disclosure by having its processor call computer-readable instructions stored in its memory.

[0105] Figure 1 A flowchart illustrating a real-time monitoring and alarm method according to an embodiment of this disclosure is shown. Figure 1 As shown, the real-time monitoring and alarm method of this disclosure embodiment may include the following steps S10-S50.

[0106] Step S10: Acquire the monitoring feature values ​​monitored during the production process multiple times according to a preset time period, and determine the currently acquired monitoring feature value as the target feature value each time the monitoring feature value is acquired.

[0107] In one possible implementation, the real-time monitoring and alarm method of this disclosure can be applied to real-time monitoring of industrial production processes in any field such as steel, metallurgy, energy, and chemicals. This allows for timely alarms when equipment malfunctions, component failures, or improper operations occur during production, prompting staff to handle emergencies and prevent accidents. Specifically, the electronic equipment can repeatedly acquire monitoring feature values ​​representing the production status at preset time intervals and determine whether an alarm is needed each time a monitoring feature value is acquired based on a preset alarm strategy. The monitoring feature values ​​can be sensor data detected by sensors during production, or numerical values ​​representing the normality of the production process obtained by processing sensor data from multiple sensors.

[0108] Optionally, the electronic device acquires monitoring feature values ​​multiple times according to a preset time period, and determines the current monitoring feature value as the target feature value each time a monitoring feature value is acquired. For example, if the monitoring feature value acquired by the electronic device in a certain period is T1, then T1 is determined as the target feature value. If the electronic device acquires a new monitoring feature value of T2 in the next period, then T1 acquired in the previous period is no longer used as the target feature value, and T2 is updated as the new target feature value.

[0109] Step S20: Determine the initial window length of the target feature value, and use the monitoring feature values ​​obtained within the window length before the target feature value as reference feature values.

[0110] In one possible implementation, the electronic device can adjust the initial window length when determining the current target feature value and the alarm status at the time of acquiring the current target feature value. Specifically, if the current target feature value is the first acquired monitoring feature value, the window length corresponding to the target feature value can be 0. If the current target feature value is not the first acquired monitoring feature value, the window length corresponding to the current target feature value can inherit the adjusted window length from the previous target feature value.

[0111] Optionally, after determining the initial window length corresponding to the current target feature value, the electronic device obtains multiple monitoring feature values ​​acquired before the current target feature value based on the window length, and uses them as reference feature values ​​corresponding to the current target feature value. For example, when the window length corresponding to the current target feature value is 3, the electronic device sequentially obtains three monitoring feature values ​​before the current target feature value as corresponding reference feature values.

[0112] Figure 2 A schematic diagram illustrating a target feature value and a reference feature value according to an embodiment of the present disclosure is shown. Figure 2 As shown, the monitoring feature value acquired at the current time k is T. k In this case, the electronic device determines the current target feature value as T. k Given an initial window length of 3 corresponding to the current target feature value, the electronic device can sequentially determine three adjacent monitoring feature values ​​T starting from the current target feature value. k-1 , T k-2 and T k-3 This is the corresponding reference feature value.

[0113] Step S30: Determine the reference action corresponding to the target feature value based on the reward value corresponding to each reference feature value.

[0114] In one possible implementation, each reference feature value has a corresponding reward value, representing the impact of the reference feature value's state of choosing to alarm or not alarm at its corresponding time on the current time. After determining multiple reference feature values ​​corresponding to the current target feature value, the electronic device determines the reward value of each reference feature value to determine a reference action for the target feature value based on the reward values ​​of the multiple reference feature values. The reference actions include alarm actions and normal actions; an alarm action indicates that an alarm may be required at the corresponding time, while a normal action indicates that an alarm may not be required at the corresponding time. The reward value corresponding to each reference feature value can be determined based on the corresponding reference action when judging the alarm state of the reference feature value at the corresponding time.

[0115] Optionally, the process by which the electronic device determines the reference action for the target feature value based on the reward value corresponding to each reference feature value may include: determining an alarm confidence upper bound and a normal confidence upper bound for the target feature value based on the reward value corresponding to each reference feature value. In response to the alarm confidence upper bound being greater than the normal confidence upper bound, the reference action corresponding to the target feature value is determined to be an alarm action. In response to the normal confidence upper bound being greater than the alarm confidence upper bound, the reference action corresponding to the target feature value is determined to be a normal action. The alarm confidence upper bound characterizes the probability that the reference action corresponding to the current target feature value is an alarm action, and the normal confidence upper bound characterizes the probability that the reference action corresponding to the current target feature value is not an alarm action. That is, the electronic device can determine that the reference action corresponding to the target feature value is an alarm action when the alarm probability is high, and determine that the reference action corresponding to the target feature value is not an alarm action when the probability of not alarming is high. The reference action is only used as a reference for whether an alarm is triggered or not at the current moment to generate a reward value for determining the alarm state at the next moment, and is not directly used to determine the final alarm state.

[0116] Furthermore, the reward value corresponding to each reference feature value can be either an alarm reward value or a normal reward value. The alarm reward value is the reward value calculated when the reference action corresponding to the reference feature value is an alarm action, and the normal reward value is the reward value calculated when the reference action corresponding to the reference feature value is a normal action. The electronic device can calculate the alarm confidence upper bound and the normal confidence upper bound by statistically analyzing the alarm reward value and the normal reward value included in the reward value corresponding to the current target feature value. That is, the electronic device can first statistically analyze the alarm reward value and the normal reward value included in the reward value corresponding to each reference feature value, and then determine the alarm confidence upper bound and the normal confidence upper bound corresponding to the target feature value based on each alarm reward value and each normal reward value.

[0117] In one possible implementation, both the alarm confidence upper bound and the normal confidence upper bound can be determined using corresponding estimates, metrics, and preset control coefficients. The estimates characterize the probability that the reference action for the target feature value is the corresponding action; the metrics characterize the uncertainty that the reference action for the target feature value is the corresponding action; and the control coefficients are preset values ​​used as weights for the metrics.

[0118] For example, the electronic device can determine the alarm estimate based on the position of each alarm reward value and its corresponding reference feature value, as well as a preset forgetting factor. It then determines the alarm metric based on the forgetting factor, dynamic probability, and the statistically obtained number of alarm reward values. The product of the alarm metric and a preset control coefficient, summed with the alarm estimate, serves as the upper bound of the alarm confidence level. The position of the reference feature value corresponding to each alarm reward value is a value representing its acquisition order; that is, the reference feature value is the nth time the electronic device acquired the monitoring feature value. For example, if the corresponding reference feature value is the i-th time the electronic device acquired the monitoring feature value, the position of the reference feature value is determined as i. The preset forgetting factor is a value between 0 and 1, used to reduce the influence of reference feature values ​​far from the target feature value on the calculation results. The dynamic probability is determined based on the window length corresponding to the target feature value, decreasing as the window length increases; for example, it can be the reciprocal of the window length.

[0119] The alarm estimate can be obtained through the formula. The calculated alarm metric value can be obtained using the formula... Calculated. λ is the forgetting factor, N K (a1) represents the number of alarm reward values ​​obtained statistically, i.e., the number of reference feature values ​​corresponding to alarm reward values. i represents the position of the reference feature value corresponding to the alarm estimate, and r... a1 Let p be the alarm reward value corresponding to the reference feature value, and p be the dynamic probability value. The alarm estimate passes... This indicates that the alarm metric value is obtained through... This means that, when the control coefficient is represented by c, the calculated upper limit of the alarm confidence can be...

[0120] Furthermore, the method for determining the normal confidence upper bound is similar to that for determining the alarm confidence upper bound. That is, the electronic device can determine the normal estimate based on the location of each normal reward value and its corresponding reference feature value, as well as the forgetting factor. The normal metric is determined based on the forgetting factor, dynamic probability, and the statistically obtained number of normal reward values. The product of the normal metric and a preset control coefficient, summed with the normal estimate, serves as the normal confidence upper bound. Here, the statistically obtained normal reward value represents the number of reference feature values ​​corresponding to the normal reward value.

[0121] The normal estimate can be obtained through the formula. The calculated normal measurement value can be obtained through the formula. Calculated. λ is the forgetting factor, N K (a2) represents the number of normal reward values ​​obtained from statistics, i represents the position of the reference feature value corresponding to the normal estimated value, and r represents the number of normal reward values ​​obtained from statistics. a2 Let p be the normal reward value corresponding to the reference feature value, and p be the dynamic probability value. The normal estimate passes... This indicates that the normal measurement value is obtained through... This means that, when the control coefficient is represented by c, the calculated upper bound of the normal confidence level can be...

[0122] Optionally, after calculating the alarm confidence upper bound and the normal confidence upper bound corresponding to the target feature value, the electronic device determines the reference action corresponding to the target feature value by comparing their magnitudes. That is, if the alarm confidence upper bound is greater than the normal confidence upper bound, the reference action for the target feature value is determined to be an alarm action; if the normal confidence upper bound is greater than the alarm confidence upper bound, the reference action for the target feature value is determined to be a normal action.

[0123] In one possible implementation, after determining the reference action corresponding to the target feature value, the electronic device can also calculate the corresponding reward value based on the target feature value and the reference action, in order to determine the reference action corresponding to the next acquired monitoring feature value. Specifically, when the reference action is an alarm action, the electronic device can use formula 1 / (1+e...) Stat -Stat_UCL The reward value corresponding to the target feature value is calculated. When the reference action is a normal action, the electronic device can use formula 1-1 / (1+e) Stat-Stat_UCL The reward value corresponding to the target feature value is calculated. Stat is the target feature value, and Stat_UCL is the preset feature value threshold. Optionally, the reward value corresponding to each reference feature value can also be determined in the same way.

[0124] Step S40: Adjust the window length according to the reference action corresponding to the target feature value, and update the corresponding reference feature value according to the adjusted window length.

[0125] In one possible implementation, after the electronic device determines the reference action corresponding to the target feature value, it adjusts the current window length according to a preset adjustment rule and the reference action, and then updates the reference feature value corresponding to the target feature value after adjusting the window length. For example, if the current target feature value is T... k When the initial window length is 3, the three reference feature values ​​determined sequentially by the electronic device are T. k-1 T k-2 and T k-3 With the window length adjusted to 2, the reference feature value corresponding to the target feature value is updated to T. k-1 and T k-2 With the window length adjusted to 4, the reference feature value corresponding to the target feature value is updated to T. k-1 T k-2 T k-3 and T k-4 .

[0126] Optionally, the electronic device can adjust the window length based on changes in the reference actions corresponding to the target feature value and adjacent feature values. That is, the electronic device can compare the reference action of the current target feature value with the reference action corresponding to the previous target feature value, where the previous target feature value is the monitoring feature value acquired in the previous cycle. If the reference action corresponding to the target feature value is the same as the reference action corresponding to the previous target feature value, the window length is increased by 1. If the reference action corresponding to the target feature value is different from the reference action corresponding to the previous target feature value, the window length is scaled according to a preset scaling parameter. Specifically, if the current reference action is a normal action and the reference action corresponding to the previous target feature value is also a normal action, or if the current reference action is an alarm action and the reference action corresponding to the previous target feature value is also an alarm action, the electronic device increases the window length by 1. If the current reference action is an alarm action and the reference action corresponding to the previous target feature value is a normal action, or if the current reference action is a normal action and the reference action corresponding to the previous target feature value is an alarm action, the electronic device scales the window length according to a preset scaling parameter. The scaling parameter can be a preset value between 0 and 1.

[0127] Furthermore, the electronic device can scale the window length using scaling parameters by taking the floor function of the product of the window length and the scaling parameters to obtain a candidate length, and then determining the adjusted window length based on the size of the candidate length. Specifically, if the candidate length is greater than the preset minimum length, the candidate length is determined as the adjusted window length. If the candidate length is less than the minimum length, the minimum length is determined as the adjusted window length.

[0128] Step S50: Determine the cumulative alarm reward and cumulative normal reward based on the reward value corresponding to the updated reference feature value, and determine the alarm status corresponding to the target feature value based on the cumulative alarm reward and cumulative normal reward.

[0129] In one possible implementation, after determining the updated reference feature value, the electronic device calculates the cumulative alarm reward and the cumulative normal reward based on the reward value corresponding to the updated reference feature value. The cumulative alarm reward is the sum of alarm reward values ​​among the reward values ​​corresponding to the updated reference feature value, and the cumulative normal reward is the sum of normal reward values ​​among the reward values ​​corresponding to the updated reference feature value. After determining the cumulative alarm reward and the cumulative normal reward, the electronic device can determine the alarm status corresponding to the target feature value based on these rewards. The alarm status can include "alarm" and "normal," used to determine whether an alarm is needed at the current moment. This determination method can involve comparing the magnitudes of the cumulative alarm reward and the cumulative normal reward; if the cumulative alarm reward is greater, the alarm status is determined to be "alarm," and if the cumulative normal reward is greater, the alarm status is determined to be "normal."

[0130] Furthermore, when the electronic device determines the alarm state corresponding to the target feature value, it selects whether to alarm or not based on the alarm state. Alternatively, the electronic device can calculate the corresponding alarm state before acquiring the target feature value, and directly trigger an alarm if the alarm state is alarm-related, and continue acquiring the target feature value if the alarm state is normal.

[0131] Based on the aforementioned technical features, this embodiment of the disclosure can determine the alarm status for the next period by setting a reward value, based on multiple monitoring feature values ​​for which alarm status has already been determined. This means considering multiple monitoring feature values ​​within a time interval for alarm purposes, avoiding false alarms or missed alarms caused by instantaneous changes in feature values. Furthermore, this embodiment of the disclosure improves the accuracy of alarm status by setting a dynamic time window to dynamically acquire the reward value corresponding to the monitoring feature value that has a significant impact on the alarm status of the next period, and reduces the computational load due to the filtering of monitoring feature values.

[0132] Figure 3 A schematic diagram of a real-time monitoring and alarm device according to an embodiment of this disclosure. (See diagram below.) Figure 3As shown, the real-time monitoring and alarm device of this disclosure embodiment may include:

[0133] The target feature value determination module 30 is used to acquire the monitoring feature values ​​monitored during the production process multiple times according to a preset time period, and determine the currently acquired monitoring feature value as the target feature value each time the monitoring feature value is acquired;

[0134] The reference feature value determination module 31 is used to determine the initial window length of the target feature value and to use the monitoring feature values ​​obtained within the window length before the target feature value as reference feature values.

[0135] The action determination module 32 is used to determine the reference action corresponding to the target feature value based on the reward value corresponding to each reference feature value. The reward value is determined based on the reference action corresponding to the reference feature value. The reference action includes alarm action and normal action.

[0136] The window adjustment module 33 is used to adjust the window length according to the reference action corresponding to the target feature value, and update the corresponding reference feature value according to the adjusted window length.

[0137] The monitoring and alarm module 34 is used to determine the cumulative alarm reward and the cumulative normal reward based on the reward value corresponding to the updated reference feature value, and to determine the alarm status corresponding to the target feature value based on the cumulative alarm reward and the cumulative normal reward, wherein the alarm status includes alarm and normal.

[0138] In one possible implementation, the reference feature value determination module 34 is further configured to:

[0139] The window length adjusted based on the reference action corresponding to the previous target feature value is determined as the initial window length for the current target feature value.

[0140] In one possible implementation, the action determination module 32 is further configured to:

[0141] The alarm confidence upper bound and the normal confidence upper bound corresponding to the target feature value are determined based on the reward value corresponding to each of the reference feature values;

[0142] In response to the alarm confidence upper bound being greater than the normal confidence upper bound, the reference action corresponding to the target feature value is determined to be an alarm action;

[0143] In response to the normal confidence upper bound being greater than the alarm confidence upper bound, the reference action corresponding to the target feature value is determined to be a normal action.

[0144] In one possible implementation, the action determination module 32 is further configured to:

[0145] The alarm reward value and the normal reward value are included in the reward value corresponding to each of the aforementioned reference feature values;

[0146] The alarm confidence upper bound and the normal confidence upper bound corresponding to the target feature value are determined based on each of the alarm reward values ​​and each of the normal reward values.

[0147] In one possible implementation, the action determination module 32 is further configured to:

[0148] The alarm estimate is determined based on the location of each alarm reward value and the corresponding reference feature value, as well as a preset forgetting factor;

[0149] The normal estimate is determined based on the position of each normal reward value and its corresponding reference feature value, as well as the forgetting factor;

[0150] An alarm metric is determined based on the forgetting factor, dynamic probability, and the number of alarm reward values ​​obtained statistically, wherein the dynamic probability is determined based on the window length corresponding to the target feature value.

[0151] The normal metric value is determined based on the forgetting factor, the dynamic probability, and the number of normal reward values ​​obtained statistically.

[0152] The product of the alarm metric and the preset control coefficient is calculated, and the sum of this product and the alarm estimate is used as the upper limit of the alarm confidence value.

[0153] The product of the normal metric and the preset control coefficient is calculated, and the sum of this product and the normal estimate is used as the upper bound of the normal confidence value.

[0154] In one possible implementation, the action determination module 32 is further configured to:

[0155] Through formula Calculate the alarm estimate, where λ is the forgetting factor and N K (a1) represents the number of alarm reward values ​​obtained statistically, i represents the position of the reference feature value corresponding to the alarm estimate, and r represents the position of the reference feature value. a1 The alarm reward value is the reference feature value.

[0156] In one possible implementation, the action determination module 32 is further configured to:

[0157] Through formula Calculate the alarm estimate, where λ is the forgetting factor and N K (a2) represents the number of normal reward values ​​obtained statistically, where i is the position of the reference feature value corresponding to the normal estimated value, and r is the number of normal reward values ​​obtained statistically. a2 This is the normal reward value corresponding to the reference feature value.

[0158] In one possible implementation, the action determination module 32 is further configured to:

[0159] Through formula Calculate the alarm metric value, where p is the dynamic probability value.

[0160] In one possible implementation, the action determination module 32 is further configured to:

[0161] Through formula Calculate the normal metric, where p is the dynamic probability value.

[0162] In one possible implementation, the window adjustment module 33 is further configured to:

[0163] In response to the fact that the reference action corresponding to the target feature value is the same as the reference action corresponding to the previous target feature value, it is determined that the window length will be extended by 1;

[0164] In response to the fact that the reference action corresponding to the target feature value is different from the reference action corresponding to the previous target feature value, the length of the window is scaled according to the preset scaling parameters.

[0165] In one possible implementation, the window adjustment module 33 is further configured to:

[0166] Calculate the product of the window length and the scaling parameter, and round it to obtain the candidate length;

[0167] In response to the candidate length being greater than the preset minimum length, the candidate length is determined to be the adjusted window length;

[0168] In response to the candidate length being less than the shortest length, the shortest length is determined to be the adjusted window length.

[0169] In one possible implementation, the device further includes:

[0170] The reward value calculation module is used to calculate the corresponding reward value based on the target feature value and the reference action.

[0171] In one possible implementation, the reward value calculation module is further configured to:

[0172] In response to the reference action being an alarm action, the formula 1 / (1+e) is used. Stat-Stat_UCL Calculate the reward value corresponding to the target feature value, where Stat is the target feature value and Stat_UCL is a preset feature value threshold.

[0173] In one possible implementation, the reward value calculation module is further configured to:

[0174] In response to the reference action being a normal action, the formula 1-1 / (1+e) is used. Stat-Stat_UCL Calculate the reward value corresponding to the target feature value.

[0175] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0176] This disclosure also proposes a computer-readable storage medium storing computer program instructions that, when executed by a processor, implement the above-described method. The computer-readable storage medium can be volatile or non-volatile.

[0177] This disclosure also proposes an electronic device, including: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.

[0178] This disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device performs the above-described method.

[0179] Figure 4 A schematic diagram of an electronic device 800 according to an embodiment of the present disclosure is shown. For example, the electronic device 800 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.

[0180] Reference Figure 4 The electronic device 800 may include one or more of the following components: processing component 802, memory 804, power supply component 806, multimedia component 808, audio component 810, input / output (I / O) interface 812, sensor component 814, and communication component 816.

[0181] Processing component 802 typically controls the overall operation of electronic device 800, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the methods described above. Furthermore, processing component 802 may include one or more modules to facilitate interaction between processing component 802 and other components. For example, processing component 802 may include a multimedia module to facilitate interaction between multimedia component 808 and processing component 802.

[0182] Memory 804 is configured to store various types of data to support the operation of electronic device 800. Examples of this data include instructions for any application or method operating on electronic device 800, contact data, phonebook data, messages, pictures, videos, etc. Memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0183] Power supply component 806 provides power to various components of electronic device 800. Power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 800.

[0184] Multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0185] Audio component 810 is configured to output and / or input audio signals. For example, audio component 810 includes a microphone (MIC) configured to receive external audio signals when electronic device 800 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 804 or transmitted via communication component 816. In some embodiments, audio component 810 also includes a speaker for outputting audio signals.

[0186] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0187] Sensor assembly 814 includes one or more sensors for providing state assessments of various aspects of electronic device 800. For example, sensor assembly 814 can detect the on / off state of electronic device 800, the relative positioning of components such as the display and keypad of electronic device 800, changes in position of electronic device 800 or a component of electronic device 800, the presence or absence of user contact with electronic device 800, orientation or acceleration / deceleration of electronic device 800, and temperature changes of electronic device 800. Sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 814 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.

[0188] Communication component 816 is configured to facilitate wired or wireless communication between electronic device 800 and other devices. Electronic device 800 can access wireless networks based on communication standards, such as WiFi, 2G, or 3G, or combinations thereof. In one exemplary embodiment, communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 816 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0189] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0190] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 804 including computer program instructions that can be executed by a processor 820 of an electronic device 800 to perform the above-described method.

[0191] Figure 5 A schematic diagram of another electronic device 1900 according to an embodiment of the present disclosure is shown. For example, the electronic device 1900 may be provided as a server or a terminal device. (Refer to...) Figure 5 The electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by memory 1932 for storing instructions, such as application programs, that can be executed by the processing component 1922. The application programs stored in memory 1932 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 1922 is configured to execute instructions to perform the methods described above.

[0192] Electronic device 1900 may also include a power supply component 1926 configured to perform power management of electronic device 1900, a wired or wireless network interface 1950 configured to connect electronic device 1900 to a network, and an input / output (I / O) interface 1958. Electronic device 1900 can operate on an operating system stored in memory 1932, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or similar.

[0193] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by a processing component 1922 of an electronic device 1900 to perform the above-described method.

[0194] The present disclosure may be a system, method and / or computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0195] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through an electrical wire.

[0196] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0197] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0198] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0199] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0200] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0201] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0202] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A real-time monitoring and alarm method, characterized in that, The method includes: The monitoring feature values ​​monitored during the production process are acquired multiple times according to a preset time period, and the currently acquired monitoring feature value is determined to be the target feature value each time the monitoring feature value is acquired. Determine the initial window length of the target feature value, and use the monitoring feature values ​​obtained within the window length before the target feature value as reference feature values; A reference action corresponding to the target feature value is determined based on the reward value corresponding to each reference feature value. The reward value is determined based on the reference action corresponding to the reference feature value. The reference action includes alarm action and normal action. The window length is adjusted according to the reference action corresponding to the target feature value, and the corresponding reference feature value is updated according to the adjusted window length; The cumulative alarm reward and cumulative normal reward are determined based on the reward value corresponding to the updated reference feature value, and the alarm status corresponding to the target feature value is determined based on the cumulative alarm reward and cumulative normal reward, wherein the alarm status includes alarm and normal. Determining the initial window length for the target feature value includes: The window length adjusted based on the reference action corresponding to the previous target feature value is determined as the initial window length for the current target feature value; The step of determining the reference action corresponding to the target feature value based on the reward value corresponding to each reference feature value includes: The alarm confidence upper bound and the normal confidence upper bound corresponding to the target feature value are determined based on the reward value corresponding to each of the reference feature values; In response to the alarm confidence upper bound being greater than the normal confidence upper bound, the reference action corresponding to the target feature value is determined to be an alarm action; In response to the normal confidence upper bound being greater than the alarm confidence upper bound, the reference action corresponding to the target feature value is determined to be a normal action.

2. The method according to claim 1, characterized in that, The step of determining the alarm confidence upper bound and the normal confidence upper bound corresponding to the target feature value based on the reward value corresponding to each reference feature value includes: The alarm reward value and the normal reward value are included in the reward value corresponding to each of the aforementioned reference feature values; The alarm confidence upper bound and the normal confidence upper bound corresponding to the target feature value are determined based on each of the alarm reward values ​​and each of the normal reward values.

3. The method according to claim 2, characterized in that, The step of determining the alarm confidence upper bound and the normal confidence upper bound corresponding to the target feature value based on each alarm reward value and each normal reward value includes: The alarm estimate is determined based on the location of each alarm reward value and the corresponding reference feature value, as well as a preset forgetting factor; The normal estimate is determined based on the position of each normal reward value and its corresponding reference feature value, as well as the forgetting factor; An alarm metric is determined based on the forgetting factor, dynamic probability, and the number of alarm reward values ​​obtained statistically, wherein the dynamic probability is determined based on the window length corresponding to the target feature value. The normal metric value is determined based on the forgetting factor, the dynamic probability, and the number of normal reward values ​​obtained statistically. The product of the alarm metric and the preset control coefficient is calculated, and the sum of this product and the alarm estimate is used as the upper limit of the alarm confidence value. The product of the normal metric and the preset control coefficient is calculated, and the sum of this product and the normal estimate is used as the upper bound of the normal confidence value.

4. The method according to claim 3, characterized in that The step of determining the alarm estimate based on the position of each alarm reward value and the corresponding reference feature value, as well as a preset forgetting factor, includes: Through formula Calculate the alarm estimate, where Forgetting factor, To count the number of alarm reward values ​​obtained, i represents the position of the reference feature value corresponding to the alarm estimate. The alarm reward value is the reference feature value.

5. The method according to claim 3 or 4, characterized in that, The step of determining the normal estimate based on each normal reward value, the corresponding reference feature value position, and the forgetting factor includes: Through formula Calculate the alarm estimate, where Forgetting factor, To count the number of normal reward values ​​obtained, i represents the position of the reference feature value corresponding to the normal estimated value. This is the normal reward value corresponding to the reference feature value.

6. The method according to claim 4, characterized in that, The step of determining the alarm metric based on the forgetting factor, dynamic probability, and the number of alarm reward values ​​obtained statistically includes: Through formula Calculate the alarm metric value, where p is the dynamic probability value.

7. The method according to claim 5, characterized in that, The step of determining the normality metric based on the forgetting factor, the dynamic probability, and the statistically obtained number of normal reward values ​​includes: Through formula Calculate the normal metric, where p is the dynamic probability value.

8. The method according to any one of claims 1-4, characterized in that, Adjusting the window length according to the reference action corresponding to the target feature value includes: In response to the fact that the reference action corresponding to the target feature value is the same as the reference action corresponding to the previous target feature value, it is determined that the window length will be extended by 1; In response to the fact that the reference action corresponding to the target feature value is different from the reference action corresponding to the previous target feature value, the length of the window is scaled according to the preset scaling parameters.

9. The method according to claim 8, characterized in that, The scaling of the window length according to preset scaling parameters includes: Calculate the product of the window length and the scaling parameter, and round it to obtain the candidate length; In response to the candidate length being greater than the preset minimum length, the candidate length is determined to be the adjusted window length; In response to the candidate length being less than the shortest length, the shortest length is determined to be the adjusted window length.

10. The method according to any one of claims 1-4, characterized in that, The method further includes: The corresponding reward value is calculated based on the target feature value and the reference action.

11. The method according to claim 10, characterized in that, The step of calculating the corresponding reward value based on the target feature value and the reference action includes: In response to the reference action being an alarm action, the formula is used. Calculate the reward value corresponding to the target feature value, where For the target feature value, This is a preset feature value threshold.

12. The method according to claim 11, characterized in that, The step of calculating the corresponding reward value based on the target feature value and the reference action further includes: In response to the reference action being a normal action, via the formula Calculate the reward value corresponding to the target feature value.

13. A real-time monitoring and alarm device, characterized in that, The device includes: The target feature value determination module is used to acquire monitoring feature values ​​monitored during the production process multiple times according to a preset time period, and determine the currently acquired monitoring feature value as the target feature value each time the monitoring feature value is acquired; The reference feature value determination module is used to determine the initial window length of the target feature value and to use the monitoring feature values ​​obtained within the window length before the target feature value as reference feature values. An action determination module is used to determine a reference action corresponding to the target feature value based on a reward value corresponding to each reference feature value. The reward value is determined based on the reference action corresponding to the reference feature value. The reference action includes an alarm action and a normal action. The window adjustment module is used to adjust the window length according to the reference action corresponding to the target feature value, and update the corresponding reference feature value according to the adjusted window length; The monitoring and alarm module is used to determine the cumulative alarm reward and the cumulative normal reward based on the reward value corresponding to the updated reference feature value, and to determine the alarm status corresponding to the target feature value based on the cumulative alarm reward and the cumulative normal reward, wherein the alarm status includes alarm and normal. The reference feature value determination module is further used for: The window length adjusted based on the reference action corresponding to the previous target feature value is determined as the initial window length for the current target feature value; The action determination module is further used for: The alarm confidence upper bound and the normal confidence upper bound corresponding to the target feature value are determined based on the reward value corresponding to each of the reference feature values; In response to the alarm confidence upper bound being greater than the normal confidence upper bound, the reference action corresponding to the target feature value is determined to be an alarm action; In response to the normal confidence upper bound being greater than the alarm confidence upper bound, the reference action corresponding to the target feature value is determined to be a normal action.

14. An electronic device, characterized in that, include: processor; a memory for storing processor-executable instructions; The processor is configured to implement the method of any one of claims 1 to 12 when executing instructions stored in the memory.

15. A non-volatile computer-readable storage medium storing computer program instructions thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 12.

Citation Information

Patent Citations

  • Depth weighted double-Q learning-based large-scale monitoring method and monitoring robot

    CN107292392A

  • Driver working state monitoring system and driver working state monitoring method

    CN107845238A