A security management system and method for intelligent cash box
By introducing a cumulative expected reward mechanism, the intelligent box security management system calculates short-term and long-term reward values and generates a selection list, which solves the problem that traditional systems cannot adapt to diverse scenarios and improves management efficiency and security.
Patent Information
- Application Number
- CN202510088562.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-01-21
AI Technical Summary
The traditional smart box safety management system lacks dynamic response capabilities to different states and risk levels, and is unable to adapt to the needs of diverse scenarios, resulting in long-term safety hazards and inefficient management.
The cumulative expected reward mechanism is introduced. By calculating short-term reward values and long-term cumulative expected returns, we judge whether to select a suitable strategy in short-term rapid response or long-term optimization. The management system obtains the probability that any security state occurs in the history of any security state and performs a certain action, and calculates the reward value of each security state and action combination through the reward function to generate a short-term and long-term selection list.
It improves the management efficiency of smart boxes, solves the conflict between short-term and long-term goals, realizes dynamic security management, and improves the security and operation stability of the system.
Smart Images

Figure CN119940939B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cash box management, and in particular to a security management system and method for an intelligent cash box. Background Art
[0002] Smart cash boxes are high-security devices designed specifically for banks and financial institutions for storing cash, bills, and other valuables. With the development of banking services, the security requirements for cash boxes are gradually increasing. Traditional cash boxes, due to their relatively backward technology, cannot meet the security, management efficiency, and operational convenience requirements of the modern financial system. The introduction of smart cash boxes not only improves the security of cash boxes, but also provides an intelligent solution for bank operations and management.
[0003] The existing technology has the following defects:
[0004] Traditional smart cash box security management systems usually make decisions based on fixed rules (such as simply triggering alarms or recording logs). They lack the ability to dynamically respond to different states and risk levels and cannot adapt to diverse scenario requirements. Decisions are often only made based on the current state (short-term goals) and ignore the optimization of long-term security and risk control. For example, in order to reduce the false alarm rate, potential threats may be ignored, resulting in long-term security risks, thereby reducing the management efficiency of smart cash boxes.
[0005] Based on this, the present invention proposes a security management system and method for smart cash boxes. By introducing cumulative expected rewards, the short-term reward value and the long-term cumulative expected benefits are calculated at the same time, and it is determined whether to choose a suitable strategy in short-term rapid response or long-term optimization, thereby solving the problem of conflict between short-term and long-term goals and improving the management efficiency of smart cash boxes. Summary of the Invention
[0006] The purpose of the present invention is to provide a security management system and method for an intelligent cash box to address the deficiencies in the background technology.
[0007] In order to achieve the above object, the present invention provides the following technical solution: a security management method for a smart cash box, the management method comprising the following steps:
[0008] The management system obtains all security states of the smart cash box, creates a state set for all security states, and obtains the corresponding action set based on the state set. It also obtains the probability of transitioning to the next security state after any security state in the state set occurs and a certain action is executed. In each security state, all executed actions are initially sorted according to the probability.
[0009] The reward function calculates the reward value for each combination of safety state and action. When the smart cash box enters a certain safety state, the management system first obtains the reward value between the safety state and each action, and then combines the reward value with the transition probability to obtain the cumulative expected reward.
[0010] Sort all actions by reward value to obtain a short-term selection list. Sort all actions by cumulative expected rewards to obtain a long-term selection list. After analyzing the historical operating status data of the smart cash box, determine whether to select the action that is ranked first in the short-term selection list or the long-term selection list in positive order and execute it.
[0011] In a preferred embodiment, the reward value for each safe state and action combination is calculated using a reward function, including the following steps:
[0012] The reward function is used to calculate the reward value for each safe state and action combination. The expression is:
[0013] , where is the reward value, For the execution success rate, For abnormal recovery speed, is the action response index, 、 、 is the adjustment coefficient, and 、 、 Both are greater than 0.
[0014] In a preferred embodiment, combining the reward value and the transition probability to obtain the cumulative expected reward includes the following steps:
[0015] Get the reward value between the current safety state and each action, get the current safety state and select When taking an action, the probability of transferring to a normal safe state and the probability of transferring to an abnormal safe state are calculated, and the cumulative expected reward is expressed as: , where To take the first The cumulative expected reward of each action, To take the The probability of an action transferring to a normal safe state, To take the The probability of an action transferring to an abnormally safe state, Indicates that the current security status is The reward value for each action, 、 is the weight coefficient, and .
[0016] In a preferred embodiment, after analyzing the historical operating status data of the smart cash box, determining whether to select the action that is first in the positive order from the short-term selection list or the long-term selection list includes the following steps:
[0017] Obtain the fault frequency of smart cash boxes over a period of history. After obtaining the daily fault frequency, calculate the average fault frequency and the standard deviation of the fault frequency. Divide the average fault frequency by the standard deviation of the fault frequency to obtain the fault impact amplitude.
[0018] Compare the acquired fault impact amplitude with the preset impact threshold. If the fault impact amplitude is less than or equal to the impact threshold, it is judged that the smart cash box has a small number of overall faults in a historical period, and the current safety state adopts the action ranked first in the long-term selection list to execute. If the fault impact amplitude is greater than the impact threshold, it is judged that the smart cash box has a large number of overall faults in a historical period, and the current safety state adopts the action ranked first in the short-term selection list to execute.
[0019] In a preferred embodiment, the calculation expressions for the average fault frequency and the standard deviation of the fault frequency are respectively:
[0020] , where is the average failure frequency, is the standard deviation of the fault frequency, is the number of days during the monitoring period, For the Failure frequency per day.
[0021] In a preferred embodiment, when the smart cash box enters a certain safety state, the management system first obtains the reward value between the safety state and each action, including the following steps:
[0022] When the smart cash box enters a certain safety state, the management system first obtains the execution success rate, abnormal recovery speed and action response index between the safety state and each action, and substitutes the execution success rate, abnormal recovery speed and action response index into the reward function to calculate the reward value between the safety state and each action. , Indicates that the current security status is The reward value for each action.
[0023] In a preferred embodiment, the probability of transitioning to the next safe state after any safe state in the state set history occurs and a certain action is executed is obtained. In each safe state, all executed actions are initially sorted according to the probability, including the following steps:
[0024] Through historical operation logs, the security state transition situation after executing a certain action in each security state is collected, and the security state S is calculated based on the transition situation. i Execute action A i,j Then transfer to the next safe state S i+1 The probability of is expressed as: , where Indicates safe state S i Execute action A i,j Then transfer to the next safe state S i+1 The probability of Indicates the A safe state, Indicates the The first step of the safe state execution Actions
[0025] For each safe state S i , sum up all the probabilities of transferring to the normal safety state to obtain the normal safety probability, sum up all the probabilities of transferring to the abnormal safety state to obtain the abnormal safety probability, divide the normal safety probability by the abnormal safety probability to obtain the ranking value, and sort all the execution actions from large to small according to the ranking value to generate the initial ranking table.
[0026] In a preferred embodiment, all actions are sorted by reward value to obtain a short-term selection list, and all actions are sorted by cumulative expected reward to obtain a long-term selection list, including the following steps:
[0027] Sort all actions by reward value from large to small to generate a short-term selection list, and sort all actions by cumulative expected reward from large to small to generate a long-term selection list.
[0028] In a preferred embodiment, the management system obtains all security states of the smart cash box, creates a state set for all security states, and obtains a corresponding action set based on the state set, including the following steps:
[0029] Analyze the cash box's operating logic and potential security risks, and define security states. Each security state describes the cash box's operating conditions and security level.
[0030] Extract the security status and its conversion relationship through historical operation logs, classify the collected data into different states, and form a discrete state set;
[0031] All the acquired security states are established as a state set S={S1,S2,...,S n}, n represents the number of safe states;
[0032] Define a corresponding action set for each security state, and the action set A under each security state i ={A i,1 ,A i,2 ,...,A i,m}, m represents the number of actions in the action set.
[0033] A security management system for an intelligent cash box includes a set partitioning module, a primary sorting module, a calculation module, a secondary sorting module, and an execution module;
[0034] Set division module: obtains all security states of the smart cash box, creates a state set for all security states, and obtains the corresponding action set based on the state set;
[0035] One-shot sorting module: Obtain the probability of transitioning to the next safe state after any safe state in the state set history occurs and a certain action is executed. In each safe state, all executed actions are initially sorted according to the probability.
[0036] Calculation module: This module uses a reward function to calculate the reward value for each combination of safety state and action. When the smart cash box enters a certain safety state, it first obtains the reward value between the safety state and each action, and then combines the reward value with the transition probability to obtain the cumulative expected reward.
[0037] Secondary sorting module: sorts all actions by reward value to obtain a short-term selection list, and sorts all actions by cumulative expected reward to obtain a long-term selection list;
[0038] Execution module: After analyzing the historical operating status data of the smart cash box, it determines whether to execute the action that is ranked first in the short-term selection list or the long-term selection list in positive order.
[0039] In the above technical solution, the technical effects and advantages provided by the present invention are:
[0040] The present invention obtains the probability of transitioning to the next safe state after any safe state in the state set occurs and a certain action is performed, and calculates the reward value for each safe state and action combination through a reward function. When a safe state appears in the smart cash box, the management system first obtains the reward value between the safe state and each action, then combines the reward value with the transition probability to obtain the cumulative expected reward, sorts all actions according to the reward value, obtains a short-term selection list, sorts all actions according to the cumulative expected reward, obtains a long-term selection list, analyzes the historical operating status data of the smart cash box, and determines whether to select the action ranked first in the short-term selection list or the long-term selection list. By introducing the cumulative expected reward, the management system simultaneously calculates the short-term reward value and the long-term cumulative expected benefit, determines whether to choose the appropriate strategy between short-term rapid response and long-term optimization, resolves the problem of conflict between short-term and long-term goals, and improves the management efficiency of the smart cash box. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments described in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0042] Figure 1 Flow chart of the method of the present invention. DETAILED DESCRIPTION
[0043] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0044] Example 1: Please refer to Figure 1 As shown, the security management method of a smart cash box described in this embodiment includes the following steps:
[0045] The management system obtains all safety states of the smart cash box, establishes a state set for all safety states, and obtains the corresponding action set based on the state set, obtains the probability of transitioning to the next safety state after any safety state in the history of the state set occurs and a certain action is executed, and in each safety state, all executed actions are initially sorted according to the probability, and the reward value of each safety state and action combination is calculated through the reward function. When the smart cash box appears in a certain safety state, the management system first obtains the reward value between the safety state and each action, and then combines the reward value with the transition probability to obtain the cumulative expected reward, sorts all actions according to the reward value, obtains a short-term selection list, sorts all actions according to the cumulative expected reward, obtains a long-term selection list, and after analyzing the historical operating status data of the smart cash box, determines whether to select the action that is ranked first in the short-term selection list or the long-term selection list in positive order to execute.
[0046] This application obtains the probability of transitioning to the next safe state after any safe state in the state set history occurs and a certain action is performed, and calculates the reward value of each safe state and action combination through the reward function. When the smart cash box appears in a certain safe state, the management system first obtains the reward value between the safe state and each action, and then combines the reward value with the transition probability to obtain the cumulative expected reward, sorts all actions according to the reward value, obtains a short-term selection list, sorts all actions according to the cumulative expected reward, obtains a long-term selection list, analyzes the historical operating status data of the smart cash box, and determines whether to select the action that is ranked first in the short-term selection list or the long-term selection list. The management system introduces the cumulative expected reward, calculates the short-term reward value and the long-term cumulative expected benefit at the same time, and determines whether to choose the appropriate strategy between short-term rapid response and long-term optimization, solves the problem of conflict between short-term and long-term goals, and improves the management efficiency of the smart cash box.
[0047] Example 2: The management system obtains all security states of the smart cash box, creates a state set for all security states, and obtains a corresponding action set based on the state set, including the following steps:
[0048] Analyze the cash box's operating logic and potential security risks, and define the possible security states that may occur in the system. Each state can describe the cash box's operating conditions and security level, and must include the following attributes:
[0049] Normal state: The system is operating normally without abnormal behavior.
[0050] Abnormal status: involving password errors, illegal operations, etc.
[0051] Alarm status: A corresponding alarm is triggered after a threat is detected.
[0052] Environmental conditions: external physical or electronic disturbances (such as vibration or power outage).
[0053] Fault status: System hardware or software malfunctions.
[0054] Real-time data collection: Sensors and monitoring devices are used to capture cash box operating parameters, such as the number of times a cash box is opened and closed, password entry history, the status of built-in hardware (such as locks, doors, and network modules), and environmental variables (vibration, temperature, humidity, and power supply). The system then identifies the threshold for each collected parameter. For example, three consecutive incorrect password entries equal an abnormal state. A strong vibration signal equals an alarm state.
[0055] Extract security status and its transition relationships from historical operation logs. For example, a transition to a locked state occurs after multiple consecutive unauthorized access attempts.
[0056] The collected data is classified into different states to form a set of discrete states that can be identified by the system.
[0057] All the acquired security states are established as a state set S={S1,S2,...,S n}, n represents the number of safe states and is given a clear meaning, for example:
[0058] S1 = normal operating state (no threat);
[0059] S2 = abnormal behavior detection status (e.g., multiple incorrect passwords);
[0060] S3 = environmental interference state (such as vibration or power failure);
[0061] S4 = threat alarm state (e.g., illegal violent attempt);
[0062] S5 = Hardware or software failure status (such as communication abnormality).
[0063] Define a corresponding action set for each security state, clarify the actions that can be taken in different states, and define the action type:
[0064] Logging action: records status information for later analysis.
[0065] Alarm action: Activate sound or signal alarm.
[0066] Recovery actions: Perform operations such as password reset and communication restart.
[0067] Blocking action: Forced lock to prevent illegal operation.
[0068] Notification Action: Sends an alert or status notification to an administrator.
[0069] Action set A in each safety state i ={A i,1 ,A i,2 ,...,Ai,m}, m represents the number of actions in the action set. The design of actions must match the state attributes and goals (such as security and stability). For example:
[0070] Security state S2 = abnormal behavior detection state, then A2 = {A 2,1 :Trigger alarm,A 2,2 :Notify the administrator, A 2,3 : Lock cash box}.
[0071] Obtain the probability of transitioning to the next safe state after any safe state occurs and a certain action is executed in the state set history. In each safe state, all executed actions are initially sorted according to the probability, including the following steps:
[0072] Through historical operation logs, the security state transition situation after executing a certain action in each security state is collected, and the security state S is calculated based on the transition situation. i Execute action A i,j Then transfer to the next safe state S i+1 The probability of is expressed as: , where Indicates safe state S i Execute action A i,j Then transfer to the next safe state S i+1 The probability of Indicates the A safe state, Indicates the The first step of the safe state execution Action, for example, safe state S2 performs action A 2,1 The total number of times is 100, the number of times transferred to S1 is 70 times, and the number of times transferred to S3 is 30 times, then the safe state S2 executes action A 2,1 The probability of transitioning to S1 is 0.7, and the safe state S2 executes action A 2,1 The probability of transferring to S3 is 0.3;
[0073] All state transition probabilities are matrixed to facilitate subsequent decision optimization. The probability matrix expression is: .
[0074] For each safe state S i The normal safety probability is obtained by summing up all the probabilities of transitioning to the normal safety state, and the abnormal safety probability is obtained by summing up all the probabilities of transitioning to the abnormal safety state. The normal safety probability is divided by the abnormal safety probability to obtain the ranking value. All execution actions are sorted from large to small according to the ranking value to generate an initial ranking table. This has the following advantages:
[0075] The safety value of each action is determined by comparing its contribution to restoring the system to a normal state (normal safety probability) with the risk of causing an abnormal state (abnormal safety probability). Actions with higher ranking values are more likely to contribute to restoring the system to a normal state, and prioritizing these actions maximizes system stability. The calculation incorporates all possible state transition probabilities, avoiding a single objective (such as restoring a specific state) and ensuring the global applicability of action selection. For example, an action may be effective in a specific scenario but may pose a greater overall safety risk. Probabilistic weighting can avoid this limitation. Actions with higher ranking values are more likely to restore the system to a normal state in a shorter period of time, reducing safety risks.
[0076] Because the ranking value also considers the probability of abnormal safety, it avoids selecting actions that may be effective in the short term but increase the probability of abnormality in the long term, thereby achieving more lasting security. Based on historical state transition probability data, the ranking value can be dynamically adjusted after each transition probability update, allowing the system to automatically adapt to changes in the operating environment or threat type. For example, in the event of increased external interference, protective actions (such as locking the cash box) may be prioritized.
[0077] After generating the initial sorting table, simply select actions in order according to the sorting, facilitating the implementation and maintenance of the management system. By prioritizing actions, redundant operations can be reduced by avoiding actions that are irrelevant or have minimal impact on system recovery. For example, actions like triggering an alarm or notifying an administrator can be prioritized, while lower-priority actions like logging can be delayed.
[0078] This approach is independent of specific smart checkbox functionality and can be extended to other security management systems requiring state transition analysis, such as industrial equipment monitoring and smart home security. When the system is expanded or new states are added, only transition probabilities need to be recalculated and ranking values updated, eliminating the need for significant adjustments to the entire solution. Normal and abnormal safety probabilities directly reflect the positive and negative impacts of actions on security, and ranking values provide quantitative metrics for comparing actions. Further optimization of ranking rules can be achieved by incorporating other evaluation metrics (such as action execution cost and time considerations) to enhance the overall effectiveness of the solution.
[0079] The reward function is used to calculate the reward value for each safe state and action combination, which includes the following steps:
[0080] The reward function is used to calculate the reward value for each safe state and action combination. The expression is:
[0081] , where is the reward value, For the execution success rate, For abnormal recovery speed, is the action response index, 、 、 is the adjustment coefficient, and 、 、 Both are greater than 0.
[0082] The execution success rate calculation logic is as follows: obtain the number of successful executions of the security status and action combination, and divide the number of successful executions by the usage time of the monitored smart cash box (i.e., the time from the time the smart cash box was put into use to the current time point) to obtain the execution success rate;
[0083] The relationship between the execution success rate and the reward value is that they jointly reflect the reliability and effectiveness of a certain safety state and action combination:
[0084] The higher the execution success rate, the greater the reward value: this indicates that under a certain safety state, executing a specific action can more stably achieve the goal. The system will give priority to such actions and give higher reward values to encourage repeated use. The lower the execution success rate, the smaller the reward value: actions with low success rates may be unstable or have higher risks, and the reward value will be reduced accordingly, thereby reducing the probability of the combination being selected. Incentive and exploration balance: Even if the success rate is low, the system may still give a certain basic reward value to ensure that these combinations still have a chance to be selected under certain specific conditions to explore their potential value.
[0085] The execution success rate directly determines the size of the reward value, which reflects the system's preference for action combinations with high reliability. The reward mechanism dynamically adjusts the reward value to guide the system to select a better strategy to ensure the safe management of smart cash boxes.
[0086] The abnormality recovery speed is obtained online through the system log of the smart cash box, that is, the time it takes for the smart cash box to return to normal after the abnormality occurs;
[0087] The relationship between the abnormality recovery speed and the reward value reflects the system's preference for fast response and effective recovery capabilities: the faster the abnormality recovery speed, the higher the reward value: a shorter recovery time indicates that the action combination is more efficient in handling the abnormality, and the system will give higher reward values to prioritize these combinations for similar abnormal situations in the future.
[0088] The slower the abnormality recovery, the lower the reward value: A longer recovery time may indicate a less effective or inefficient action combination, and the system will accordingly reduce its reward value and lower its priority in similar scenarios. In some scenarios with extremely high timeliness requirements (such as abnormal conditions with high security risks), the impact of recovery speed on reward value may be magnified to ensure a quick system response.
[0089] The speed of abnormality recovery directly impacts the reward value. Together with the execution success rate, it forms a key indicator for optimizing action selection within the intelligent cash box security management system. Through this reward mechanism, the system can dynamically optimize and select more effective strategies, improving overall management efficiency and security performance.
[0090] The calculation logic of the action response index is as follows: obtain the action response duration and the number of action conflicts, normalize the action response duration and the number of action conflicts so that the value range of the action response duration and the number of action conflicts is mapped to [0,1], obtain the normalized value of the action response duration and the normalized value of the number of action conflicts, and sum the normalized values of the action response duration and the number of action conflicts to obtain the action response index.
[0091] The magnitude of the action response index has the following relationship with the reward value of the safety state and action combination, reflecting the combined effect of action efficiency and execution competition:
[0092] A smaller action response index indicates a shorter response time, fewer conflicts, and higher execution efficiency and independence. The system prioritizes these action combinations and assigns higher rewards to improve overall response efficiency. A larger action response index indicates a lower reward: A larger action response index typically indicates a longer response time or more conflicts, which can lead to decreased system efficiency or increased risk. The reward for such combinations is lower, reducing their likelihood of selection in similar future scenarios. Multi-objective balance: The calculation of the action response index takes response time and the number of conflicts into consideration. This allows the system to balance speed and stability when allocating rewards, avoiding bias caused by a single metric.
[0093] The relationship between the action response index and the reward value reflects the system's optimization goals for action efficiency and reliability. By dynamically adjusting the reward value, the system prioritizes action combinations with fast response and minimal conflict, helping to improve exception handling and overall operational stability.
[0094] When the smart cash box enters a certain safety state, the management system first obtains the reward value between the safety state and each action, including the following steps:
[0095] When the smart cash box enters a certain safety state, the management system first obtains the execution success rate, abnormal recovery speed and action response index between the safety state and each action, and substitutes the execution success rate, abnormal recovery speed and action response index into the reward function to calculate the reward value between the safety state and each action. , Indicates that the current security status is The reward value for each action.
[0096] Combining the reward value and the transition probability to obtain the cumulative expected reward includes the following steps:
[0097] Get the reward value between the current safety state and each action, get the current safety state and select When taking an action, the probability of transferring to a normal safe state and the probability of transferring to an abnormal safe state are calculated, and the cumulative expected reward is expressed as: , where To take the first The cumulative expected reward of each action, To take the The probability of an action transferring to a normal safe state, To take the The probability of an action transferring to an abnormally safe state, Indicates that the current security status is The reward value for each action, 、 is the weight coefficient, and .
[0098] Sort all actions by reward value to get a short-term selection list. Sort all actions by cumulative expected reward to get a long-term selection list, including the following steps:
[0099] Sort all actions by reward value from large to small to generate a short-term selection list. The higher the action is ranked in the short-term selection list, the better the short-term performance of the action is. Sort all actions by cumulative expected reward from large to small to generate a long-term selection list. The higher the action is ranked in the long-term selection list, the better the long-term performance of the action is.
[0100] After analyzing the historical operating status data of the smart cash box, it is determined that the action that is first in the positive order from the short-term selection list or the long-term selection list is to be executed, including the following steps:
[0101] Obtain the fault frequency of the smart cash box over a historical period (7 or 14 days). After obtaining the daily fault frequency, calculate the average fault frequency and the standard deviation of the fault frequency. Divide the average fault frequency by the standard deviation of the fault frequency to obtain the fault impact amplitude. A larger fault impact amplitude indicates a greater number of overall faults in the smart cash box over the historical period.
[0102] Compare the acquired fault impact amplitude with the preset impact threshold. If the fault impact amplitude is less than or equal to the impact threshold, it is determined that the smart cash box has experienced a small number of faults in the past period of time, and the current safety state adopts the action ranked first in the long-term selection list. If the fault impact amplitude is greater than the impact threshold, it is determined that the smart cash box has experienced a large number of faults in the past period of time, and the current safety state adopts the action ranked first in the short-term selection list.
[0103] When the fault impact amplitude is large: it indicates that the smart cash box has a large number of historical faults and they occur frequently, which means that the system operation status is highly unstable.
[0104] When the fault impact amplitude is small: it indicates that the smart cash box has fewer historical faults and operates relatively smoothly.
[0105] The essence of the fault impact amplitude: It is used to measure the fault frequency of the smart cash box under historical operating conditions and the impact of fluctuations on the overall system operation.
[0106] The impact threshold is a preset reference standard for the system, used to determine whether the fault impact magnitude exceeds the allowable range. If the fault impact magnitude ≤ the impact threshold, the system has experienced a low number of faults and is operating well overall. If the fault impact magnitude > the impact threshold, the system has experienced a high number of faults, indicating recent poor operating conditions and requiring immediate resolution.
[0107] (1) When the fault impact magnitude is small (≤ impact threshold): Use the long-term selection list; the number of faults is small, the system is stable, and the current anomaly is less urgent. Prioritizing actions from the long-term selection list helps the system make better decisions from the perspective of global optimization and long-term benefits. Long-term actions typically bring greater overall benefits when the system is in good health. The system is running smoothly and requires less maintenance, such as scheduled maintenance or non-urgent performance optimization operations.
[0108] (2) When the fault impact magnitude is large (> impact threshold): use the short-term selection list;
[0109] The system has a high number of failures and unstable operating status, indicating a high level of urgency. Prioritize actions from the short-term selection list to quickly respond to the current situation and address the anomaly promptly to prevent further escalation. Short-term actions can quickly improve system security and stability and reduce risk. Frequent system failures, such as hardware failures and illegal operation detection, require immediate attention.
[0110] Long-term selection list: Suitable for scenarios with low urgency and low failure frequency, focusing on global optimization and future benefits. It improves the overall performance of the system through strategies that maximize long-term benefits.
[0111] Short-term selection list: Suitable for high-urgency, high-failure-frequency scenarios, focusing on rapid problem resolution and immediate response. This ensures system security and stability by rapidly increasing short-term benefits.
[0112] By leveraging historical fault data, the system can dynamically switch between short-term and long-term decisions based on actual operating conditions, avoiding overreliance on a single strategy. When the system is stable, long-term actions are selected to optimize future returns; when the system is in poor condition, short-term actions are selected to quickly improve current performance. This approach enables the system to intelligently evaluate historical data and current conditions, enabling more efficient safety management. When the fault magnitude is high, the short-term list prioritizes issues to prevent escalation. When the fault magnitude is low, the long-term list optimizes the system to reduce resource waste.
[0113] Through this dynamic selection mechanism, the system can choose appropriate actions under different operating conditions, which can not only ensure rapid response to high-frequency faults, but also optimize the long-term performance of the system under stable conditions, thereby improving the safety and operating efficiency of smart cash boxes.
[0114] The calculation expressions for the average failure frequency and the standard deviation of the failure frequency are:
[0115] , where is the average failure frequency, is the standard deviation of the fault frequency, is the number of days during the monitoring period, For the Failure frequency per day.
[0116] Example 3: The security management system of a smart cash box described in this embodiment includes a set partitioning module, a primary sorting module, a calculation module, a secondary sorting module, and an execution module;
[0117] Set partitioning module: obtains all security states of the smart cash box, creates a state set for all security states, and obtains the corresponding action set based on the state set. The state set and action set are sent to the primary sorting module and the secondary sorting module;
[0118] The primary sorting module obtains the probability of transitioning to the next safe state after any safe state in the state set history occurs and a certain action is executed. In each safe state, all executed actions are initially sorted according to the probability, and the transition probability is sent to the calculation module.
[0119] Calculation module: This module calculates the reward value for each combination of safety state and action through a reward function. When the smart cash box enters a certain safety state, it first obtains the reward value between the safety state and each action. Then, it combines the reward value with the transition probability to obtain the cumulative expected reward. The reward value and the cumulative expected reward are sent to the secondary sorting module.
[0120] Secondary sorting module: sorts all actions by reward value to obtain a short-term selection list, sorts all actions by cumulative expected reward to obtain a long-term selection list, and sends the short-term selection list and long-term selection list to the execution module;
[0121] Execution module: After analyzing the historical operating status data of the smart cash box, it determines whether to execute the action that is ranked first in the short-term selection list or the long-term selection list in positive order.
[0122] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.
[0123] Throughout this specification, references to terms such as "one embodiment," "example," or "specific example" indicate that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0124] The preferred embodiments of the present invention disclosed above are intended only to help illustrate the present invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the present invention to specific embodiments. Obviously, many modifications and variations are possible based on the contents of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the present invention, thereby enabling those skilled in the art to better understand and utilize the present invention. The present invention is limited only by the claims and their full scope and equivalents.
Claims
1. A method for security management of an intelligent cash box, characterized by: The management method comprises the following steps: The management system obtains all security states of the smart cash box, creates a state set for all security states, and obtains the corresponding action set based on the state set. It also obtains the probability of transitioning to the next security state after any security state in the state set occurs and a certain action is executed. In each security state, all executed actions are initially sorted according to the probability. The reward function calculates the reward value for each combination of safety state and action. When the smart cash box enters a certain safety state, the management system first obtains the reward value between the safety state and each action, and then combines the reward value with the transition probability to obtain the cumulative expected reward. Sort all actions by reward value to obtain a short-term selection list. Sort all actions by cumulative expected reward to obtain a long-term selection list. After analyzing the historical operating status data of the smart cash box, determine whether to select the action that is ranked first in the short-term selection list or the long-term selection list in positive order and execute it. The reward function is used to calculate the reward value for each safe state and action combination, which includes the following steps: The reward function is used to calculate the reward value for each safe state and action combination. The expression is: , where is the reward value, For the execution success rate, For abnormal recovery speed, is the action response index, 、 、 is the adjustment coefficient, and 、 、 All greater than 0; Combining the reward value and the transition probability to obtain the cumulative expected reward includes the following steps: Get the reward value between the current safety state and each action, get the current safety state and select When taking an action, the probability of transferring to a normal safe state and the probability of transferring to an abnormal safe state are calculated, and the cumulative expected reward is expressed as: , where To take the first The cumulative expected reward of each action, To take the The probability of an action transferring to a normal safe state, To take the The probability of an action transferring to an abnormally safe state, Indicates that the current security status is The reward value for each action, 、 is the weight coefficient, and ; After analyzing the historical operating status data of the smart cash box, it is determined that the action that is first in the positive order from the short-term selection list or the long-term selection list is to be executed, including the following steps: Obtain the fault frequency of smart cash boxes over a period of history. After obtaining the daily fault frequency, calculate the average fault frequency and the standard deviation of the fault frequency. Divide the average fault frequency by the standard deviation of the fault frequency to obtain the fault impact amplitude. Compare the acquired fault impact amplitude with the preset impact threshold. If the fault impact amplitude is less than or equal to the impact threshold, it is determined that the smart cash box has experienced a small number of faults in the past period of time, and the current safety state adopts the action ranked first in the long-term selection list. If the fault impact amplitude is greater than the impact threshold, it is determined that the smart cash box has experienced a large number of faults in the past period of time, and the current safety state adopts the action ranked first in the short-term selection list. Obtain the probability of transitioning to the next safe state after any safe state occurs and a certain action is executed in the state set history. In each safe state, all executed actions are initially sorted according to the probability, including the following steps: Through historical operation logs, the security state transition situation after executing a certain action in each security state is collected, and the security state S is calculated based on the transition situation. i Execute action A i,j Then transfer to the next safe state S i+1 The probability of is expressed as: , where Indicates safe state S i Execute action A i,j Then transfer to the next safe state S i+1 The probability of Indicates the A safe state, Indicates the The first step of the safe state execution Actions For each safe state S i , sum up all the probabilities of transferring to the normal safety state to obtain the normal safety probability, sum up all the probabilities of transferring to the abnormal safety state to obtain the abnormal safety probability, divide the normal safety probability by the abnormal safety probability to obtain the ranking value, and sort all the execution actions from large to small according to the ranking value to generate the initial ranking table.
2. The method for security management of a smart cash box according to claim 1, characterized in that: The calculation expressions for the average failure frequency and the standard deviation of the failure frequency are: , where is the average failure frequency, is the standard deviation of the fault frequency, is the number of days during the monitoring period, For the Failure frequency per day.
3. The method for security management of a smart cash box according to claim 2, characterized in that: When the smart cash box enters a certain safety state, the management system first obtains the reward value between the safety state and each action, including the following steps: When the smart cash box enters a certain safety state, the management system first obtains the execution success rate, abnormal recovery speed and action response index between the safety state and each action, and substitutes the execution success rate, abnormal recovery speed and action response index into the reward function to calculate the reward value between the safety state and each action. , Indicates that the current security status is The reward value for each action.
4. The method for security management of a smart cash box according to claim 3, characterized in that: Sort all actions by reward value to get a short-term selection list. Sort all actions by cumulative expected reward to get a long-term selection list, including the following steps: Sort all actions by reward value from large to small to generate a short-term selection list, and sort all actions by cumulative expected reward from large to small to generate a long-term selection list.
5. The method for security management of a smart cash box according to claim 4, characterized in that: The management system obtains all security states of the smart cash box, creates a state set for all security states, and obtains the corresponding action set based on the state set, including the following steps: Analyze the cash box's operating logic and potential security risks, and define security states. Each security state describes the cash box's operating conditions and security level. Extract the security status and its conversion relationship through historical operation logs, classify the collected data into different states, and form a discrete state set; All the acquired security states are established as a state set S={S1,S2,...,S n }, n represents the number of safe states; Define a corresponding action set for each security state, and the action set A under each security state i ={A i,1 ,A i,2 ,...,A i,m }, m represents the number of actions in the action set.
6. A security management system for a smart cash box, used to implement the management method according to any one of claims 1 to 5, characterized in that: It includes set partitioning module, primary sorting module, calculation module, secondary sorting module and execution module; Set division module: obtains all security states of the smart cash box, creates a state set for all security states, and obtains the corresponding action set based on the state set; One-shot sorting module: Obtain the probability of transitioning to the next safe state after any safe state in the state set history occurs and a certain action is executed. In each safe state, all executed actions are initially sorted according to the probability. Calculation module: This module uses a reward function to calculate the reward value for each combination of safety state and action. When the smart cash box enters a certain safety state, it first obtains the reward value between the safety state and each action, and then combines the reward value with the transition probability to obtain the cumulative expected reward. Secondary sorting module: sorts all actions by reward value to obtain a short-term selection list, and sorts all actions by cumulative expected reward to obtain a long-term selection list; Execution module: After analyzing the historical operating status data of the smart cash box, it determines whether to execute the action that is ranked first in the short-term selection list or the long-term selection list in positive order.
Citation Information
Patent Citations
Determining action selection policies of execution device
CN112292699A
Collaborative storage method and system for multi-access edge computing and electronic equipment
CN115134418A