Safety management system and method for intelligent money box
By introducing a cumulative expected reward mechanism into the security management system of smart box, short-term and long-term reward values are calculated, and the problems of conflict between short-term and long-term goals in traditional systems are solved, and management efficiency and security are improved.
Patent Information
- Application Number
- CN202510088562.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-21
AI Technical Summary
The traditional smart box safety management system lacks dynamic response capabilities and is unable to adapt to the needs of diverse scenarios. It conflicts with short-term and long-term goals, resulting in inefficient management.
By introducing cumulative expected rewards, calculate short-term reward values and long-term cumulative expected returns, judge whether to choose a suitable strategy in short-term rapid response or long-term optimization, and solve the problem of conflict between short-term and long-term goals.
The management efficiency of smart boxes is improved, and long-term security and risk control are optimized through dynamic selection of strategies to avoid problems such as high false positive rates and potential threat ignorance.
Smart Images

Figure CN119940939A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cash box management, and in particular to a security management system and method for an intelligent cash box. Background Art
[0002] Smart cash boxes are high-security devices designed specifically for banks and financial institutions for storing cash, bills and other valuables. With the development of banking business, the security requirements for cash boxes are gradually increasing. Traditional cash boxes are relatively backward in technology and cannot meet the requirements of the modern financial system for security, management efficiency and ease of operation. The introduction of smart cash boxes not only improves the security of cash boxes, but also provides intelligent solutions for the operation and management of banks.
[0003] The prior art has the following defects:
[0004] Traditional smart cash box security management systems usually make decisions based on fixed rules (such as simply triggering alarms or recording logs). They lack the ability to dynamically respond to different states and risk levels and cannot adapt to diverse scenario requirements. Decisions are often only made based on the current state (short-term goals) and ignore the optimization of long-term security and risk control. For example, in order to reduce the false alarm rate, potential threats may be ignored, resulting in long-term safety hazards, thereby reducing the management efficiency of smart cash boxes.
[0005] Based on this, the present invention proposes a security management system and method for smart cash boxes. By introducing cumulative expected rewards, short-term reward values and long-term cumulative expected returns are calculated simultaneously, and it is determined whether to choose a suitable strategy in short-term rapid response or long-term optimization, so as to solve the problem of conflict between short-term and long-term goals and improve the management efficiency of smart cash boxes. Summary of the invention
[0006] The purpose of the present invention is to provide a security management system and method for an intelligent cash box to solve the deficiencies in the background technology.
[0007] In order to achieve the above object, the present invention provides the following technical solution: a security management method for a smart cash box, the management method comprising the following steps:
[0008] The management system obtains all security states of the smart cash box, establishes a state set for all security states, and obtains the corresponding action set based on the state set. It obtains the probability of transitioning to the next security state after any security state in the state set history occurs and a certain action is executed. In each security state, all executed actions are initially sorted according to the probability.
[0009] The reward function is used to calculate the reward value for each combination of safety status and action. When the smart cash box is in a certain safety status, the management system first obtains the reward value between the safety status and each action, and then combines the reward value with the transition probability to obtain the cumulative expected reward.
[0010] Sort all actions by reward value to obtain a short-term selection list. Sort all actions by cumulative expected rewards to obtain a long-term selection list. After analyzing the historical operating status data of the smart cash box, determine whether to select the first action in the short-term selection list or the long-term selection list in positive order to execute.
[0011] In a preferred embodiment, the reward value of each safety state and action combination is calculated by a reward function, including the following steps:
[0012] The reward function is used to calculate the reward value for each safe state and action combination. The expression is:
[0013] , where is the reward value, For the execution success rate, is the abnormal recovery speed, is the action response index, , , is the adjustment coefficient, and , , Both are greater than 0.
[0014] In a preferred embodiment, combining the reward value and the transition probability to obtain the cumulative expected reward includes the following steps:
[0015] Get the reward value between the current safety state and each action, get the current safety state and select When taking an action, the probability of transferring to a normal safe state and the probability of transferring to an abnormal safe state are calculated, and the cumulative expected reward is expressed as: , where To take the first step in the current security situation The cumulative expected reward of each action is To take the The probability of an action transferring to a normal safe state, To take the The probability of an action transferring to an abnormally safe state, Indicates that the current security status is The reward value for each action, , is the weight coefficient, and .
[0016] In a preferred embodiment, after analyzing the historical operation status data of the smart cash box, determining to execute the action that is first in the positive order from the short-term selection list or the long-term selection list includes the following steps:
[0017] Obtain the fault frequency of the smart cash box in a certain period of history. After obtaining the daily fault frequency, calculate the average fault frequency and the standard deviation of the fault frequency. Divide the average fault frequency by the standard deviation of the fault frequency to obtain the fault impact amplitude.
[0018] The acquired fault impact amplitude is compared with the preset impact threshold. If the fault impact amplitude is less than or equal to the impact threshold, it is judged that the overall number of failures of the smart cash box in the historical period is small, and the current safety state adopts the action ranked first in the long-term selection list to execute. If the fault impact amplitude is greater than the impact threshold, it is judged that the overall number of failures of the smart cash box in the historical period is large, and the current safety state adopts the action ranked first in the short-term selection list to execute.
[0019] In a preferred embodiment, the calculation expressions for the average fault frequency and the fault frequency standard deviation are respectively:
[0020] , where is the average failure frequency, is the fault frequency standard deviation, is the number of days during the monitoring period, For the Failure frequency per day.
[0021] In a preferred embodiment, when the smart cash box is in a certain safety state, the management system first obtains the reward value between the safety state and each action, including the following steps:
[0022] When the smart cash box is in a certain safety state, the management system first obtains the execution success rate, abnormal recovery speed and action response index between the safety state and each action, and substitutes the execution success rate, abnormal recovery speed and action response index into the reward function to calculate the reward value between the safety state and each action. , Indicates that the current security status is The reward value for an action.
[0023] In a preferred embodiment, the probability of transitioning to the next safe state after any safe state in the state set history occurs and a certain action is executed is obtained, and in each safe state, all executed actions are initially sorted according to the probability, including the following steps:
[0024] Through historical operation logs, the security state transition after executing a certain action in each security state is collected, and the security state S is calculated based on the transition situation. i Execute action A i,j Then transfer to the next safe state S i+1 The probability of is expressed as: , where Indicates safe state S i Execute action A i,j Then transfer to the next safe state S i+1 The probability of Indicates A safe state, Indicates The first safe state execution Actions;
[0025] For each safety state S i , sum up all the probabilities of transferring to the normal safety state to obtain the normal safety probability, sum up all the probabilities of transferring to the abnormal safety state to obtain the abnormal safety probability, divide the normal safety probability by the abnormal safety probability to obtain the ranking value, sort all the execution actions from large to small according to the ranking value, and generate the initial sorting table.
[0026] In a preferred embodiment, all actions are sorted according to the reward value to obtain a short-term selection list, and all actions are sorted according to the cumulative expected reward to obtain a long-term selection list, including the following steps:
[0027] Sort all actions by reward value from large to small to generate a short-term selection list, and sort all actions by cumulative expected rewards from large to small to generate a long-term selection list.
[0028] In a preferred embodiment, the management system obtains all security states of the smart cash box, establishes a state set for all security states, and obtains a corresponding action set based on the state set, including the following steps:
[0029] Analyze the cash box's operating logic and potential security risks, and define security states. Each security state describes the cash box's operating conditions and security level.
[0030] Extract the safety status and its conversion relationship through historical operation logs, classify the collected data into different states, and form a discrete state set;
[0031] All the acquired security states are established as a state set S={S1,S2,...,S n}, n represents the number of safe states;
[0032] Define a corresponding action set for each security state. The action set A in each security state i ={A i,1 ,A i,2 ,...,A i,m}, m represents the number of actions in the action set.
[0033] A security management system for an intelligent cash box, comprising a set division module, a primary sorting module, a calculation module, a secondary sorting module, and an execution module;
[0034] Set division module: obtain all security states of the smart cash box, establish a state set for all security states, and obtain the corresponding action set based on the state set;
[0035] One-time sorting module: obtains the probability of transitioning to the next safe state after any safe state in the state set history occurs and a certain action is executed. In each safe state, all executed actions are initially sorted according to the probability;
[0036] Calculation module: Calculate the reward value of each safety state and action combination through the reward function. When the smart cash box appears in a certain safety state, first obtain the reward value between the safety state and each action, and then combine the reward value with the transition probability to obtain the cumulative expected reward;
[0037] Secondary sorting module: sort all actions according to reward value to obtain a short-term selection list, sort all actions according to cumulative expected rewards to obtain a long-term selection list;
[0038] Execution module: After analyzing the historical operating status data of the smart cash box, determine the first action to be executed in the positive order from the short-term selection list or the long-term selection list.
[0039] In the above technical solution, the technical effects and advantages provided by the present invention are:
[0040] The present invention obtains the probability of transferring to the next safe state after any safe state in the state set history occurs and a certain action is executed, and calculates the reward value of each combination of safe state and action through the reward function. When a certain safe state appears in the smart cash box, the management system first obtains the reward value between the safe state and each action, and then obtains the cumulative expected reward by combining the reward value and the transfer probability, sorts all actions according to the reward value, obtains a short-term selection list, sorts all actions according to the cumulative expected reward, obtains a long-term selection list, and after analyzing the historical operating state data of the smart cash box, determines whether to select the first action in the positive order from the short-term selection list or the long-term selection list. The management system introduces the cumulative expected reward, calculates the short-term reward value and the long-term cumulative expected benefit at the same time, determines whether to choose the appropriate strategy in the short-term rapid response or long-term optimization, solves the problem of conflict between short-term and long-term goals, and improves the management efficiency of the smart cash box. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0042] Figure 1 The figure is a flow chart of the method of the present invention. DETAILED DESCRIPTION
[0043] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0044] Example 1: Please refer to Figure 1 As shown, the security management method of a smart cash box described in this embodiment includes the following steps:
[0045] The management system obtains all safety states of the smart cash box, establishes a state set for all safety states, and obtains the corresponding action set based on the state set, obtains the probability of transferring to the next safety state after any safety state in the state set history occurs and a certain action is executed, and in each safety state, all executed actions are initially sorted according to the probability, and the reward value of each safety state and action combination is calculated through the reward function. When the smart cash box appears in a certain safety state, the management system first obtains the reward value between the safety state and each action, and then combines the reward value with the transition probability to obtain the cumulative expected reward, sorts all actions according to the reward value, obtains a short-term selection list, sorts all actions according to the cumulative expected reward, obtains a long-term selection list, and after analyzing the historical operating status data of the smart cash box, determines whether to select the action that is ranked first in the positive order from the short-term selection list or the long-term selection list for execution.
[0046] This application obtains the probability of transferring to the next safe state after any safe state in the history of the state set occurs and a certain action is executed, and calculates the reward value of each combination of safe state and action through the reward function. When a certain safe state appears in the smart cash box, the management system first obtains the reward value between the safe state and each action, and then combines the reward value with the transition probability to obtain the cumulative expected reward, sorts all actions according to the reward value, obtains a short-term selection list, sorts all actions according to the cumulative expected reward, obtains a long-term selection list, and after analyzing the historical operating status data of the smart cash box, determines whether to select the first action in the positive order from the short-term selection list or the long-term selection list. The management system introduces the cumulative expected reward, calculates the short-term reward value and the long-term cumulative expected benefit at the same time, determines whether to choose the appropriate strategy in the short-term rapid response or long-term optimization, solves the problem of conflict between short-term and long-term goals, and improves the management efficiency of the smart cash box.
[0047] Embodiment 2: The management system obtains all security states of the smart cash box, establishes a state set for all security states, and obtains a corresponding action set based on the state set, including the following steps:
[0048] Analyze the cash box's operating logic and potential security risks, and define the possible security states that may occur in the system. Each state can describe the cash box's operating conditions and security level, and must include the following attributes:
[0049] Normal state: The system is running normally without abnormal behavior.
[0050] Abnormal status: involving password errors, illegal operations, etc.
[0051] Alarm status: The corresponding alarm is triggered after the threat is detected.
[0052] Environmental conditions: External physical or electronic disturbances (such as vibration or power outage).
[0053] Fault status: System hardware or software malfunctions.
[0054] Real-time data collection: Use sensors and monitoring equipment to obtain cash box operating parameters, such as the number of opening and closing times, password input records, the status of built-in hardware (such as locks, doors, network modules), and environmental variables (vibration, temperature and humidity, power supply, etc.), and identify the state threshold of each collected parameter. For example, more than 3 consecutive password input errors = abnormal state. Detection of strong vibration signals = alarm state.
[0055] Extract security status and its transition relationship through historical operation logs. For example, after multiple consecutive unauthorized access attempts, the state is transferred to the locked state.
[0056] The collected data is classified into different states to form a set of discrete states that can be recognized by the system.
[0057] All the acquired security states are established as a state set S={S1,S2,...,S n}, n represents the number of safe states and is given a clear meaning, for example:
[0058] S1 = normal operating status (no threat);
[0059] S2 = abnormal behavior detection status (such as multiple incorrect passwords);
[0060] S3 = environmental disturbance state (such as vibration or power failure);
[0061] S4 = threat alarm status (e.g., illegal violent attempt);
[0062] S5 = Hardware or software failure status (such as communication abnormality).
[0063] Define a corresponding action set for each security state, clarify the operations that can be taken in different states, and define the action type:
[0064] Logging actions: Record status information for later analysis.
[0065] Alarm action: Activate sound or signal alarm.
[0066] Recovery actions: Perform operations such as password reset and communication restart.
[0067] Blocking action: Forced lock to prevent illegal operation.
[0068] Notification Action: Send an alert or status notification to an administrator.
[0069] The action set A in each safety state i ={A i,1 ,A i,2 ,...,Ai,m}, m represents the number of actions in the action set. The design of the action needs to match the state attributes and goals (such as security and stability), for example:
[0070] Security state S2 = abnormal behavior detection state, then A2 = {A 2,1 :Trigger alarm, A 2,2 :Notify the administrator, A 2,3 : Lock the cash box}.
[0071] Obtain the probability of transitioning to the next safe state after any safe state in the state set history occurs and a certain action is executed. In each safe state, all execution actions are initially sorted according to the probability, including the following steps:
[0072] Through historical operation logs, the security state transition after executing a certain action in each security state is collected, and the security state S is calculated based on the transition situation. i Execute action A i,j Then transfer to the next safe state S i+1 The probability of is expressed as: , where Indicates safe state S i Execute action A i,j Then transfer to the next safe state S i+1 The probability of Indicates A safe state, Indicates The first safe state execution For example, safe state S2 performs action A. 2,1 The total number of times is 100, the number of transfers to S1 is 70 times, and the number of transfers to S3 is 30 times, then the safe state S2 executes action A 2,1 The probability of transitioning to S1 is 0.7, and the safe state S2 executes action A 2,1 The probability of transferring to S3 is 0.3;
[0073] All state transition probabilities are matrixed to facilitate subsequent decision optimization. The probability matrix expression is: .
[0074] For each safety state S i , sum up all the probabilities of transferring to the normal safety state to obtain the normal safety probability, sum up all the probabilities of transferring to the abnormal safety state to obtain the abnormal safety probability, divide the normal safety probability by the abnormal safety probability to obtain the ranking value, sort all the execution actions from large to small according to the ranking value, and generate the initial sorting table, which has the following advantages:
[0075] By comparing the contribution of the action to the recovery to the normal state (normal safety probability) with the risk of causing the abnormal state (abnormal safety probability), the safety value of each action can be clarified. The larger the ranking value, the more conducive the action is to recovering the normal state. Prioritizing these actions can maximize the stability of the system. All possible state transition probabilities are included in the calculation to avoid simply considering a single goal (such as restoring a specific state) to ensure the applicability of action selection in a global scope. For example: an action may work better in a specific scenario, but it may cause greater safety risks overall. This limitation can be avoided by probability weighting. Actions with high ranking values are more likely to restore the system to a normal state in a short period of time, reducing safety risks.
[0076] Since the ranking value also considers the abnormal safety probability, it avoids selecting actions that may be effective in the short term but increase the abnormal probability in the long term, thereby achieving more lasting security. Based on the historical probability data of state transitions, the ranking value can be dynamically adjusted after each update of the transition probability, so that the system can automatically adapt when the operating environment or threat type changes. For example, in the case of increased external interference, protective actions (such as locking the cash box) may be performed with higher priority.
[0077] After the initial sorting table is generated, you only need to select actions in order according to the sorting, which is convenient for the implementation and maintenance of the management system. By sorting the actions by priority, you can avoid executing actions that are irrelevant to system recovery or have little contribution, thereby reducing redundant operations. For example, actions such as triggering alarms and notifying administrators are executed first, while low-priority operations such as logging are delayed.
[0078] This method does not rely on specific smart box functions and can be extended to other security management systems that require state transition analysis, such as industrial equipment monitoring, smart home security, etc. When the system is expanded or a new state is added, it is only necessary to recalculate the transition probability and update the ranking value, without making major adjustments to the entire solution. The normal safety probability and abnormal safety probability directly reflect the positive and negative impact of the action on security. The ranking value provides a quantitative indicator to facilitate the comparison of the pros and cons of the actions. Further combined with other evaluation dimensions (such as action execution cost, time cost, etc.) to optimize the sorting rules and improve the overall effect of the solution.
[0079] The reward function is used to calculate the reward value for each safe state and action combination, including the following steps:
[0080] The reward function is used to calculate the reward value for each safe state and action combination. The expression is:
[0081] , where is the reward value, For the execution success rate, is the abnormal recovery speed, is the action response index, , , is the adjustment coefficient, and , , Both are greater than 0.
[0082] The calculation logic of the execution success rate is as follows: obtain the number of successful executions of the security status and action combination, and divide the number of successful executions by the usage time of the monitored smart cash box (i.e., the time from the time the smart cash box is put into use to the current time point) to obtain the execution success rate;
[0083] The relationship between the execution success rate and the reward value is that they jointly reflect the reliability and effectiveness of a certain security state and action combination:
[0084] The higher the execution success rate, the greater the reward value: This means that in a certain safety state, executing a specific action can achieve the goal more stably. The system will give priority to such actions and give higher reward values to encourage repeated use. The lower the execution success rate, the smaller the reward value: actions with low success rates may be unstable or have higher risks, and the reward value will be reduced accordingly, thereby reducing the probability of the combination being selected. Incentive and exploration balance: Even if the success rate is low, the system may still give a certain basic reward value to ensure that these combinations still have a chance to be selected under certain specific conditions to explore their potential value.
[0085] The execution success rate directly determines the size of the reward value, which reflects the system's preference for action combinations with high reliability. The reward mechanism dynamically adjusts the reward value to guide the system to choose a better strategy to ensure the safe management of smart cash boxes.
[0086] The abnormal recovery speed is obtained online through the system log of the smart cash box, that is, the time it takes for the smart cash box to recover from the abnormality.
[0087] The relationship between the abnormality recovery speed and the reward value reflects the system's preference for fast response and effective recovery capabilities: the faster the abnormality recovery speed, the higher the reward value: a shorter recovery time indicates that the action combination is more efficient in handling abnormalities, and the system will give higher reward values to prioritize these combinations for similar abnormal situations in the future.
[0088] The slower the abnormal recovery speed, the lower the reward value: a longer recovery time may mean that the action combination is less effective or inefficient, and the system will reduce its reward value accordingly, reducing its priority in similar scenarios. In some scenarios with extremely high timeliness requirements (such as abnormal states with greater security risks), the impact of recovery speed on reward value may be amplified to ensure that the system responds quickly.
[0089] The abnormal recovery speed directly affects the reward value, which, together with the execution success rate, constitutes the key indicator for optimizing action selection in the smart cash box safety management system. Through the reward mechanism, the system can dynamically optimize and select more efficient strategies to improve overall management efficiency and safety performance.
[0090] The calculation logic of the action response index is as follows: obtain the action response time and the number of action conflicts, normalize the action response time and the number of action conflicts, map the value range of the action response time and the number of action conflicts to [0,1], obtain the normalized value of the action response time and the normalized value of the action conflicts, and sum the normalized value of the action response time and the normalized value of the action conflicts to obtain the action response index.
[0091] The size of the action response index has the following relationship with the reward value of the safety state and action combination, reflecting the combined impact of action efficiency and execution competition:
[0092] A smaller action response index indicates that the action has a short response time, a small number of conflicts, and higher execution efficiency and independence. The system will give priority to this type of action combination and give a higher reward value to improve the overall response efficiency. The larger the action response index, the lower the reward value: a larger action response index usually means a long response time or more action conflicts, which may lead to a decrease in system operating efficiency or an increase in risk. The reward value of this type of combination will be lower, reducing its probability of selection in similar scenarios in the future. Multi-objective balance: The calculation of the action response index takes into account the response time and the number of conflicts. The system can weigh speed and stability when allocating reward values to avoid deviations caused by a single indicator.
[0093] The relationship between the action response index and the reward value reflects the system's optimization goal of action efficiency and reliability. By dynamically adjusting the reward value, the system gives priority to action combinations with fast response and less conflict, which helps improve exception handling and overall operational stability.
[0094] When the smart cash box is in a certain safety state, the management system first obtains the reward value between the safety state and each action, including the following steps:
[0095] When the smart cash box is in a certain safety state, the management system first obtains the execution success rate, abnormal recovery speed and action response index between the safety state and each action, and substitutes the execution success rate, abnormal recovery speed and action response index into the reward function to calculate the reward value between the safety state and each action. , Indicates that the current security status is The reward value for an action.
[0096] Combining the reward value with the transition probability to obtain the cumulative expected reward includes the following steps:
[0097] Get the reward value between the current safety state and each action, get the current safety state and select When taking an action, the probability of transferring to a normal safe state and the probability of transferring to an abnormal safe state are calculated, and the cumulative expected reward is expressed as: , where To take the first step in the current security situation The cumulative expected reward of each action is To take the The probability of an action transferring to a normal safe state, To take the The probability of an action transferring to an abnormally safe state, Indicates that the current security status is The reward value for each action, , is the weight coefficient, and .
[0098] Sort all actions by reward value to get a short-term selection list. Sort all actions by cumulative expected reward to get a long-term selection list, including the following steps:
[0099] Sort all actions by reward value from large to small to generate a short-term selection list. The higher the action is ranked in the short-term selection list, the better the short-term performance of the action is. Sort all actions by cumulative expected reward from large to small to generate a long-term selection list. The higher the action is ranked in the long-term selection list, the better the long-term performance of the action is.
[0100] After analyzing the historical operation status data of the smart cash box, it is determined that the action that is first in the positive order from the short-term selection list or the long-term selection list is to be executed, including the following steps:
[0101] Obtain the fault frequency of the smart cash box within a historical period (7 days or 14 days). After obtaining the daily fault frequency, calculate the average fault frequency and the standard deviation of the fault frequency. Divide the average fault frequency by the standard deviation of the fault frequency to obtain the fault impact amplitude. The larger the fault impact amplitude, the more faults the smart cash box has experienced in the historical period.
[0102] Compare the acquired fault impact amplitude with the preset impact threshold. If the fault impact amplitude is less than or equal to the impact threshold, it is judged that the number of overall faults of the smart cash box in the historical period is small, and the current safety state adopts the action ranked first in the long-term selection list to execute. If the fault impact amplitude is greater than the impact threshold, it is judged that the number of overall faults of the smart cash box in the historical period is large, and the current safety state adopts the action ranked first in the short-term selection list to execute;
[0103] When the fault impact amplitude is large: it indicates that the smart cash box has a large number of historical faults and they occur frequently, which means that the system operation status is highly unstable.
[0104] When the fault impact amplitude is small: it indicates that the smart cash box has fewer historical faults and runs relatively smoothly.
[0105] The nature of the fault impact amplitude: It is used to measure the fault frequency of the smart box under historical operating conditions and the impact of fluctuations on the overall system operation.
[0106] The impact threshold is a preset reference standard for the system, which is used to determine whether the fault impact amplitude exceeds the allowable range. Fault impact amplitude ≤ impact threshold: The system has few faults and the overall operating status is good. Fault impact amplitude > impact threshold: The system has many faults, indicating that the recent operating status is poor and the current problem needs to be solved first.
[0107] (1) When the fault impact magnitude is small (≤ impact threshold): Use the long-term selection list; the number of faults is small and the system is running stably, and the urgency of the current anomaly is low. Prioritizing actions from the long-term selection list helps the system make better decisions from the perspective of global optimization and long-term benefits. Long-term actions usually bring greater overall benefits when the system is in good health. The system runs smoothly and requires less maintenance, such as scheduled maintenance or non-urgent performance optimization operations.
[0108] (2) When the fault impact magnitude is large (> impact threshold): use the short-term selection list;
[0109] There are many failures and the operation status is unstable. The current exception has a high urgency. Prioritize the selection of actions from the short-term selection list to quickly respond to the current status and handle the exception in time to prevent the problem from further deteriorating. Short-term actions can quickly improve the security and stability of the system and reduce risks. The system frequently fails during operation, such as hardware failures, illegal operation detection, and other problems that need to be handled immediately.
[0110] Long-term selection list: suitable for low-urgency, low-failure-frequency scenarios, focusing on global optimization and future benefits. Improve the overall performance of the system through strategies that maximize long-term benefits.
[0111] Short-term selection list: Applicable to scenarios with high urgency and high failure frequency, focusing on rapid problem solving and immediate response. Ensure the security and stability of the system through rapid improvement of short-term benefits.
[0112] Through historical fault data, the system can dynamically switch short-term or long-term decisions based on the actual operating status to avoid over-reliance on a single strategy. When the system status is stable, long-term actions are selected to optimize future benefits; when the system status is poor, short-term actions are selected to quickly improve current operating performance. This approach allows the system to intelligently evaluate historical data and current status, thereby achieving more efficient safety management. When the fault amplitude is high, the problem is handled first through the short-term list to avoid further escalation of the problem. When the fault amplitude is low, the system is optimized through the long-term list to reduce resource waste.
[0113] Through this dynamic selection mechanism, the system can choose appropriate actions under different operating conditions, which can not only ensure rapid response to high-frequency faults, but also optimize the long-term performance of the system under stable conditions, thereby improving the safety and operating efficiency of smart cash boxes.
[0114] The calculation expressions of the average failure frequency and the standard deviation of the failure frequency are:
[0115] , where is the average failure frequency, is the fault frequency standard deviation, is the number of days during the monitoring period, For the Failure frequency per day.
[0116] Embodiment 3: The security management system of a smart cash box described in this embodiment includes a set division module, a primary sorting module, a calculation module, a secondary sorting module, and an execution module;
[0117] Set division module: obtains all security states of the smart cash box, establishes a state set for all security states, and obtains the corresponding action set based on the state set. The state set and action set are sent to the primary sorting module and the secondary sorting module;
[0118] One-time sorting module: obtains the probability of transitioning to the next safe state after any safe state in the state set history occurs and a certain action is executed. In each safe state, all executed actions are initially sorted according to the probability, and the transition probability is sent to the calculation module;
[0119] Calculation module: Calculate the reward value of each safety state and action combination through the reward function. When the smart cash box appears in a certain safety state, first obtain the reward value between the safety state and each action, and then combine the reward value with the transition probability to obtain the cumulative expected reward. The reward value and the cumulative expected reward are sent to the secondary sorting module;
[0120] Secondary sorting module: sort all actions according to the reward value, obtain the short-term selection list, sort all actions according to the cumulative expected reward, obtain the long-term selection list, and send the short-term selection list and the long-term selection list to the execution module;
[0121] Execution module: After analyzing the historical operating status data of the smart cash box, it determines the execution of the action that is selected first in the positive order from the short-term selection list or the long-term selection list.
[0122] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters in the formula are set by technicians in this field according to actual conditions.
[0123] In the description of this specification, the description with reference to the terms "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0124] The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to only specific implementation methods. Obviously, many modifications and changes can be made according to the content of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and use the present invention well. The present invention is limited only by the claims and their full scope and equivalents.
Claims
1. A method for the safe management of an intelligent cash box, characterized in that: The management method comprises the following steps: The management system obtains all security states of the smart cash box, establishes a state set for all security states, and obtains the corresponding action set based on the state set. It obtains the probability of transitioning to the next security state after any security state in the state set history occurs and a certain action is executed. In each security state, all executed actions are initially sorted according to the probability. The reward function is used to calculate the reward value for each combination of safety status and action. When the smart cash box is in a certain safety status, the management system first obtains the reward value between the safety status and each action, and then combines the reward value with the transition probability to obtain the cumulative expected reward. Sort all actions by reward value to obtain a short-term selection list. Sort all actions by cumulative expected rewards to obtain a long-term selection list. After analyzing the historical operating status data of the smart cash box, determine whether to select the first action in the short-term selection list or the long-term selection list in positive order to execute.
2. The method for security management of a smart cash box according to claim 1, characterized in that: The reward function is used to calculate the reward value for each safe state and action combination, including the following steps: The reward function is used to calculate the reward value for each safe state and action combination. The expression is: , where is the reward value, For the execution success rate, is the abnormal recovery speed, is the action response index, , , is the adjustment coefficient, and , , Both are greater than 0.
3. The method for security management of a smart cash box according to claim 2, characterized in that: Combining the reward value with the transition probability to obtain the cumulative expected reward includes the following steps: Get the reward value between the current safety state and each action, get the current safety state and select When taking an action, the probability of transferring to a normal safe state and the probability of transferring to an abnormal safe state are calculated, and the cumulative expected reward is expressed as: , where To take the first step in the current security situation The cumulative expected reward of each action is To take the The probability of an action transferring to a normal safe state, To take the The probability of an action transferring to an abnormally safe state, Indicates that the current security status is The reward value for each action, , is the weight coefficient, and .
4. The method for security management of a smart cash box according to claim 3 is characterized in that: After analyzing the historical operation status data of the smart cash box, it is determined that the action that is first in the positive order from the short-term selection list or the long-term selection list is to be executed, including the following steps: Obtain the fault frequency of the smart cash box in a certain period of history. After obtaining the daily fault frequency, calculate the average fault frequency and the standard deviation of the fault frequency. Divide the average fault frequency by the standard deviation of the fault frequency to obtain the fault impact amplitude. The acquired fault impact amplitude is compared with the preset impact threshold. If the fault impact amplitude is less than or equal to the impact threshold, it is judged that the overall number of failures of the smart cash box in the historical period is small, and the current safety state adopts the action ranked first in the long-term selection list to execute. If the fault impact amplitude is greater than the impact threshold, it is judged that the overall number of failures of the smart cash box in the historical period is large, and the current safety state adopts the action ranked first in the short-term selection list to execute.
5. The method for security management of a smart cash box according to claim 4, characterized in that: The calculation expressions of the average failure frequency and the standard deviation of the failure frequency are: , where is the average failure frequency, is the fault frequency standard deviation, is the number of days during the monitoring period, For the Failure frequency per day.
6. The method for security management of a smart cash box according to claim 5, characterized in that: When the smart cash box is in a certain safety state, the management system first obtains the reward value between the safety state and each action, including the following steps: When the smart cash box is in a certain safety state, the management system first obtains the execution success rate, abnormal recovery speed and action response index between the safety state and each action, and substitutes the execution success rate, abnormal recovery speed and action response index into the reward function to calculate the reward value between the safety state and each action. , Indicates that the current security status is The reward value for an action.
7. The method for security management of a smart cash box according to claim 6, characterized in that: Obtain the probability of transitioning to the next safe state after any safe state in the state set history occurs and a certain action is executed. In each safe state, all execution actions are initially sorted according to the probability, including the following steps: Through historical operation logs, the security state transition after executing a certain action in each security state is collected, and the security state S is calculated based on the transition situation. i Execute action A i,j Then transfer to the next safe state S i+1 The probability of is expressed as: , where Indicates safe state S i Execute action A i,j Then transfer to the next safe state S i+1 The probability of Indicates A safe state, Indicates The first safe state execution Actions; For each safety state S i , sum up all the probabilities of transferring to the normal safety state to obtain the normal safety probability, sum up all the probabilities of transferring to the abnormal safety state to obtain the abnormal safety probability, divide the normal safety probability by the abnormal safety probability to obtain the ranking value, sort all the execution actions from large to small according to the ranking value, and generate the initial sorting table.
8. The method for security management of a smart cash box according to claim 7, characterized in that: Sort all actions by reward value to get a short-term selection list. Sort all actions by cumulative expected reward to get a long-term selection list, including the following steps: Sort all actions by reward value from large to small to generate a short-term selection list, and sort all actions by cumulative expected rewards from large to small to generate a long-term selection list.
9. A method for security management of a smart cash box according to claim 8, characterized in that: The management system obtains all security states of the smart cash box, establishes a state set for all security states, and obtains a corresponding action set based on the state set, including the following steps: Analyze the cash box's operating logic and potential security risks, and define security states. Each security state describes the cash box's operating conditions and security level. Extract the safety status and its conversion relationship through historical operation logs, classify the collected data into different states, and form a discrete state set; All the acquired security states are established as a state set S={S1,S2,...,S n }, n represents the number of safe states; Define a corresponding action set for each security state. The action set A in each security state i ={A i,1 ,A i,2 ,...,A i,m }, m represents the number of actions in the action set.
10. A security management system for a smart cash box, used to implement the management method according to any one of claims 1 to 9, characterized in that: It includes a set partitioning module, a primary sorting module, a calculation module, a secondary sorting module, and an execution module; Set division module: obtain all security states of the smart cash box, establish a state set for all security states, and obtain the corresponding action set based on the state set; One-time sorting module: obtains the probability of transitioning to the next safe state after any safe state in the state set history occurs and a certain action is executed. In each safe state, all executed actions are initially sorted according to the probability; Calculation module: The reward function is used to calculate the reward value of each combination of safety state and action. When the smart cash box is in a certain safety state, the reward value between the safety state and each action is obtained first, and then the cumulative expected reward is obtained by combining the reward value with the transition probability. Secondary sorting module: sort all actions according to reward value to obtain a short-term selection list, sort all actions according to cumulative expected rewards to obtain a long-term selection list; Execution module: After analyzing the historical operating status data of the smart cash box, determine the first action to be executed in the positive order from the short-term selection list or the long-term selection list.
Citation Information
Patent Citations
Determining action selection policies of execution device
CN112292699A
Intrusion detection algorithm based on depth strategy gradient
CN113344071A
Collaborative storage method and system for multi-access edge computing and electronic equipment
CN115134418A
A method and system for diversifying database behavior monitoring systems
CN116804963A
Satellite network intelligent resource scheduling method based on reinforcement learning
CN117314049A