Power Internet of Things Network Security Risk Prediction Method Based on Levenshtein Distance Algorithm
By building a causal database in the power Internet of Things network and using the Levenshtein distance algorithm, the problem of poor performance in medium and long-term security risk prediction in the existing technology is solved, and medium- and long-term risk prediction and accurate prediction of future alarm events are achieved.
Patent Information
- Application Number
- CN202111035820.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-03
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-09-03
AI Technical Summary
The existing technology has poor long-term security risk prediction in power Internet of Things networks, and the method based on Markov model cannot be applied to medium- and long-term prediction in the system.
The Levenshtein distance algorithm is used to construct a causal database, store alarm events through linked list hashing, filter low-frequency data, and use the Levenshtein distance algorithm to calculate the similarity of alarm events to predict future risks.
It realizes medium- and long-term security risk prediction in the power Internet of Things network, enriches the causal database, and facilitates the system to predict the risk level of future alarm events in the medium and long term.
Smart Images

Figure CN113886811B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security prediction, and specifically refers to a method for predicting the network security risk of the power Internet of Things based on the Levenshtein distance algorithm. Background Art
[0002] When the scale of the power Internet of Things network gradually increases, the number of accompanying network attack events also gradually rises. It is very necessary to study network security. Traditional protection methods represented by intrusion detection technology and firewalls are no longer able to meet the requirements of large-scale networks for security protection. Network security protection is based on security situation analysis and security risk prediction. The risk prediction link is in the final stage of the network security situation awareness system. Only by predicting possible warning events can we prevent problems before they occur and better maintain the network security situation.
[0003] Accurately predicting the security risk probability in the network is of great significance for improving network security. In recent years, researchers have conducted many studies in the field of network security risk prediction. The commonly used technology is to adopt a network security risk prediction method based on the hidden Markov model. Although the strategy based on the Markov model has good prediction effects, it is not suitable for long-term prediction in the system. Whether it is a fault or maintenance, it is assumed that the probability of state change is fixed. Summary of the Invention
[0004] Based on the above problems, the present invention provides a method for predicting the network security risk of the power Internet of Things based on the Levenshtein distance algorithm, which solves the problem of poor long-term prediction effect of the existing technology for the system.
[0005] To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0006] A method for predicting the network security risk of the power Internet of Things based on the Levenshtein distance algorithm includes the following steps:
[0007] Step 1: Take the attack source IP, attack behavior, and attack target IP in a single warning event as a piece of valid warning information. For each piece of warning information, taking the current warning event as the result, find the six closest warning events before the occurrence time of this warning event as the cause, and thus construct a piece of causal data. Store the causal data in the database to form a causal database;
[0008] Step 2: Filter the causal database;
[0009] Step 3: Use the Levenshtein distance algorithm to predict the warning event.
[0010] Further, in the above step 1, it specifically includes the following:
[0011] First, sort all alarm events by start time. Then, taking the current alarm event as the result and using the start time of this alarm event as the reference point to move forward, select the six alarm events that occurred closest to and before the reference point start time as the causes, thus constructing a causal data, and store the causal data in the database to form a causal database.
[0012] Further, in step 1, it also includes:
[0013] The causal database stores causal data in the way of chained hashing. Whenever new causal data is added to the database, count the number of occurrences of the causal data, and use x and y to count the number. Among them, whenever a completely new causal data is added, add it to the array in order. At the same time, compare all the array data with the same result before it, calculate the Levenshtein similarity before the cause, and set the tolerance TOL. If there is an item with a tolerance greater than TOL, connect this item to the linked list of the new data, and also connect the new data to the linked list of this item, then update the data. The y value of the new data is equal to the x value before update plus the x values of all data on its own linked list. At the same time, update the y values of these linked list data in the array, and its y value is equal to the x value before update plus the x values of all data on its own linked list. Whenever an existing causal data in the array is added, directly find this data, let x = x + 1, and at the same time update the y value to be equal to the x value before update plus the x values of all data on its own linked list. Finally, update all data on the linked list of this data in the array.
[0014] Further, in step 2, the method for filtering the causal database is specifically as follows:
[0015] It is necessary to filter out low-frequency causal data. Causal data that appears only once is directly deleted, and a threshold is set to delete causal data below the threshold. The number y is used when filtering the causal database.
[0016] Further, the formula for the threshold is:
[0017] .
[0018] Further, step 3 specifically includes the following steps:
[0019] Step 31: Sort the current existing alarm events in the order of start time;
[0020] Step 31: After an alarm event occurs, select the three alarm events with the closest start time including this alarm event, and at the same time select the first six alarm events that include themselves for these three alarm events as their cause sequences respectively;
[0021] Step 33: Use the Levenshtein distance algorithm to calculate the similarity between all constructed cause sequences and all cause sequences in the filtered cause-and-effect library;
[0022] Step 34: Use the effect in the matched cause-and-effect data as the prediction result, denoted as the predicted alarm event, and calculate the risk of occurrence of the predicted alarm event obtained using a certain cause sequence;
[0023] Step 35: Calculate the risk level that the predicted alarm event may occur after the current alarm event occurs according to the result of Step 34.
[0024] Furthermore, in Step 34, the formula for calculating the risk of occurrence of the predicted alarm event obtained using a certain cause sequence is:
[0025] ;
[0026] wherein, is the risk of occurrence of alarm event t obtained after calculation using the num-th cause sequence, Similarity is the similarity degree calculated between the current "cause sequence" and the "cause sequence" in the cause-and-effect library, m is the number of times the effect corresponding to the "cause sequence" in the cause-and-effect library appears in the initial database, x is the number of times this piece of cause-and-effect data appears, is the sum of calculations for all prediction results of alarm event t.
[0027] Furthermore, in Step 35, the formula for calculating the risk level that the predicted alarm event may occur after the current alarm event occurs is:
[0028] ;
[0029] wherein, is the risk level that alarm event t may occur after the current alarm event occurs.
[0030] Compared with the prior art, the beneficial effects of the present invention are as follows: associating the occurrence events of alarm events, directly constructing a cause-and-effect database, and at the same time combining the current state, predicting the risk levels of various alarm events occurring through the cause-and-effect database, predicting future situations with reference to past data conditions, and as new alarm events are continuously generated, the cause-and-effect library can also be enriched simultaneously, which is convenient for medium- and long-term prediction in the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 is the flowchart of this embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0032] The present invention will be further described below with reference to the drawings. The implementation manners of the present invention include but are not limited to the following embodiments.
[0033] As Figure 1 shown in the power Internet of Things network security risk prediction method based on the Levenshtein distance algorithm, which includes the following steps:
[0034] Step 1: Construct a causal database.
[0035] Among them, a single alarm event in the power Internet of Things specifically includes the source IP of the attack, the attack behavior, the attacked IP, the start time, and the end time. The source IP of the attack, the attack behavior, and the attacked IP in a single alarm event are used as a valid alarm message.
[0036] In addition, for each alarm message, a causal database is constructed. First, all alarm events are sorted by the start time, and then, taking the current alarm event as the result and using the start time of this alarm event as the reference point, six alarm events that are closest to the start time of the reference point and occurred before it are used as the causes. Thus, a causal data is constructed and stored in the database to form a causal database.
[0037] In addition, the causal database stores causal data in the way of linked list hashing. Whenever new causal data is added to the database, the number of occurrences of the causal data is counted, and the counts are statistically analyzed using x and y. Among them, whenever a completely new causal data is added, it is sequentially added to the array. At the same time, compare all the array data with the same result before it, calculate the Levenshtein similarity before the cause, and set the tolerance TOL. If there is an item with a tolerance greater than TOL, connect this item to the linked list of the new data, and also connect the new data to the linked list of this item. Then update the data. The y value of the new data is equal to the x value before the update plus the x values of all the data on its own linked list. At the same time, update the y values of these linked list data in the array, and its y value is equal to the x value before the update plus the x values of all the data on its own linked list. Whenever a causal data that already exists in the array is added, directly find this data, set x = x + 1, and at the same time update the y value to be equal to the x value before the update plus the x values of all the data on its own linked list. Finally, update all the data on this data's linked list in the array. Based on this, an example of the causal database is shown in the following table (alarm events are represented by capital letters):
[0038]
[0039] Step 2: Filter the causal database.
[0040] Among them, it is necessary to filter out low-frequency causal data. Causal data that appears only once is directly deleted, and a threshold is set to delete causal data with a frequency lower than the threshold. The frequency y is used when filtering the causal database. In addition, the formula for the threshold is: (highest frequency + lowest frequency) / 2.
[0041] Step 3: Use the Levenshtein distance algorithm to predict alarm events.
[0042] Specifically, it includes the following steps:
[0043] Step 31: Sort the existing alarm events in chronological order of the start time;
[0044] Step 31: After an alarm event occurs, select the three alarm events with the most recent start time including this alarm event, and at the same time select the first six alarm events that these three alarm events contain themselves as their cause sequences respectively;
[0045] For example, when the alarm events are A, B, C, D, AD, F, G, HI, select HI, G, F to construct the cause sequences. The cause sequence of HI is C, D, AD, F, G, HI, the cause sequence of G is B, C, D, AD, F, G, and the cause sequence of F is A, B, C, D, AD, F; Let them be the first, second, and third cause sequences respectively;
[0046] Step 33: Use the Levenshtein distance algorithm to calculate the similarity between all the constructed cause sequences and all the cause sequences in the filtered causal library;
[0047] Step 34: Take the result in the matched causal data as the prediction result, denoted as the predicted alarm event, and calculate the risk level of the predicted alarm event occurring using a certain cause sequence. The formula is:
[0048] ;
[0049] Among them, is the risk level of the alarm event t occurring after calculation using the num-th cause sequence, Similarity is the similarity degree calculated between the current "cause sequence" and the "cause sequence" in the causal library, m is the number of times the result corresponding to the "cause sequence" in the causal library appears in the initial database, x is the number of times this piece of causal data appears, is the sum of calculations for all prediction results of the alarm event t;
[0050] Step 35: Calculate the risk level that the predicted alarm event may occur after the current alarm event according to the result of Step 34. The formula is:
[0051] ;
[0052] Among them, the credibility of the first cause sequence is the highest, and the credibility of the second and third cause sequences decreases in turn. It is the risk level that warning event t may occur after the current warning event occurs.
[0053] The above are the embodiments of the present invention. The above embodiments and the specific parameters in the embodiments are only for clearly expressing the inventor's invention verification process, and are not used to limit the patent protection scope of the present invention. The patent protection scope of the present invention still depends on its claims. Any equivalent structural changes made by using the content of the specification and drawings of the present invention should similarly be included in the protection scope of the present invention.
Claims
1. A method for predicting power Internet of Things network security risks based on the Levenshtein distance algorithm, characterized in that, It includes the following steps: Step 1: Take the attack source IP, attack behavior, and attack target IP in a single alarm event as a valid alarm message. For each alarm message, taking the current alarm event as the result, find the six closest alarm events before the occurrence time of this alarm event as the cause, thereby constructing a causal data, and storing the causal data in the database to form a causal database; Step 2: Filter the causal database; Step 3: Use the Levenshtein distance algorithm to predict alarm events; In the above Step 1, it also includes: The causal database stores causal data in a chained hash manner. Whenever new causal data is added to the database, count the number of occurrences of the causal data, and use x and y to count the number of times. Among them, whenever a completely new causal data is added, it is added to the array in order, and at the same time, compare all the array data with the same result before it, calculate the Levenshtein similarity before the cause, and set the tolerance TOL. If there is an item with a tolerance greater than TOL, connect this item to the linked list of the new data, and at the same time connect the new data to the linked list of this item, and then update the data. The y value of the new data is equal to the x value before the update plus the x values of all the data on its own linked list, and at the same time update the y values of these linked list data in the array, and its y value is equal to the x value before the update plus the x values of all the data on its own linked list. Whenever a causal data that already exists in the array is added, directly find this data, make x = x + 1, and at the same time update the y value to be equal to the x value before the update plus the x values of all the data on its own linked list, and finally update all the data on this data's linked list in the array; The specific content of the above Step 3 includes the following steps: Step 31: Sort the existing alarm events in chronological order; Step 32: After an alarm event occurs, select the three alarm events with the most recent start time including this alarm event, and at the same time select the first six alarm events that contain themselves among these three alarm events as their cause sequences respectively; Step 33: Use the Levenshtein distance algorithm to calculate the similarity between all the constructed cause sequences and all the cause sequences in the filtered causal library; Step 34: Take the result in the causal data obtained by the matching as the prediction result, expressed as a predicted alarm event, and calculate the risk level of the occurrence of the predicted alarm event obtained by using a certain cause sequence; Step 35: Calculate the risk level that the predicted alarm event may occur after the current alarm event according to the result of Step 34.
2. The method for predicting power Internet of Things network security risks based on the Levenshtein distance algorithm according to claim 1, wherein In the above Step 1, it specifically includes the following: First, sort all alarm events in chronological order, then take the current alarm event as the result, and use the start time of this alarm event as the reference point to move forward. Take the six alarm events that are closest to the start time of the reference point and occurred before it as the cause, thereby constructing a causal data, and storing the causal data in the database to form a causal database.
3. The power Internet of Things network security risk prediction method based on the Levenshtein distance algorithm according to claim 1, characterized in that, In the above Step 2, the method for filtering the causal database is specifically as follows: It is necessary to filter the causal data with low frequency. The causal data that appears only once is directly deleted, and a threshold is set to delete the causal data below the threshold. The number of occurrences y is used when filtering the causal database.
4. The power Internet of Things network security risk prediction method based on the Levenshtein distance algorithm according to claim 3, wherein The formula for the threshold is as follows: 。 5. The method for predicting power Internet of Things network security risks based on the Levenshtein distance algorithm according to claim 4, wherein In step 34, the formula for calculating the risk level of the predicted alarm event occurring using a certain cause sequence is: ; Among them, is the risk level of the occurrence of the warning event t calculated using the num-th cause sequence. Similarity is the similarity degree calculated between the current "cause sequence" and the "cause sequence" in the cause-and-effect library. m is the number of times the result corresponding to the "cause sequence" in the cause-and-effect library appears in the initial database, and x is the number of times this piece of cause-and-effect data appears. is the sum of calculations for all prediction results that are warning event t.
6. The method for predicting power Internet of Things network security risks based on the Levenshtein distance algorithm according to claim 5, characterized in that, In step 35, the formula for calculating the risk level that the predicted alarm event may occur after the current alarm event occurs is: ; Among them, is the risk level that the alarm event t may occur after the current alarm event occurs.
Citation Information
Patent Citations
Causal knowledge-based power information network attack scene reconstruction method and system
CN111541661A
Power Internet of Things network security risk prediction method based on Levenshtein distance algorithm
CN113886811A