Abnormal operation detection method, device, storage medium and computer program product

By constructing the state transfer matrix and using the relative deviation analysis method, the problem of insufficient accuracy of abnormal operation detection is solved, more efficient abnormal operation recognition is achieved, and the security of computer system and user information is improved.

CN116841842BActive Publication Date: 2025-08-01SHENZHEN HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310628511.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-30
Publication Date
2025-08-01
Estimated Expiration
2043-05-30

AI Technical Summary

Technical Problem

In the prior art, the accuracy of abnormal operation detection is insufficient, making it difficult to effectively identify abnormal operations that pose a threat to the security of computer systems and user information.

Method used

By obtaining the historical operation data of the target object and the operation data to be detected, a historical state transfer matrix and the state transfer matrix to be detected are constructed, and the relative deviation analysis method is used to determine the deviation of the operation data to be detected relative to the historical operation data to be detected, and whether there are abnormal operations are judged.

Benefits of technology

It improves the accuracy of abnormal operation detection, can identify abnormal operation more accurately, reduce misjudgment, and enhances the security of computer system and user information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116841842B_ABST
    Figure CN116841842B_ABST
Patent Text Reader

Abstract

The present application discloses an abnormal operation detection method, device, storage medium and computer program product, belonging to the field of computer security technology. This solution constructs a state transition matrix of historical operation data and a state transition matrix of operation data to be detected, determines the relative deviation of the state transition matrix to be detected relative to the historical state transition matrix, and thus determines the abnormal detection result based on the relative deviation. Since the state transition matrix can characterize the characteristics of the corresponding data changing with the operations included over time, the relative deviation of the state transition characteristics of the abnormal operation sequence relative to the state transition characteristics of the normal operation sequence is relatively large, and the relative deviation can more represent the degree of the operation data to be detected relative to the average situation of the historical operation data compared to the absolute deviation. Therefore, the accuracy of the abnormal operation detection in this solution is higher.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer security technology, and in particular to an abnormal operation detection method, device, storage medium and computer program product. Background Art

[0002] Currently, when using a computer, any object may generate a large number of operations. For example, when a user uses a mobile phone, computer, or other device, a series of user operations are generated. When a device such as a smart home appliance is running, a series of terminal operations are generated. When a server is running, a series of background operations are generated. These operations are generally normal operations that do not threaten the security of the computer system or user information. However, there are some abnormal operations that may pose a threat to the security of computer systems and user information. To ensure the security of computer systems and user information, it is necessary to detect abnormal operations. How to improve the accuracy of abnormal operation detection is a hot research topic in the industry. Summary of the Invention

[0003] This application provides a method, device, storage medium, and computer program product for detecting abnormal operations, which can improve the accuracy of abnormal operation detection. The technical solution is as follows:

[0004] In a first aspect, a method for detecting abnormal operations is provided, the method comprising:

[0005] Obtain historical operation data and operation data to be detected of the target object, where the historical operation data includes a normal operation sequence; determine a historical state transfer matrix and a state transfer matrix to be detected, where the historical state transfer matrix is the state transfer matrix of the historical operation data, and the state transfer matrix to be detected is the state transfer matrix of the operation data to be detected, and the state transfer matrix characterizes the time-varying characteristics of the operations contained in the corresponding data; determine the relative deviation of the state transfer matrix to be detected relative to the historical state transfer matrix; and determine the abnormality detection result of the operation data to be detected based on the relative deviation.

[0006] Since the state transition matrix can characterize the characteristics of the corresponding data as the operations included change over time, the relative deviation of the state transition characteristics of the abnormal operation sequence is relatively large compared with the state transition characteristics of the normal operation sequence, and the relative deviation can better characterize the degree to which the operation data to be detected is relative to the average situation of the historical operation data than the absolute deviation. Therefore, the accuracy of abnormal operation detection in this scheme is higher.

[0007] Optionally, the historical operation data includes historical data for K time periods, the operation data to be detected includes the data to be detected for one time period, each time period includes M sub-time periods, both K and M are integers not less than 1, the historical state transition matrix includes K groups of historical matrices corresponding one-to-one to the historical data for these K time periods, each group of historical matrices includes M historical matrices corresponding one-to-one to the historical data for the corresponding M sub-time periods, and the state transition matrix to be detected includes M matrices to be detected corresponding one-to-one to the data to be detected for the corresponding M sub-time periods; determining the relative deviation of the state transition matrix to be detected with respect to the historical state transition matrix includes: determining the relative deviation of each matrix to be detected among these M matrices to be detected with respect to each group of historical matrices among these K groups of historical matrices, obtaining K×M relative deviations; based on these K×M relative deviations, determining the relative deviation of the state transition matrix to be detected with respect to the historical state transition matrix. In this way, by dividing the time periods into sub-time periods to determine the state transition matrix for each sub-time period, it is considered that the research object usually performs a series of related operations within a short time period, which can more accurately determine the deviation of the operation data to be detected with respect to the historical operation data, and is beneficial to further improving the accuracy of abnormal operation detection.

[0008] Optionally, based on these K×M relative deviations, determining the relative deviation of the state transition matrix to be detected with respect to the historical state transition matrix includes: determining the minimum value among these K×M relative deviations as the relative deviation of the state transition matrix to be detected with respect to the historical state transition matrix. It should be understood that this minimum value represents the operation sequence deviation of the time period corresponding to the operation data to be detected from the operation sequence of the time period among the K historical time periods that is the closest, and using the closest operation sequence deviation as the basis for determining abnormal operations can improve the accuracy of abnormal operation detection.

[0009] Optionally, determining the relative deviation of each matrix to be detected among these M matrices to be detected with respect to each group of historical matrices among these K groups of historical matrices includes: for a reference historical group among these K groups of historical matrices, determining the N-norm of each historical matrix in the reference historical group, obtaining M N-norms, where N is an integer greater than 1, and the reference historical group is any one of these K groups of historical matrices; for the first matrix to be detected among these M matrices to be detected, determining the N-distance between the first matrix to be detected and each historical matrix in the reference historical group, obtaining M N-distances, and the first matrix to be detected is any one of these M matrices to be detected; determining the ratio of each of these M N-distances to the corresponding N-norm among these M N-norms, obtaining M ratios; determining the minimum value among these M ratios as the relative deviation of the first matrix to be detected with respect to the reference historical group. That is, the relative deviation is determined according to the calculation method of matrix norm.

[0010] Optionally, based on the relative deviation, an anomaly detection result of the operation data to be detected is determined, including: if the relative deviation of the state transition matrix to be detected relative to the historical state transition matrix exceeds the anomaly detection threshold, it is determined that there is an anomaly in the operation data to be detected. That is, the anomaly detection result is determined by comparing with a threshold.

[0011] In a second aspect, an anomaly operation detection device is provided. The anomaly operation detection device has the function of implementing the behavior of the anomaly operation detection method in the first aspect above. The anomaly operation detection device includes one or more modules, and the one or more modules are used to implement the anomaly operation detection method provided in the first aspect above.

[0012] In a third aspect, a computing device is provided. The computing device includes a processor and a memory. The memory is used to store a program for executing the anomaly operation detection method provided in the first aspect above, and to store data involved in implementing the anomaly operation detection method provided in the first aspect above. The processor is configured to execute the program stored in the memory. The computing device may further include a communication bus, and the communication bus is used to establish a connection between the processor and the memory.

[0013] In a fourth aspect, a computing device cluster is provided, including at least one computing device, and each computing device includes a processor and a memory; the processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the anomaly operation detection method provided in the first aspect above.

[0014] In a fifth aspect, a computer-readable storage medium is provided. Computer program instructions are stored in the computer-readable storage medium. When the computer program instructions run on a processor, the processor is caused to execute the anomaly operation detection method provided in the first aspect above.

[0015] In a sixth aspect, a computer program product including instructions is provided. When the instructions run on a processor, the anomaly operation detection method provided in the first aspect above is implemented.

[0016] The technical effects obtained in the second aspect, third aspect, fourth aspect, fifth aspect and sixth aspect above are similar to the technical effects obtained by the corresponding technical means in the first aspect, and will not be elaborated here. Description of the Drawings

[0017] Figure 1 is a flowchart of an anomaly operation detection method provided by an embodiment of the present application;

[0018] Figure 2 is an operation sequence diagram provided by an embodiment of the present application;

[0019] Figure 3 is another operation sequence diagram provided by an embodiment of the present application;

[0020] Figure 4 is yet another operation sequence diagram provided by an embodiment of the present application;

[0021] Figure 5 is yet another operation sequence diagram provided by an embodiment of the present application;

[0022] Figure 6 is a schematic diagram of the relative deviation of the data to be detected with respect to historical data provided by an embodiment of the present application;

[0023] Figure 7 is yet another operation sequence diagram provided by an embodiment of the present application;

[0024] Figure 8 is yet another operation sequence diagram provided by an embodiment of the present application;

[0025] Figure 9 is yet another operation sequence diagram provided by an embodiment of the present application;

[0026] Figure 10 is yet another schematic diagram of the relative deviation of the data to be detected with respect to historical data provided by an embodiment of the present application;

[0027] Figure 11 is yet another operation sequence diagram provided by an embodiment of the present application;

[0028] Figure 12 is yet another operation sequence diagram provided by an embodiment of the present application;

[0029] Figure 13 is a flowchart of yet another abnormal operation detection method provided by an embodiment of the present application;

[0030] Figure 14 is a schematic structural diagram of an abnormal operation detection device provided by an embodiment of the present application;

[0031] Figure 15 is a schematic structural diagram of a computing device provided by an embodiment of the present application;

[0032] Figure 16 is a schematic structural diagram of a computing device cluster provided by an embodiment of the present application;

[0033] Figure 17 is a schematic structural diagram of yet another computing device cluster provided by an embodiment of the present application. Detailed Embodiments

[0034] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0035] For ease of understanding, the application scenarios and background knowledge involved in the embodiments of the present application will be briefly introduced first.

[0036] At present, when computers and computer systems (or computer networks) are widely used, it is of great significance to use automated means to ensure the safe and reliable operation of computer systems. However, due to various reasons, computer systems and networks may be threatened by abnormal users. For example, after an abnormal user accesses a computer system, the login accounts or authorization information of normal users may be leaked. These leaks may pose potential risks to computer systems and user information, bringing risks to the safe operation of computer networks. How to prevent events that may threaten computer networks is an important research topic in the field of computer security research. Compared with traditional security protection methods, using a security protection scheme based on user behavior analysis is one of the current hotspots in theoretical research. Among these studies, the research on the principles and methods of abnormal operation detection is an important one.

[0037] In order to provide necessary backtracking information when a computer system fails, various logs have become information widely existing in current computer networks. These logs contain a large amount of information. For example, the login time of users, the operations of users, the automated operations of intelligent terminals, the background operations of servers, etc. These data that record the operation history in chronological order objectively reflect important information such as the state and trend of a research object changing over time.

[0038] It should be noted that during the process of a target object operating a computer system, most operations actually have internal correlations. This internal correlation is usually directly related to the purpose of the target object using the computer system. By tracking the correlations between operation sequences, the purpose of the target object using the computer system can be inferred. Based on this, the embodiments of the present application propose an abnormal operation detection method, in which the internal behavioral characteristics of the operation sequence are represented based on the idea of state transition. This method can depict the operation sequence of the target object with a state transition matrix, and further analyze the behavior of the target object based on this to determine whether the target object has abnormal operations.

[0039] Next, some concepts and terms involved in the embodiments of the present application will be introduced.

[0040] 1. Time window

[0041] When studying the behavior of a research object (i.e., the target object) in a computer network, it is necessary to extract the operations of the research object over a period of time for analysis. This period of time is called a time window, which can be represented by the symbol T. The time window can be appropriately selected according to different scenarios and the characteristics of the problem. In the embodiments of the present application, the time window is set to a constant.

[0042] 2. Key Operations

[0043] When using a computer system, any research object may generate a large number of operations. In an actual production environment, these operations are often designed according to the need to trace the production process, and their quantity and types are different in different production environments. However, usually not all operations are meaningful for the analysis behavior of the research object. Therefore, in the embodiments of the present application, the operations that are meaningful for behavior analysis are called key operations.

[0044] 3. Markov Process

[0045] The Markov process is a class of stochastic processes proposed by the Russian mathematician Markov in 1907. This process discretizes the change of the continuous system state in time, transforming the continuous process into a series of discrete states with internal correlations. The change process of the system can be regarded as a kind of conversion between these discrete states, and this conversion is called a state transition.

[0046] These discrete states can be finite or infinite, and the set they form is called the state space. For the embodiments of the present application, since the main discussion is about the relationship between operation sequences, the state space includes a finite number of states.

[0047] In the study of the Markov process, the current state can depend on the state at the previous moment or on multiple states over a previous period of time. For the embodiments of the present application, to comprehensively consider the computational efficiency and performance, a basic Markov model is adopted, that is, the current state only depends on the state at the previous moment. Of course, in some other embodiments, a more complex Markov model can also be adopted according to the actual situation to implement abnormal operation detection, that is, the current state depends on multiple states over a previous period of time. It should be understood that in the embodiments of the present application, one state corresponds to one key operation, that is, each key operation has a state value.

[0048] Generally, let X(t) represent a stochastic process, where t ∈ T and E is the state space of the stochastic process. Suppose the continuous time t can be discretized as t1 < t2 < … < t n <…, and X(t i ) = x i, i = 1, 2, …, n, …. Where n is a positive integer. When t ∈ (t n , t n+1 , if F(x, t|x n , x n-1 , …, x2, x1; t n , t n-1 , …, t2, t1) = F(x, t|x n , t n ) or P{X(t) ≤ x|X(t n ) = x n , …, X(t1) = x1} = P{X(t) ≤ x|X(t n ) = x n}, then the stochastic process is said to have the Markov property, also known as the lack of aftereffect or memorylessness. A stochastic process that satisfies the Markov property is called a Markov process. For the case where both time and state are discrete, a Markov process is also called a Markov chain.

[0049] Let p ij denote the probability that the previous state is i and the next state is j. p ij is then called the state transition probability from state i to state j. For a state space E containing a finite number of elements, the transition probabilities between any two states can be listed in full and presented in the form of a matrix [p ij . This matrix is called the state transition matrix and is denoted as P.

[0050] If the previous state is denoted by x n-1 , then the predicted value of x n can be expressed using the state transition matrix P as x n = x n-1 ·P.

[0051] It can be seen from the representation form of the Markov chain that if the operation sequence of the research object changing with time is characterized by the state transition matrix, then the values of the various transition probabilities in the state transition matrix give important characteristics of the behavior of the research object. Different state transition matrices can be used to represent different operation behaviors.

[0052] 4. Matrix Norm

[0053] Matrix norm is a scalar-valued function defined on a matrix set and can often be used as a characteristic of a matrix. Using the norm, the distance between any two elements in the space can be constructed. According to different needs, different matrix norms can be constructed on the same matrix set.

[0054] In the embodiments of the present application, for the state transition matrix with real number elements, the Frobenius matrix norm defined by Equation (1) is used:

[0055]

[0056] where P is a matrix with m rows and n columns, and the element p ij is a real number. Using the Frobenius norm, the distance d(A, B) between two real coefficient matrices A and B with m rows and n columns can be defined as shown in Equation (2).

[0057]

[0058] Next, the abnormal operation detection method provided by the embodiments of the present application will be introduced.

[0059] Figure 1 is a flowchart of an abnormal operation detection method provided by the embodiments of the present application. This method can be executed by a computing device or a cluster of computing devices. Next, an example will be given where this method is executed by a single computing device. Please refer to Figure 1 and this method includes the following steps.

[0060] Step 101: Obtain the historical operation data and the operation data to be detected of the target object, where the historical operation data includes normal operation sequences.

[0061] The computing device obtains the historical operation data and the operation data to be detected of the target object from the stored log data. Among them, the target object can be an object of types such as a target user, a target terminal, a target server, etc.

[0062] Since the log data usually contains a relatively large variety of information types, including the data required for the abnormal operation detection method and other data, the computing device can screen the log data to obtain the historical operation data and the operation data to be detected of the target object. Among them, there are many ways of data screening, and the embodiments of the present application will not introduce them in detail.

[0063] In the embodiments of the present application, data screening includes retaining data related to key operations and removing data unrelated to key operations. Both the historical operation data and the operation data to be detected are data related to key operations. A key operation refers to an operation related to the security of the computer system and / or user information, that is, an operation with research significance. In some other embodiments, data screening also includes removing abnormal data, such as removing outliers, null values, etc.

[0064] Optionally, the computing device periodically executes the steps of the abnormal operation detection method. The computing device obtains the operation data of the target object in the current time period to obtain the operation data to be detected, and obtains the operation data of the target object in the historical time period before the current time period to obtain the historical operation data.

[0065] Among them, the duration of the historical time period can be an integer multiple of the duration of the current time period. In other words, the computing device periodically obtains the operation data of the target object in the latest time period as the operation data to be detected, and uses the operation data of K time periods obtained before the latest time period as the historical operation data, where K is an integer not less than 1. In this way, the historical operation data includes the historical data of K time periods, and the operation data to be detected includes the data to be detected of one time period. Of course, the duration of the historical time period may not be an integer multiple of the duration of the current time period. For example, the historical time period is 7.5 days and the current time period is 1 day. In the embodiments of the present application, the case where the duration of the historical time period is an integer multiple of the duration of the current time period is taken as an example for introduction.

[0066] Exemplarily, the target object is User 1, the duration of each time period is one day, and K is 10. The computing device obtains the log data generated by User 1 today, filters the obtained log data, and obtains the data related to the key operations generated by User 1 today, and uses the filtered data as the operation data to be detected. The computing device uses the operation data of the 10 days before today obtained before today as the historical operation data. In this way, the historical operation data includes the operation data of the 10 days before today, and the operation data to be detected includes the operation data of one day today.

[0067] Step 102: Determine the historical state transition matrix and the state transition matrix to be detected.

[0068] Among them, the historical state transition matrix is the state transition matrix of the historical operation data, and the state transition matrix to be detected is the state transition matrix of the operation data to be detected. The state transition matrix characterizes the characteristics of the operations included in the corresponding data changing over time. That is, after the computing device obtains the historical operation data of the target object, it determines the state transition matrix of the historical operation data to obtain the historical state transition matrix. After the computing device obtains the operation data to be detected of the target object, it determines the state transition matrix of the operation data to be detected to obtain the state transition matrix to be detected.

[0069] As described above, the state transition matrix to be detected can be the state transition matrix for a time period, and the historical state transition matrix can be the state transition matrix for K time periods. Based on this, the computing device can determine a state transition matrix for the operation data of each time period. That is, the computing device determines a state transition matrix for the operation data to be detected and K state transition matrices for the historical operation data. Alternatively, the computing device can divide each time period into M sub-time periods and determine a state transition matrix for the operation data of each sub-time period. In this way, each time period contains M sub-time periods. Here, M is an integer not less than 1. That is, the computing device determines M state transition matrices for the operation data to be detected and K×M state transition matrices for the historical operation data. It can be seen that the computing device can determine the state transition matrix according to a coarse-grained time window (i.e., the granularity of a time period), or can also determine the state transition matrix according to a fine-grained time window (i.e., the granularity of a sub-time period).

[0070] Next, the specific implementation process of determining the state transition matrix according to the fine-grained time window will be introduced first.

[0071] In the embodiment of the present application, the computing device determines the types of critical operations, and determines the state values of each critical operation according to the total number of types of critical operations. Different critical operations correspond to different state values.

[0072] Exemplarily, there are 8 types of critical operations in total, and the computing device respectively records the state values of these 8 types of keyword operations as 0-7.

[0073] The computing device divides the historical operation data into historical data of K×M sub-time periods according to the duration of the time period and the sub-time period, and divides the operation data to be detected into detected data of M sub-time periods. The computing device sorts the critical operations included in the historical data of each sub-time period in the historical operation data according to the time order of the critical operations, so as to obtain a total of K×M historical operation sequences. Similarly, the computing device sorts the critical operations included in the detected data of each sub-time period in the operation data to be detected according to the time order of the critical operations, so as to obtain a total of M detected operation sequences. Among them, the historical operation sequence and the detected operation sequence can be represented by a sequence composed of the corresponding state values.

[0074] The computing device counts the number of times each operation in each of the K×M historical operation sequences transfers to another operation, determines the state transition matrix of each historical operation sequence based on the statistical results, and obtains K×M historical matrices. These K×M historical matrices are divided into K groups, that is, the historical state transition matrix contains K groups of historical matrices corresponding one-to-one to the historical data of these K time periods, and each group of historical matrices contains M historical matrices corresponding one-to-one to the historical data of the corresponding M sub-time periods. Similarly, the computing device counts the number of times each operation in each of the M to-be-detected operation sequences transfers to another operation, determines the state transition matrix of the to-be-detected operation sequences based on the statistical results, and obtains M to-be-detected matrices. In this way, the to-be-detected state transition matrix contains M to-be-detected matrices corresponding one-to-one to the to-be-detected data of the corresponding M sub-time periods.

[0075] It can be seen that in the implementation process of determining the state transition matrix according to the coarse-grained time window, the statistical duration of the computing device is the length of a relatively short time window, that is, the statistical duration is the duration of a sub-time period, a sub-time period is a time window, and one state transition matrix is obtained for one time window.

[0076] Next, the specific implementation process of determining the state transition matrix according to the coarse-grained time window is introduced.

[0077] The main difference between the implementation process of the computing device determining the state transition matrix according to the coarse-grained time window and the implementation process of determining the state transition matrix according to the fine-grained time window lies in the different statistical durations. In the implementation process of determining the state transition matrix according to the fine-grained time window, the statistical duration is the duration of a time period, a time period is a time window, and one state transition matrix is obtained for one time window.

[0078] The computing device divides the historical operation data into historical data of K time periods according to the duration of the above-mentioned one time period. The computing device sorts the key operations included in the historical data of each time period according to the time order of the key operations included in the historical data of each time period, so as to obtain a total of K historical operation sequences. Similarly, the computing device sorts the key operations included in the to-be-detected operation data according to the time order of the key operations included in the to-be-detected operation data, so as to obtain a total of one to-be-detected operation sequence. Among them, the historical operation sequence and the to-be-detected operation sequence can be represented by a sequence composed of corresponding state values.

[0079] The computing device counts the number of times each operation in each of the K historical operation sequences is transferred to another operation, determines the state transition matrix of each historical operation sequence based on the statistical results, and obtains K historical matrices. These K historical matrices correspond one-to-one with the historical data of these K time periods. Similarly, the computing device counts the number of times each operation in the one to-be-detected operation sequence is transferred to another operation, determines the state transition matrix of the to-be-detected operation sequence based on the statistical results, and obtains a to-be-detected matrix. Thus, the to-be-detected state transition matrix contains a to-be-detected matrix.

[0080] In addition to determining the state transition matrix according to the time windows of the above two granularities, the computing device can also determine the state transition matrix according to other granularities or other methods. For example, in an embodiment where the duration of the historical time period is not an integer multiple of the current time period, the computing device can determine the multiple historical matrices included in the historical state transition matrix in a manner of combining time windows of multiple granularities, and / or determine the multiple to-be-detected matrices included in the to-be-detected operation sequence in a manner of combining time windows of multiple granularities.

[0081] To verify the effectiveness of this solution, Example 1 and Example 2 will be given below based on the actual data on the network of a certain company to illustrate the implementation process of determining the state transition matrix by way of examples.

[0082] Example 1:

[0083] Figure 2 It is a sequence diagram corresponding to the key operations of a certain user within a time window of length T. There are 9 key operations in total, and the state values of these 9 key operations are respectively denoted as 0, 1, …, 8. According to Figure 2 the shown operation sequence, the number of times transferred from state i to state j can be counted to obtain the state transition count matrix S shown in Table 1, where i, j = 0, 1, …, 8.

[0084] Table 1

[0085]

[0086]

[0087] The first column and the first row in Table 1 respectively represent the previous state value and the current state value. The data in Table 1 except for the first row and the first column represent the element values in the state transition count matrix S. The first row of the state transition count matrix S represents the number of times the state with the previous state value of 0 is transferred to other states within this time window. The other rows are similar. The first column of this matrix represents the number of previous states when the current state value is 0. The magnitude of this number gives a characteristic of the user's operations changing over time.

[0088] According to the state transition count matrix S, the computing device can calculate the state transition matrix P shown in Table 2 according to the following formula (3).

[0089]

[0090] In formula (3), n = 8, indicating that there are 8 key operations in total, that is, 8 states. The first column and the first row in Table 2 represent the previous state value and the current state value respectively, and the data except the first row and the first column in Table 2 represent the element values in the state transition matrix P.

[0091] Table 2

[0092]

[0093] Similar to Figure 2 the key operation sequence diagram of the given user, Figure 3 is the sequence diagram corresponding to the key operations of another user within a time window of length T. Similar to the determination method of the state transition count matrix S introduced above, Figure 3 the state transition count matrix corresponding to the shown user is as shown in Table 3.

[0094] Table 3

[0095]

[0096]

[0097] Similar to the determination method of the state transition matrix P introduced above, Figure 3 the state transition matrix corresponding to the shown user is as shown in Table 4.

[0098] Table 4

[0099]

[0100] As can be seen from Example 1, the state transition matrices corresponding to different operation sequences may be different, and these differences reflect different operation habits of users, that is, differences in operation behaviors. Compared with the behavioral characteristics based on time series, the method of using the state transition matrix to characterize the operation sequence of the research object not only pays attention to the change situation of the operation sequence of the research object, but also eliminates the influence caused by the possible inconsistency of the length of the original operation sequence. Among them, the reason for the inconsistency of the length of the original operation sequence may be that the number of key operations in different time periods within the same length of time window may be different.

[0101] Example 2:

[0102] As can be seen from the above, compared with the method of characterizing the behavior characteristics of the research object based on the original operation sequence (i.e., time series), the behavior characteristic analysis method based on the state transition matrix pays more attention to the conversion between the two operation types before and after. Figure 4 and Figure 5 respectively give two different key operation sequences. Using the method in Example 1, it can be obtained that the state transition count matrices corresponding to these two key operation sequences are the same. Therefore, the state transition matrices of these two key operation sequences are the same, and the state transition matrices of these two key operation sequences are shown in Table 5.

[0103] Table 5

[0104]

[0105]

[0106] As can be seen from Example 2, from the perspective of state transition, different operation sequences may also produce the same state transition matrix. Many computer network security cases show that the operations of external intrusion often have the same operation logic, that is, the key operations have the same order, while the specific execution time of the key operations may vary randomly. Therefore, since this solution focuses on the state transition, it can discover abnormal operations that are impossible or difficult to discover by traditional time series-based methods. Conversely, the normal operation sequences of the research object may also appear randomly, and this solution can also be used to discriminate normal operation sequences.

[0107] Step 103: Determine the relative deviation of the state transition matrix to be detected relative to the historical state transition matrix.

[0108] After calculating the device determines the historical state transition matrix and the state transition matrix to be detected, it determines the relative deviation of the state transition matrix to be detected relative to the historical state transition matrix. Since the relative deviation can better characterize the degree of the average situation of the operation data to be detected deviating from the historical operation data compared with the absolute deviation, it can more accurately determine whether the target object has performed abnormal operations according to the relative deviation in the following.

[0109] In an embodiment of determining a state transition matrix according to a fine-grained time window, the historical state transition matrix includes K groups of historical matrices corresponding one-to-one to the historical data of K time periods, each group of historical matrices includes M historical matrices corresponding one-to-one to the historical data of corresponding M sub-time periods, and the state transition matrix to be detected includes M matrices to be detected corresponding one-to-one to the data to be detected of corresponding M sub-time periods. Based on this, the computing device determines the relative deviation of each matrix to be detected in these M matrices to be detected with respect to each group of historical matrices in these K groups of historical matrices, obtaining K×M relative deviations. Then, the computing device determines the relative deviation of the state transition matrix to be detected with respect to the historical state transition matrix based on these K×M relative deviations.

[0110] Among them, an implementation manner for the computing device to determine the relative deviation of each matrix to be detected in these M matrices to be detected with respect to each group of historical matrices in these K groups of historical matrices is as follows: for a reference historical group in these K groups of historical matrices, determine the N-order norm of each historical matrix in the reference historical group, obtaining M N-order norms, where N is an integer greater than 1, and the reference historical group is any one of these K groups of historical matrices; for the first matrix to be detected in these M matrices to be detected, determine the N-order distance between the first matrix to be detected and each historical matrix in the reference historical group, obtaining M N-order distances, and the first matrix to be detected is any one of these M matrices to be detected; determine the ratio of each of these M N-order distances to the corresponding N-order norm in these M N-order norms, obtaining M ratios; and determine the minimum value among these M ratios as the relative deviation of the first matrix to be detected with respect to the reference historical group.

[0111] An implementation manner for the computing device to determine the relative deviation of the state transition matrix to be detected with respect to the historical state transition matrix based on these K×M relative deviations is as follows: determine the minimum value among these K×M relative deviations as the relative deviation of the state transition matrix to be detected with respect to the historical state transition matrix. In this way, any research object that may have abnormal operations can be not missed.

[0112] In some other embodiments, the computing device determines the second smallest value among the above K×M relative deviations as the relative deviation of the state transition matrix to be detected with respect to the historical state transition matrix. In this way, the determination basis for detecting abnormal operations can be slightly relaxed.

[0113] As can be seen from the above, in an embodiment of determining a state transition matrix according to a fine-grained time window, there are a total of K groups of historical matrices and M matrices to be detected. The computing device can calculate the relative deviation f(P) of each matrix to be detected P with respect to each group of historical matrices P t according to formula (4). The larger the relative deviation, the greater the difference between the corresponding matrix to be detected and a corresponding group of historical matrices.

[0114]

[0115] In formula (4), t represents a sub - time period, m represents the total number of sub - time periods included in a time period, and in the embodiments of the present application, m = M. δ represents an anomaly detection threshold, which can be set as needed.

[0116] In some other embodiments, the computing device can also calculate according to the formula calculate the relative deviation of each of the above - mentioned M matrices to be detected with respect to each of the above - mentioned K×M historical matrices, obtaining a total of M×K×M relative deviations. Then, the computing device determines the minimum value among these M×K×M relative deviations as the relative deviation of the matrix to be detected with respect to the historical state - transition matrix. Alternatively, the computing device determines the second - smallest value among these M×K×M relative deviations as the relative deviation of the matrix to be detected with respect to the historical state - transition matrix.

[0117] In the embodiments where the state - transition matrix is determined according to a coarse - grained time window, the historical state - transition matrix includes K historical matrices, and the matrix to be detected includes one matrix to be detected. Based on this, the computing device determines the relative deviation of this matrix to be detected with respect to each of these K historical matrices, obtaining K relative deviations. The computing device determines the minimum value among these K relative deviations as the relative deviation of the matrix to be detected with respect to the historical state - transition matrix. Alternatively, the computing device determines the second - smallest value among these K relative deviations as the relative deviation of the matrix to be detected with respect to the historical state - transition matrix.

[0118] Step 104: Based on this relative deviation, determine the anomaly detection result of the operation data to be detected.

[0119] In the embodiments of the present application, if the relative deviation of the matrix to be detected with respect to the historical state - transition matrix exceeds the anomaly detection threshold, the computing device determines that the operation data to be detected is abnormal, that is, the research object has performed an abnormal operation during the time period corresponding to the operation data to be detected. If the relative deviation of the matrix to be detected with respect to the historical state - transition matrix does not exceed the anomaly detection threshold, the computing device determines that the operation data to be detected is not abnormal.

[0120] Among them, the smaller the anomaly detection threshold, the greater the possibility that the operation data to be detected is determined to be abnormal, and the larger the anomaly detection threshold, the smaller the possibility that the operation data to be detected is determined to be abnormal.

[0121] In some other embodiments, the computing device may determine the anomaly detection result of the operation data to be detected in other ways based on the relative deviation of the state transition matrix to be detected with respect to the historical state transition matrix. For example, the computing device determines the anomaly detection result of the operation data to be detected through a neural network based on the relative deviation of the state transition matrix to be detected with respect to the historical state transition matrix.

[0122] When it is determined that there is an anomaly in the operation data to be detected, the computing device may generate an alarm message, which is used to prompt that there may be abnormal operations of the target object during the detection time period (such as the current time period).

[0123] To verify the effectiveness of the method, based on the data of a certain company's actual network in the above text, Example 3 and Example 4 are given to verify the performance of this solution.

[0124] Example 3:

[0125] Taking the data of a certain company's actual network within 14 days as an example, this example gives the practice of abnormal behavior judgment. The log data of this network records approximately 8,270,000,000 pieces within these 14 days, with a size of approximately 720 GB. The research object is users, and the number of users is 1973. If one day is taken as a time period and the time window T is set to a fine-grained 1 hour, then the log data of each user within these 14 days can be divided into 336 operation sequences. It should be understood that in this Example 3, K = 14 and M = 24, indicating that there are 14 time periods of historical operation data, and each time period contains 24 sub-time periods.

[0126] Based on the log data from February 10, 2021 to February 23, 2021, the historical operation data is determined, and the operation data to be detected is obtained based on the log data of February 24, 2021. The time window T = 1 hour. In the time period timeslot = 2021022400, there is an abnormal user User1, and the total number of valid historical data of this user contains 336 operation sequences.

[0127] Figure 6 The distribution of f(P) of user User1 calculated according to formula (4) is given. From Figure 6 It can be seen that the maximum value of f(P) of user User1 within these 14 days is 1.45889, and the corresponding sub-time period is timeslot = 2021021215, and the minimum value is 0.56526, and the corresponding sub-time period is timeslot = 2021021700. Taking the anomaly detection threshold as 0.5, since this minimum value is greater than 0.5, the operation data of User1 in the sub-time period timeslot = 2021022400 will be determined to contain abnormal operation sequences, that is, there is an anomaly in the operation data to be detected.

[0128] Figure 7 The operation sequence diagram of user User1 within the sub - time slot timeslot = 2021022400 is given, and its corresponding state transition matrix is shown in Table 6.

[0129] Table 6

[0130]

[0131] Figure 8 The operation sequence diagram of the sub - time slot timeslot = 2021021700, which is closest to the operation characteristics of user User1 in the sub - time slot timeslot = 2021022400, is given, and its corresponding state transition matrix is shown in Table 7.

[0132] Table 7

[0133]

[0134] Comparison Figure 7 and Figure 8 It can be seen that from the perspective of the time series, the state transitions of the operation sequences of user User1 in these two sub - time slots are not the same. However, by comparing the two corresponding state transition matrices of these two diagrams (i.e., Table 6 and Table 7), it can be seen that from a statistical sense, the numerical values of the state transition matrices corresponding to these two sub - time slots are very close. The absolute deviations of these two state transition matrices are shown in Table 8.

[0135] Table 8

[0136]

[0137] Figure 9 The operation sequence diagram of user User1 within the sub - time slot timeslot = 2021021215 is given, and the state transition matrix corresponding to the operation sequence of this sub - time slot is shown in Table 9.

[0138] Table 9

[0139]

[0140] Comparison Figure 7 and Figure 9 It can be seen that within these two sub - time slots, from the perspective of the time series, there are significant differences in the state changes of the operation sequences of user User1. From the two corresponding state transition matrices of these two sub - time slots (i.e., Table 6 and Table 9), it can be seen that from a statistical sense, the differences in the values of the state transition matrices corresponding to the operation sequences within these two sub - time slots are also very obvious. The absolute deviations of the state transition matrices corresponding to these two sub - time slots, that is, the difference matrix obtained by subtraction operation, are shown in Table 10.

[0141] Table 10

[0142]

[0143]

[0144] Although the numerical deviation (i.e., the absolute deviation) of the state transition matrices corresponding to the sub - time slots timeslot = 2021022400 and timeslot = 2021021700 for user User1 is small, since most of the elements in these two state transition matrices are zero (only 10 out of 81 elements are non - zero), from the perspective of matrix norm, the relative deviation of these two state transition matrices is still large. Thus, detecting abnormal operations according to the relative deviation rather than the absolute deviation is more accurate.

[0145] Example 4:

[0146] Example 4 uses the log data in the same time period as Example 3 above. During the time period timeslot = 2021022400, there is a normal user User2, and the total number of valid historical operation data of this user contains 336 historical operation sequences.

[0147] Figure 10 The distribution of f(P) of user User2 calculated by formula (4) is given. Figure 10 It can be seen that the maximum value of f(P) of user User2 in these 14 days is 1.06547, and the corresponding sub - time slot is timeslot = 2021021323, and the minimum value is 0.02527, and the corresponding sub - time slot is timeslot = 2021021709. Taking the abnormal detection threshold of 0.5 as an example, since this minimum value is less than 0.5, the operation data of User2 in the sub - time slot timeslot = 2021022400 will be determined as normal operation data, that is, it does not contain abnormal operations.

[0148] Figure 11 The operation sequence diagram of user User2 in the time period timeslot = 2021022400 is given, and its corresponding state transition matrix is shown in Table 11.

[0149] Table 11

[0150]

[0151] Figure 12The operation sequence diagram of the sub - time slot timeslot = 2021021709, which is closest to the operation characteristics of user User2 in the sub - time slot timeslot = 2021022400, is given, and its corresponding state transition matrix is shown in Table 12.

[0152] Table 12

[0153]

[0154]

[0155] Comparison Figure 11 and Figure 12 It can be seen that from the perspective of the time series, there is a certain similarity in the state transition of the operation sequences of user User2 in these two sub - time slots, but they are not exactly the same. By comparing the corresponding state transition matrices of these two diagrams (i.e., Table 11 and Table 12), it can be seen that from a statistical sense, the state transition matrices corresponding to these two sub - time slots are very close in value, and the absolute deviation of these two state transition matrices is shown in Table 13.

[0156] Table 13

[0157]

[0158] It can be seen that there are operation sequence transfer situations in the historical records of user User2 that are very close to the current detection period, and the value used to measure the "closeness" (i.e., the minimum value of f(P)) is less than the anomaly detection threshold. Therefore, it is determined that the operation data to be detected of user User2 is normal, that is, it does not contain abnormal operations.

[0159] As can be seen from the embodiments of the anomaly detection method introduced above, this solution mainly characterizes the key operation behaviors of the research object accessing the computer system as a Markov state transition model, and realizes the detection of abnormal behaviors by comparing the state transition characteristics shown in multiple accesses. Since a large amount of information related to the operations of the research object is recorded in the log data of the computer system, but not all information is key information for the detection of abnormal behaviors, usually the log information can be screened first. The data retained after screening can be called the key data of the system log. The key data of the system log contains the key operations of the research object, which can be sorted according to the time sequence of the key operations, and the state transition matrix in the Markov process can be constructed accordingly. By determining the relative deviation of the state transition matrix of the operation data to be detected with respect to the historical operation data, a judgment can be given on whether the operation behavior of the research object is abnormal.

[0160] Based on the above discussion, it can be known that this solution mainly includes two stages. Please combine Figure 13An explanation of these two stages is provided below. Refer to Figure 13 , the first stage is used to screen the data containing key operations from the original log data, and construct the state transition matrix in the Markov chain based on the screened data, including the state transition matrix of historical operation data and the state transition matrix of the operation data to be detected. The second stage is used to perform operation sequence detection, that is, according to the state transition matrices of historical operation data and the operation data to be detected, detect whether the operation data to be detected may be abnormal, that is, obtain the abnormal detection result.

[0161] In summary, in the embodiments of the present application, by constructing the state transition matrix of historical operation data and the state transition matrix of the operation data to be detected, the relative deviation of the state transition matrix to be detected relative to the historical state transition matrix is determined, so as to determine the abnormal detection result based on the relative deviation. Since the state transition matrix can characterize the characteristics of the corresponding data changing with the operations included over time, the relative deviation of the state transition characteristics of the abnormal operation sequence relative to the state transition characteristics of the normal operation sequence is relatively large, and the relative deviation can better characterize the degree of the operation data to be detected relative to the average situation of the historical operation data compared to the absolute deviation. Therefore, the accuracy of the abnormal operation detection in this solution is higher.

[0162] Figure 14 FIG. 1400 is a schematic structural diagram of an abnormal operation detection device 1400 provided by an embodiment of the present application. The device 1400 can be implemented as part or all of a computing device by software, hardware, or a combination of both. The computing device can be the computing device in the above method embodiment, or Figures 15 to 17 any of the computing devices shown in Figure 14 , the device 1400 includes: an acquisition module 1401, a first determination module 1402, a second determination module 1403, and a third determination module 1404.

[0163] The acquisition module 1401 is configured to acquire the historical operation data and the operation data to be detected of the target object, and the historical operation data includes a normal operation sequence;

[0164] The first determination module 1402 is configured to determine the historical state transition matrix and the state transition matrix to be detected. The historical state transition matrix is the state transition matrix of the historical operation data, and the state transition matrix to be detected is the state transition matrix of the operation data to be detected. The state transition matrix characterizes the characteristics of the operations included in the corresponding data changing with time;

[0165] The second determination module 1403 is configured to determine the relative deviation of the state transition matrix to be detected relative to the historical state transition matrix;

[0166] The third determination module 1404 is configured to determine the abnormal detection result of the operation data to be detected based on the relative deviation.

[0167] Optionally, the historical operation data includes historical data for K time periods, the operation data to be detected includes data to be detected for one time period, each time period includes M sub - time periods, both K and M are integers not less than 1, the historical state transition matrix includes K groups of historical matrices corresponding one - to - one with the historical data of these K time periods, each group of historical matrices includes M historical matrices corresponding one - to - one with the historical data of the corresponding M sub - time periods, and the state transition matrix to be detected includes M matrices to be detected corresponding one - to - one with the data to be detected of the corresponding M sub - time periods;

[0168] The second determination module 1403 includes:

[0169] A first determination sub - module, configured to determine the relative deviation of each matrix to be detected in these M matrices to be detected with respect to each group of historical matrices in these K groups of historical matrices, obtaining K×M relative deviations;

[0170] A second determination sub - module, configured to determine the relative deviation of the state transition matrix to be detected with respect to the historical state transition matrix based on these K×M relative deviations.

[0171] Optionally, the second determination sub - module is specifically configured to: determine the minimum value among these K×M relative deviations as the relative deviation of the state transition matrix to be detected with respect to the historical state transition matrix.

[0172] Optionally, the first determination sub - module is specifically configured to:

[0173] For a reference historical group among these K groups of historical matrices, determine the N - norm of each historical matrix in the reference historical group, obtaining M N - norms, where N is an integer greater than 1, and the reference historical group is any one of these K groups of historical matrices;

[0174] For a first matrix to be detected among the M matrices to be detected, determine the N - distance between the first matrix to be detected and each historical matrix in the reference historical group, obtaining M N - distances, where the first matrix to be detected is any one of the M matrices to be detected;

[0175] Determine the ratio of each of these M N - distances to the corresponding N - norm among these M N - norms, obtaining M ratios;

[0176] Determine the minimum value among these M ratios as the relative deviation of the first matrix to be detected with respect to the reference historical group.

[0177] Optionally, the third determination module 1404 includes:

[0178] A third determination sub - module, configured to determine that there is an abnormality in the operation data to be detected if the relative deviation of the state transition matrix to be detected with respect to the historical state transition matrix exceeds the anomaly detection threshold.

[0179] Among them, the above-mentioned acquisition module, first determination module, second determination module, and third determination module can all be implemented by software or by hardware. Exemplarily, next, taking the first determination module as an example, the implementation manner of the first determination module will be introduced. Similarly, the implementation manners of the acquisition module, second determination module, and third determination module can refer to the implementation manner of the first determination module.

[0180] As an example of a software functional unit, the first determination module may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above-mentioned computing instance may be one or more. For example, the first determination module may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running this code may be distributed in the same region, or may be distributed in different regions. Further, the multiple hosts / virtual machines / containers for running this code may be distributed in the same availability zone (AZ), or may be distributed in different AZs, and each AZ includes one data center or multiple geographically proximate data centers. Among them, generally one region may include multiple AZs.

[0181] Similarly, the multiple hosts / virtual machines / containers for running this code may be distributed in the same virtual private cloud (VPC), or may be distributed in multiple VPCs. Among them, generally one VPC is set within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set in each VPC, and the interconnection between VPCs is realized through the communication gateway.

[0182] As an example of a hardware functional unit, the first determination module may include at least one computing device, such as a server. Alternatively, the first determination module may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). Among them, the above PLD may be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0183] The multiple computing devices included in the first determination module may be distributed in the same region or in different regions. The multiple computing devices included in the first determination module may be distributed in the same availability zone (AZ) or in different AZs. Similarly, the multiple computing devices included in the first determination module may be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Among them, the multiple computing devices may be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0184] It should be understood that in other embodiments, the first determination module may be used to execute any step in the abnormal operation detection method, and the acquisition module, the second determination module, and the third determination module may all be used to execute any step in the abnormal operation detection method. The steps to be implemented by the acquisition module, the first determination module, the second determination module, and the third determination module can be specified as needed. The abnormal operation detection device's all functions are realized by the acquisition module, the first determination module, the second determination module, and the third determination module respectively implementing different steps in the abnormal operation detection method.

[0185] In the embodiments of the present application, since the state transition matrix can characterize the characteristics of the corresponding data changing with time along with the included operations, the relative deviation of the state transition characteristics of the abnormal operation sequence with respect to the state transition characteristics of the normal operation sequence is relatively large, and the relative deviation can more accurately characterize the degree of the average situation of the operation data to be detected with respect to the historical operation data compared to the absolute deviation. Therefore, the accuracy of the abnormal operation detection in this solution is higher.

[0186] It should be noted that: when detecting abnormal operations, the abnormal operation detection device provided in the above embodiments is only exemplified by the division of the above functional modules. In actual applications, the above functions can be allocated to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the abnormal operation detection device provided in the above embodiments and the embodiments of the abnormal operation detection method belong to the same concept. For the specific implementation process, please refer to the method embodiments and will not be elaborated here.

[0187] An embodiment of the present application also provides a computing device 100. As Figure 15 shown, the computing device 100 includes: a bus 102, a processor 104, a memory 106, and a communication interface 108. The processor 104, the memory 106, and the communication interface 108 communicate with each other through the bus 102. The computing device 100 may be a server or a terminal device. It should be understood that the number of processors and memories in the computing device 100 in the embodiments of the present application is not limited.

[0188] The bus 102 may be a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 15 only one line is shown in the figure, but it does not mean that there is only one bus or one type of bus. The bus 102 may include a path for transmitting information between various components of the computing device 100 (for example, the memory 106, the processor 104, the communication interface 108).

[0189] The processor 104 may include any one or more of processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0190] The memory 106 may include volatile memory, such as random access memory (RAM). The processor 104 may also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0191] Executable program code is stored in the memory 106, and the processor 104 executes the executable program code to implement the functions of the foregoing acquisition module, first determination module, second determination module, and third determination module respectively, so as to implement the abnormal operation detection method. That is, instructions for executing the abnormal operation detection method are stored on the memory 106.

[0192] The communication interface 108 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 100 and other devices or a communication network.

[0193] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device may be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device may also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0194] As Figure 16 shown, the computing device cluster includes at least one computing device 100. Instructions for executing the abnormal operation detection method may be stored in the memory 106 of one or more of the computing devices 100 in the computing device cluster.

[0195] In some possible implementation manners, partial instructions for executing the abnormal operation detection method may also be stored in the memory 106 of one or more of the computing devices 100 in the computing device cluster. In other words, a combination of one or more computing devices 100 may jointly execute the instructions for executing the abnormal operation detection method.

[0196] It should be noted that the memories 106 in different computing devices 100 in the computing device cluster may store different instructions, which are respectively used to execute partial functions of the abnormal operation detection device. That is, the instructions stored in the memories 106 of different computing devices 100 may implement the functions of one or more of the acquisition module, the first determination module, the second determination module, and the third determination module.

[0197] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. Among them, the network can be a wide area network or a local area network, etc. Figure 17 shows a possible implementation. As Figure 17 shown, two computing devices 100A and 100B are connected via a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this type of possible implementation, the memory 106 in the computing device 100A stores instructions for executing the functions of the acquisition module in the exception detection device. At the same time, the memory 106 in the computing device 100B stores instructions for executing the functions of the first determination module, the second determination module, and the third determination module in the exception detection device.

[0198] Figure 17 The connection method between the computing device clusters shown can be considered. Since the exception operation detection method provided in this application requires a large amount of data storage and data processing, etc., it is considered to hand over the functions implemented by the first determination module, the second determination module, and the third determination module to the computing device 100B for execution.

[0199] It should be understood that Figure 17 the functions of the computing device 100A shown in

[0200] can also be completed by multiple computing devices 100. Similarly, the functions of the computing device 100B can also be completed by multiple computing devices 100. Figure 16 and Figure 17 shown in the connection method of the computing device cluster. The difference is that the memory 106 in one or more computing devices 100 in this computing device cluster can store the same instructions for executing the exception operation detection method.

[0201] In some possible implementations, the memory 106 in one or more computing devices 100 in this computing device cluster can also store some instructions for executing the exception operation detection method respectively. In other words, a combination of one or more computing devices 100 can jointly execute the instructions for executing the exception operation detection method.

[0202] This application embodiment also provides a computer program product containing instructions. This computer program product can be software or a program product containing instructions that can run on a computing device or be stored in any available medium. When this computer program product runs on at least one computing device, it causes at least one computing device to execute the steps of the exception operation detection method.

[0203] Embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium may be any available medium that can be stored by a computing device or a data storage device such as a data center that includes one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state drive), etc. The computer-readable storage medium includes instructions that direct the computing device to perform the steps of the abnormal operation detection method.

[0204] The system architecture and business scenarios described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those of ordinary skill in the art will know that with the evolution of the network architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0205] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer, or a data storage device such as a server or a data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a digital versatile disc (DVD)), or a semiconductor medium (e.g., a solid state disk (SSD)), etc. It should be noted that the computer-readable storage medium mentioned in the embodiments of the present application may be a non-volatile storage medium, in other words, a non-transitory storage medium.

[0206] It should be understood that the "at least one" mentioned herein refers to one or more, and the "multiple" refers to two or more. In the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B may mean A or B. The "and / or" herein is merely a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in order to clearly describe the technical solutions of the embodiments of the present application, in the embodiments of the present application, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and effects. Those skilled in the art can understand that the terms such as "first" and "second" do not limit the quantity and execution order, and the terms such as "first" and "second" do not necessarily limit to be different.

[0207] It should be noted that the information (including but not limited to user equipment information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the embodiments of the present application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data need to comply with the relevant laws, regulations and standards of the relevant countries and regions. For example, the log data involved in the embodiments of the present application are all obtained under sufficient authorization.

[0208] The above are the embodiments provided by the present application, which are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. An abnormal operation detection method, characterized in that, The method includes: Filtering the log data of the target object to obtain the historical operation data of the target object in the historical period before the current period and the operation data to be detected in the current period. The historical operation data includes a normal operation sequence. Both the historical operation data and the operation data to be detected include one or more key operations. The key operation refers to an operation related to the security of the computer system and / or user information. The historical operation data includes historical data of K time periods, and the operation data to be detected includes to-be-detected data of one time period. Each time period includes M sub-time periods, and both K and M are integers not less than 1. Determining a historical state transition matrix and a to-be-detected state transition matrix according to the granularity of the sub-time periods. The historical state transition matrix is the state transition matrix of the historical operation data, and the to-be-detected state transition matrix is the state transition matrix of the operation data to be detected. The state transition matrix characterizes the feature of the operations included in the corresponding data changing over time. The state transition matrix includes multiple elements, and the multiple elements characterize the probability of each key operation in the corresponding data being transferred to another key operation. The probability is determined based on the number of times that the corresponding key operation is transferred to another key operation. Among them, the historical state transition matrix includes K groups of historical matrices corresponding one by one to the historical data of the K time periods, and each group of historical matrices includes M historical matrices corresponding one by one to the historical data of the corresponding M sub-time periods. The to-be-detected state transition matrix includes M to-be-detected matrices corresponding one by one to the to-be-detected data of the corresponding M sub-time periods. Determining the relative deviation of each to-be-detected matrix in the M to-be-detected matrices with respect to each group of historical matrices in the K groups of historical matrices, and obtaining K×M relative deviations. Determining the minimum value or the second minimum value among the K×M relative deviations as the relative deviation of the to-be-detected state transition matrix with respect to the historical state transition matrix. Determining the anomaly detection result of the operation data to be detected based on the relative deviation of the to-be-detected state transition matrix with respect to the historical state transition matrix.

2. The method according to claim 1, characterized in that, The determining the relative deviation of each to-be-detected matrix in the M to-be-detected matrices with respect to each group of historical matrices in the K groups of historical matrices includes: For a reference historical group in the K groups of historical matrices, determining the N-order norm of each historical matrix in the reference historical group to obtain M N-order norms. N is an integer greater than 1, and the reference historical group is any one of the K groups of historical matrices. For a first to-be-detected matrix in the M to-be-detected matrices, determining the N-order distance between the first to-be-detected matrix and each historical matrix in the reference historical group to obtain M N-order distances. The first to-be-detected matrix is any one of the M to-be-detected matrices. Determining the ratio of each N-order distance in the M N-order distances to the corresponding N-order norm in the M N-order norms to obtain M ratios. Determining the minimum value among the M ratios as the relative deviation of the first to-be-detected matrix with respect to the reference historical group.

3. The method according to claim 1 or 2, characterized in that, Determining the anomaly detection result of the to-be-detected operation data based on the relative deviation of the to-be-detected state transition matrix with respect to the historical state transition matrix includes: If the relative deviation of the to-be-detected state transition matrix with respect to the historical state transition matrix exceeds the anomaly detection threshold, it is determined that there is an anomaly in the to-be-detected operation data.

4. An abnormal operation detection device, characterized in that, The device includes: An acquisition module, configured to screen the log data of a target object to obtain historical operation data of the target object in a historical time period before the current time period and to-be-detected operation data in the current time period. The historical operation data includes a normal operation sequence. Both the historical operation data and the to-be-detected operation data include one or more key operations. The key operation refers to an operation related to the security of a computer system and / or user information. The historical operation data includes historical data of K time periods, and the to-be-detected operation data includes to-be-detected data of one time period. Each time period includes M sub-time periods, and both K and M are integers not less than 1; A first determination module, configured to determine a historical state transition matrix and a to-be-detected state transition matrix according to the granularity of the sub-time periods. The historical state transition matrix is the state transition matrix of the historical operation data, and the to-be-detected state transition matrix is the state transition matrix of the to-be-detected operation data. The state transition matrix characterizes the feature of the operations included in the corresponding data changing over time. The state transition matrix includes multiple elements, and the multiple elements characterize the probability of each key operation in the corresponding data being transferred to another key operation. The probability is determined based on the number of times each key operation is transferred to another key operation; wherein, the historical state transition matrix includes K groups of historical matrices corresponding one-to-one to the historical data of the K time periods, and each group of historical matrices includes M historical matrices corresponding one-to-one to the historical data of the corresponding M sub-time periods. The to-be-detected state transition matrix includes M to-be-detected matrices corresponding one-to-one to the to-be-detected data of the corresponding M sub-time periods; A second determination module, configured to determine the relative deviation of the to-be-detected state transition matrix with respect to the historical state transition matrix; A third determination module, configured to determine the anomaly detection result of the to-be-detected operation data based on the relative deviation of the to-be-detected state transition matrix with respect to the historical state transition matrix; Wherein, the second determination module includes: A first determination sub-module, configured to determine the relative deviation of each to-be-detected matrix in the M to-be-detected matrices with respect to each group of historical matrices in the K groups of historical matrices, obtaining K×M relative deviations; A second determination sub-module, configured to determine the minimum value or the second minimum value among the K×M relative deviations as the relative deviation of the to-be-detected state transition matrix with respect to the historical state transition matrix.

5. The device according to claim 4, characterized in that, The first determination sub-module is specifically configured to: For a reference historical group in the K groups of historical matrices, determine the N-norm of each historical matrix in the reference historical group, obtaining M N-norms. N is an integer greater than 1, and the reference historical group is any one of the K groups of historical matrices; For the first matrix to be detected among the M matrices to be detected, determine the N-th order distance between the first matrix to be detected and each historical matrix in the reference historical group, obtaining M N-th order distances, where the first matrix to be detected is any one of the M matrices to be detected; Determine the ratio of each N-th order distance among the M N-th order distances to the corresponding N-th order norm among the M N-th order norms, obtaining M ratios; Determine the minimum value among the M ratios as the relative deviation of the first matrix to be detected with respect to the reference historical group.

6. The device according to claim 4 or 5, characterized in that, The third determination module includes: A third determination sub-module, configured to determine that the operation data to be detected is abnormal if the relative deviation of the state transition matrix to be detected with respect to the historical state transition matrix exceeds the anomaly detection threshold.

7. A computing device, characterized in that, The computing device includes a processor and a memory; The processor is configured to execute the instructions stored in the memory, so that the computing device executes the steps of the method according to any one of claims 1-3.

8. A cluster of computing devices, characterized in that, Including at least one computing device, each computing device includes a processor and a memory; The processor of the at least one computing device is configured to execute the instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the steps of the method according to any one of claims 1-3.

9. A computer-readable storage medium, characterized in that, Including computer program instructions, which implement the method according to any one of claims 1-3 when executed by a processor.

10. A computer program product comprising instructions, characterized in that, When the instructions are executed by a processor, the method according to any one of claims 1-3 is implemented.

Citation Information

Patent Citations

  • Log detection method and device and computer readable storage medium

    CN114817189A