Abnormal access behavior detection method and electronic device

By generating directed graphs and using multi-models of unsupervised algorithms to detect anomaly access behavior, the problems of insufficient flexibility of baseline rules and high dependence on machine learning data in the prior art are solved, and higher detection accuracy and wider anomaly access coverage are achieved.

CN114143015BActive Publication Date: 2025-07-04HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010808974.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-12
Publication Date
2025-07-04
Estimated Expiration
2040-08-12

AI Technical Summary

Technical Problem

In the existing abnormal access behavior detection methods, the baseline rules are not flexible enough and rely on manual settings. The machine learning method has high requirements for training data, making it difficult to effectively establish a model with few samples, resulting in low detection accuracy.

Method used

Based on the first log data, the first, second and third models of the unsupervised algorithm are used to identify abnormal access behavior, including identifying multi-node jump login, cross-service group access, and abnormal access that does not match the historical behavior, and detecting through topological sorting of directed graphs, community discovery and embedding models.

Benefits of technology

No historical attack sample support is required, which reduces the requirements for data source quality, improves the accuracy of abnormal access behavior detection, and covers multiple abnormal access scenarios through correlation analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114143015B_ABST
    Figure CN114143015B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides an abnormal access behavior detection method and an electronic device, which relate to the field of communication technologies and include: generating a directed graph according to first log data; using one or more of a first model, a second model, or a third model to identify abnormal access behaviors in the directed graph and determining an abnormal detection result; wherein, the first model is used to identify abnormal access behaviors of multi-node jump logins according to the directed graph; the second model is used to identify abnormal access behaviors of cross-business group access according to the directed graph; the third model is used to identify abnormal access behaviors that do not conform to historical access behaviors according to the directed graph. The embodiment of the present application can generate a directed graph based on the first log data, identify abnormal access behaviors in the graph from the perspective of the directed graph, without the support of historical attack samples, thereby reducing the requirements for the quality of data sources and achieving an improvement in the accuracy of abnormal access behavior detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to communication technologies, and in particular to a method for detecting abnormal access behaviors and an electronic device. Background Art

[0002] With the continuous development of Internet technologies, the security of network devices has become increasingly important. For example, due to the interconnection between devices, after an attacker invades a server through methods such as weak passwords, security vulnerabilities, and system backdoors, other devices interacting with the server may be exposed to security risks. Therefore, it is necessary to detect abnormal accesses in network devices to enhance security.

[0003] Currently, there are two commonly used methods for detecting abnormal accesses. The first one is baseline detection based on statistics and profiling. For example, defense personnel establish an access baseline through historical interconnection records, define normal access behaviors, and determine accesses that occur for the first time, exceed thresholds, or deviate from the profile as abnormal access behaviors. The second one is machine learning detection based on pattern features. For example, defense personnel collect access logs over a period of time, extract features, learn normal behavior patterns, and use machine learning algorithms to build a detection model.

[0004] However, in the first method for detecting abnormal access behaviors mentioned above, there is a problem that baseline rules are usually defined manually, resulting in inflexible detection. In the second method for detecting abnormal access behaviors mentioned above, machine learning algorithms require a large amount of training data, and it is difficult to effectively build a model when the types of samples are few, which has certain limitations. Summary of the Invention

[0005] Embodiments of this application provide a method for detecting abnormal access behaviors and an electronic device. A directed graph is generated based on first log data, and abnormal access behaviors in the graph are identified from the perspective of the directed graph, without the support of historical attack samples, thereby reducing the requirements for the quality of data sources and achieving an improvement in the accuracy of detecting abnormal access behaviors.

[0006] In a first aspect, an embodiment of the present application provides an abnormal access behavior detection method, including: generating a directed graph according to first log data; wherein, the directed graph includes: a plurality of nodes for identifying devices, and directed access relationships between the plurality of nodes; using one or more of a first model, a second model, or a third model to identify abnormal access behaviors in the directed graph and determine an abnormal detection result; wherein, the first model is used to identify abnormal access behaviors of multi-node jump logins according to the directed graph; the second model is used to identify abnormal access behaviors of cross-business group access according to the directed graph; the third model is used to identify abnormal access behaviors that do not conform to historical access behaviors according to the directed graph. In this way, the embodiment of the present application can generate a directed graph based on the first log data, identify abnormal access behaviors in the graph from the perspective of the directed graph, without the support of historical attack samples, thereby reducing the requirements for the quality of the data source and achieving an improvement in the accuracy of detecting abnormal access behaviors.

[0007] In a possible implementation manner, the first model, the second model, and the third model are all implemented using unsupervised algorithms. In this way, the embodiment of the present application can be independent of prior knowledge and specific feature thresholds input by humans and does not require the support of historical attack samples, thereby avoiding the requirements for the diversity of source data in existing detection models.

[0008] In a possible implementation manner, the embodiment of the present application uses the first model to identify abnormal access behaviors in the directed graph, including: for the source node and the destination node among the plurality of nodes, calculating the maximum number of hops from the source node to the destination node; wherein, the maximum number of hops is used to represent the number of hops for continuous access with the destination node as a springboard; identifying an access behavior with a maximum number of hops greater than a first threshold as an abnormal access behavior. In this way, the first model can be used to identify abnormal access behaviors of multi-node jump logins according to the directed graph.

[0009] In a possible implementation manner, the embodiment of the present application uses the second model to identify abnormal access behaviors in the directed graph, including: classifying a plurality of nodes into the community where the neighbor node with the largest gain is located; compressing the nodes classified into the same community into a first node until the result of the classification no longer changes; identifying the access behavior corresponding to the first node with cross-community access as an abnormal access behavior. In this way, the second model can be used to identify abnormal access behaviors of cross-business group access according to the directed graph.

[0010] In a possible implementation, the embodiment of the present application uses a third model to identify abnormal access behaviors in a directed graph, including: converting the nodes in the directed graph into embedding vectors; for the source node and the destination node among multiple nodes, normalizing the embedding vector matrix corresponding to the set of predecessor nodes of the destination node to obtain a set of normalized unit vectors; wherein, the set of predecessor nodes of the destination node is the set of nodes in the directed graph that point to the destination node; calculating the cosine similarity between the embedding vector corresponding to the source node and the set of normalized unit vectors; and identifying the access behavior with a cosine similarity less than a second threshold as an abnormal access behavior. In this way, the third model can be used to identify abnormal access behaviors that do not conform to historical access behaviors according to the directed graph.

[0011] In a possible implementation, the embodiment of the present application normalizes the embedding vector matrix corresponding to the set of predecessor nodes of the destination node, including: training the embedding vector matrix corresponding to the set of predecessor nodes of the destination node using a one-class support vector machine; and normalizing the trained embedding vector matrix.

[0012] In a possible implementation, using one or more of the first model, the second model, or the third model to identify abnormal access behaviors in a directed graph and determine an anomaly detection result, including: using one or more of the first model, the second model, or the third model to identify abnormal access behaviors in the directed graph to obtain multiple identification results; and performing correlation analysis on the multiple identification results to obtain an anomaly detection result. In this way, anomaly access behavior detection can be performed from multiple dimensions, thereby comprehensively monitoring multiple types of abnormal access behaviors.

[0013] In a possible implementation, the embodiment of the present application performs correlation analysis on multiple identification results to obtain an anomaly detection result, including: constructing a linear function between the multiple identification results; and determining the access behavior corresponding to the linear function as the anomaly detection result when the value of the linear function meets a preset anomaly condition. In this way, the weights of the detection results of the first model, the second model, and the third model in the correlation analysis process and the preset anomaly conditions that the linear function values meet can be set, so as to cover multiple abnormal access scenarios.

[0014] In a possible implementation, the method further includes: updating the directed graph according to the anomaly detection result; and updating the hyperparameters of the first model, the hyperparameters of the second model, and / or the hyperparameters of the third model according to the anomaly detection result. In this way, the model can be optimized and updated in real time according to the detection result.

[0015] In a possible implementation, the method further includes: obtaining second log data; wherein, the generation time of the second log data is earlier than the generation time of the first log data; generating a directed graph of the second log data; loading the directed graph of the second log data, and the model parameters related to the first model, the second model, or the third model, into the first model, the second model, or the third model. In this way, before detecting abnormal access behavior for the first log data, the first model, the second model, and the third model have completed model initialization, so as to achieve real-time detection of the first log data.

[0016] In a possible implementation, the embodiment of the present application generates a directed graph according to the first log data, including: regularly obtaining the first log data; generating a first directed graph corresponding to the first log data; filtering out the part of the first directed graph that overlaps with the third directed graph of the third log data to obtain a directed graph; wherein, the generation time of the third log data is earlier than the generation time of the first log data, and the difference between the generation time of the third log data and the generation time of the first log data is less than a time threshold. In this way, the part of the directed graph of the first log that overlaps with the directed graph of the third log can be filtered out, so that it is not necessary to repeatedly perform abnormal access detection on the overlapping part of the log data, reducing the computational amount of the model.

[0017] In a possible implementation, the node is the Internet Protocol (IP) address of a device, wherein, among the directed access relationships between multiple nodes, there is one or more of the following: the maximum number of accesses per hour between multiple nodes, the total historical number of accesses, or the latest access time.

[0018] In a possible implementation, the method further includes: sending an alarm message to a target object according to the abnormal detection result. In this way, risk control can be performed according to the alarm message and closed in a timely manner.

[0019] In a possible implementation, the alarm message includes one or more of the following: the log information corresponding to the abnormal detection result, the cause of the alarm, or the recommended handling method.

[0020] Second aspect, an embodiment of the present application provides an abnormal access behavior detection device, which can be a terminal device, or a chip or chip system in the terminal device. The abnormal access behavior detection device may include a processing unit. When the abnormal access behavior detection device is a terminal device, the processing unit may be a processor. The abnormal access behavior detection device may further include a storage unit, which may be a memory. The storage unit is used to store instructions, and the processing unit executes the instructions stored in the storage unit to enable the terminal device to implement an abnormal access behavior detection method described in the first aspect or any possible implementation manner of the first aspect. When the abnormal access behavior detection device is a chip or chip system in the terminal device, the processing unit may be a processor. The processing unit executes the instructions stored in the storage unit to enable the terminal device to implement an abnormal access behavior detection method described in the first aspect or any possible implementation manner of the first aspect. The storage unit may be a storage unit in the chip (for example, registers, caches, etc.), or a storage unit outside the chip in the terminal device (for example, read-only memory, random access memory, etc.).

[0021] Exemplarily, the processing unit is configured to generate a directed graph according to the first log data; wherein, the directed graph includes: a plurality of nodes for identifying devices, and directed access relationships between the plurality of nodes; the processing unit is further configured to use one or more of the first model, the second model, or the third model to identify abnormal access behaviors in the directed graph and determine an abnormal detection result; wherein, the first model is used to identify abnormal access behaviors of multi-node jump login according to the directed graph; the second model is used to identify abnormal access behaviors of cross-business group access according to the directed graph; the third model is used to identify abnormal access behaviors that do not conform to historical access behaviors according to the directed graph.

[0022] In a possible implementation manner, the first model, the second model, and the third model are all implemented using unsupervised algorithms.

[0023] In a possible implementation manner, the processing unit is specifically configured to calculate the maximum number of hops from the source node to the destination node according to the source node and the destination node among the plurality of nodes; wherein, the maximum number of hops is used to represent the number of hops for continuous access with the destination node as a springboard; the processing unit is specifically further configured to identify an access behavior with the maximum number of hops greater than a first threshold as an abnormal access behavior.

[0024] In a possible implementation manner, the processing unit is specifically configured to calculate the maximum number of hops from the source node to the destination node according to the source node and the destination node among the plurality of nodes; wherein, the maximum number of hops is used to represent the number of hops for continuous access with the destination node as a springboard; the processing unit is specifically further configured to identify an access behavior with the maximum number of hops greater than a first threshold as an abnormal access behavior.

[0025] In a possible implementation, the processing unit is specifically configured to convert the nodes in the directed graph into embedding vectors; for the source node and the destination node among the multiple nodes, the processing unit is further specifically configured to normalize the embedding vector matrix corresponding to the set of predecessor nodes of the destination node to obtain a set of normalized unit vectors; wherein, the set of predecessor nodes of the destination node is the set of nodes in the directed graph that point to the destination node; the processing unit is further specifically configured to calculate the cosine similarity between the embedding vector corresponding to the source node and the set of normalized unit vectors; the processing unit is further configured to identify an access behavior with a cosine similarity less than a second threshold as an abnormal access behavior.

[0026] In a possible implementation, the processing unit is specifically configured to train the embedding vector matrix corresponding to the set of predecessor nodes of the destination node by using a one-class support vector machine; the processing unit is further specifically configured to normalize the trained embedding vector matrix.

[0027] In a possible implementation, the processing unit is specifically configured to identify abnormal access behaviors in the directed graph by using one or more of a first model, a second model, or a third model to obtain multiple recognition results; the processing unit is further specifically configured to perform correlation analysis on the multiple recognition results to obtain an anomaly detection result.

[0028] In a possible implementation, the processing unit is specifically configured to construct a linear function between the multiple recognition results; the processing unit is further configured to determine the access behavior corresponding to the linear function as an anomaly detection result when the value of the linear function meets a preset anomaly condition.

[0029] In a possible implementation, the processing unit is specifically configured to update the directed graph according to the anomaly detection result; the processing unit is further specifically configured to update the hyperparameters of the first model, the hyperparameters of the second model, and / or the hyperparameters of the third model according to the anomaly detection result.

[0030] In a possible implementation, the processing unit is specifically configured to obtain second log data; the generation time of the second log data is earlier than the generation time of the first log data; the processing unit is further configured to generate a directed graph of the second log data; the processing unit is further configured to load the directed graph of the second log data and the model parameters related to each of the first model, the second model, or the third model into the first model, the second model, or the third model.

[0031] In a possible implementation, the processing unit is further configured to periodically obtain first log data; the processing unit is further configured to generate a first directed graph corresponding to the first log data; the processing unit is further configured to filter out the overlapping part of the first directed graph with the third directed graph of the third log data to obtain a directed graph; the generation time of the third log data is earlier than the generation time of the first log data, and the difference between the generation time of the third log data and the generation time of the first log data is less than a time threshold.

[0032] In a possible implementation, the node is the Internet Protocol (IP) address of a device, and the directed access relationships among multiple nodes include one or more of the following: the maximum number of accesses per hour among multiple nodes, the total number of historical accesses, or the latest access time.

[0033] In a possible implementation, the abnormal access behavior detection device may also include a communication unit. When the abnormal access behavior detection device is a terminal device, the communication unit may be a communication interface or an interface circuit. When the abnormal access behavior detection device is a chip or a chip system within a terminal device, the communication unit may be a communication interface. For example, the communication interface may be an input / output interface, a pin, or a circuit, etc.

[0034] Exemplarily, the communication unit is configured to send an alarm message to a target object according to the abnormal detection result.

[0035] In a possible implementation, the alarm message includes one or more of the following: log information corresponding to the abnormal detection result, the cause of the alarm generation, or a recommended handling method.

[0036] In a third aspect, an embodiment of the present application provides an electronic device, including: a processor, and the processor is configured to run code instructions to implement any method in the first aspect or any possible implementation manner of the first aspect.

[0037] In a fourth aspect, an embodiment of the present application provides an electronic device, including: a processor and an interface circuit, the interface circuit is configured to communicate with other devices; the processor is configured to run code instructions to implement any method in the first aspect or any possible implementation manner of the first aspect.

[0038] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, and the computer-readable storage medium stores instructions, which when executed, implement any method in the first aspect or any possible implementation manner of the first aspect.

[0039] It should be understood that the second aspect to the fifth aspect of the present application correspond to the technical solutions of the first aspect of the present application, and the beneficial effects obtained by each aspect and the corresponding feasible implementation manners are similar and will not be described in detail. Description of the Drawings

[0040] Figure 1 Schematic diagram of an architecture for an abnormal access behavior detection technology provided by an embodiment of the present application;

[0041] Figure 2 Schematic diagram of an existing baseline detection of abnormal access behavior based on statistics and portraits;

[0042] Figure 3 Schematic diagram of an existing detection of abnormal access behavior based on machine learning;

[0043] Figure 4 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application;

[0044] Figure 5 Schematic diagram of an application scenario of an abnormal access behavior detection method provided by an embodiment of the present application;

[0045] Figure 6 Functional framework diagram of a UEBA product provided by an embodiment of the present application;

[0046] Figure 7 Architecture diagram of an enterprise's security monitoring of business devices in a production environment provided by an embodiment of the present application;

[0047] Figure 8 Another schematic diagram of an application scenario of an abnormal access behavior detection method provided by an embodiment of the present application;

[0048] Figure 9 Architecture diagram of a cloud service provider's security monitoring of internal tenants and / or external tenants provided by an embodiment of the present application;

[0049] Figure 10 Architecture diagram of a community network described by a directed graph provided by an embodiment of the present application;

[0050] Figure 11 Flowchart of using a second model to identify abnormal access behavior in a directed graph provided by an embodiment of the present application;

[0051] Figure 12 Flowchart of using a third model to identify abnormal access behavior in a directed graph provided by an embodiment of the present application;

[0052] Figure 13 Schematic diagram of an abnormal access behavior detection method provided by an embodiment of the present application;

[0053] Figure 14 Directed graph constructed by an embodiment of the present application;

[0054] Figure 15Flowchart of an abnormal access behavior detection method provided by an embodiment of the present application;

[0055] Figure 16 System architecture diagram of an abnormal access behavior detection system provided by an embodiment of the present application;

[0056] Figure 17 Structural schematic diagram of an abnormal access behavior detection device provided by an embodiment of the present application;

[0057] Figure 18 Hardware structural schematic diagram of an abnormal access behavior detection device provided by an embodiment of the present application. Detailed implementation manners

[0058] For the convenience of clearly describing the technical solutions of the embodiments of the present application, in the embodiments of the present application, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and effects. For example, the first log and the second log are only used to distinguish network logs within different time windows, and do not limit their sequence. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity, and terms such as "first" and "second" do not necessarily mean different.

[0059] It should be noted that in the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific manner.

[0060] In the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B may be singular or plural.

[0061] Abnormal access behaviors in the field of security may include one or more of targeted attack behaviors, abnormal login behaviors, or unauthorized access behaviors.

[0062] Exemplarily, the targeted attack behavior may refer to an attack launched by an attacker against a specified target. For example, the attacker conducts port scanning, distributed denial of service attack (DDoS), structured query language (SQL) injection, cross site script attack (XSS), etc. on a machine exposed to the public network for vulnerability mining.

[0063] Exemplarily, the abnormal login behavior may refer to a user attempting to log in to a website application or server using a non - personal account. This abnormal login behavior may be caused by the leakage, theft, or successful brute - force cracking of the account password, and is used to bypass the identification and verification of the user's identity. The data that can be abnormally accessed (or may be called abnormal points) in the abnormal login behavior includes the Internet Protocol (IP) interconnected between the login networks, the login location, the login time, the number of login failures, the login operation, etc.

[0064] Exemplarily, the unauthorized access behavior may refer to an attacker initiating a connection request to a machine that they have no right to access or no need to access. The unauthorized access behavior may be a lateral penetration and / or vertical penetration initiated after an internal threat or a successful external network intrusion, and is used to increase the user's privilege in the server and expand the attack surface. In a possible understanding, lateral penetration may refer to an attacker attempting to access the resources of a user with the same privilege as theirs, and vertical penetration may refer to a low - level attacker attempting to access the resources of a high - level user.

[0065] The data that can be abnormally accessed (or may be called abnormal points) in the unauthorized access behavior includes the IP address, user information, service attributes, etc. In a possible understanding, user information may refer to the user's login account, login time, and number of logins, etc., and service attributes may refer to the characteristics shown under the support of the network or terminal capabilities and their hierarchical functions.

[0066] For the above - mentioned unauthorized access behavior, possible abnormal access behavior detection methods include: a baseline detection method based on statistics and profiling, and a machine - learning detection method based on pattern features.

[0067] Exemplarily, Figure 1 shows a general architecture schematic diagram of the abnormal access behavior detection method.

[0068] As Figure 1As shown, when an electronic device detects abnormal access behavior, it can perform data collection, data cleaning, etc. on the collected network logs and store the processed data. The electronic device uses an abnormal access behavior detection method to detect the stored data, passes the detection result to the user, and provides disposal suggestions. Among them, the abnormal access behavior detection method can include one of a baseline detection method based on statistics and profiling or a machine learning detection method based on pattern features.

[0069] In a possible implementation, the baseline detection method is implemented based on statistical methods and user profiling. Defense personnel establish an access baseline through historical access records, define normal access behavior, and determine access that appears for the first time, exceeds the threshold, or deviates from the profile as abnormal. In a possible understanding, user profiling can refer to a rule base established based on network information, terminal information, service information, operation information, etc. Among them, network information can include IP segments, IP geographical locations, port numbers, etc., terminal information can include device types, operating systems, etc., service information can include device-borne applications, subordinate business regions, etc., and operation information can include login methods, system instructions initiated during access, etc.

[0070] Exemplarily, a possible implementation of detecting abnormal access behavior based on user profiling is: The electronic device depicts the normal profile of each machine according to the user profile information and detects abnormal access behavior according to the normal profile. For example, access behavior that does not conform to the normal profile is determined as abnormal access behavior.

[0071] In a possible understanding, the normal profile can be a rule base established based on the network information, terminal information, service information, operation information, etc. of users who have access rights to the machine.

[0072] Exemplarily, Figure 2 shows a schematic diagram of a baseline detection method based on statistics and profiling.

[0073] As Figure 2 As shown, when the electronic device detects abnormal access behavior, it extracts parameters in the historical access logs that can reflect access patterns. Exemplarily, the extracted parameters can include request frequency, access duration, number of access IPs, etc. Among them, the form of the extracted parameters can be single variables, composite variables, statistical variables, etc. The electronic device performs outlier mining and statistical value calculation on the extracted data and conducts abnormal access behavior detection according to the manually set abnormal threshold.

[0074] However, when using a baseline to detect abnormal access behavior, since the rules of the baseline are usually artificially defined and not flexible enough, attackers will summarize the baseline rules and bypass the rules to perform abnormal access. Exemplarily, for the frequency threshold, attackers can perform abnormal access by reducing the access frequency. At the same time, when using a baseline to detect abnormal access behavior, there are many false positives. For example, when an internal user switches the IP or port due to business needs, false abnormal access behavior alerts may be generated.

[0075] Exemplarily, Figure 3 Fig. shows a schematic diagram of detecting abnormal access behavior based on machine learning.

[0076] As Figure 3 shown, detecting abnormal access behavior based on machine learning can include a training stage and an abnormal access detection stage.

[0077] In the training stage, a sample data set is constructed, and the differences between the two types of data, namely normal access samples and malicious samples captured historically, are found through a classification model or a clustering model. Exemplarily, when modeling, a feature set can be constructed from the time dimension and the space dimension. The time dimension includes time series features such as access frequency and interval time, and the space dimension can include interaction data such as network quintuples and traffic payloads. In a possible understanding, the network quintuple can include: source IP address, destination IP address, protocol number, source port, and destination port.

[0078] In the abnormal access detection stage, real data (such as network log data) is obtained, the real data is feature-extracted and then input into the detection model to determine whether there is abnormal access behavior, and the detection result is output.

[0079] In a possible implementation, the detection model can be a classification detection model based on a supervised algorithm. When the classification detection model based on a supervised algorithm detects abnormal access behavior, it is necessary to label the normal access samples and the malicious samples captured historically respectively, and find the differences between the two types of data through features. Common algorithms include support vector machine (SVM), k-nearest neighbor (KNN), linear regression, etc.

[0080] However, when using machine learning for detecting abnormal access behaviors, although it makes up for the limitations of the baseline detection method to a certain extent and no longer relies on manually set baselines to detect abnormal access behaviors, the machine learning detection method is applicable to the detection of abnormal access behaviors of a single device. During the detection process, a learning model is established for each machine separately, and it is impossible to globally identify abnormal access behaviors. At the same time, when using the machine learning method to detect abnormal access behaviors, the requirements for training data are too high. It is difficult to effectively establish a model when the types of samples are too few or the scenarios are too one-sided. At the same time, the selection of the feature set also has high requirements for the prior knowledge of the defense personnel.

[0081] Due to the above problems existing in the baseline detection method based on statistics and profiling and the machine learning detection method based on pattern features, the embodiments of the present application provide an abnormal access behavior detection method, which generates a directed graph based on the first log data and identifies abnormal access behaviors in the graph from the perspective of the directed graph. The embodiments of the present application do not require historical attack sample support, thereby reducing the requirements for the quality of the data source and achieving an improvement in the accuracy of detecting abnormal access behaviors.

[0082] The abnormal access behavior detection method of the embodiments of the present application can be applied to the security monitoring of the network behaviors of internal employees by an enterprise, and can also be applied to the security monitoring of the business devices in the production environment by an enterprise, or the security monitoring of internal tenants and / or external tenants by a cloud service provider, etc.

[0083] Exemplarily, Figure 4 A schematic structural diagram of an electronic device to which the abnormal access behavior detection method of the embodiments of the present application is applicable is as Figure 4 shown. The electronic device 401 may include: a processor 402, an external memory interface 403, an internal memory 404, a display screen 405, etc. It can be understood that the structure schematically shown in this embodiment does not constitute a specific limitation on the electronic device 401. In the embodiments of the present application, the electronic device 401 may include more or fewer components than shown, or combine certain components, or split certain components, or arrange different components.

[0084] The processor 402 may include one or more processing units. For example, the processor 102 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, a display process unit (DPU), and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors. In this embodiment, the electronic device 401 may also include one or more processors 402. Among them, the controller may be the nerve center and command center of the electronic device 401. A memory may also be provided in the processor 402 for storing instructions and data.

[0085] The external memory interface 403 may be used to connect an external memory card, such as a Micro SD card, to implement the storage capacity expansion of the electronic device 401. The external memory card communicates with the processor 402 through the external memory interface 403 to implement the data storage function. Exemplarily, the electronic device 401 may save data files such as IP addresses and access times in the external memory card.

[0086] The internal memory 404 may be used to store one or more computer programs, and the one or more computer programs may include instructions. The processor 402 may execute various functional applications and data processing, etc. of the electronic device 401 by running the above instructions stored in the internal memory 404. The internal memory 404 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system; the program storage area may also store one or more application programs, etc. The data storage area may store the data created during the use of the electronic device 401 (such as IP addresses and access times), etc.

[0087] The display screen 405 is used to display images, videos, etc. The display screen 405 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a quantum dot light-emitting diode (QLED), etc. In this embodiment, the electronic device 405 may include one or N display screens 405, where N is a positive integer greater than 1.

[0088] Exemplarily, Figure 5 A schematic diagram of an application scenario of the abnormal access behavior detection method provided in the embodiments of the present application.

[0089] In a possible implementation manner, the abnormal access behavior detection method in the embodiments of the present application can be used for an enterprise to perform security monitoring on the network behaviors of internal employees. Refer to Figure 5 , the application scenario includes an electronic device 501 in the internal monitoring center of the company and terminal devices 502 deployed at the internal user side (for example, including terminal devices 5021 to 502N, etc., where N is a natural number). The number of the electronic device 501 and the terminal devices 502 can be one or multiple. The structure of the electronic device applicable to the embodiments of the present application can be as Figure 4 shown, which will not be elaborated here. In a possible understanding manner, the electronic device 501 can be a terminal or a server, and the terminal device 502 can be any form of terminal of an internal user. For example, it can include a mobile phone, a computer, a tablet, etc. The embodiments of the present application do not make specific limitations on the terminal device.

[0090] The electronic device 501 can provide an internal memory, and the internal memory can include a storage program area and a storage data area. Among them, the storage program area can store one or more application programs, such as an application program used to detect abnormal access behaviors. The storage data area can store log data, etc. Exemplarily, the log data can reflect information such as the source IP, the destination IP, and the number of accesses. By acquiring network access log data and using the abnormal access behavior detection method in this embodiment for detection, and transmitting the detection result to the monitoring center, the monitoring of the internal network behaviors of the company is realized.

[0091] Based on this application scenario, the method of the embodiments of the present application can be used in user and entity behavior analytics (UEBA) products. Exemplarily, Figure 6 FIG. Figure 6 is a schematic diagram of a functional framework of a UEBA product applicable to the embodiments of the present application. Refer to Figure 6 , a UEBA product may include: a data reception and processing component, a data storage component, and an analysis component.

[0092] Exemplarily, the data reception and processing component obtains access logs between users and performs preprocessing, and stores the preprocessed log data in the data storage component as data for submission to the analysis component for detection. In the analysis component, the abnormal access behavior detection method of the embodiments of the present application is used to analyze user behavior and focus on internal threats within the enterprise.

[0093] In one possible understanding, an internal threat may be an abnormal access behavior of an internal user, which can be divided into two parts: abnormal access behavior within the enterprise and abnormal access behavior outside the enterprise. Among them, the abnormal access behavior within the enterprise may include deliberate data collection, abnormal and illegal access, account abuse, etc., and the abnormal access behavior outside the enterprise may include data leakage, continuous data transmission, abnormal website access, etc.

[0094] In a possible implementation, the analysis of user behavior by a UEBA product may include: operation behavior analysis, business process analysis, and user relationship analysis. Exemplarily, a possible implementation of the UEBA to analyze the operation behavior of a user is: the UEBA analyzes the web access logs of the user, and understands the user's web access behavior through a directed graph to prevent abnormal network access within the enterprise. Exemplarily, a possible implementation of the UEBA product to analyze the business process of a user is: the UEBA product extracts elements or buttons on the browser page and defines them as business tags. When the user triggers a corresponding event on the page, a corresponding business operation behavior log record will be generated, so as to understand the real business behavior of the employee and the behavior track of the business operation, and prevent the leakage of internal business data of the company. Exemplarily, a possible implementation of the UEBA product to analyze the user relationship is: the UEBA product collects the access logs of the user, and potential relationships and work intersections between users can be found through the directed graph corresponding to the access logs.

[0095] In a possible implementation, the abnormal access behavior detection method of the embodiments of the present application can be used for an enterprise to perform security monitoring on business equipment in a production environment. Refer to Figure 5, the application scenario includes the electronic device 501 and the terminal device 502 in the enterprise internal monitoring center. In this application scenario, the number of terminal devices 502 can be one or more (for example, including terminal devices 5021 to 502N, etc., where N is a natural number). In a possible understanding, the terminal device can be a machine for the enterprise to provide services externally.

[0096] Exemplarily, Figure 7 This is an architecture diagram applicable to the enterprise's security monitoring of production environment business devices in the embodiments of the present application. Refer to Figure 7 , the enterprise's security monitoring of production environment business devices may include: an information source component, an analysis engine component, and a response component. In combination with Figure 5 , the information source component, the analysis engine component, and the response component can be set in the electronic device 501.

[0097] In a possible understanding, the information source component may be responsible for collecting log data. Exemplarily, the log data may include information such as source IP, destination IP, and access times. The information source component can provide the obtained log data to the analysis engine component. In the analysis engine component, the abnormal access behavior detection method of the embodiments of the present application is used to determine whether there is an abnormal access behavior, and the detection result is output to the response component to implement the enterprise's security monitoring of production environment business devices.

[0098] Based on this application scenario, the abnormal access behavior detection method provided by the embodiments of the present application can be applied to an intrusion detection system, and batch deployed on devices to be detected to detect intrusion events.

[0099] Exemplarily, Figure 8 Another application scenario schematic diagram of the abnormal access behavior detection method provided by the embodiments of the present application.

[0100] In a possible implementation, the abnormal access behavior detection method of the embodiments of the present application can be used by a cloud service provider to perform security monitoring on internal tenants and / or external tenants. Refer to Figure 8 , this application scenario includes a cloud service provider 801, internal tenants 802 (for example, including internal tenants 8021 to 802N, etc., where N is a natural number) and external tenants 803 (for example, including external tenants 8031 to 803N, etc., where N is a natural number). Among them, the number of internal tenants 802 and external tenants 803 can be one or more. In a possible understanding, internal tenants are users for internal business deployment in the company, and external tenants are users who lease external company's machines for use.

[0101] Exemplarily, Figure 9 This is a framework schematic diagram for a cloud service provider to perform security monitoring on internal tenants and / or external tenants.

[0102] See also Figure 9 , the cloud service provider 801 may include: a log collection probe, an analysis platform, and an association rule engine.

[0103] Among them, the log collection probe can collect the network logs of the internal tenant 802 and / or the external tenant 803, and pre-process the collected network logs to obtain the first log data. In a possible understanding, the first log data can reflect the source IP, destination IP, number of visits and other information of the access behavior of the internal tenant 802 and / or the external tenant 803. The analysis platform can store the first log data submitted by the log probe, and then use the abnormal access behavior detection method of the embodiment of the present application to detect whether there is abnormal access behavior. The association rule engine can match the association rules to the detection results and generate associated alarms for abnormal access behaviors. Finally, the alarm information is passed to the cloud service provider 801 to achieve security monitoring of the internal tenant 802 and / or the external tenant 803.

[0104] Based on this application scenario, the method of the embodiment of the present application can be applied to cloud server (elastic compute service, ECS) security monitoring software to ensure service security while supervising whether tenants have any illegal operations. It can also be applied to situational awareness platforms, using probes to collect terminal data, and the security operations center (security operations center, SOC) and other functional departments can uniformly monitor and analyze it.

[0105] Some words in the embodiments of the present application are explained below. Among them, the words explained in the embodiments of the present application are for the convenience of understanding of those skilled in the art and do not constitute a limitation on the embodiments of the present application.

[0106] The first log data described in the embodiment of the present application may include information such as source IP, destination IP, number of accesses and access time.

[0107] The directed graph described in the embodiment of the present application includes multiple nodes for identifying devices and directed access relationships between multiple nodes. Exemplarily, the directed access relationship between multiple nodes can be understood as one or more of the maximum number of visits per hour, the total number of historical visits, or the latest visit time between multiple nodes.

[0108] The node described in the embodiment of the present application may be an Internet Protocol IP address of a device.

[0109] The predecessor node described in the embodiments of the present application is a node in a directed graph that directly points to the current target. Exemplarily, if there are access behaviors with destination IP A and destination IP B in the access log, and these access behaviors are represented by a directed line segment with two points and one line, that is, A points to B, then the predecessor node of B in the directed graph is A.

[0110] The first model described in the embodiments of the present application can also be referred to as a path model, which is used to identify abnormal access behaviors of multi-node jump logins according to a directed graph. For example, by using a graph traversal algorithm to calculate the maximum access path, abnormal behaviors of continuously accessing multiple devices in the form of a link are identified. The graph traversal algorithm can be a depth first search (DFS) algorithm, a breadth first search (BFS) algorithm, etc.

[0111] In a possible implementation, the embodiments of the present application load the directed graph into the first model, and in the first model, with the help of a heap data structure, a corresponding topological sorting table is generated starting from the source node in the directed graph using the DFS algorithm. The embodiments of the present application use this topological sorting table to calculate the maximum unweighted path length from the source IP to the destination IP, or it can be called the maximum number of hops. In a possible understanding, this maximum number of hops can be used to represent the number of hops for continuous access with the destination node as a springboard. Based on the assumption that it rarely occurs to use an access device as a springboard for continuous access in a normal scenario, access behaviors with a maximum number of hops greater than the first threshold are identified as abnormal access behaviors.

[0112] The first threshold described in the embodiments of the present application can be the threshold used when detecting abnormal access behaviors using the first model, and this threshold can be adjusted according to the node scale.

[0113] The second model described in the embodiments of the present application can also be referred to as a community model, which is used to identify abnormal access behaviors across communities according to a directed graph. For example, by using a community discovery algorithm to cluster nodes, abnormal behaviors of cross-community access are identified. The community discovery algorithm can be a Louvain algorithm, a label propagation algorithm (LPA) algorithm, etc.

[0114] Exemplarily, Figure 10 is a network composed of multiple business groups described by a directed graph. In a possible understanding, a business group can also be a community. As Figure 10 shown, a community can be a sub-directed graph containing vertices and edges. The community discovery algorithm can be used to discover the community structure in this network. For example, the connections between nodes within the same community are very tight, while the connections between different communities are relatively sparse. By discovering the community structure in the directed graph, abnormal access behaviors of cross-business group access are identified.

[0115] In a possible implementation, Figure 11 The flowchart for identifying abnormal access behaviors in a directed graph using a second model includes the following steps:

[0116] S1101: Load the directed graph into the second model.

[0117] S1102: In the second model, use the Louvain algorithm based on modularity to aggregate the nodes in the directed graph.

[0118] Exemplarily, the aggregation of nodes in the directed graph in the embodiments of the present application may include two stages.

[0119] The first stage: According to the modularity gain, classify the nodes into the community where the neighbor node with the largest gain is located.

[0120] Among them, the relative modularity gain where k i,in represents the sum of the weights of the incident clusters C by node i, ∑tot represents the total weight of the incident cluster C, k i represents the weighted degree sum of node i, and m represents the total weight of all edges in the directed graph.

[0121] The second stage: The second model compresses the directed graph and compresses the nodes in the same community into the first node.

[0122] S1103: Iterate the two stages in S1102 until the graph modularity Q no longer changes.

[0123] Among them, A ij represents the weight of the edge between node i and node j, c i represents the community number of node i, σ(c i , c j ) function returns 1 if node i and node j belong to the same community, otherwise returns 0.

[0124] S1104: Based on the assumption that devices that mutually access in a normal scenario usually belong to the same community, identify the access behavior corresponding to the first node with cross-community access as an abnormal access behavior.

[0125] The third model described in the embodiments of the present application can also be referred to as an embedding model, which is used to identify abnormal access behaviors that do not conform to historical access behaviors based on the distances between nodes in a directed graph. For example, the embedding model can adopt a graph embedding algorithm, learn the embedding vectors of each node through a specific random walk strategy, and finally discover abnormal behaviors that do not conform to historical access behaviors based on the node distances. Among them, the embedding model can learn the low-dimensional latent representations of the nodes in the directed graph through the graph embedding algorithm, and the learned feature representations can be used as features for various tasks based on the directed graph, such as classification, clustering, link prediction, and visualization. The graph embedding algorithm can be the node2vec algorithm, the DeepWalk algorithm, etc.

[0126] In a possible implementation, Figure 12 The flowchart for using the third model to identify abnormal access behaviors in a directed graph includes the following steps:

[0127] S1201: Load the directed graph into the third model.

[0128] S1202: In the third model, use the graph embedding algorithm to convert the nodes in the directed graph into embedding vectors.

[0129] Exemplarily, for node i, a corresponding adjacent node sequence is generated through a random walk strategy based on the second-order transition probability, and these sequences are regarded as texts and sent into the word2vec model to obtain the corresponding vectors where q is the sequence length. The transition probability between two nodes can be expressed as, π vx = α pq (t,x)·w vx , where, v is the current node, x is the next node, t is the previous node, w is the weight of the edge between the two nodes, and α pq is defined as p and q are hyperparameters used to control the random walk strategy, and d tx is the shortest path distance between node t and node x, where, d tx = 0 indicates that node x coincides with node t, that is, the shortest path distance is 0, and d tx = 1 indicates that node x is connected to node t, that is, the shortest path distance is 1, and d tx = 2 indicates that node x is not connected to node t, that is, the distance is greater than 1.

[0130] S1203: In the third model, normalize the embedding vector matrix corresponding to the set of predecessor nodes of the destination node to obtain a set of normalized unit vectors.

[0131] Optionally, in the third model, for the vector matrix Matrix corresponding to the set of predecessor nodes of the destination node in the directed graph pred = [r (1), r (2) , …, r (n) is normalized, where n is the in-degree of the destination IP. The in-degree of the destination IP can be the number of directed line segments starting from the destination IP and terminating at the destination IP, and r (i) is the embedding vector corresponding to the predecessor node i. Finally, calculate the vector r corresponding to the source IP (src) and the cosine similarity of the normalized unit vector set.

[0132] Exemplarily, when the number of devices to be monitored is more than a thousand, node scale and time cost need to be considered when performing embedding operations on the directed graph and calculating similarity. Therefore, the applicable scenario of this method is anomaly access detection under a large-scale device cluster.

[0133] Optionally, in the third model, the one class support vector machine (OneClassSVM) algorithm can also be used to train the vector matrix Matrix corresponding to the set of predecessor nodes of the destination nodes in the directed graph pred = [r (1) , r (2) , …, r (n) , and use the trained model to calculate the vector r corresponding to the source IP (src) for normalization. Exemplarily, this method can be applied to anomaly access behavior detection in a small range. In this scenario, since there are fewer detection devices, when the third model performs embedding operations on the directed graph and calculates similarity, the requirements for node scale and time cost are relatively low. Therefore, when using the node2vec graph embedding algorithm, the learning parameters can be adjusted according to the actual interconnection situation of the nodes. At the same time, when calculating similarity, an anomaly detection algorithm that can identify outliers more accurately can be used.

[0134] S1204: Calculate the cosine similarity between the embedding vector corresponding to the node and the normalized unit vector set.

[0135] S1205: Based on the assumption that source IPs accessing the same device under normal scenarios usually have similar behaviors, identify the access behavior with a cosine similarity less than the second threshold as an anomaly access behavior.

[0136] The second threshold described in the embodiments of the present application can be the threshold used when detecting anomaly access behavior using the third model. Similarly, this threshold can also be adjusted according to the node scale.

[0137] The association analysis described in the embodiments of the present application can be to construct a linear function between multiple recognition results of the first model, the second model, or the third model. When the value of the linear function meets the preset anomaly condition, determine the access behavior corresponding to the linear function as the anomaly detection result.

[0138] Exemplarily, in the output results of the first model, the second model, and the third model, if there is an abnormal access behavior, the outputs of the first model, the second model, and the third model are 1; otherwise, they are 0. In the embodiments of the present application, a linear summation function between the three models is constructed. If the value of the linear summation function is greater than or equal to 1, there is an abnormal access behavior in the directed graph. Exemplarily, if the output of the first model is 1, the abnormal access behavior may be an abnormal access behavior of multi-node jump login; if the output of the second model is 1, the abnormal access behavior may be an abnormal access behavior of cross-business group access; if the output of the third model is 1, the abnormal access behavior may be an abnormal access behavior that does not conform to historical behavior access.

[0139] Figure 13 The present application provides an abnormal access behavior detection method for embodiments, including the following steps:

[0140] S1301: The electronic device generates a directed graph according to the first log data.

[0141] In a possible implementation manner, generating a directed graph according to the first log data may include two stages: data preprocessing and directed graph construction.

[0142] Exemplarily, in the data preprocessing stage, the electronic device collects the network access logs of all devices within a time window. The size of the time window can be set as the most recent N days (N can be any value greater than 2, such as 30, etc.). The electronic device may be a terminal device deployed at the user end, such as a mobile phone, a computer, a tablet, etc. The embodiments of the present application do not make specific limitations on the electronic device.

[0143] The electronic device may preprocess the obtained network access logs. For example, the preprocessing may include removing irrelevant information from the network access logs, aggregating the same access records, etc. The data obtained after the network access logs are preprocessed is the first log data. Exemplarily, a possible implementation of removing irrelevant information from the network access logs is: in the network access logs, remove other information except for the required source IP, destination IP, access times, etc. For example, other information may include network environment information, etc. Exemplarily, a possible implementation of removing irrelevant information from the network access logs is: in the network access logs, remove the access records of whitelist users. In a possible understanding manner, the whitelist users may be internal users of an enterprise. Exemplarily, a possible implementation of aggregating the same access records is: in the network access logs, aggregate the same access records into one access record to reduce the model calculation amount.

[0144] Exemplarily, in the directed graph construction stage, in the embodiments of the present application, an access behavior is represented by a directed line segment composed of points and lines to construct a directed graph. For example, asFigure 14 As shown, the directed graph includes multiple nodes for identifying devices and directed access relationships between multiple nodes. Among them, an IP address can correspond to a point in the directed graph. Exemplarily, the source IP can correspond to the starting point of the directed line segment, and the destination IP can correspond to the end point of the directed graph. The weight of the edge in the directed graph is the maximum number of accesses per hour.

[0145] S1302: Identify abnormal access behaviors in the directed graph using one or more of the first model, the second model, or the third model, and determine the abnormal detection result.

[0146] In a possible implementation, the embodiment of the present application uses the first model to identify abnormal access behaviors in the directed graph. As Figure 14 shown, in addition to single accesses between two devices in the directed graph, there are also continuous accesses in the form of a link. In one possible understanding, the link-type access uses the destination IP of one access behavior as the source IP to continue accessing another device. By repeating this operation, continuous penetration access to multiple devices in the form of a link is achieved. In normal network access, abnormal access behaviors such as using an accessed device as a springboard for continuous access rarely occur. Therefore, in the embodiment of the present application, by using the first model, starting from the source IP, the DFS algorithm is used to traverse the nodes in the directed graph, and the maximum unweighted path length from the source IP to the destination IP, which can also be called the maximum number of hops, is calculated according to the topological sorting table obtained after traversal. By comparing the relationship between the maximum number of hops and the first threshold, it is detected whether there are abnormal access behaviors in the directed graph. If the first model detects that there are abnormal access behaviors in the directed graph, the abnormal access behavior can be the behavior of multi-node jump login.

[0147] In a possible implementation, the embodiment of the present application uses the second model to identify abnormal access behaviors in the directed graph. Based on the assumption that devices that access each other under normal scenarios usually belong to the same business group, as Figure 14 shown, the subgraph corresponding to the subset of nodes with close connections in the directed graph can be called a business group, or also called a community. Among them, the connections between nodes within the community are relatively close, but the connections between each community are relatively sparse. In the second model, by using the Louvain algorithm for node aggregation, the nodes in the same community are compressed into a new node, which can represent a community. By judging whether there are access behaviors between the new nodes after aggregation, it is detected whether there are abnormal access behaviors in the directed graph. If the second model detects that there are abnormal access behaviors in the directed graph, the abnormal access behavior can be the abnormal access behavior of cross-business group access.

[0148] In a possible implementation, the embodiments of the present application use a third model to identify abnormal access behaviors in a directed graph. Based on the assumption that source IPs accessing the same device under normal scenarios usually have similar access behaviors, in the third model, the node2vec algorithm is used to learn the embedding vectors of each node through a specific random walk strategy, and abnormal access behaviors inconsistent with historical access behaviors are discovered based on node distances. If abnormal access behaviors are detected in the directed graph, the abnormal access behaviors can be behaviors inconsistent with historical access behaviors.

[0149] The above embodiments introduce a method for detecting a single abnormal access behavior in a directed graph. At the same time, graph algorithms such as DFS, Louvain, and node2vec used by the first model, the second model, and the third model respectively are all unsupervised structures, and do not require prior definition of attack behaviors, sample support, or training, and can be directly applied to the detection of abnormal access behaviors.

[0150] In different network access scenarios, in addition to single abnormal access behaviors, multiple abnormal access behaviors can also be covered. For example Figure 14 As shown, a directed graph can include directed access relationships between multiple nodes. In different access scenarios, the generated directed graphs can also be different. To comprehensively detect abnormal access behaviors in a directed graph, in addition to the method for detecting a single abnormal access behavior, the embodiments of the present application provide multiple methods for detecting abnormal access behaviors, covering multiple abnormal access scenarios.

[0151] In a possible implementation, the embodiments of the present application use the first model and the second model to identify abnormal access behaviors in a directed graph. By performing correlation analysis on the detection results of the first model and the second model, the embodiments of the present application can include the following several results.

[0152] The first result includes: If the first model detects abnormal access behaviors in the directed graph, while the second model detects no abnormal access behaviors in the directed graph, then by performing correlation analysis on the detection results of the first model and the second model, there may be abnormal access behaviors of multi-node jump login in the directed graph.

[0153] The second result includes: If the first model detects no abnormal access behaviors in the directed graph, while the second model detects abnormal access behaviors in the directed graph, then by performing correlation analysis on the detection results of the first model and the second model, there may be abnormal access behaviors of cross-business group access in the directed graph.

[0154] The third result includes: If the first model detects abnormal access behaviors in the directed graph, while the second model detects abnormal access behaviors in the directed graph, then by performing correlation analysis on the detection results of the first model and the second model, there are abnormal access behaviors of multi-node jump login and cross-business group access in the directed graph.

[0155] The fourth result includes: If the first model detects no abnormal access behavior in the directed graph and the second model also detects no abnormal access behavior in the directed graph, then the detection results of the first model and the second model are analyzed associatively, and there is no abnormal access behavior in the directed graph.

[0156] In a possible implementation, the embodiments of the present application use the first model and the third model to identify abnormal access behavior in a directed graph, and the detection results of the first model and the third model are analyzed associatively. The embodiments of the present application may include the following results.

[0157] The first result includes: If the first model detects abnormal access behavior in the directed graph and the third model detects no abnormal access behavior in the directed graph, then the detection results of the first model and the second model are analyzed associatively, and there may be abnormal access behavior of multi-node jump login in the directed graph.

[0158] The second result includes: If the first model detects no abnormal access behavior in the directed graph and the third model detects abnormal access behavior in the directed graph, then the detection results of the first model and the third model are analyzed associatively, and there may be abnormal access behavior inconsistent with historical access behavior in the directed graph.

[0159] The third result includes: If the first model detects abnormal access behavior in the directed graph and the third model also detects abnormal access behavior in the directed graph, then the detection results of the first model and the third model are analyzed associatively, and there are abnormal access behaviors of multi-node jump login and inconsistent with historical access behavior in the directed graph.

[0160] The fourth result includes: If the first model detects no abnormal access behavior in the directed graph and the third model detects no abnormal access behavior in the directed graph, then the detection results of the first model and the third model are analyzed associatively, and there is no abnormal access behavior in the directed graph.

[0161] In a possible implementation, the embodiments of the present application use the second model and the third model to identify abnormal access behavior in a directed graph. The detection results of the first model and the second model are analyzed associatively. The embodiments of the present application may include the following results.

[0162] The first result includes: If the second model detects abnormal access behavior in the directed graph and the third model detects no abnormal access behavior in the directed graph, then the detection results of the second model and the third model are analyzed associatively, and there may be abnormal access behavior of cross-business group access in the directed graph.

[0163] The second result includes: If the second model detects no abnormal access behavior in the directed graph and the third model detects abnormal access behavior in the directed graph, then the detection results of the second model and the third model are analyzed associatively, and there may be abnormal access behavior inconsistent with historical access behavior in the directed graph.

[0164] The third result includes: If the second model detects abnormal access behavior in the directed graph and the third model also detects abnormal access behavior in the directed graph, then the detection results of the second model and the third model are analyzed associatively. There are cross-business-group access and abnormal access behavior that does not conform to historical behavior in the directed graph.

[0165] The fourth result includes: If the second model does not detect abnormal access behavior in the directed graph and the third model also does not detect abnormal access behavior in the directed graph, then the detection results of the second model and the third model are analyzed associatively. There is no abnormal access behavior in the directed graph.

[0166] In a possible implementation, the embodiments of the present application use the first model, the second model, and the third model to identify abnormal access behavior in a directed graph. By analyzing the detection results of the first model, the second model, and the third model associatively, the embodiments of the present application may include the following results.

[0167] The first result includes: If the first model detects abnormal access behavior in the directed graph, the second model does not detect abnormal access behavior in the directed graph, and the third model also does not detect abnormal access behavior in the directed graph, then the detection results of the first model, the second model, and the third model are analyzed associatively. There may be abnormal access behavior of multi-node jump login in the directed graph.

[0168] The second result includes: If the first model does not detect abnormal access behavior in the directed graph, the second model detects abnormal access behavior in the directed graph, and the third model does not detect abnormal access behavior in the directed graph, then the detection results of the first model, the second model, and the third model are analyzed associatively. There may be abnormal access behavior of cross-business-group access in the directed graph.

[0169] The third result includes: If the first model does not detect abnormal access behavior in the directed graph, the second model does not detect abnormal access behavior in the directed graph, and the third model detects abnormal access behavior in the directed graph, then the detection results of the first model, the second model, and the third model are analyzed associatively. There may be abnormal access behavior that does not conform to historical access behavior in the directed graph.

[0170] The fourth result includes: If the first model detects abnormal access behavior in the directed graph, the second model detects abnormal access behavior in the directed graph, and the third model does not detect abnormal access behavior in the directed graph, then the detection results of the first model, the second model, and the third model are analyzed associatively. There may be abnormal access behavior of multi-node jump login and cross-business-group access in the directed graph.

[0171] The fifth result includes: If the first model detects abnormal access behavior in the directed graph, the second model detects no abnormal access behavior in the directed graph, and at the same time the third model detects abnormal access behavior in the directed graph, then the detection results of the first model, the second model, and the third model are analyzed associatively. There can be multi-node jump logins and abnormal access behaviors that do not conform to historical access behaviors in the directed graph.

[0172] The sixth result includes: If the first model detects no abnormal access behavior in the directed graph, the second model detects abnormal access behavior in the directed graph, and at the same time the third model detects abnormal access behavior in the directed graph, then the detection results of the first model, the second model, and the third model are analyzed associatively. There can be cross-business group accesses and abnormal access behaviors that do not conform to historical access behaviors in the directed graph.

[0173] The seventh result includes: If the first model detects abnormal access behavior in the directed graph, the second model detects abnormal access behavior in the directed graph, and at the same time the third model detects abnormal access behavior in the directed graph, then the detection results of the first model, the second model, and the third model are analyzed associatively. There can be three types of abnormal access behaviors in the directed graph: multi-node jump logins, cross-business group accesses, and abnormal access behaviors that do not conform to historical access behaviors.

[0174] The eighth result includes: If the first model detects no abnormal access behavior in the directed graph, the second model detects no abnormal access behavior in the directed graph, and at the same time the third model detects no abnormal access behavior in the directed graph, then the detection results of the first model, the second model, and the third model are analyzed associatively. There is no abnormal access behavior in the directed graph.

[0175] The embodiments of the present application provide an abnormal access behavior detection method, which can be independent of historical samples, generate a directed graph based on the first log, and identify abnormal access behaviors in the directed graph from a graph perspective. At the same time, the embodiments of the present application do not require historical attack sample support, thereby reducing the requirements for the quality of data sources and achieving an improvement in the accuracy of abnormal access behavior detection.

[0176] Figure 15 An abnormal access behavior detection method provided by the embodiments of the present application can be executed by a model initialization unit, a real-time detection unit, and a model update unit.

[0177] Exemplarily, the model initialization unit is responsible for training the first model, the second model, and the third model respectively with the directed graph corresponding to the obtained second log as samples to obtain the model parameters of the first model, the second model, and the third model respectively. S1501 to S1503 in the embodiments of the present application can be executed by the model initialization unit.

[0178] Exemplarily, the real-time detection unit performs abnormal access behavior detection based on the directed graph corresponding to the first log obtained in real time, using one or more of the first model, the second model, or the third model, and performs correlation analysis on one or more detection results obtained by the first model, the second model, or the third model, covering multiple attack scenarios, and realizing the global identification of abnormal access behavior. In the embodiments of the present application, S1504 to S1515 may be executed by the model detection unit.

[0179] Exemplarily, the model update unit periodically collects the latest access logs to update the graph model, and optimizes the model parameters according to the detection and review results to maintain the consistency between the model and the directed graph state. In the embodiments of the present application, S1516 to S1520 may be executed by the model update unit.

[0180] As Figure 15 shown, the abnormal access behavior detection method may include the following steps:

[0181] S1501: The electronic device collects the access logs within the time window.

[0182] Among them, the size of the time window may be set to the last N days (N may be any value greater than 2, such as 30, etc.).

[0183] S1502: Preprocess the access logs to generate second log data.

[0184] Exemplarily, in the embodiments of the present application, the preprocessing may include extracting the source IP and the destination IP, removing the irrelevant information of the access logs, aggregating the same access records, etc.

[0185] S1501 and S1502 in the embodiments of the present application may correspond to Figure 13 the description of S1301 therein, which will not be elaborated here.

[0186] Among them, the preprocessing of the access logs in the embodiments of the present application can play a role in reducing the model calculation amount. In a possible implementation manner, the electronic device may also not preprocess the access logs collected in S1501, but use the access logs collected in S1501 as the second log data, that is, S1502 may be an optional step, and the embodiments of the present application do not make specific limitations on this.

[0187] S1503: According to the second log data, construct a directed graph corresponding to the second log data, and store the directed graph file locally.

[0188] Exemplarily, in the embodiments of the present application, a single access behavior in the second log data is represented by a directed line segment composed of points and lines to construct a directed graph. The directed graph includes multiple nodes for identifying devices and directed access relationships between multiple nodes. Among them, an IP address can correspond to a point in the directed graph. Exemplarily, the source IP can correspond to the starting point of the directed line segment, and the destination IP can correspond to the end point of the directed graph. The weight of the edge in the directed graph is the maximum number of accesses per hour. Then, the generated directed graph file is stored locally.

[0189] A possible implementation of storing the generated directed graph file locally is as follows: In the second log data, there may be an access record of a node with an IP address of ip x to a node with an IP address of ip y , where the highest access frequency per hour is freq xy . Then, the directed graph file generated from this access record can be stored locally in JSON format. Exemplarily, the storage form of this directed graph file is {"nodes": [{"id": ip x}, {"id": ip y}], "links": [{"source": ip x , "target": ip y , "frequency": freq xy}]。

[0190] In the embodiments of the present application, the directed graphs corresponding to the obtained second logs are respectively used as training samples for the first model, the second model, and the third model to update the model parameters related to the first model, the second model, and the third model respectively, and then the updated model parameters are stored locally.

[0191] In a possible understanding manner, the model files stored locally may include: the directed graphs corresponding to the second logs and the model parameters related to the first model, the second model, and the third model respectively.

[0192] In a possible implementation manner, if the first model, the second model, and the third model are running for the first time after deployment and the local model parameter configuration file is empty at this time, then the preset hard-coded parameters are used.

[0193] S1504: The electronic device periodically obtains access logs within a time threshold.

[0194] Among them, the size of the time threshold can be set to N hours (N can be any value greater than 0, such as 1, etc.).

[0195] In a possible implementation, the time period for the electronic device to obtain access logs can be the same as the time threshold. Exemplarily, if the time threshold is set to 1 hour, the time period for the electronic device to obtain access logs is also 1 hour, that is, the electronic device obtains the latest 1-hour access logs every 1 hour.

[0196] S1505: The electronic device preprocesses the access logs obtained within the time threshold to obtain first log data.

[0197] Exemplarily, the preprocessing may include extracting the source IP and destination IP, removing irrelevant information from the access logs, aggregating the same access records, etc.

[0198] In the embodiment of the present application, S1505 may correspond to Figure 13 the description of S1301 in, which will not be elaborated here.

[0199] Among them, the generation time of the second log data obtained by S1502 is earlier than the generation time of this first log data.

[0200] In a possible understanding, the second log data is the access logs obtained within the time window. In the embodiment of the present application, the directed graph corresponding to the second log is used as a sample for training the model to obtain the model parameters of the first model, the second model, and the third model respectively. And the first log data is the data for real-time detection. Before detecting abnormal access behaviors, the first model, the second model, and the third model have completed model initialization. Therefore, the generation time of the second log data is earlier than the generation time of this first log data.

[0201] S1506: The electronic device generates a directed graph corresponding to the first log according to the first log.

[0202] In the embodiment of the present application, S1506 may correspond to the description of constructing a directed graph in the above S1503, which will not be elaborated here.

[0203] S1507: The electronic device collects the access logs that have appeared in the historical records to obtain third log data.

[0204] Among them, the generation time of the third log data is earlier than the generation time of the first log, and the difference between the generation time of the third log data and the generation time of the first log is less than the time threshold.

[0205] In a possible understanding, the access behaviors in the third log data have been detected for abnormal access behaviors recently, and there are no abnormal access behaviors. Then, it is not necessary to repeat the abnormal access detection for this part of the data, reducing the calculation amount.

[0206] S1508: The electronic device generates a corresponding third directed graph according to the third log.

[0207] S1508 in the embodiments of the present application can correspond to the description of constructing a directed graph in S1503 above, which will not be elaborated here.

[0208] S1509: The electronic device filters the overlapping part between the directed graph corresponding to the first log and the directed graph corresponding to the third log in the directed graph corresponding to the first log, to obtain a directed graph for detection.

[0209] In a possible implementation, the electronic device obtains the access log of the latest 1 hour (for example, from 11:00 am to 12:00 pm) as the first log, and the historical record stored locally, that is, the third log, can be the access log earlier than the latest 1 hour (for example, from 10:00 am to 11:00 am). Before detecting abnormal access behaviors in the access log of the latest 1 hour, the electronic device can filter the overlapping part between the directed graph corresponding to the first log and the directed graph corresponding to the third log, and screen out the access behaviors that occurred during the period earlier than the latest 1 hour in the latest 1 hour, to obtain a filtered directed graph.

[0210] Among them, filtering the access behaviors that appear in the directed graph corresponding to the third log in the directed graph corresponding to the first log in the embodiments of the present application can reduce the model load and achieve the effect of reducing the model calculation amount. In a possible implementation, the electronic device can also directly perform abnormal access behavior detection without filtering the directed graph generated in S1506, that is, S1507 to S1509 can be optional steps, and the embodiments of the present application do not make specific limitations on this.

[0211] S1510: In the electronic device, load the directed graph and the model parameters stored locally.

[0212] Optionally, the directed graph loaded in the electronic device can be the directed graph corresponding to the first log.

[0213] Optionally, the directed graph loaded in the electronic device can be the directed graph obtained by filtering the overlapping part between the directed graph corresponding to the first log and the directed graph corresponding to the third log.

[0214] In a possible understanding, the model parameters can respectively refer to the model parameters related to the first model, the second model, and the third model. The model parameters stored locally loaded in the embodiments of the present application are the model parameters obtained in S1503, and the specific implementation can correspond to the description of S1503, which will not be elaborated here.

[0215] S1511: Use the first model to identify abnormal access behaviors in the directed graph, and determine the abnormal access behavior detection result.

[0216] Specifically, reference can be made to the description of the first model detecting abnormal access behaviors in the glossary part, which will not be elaborated here.

[0217] S1512: Identify abnormal access behaviors in the directed graph using the second model, and determine the detection results of abnormal access behaviors.

[0218] Specifically, reference can be made to Figure 11 the description of the second model for detecting abnormal behaviors, which will not be elaborated here.

[0219] S1513: Identify abnormal access behaviors in the directed graph using the third model, and determine the detection results of abnormal access behaviors.

[0220] Specifically, reference can be made to Figure 12 the description of the third model for detecting abnormal behaviors, which will not be elaborated here.

[0221] S1514: Perform correlation analysis on the detection results of abnormal access behaviors determined by one or more of the above first model, second model, or third model to determine the detection results of abnormal access behaviors.

[0222] In a possible implementation, when the electronic device uses one of the first model, second model, or third model to detect abnormal access behaviors, the electronic device may also not perform correlation analysis on the detection results of abnormal access detected by one of the first model, second model, or third model. That is, S1514 can be an optional step, and the embodiments of the present application do not make specific limitations on this.

[0223] S1515: The electronic device sends an alarm message to the target object according to the detection results of abnormal access behaviors.

[0224] Among them, the alarm message may include one or more of the log information corresponding to the detection results of abnormal access behaviors, the cause of the alarm, or the recommended handling method.

[0225] Exemplarily, a possible implementation where the alarm message is the log information corresponding to the detection results of abnormal access behaviors is as follows: If it is determined that there are abnormal access behaviors in the directed graph, then by searching the directed graph file stored locally, log information such as the source IP and destination IP corresponding to the abnormal access behaviors can be found.

[0226] Exemplarily, a possible implementation of the alarm information for the cause of the alarm is as follows: Output files corresponding to the causes of the alarm for the models can be preset in the first model, the second model, and the third model. For example, when the first model detects an abnormal access behavior, the output cause of the alarm can be an abnormal access behavior of multi-node jumping login; when the second model detects an abnormal access behavior, the output cause of the alarm can be an abnormal access behavior of cross-community access; when the third model detects an abnormal access behavior, the output cause of the alarm can be an abnormal access behavior that does not conform to the historical access behavior. In a possible implementation manner, when multiple of the first model, the second model, or the third model are used to detect abnormal access behaviors, the output cause of the alarm is a combination of the causes of the alarm corresponding to multiple of the first model, the second model, or the third model used. For example, when the first model and the second model are used to detect abnormal access behaviors, if the first model detects an abnormal access behavior and at the same time the second model detects an abnormal access behavior, the output cause of the alarm is an abnormal access behavior of multi-node jumping login and an abnormal access behavior of cross-business-group access.

[0227] Exemplarily, a possible implementation of the alarm information for the recommended handling method is as follows: According to the detection result of the abnormal access behavior, the user is notified to check whether there is an abnormal access behavior.

[0228] S1516: The electronic device collects incremental logs and detection results.

[0229] Exemplarily, the electronic device uses the collected detection results for auditing, and uses the audited detection results and the collected incremental logs to update the model.

[0230] In a possible understanding manner, the incremental log may refer to the first log data obtained within the latest time threshold.

[0231] S1517: The auditing module audits the detected abnormal access behaviors.

[0232] In a possible understanding manner, the auditing module may be composed of machines and / or humans.

[0233] In a possible implementation, if the audit module confirms that the detection result of the abnormal access behavior is correct, risk control is performed according to the alarm information corresponding to the abnormal access behavior and the loop is closed in a timely manner. In a possible understanding, risk control may include that the risk controller reduces the losses caused by the abnormal access behavior or reduces the possibility of the abnormal access behavior occurring. Exemplarily, when there is an abnormal access behavior of multi-node jumping login within an enterprise, the log information of the corresponding abnormal access behavior can be found through a directed graph, the source IP and the destination IP are determined, and at the same time, the abnormal access behavior is analyzed to check whether there are malicious data packets entering the enterprise internal network, so as to avoid the impact on the normal operation of the internal server.

[0234] In a possible implementation, if the audit module confirms that the detection result of the abnormal access behavior is incorrect, the abnormal access behavior is recorded as a false alarm sample for model optimization.

[0235] S1518: The model optimization module updates the model parameters according to the false alarm and / or missed alarm detection results of the abnormal access behavior.

[0236] In a possible implementation, according to the audit result obtained from S1517, the false alarm and / or missed alarm results of one or more of the first model, the second model, or the third model are obtained. Using the false alarm and / or missed alarm results, one or more of the first model, the second model, or the third model are optimized, and the locally stored model parameters are updated.

[0237] S1519: The electronic device updates the directed graph based on the incremental log.

[0238] In a possible understanding, the incremental log may refer to the first log collected within the latest time threshold.

[0239] In a possible implementation, the directed graph corresponding to the second log is updated according to the directed graph generated from the first log data in the embodiments of the present application. Among them, the generation time of the second log data is earlier than the generation time of the first log data.

[0240] Optionally, the obtained directed graph after update may include the directed graph corresponding to the second log and the directed graph of the first log.

[0241] Optionally, the obtained directed graph after update may include the directed graph corresponding to the second log and the directed graph obtained by filtering the part overlapping with the third log in the first log.

[0242] S1520: The electronic device stores the updated directed graph file and model parameters locally.

[0243] Optionally, in a possible implementation, the embodiments of the present application store the updated model parameters obtained in S1518 and the updated directed graph obtained in S1519 locally. Subsequently, the iterative update process can be performed in the model update unit to obtain a more accurate model.

[0244] The embodiments of the present application provide an abnormal access behavior detection method, which can be independent of historical samples, generate a directed graph based on the first log, and identify abnormal access behaviors in the directed graph from the perspective of the graph. At the same time, the embodiments of the present application do not require historical attack sample support, thereby reducing the requirements for the quality of the data source, enabling rapid deployment, and adding an online learning mechanism to optimize the model in real time to improve the accuracy of abnormal access behavior detection.

[0245] The embodiments of the present application will be described in detail in combination with flowcharts for an abnormal access behavior detection method. However, it should be understood that the relevant descriptions of these flowcharts and their corresponding embodiments are only examples for easy understanding and should not constitute any limitation to the present application. Each step in each flowchart is not necessarily required to be executed. For example, some steps can be omitted. Moreover, the execution order of each step is not fixed and is not limited to that shown in the figure. The execution order of each step should be determined according to its function and internal logic.

[0246] To facilitate the understanding of the embodiments of the present application, the embodiments of the present application provide a system architecture for abnormal access behavior detection. Among them, the abnormal access behavior detection system can be a single electronic device with the function of abnormal access behavior detection. It can also be a combination of at least two electronic devices, that is, at least two electronic devices are combined into an integrated system with the function of abnormal access behavior detection. When the abnormal access behavior detection system is a combination of at least two electronic devices, the two electronic devices in the abnormal access behavior detection system can communicate through one of the communication methods of Bluetooth, wired connection, or wireless transmission.

[0247] Among them, the specific system architecture of abnormal access behavior detection is not limited to the following structure.

[0248] As Figure 16 shown, the abnormal access behavior detection system may include: a data processing module 1601, a multi-dimensional detection module 1602, an association analysis module 1603, and an online learning module 1604.

[0249] The data processing module 1601 can be used to perform steps such as extracting fields, filtering logs, and generating a directed graph. For example, the data processing module 1601 is used to preprocess all network access logs within the detection range to obtain first log data, and generate a directed graph according to the first log data.

[0250] Exemplarily, the preprocessing may include removing irrelevant information from the network access logs, aggregating the same access records, etc. Among them, a possible implementation of removing the irrelevant information from the network access logs is: in the network access logs, removing other information except for the required information such as the source IP, destination IP, access times, etc. For example, the other information may include network environment information, etc. Exemplarily, a possible implementation of removing the irrelevant information from the network access logs is: in the network access logs, removing the access records of the whitelist users. In one possible understanding, the whitelist users may be the internal users of the enterprise. In one possible understanding, the same access records may be two or more access behaviors with the same source IP and destination IP in the first log data.

[0251] The multi-dimensional detection module 1602 is used to identify abnormal access behaviors in the directed graph by using one or more of the first model, the second model, or the third model.

[0252] Among them, the specific implementation manner of using one or more of the first model, the second model, or the third model to identify abnormal access behaviors in the directed graph may correspond to Figure 13 the description in S1302 therein, which will not be elaborated here.

[0253] The first model, the second model, and the third model used in the multi-dimensional detection module 1602 are implemented based on unsupervised algorithms and historical logs, and may not rely on the prior knowledge input by humans and specific metric thresholds, without the need to select features to construct rule templates, avoiding the diversity requirements of the existing detection models for the source data, and having good generality.

[0254] The correlation analysis module 1603 is used to comprehensively judge the preliminary detection results in combination with the historical records and formulate an alarm information sending strategy. The correlation analysis module 1603 performs correlation analysis on the detection results of the first model, the second model, and the third model, detects abnormal access behaviors from different dimensions, and comprehensively monitors various types of abnormal access behaviors.

[0255] In a possible implementation manner, the correlation analysis is to construct a linear function between multiple recognition results of the first model, the second model, or the third model, and when the value of the linear function meets the preset abnormal conditions, determine that the access behavior corresponding to the linear function is the abnormal detection result.

[0256] Exemplarily, a possible implementation of the embodiment of the present application for performing correlation analysis on the detection result by constructing a linear function between various recognition results of the first model, the second model, or the third model is as follows: the output results corresponding to the first model, the second model, and the third model are x1, x2, and x3 respectively. Among them, the abnormal access behavior detection result determined after the correlation analysis of x1, x2, and x3 is y. The embodiment of the present application constructs a linear function y = k1x1 + k2x2 + k3x3 between the first model, the second model, and the third model. Among them, k1, k2, and k3 are the weights of the respective detection results of the first model, the second model, and the third model in the correlation analysis process. The magnitudes of k1, k2, and k3 can be respectively set to N (N can be any value greater than 0, such as 1, etc.). When the output value y of the linear function y = k1x1 + k2x2 + k3x3 meets the preset abnormal conditions, the access behavior corresponding to the linear function is determined as the abnormal detection result. Among them, the weights k1, k2, and k3 of the respective detection results of the first model, the second model, and the third model in the correlation analysis process and the preset abnormal conditions of the output value y of the linear function can be set by humans and / or machines.

[0257] In a possible implementation manner, the embodiment of the present application constructs a linear function between the first model, the second model, and the third model. Among them, the weights k1, k2, and k3 of the respective detection results of the first model, the second model, and the third model in the correlation analysis process are respectively 1. At this time, the linear function between the first model, the second model, and the third model is y = x1 + x2 + x3. Among them, in a possible implementation manner, if there is an abnormal access behavior in the detection result, the outputs corresponding to the first model, the second model, and the third model are 1, otherwise 0. If the value y of the linear function is greater than or equal to 1, there is an abnormal access behavior in the directed graph. Exemplarily, if the output of the first model is 1, the abnormal access behavior can be the abnormal access behavior of multi-node jump login; if the output of the second model is 1, the abnormal access behavior can be the abnormal access behavior of cross-business group access; if the output of the third model is 1, the abnormal access behavior can be the abnormal access behavior that does not conform to the historical behavior access.

[0258] In a possible understanding manner, the historical record can be the access record with no abnormal access behavior detected in the recent period.

[0259] Exemplarily, a possible implementation of formulating an alarm information sending strategy is as follows: based on the detection result obtained by correlating and analyzing one or more of the three models, alarm information is sent to the target object according to the abnormal access behavior detection result. Among them, the alarm information includes one or more of the following: the log information corresponding to the abnormal access behavior detection result, the cause of the alarm generation, or the recommended handling method.

[0260] Among them, the specific bubbling strategy can correspond to the description of S1512 in Figure 15 , which will not be elaborated here.

[0261] The online learning module 1604 may include a result feedback part, a result review part, a model update part, etc., and is used to update the model parameters by collecting detection and review results.

[0262] Exemplarily, the specific implementation of the online learning module 1604 can correspond to the description of S1514 to S1517 in Figure 15 , which will not be elaborated here.

[0263] Among them, the system architecture described in the embodiments of the present application is to more clearly illustrate the technical solutions of the embodiments of the present application, and does not constitute a limitation on the technical solutions provided by the embodiments of the present application. Further, those skilled in the art can know that with the evolution of the network architecture and the emergence of new service scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0264] The method for detecting abnormal access behavior in the embodiments of the present application has been described above. Next, the device for executing the method for detecting abnormal access behavior provided by the embodiments of the present application will be described. Those skilled in the art can understand that the method and the device can be combined and cited with each other, and the abnormal access behavior detection device provided by the embodiments of the present application can execute the steps in the above-mentioned method for detecting abnormal access behavior.

[0265] As Figure 17 shown, Figure 17 FIG. shows a schematic structural diagram of an abnormal access behavior detection device provided by an embodiment of the present application. The abnormal access behavior detection device may be the terminal device in the embodiment of the present application, or a chip or a chip system in the terminal device. The abnormal access behavior detection device includes: a processing unit 1701. Among them, the processing unit 1701 is used to generate a directed graph according to the first log data; wherein, the directed graph includes: a plurality of nodes for identifying devices, and directed access relationships between the plurality of nodes; the processing unit 1701 is further used to identify abnormal access behaviors in the directed graph by using one or more of the first model, the second model, or the third model, and determine an abnormal detection result; wherein, the first model is used to identify abnormal access behaviors of multi-node jump logins according to the directed graph; the second model is used to identify abnormal access behaviors of cross-business group access according to the directed graph; the third model is used to identify abnormal access behaviors that do not conform to historical access behaviors according to the directed graph.

[0266] Exemplarily, taking the abnormal access behavior detection device as the terminal device or applied to a chip or a chip system in the terminal device as an example, the processing unit 1701 is used to support the abnormal access detection device to execute S1301 and / or S1302, etc. in the above embodiments.

[0267] In a possible implementation, the abnormal access behavior detection device may further include: a storage unit 1702. The storage unit 1702 may include one or more memories, and the memory may be a device or a component in one or more circuits for storing programs or data.

[0268] The storage unit 1702 may exist independently and be connected to the processing unit 1701 through a communication bus. The storage unit 1702 may also be integrated with the processing unit 1701.

[0269] Taking the abnormal access behavior detection device as an example of the chip or chip system of the terminal device in the embodiments of the present application, the storage unit 1702 may store computer-executable instructions of the method of the terminal device, so that the processing unit 1701 executes the method of the terminal device in the above embodiments. The storage unit 1702 may be a register, a cache, or a random access memory (RAM), etc. The storage unit 1702 may be integrated with the processing unit 1701. The storage unit 1702 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, and the storage unit 1702 may be independent of the processing unit 1701.

[0270] In a possible implementation, the first model, the second model, and the third model are all implemented using unsupervised algorithms.

[0271] In a possible implementation, the processing unit is specifically configured to calculate the maximum number of hops from the source node to the destination node according to the source node and the destination node among multiple nodes; wherein, the maximum number of hops is used to represent the number of hops for continuous access with the destination node as a springboard; the processing unit is further specifically configured to identify an access behavior with the maximum number of hops greater than a first threshold as an abnormal access behavior.

[0272] In a possible implementation, the processing unit is specifically configured to calculate the maximum number of hops from the source node to the destination node according to the source node and the destination node among multiple nodes; wherein, the maximum number of hops is used to represent the number of hops for continuous access with the destination node as a springboard; the processing unit is further specifically configured to identify an access behavior with the maximum number of hops greater than a first threshold as an abnormal access behavior.

[0273] In a possible implementation, the processing unit is specifically configured to convert the nodes in the directed graph into embedding vectors; for the source node and the destination node among the multiple nodes, the processing unit is further specifically configured to normalize the embedding vector matrix corresponding to the set of predecessor nodes of the destination node to obtain a set of normalized unit vectors; wherein, the set of predecessor nodes of the destination node is the set of nodes in the directed graph that point to the destination node; the processing unit is further specifically configured to calculate the cosine similarity between the embedding vector corresponding to the source node and the set of normalized unit vectors; the processing unit is further configured to identify an access behavior with a cosine similarity less than a second threshold as an abnormal access behavior.

[0274] In a possible implementation, the processing unit is specifically configured to train the embedding vector matrix corresponding to the set of predecessor nodes of the destination node by using a one-class support vector machine; the processing unit is further specifically configured to normalize the trained embedding vector matrix.

[0275] In a possible implementation, the processing unit is specifically configured to identify abnormal access behaviors in the directed graph by using one or more of a first model, a second model, or a third model to obtain multiple recognition results; the processing unit is further specifically configured to perform correlation analysis on the multiple recognition results to obtain an anomaly detection result.

[0276] In a possible implementation, the processing unit is specifically configured to construct a linear function between the multiple recognition results; the processing unit is further configured to determine that the access behavior corresponding to the linear function is an anomaly detection result when the value of the linear function meets a preset anomaly condition.

[0277] In a possible implementation, the processing unit is specifically configured to update the directed graph according to the anomaly detection result; the processing unit is further specifically configured to update the hyperparameters of the first model, the hyperparameters of the second model, and / or the hyperparameters of the third model according to the anomaly detection result.

[0278] In a possible implementation, the processing unit is specifically configured to obtain second log data; the generation time of the second log data is earlier than the generation time of the first log data; the processing unit is further configured to generate a directed graph of the second log data; the processing unit is further configured to load the directed graph of the second log data and the model parameters related to the first model, the second model, or the third model into the first model, the second model, or the third model.

[0279] In a possible implementation, the processing unit is further configured to periodically obtain first log data; the processing unit is further configured to generate a first directed graph corresponding to the first log data; the processing unit is further configured to filter out the part of the first directed graph that overlaps with the third directed graph of the third log data to obtain a directed graph; wherein, the generation time of the third log data is earlier than the generation time of the first log data, and the difference between the generation time of the third log data and the generation time of the first log data is less than a time threshold.

[0280] In a possible implementation, the node is the Internet Protocol (IP) address of a device, and the directed access relationships among multiple nodes include one or more of the following: the maximum number of accesses per hour among multiple nodes, the total number of historical accesses, or the latest access time.

[0281] In a possible implementation, the abnormal access detection device may further include: a communication unit 1703. The communication unit 1703 is configured to support the abnormal access behavior detection device to perform steps of information sending or receiving. Exemplarily, when the abnormal access behavior detection device is a terminal device, the communication unit 1703 may be a communication interface or an interface circuit. When the abnormal access behavior detection device is a chip or a chip system within a terminal device, the communication unit 1703 may be a communication interface. For example, the communication interface may be an input / output interface, a pin, or a circuit, etc.

[0282] Exemplarily, the communication unit 1703 is configured to send an alarm message to a target object according to the abnormal detection result.

[0283] In a possible implementation, the alarm message includes one or more of the following: log information corresponding to the abnormal detection result, cause of alarm generation, or recommended handling method.

[0284] The device in this embodiment can correspondingly be used to execute the steps performed in the above method embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here.

[0285] Figure 18 This is a schematic hardware structure diagram of an abnormal access behavior detection device provided by an embodiment of the present application. Please refer to Figure 18 , the network management device includes: a memory 1801 and a processor 1802. The communication device may further include an interface circuit 1803. Among them, the memory 1801, the processor 1802, and the interface circuit 1803 can communicate; Exemplarily, the memory 1801, the processor 1802, and the interface circuit 1803 can communicate through a communication bus. The memory 1801 is used to store computer execution instructions and is controlled by the processor 1802 to execute, so as to implement the abnormal access behavior detection method provided in the following embodiments of the present application.

[0286] In a possible implementation, the computer-executable instructions in the embodiments of the present application may also be referred to as application code, and the embodiments of the present application do not make specific limitations thereon.

[0287] Optionally, the interface circuit 1803 may further include a transmitter and / or a receiver.

[0288] Optionally, the above-mentioned processor 1802 may include one or more CPUs, and may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the present application may be directly embodied as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor.

[0289] The embodiments of the present application further provide a computer-readable storage medium. The methods described in the above embodiments may be implemented in whole or in part by software, hardware, firmware, or any combination thereof. If implemented in software, the functions may be stored as one or more instructions or codes on a computer-readable medium or transmitted on a computer-readable medium. The computer-readable medium may include a computer storage medium and a communication medium, and may also include any medium that can transfer a computer program from one place to another. The storage medium may be any target medium accessible by a computer.

[0290] In one possible implementation, the computer-readable medium may include RAM, ROM, compact disc read-only memory (CD-ROM), or other optical disc storage, magnetic disk storage, or any other medium targeted to carry or store the required program code in the form of instructions or data structures and accessible by a computer. Moreover, any connection is properly termed a computer-readable medium. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, Digital Subscriber Line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. As used herein, disk and disc include optical disc, laser disc, optical disc, Digital Versatile Disc (DVD), floppy disk, and Blu-ray disc, where disks typically reproduce data magnetically, while discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0291] Embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processing unit of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processing unit of the computer or other programmable data processing device generate means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0292] The above specific implementation manners further elaborate on the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific implementation manners of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present invention should be included within the protection scope of the present invention.

[0293] It should be noted that the abnormal access behavior described in the present application may also adopt other definitions or names in specific applications. Exemplarily, the abnormal access behavior may be referred to as an abnormal attack behavior, abnormal access, etc. Or the above abnormal access behavior may also be defined with other names according to the actual application scenario, and the embodiments of the present application do not make specific limitations thereto.

Claims

1. An abnormal access behavior detection method, applied to an electronic device, characterized in that, including: generating a directed graph according to first log data; the directed graph includes: a plurality of nodes for identifying devices, and a directed access relationship between the plurality of nodes; identifying abnormal access behaviors in the directed graph by using one or more of a first model, a second model, or a third model to determine an abnormal detection result; wherein, the first model is used to identify an abnormal access behavior of multi-node jump login according to the directed graph, and the abnormal access behavior is an access behavior in which the maximum number of hops from a source node to a destination node in the directed graph is greater than a first threshold; the second model is used to identify an abnormal access behavior of cross-business-group access according to the directed graph, and the abnormal access behavior is an access behavior corresponding to a node with cross-business-group access, and the business group is formed by classifying a plurality of the nodes into the business group where the neighbor node with the maximum gain is located; the third model is used to identify an abnormal access behavior that does not conform to historical access behaviors according to the directed graph, and the abnormal access behavior is an access behavior in which the cosine similarity between the embedding vector corresponding to the source node and the set of normalized unit vectors is less than a second threshold, and the set of normalized unit vectors is obtained by normalizing the embedding vector matrix corresponding to the set of precursor nodes of the destination node, and the set of precursor nodes of the destination node is the set of nodes in the directed graph that point to the destination node.

2. The method according to claim 1, wherein The first model, the second model, and the third model are all implemented by using unsupervised algorithms.

3. The method according to claim 1 or 2, characterized in that, During the process of using the second model to identify abnormal access behaviors in the directed graph, it further includes: compressing the nodes classified into the same community into a first node until the result of the classification no longer changes, and the business group is the community; identifying the access behavior corresponding to the first node with cross-community access as an abnormal access behavior.

4. The method according to claim 1, characterized in that The normalization of the embedding vector matrix corresponding to the set of precursor nodes of the destination node includes: training the embedding vector matrix corresponding to the set of precursor nodes of the destination node by using a one-class support vector machine; normalizing the trained embedding vector matrix.

5. The method according to any one of claims 1-2 and 4, characterized in that, Using one or more of the first model, the second model, or the third model to identify abnormal access behaviors in the directed graph to determine an abnormal detection result includes: using one or more of the first model, the second model, or the third model to identify abnormal access behaviors in the directed graph to obtain multiple identification results; performing correlation analysis on the multiple identification results to obtain the abnormal detection result.

6. The method according to claim 5, wherein The performing correlation analysis on the multiple identification results to obtain the abnormal detection result includes: constructing a linear function between the multiple identification results; when the value of the linear function meets a preset abnormal condition, determining the access behavior corresponding to the linear function as the abnormal detection result.

7. The method according to any one of claims 1-2, 4, and 6, characterized in that, It further includes: updating the directed graph according to the abnormal detection result; updating the hyperparameters of the first model, the hyperparameters of the second model, and / or the hyperparameters of the third model according to the abnormal detection result.

8. The method according to any one of claims 1-2, 4, and 6, characterized in that It further includes: obtaining second log data; the generation time of the second log data is earlier than the generation time of the first log data; Generate a directed graph of the second log data; Load the directed graph of the second log data, and the model parameters related to the first model, the second model, or the third model, into the first model, the second model, or the third model.

9. The method according to any one of claims 1-2, 4, and 6, characterized in that, The generating a directed graph according to the first log data includes: Regularly obtain the first log data; Generate a first directed graph corresponding to the first log data; Filter out the overlapping part between the first directed graph and the third directed graph of the third log data to obtain the directed graph; the generation time of the third log data is earlier than the generation time of the first log data, and the difference between the generation time of the third log data and the generation time of the first log data is less than a time threshold.

10. The method according to any one of claims 1-2, 4, and 6, characterized in that The node is the Internet Protocol (IP) address of a device, and the directed access relationships between multiple nodes include one or more of the maximum number of accesses per hour, the total number of historical accesses, or the latest access time between multiple nodes.

11. The method according to any one of claims 1-2, 4, and 6, characterized in that, It further includes: Send an alarm message to a target object according to the anomaly detection result.

12. The method according to claim 11, wherein The alarm message includes one or more of the following: the log information corresponding to the anomaly detection result, the cause of the alarm, or the recommended handling method.

13. An electronic device, characterized in that, It includes: A unit for performing each step of any one of claims 1-12.

14. An electronic device, characterized in that, It includes: a processor for calling a program in a memory to execute the method of any one of claims 1-12.

15. An electronic device, characterized in that, It includes: A processor and an interface circuit, the interface circuit is used to communicate with other devices, and the processor is used to execute the method of any one of claims 1 to 12.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions, and when the instructions are executed, the computer executes the method of any one of claims 1-12.

17. A computer program product, characterized in that, The computer program product stores instructions, and when the instructions are executed, the computer executes the method of any one of claims 1-12.

Citation Information

Patent Citations

  • Method, device, medium and device for detecting abnormal behavior access of world wide web

    CN109040073A