A method and system for identifying malicious websites based on IP address feature analysis
By building an IP dynamic graph structure and combining the graph attention network and long-term memory network, complex behavior patterns between IP addresses are captured, and the problems of inefficient identification and lack of adaptive detection mechanisms in the prior art are solved, and efficient identification and dynamic detection balance of complex attack methods are achieved.
Patent Information
- Application Number
- CN202510272470.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-03-10
AI Technical Summary
The existing technology is difficult to capture the collaborative behavior patterns that evolve in real-time between IP addresses, and the recognition efficiency is inefficient; the fusion of space-time features is missing, so it is impossible to effectively detect the day-night mode switching and cross-autonomous system migration designed by attackers; the adaptive detection mechanism is lacking, so it is impossible to dynamically balance the detection accuracy and false alarm rate.
By constructing an IP dynamic graph structure, combining the graph attention network and long-term memory network, the complex relationships and behavior patterns between IP addresses are captured; the spatial topological relationships and time series behavior characteristics of IP addresses are integrated; dynamic anomaly detection threshold is used to dynamically adjust the detection threshold according to changes in the network environment.
It effectively improves the recognition efficiency of complex attack methods such as distributed proxy pool attacks; overcomes the shortcomings of methods based on rules engines or single-dimensional statistical models in detecting attackers' evasion strategies; realizes dynamic balance of detection accuracy and false alarm rates when network environment changes, and responds to sudden large-scale IP fission attacks in a timely manner.
Smart Images

Figure CN119788427B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of network security technology, and more specifically, to a method and system for identifying malicious URLs based on IP address feature analysis. Background Art
[0002] While the widespread use of the Internet has brought convenience to people's lives, it has also provided a breeding ground for online fraud. In recent years, online fraud cases have occurred frequently, and fraud methods have been constantly updated, causing huge economic losses to society. Malicious URLs, as a key carrier of online fraud, disguise themselves as legitimate URLs to trick users into visiting and committing fraudulent acts. With the continuous evolution of online fraud technology, attackers use dynamic IP pool rotation, cross-domain communication disguises, and other means to achieve rapid fission and covert propagation of fraudulent URLs. Traditional detection methods face significant technical bottlenecks in dealing with such dynamic and organized threats. Existing technologies mainly have the following limitations:
[0003] 1. Insufficient dynamic correlation analysis capabilities: Traditional methods rely on static IP reputation databases or isolated node analysis, making it difficult to capture the collaborative behavior patterns that evolve in real time between IP addresses, resulting in inefficient identification of distributed proxy pool attacks.
[0004] 2. Lack of spatiotemporal feature fusion: Methods based on rule engines or single-dimensional statistical models cannot effectively integrate the spatial topological relationships and time series behavioral characteristics of IP addresses, and lack the ability to detect evasion strategies designed by attackers, such as daytime and nighttime mode switching and cross-AS migration.
[0005] 3. Lack of adaptive detection mechanism: Fixed threshold strategies are easily affected by fluctuations in the network environment and cannot dynamically balance detection accuracy and false alarm rate. In particular, they are slow to respond when dealing with sudden large-scale IP fission attacks.
[0006] Existing technologies, such as the Chinese patent application with publication number "CN113779481A", disclose a method for identifying fraudulent websites. The method includes: obtaining the web page source code of the website to be identified, and obtaining a target feature vector based on the web page source code; judging whether there is a target standard feature vector matching the target feature vector in a preset fraudulent website feature library; if so, determining the website to be identified as a fraudulent website, and determining the fraudulent website classification of the website to be identified based on the fraud type label of the standard feature vector set corresponding to the target standard feature vector.
[0007] The problems with the above-mentioned existing technologies are that, by matching feature vectors obtained from webpage source code with a pre-set feature library to identify fraudulent websites, this method struggles to capture the real-time evolving collaborative behavior patterns between IP addresses. It also has limited ability to identify attacks that rely on dynamic associations of IP addresses, such as distributed proxy pool attacks. Furthermore, it lacks the ability to detect evasion strategies that exploit temporal and spatial transformations. The lack of a mechanism for dynamic adjustment based on the network environment makes it difficult to balance detection accuracy and false positive rates when the network environment changes, leading to a delayed response to new or changing fraudulent websites. Summary of the Invention
[0008] To solve the above technical problems, the present invention proposes a method and system for identifying malicious URLs based on IP address feature analysis.
[0009] The technical solutions of the present invention are as follows:
[0010] The present invention proposes a method for identifying malicious websites based on IP address feature analysis, comprising the following steps:
[0011] Step S1, collecting network traffic data of the IP address through network traffic logs, firewall logs and DNS query records, and preprocessing the collected network traffic data;
[0012] Step S2, based on the pre-processed network traffic data, the IP addresses are used as nodes and the communication relationships between the IP addresses are used as edges to construct an IP dynamic graph structure;
[0013] Step S3: Encode the IP dynamic graph structure through the graph attention network to capture the complex relationship between nodes, and combine it with the long short-term memory network to capture the time series characteristics of IP address behavior and construct the node behavior pattern;
[0014] Step S4: Calculate the node's reconstruction error and local outlier factor (LOF) based on the node's behavior pattern, and calculate the IP address's anomaly score. When the IP address's anomaly score is higher than the dynamic anomaly detection threshold, the IP address is determined to be abnormal. The abnormal IP address is then deeply analyzed in combination with the URL's domain name features, web page content features, and network access behavior features to identify malicious URLs.
[0015] As a preferred implementation, the network traffic data includes request frequency, communication object, traffic size, protocol type, port number, and geographic location information.
[0016] As a preferred implementation, the IP dynamic graph structure is updated based on network traffic data collected in real time.
[0017] As a preferred embodiment, the step of encoding the IP dynamic graph structure through the graph attention network to capture the complex relationship between nodes includes:
[0018] Calculate the correlation between nodes. The calculation formula is:
[0019] ;
[0020] Based on the normalization of the correlation between the node and its neighboring nodes, the node's attention coefficient is obtained:
[0021] ;
[0022] Aggregate neighbor node features according to the attention coefficient as the complex relationship between nodes:
[0023] ;
[0024] Where: is the association degree between node i and its neighbor node j; is the rectified linear unit activation function; is the transpose of the attention mechanism parameter vector; W is the graph attention network weight matrix; is the feature vector of node i; is the feature vector of node j; For splicing operation; is the time decay factor; is the last communication time interval between node i and node j; is the diversity factor, which is determined by the number of communication partners between nodes i and j; The risk factor of the agreement is assigned different risk factors according to the type of agreement; is the attention coefficient; is the set of neighbor nodes of node i; is the association degree between node i and its neighbor node q; is the feature vector of node i after being processed by the graph attention network; is the activation function.
[0025] As a preferred embodiment, in the process of constructing the behavior pattern of the node, the behavior pattern of the node is generated by a gated attention fusion mechanism, wherein:
[0026] The calculation formula of the gating signal is:
[0027]
[0028] Where: g is the gate signal; is a learnable parameter; is the Sigmoid function; For splicing operation;
[0029] ;
[0030] Where: is the final behavior pattern representation of node i; is the feature vector of node i output after processing by the graph attention network; is the time series feature of node i output by the long short-term memory network.
[0031] As a preferred embodiment, the reconstruction error and local outlier factor (LOF) of the node are calculated according to the behavior pattern of the node to obtain the anomaly score of the IP address. The specific calculation process of the anomaly score is as follows:
[0032] The reconstruction error is calculated based on the autoencoder of the graph attention network. The specific formula is:
[0033] ;
[0034] Where: is the reconstruction error; is the original feature vector of node i; For the autoencoder pair Reconstructed feature vector of node i;
[0035] The local outlier factor LOF is calculated as follows:
[0036] ;
[0037] Where: is the local outlier factor of node i; 、 are the k-local reachability densities of node m and node i respectively; is the k-neighboring set of node i;
[0038] The specific calculation formula for the anomaly score is:
[0039] ;
[0040] in: is the abnormality score of node i; is the normalization operation; is the weight coefficient, which is dynamically adjusted based on the historical anomaly score. The specific calculation formula is:
[0041] ;
[0042] Where: The variance of the historical anomaly scores; is the adjustment parameter.
[0043] As a preferred embodiment, the dynamic anomaly detection threshold is specifically calculated as follows:
[0044] ;
[0045] Where: is the dynamic anomaly detection threshold at time t; is the mean of the anomaly scores within the historical time window; is the standard deviation over the same period; 、 is a dynamic weight coefficient, which is dynamically adjusted based on the historical score variance; An exponential moving average of historical anomaly scores.
[0046] On the other hand, the present invention also provides a malicious website identification system based on IP address feature analysis, comprising:
[0047] The data collection and preprocessing module collects network traffic data of IP addresses through network traffic logs, firewall logs and DNS query records, and preprocesses the collected network traffic data;
[0048] The dynamic graph structure construction module uses the pre-processed network traffic data, takes IP addresses as nodes, and the communication relationships between IP addresses as edges to construct the IP dynamic graph structure;
[0049] The feature extraction and fusion module encodes the IP dynamic graph structure through the graph attention network to capture the complex relationships between nodes. It also combines the long short-term memory network to capture the time series characteristics of IP address behavior and construct the node behavior model.
[0050] The malicious URL identification module calculates the node's reconstruction error and local outlier factor (LOF) based on the node's behavior pattern to obtain the IP address's anomaly score. When the IP address's anomaly score is higher than the dynamic anomaly detection threshold, the IP address is judged to be abnormal. The module then conducts an in-depth analysis of the abnormal IP address combined with the URL's domain name features, web page content features, and network access behavior features to identify malicious URLs.
[0051] On the other hand, the present invention further provides an electronic device having a computer program stored thereon, wherein when the computer program is executed by a processor, the method for identifying malicious URLs based on IP address feature analysis as described in any embodiment of the present invention is implemented.
[0052] On the other hand, the present invention also provides a computer-readable medium for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement a malicious URL identification method based on IP address feature analysis as described in any embodiment of the present invention.
[0053] The present invention has the following beneficial effects:
[0054] 1. By constructing an IP dynamic graph structure and combining it with a graph attention network and a long short-term memory network, we can capture the real-time evolving collaborative behavior patterns between IP addresses, effectively improving the recognition efficiency of complex attack methods such as distributed proxy pool attacks.
[0055] 2. The integration of the spatial topological relationship and time series behavioral characteristics of IP addresses overcomes the shortcomings of methods based on rule engines or single-dimensional statistical models in detecting attackers' evasion strategies such as day-night mode switching and cross-autonomous system migration.
[0056] 3. The use of dynamic anomaly detection thresholds can dynamically balance detection accuracy and false alarm rate according to changes in the network environment, providing a more timely response when dealing with sudden large-scale IP fission attacks. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 Schematic diagram of the method flow of embodiment 1. DETAILED DESCRIPTION
[0058] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0059] It should be understood that the step numbers used herein are only for convenience of description and are not intended to limit the order in which the steps are to be executed.
[0060] It should be understood that the terms used in the present specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, the singular forms "a", "an" and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0061] The terms “include” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0062] The term "and / or" refers to and includes any and all possible combinations of one or more of the associated listed items.
[0063] Example 1:
[0064] In order to make the purpose, technical solutions and advantages of the present invention more clear, the following will be combined with the specific embodiments of the present application and refer to the attached Figure 1 , clearly and completely describe the technical solution of the present invention.
[0065] To solve the problems of the prior art, this embodiment provides a method for identifying malicious websites based on IP address feature analysis, comprising the following steps:
[0066] Step S1, collecting network traffic data of the IP address through network traffic logs, firewall logs and DNS query records, and preprocessing the collected network traffic data;
[0067] The pre-processing process includes: data cleaning, deduplication and formatting to ensure data quality and consistency;
[0068] The network traffic data includes:
[0069] Request frequency: The number of network requests sent by an IP address within a given period of time. It can be used to measure the activity intensity of an IP address or a group of IP addresses on the network.
[0070] Communication partner: During network communication, another entity (usually another IP address or domain name) that interacts with a specific IP address.
[0071] Traffic size: The amount of data in each request or response.
[0072] Protocol type: network protocols such as HTTP, HTTPS, FTP, etc.
[0073] Port number: Service port related to communication.
[0074] Geographic location information: Geographic coordinates (longitude and latitude) obtained based on IP address query.
[0075] Step S2, based on the pre-processed network traffic data, the IP addresses are used as nodes and the communication relationships between the IP addresses are used as edges to construct an IP dynamic graph structure;
[0076] In the IP dynamic graph structure, nodes and edges are assigned corresponding features based on preprocessed network traffic data. Node features include: request frequency (the number of requests initiated per unit time); geographic location (geographic coordinates obtained from IP address queries); historical behavior (activity patterns over a period of time, such as frequent visits to specific types of websites or unusual activity within a certain time period); associated entities (other entities that the IP address frequently interacts with, such as other IP addresses or domain names); and time series features (reflecting behavioral trends at different times of the day and differences in behavior between weekends and weekdays). Edge communication relationships are derived by analyzing network traffic data. Edge features include: connection frequency (the number of connections established between two IP addresses); interaction strength (the strength of interaction between two nodes, measured by the amount of data transferred); duration (the duration of a complete communication session); latency (the time from request initiation to response reception); error rate (the proportion of errors occurring during communication, such as timeouts or retransmissions); and encryption status (whether encryption is used in communication and the type of encryption algorithm used). Thus, the dynamic graph structure based on network traffic data can reflect the interaction relationships between IP addresses and changes in their behavioral patterns in real time.
[0077] The IP dynamic graph structure can be updated according to the network traffic data collected in real time to reflect the latest network status.
[0078] Step S3: Encode the IP dynamic graph structure through the graph attention network to capture the complex relationship between nodes, and combine the temporal characteristics of the IP address encoded by the long short-term memory network to construct the node behavior pattern;
[0079] The graph attention network encodes the node features and edge features in the IP dynamic graph structure to capture the complex relationships between nodes. The specific steps are as follows:
[0080] Calculate the correlation between nodes. The calculation formula is:
[0081] ;
[0082] in:
[0083] ;
[0084] ;
[0085] Based on the normalization of the correlation between the node and its neighboring nodes, the node's attention coefficient is obtained:
[0086] ;
[0087] Aggregate neighbor node features according to the attention coefficient as the complex relationship between nodes:
[0088] ;
[0089] Where: is the association degree between node i and its neighbor node j; is the rectified linear unit activation function; is the transpose of the attention mechanism parameter vector; W is the graph attention network weight matrix; is the feature vector of node i; is the feature vector of node j; For splicing operation; is the time decay factor; is the last communication time interval between node i and node j; is the diversity factor, the number of communication objects between nodes i and j Sure; The protocol risk factor is assigned different risk weights based on the protocol type (e.g. HTTP=1.2, HTTPS=0.8) to reflect the security of the protocol. is the basic time decay factor, take 0.1; is the average of the historical communication intervals between nodes i and j; is the learning rate, take 0.01; is the attention coefficient; is the set of neighbor nodes of node i; is the association degree between node i and its neighbor node q; is the feature vector of node i after being processed by the graph attention network, which can reflect the relative position of the node in the IP dynamic graph structure and the characteristics of its interaction relationship with other nodes; is the activation function.
[0090] While using the graph attention network to encode the IP dynamic graph structure, the long short-term memory network is introduced to capture the time series characteristics of IP address behavior. For the time series data of each node, LSTM can be applied to extract long-term dependencies. Assume that the maximum time step of the data sequence is , then , the workflow of the LSTM unit is as follows:
[0091] ;
[0092] Where, is the output of the input gate, between [0, 1], which is used to control the extent to which new information is added to the cell state. is the weight matrix of the input gate, which determines the hidden state at the previous moment and current moment input Impact on the input gate. is the bias term of the input gate, which adjusts the default behavior of the input gate. It is the sigmoid activation function, which compresses the calculation results into the (0,1) interval to determine which information should be considered important.
[0093] ;
[0094] Where, is the output of the forget gate, between [0, 1], which determines which information in the cell state at the previous moment should be forgotten. is the weight matrix of the forget gate, which determines the hidden state of the previous moment and current moment input Impact on the forget gate. is the bias term of the forget gate. It is the sigmoid activation function, which compresses the calculation results into the (0,1) interval to determine which information should be forgotten.
[0095] ;
[0096] Where, is the output of the output gate, between [0,1], which determines which parts of the current cell state should be output as the hidden state of the current time step. is the weight matrix of the output gate, which determines the hidden state at the previous moment and current moment input Impact on the output gate. is the bias term of the output gate. It is the sigmoid activation function, which compresses the calculation results into the (0,1) interval to determine which information should be output.
[0097] ;
[0098] Where, Candidate cell states represent new information that may be added to the cell state. Its value range is between [-1, 1] and is determined by the hyperbolic tangent function. Decide. is the weight matrix for generating the weight matrix of candidate cell states, is the bias term of the candidate cell state.
[0099] ;
[0100] Where, is the cell state at the current time step, which is the cell state at the previous moment controlled by the forget gate and new candidate cell states controlled by the input gate The weighted sum of . The output of the forget gate determines the cell state at the previous moment Which parts should be discarded. Is the output of the input gate, which determines the state of the new candidate cell Which parts should be added to the current cell state Go among them.
[0101] ;
[0102] Where, is the hidden state of the current time step, which is based on the current cell state The result controlled by the output gate. is the output of the output gate, is the hyperbolic tangent function.
[0103] The node representation obtained by the graph attention network and the final hidden state obtained after processing by the long short-term memory network The gated attention fusion mechanism is used to fuse and generate the behavior pattern of the node, where:
[0104] The calculation formula of the gating signal is:
[0105] ;
[0106] Where: g is the gate signal; is a learnable parameter; is the Sigmoid function;
[0107] ;
[0108] Where: is the final behavior pattern representation of node i; For splicing operation.
[0109] Step S4: Calculate the node's reconstruction error and local outlier factor (LOF) based on the node's behavior pattern to obtain the IP address's anomaly score. When the IP address's anomaly score is higher than the dynamic anomaly detection threshold, the IP address is determined to be abnormal. The abnormal IP address is then deeply analyzed in combination with the URL's domain name features, web page content features, and network access behavior features to identify malicious URLs.
[0110] The reconstruction error of the node is calculated based on the autoencoder of the graph attention network, and the specific formula is:
[0111] ;
[0112] Where: is the reconstruction error; is the original feature vector of node i; For the autoencoder pair The reconstructed feature vector of node i. The graph attention network autoencoder uses the selected loss function MSE to train the autoencoder. During the training process, the autoencoder adjusts its weights and biases so that for each input , its reconstructed output As close to the input as possible This means the autoencoder is learning how to effectively represent the normal patterns in the data. After training, we can evaluate the model's performance by comparing the reconstruction error on the test set. For normal data, the reconstruction error should be small, while for abnormal data, the reconstruction error should be relatively large. This allows us to characterize the behavior under normal patterns.
[0113] The local outlier factor LOF is specifically calculated as follows:
[0114] ;
[0115] Where: is the local outlier factor of node i; 、 are the k-local reachability densities of node m and node i respectively; is the k-neighboring set of node i.
[0116] The specific calculation formula for the anomaly score is:
[0117] ;
[0118] in: is the weight coefficient, which is dynamically adjusted based on the historical anomaly score. The specific calculation formula is:
[0119] ;
[0120] Where: is the abnormality score of node i; is the normalization operation; The variance of the historical anomaly scores; To adjust the parameters, in scenarios where network traffic is relatively stable and anomaly score fluctuations are small, the The value of The impact of the network environment is relatively mild; however, in scenarios where the network environment is complex and changeable and the abnormal score fluctuates greatly, the , enhance the regulatory effect of the variance pair so as to more flexibly adjust the weights of the reconstruction error and the local outlier factor LOF in the anomaly score calculation.
[0121] Dynamic anomaly detection threshold, the specific calculation formula is:
[0122] ;
[0123] in:
[0124] ;
[0125] Where: is the dynamic anomaly detection threshold at time t; is the mean of the anomaly scores within the historical time window; is the standard deviation over the same period; 、 is the dynamic weight coefficient, ; An exponential moving average to score historical anomalies; is the variance of the anomaly score.
[0126] We conduct in-depth analysis of abnormal IP addresses combined with domain name characteristics of URLs (such as whether there are spelling errors similar to well-known websites, use of rare or suspicious top-level domain names, etc.), web page content characteristics (such as whether it contains false winning information, links that induce downloading malware, a large number of pop-up ads, etc.), and network access behavior characteristics (such as whether there are a large number of abnormal access requests in a short period of time, frequent jumps to other suspicious URLs, etc.) to identify malicious URLs.
[0127] Example 2:
[0128] This embodiment provides a malicious website identification system based on IP address feature analysis, including:
[0129] The data collection and preprocessing module collects network traffic data of IP addresses through network traffic logs, firewall logs and DNS query records, and preprocesses the collected network traffic data;
[0130] The dynamic graph structure construction module uses the pre-processed network traffic data, takes IP addresses as nodes, and the communication relationships between IP addresses as edges to construct the IP dynamic graph structure;
[0131] The feature extraction and fusion module encodes the IP dynamic graph structure through the graph attention network to capture the complex relationships between nodes. It also combines the long short-term memory network to capture the time series characteristics of IP address behavior and construct the node behavior model.
[0132] The malicious URL identification module calculates the node's reconstruction error and local outlier factor (LOF) based on the node's behavior pattern to obtain the IP address's anomaly score. When the IP address's anomaly score is higher than the dynamic anomaly detection threshold, the IP address is judged to be abnormal. The module then conducts an in-depth analysis of the abnormal IP address combined with the URL's domain name features, web page content features, and network access behavior features to identify malicious URLs.
[0133] Example 3:
[0134] This embodiment provides an electronic device having a computer program stored thereon. When the computer program is executed by a processor, the method for identifying malicious URLs based on IP address feature analysis as described in any embodiment of the present invention is implemented.
[0135] Example 4:
[0136] This embodiment provides a computer-readable medium for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement a malicious URL identification method based on IP address feature analysis as described in any embodiment of the present invention.
[0137] In the embodiments of the present application, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. Among them, A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b and c can represent: a, b, c, a and b, a and c, b and c or a and b and c, where a, b, c can be single or multiple.
[0138] Those skilled in the art will appreciate that the various units and algorithm steps described in the embodiments disclosed herein can be implemented using a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0139] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0140] In the several embodiments provided in this application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of this application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory; hereinafter referred to as: ROM), random access memory (Random Access Memory; hereinafter referred to as: RAM), magnetic disk or optical disk, and other media that can store program code.
[0141] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention's description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for identifying malicious websites based on IP address feature analysis, characterized in that: The following steps are involved: Step S1, collecting network traffic data of IP addresses through network traffic logs, firewall logs and DNS query records, and preprocessing the collected network traffic data; Step S2, based on the pre-processed network traffic data, the IP addresses are used as nodes and the communication relationships between the IP addresses are used as edges to construct an IP dynamic graph structure; Step S3, encode the IP dynamic graph structure through the graph attention network to capture the complex relationship between nodes, and combine the long short-term memory network LSTM to capture the time series characteristics of IP address behavior and construct the behavior pattern of the node; The step of encoding the IP dynamic graph structure through the graph attention network and capturing the complex relationship between nodes includes: Calculate the correlation between nodes. The calculation formula is: ; Based on the normalization of the correlation between the node and its neighboring nodes, the attention coefficient of the node is obtained: ; Aggregate neighbor node features according to the attention coefficient as the complex relationship between nodes: ; Where: is the association degree between node i and its neighbor node j; is the rectified linear unit activation function; is the transpose of the attention mechanism parameter vector; W is the graph attention network weight matrix; is the feature vector of node i; is the feature vector of node j; For splicing operation; is the time decay factor; is the last communication time interval between node i and node j; is the diversity factor, which is determined by the number of communication partners between nodes i and j; It is the risk factor of the agreement, and different risk factors are assigned according to the type of agreement; is the attention coefficient; is the set of neighbor nodes of node i; is the association degree between node i and its neighbor node q; is the feature vector of node i after being processed by the graph attention network; is the activation function; In the process of constructing the behavior pattern of the node, the behavior pattern of the node is generated through a gated attention fusion mechanism, wherein: The calculation formula of the gating signal is: ; Where: g is the gate signal; is a learnable parameter; is the Sigmoid function; For splicing operation; ; Where: is the final behavior pattern representation of node i; is the feature vector of node i output after processing by the graph attention network; is the time series feature of node i output by the long short-term memory network; Step S4, calculate the node reconstruction error and local outlier factor LOF according to the node behavior pattern, and calculate the anomaly score of the IP address. When the anomaly score of the IP address is higher than the dynamic anomaly detection threshold, the IP address is determined to be abnormal, and the IP address with abnormality is deeply analyzed in combination with the domain name characteristics of the website, the web page content characteristics and the network access behavior characteristics to identify malicious websites; The dynamic anomaly detection threshold is specifically calculated as follows: ; Where: is the dynamic anomaly detection threshold at time t; is the mean of the anomaly scores within the historical time window; is the standard deviation during the same period; , is a dynamic weight coefficient, which is dynamically adjusted based on the historical score variance; An exponential moving average of historical anomaly scores.
2. According to claim 1, a method for identifying malicious websites based on IP address feature analysis is characterized in that: The network traffic data includes: request frequency, communication object, traffic size, protocol type, port number, and geographic location information.
3. The method for identifying malicious websites based on IP address feature analysis according to claim 1, characterized in that: The IP dynamic graph structure is updated according to the network traffic data collected in real time.
4. The method for identifying malicious websites based on IP address feature analysis according to claim 1, characterized in that: The node reconstruction error and the local outlier factor LOF are calculated according to the node behavior pattern to obtain the anomaly score of the IP address. The specific calculation process of the anomaly score is as follows: The reconstruction error is calculated based on the autoencoder of the graph attention network. The specific formula is: ; Where: is the reconstruction error; is the original feature vector of node i; For the autoencoder pair Reconstructed feature vector of node i; The local outlier factor LOF is calculated as follows: ; Where: is the local outlier factor of node i; , are the k-local reachability densities of node m and node i respectively; is the k-neighboring set of node i; The specific calculation formula for the anomaly score is: ; in: is the abnormality score of node i; It is a normalization operation; is the weight coefficient, which is dynamically adjusted based on the historical anomaly score. The specific calculation formula is: ; Where: The variance of the historical anomaly scores; To adjust the parameters.
5. A malicious website identification system based on IP address feature analysis, characterized in that: include: The data collection and preprocessing module collects network traffic data of IP addresses through network traffic logs, firewall logs and DNS query records, and preprocesses the collected network traffic data; The dynamic graph structure building module uses the IP addresses as nodes and the communication relationships between IP addresses as edges to build the IP dynamic graph structure based on the pre-processed network traffic data; The feature extraction and fusion module encodes the IP dynamic graph structure through the graph attention network to capture the complex relationship between nodes, and combines the long short-term memory network to capture the time series characteristics of IP address behavior and construct the behavior pattern of the node; The step of encoding the IP dynamic graph structure through the graph attention network and capturing the complex relationship between nodes includes: Calculate the correlation between nodes. The calculation formula is: ; Based on the normalization of the correlation between the node and its neighboring nodes, the attention coefficient of the node is obtained: ; Aggregate neighbor node features according to the attention coefficient as the complex relationship between nodes: ; Where: is the association degree between node i and its neighbor node j; is the rectified linear unit activation function; is the transpose of the attention mechanism parameter vector; W is the graph attention network weight matrix; is the feature vector of node i; is the feature vector of node j; For splicing operation; is the time decay factor; is the last communication time interval between node i and node j; is the diversity factor, which is determined by the number of communication partners between nodes i and j; It is the risk factor of the agreement, and different risk factors are assigned according to the type of agreement; is the attention coefficient; is the set of neighbor nodes of node i; is the association degree between node i and its neighbor node q; is the feature vector of node i after being processed by the graph attention network; is the activation function; In the process of constructing the behavior pattern of the node, the behavior pattern of the node is generated through a gated attention fusion mechanism, wherein: The calculation formula of the gating signal is: ; Where: g is the gate signal; is a learnable parameter; is the Sigmoid function; For splicing operation; ; Where: is the final behavior pattern representation of node i; is the feature vector of node i output after processing by the graph attention network; is the time series feature of node i output by the long short-term memory network; The malicious URL identification module calculates the node reconstruction error and local outlier factor LOF according to the node behavior pattern to obtain the anomaly score of the IP address. When the anomaly score of the IP address is higher than the dynamic anomaly detection threshold, the IP address is judged to be abnormal. The abnormal IP address is deeply analyzed in combination with the domain name characteristics of the URL, the web page content characteristics, and the network access behavior characteristics to identify malicious URLs. The dynamic anomaly detection threshold is specifically calculated as follows: ; Where: is the dynamic anomaly detection threshold at time t; is the mean of the anomaly scores within the historical time window; is the standard deviation during the same period; , is a dynamic weight coefficient, which is dynamically adjusted based on the historical score variance; An exponential moving average of historical anomaly scores.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method for identifying malicious URLs based on IP address feature analysis as described in any one of claims 1 to 4 is implemented.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for identifying malicious URLs based on IP address feature analysis as described in any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Fraud website identification method and device, equipment and storage medium
CN113779481A
Phishing domain name detection method based on registration information and analysis relation
CN117729034A
Bad flow monitoring method and system based on LSTM and clustering algorithm
CN119232489A