An illegal website gang recognition method and system based on multi-dimensional feature fusion

By capturing network traffic data in real time, extracting multi-dimensional features and performing fusion analysis, building machine learning models and network topology structures, the problem of inefficient identification of illegal website gangs in traditional methods is solved, and efficient and accurate detection and dynamic update of illegal website gangs is achieved.

CN119449481BActive Publication Date: 2025-07-22HARBIN INST OF TECH AT WEIHAI +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510013317.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-07-22
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

Traditional illegal website gang detection methods are inefficient and insufficiently accurate, making it difficult to cope with complex and hidden illegal website operating models, and the existing technology is difficult to achieve efficient and accurate identification and crackdown.

Method used

Through traffic mirroring technology and network probes, network traffic data is captured in real time, multi-dimensional features are extracted and fusion analysis is performed, machine learning models are built, illegal gang identification and network topology construction are constructed, and model updates are combined with abnormal detection.

Benefits of technology

It realizes efficient and accurate identification of illegal website gangs, improves the real-time and comprehensiveness of network security detection, can discover hidden illegal gang behaviors and relationships, and dynamically updates the model to adapt to changes in the network environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119449481B_ABST
    Figure CN119449481B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for identifying illegal website gangs based on multi-dimensional feature fusion, which relates to the field of network security. The method includes: capturing network traffic data in real time through traffic mirroring technology or network probes to collect passive traffic data; saving the collected data and extracting multi-dimensional features from the target website through active requests to achieve feature extraction; performing feature fusion and multi-dimensional analysis on the extracted multi-dimensional features to obtain fusion features and analysis results; training and optimizing a machine learning model according to the fusion features and analysis results to obtain a trained model; and identifying and associating new website data by using the trained model to construct the network topology of illegal gangs. The present invention realizes efficient and accurate identification of illegal website gangs, and improves the real-time performance and comprehensiveness of network security monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technology, and in particular to an illegal website group identification method and system based on multi-dimensional feature fusion. Background Art

[0002] In recent years, the activities of illegal website gangs have become increasingly rampant, especially in the fields of cybercrime and data theft, which have brought great harm to social security and economic development. Traditional detection methods mainly rely on rule matching and manual review, which are not only inefficient, but also have problems such as insufficient accuracy, slow response speed, and high false positive rate. It is difficult to cope with the increasingly complex and covert operation mode of illegal website gangs.

[0003] With the development of big data analysis and artificial intelligence technology, how to accurately identify and crack down on illegal website groups through automated and intelligent means, combined with multi-dimensional feature fusion and machine learning technology, has become a key issue that needs to be urgently addressed in the current network security field. Summary of the invention

[0004] The technical problem to be solved by the present invention is to provide a method and system for identifying illegal website groups based on multi-dimensional feature fusion. The present invention realizes efficient and accurate identification of illegal website groups and improves the real-time and comprehensiveness of network security detection.

[0005] In order to solve the above technical problems, the technical solution of the present invention is as follows:

[0006] In a first aspect, a method for identifying illegal website groups based on multi-dimensional feature fusion is provided, the method comprising:

[0007] Capture network traffic data in real time through traffic mirroring technology or network probes to collect passive traffic data;

[0008] The collected data is saved, and multi-dimensional features are extracted from the target website through active request to achieve feature extraction;

[0009] The extracted multi-dimensional features are subjected to feature fusion and multi-dimensional analysis to obtain fused features and analysis results;

[0010] According to the fusion features and analysis results, the machine learning model is trained and optimized to obtain a trained model;

[0011] By using the trained model to identify and associate new website data with illegal groups, the network topology of illegal groups can be constructed.

[0012] According to the network topology of the illegal gang, anomaly detection is performed on unknown illegal websites, and the model is updated based on the anomaly detection to identify the illegal website gang.

[0013] Preferably, network traffic data is captured in real time through traffic mirroring technology or network probes to collect passive traffic data, including:

[0014] Deploy traffic mirroring technology or network probes at the core nodes of the network environment;

[0015] Capture network traffic data in real time through traffic mirroring technology or network probes, and combine with preliminary filtering rules to collect passive traffic data.

[0016] Preferably, the collected data is saved, and multi-dimensional features are extracted from the target website through active requests to achieve feature extraction, including:

[0017] Store the collected data in structured and unstructured forms;

[0018] Extract multi-dimensional features from the target website through active requests to achieve feature extraction, including structural features, domain name and IP features extraction, traffic features and behavior pattern features.

[0019] Preferably, the extracted multi-dimensional features are subjected to feature fusion and multi-dimensional analysis to obtain fusion features and analysis results, including:

[0020] Perform data cleaning and fusion on the extracted multi-dimensional features to obtain fusion features;

[0021] Analyze the potential correlation between target websites through feature dimensionality reduction and data mining techniques to obtain analysis results.

[0022] Preferably, according to the fusion features and analysis results, the machine learning model is trained and optimized to obtain a trained model, including:

[0023] Construct a classification model based on the labeled data set;

[0024] According to the fusion features and analysis results, perform model training through machine learning algorithms to obtain a trained machine learning model;

[0025] Tune and validate the trained machine learning model to obtain a trained model.

[0026] Preferably, by using the trained model to identify and associate analyze new website data to construct the network topology of illegal groups, including:

[0027] Perform real-time analysis on new website data by using the trained model to predict its illegal probability to obtain a prediction result;

[0028] According to the prediction result, mine its correlation with known illegal groups to obtain a correlation analysis result;

[0029] Based on the relevance analysis result, construct the network topology of the illegal gang.

[0030] Preferably, based on the network topology of the illegal gang, perform anomaly detection on unknown illegal websites and update the model according to the anomaly detection to realize the identification of illegal website gangs, including:

[0031] Perform anomaly detection on unknown illegal websites through unsupervised learning algorithms to obtain detection results;

[0032] Update the sample data set according to the detection results and retrain the model to update the model;

[0033] Based on the updated model, realize the identification of illegal website gangs.

[0034] In a second aspect, an illegal website gang identification system based on multi-dimensional feature fusion includes:

[0035] An acquisition module for capturing network traffic data in real time through traffic mirroring technology or network probes to collect passive traffic data; extracting multi-dimensional features from the target website through active requests to achieve feature extraction; performing feature fusion and multi-dimensional analysis on the extracted multi-dimensional features to obtain fusion features and analysis results;

[0036] A processing module for training and optimizing a machine learning model according to the fusion features and analysis results to obtain a trained model; performing illegal gang identification and correlation analysis on new website data by using the trained model to construct the network topology of the illegal gang; performing anomaly detection on unknown illegal websites according to the network topology of the illegal gang and updating the model according to the anomaly detection to realize the identification of illegal website gangs.

[0037] In a third aspect, a computing device includes:

[0038] One or more processors;

[0039] A storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described above.

[0040] In a fourth aspect, a computer-readable storage medium stores a program, and when the program is executed by a processor, the method described above is implemented.

[0041] The above solution of the present invention has at least the following beneficial effects:

[0042] The present invention adopts multi-dimensional feature fusion technology and realizes the accurate identification of illegal websites through a machine learning model. Compared with traditional detection methods, it significantly improves the accuracy and efficiency of identification. By comprehensively analyzing multiple dimensions such as structural features, domain name features, traffic features, and behavior features, it can discover the illegal group behavior hidden in the complex network structure, reveal the potential connections and organizational patterns among group members. Through the anomaly detection and feedback mechanism, it continuously discovers unknown illegal websites and their associated information, dynamically updates the identification model, improves the coverage rate and robustness of the system, and ensures adaptation to the rapid changes in the network environment.

[0043] The illegal website gang network topology structure constructed by the present invention provides a panoramic view for regulatory agencies, helps accurately locate the core members and key nodes of illegal website gangs, and improves the ability to combat and manage illegal activities. It is not only applicable to the identification of illegal website gangs, but also can be applied to the detection of domain name-based network attack events such as botnets and malware. It is a practical technical solution applicable to complex network security scenarios. Brief Description of the Drawings

[0044] Figure 1 It is a flowchart showing a method for identifying an illegal website gang based on multi-dimensional feature fusion provided by an embodiment of the present invention.

[0045] Figure 2 It is a schematic diagram of a system for identifying an illegal website gang based on multi-dimensional feature fusion provided by an embodiment of the present invention.

[0046] Figure 3 It is a schematic diagram of the overall functional structure implemented by the present invention. Detailed Embodiment

[0047] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.

[0048] As Figure 1 shown, an embodiment of the present invention proposes a method for identifying an illegal website gang based on multi-dimensional feature fusion, and the method includes the following steps:

[0049] Step 11, capturing network traffic data in real time through traffic mirroring technology or network probes to collect passive traffic data;

[0050] Step 12, saving the collected data and extracting multi-dimensional features from the target website through an active request to achieve feature extraction;

[0051] Step 13: Perform feature fusion and multi-dimensional analysis on the extracted multi-dimensional features to obtain fused features and analysis results;

[0052] Step 14: Train and optimize the machine learning model based on the fused features and analysis results to obtain a trained model;

[0053] Step 15: Identify illegal groups and perform correlation analysis on new website data by using the trained model to construct the network topology of illegal groups;

[0054] Step 16: Perform anomaly detection on unknown illegal websites based on the network topology of illegal groups and update the model according to the anomaly detection to achieve the identification of illegal website groups.

[0055] In the embodiments of the present invention, network traffic data is captured in real time to ensure the timeliness and freshness of the data, providing the possibility for timely detection of the activities of illegal website groups; through traffic mirroring or network probes at the core nodes, all relevant traffic flowing through the nodes can be captured. Traffic mirroring and network probe technologies generally do not require direct access to or modification of the target websites, reducing the impact on the operation of the target websites and improving the security and reliability of data collection. By actively requesting and extracting multi-dimensional data such as structural features, domain name information, traffic features, and behavior pattern features, a rich information source is provided for the identification of illegal website groups; multi-dimensional feature extraction helps to deeply understand the behavior patterns and features of the target websites; the active request tool can automatically extract features, reducing manual intervention and improving the efficiency and accuracy of data processing. Feature fusion and multi-dimensional analysis can combine features from different dimensions to form a more comprehensive feature representation, helping to improve the accuracy of illegal website group identification; during the feature fusion process, data can be cleaned and standardized to eliminate noise and redundant information, improving data quality. Multi-dimensional analysis helps to discover patterns and correlation relationships hidden in the data, providing important clues for group identification. By training and optimizing machine learning models, the identification ability and robustness of the models for illegal website groups can be improved; the optimized models can better adapt to the changing network environment and the behavior patterns of illegal website groups; the trained models can automatically identify illegal groups for newly input website data, reducing manual intervention and improving processing efficiency. The trained models can identify illegal groups for newly input website data in real time to ensure timely discovery and response to illegal activities. Through association analysis, the association relationships between illegal website groups can be revealed, helping to locate the core members and key nodes of the groups. Based on the results of association analysis, the network topology structure of illegal groups is constructed, providing an intuitive network view for regulatory authorities to facilitate the understanding and response to illegal activities. By performing anomaly detection on unknown illegal websites, potential illegal activities can be discovered, improving the coverage rate and detection ability of the system; the model is dynamically updated according to the anomaly detection results to ensure that the model can adapt to the changing behavior patterns of illegal website groups and improve the accuracy and robustness of identification; by continuously updating the model and the dataset, the performance of the illegal website group identification system can be continuously improved, enhancing its effectiveness and value in practical applications.

[0056] In a preferred embodiment of the present invention, in step 11 above, network traffic data is captured in real time through traffic mirroring technology or network probes to collect passive traffic data, including:

[0057] Step 111, deploy traffic mirroring technology or network probes at the core nodes of the network environment;

[0058] Step 112, capture network traffic data in real time through traffic mirroring technology or network probes, and combine with preliminary filtering rules to collect passive traffic data.

[0059] In an embodiment of the present invention, deploying traffic mirroring technology or network probes at the core node can capture all network traffic flowing through this node, ensuring comprehensive monitoring of the target website and avoiding omission of key information; the core node is usually a convergence point of network traffic, and traffic mirrors or probes deployed here can efficiently collect a large amount of traffic data; traffic mirroring technology and network probes have flexible deployment methods, and appropriate deployment locations can be selected according to the network architecture and requirements to ensure the effectiveness and reliability of data collection. Capturing network traffic data in real time ensures the timeliness and freshness of the data, enabling timely discovery and handling of illegal website gang activities; combining with preliminary filtering rules, the captured traffic data is preliminarily screened and filtered to remove irrelevant or low-priority traffic; through the filtering rules, it is ensured that the collected passive traffic data is highly relevant to the illegal website gang identification task, improving the accuracy and pertinence of data analysis; the preliminary filtering rules help reduce unnecessary traffic data collection and processing, optimize system resource usage, and reduce operating costs.

[0060] In a specific embodiment of the present invention, step 11 is specifically implemented as follows: (1) Gateway / Switch (core node), this device is located at the core node of the network environment, responsible for transmitting data and mirroring traffic to the traffic collection module, and it is the entry point for passive traffic collection; (2) Passive traffic collection, this part captures the network traffic data of the target website in real time through traffic mirroring technology or network probes, and only collects the traffic related to the target website through preliminary filtering rules to ensure the accuracy and efficiency of subsequent analysis; (3) Active request tool (web crawler), this tool actively requests the target website, simulates user behavior and grabs the website page content, and extracts structural features and hyperlink information; (4) Active detection traffic, uses the active request tool to obtain the structural features of the website (such as page layout, hyperlink structure), DNS record data, and domain name information (such as registration time, operator attribution, etc.).

[0061] In another specific embodiment of the present invention, a traffic collection module is deployed at the core node of the network environment (such as a gateway, switch, or server outlet), and network traffic data packets passing through are captured in real time using traffic mirroring technology or network probes. The traffic collection module can be implemented based on software (such as Wireshark, Suricata, Zeek) or hardware devices; the target network traffic is copied to the traffic analysis device through a mirror port (PortMirroring); without affecting the normal operation of the target website, the meta-information of the transmitted data packets is obtained non-invasively; preliminary filtering rules are deployed at the traffic mirror port, which can reduce the introduction of irrelevant traffic in the first step of data collection and reduce the complexity of subsequent data processing. Among them, the preliminary filtering rules include: screening specific network protocol traffic through protocol type identifiers (such as protocol headers), for example, restricting traffic to common protocols such as HTTP, HTTPS, DNS, FTP, etc., and excluding irrelevant or low-priority protocols (such as IGMP, NetBIOS, etc.); screening according to the source IP address or target IP address of the data packet, for example, only capturing traffic from a specific subnet (such as 192.168.1.0 / 24) or only collecting traffic to the target server (such as the target IP is 10.0.0.1); screening based on the source port or target port number of the data packet.

[0062] In a preferred embodiment of the present invention, in step 12 above, the collected data is saved, and multi-dimensional features are extracted from the target website through an active request to achieve feature extraction, including:

[0063] Step 121, storing the collected data in structured and unstructured forms;

[0064] Step 122, extracting multi-dimensional features from the target website through an active request to achieve feature extraction, including structural features, domain name and IP features extraction, traffic features, and behavior pattern features.

[0065] In the embodiments of the present invention, the storage methods of structured and unstructured data can support the preservation of different types of data, including various formats such as text, images, videos, logs, etc., ensuring the integrity and richness of data information. Structured data is stored through relational databases or NoSQL databases, facilitating efficient querying, analysis, and management; unstructured data can be stored through file systems, object storage, etc., to meet the processing requirements of large amounts of data. Extracting multi-dimensional features can comprehensively cover all aspects of the target website, including page structure, domain name and IP information, traffic behavior, user interaction patterns, etc., providing comprehensive data support for the identification of illegal website groups; the extraction of multi-dimensional features helps to deeply mine the feature information of the target website, reveal its potential behavior patterns and correlation relationships, and improve the accuracy and reliability of the identification of illegal website groups; the extracted multi-dimensional features can be flexibly combined according to specific needs to construct different feature vectors to adapt to different identification tasks and analysis scenarios; rich feature data can provide more training information for machine learning models, helping to improve the classification performance, generalization ability, and robustness of the models, and enhancing the efficiency and effectiveness of the identification of illegal website groups.

[0066] In a specific embodiment of the present invention, step 12 specifically includes the following implementation steps:

[0067] Step 1, network data collection. By deploying web crawler technology, publicly available data of websites is crawled regularly or in real time, including web page content, HTML structure, domain name information, access records, etc. The collected data is stored in structured and unstructured forms. Let the set of original data collected be: , where each is the data of website i , including web page content, HTML structure, domain name information, etc. The network data collection module can combine multi-threaded crawling technology or a distributed crawler framework (such as Scrapy) to improve the collection efficiency, especially suitable for continuous monitoring of large-scale target websites. In addition, preferably, the collection period can be dynamically adjusted according to the update frequency of the target website to ensure the timeliness and integrity of the data.

[0068] Step 2, structural feature extraction. Use data parsing tools to parse the crawled web page content and extract information such as page content layout, page link structure, URL path features, etc. For example, use an HTML parsing library to extract the number of hyperlinks, tag nesting levels, resource loading paths, etc. in the web page, and parse the web page content and extract its set of structural features as: , where Represent features related to the number of hyperlinks, tag nesting levels, resource loading paths, etc. During the structural feature extraction process, web semantic analysis techniques can be combined to extract higher-level features, such as the distribution of page function modules and the complexity of the user interface. Meanwhile, preferably, the parsing efficiency can be improved through HTML content chunking techniques (such as DOM tree parsing), which is particularly applicable to the processing scenarios of complex web page structures.

[0069] Step 3, traffic feature extraction. Analyze the access behavior of the website from passive traffic, including user access frequency, access time interval, source IP distribution, etc. Through traffic collection tools (such as Wireshark or custom traffic analysis modules), obtain and organize relevant traffic data. The feature set is: , where each includes the accessed target, access IP, access frequency, etc. During the traffic feature extraction process, techniques for dividing time windows (such as sliding windows) can be combined to analyze short-term and long-term traffic behavior characteristics. In addition, preferably, a distributed traffic capture architecture can be adopted to achieve efficient collection and processing in a high-concurrency network environment, thereby supporting large-scale traffic feature analysis tasks.

[0070] Step 4, domain name and IP feature extraction. Extract the registrant information, registration time, registrar, etc. of the domain name by querying domain name registration information (such as WHOIS data); obtain the geographical location, affiliated operator, etc. of the IP through an IP address location analysis tool. The feature set is: . Where each includes registration time, geographical location, operator information, etc. During the domain name and IP feature extraction process, historical WHOIS data and IP geographical distribution change records can be combined to further analyze the evolution trend of the target website. In addition, preferably, the data accuracy and coverage can be improved by integrating multiple public domain name information databases (such as RDAP, ARIN, APNIC).

[0071] Step 5, behavior pattern feature extraction. Extract the operation frequency, user interaction pattern, abnormal behavior pattern, etc. of the website based on access logs and request-response data. Mark behaviors with multiple identical operations or batch requests as potential abnormal features. The feature set is: , where each includes repeated requests, batch access patterns, abnormal path requests, etc. During the behavior pattern feature extraction process, anomaly detection algorithms (such as Isolation Forest) can be combined to automatically mark high-frequency operation and batch request behaviors. In addition, preferably, for behavior patterns across time periods, time series analysis techniques (such as LSTM) can be introduced to predict potential abnormal operation trends.

[0072] In a preferred embodiment of the present invention, in step 13, the extracted multi-dimensional features are subjected to feature fusion and multi-dimensional analysis to obtain fused features and analysis results, including:

[0073] Step 131, by performing data cleaning and fusion on the extracted multi-dimensional features to obtain fused features;

[0074] Step 132, by means of feature dimensionality reduction and data mining techniques, analyzing the potential correlation between target websites to obtain analysis results.

[0075] In the embodiments of the present invention, the data cleaning process can remove noise data, invalid data, and outliers, improving the accuracy and reliability of the data; fusing features of different dimensions can form a more comprehensive feature representation, helping to capture complex associations and potential patterns between target websites and improving the accuracy of recognition; through feature fusion, redundant information between features can be reduced, the complexity of subsequent analysis can be lowered, and the computing efficiency can be improved; the fused features are more easily understood and utilized by machine learning models, helping to improve the classification and prediction performance of the models. Feature dimensionality reduction techniques can reduce the dimensionality of the data while retaining the main information, reducing the amount of computation and storage space and improving the analysis efficiency. Data mining techniques can mine hidden association rules, patterns, and information from a large amount of data, revealing the potential correlation between target websites and providing important clues for the identification of illegal website groups.

[0076] In a specific embodiment of the present invention, step 13 is specifically implemented as follows:

[0077] Step 1, feature cleaning and standardization. Clean the extracted multi-dimensional feature data, deleting noise data and invalid data (such as irrelevant page content or incorrect traffic records). Perform normalization processing on numerical features to convert the data to a unified dimension range. Automated data cleaning tools (such as Pandas, Dask) can be introduced during the feature cleaning process to improve the processing efficiency. In addition, preferably, the handling of outliers can be combined with statistical analysis methods (such as the Z-score method or the IQR method) to ensure more accurate standardization results of the data.

[0078] Step 2, feature fusion. Use feature engineering methods to fuse multi-dimensional features. Perform one-hot encoding on categorical features and perform bucketing or discretization on continuous features. Combine feature importance analysis to select key features to reduce redundancy. The specific method for performing feature fusion is:

[0079] (1) Categorical feature encoding

[0080] Traverse the categorical feature set For each categorical feature , perform one-hot encoding to convert each categorical variable into a binary vector. For example, if the input is "A","B" , the corresponding output is . Through this method, the encoded categorical feature matrix can be obtained.

[0081] (2) Encoding of continuous features

[0082] Traverse the set of continuous features . For each continuous feature , perform bucketing or discretization. Optional operations include equal-width bucketing and equal-frequency bucketing. Among them, equal-width bucketing divides the data into intervals of equal width: , and equal-frequency bucketing divides evenly into intervals, ensuring that the amount of data in each interval is roughly the same; discretize the processed continuous features into categorical variables. The bucketed continuous feature matrix is obtained.

[0083] (3) Feature importance analysis

[0084] Combine the processed feature set (including and ) into a preliminary feature matrix and perform feature importance analysis. By combining Lasso regression and decision tree feature importance analysis, comprehensively evaluate the features and screen out the key features. The specific method is as follows:

[0085] Lasso regression (L1 regularization): Perform Lasso regression on to obtain the weight of each feature. Screen out the features with weights greater than the threshold : ;

[0086] Decision tree feature importance: Use the decision tree model to calculate the importance score of each feature. Screen out the features with importance higher than the threshold : .

[0087] (4) Feature fusion

[0088] Combine the selected important features with the original multi-dimensional feature set for fusion, combine all processed features to obtain , and finally generate the unified feature vector .

[0089] (5)Embedded Feature Method (for Handling High-Dimensional Sparse Features)

[0090] For high-dimensional sparse features (such as categorical features after One-Hot encoding):

[0091] Use an embedding layer to map high-dimensional features to a low-dimensional space, where the embedded low-dimensional features are concatenated with other features (such as structural features, traffic features, domain name features, behavioral features) to form the final feature vector : , is the multi-dimensional feature after feature selection, is the low-dimensional representation of the sparse feature obtained by mapping through the embedding layer.

[0092] Step 3: Feature Dimensionality Reduction. Use dimensionality reduction methods such as principal component analysis (PCA) or linear discriminant analysis (LDA) to reduce the feature dimension, retain key features, and reduce computational complexity. Among them, , where W is the dimensionality reduction transformation matrix; during the process, a clustering dimensionality reduction method based on feature correlation (such as t-SNE, UMAP) can be combined to process non-linearly distributed data. In addition, preferably, the dimensionality reduction target dimension can be dynamically adjusted, and a suitable dimensionality reduction scheme can be selected according to the data scale and the need to balance computational complexity.

[0093] Step 4: Preliminary Analysis of Illegal Association Behaviors. For the fused feature data, use data mining algorithms (such as association rule mining) for analysis to find potential associations between websites, such as similar registrant information, common access sources, etc. At the same time, combine association rule mining in the time dimension with weighted association rule mining to achieve dynamic behavior pattern recognition and quantification of association strength. The following are the specific methods:

[0094] (1) Multi-Dimensional Association Rule Mining

[0095] Extract the fused feature data and use the association rule mining algorithm for analysis to identify high-correlation feature item sets between websites. By calculating the following indicators, filter out high-correlation rules:

[0096] Support: Measures the proportion of the feature item set appearing simultaneously in all samples, ;

[0097] Confidence: Measures the probability that the feature item appears simultaneously given that the feature item appears: ;

[0098] Lift: Measures the feature item and the strength of the association between: ;

[0099] By screening the rules that meet the minimum support, confidence, and lift thresholds, high-correlation rules are obtained.

[0100] (2) Dynamic association analysis in the time dimension

[0101] Divide the sample data into multiple time windows in chronological order: , where and represent the start time and end time of the time window respectively; within each time window, calculate the support and confidence independently, and extract the website association rules in the current time period: ; Compare the changes in rule strength in adjacent time windows to identify the dynamic evolution trend of the association rules: Support ; Based on the dynamic analysis in the time dimension, capture the behavior evolution pattern and time sensitivity characteristics of illegal website gangs.

[0102] (3) Weighted association rule modeling

[0103] Regarding the importance of different samples and features, introduce weights , and perform weighted modeling on the association rules to improve the accuracy of the rule results. Calculate the weighted support and weighted confidence respectively: Weighted support: , where is the weight of sample , is the indicator function. Weighted confidence: .

[0104] (4) Output of association rule results

[0105] Output high-correlation rules that meet the support, confidence, and lift requirements: ; Combine the time dimension to generate the dynamic evolution trend of website association rules: ; Output weighted association rules to quantify the association strength between websites: .

[0106] In a preferred embodiment of the present invention, in step 14 above, according to the fusion features and analysis results, train and optimize the machine learning model to obtain a trained model, including:

[0107] Step 141, construct a classification model based on the labeled data set;

[0108] Step 142: Based on the fusion features and analysis results, perform model training through a machine learning algorithm to obtain a trained machine learning model;

[0109] Step 143: Optimize and validate the trained machine learning model to obtain a well-trained model.

[0110] In the embodiment of the present invention, constructing a classification model based on a labeled data set can clarify the training objectives and classification criteria of the model, ensuring the applicability and effectiveness of the model in the task of identifying illegal website groups; the labeled data set provides rich sample data for model training, including known illegal website group samples and legal website samples, which helps the model learn effective classification rules; by constructing a classification model, the structure and parameters of the model can be initially determined. Using the fusion features and analysis results as the input for model training can make full use of the information of multi-dimensional features, improving the recognition ability and accuracy of the model; selecting a suitable machine learning algorithm for training, such as support vector machine, random forest, deep neural network, etc., can give full play to the advantages of the algorithm and improve the classification performance of the model; through the training process, the model can learn the association rules and feature patterns between the target website and the illegal website group. Through the optimization and validation process, the hyperparameters and structure of the model can be optimized, improving the classification accuracy and generalization ability of the model, making it better adapt to the actual application scenario; the validation process can evaluate the performance of the model on different data sets, ensuring the stability and reliability of the model, reducing the possibility of false positives and false negatives; the optimized model can better adapt to the changes in the network environment and the evolution of the behavior patterns of illegal website groups, maintaining long-term effectiveness and practicality.

[0111] In a specific embodiment of the present invention, step 14 is specifically implemented as follows:

[0112] Step 1: Sample data annotation. Construct a labeled data set through a combination of manual and automatic annotation, including known illegal website group samples and legal website samples. The illegal sample annotation is based on existing blacklists or expert analysis results. Define the data set as: ; In order to improve the annotation efficiency and accuracy, semi-supervised learning methods can be introduced to perform predictive annotation on unlabeled samples, and at the same time, use expert feedback to further correct the sample labeling results. In addition, preferably, the labeled data set can be continuously expanded by combining a dynamically updated illegal website blacklist or the latest detection results to improve the model's recognition ability for new illegal group behaviors.

[0113] Step 2: Model selection and construction. Select a machine learning model suitable for multi-classification problems, including but not limited to support vector machine, random forest, deep neural network, etc., and construct an initial model based on multi-dimensional features. Input feature set X and output the prediction result: ; In the initial model construction, a model structure suitable for high-dimensional data can be preferentially selected. For example, a deep neural network can combine hierarchical feature processing to improve the expression ability of complex feature interactions. Meanwhile, preferably, key features (such as domain name registration time, traffic pattern, abnormal behavior pattern) are selected through feature importance analysis to reduce the impact of redundant features on the model performance.

[0114] Step 3, model training: Divide the labeled dataset into a training set and a validation set, and use the training set to train the model. Improve the classification ability by optimizing the loss function of the model (such as cross-entropy loss). Introduce data augmentation techniques in model training, perform noise simulation, feature perturbation, etc. on the existing data, so as to improve the robustness of the model to noisy data and diverse illegal website behaviors. In addition, preferably, the model training process can be accelerated through a distributed training framework (such as TensorFlow, PyTorch) to support the processing of large-scale datasets.

[0115] Step 4, model optimization: Use grid search and cross-validation techniques to adjust the model hyperparameters. The hyperparameters to be adjusted include but are not limited to learning rate, tree depth, regularization coefficient, Dropout ratio, etc. The specific selection is based on the structure and characteristics of the target model. For example, for a deep neural network, it is preferable to adjust the learning rate and regularization coefficient to prevent overfitting; for a random forest model, it is preferable to adjust the tree depth and number to balance the model complexity and performance. In addition, preferably, the early stopping technique (EarlyStopping) can be combined during the optimization process to avoid unnecessary computational overhead.

[0116] Step 5, model evaluation: Use the validation set to evaluate the model performance. Preferably, use the validation set to evaluate the performance of the optimized model; among them, the evaluation metrics include but are not limited to accuracy, recall rate, F1 score, etc., and preferably ensure the applicability of the model in the illegal website gang identification task. The metrics can be weighted according to the task requirements during the evaluation. For example, a higher weight is given to the recall rate of illegal websites to reduce false negatives. In addition, preferably, the global performance of the model can be evaluated by combining the AUC curve, etc., to further quantify the classification ability of the model at different thresholds. To improve the accuracy of the evaluation, a stratified sampling method can also be adopted to ensure that the validation set is consistent with the actual scenario in terms of data distribution.

[0117] In a preferred embodiment of the present invention, in the above step 15, by using the trained model to perform illegal gang identification and association analysis on new website data to construct the network topology structure of illegal gangs, including:

[0118] Step 151, perform real-time analysis on the new website data by using the trained model, predict its illegal probability to obtain a prediction result;

[0119] Step 152: According to the prediction results, explore their relevance to known illegal groups to obtain the relevance analysis results;

[0120] Step 153: According to the relevance analysis results, construct the network topology of illegal groups.

[0121] In the embodiments of the present invention, the real-time prediction module of the present invention supports traffic analysis in a high-concurrency environment, can process a large amount of newly input website feature data while maintaining a low latency. At the same time, the system can perform hierarchical processing according to the probability threshold of the prediction results. For example, when the illegal probability exceeds the preset high-risk threshold, an alarm mechanism is immediately triggered, and the relevant data is pushed to the regulatory department. In addition, the present invention allows for further improving its accuracy and adaptability in real-time prediction by dynamically adjusting the parameters of the prediction model or introducing incremental learning algorithms. In the clustering analysis of the present invention, the calculation method of edge weights can be combined to quantify the strength of the association between websites. For example, weights are assigned to the association relationships based on the number of shared servers, the overlap degree of IP addresses, or the similarity of behavior patterns. In addition, preferably, by dynamically adjusting the clustering parameters (such as the number of clusters, the proximity threshold, etc.), it can adapt to illegal group networks of different scales, thereby improving the accuracy and robustness of the association analysis. Preferably, the clustering analysis method combined with the time dimension can track the evolution process of the behavior of illegal groups and provide support for the subsequent dynamic network construction. The network construction process of the present invention supports dynamic updates and can adjust the network structure in real-time as the illegal website data increases or the behavior characteristics change. For example, the gang relationship graph is updated by adding new nodes and edges to ensure that it can reflect the latest behavior patterns of illegal groups. In addition, preferably, the way of combining edge weights can be used to quantify the association strength between members (such as the number of times of sharing servers, the proportion of overlapping IP addresses, etc.), thereby highlighting the core nodes and key paths and providing more fine-grained information support for accurately cracking down on illegal groups.

[0122] In a specific embodiment of the present invention, step 15 is specifically implemented as follows:

[0123] Step 1: Real-time prediction. Deploy the trained model to the illegal website detection system for real-time analysis of newly input website data. The model makes predictions based on the input multi-dimensional features and outputs its illegal probability and relevance to known illegal groups. Among them, .

[0124] Step 2: Gang association analysis. Use clustering algorithms (such as K-Means or DBSCAN) to further analyze the predicted illegal websites and explore the connections between websites, such as whether they use the same server, share IP addresses, or have similar behavior patterns. Among them, , where represents the input feature set, C is the clustering result, including the associated group information among illegal websites.

[0125] Step 3, Illegal Gang Network Construction: Based on the results of gang association analysis, use the graph data structure to construct the network topology of illegal website gangs, and visually display the association relationships among gang members in the form of a graph. This network topology, through the representation of nodes and edges, takes illegal websites as nodes and the association between members as edges, helping to reveal the organizational form and core nodes of gang behavior, and further providing a panoramic view and decision-making basis for regulatory authorities.

[0126] In a preferred embodiment of the present invention, in the above step 16, according to the network topology of illegal gangs, perform anomaly detection on unknown illegal websites and update the model according to the anomaly detection to achieve the identification of illegal website gangs, including:

[0127] Step 161, perform anomaly detection on unknown illegal websites through an unsupervised learning algorithm to obtain the detection result;

[0128] Step 162, update the sample data set according to the detection result and retrain the model to update the model;

[0129] Step 163, according to the updated model, achieve the identification of illegal website gangs.

[0130] In the embodiments of the present invention, the unsupervised learning algorithm can independently discover abnormal behaviors or patterns in unknown illegal websites without relying on labeled data, expanding the scope and ability of identifying illegal website gangs; through anomaly detection, potential threats can be discovered in the initial stage of illegal website gang activities; the unsupervised learning algorithm has strong adaptability to unknown and changing data and can cope with the diversity and concealment of illegal website gang behaviors. Dynamically updating the sample data set according to the anomaly detection results and incorporating newly discovered illegal websites into the training data can continuously optimize the identification ability and accuracy of the model; retraining the model enables it to learn the latest characteristics and patterns of illegal website gangs, maintaining the effectiveness and timeliness of the model and adapting to changes in the network environment; through continuous updating and training, the model can accumulate more experience and knowledge, and has stronger robustness and identification ability for complex and concealed illegal website gang behaviors. The updated model has stronger identification ability, can more accurately identify illegal website gangs and their members, reducing false positives and false negatives; after the model is updated, its processing speed and efficiency are usually also improved, enabling it to analyze a large amount of data faster and respond to network security threats in a timely manner; through continuous updating and optimizing the model, comprehensive coverage and continuous monitoring of illegal website gang activities can be achieved.

[0131] In a specific embodiment of the present invention, step 16 is specifically implemented as follows:

[0132] Step 1: Abnormal mode detection. Monitor potential illegal websites in the network environment that are not marked through an abnormal detection module. Unsupervised learning algorithms (such as Isolation Forest or Autoencoder) are used for abnormal detection to mark behaviors that are significantly different from the normal mode. The specific process is as follows:

[0133] (1) Construct a normal time series benchmark template

[0134] Extract the time series data of normal website access behaviors , which is used as the normal behavior template. Time series with inconsistent lengths are preprocessed through interpolation or linear extension.

[0135] (2) Dynamic Time Warping (DTW) static alignment

[0136] According to the time series to be detected and the normal behavior template . Calculate the DTW distance. Use the DTW algorithm to calculate and between the minimum alignment path: , where is the alignment path, is the distance between two time points (such as Euclidean distance). Align to the same time axis as through the DTW path to obtain the aligned time series .

[0137] (3) RNN model modeling and dynamic prediction

[0138] Use the normal time series to train a Recurrent Neural Network (RNN) model to learn the time-dependent features of normal behaviors: , where : The behavior value at the moment predicted by the RNN, : The input of the aligned time series, : The RNN model parameters. Input the time series to be detected into the trained RNN model, and calculate the error between the actual value and the predicted value : ; Set the error threshold , and determine whether the moment is abnormal: ; Finally, calculate the abnormal score of the entire time series: , where II is an indicator function. If the anomaly score is less than , then mark this website as an abnormal behavior.

[0139] Step 2, feedback and model update. For the newly detected illegal website data, add it as a new sample to the dataset and retrain the model to improve the recognition ability. This feedback mechanism can ensure that the model adapts to the ever-changing network environment. Preferably, data screening rules (such as feature similarity or confidence threshold) can be incorporated into the feedback mechanism to automatically filter mislabeled data and ensure the quality of the added data samples. In addition, preferably, online learning algorithms (such as online gradient descent or incremental random forest) can be used to achieve real-time update of the model to reduce the time delay caused by retraining.

[0140] Step 3, alarm and response. Once suspected illegal website or illegal gang behavior is detected, immediately trigger the alarm mechanism and push the relevant information to the regulatory department or the network security team for further actions. Preferably, the alarm mechanism can perform hierarchical responses according to the risk level (such as high, medium, low). High-risk behaviors can trigger real-time alarms and start early warning procedures. Preferably, to improve the operability of the alarm information, visualization tools (such as real-time risk dashboards) can be combined to intuitively display the behavior characteristics of illegal websites and the network relationship of gangs to users.

[0141] An illegal website gang recognition system 20 based on multi-dimensional feature fusion, comprising:

[0142] An acquisition module 21, configured to capture network traffic data in real time through traffic mirroring technology or network probes to collect passive traffic data; extract multi-dimensional features from the target website through active requests to achieve feature extraction; perform feature fusion and multi-dimensional analysis on the extracted multi-dimensional features to obtain fusion features and analysis results;

[0143] A processing module 22, configured to train and optimize a machine learning model according to the fusion features and analysis results to obtain a trained model; perform illegal gang recognition and association analysis on new website data by using the trained model to construct the network topology structure of illegal gangs; perform anomaly detection on unknown illegal websites according to the network topology structure of illegal gangs and update the model according to the anomaly detection to achieve the recognition of illegal website gangs.

[0144] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. An illegal website gang identification method based on multi-dimensional feature fusion, characterized in that, The method includes: Real-time capturing network traffic data through traffic mirroring technology or network probes to collect passive traffic data; Saving the collected data and extracting multi-dimensional features from the target website through active requests to achieve feature extraction, including: storing the collected data in structured and unstructured forms; extracting multi-dimensional features from the target website through active requests to achieve feature extraction, including structural features, domain name and IP features, traffic features, and behavior pattern features; Performing feature fusion and multi-dimensional analysis on the extracted multi-dimensional features to obtain fusion features and analysis results, including: cleaning and fusing the extracted multi-dimensional features to obtain fusion features; analyzing the potential correlations between target websites through feature dimensionality reduction and data mining techniques to obtain analysis results; Training and optimizing a machine learning model based on the fusion features and analysis results to obtain a trained model; Identifying and associating illegal groups in new website data by using the trained model to construct the network topology of illegal groups; Performing anomaly detection on unknown illegal websites based on the network topology of illegal groups and updating the model according to the anomaly detection to achieve the identification of illegal website groups, including: performing anomaly detection on unknown illegal websites through unsupervised learning algorithms to obtain detection results; updating the sample data set and retraining the model according to the detection results to update the model; achieving the identification of illegal website groups based on the updated model.

2. The method for identifying illegal website groups based on multi-dimensional feature fusion according to claim 1, wherein Real-time capturing network traffic data through traffic mirroring technology or network probes to collect passive traffic data, including: Deploying traffic mirroring technology or network probes at the core nodes of the network environment; Real-time capturing network traffic data through traffic mirroring technology or network probes and combining preliminary filtering rules to collect passive traffic data.

3. The method for identifying an illegal website gang based on multi-dimensional feature fusion according to claim 2, wherein, Training and optimizing a machine learning model based on the fusion features and analysis results to obtain a trained model, including: Constructing a classification model based on a labeled data set; Training the model through machine learning algorithms according to the fusion features and analysis results to obtain a trained machine learning model; Tuning and validating the trained machine learning model to obtain a trained model.

4. The method for identifying illegal website groups based on multi-dimensional feature fusion according to claim 3, characterized in that Identifying and associating illegal groups in new website data by using the trained model to construct the network topology of illegal groups, including: Performing real-time analysis on new website data by using the trained model to predict its illegal probability to obtain a prediction result; Mining its correlation with known illegal groups according to the prediction result to obtain a correlation analysis result; Constructing the network topology of illegal groups according to the correlation analysis result.

5. An illegal website gang recognition system based on multi-dimensional feature fusion, characterized in that, Including: An acquisition module for real-time capturing network traffic data through traffic mirroring technology or network probes to collect passive traffic data; Extracting multi-dimensional features from the target website through active requests to achieve feature extraction; performing feature fusion and multi-dimensional analysis on the extracted multi-dimensional features to obtain fusion features and analysis results; A processing module, configured to train and optimize a machine learning model according to the fusion features and analysis results to obtain a trained model; identify and perform correlation analysis on illegal groups for new website data by using the trained model to construct a network topology of the illegal groups; perform anomaly detection on unknown illegal websites according to the network topology of the illegal groups, and update the model according to the anomaly detection to implement the identification of illegal website groups.

6. A computing device, characterized in that, Comprising: One or more processors; A storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the method according to any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, A program is stored in the computer-readable storage medium, and when the program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Malicious URL detection method based on sparse auto-encoder

    CN116318902A

  • Fraud website identification method and device, electronic equipment and storage medium

    CN117454374A