A method and apparatus for fraud-related entity clustering based on network traffic and metadata

By acquiring network traffic and metadata from multiple data sources, using tag propagation and clustering algorithms for precise labeling and clustering, and combining situational awareness technology, the problem of difficulty in identifying and tracking online fraud entities in existing technologies has been solved, achieving efficient and real-time monitoring and assessment of fraud activities.

CN119094190BActive Publication Date: 2025-12-30BEIJING FULE TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411190742.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-28
Publication Date
2025-12-30
Estimated Expiration
2044-08-28

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively identify and track online fraud entities, especially given the decreasing rate of domain name reuse and the frequent emergence of new domain names. Traditional methods are inefficient and prone to missing crucial information.

Method used

By acquiring network traffic and metadata from multiple data sources, using pre-defined label propagation and clustering algorithms for precise labeling and clustering analysis, and combining situational awareness technology for real-time monitoring and tracking, potential fraudulent activities can be identified and assessed.

Benefits of technology

It enables rapid identification and tracking of online fraud entities, improves analysis efficiency, reduces the cost of manual intervention, provides real-time monitoring and risk assessment capabilities, and adapts to changes in the online environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119094190B_ABST
    Figure CN119094190B_ABST
Patent Text Reader

Abstract

The application discloses a method and device for fraud-related entity clustering based on network traffic and metadata, and relates to the field of network security. In the method, traffic logs and metadata are obtained from multiple data sources, including CDNs, public clouds, private clouds and internal networks in target regions; traffic in the traffic logs is marked to distinguish CDN traffic, public cloud traffic, private cloud traffic, Whois privacy protection traffic, malicious traffic and business traffic through a preset label propagation algorithm; the marked data is subjected to clustering analysis to identify potential target network entities related to fraud through a preset clustering algorithm; the metadata is subjected to aggregated analysis to identify and track target network entities from the potential target network entities through a situational awareness technology, and risk assessment is performed. The technical solution provided by the application can accurately identify and track fraud-related network entities, and improves the network security protection capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of network security, specifically to a method and apparatus for clustering fraudulent entities based on network traffic and metadata. Background Technology

[0002] With the rapid development of network technology, network security issues have become increasingly prominent, especially the frequent occurrence of online fraud, which has caused huge economic losses to individuals and businesses.

[0003] Existing methods for combating and tracking fraudulent entities largely rely on manual aggregation and analysis of domain names. However, due to the rapid changes and one-time use nature of domain names, this method is not only inefficient but also prone to missing crucial information. Some solutions may employ basic data analysis and pattern matching to track fraudulent activities, but they typically lack the ability to handle large-scale data and the depth of understanding and processing of complex fraud patterns. More specifically, current technologies primarily rely on collecting data post-incidentally to create a "sewage pool" in an attempt to identify fraudulent information. However, given the decreasing reuse rate of domain names involved in cases and the continuous emergence of new domain names, especially since many domain names are used only once or for a single day, this method is no longer effective in collecting these short-lived domain names.

[0004] Therefore, how to provide a technical solution that can accurately identify and track fraudulent online entities has become an urgent problem for those skilled in the art. Summary of the Invention

[0005] This application provides a method and apparatus for clustering fraudulent entities based on network traffic and metadata. By performing cluster analysis on labeled data using a preset clustering algorithm, it can effectively identify potential target network entities involved in fraud and automatically discover network entities with similar characteristics or behavioral patterns, thereby quickly locating potential fraudulent activities. Combined with situational awareness technology, it aggregates and analyzes metadata to further identify and track specific target network entities from potential targets. This real-time monitoring and tracking capability enables cybersecurity teams to respond quickly and effectively intervene in and combat fraudulent activities.

[0006] The first aspect of this application provides a method for clustering fraudulent entities based on network traffic and metadata, applied to a fraudulent network entity clustering and tracking platform, the method comprising:

[0007] Traffic logs and metadata are obtained from multiple data sources, including CDN, public cloud, private cloud, and the internal network of the target region;

[0008] Traffic in the traffic logs is marked using a preset tag propagation algorithm to distinguish between CDN traffic, public cloud traffic, private cloud traffic, Whois privacy-protected traffic, malicious traffic, and business traffic.

[0009] The labeled data is clustered using a pre-defined clustering algorithm to identify potential target network entities involved in fraud.

[0010] The metadata is aggregated and analyzed using situational awareness technology to identify and track target network entities from the potential target network entities, and to conduct risk assessment.

[0011] By employing the aforementioned technical solutions, traffic logs and metadata are obtained from multiple data sources, including CDN, public cloud, private cloud, and the target region's internal network, ensuring the comprehensiveness and diversity of the data. This multi-source data fusion approach can more comprehensively reflect the true situation of network traffic and metadata, providing a solid foundation for subsequent clustering analysis and situational awareness. Through a pre-defined label propagation algorithm, traffic in the traffic logs is accurately labeled, effectively distinguishing between CDN traffic, public cloud traffic, private cloud traffic, Whois privacy-protected traffic, malicious traffic, and business traffic. This refined traffic classification helps to more accurately identify fraudulent network entities during subsequent clustering analysis. A pre-defined clustering algorithm is used to perform clustering analysis on the labeled data to identify potential fraudulent target network entities. The clustering algorithm can automatically discover potential patterns and structures in the data, grouping data points with similar characteristics into one category, thereby effectively identifying fraudulent network entities. This method not only improves analysis efficiency but also reduces the cost of manual intervention. Situational awareness technology is used to aggregate and analyze metadata, further identifying and tracking target network entities from potential target network entities and conducting risk assessments. Situational awareness technology can monitor changes in network status in real time, enabling timely detection and response to potential security threats. Furthermore, by tracking target network entities, it can provide in-depth understanding of their behavioral patterns and characteristics, offering strong support for subsequent strike and defense operations.

[0012] Optionally, the step of marking the traffic in the traffic log using a preset label propagation algorithm includes:

[0013] Initial labels are set for nodes of known categories based on configuration information. The known categories of nodes include CDN nodes, public cloud nodes, and Whois privacy-protected nodes.

[0014] The final label of each node is determined based on the connection relationships between nodes, including traffic or network connections recorded in the traffic log;

[0015] Traffic is labeled based on the node labels, traffic characteristics, and traffic behavior patterns of the nodes it passes through. The traffic characteristics include protocol type, port number, and request URL, and the traffic behavior patterns include access frequency and request time distribution.

[0016] By adopting the above technical solution, initial labels are set for nodes of known categories (such as CDN nodes, public cloud nodes, and Whois privacy-protected nodes), ensuring the accuracy of label propagation. These nodes serve as "anchors," providing a reliable benchmark for label allocation across the entire network traffic. The final label for each node is determined based on the connection relationships between nodes (such as traffic or network connections recorded in traffic logs), fully considering the actual flow of network traffic and the interactions between nodes. Through the propagation algorithm, nodes with unknown labels can infer their own labels based on information from connected nodes with known labels, thereby improving the accuracy of label allocation. When labeling traffic, not only are basic traffic characteristics (such as protocol type, port number, and request URL) considered, but also traffic behavior patterns (such as access frequency and request time distribution). This multi-dimensional approach makes traffic labeling more refined, more accurately reflecting the true attributes and potential risks of traffic. By introducing traffic behavior patterns (such as access frequency and request time distribution), a deeper understanding of traffic behavior characteristics can be achieved. These behavior patterns are often closely related to fraudulent activities; therefore, by analyzing and labeling them, potential fraudulent activities can be identified more effectively. By employing a pre-defined label propagation algorithm, network traffic can be automatically labeled, reducing the cost and error rate of manual intervention. This automated label allocation method improves the efficiency and accuracy of network traffic analysis. Combining traffic characteristics with traffic behavior patterns for labeling enables the system to intelligently identify potential fraudulent activities. This intelligent identification capability provides strong support for cybersecurity teams, helping them to promptly detect and respond to online fraud activities.

[0017] Optionally, the step of performing cluster analysis on the labeled data using a preset clustering algorithm to identify potential target network entities involved in fraud includes:

[0018] Based on the requirements for identifying fraudulent activities, target features are selected, and the target features are preprocessed and normalized. The target features include domain age, SSL certificate information, IP location similarity, domain Hamming distance, IP service and port information.

[0019] Based on the characteristics of the fraudulent behavior, the neighborhood radius and minimum number of points of the preset clustering algorithm are set, and the labeled data are clustered according to the preset clustering algorithm and the target features to generate clustering results;

[0020] Based on the clustering results, the characteristics and commonalities of different clusters are analyzed to identify potential target network entities involved in fraud.

[0021] By adopting the above technical solution, a series of features related to target network entities were selected based on the identification needs of fraudulent activities, such as domain age, SSL certificate information, IP location similarity, domain Hamming distance, IP service and port information, etc. The selection of these features is based on a deep understanding of online fraud and can comprehensively reflect the potential risks of network entities. By preprocessing and normalizing these target features, the dimensional differences between different features were eliminated, improving the accuracy and reliability of clustering analysis. Based on the characteristics of fraudulent activities, the neighborhood radius and minimum number of points of the preset clustering algorithm were customized. This customization makes the clustering algorithm more adaptable to fraudulent activity identification scenarios, enabling it to more accurately discover network entities with similar characteristics and behavioral patterns. The preset clustering algorithm automatically performs clustering analysis on the labeled data, generating clustering results without manual intervention. This automated processing method greatly improves the efficiency of fraudulent activity identification and reduces labor costs. Based on the clustering results, the system can automatically analyze the characteristics and commonalities of different clusters, thereby identifying potential target network entities involved in fraud. This intelligent identification method enables the system to promptly detect and respond to potential fraud threats, improving network security defense capabilities. Cluster analysis considers multiple dimensions of features, such as domain names, SSL certificates, and IP locations, providing a comprehensive basis for risk assessment. By comprehensively analyzing these features, the potential risks of network entities can be assessed more accurately.

[0022] Optionally, the step of performing cluster analysis on the labeled data according to the preset clustering algorithm and the target features to generate clustering results includes:

[0023] The labeled data are grouped according to the neighborhood radius, the minimum number of points, and the target feature.

[0024] Traverse all data points in the marked data. When the number of data points within the radius of the target data point is greater than or equal to the minimum number of points, set the target data point as the core point.

[0025] Starting from the core point, adjacent data points are connected into a cluster based on density reachability, and clustering results are generated based on the clusters.

[0026] By employing the above technical solution and using density-based clustering algorithms (such as DBSCAN), the density of clusters is controlled by two parameters: neighborhood radius and minimum number of points. This method can effectively identify clusters of arbitrary shapes, thereby improving the accuracy of cluster analysis. During the traversal of data points, the number of data points within the neighborhood radius is used to determine whether a point is a core point. This determination method ensures that only areas with sufficiently high density can be identified as the core of a cluster, further improving the accuracy of clustering. Starting from the core points, adjacent data points are connected into a cluster based on density reachability. This connection method considers not only the spatial distance between data points but also the density characteristics of the data points, making the clustering results more stable and reliable. Even if there are noisy or outlier points in the dataset, their impact on the clustering results can be avoided through density reachability determination. The neighborhood radius and minimum number of points, as key parameters of the clustering algorithm, can be flexibly adjusted according to specific fraud detection needs. This configurability allows the clustering algorithm to be applied to datasets of different sizes and characteristics, improving the applicability and flexibility of the algorithm. Clustering algorithms can efficiently process large-scale datasets by traversing data points and connecting them based on density reachability. Furthermore, because the identification of core points is relatively simple and direct, the computational complexity of the entire clustering process is low, enabling the generation of clustering results in a short time.

[0027] Optionally, the step of aggregating and analyzing the metadata using situational awareness technology to identify and track target network entities from the potential target network entities includes:

[0028] Data aggregation technology is used to aggregate preprocessed metadata in order to reveal potential fraud patterns and trends;

[0029] Natural language processing and semantic analysis techniques are used to understand the aggregated metadata in order to reveal the logical relationships between the data. Through association reasoning algorithms, the connection between different fraudulent behaviors and online assets is determined.

[0030] Using time series analysis and deep learning models, we predict the trends of fraudulent activities and the risks to online assets, and create a fraud trend prediction chart.

[0031] By employing the aforementioned technical solutions, data aggregation technology effectively integrates preprocessed metadata, thereby revealing potential fraud patterns and trends. Data aggregation not only helps identify anomalies in individual data points but also reveals commonalities and patterns in fraudulent activities across time, space, and behavioral modes through pattern matching and trend analysis. This helps security teams better understand the nature and evolution of fraudulent activities. Natural language processing and semantic analysis technologies are used to deeply understand and analyze the aggregated metadata. These technologies can identify keywords, phrases, and sentences in the data and understand the logical relationships between them. Through association reasoning algorithms, the inherent connections between different fraudulent activities and online assets can be further determined, providing strong support for subsequent tracking and response. Based on the results of semantic understanding and association reasoning, security teams can more accurately determine the nature and scope of impact of fraudulent activities, thereby formulating more precise and effective response strategies. This helps improve the intelligence level of security decision-making and reduce the risk of human misjudgment and omissions. Through the application of time series analysis and deep learning models, the trends of fraudulent activities and the risks of online assets can be accurately predicted. These models can capture subtle changes and potential patterns in the data and predict future development trends accordingly. By creating fraud trend prediction maps, security teams can intuitively understand the evolving trends and potential risks of fraudulent activities, providing strong guidance for future security prevention efforts.

[0032] Optionally, the use of data aggregation technology to aggregate the preprocessed metadata to reveal potential fraud patterns and trends includes:

[0033] The types of fraudulent activities are used as keywords for aggregation. By statistically analyzing the distribution of each type in the overall fraudulent activities, the most frequent types of fraudulent activities can be identified.

[0034] Using the geographical location information of fraudulent activities as the aggregation keywords, a heat map or distribution map of fraudulent activities is drawn by counting the number of fraudulent activities in each geographical location.

[0035] Using the time information of fraudulent acts as the aggregation keyword, a timeline chart of fraudulent acts is drawn by statistically analyzing the number of fraudulent acts that occur in each time period.

[0036] By cross-combining the analysis results of a single dimension, a multi-dimensional cross-aggregated view is formed, where the single dimension includes type, location, and time.

[0037] By employing the aforementioned technical solutions and using fraud type as the aggregation keyword, the distribution of each type within the overall fraud activity can be statistically analyzed to visually identify the most frequent fraud types. This precise identification helps security teams prioritize limited resources for preventing and combating high-incidence fraud, improving prevention efficiency. Using geographic location information as the aggregation keyword, heatmaps or distribution maps of fraud activities can be created to visually demonstrate their distribution across different geographical areas. This visualization helps security teams quickly locate high-risk areas and take appropriate preventative measures to reduce fraud occurrences. Using time information as the aggregation keyword, the number of fraud activities occurring within each time period can be statistically analyzed and timeline graphs can be drawn to clearly show the temporal distribution patterns of fraud activities. This analysis helps security teams predict peak periods for future fraud and prepare in advance, effectively curbing the spread of fraud. By cross-combining the results of single-dimensional analyses to form a multi-dimensional cross-aggregated view, a more comprehensive understanding of the complexity and diversity of fraud activities can be achieved. This multi-dimensional analysis not only helps to reveal the intrinsic connections and mutual influences between fraudulent activities, but also provides security teams with more accurate decision support, enabling them to develop more comprehensive and effective prevention strategies.

[0038] Optionally, the step of employing natural language processing and semantic analysis techniques to understand the aggregated metadata in order to reveal the logical relationships between the data, and determining the connection between different fraudulent behaviors and online assets through association reasoning algorithms, includes:

[0039] Entities are extracted from aggregated metadata using named entity recognition technology. These entities include the name of the fraud method, the type of victim, the IP address, and the domain name.

[0040] By using relation extraction technology, the relationships between the entities are identified and extracted. Combined with contextual information, semantic understanding is performed on the extracted entities and the relationships to reveal the underlying logical relationships.

[0041] Different data points are associated according to preset rules, and the associated data points are represented in the form of a network graph, where nodes represent entities and edges represent the relationships between entities.

[0042] By using association reasoning algorithms, path search and influence analysis are performed in the network graph to reveal the connections between different fraudulent activities and online assets.

[0043] By employing the aforementioned technical solutions, key entities, such as fraud method names, victim types, IP addresses, and domain names, can be accurately extracted from massive and complex metadata. These entities are fundamental to understanding the relationship between fraudulent activities and online assets; accurate extraction helps improve the accuracy and efficiency of subsequent analysis. Relationship extraction technology can identify and extract relationships between entities and combine this with contextual information for semantic understanding. This not only reveals direct, explicit relationships between entities but also uncovers implicit, deep-seated logical relationships through semantic reasoning. This in-depth understanding helps to more comprehensively grasp the complex relationship network between fraudulent activities and online assets. Representing related data points in the form of a network graph, where nodes represent entities and edges represent relationships between entities, makes the complex relationship network of fraudulent activities and online assets intuitive and easy to understand. This helps security teams quickly locate key nodes and relationship paths, providing strong support for subsequent analysis and decision-making. Path searching and influence analysis within the network graph can efficiently reveal the connections between different fraudulent activities and online assets. Through association reasoning algorithms, it is possible to identify how fraudulent activities spread and proliferate through online assets; this information is crucial for developing targeted prevention strategies.

[0044] A second aspect of this application provides a system for clustering fraudulent entities based on network traffic and metadata, including a collection module, a tagging module, a clustering module, and an execution module, wherein:

[0045] The acquisition module is configured to acquire traffic logs and metadata from multiple data sources, including CDN, public cloud, private cloud, and the internal network of the target area.

[0046] The tagging module is configured to tag traffic in the traffic logs using a preset tag propagation algorithm to distinguish between CDN traffic, public cloud traffic, private cloud traffic, Whois privacy-protected traffic, malicious traffic, and business traffic.

[0047] The clustering module is configured to perform cluster analysis on labeled data using a preset clustering algorithm to identify potential target network entities involved in fraud.

[0048] The execution module is configured to perform aggregate analysis of the metadata using situational awareness technology to identify and track target network entities from the potential target network entities and to conduct risk assessment.

[0049] A third aspect of this application provides an electronic device including a processor, a memory, a user interface, and a network interface, wherein the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any of the foregoing.

[0050] A fourth aspect of this application provides a computer-readable storage medium storing instructions that, when executed, perform the method described in any of the preceding descriptions.

[0051] In summary, one or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:

[0052] 1. Traffic logs and metadata are obtained from multiple data sources, including CDN, public cloud, private cloud, and internal networks of the target region. This wide range and diversity of data sources ensures the comprehensiveness and accuracy of the data, providing a solid foundation for subsequent analysis.

[0053] 2. By using a pre-defined tag propagation algorithm to label traffic in traffic logs, CDN traffic, public cloud traffic, private cloud traffic, Whois privacy-protected traffic, malicious traffic, and business traffic are effectively distinguished. This refined classification not only improves the accuracy of data analysis but also provides clear boundary conditions for subsequent cluster analysis.

[0054] 3. By using a pre-defined clustering algorithm to perform cluster analysis on the labeled data, potential fraud-related network entities can be automatically identified. This data-driven method is more efficient and accurate than traditional manual analysis, and can quickly locate fraud targets, providing strong support for subsequent processing.

[0055] 4. By aggregating and analyzing metadata through situational awareness technology, not only are potential fraud patterns and trends revealed, but also the logical relationships between data are further understood through natural language processing and semantic analysis, time series analysis and deep learning models. This helps the security team to unravel complex data, accurately identify and track target network entities.

[0056] 5. After identifying and tracking the target network entity, this method also conducts a risk assessment. By comprehensively considering multiple factors (such as the type of fraudulent behavior, frequency of occurrence, and scope of impact), it can accurately assess the risk level of the target network entity, providing a scientific basis for subsequent response measures.

[0057] 6. It has the ability to process network traffic and metadata in real time, enabling it to promptly detect and respond to potential fraudulent activities. Furthermore, by dynamically adjusting the parameters of clustering algorithms and situational awareness technologies, it can adapt to changes in the network environment and maintain high efficiency and accuracy. Attached Figure Description

[0058] Figure 1This is a flowchart illustrating the method for clustering fraudulent entities based on network traffic and metadata disclosed in an embodiment of this application.

[0059] Figure 2 This is a schematic diagram of the overall architecture of the method for clustering fraudulent entities based on network traffic and metadata disclosed in the embodiments of this application;

[0060] Figure 3 This is a schematic diagram of the modules of the system for clustering fraudulent entities based on network traffic and metadata disclosed in the embodiments of this application;

[0061] Figure 4 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of this application.

[0062] Explanation of reference numerals in the attached diagram: 301, acquisition module; 302, tagging module; 303, clustering module; 304, execution module; 401, processor; 402, communication bus; 403, user interface; 404, network interface; 405, memory. Detailed Implementation

[0063] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0064] In the description of the embodiments in this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Specifically, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a concrete manner.

[0065] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0066] This embodiment discloses a method for clustering fraudulent entities based on network traffic and metadata, which is applied to a fraudulent network entity clustering and tracking platform. Figure 1This is a flowchart illustrating the method for clustering fraudulent entities based on network traffic and metadata disclosed in an embodiment of this application. Figure 1 As shown, the method includes the following steps:

[0067] S110. Obtain traffic logs and metadata from multiple data sources, including CDN, public cloud, private cloud, and the internal network of the target area;

[0068] Figure 2 This is a schematic diagram of the overall architecture of the method for clustering fraudulent entities based on network traffic and metadata disclosed in the embodiments of this application, combined with... Figure 2 The embodiments of this application will be described. For example... Figure 2 As shown, after data collection, a metadata set and collected tag data are constructed, and then the data is stored and processed.

[0069] A CDN (Content Delivery Network) is an intelligent virtual network built on top of an existing network. A public cloud refers to cloud computing services owned, managed, and operated by a third-party cloud service provider, while a private cloud allows an enterprise or organization to own and manage its own computing resources and services. Deploy network monitoring tools (such as Wireshark, tcpdump, Snort, etc.) within the target region's internal network to capture real-time network traffic. Design and implement a distributed web crawler to scrape data from data sources that support web access (CDN, public cloud, private cloud, and the target region's internal network). Use libraries such as Scrapy and BeautifulSoup to parse web pages and extract the required information. For data sources that provide API interfaces, write API client programs to periodically call the API and scrape data. Store the scraped data in a distributed file system (such as HDFS or Ceph) for subsequent processing and analysis. Additionally, NoSQL databases (such as Cassandra or MongoDB) can be used to store structured or semi-structured data. Utilize message queues (such as Kafka and RabbitMQ) and stream processing frameworks (such as Apache Flink and Apache Spark Streaming) to process newly crawled data in real time and update information in the metadata engine. For time-sensitive metadata, design a mechanism to periodically record historical snapshots for retrospective analysis and auditing. Integrate data from different data sources into a unified metadata set, including PassiveDNS, ICP, Domain, Host, Whois, SSL Certificate, IP ASN, IP Location, and other information. Collect data on tagged fraud cases from law enforcement agencies, security organizations, etc., including information such as fraudulent hosts, domains, and fraud types. For IPs and Domains involved in fraud, record relevant metadata snapshots for in-depth analysis and tracking. Use encryption technology during data transmission and storage to ensure data security and privacy. Ensure that the data crawling and processing process complies with relevant laws and regulations, such as GDPR and CCPA.

[0070] S120. The traffic in the traffic log is marked by a preset tag propagation algorithm to distinguish CDN traffic, public cloud traffic, private cloud traffic, Whois privacy protection traffic, malicious traffic and business traffic.

[0071] like Figure 2As shown, the data processing and security steps include hashing or encrypting sensitive information, tagging and classifying traffic using a label propagation algorithm, and cluster analysis. After collection, the data is processed by a dedicated cleaning engine. This engine is responsible for identifying and removing invalid, duplicate, or malformed data, ensuring the quality of data for subsequent analysis. Sensitive information (such as personally identifiable information and financial information) is hashed or encrypted during the cleaning process to prevent data leakage and misuse. This processing method preserves data usability (for analysis and statistics) while ensuring data privacy. To better understand and utilize traffic data, a label propagation algorithm is introduced. This algorithm can automatically tag data into different categories based on its characteristics and attributes, such as CDN traffic, public cloud traffic, private cloud traffic, Whois privacy-protected traffic, malicious traffic, and business traffic. For CDN traffic, the algorithm identifies IP addresses accelerated by the CDN and tags them accordingly. The distinction between public cloud traffic and private cloud traffic is based on information such as IP address, port number, and protocol type. This distinction helps enterprises understand their resource usage in different cloud environments and make corresponding resource allocation and cost optimizations. Whois privacy-protected data refers to domain data that hides registrant information through the Whois service. The algorithm identifies these domains and tags them with privacy protection labels to safeguard the registrant's privacy rights. Identifying malicious traffic is a crucial aspect of network security. The algorithm combines various characteristics (such as request frequency, request content, and source IP address) to determine whether traffic is malicious and tags it accordingly. This helps enterprises promptly detect and prevent malicious attacks, ensuring network security. Business traffic refers to traffic generated by normal business activities. The algorithm considers any remaining traffic not explicitly classified into any of the above categories as business traffic and tags it accordingly. Through this tag propagation algorithm, inappropriate clustering of unrelated network assets can be effectively avoided. This is because each traffic data point is given a clear label, allowing data analysts to clearly understand the attributes and affiliation of each data point. Cluster optimization not only improves the accuracy and efficiency of data analysis but also provides strong support for enterprises to develop targeted network strategies. For example, enterprises can optimize their network architecture and resource allocation based on the distribution of CDN traffic and public cloud traffic; and strengthen network security protection based on the results of malicious traffic identification.

[0072] Optionally, the step of marking the traffic in the traffic log using a preset label propagation algorithm includes:

[0073] Initial labels are set for nodes of known categories based on configuration information. The known categories of nodes include CDN nodes, public cloud nodes, and Whois privacy-protected nodes.

[0074] The final label of each node is determined based on the connection relationships between nodes, including traffic or network connections recorded in the traffic log;

[0075] Traffic is labeled based on the node labels, traffic characteristics, and traffic behavior patterns of the nodes it passes through. The traffic characteristics include protocol type, port number, and request URL, and the traffic behavior patterns include access frequency and request time distribution.

[0076] Based on known configuration information or prior knowledge, initial labels are assigned to nodes of specific categories. These known node categories include, but are not limited to, CDN nodes, public cloud nodes, and Whois privacy-protected nodes. These initial labels are the starting point for the label propagation algorithm, providing the foundation for subsequent propagation processes. The algorithm analyzes the connectivity relationships between nodes, primarily derived from traffic or network connections recorded in traffic logs. By considering interactions and traffic flow between nodes, the algorithm infers potential associations and similarities between nodes. Based on these connectivity relationships, the algorithm iteratively updates the label of each node. In each iteration, a node updates its own label based on the labels of its neighboring nodes and its own characteristics. This process continues until the labels of all nodes tend to stabilize (i.e., converge) or the preset maximum number of iterations is reached. After the label propagation process is complete, the algorithm further labels traffic based on the final labels of each node and the traffic characteristics and behavioral patterns in the traffic logs. Traffic characteristics include, but are not limited to, protocol type, port number, and request URL, while traffic behavior patterns may include access frequency, request time distribution, etc.

[0077] The specific markings are as follows:

[0078] CDN traffic: Based on the tags and traffic characteristics of CDN nodes (such as CDN IP address, common request headers, cache hits, etc.), traffic belonging to the CDN acceleration category is marked.

[0079] Public cloud traffic: By identifying the tags and traffic characteristics of public cloud nodes (such as public cloud IP address ranges, public cloud endpoints, API calls, etc.), traffic originating from the public cloud is marked.

[0080] Private cloud traffic: Similarly, traffic from the enterprise private cloud environment is tagged based on the labels and characteristics of the private cloud nodes (such as private cloud IP address ranges, internal endpoints, enterprise-specific APIs, etc.).

[0081] Whois Privacy Protection Traffic: Identifies traffic related to domain addresses that use the Whois privacy protection service.

[0082] Malicious Traffic: By combining external threat intelligence and traffic characteristics (such as abnormal access patterns, malware signatures, etc.), potential malicious IPs and suspicious traffic are flagged.

[0083] Normal business traffic: For traffic that has been clearly identified as legitimate and normal business traffic, maintain its whitelist status and conduct further analysis as needed.

[0084] Intranet traffic and VPN traffic: Based on the intranet configuration and VPN protocol characteristics, legitimate traffic and VPN traffic in the intranet are identified.

[0085] By employing a pre-defined label propagation algorithm, initial labels can be set based on nodes of known categories (such as CDN nodes, public cloud nodes, and Whois privacy-protected nodes), and the labels are iteratively updated using the connection relationships between nodes (such as traffic or network connections recorded in traffic logs) until convergence. This method is more intelligent and accurate than traditional rule-based or simple statistical classification methods, and can better capture the inherent characteristics and correlations of traffic data. Accurate traffic classification provides a solid foundation for subsequent data analysis. By labeling different categories of traffic, it is easier to identify critical traffic, abnormal traffic, and potential security threats in the network. This helps network administrators better understand the network status, formulate targeted network policies, optimize network resource allocation, and improve network performance and security. In addition to basic traffic classification, this application also combines traffic characteristics and traffic behavior patterns (such as protocol type, port number, request URL, access frequency, request time distribution, etc.) to further label traffic. This refined labeling method enables network administrators to implement more precise and effective traffic control and management measures, such as bandwidth allocation, access control, and security policies, based on the characteristics and needs of different traffic types.

[0086] S130. Perform cluster analysis on the labeled data using a preset clustering algorithm to identify potential target network entities involved in fraud;

[0087] like Figure 2As shown, clustering analysis includes selecting parameters for the clustering algorithm, processing noise, assigning feature weights, and feature engineering. After the data in the traffic logs is effectively labeled and classified through a preset label propagation algorithm, a preset clustering algorithm is used to perform in-depth clustering analysis on these labeled data. The main purpose of this step is to identify potential target network entities that may be involved in fraud, attacks, or other malicious activities. Clustering analysis is an unsupervised learning technique that can divide labeled traffic data into several groups or clusters, making objects within the same cluster similar to each other, while objects in different clusters are quite different. In this embodiment, clustering can help identify network entities with abnormal behavior patterns that may be related to fraudulent activities. Choosing a suitable clustering algorithm is crucial for identifying potential target network entities. Commonly used clustering algorithms include K-means, DBSCAN, and hierarchical clustering. Each algorithm has its own characteristics and applicable scenarios. For example, the K-means algorithm is suitable for discovering spherical clusters, but the number of clusters needs to be specified in advance; the DBSCAN algorithm does not require the number of clusters to be specified in advance and can identify clusters of arbitrary shapes, but it is more sensitive to the selection of parameters. The preset clustering algorithm in the embodiments of this application may include any of the above-mentioned clustering algorithms.

[0088] Optionally, the step of performing cluster analysis on the labeled data using a preset clustering algorithm to identify potential target network entities involved in fraud includes:

[0089] Based on the requirements for identifying fraudulent activities, target features are selected, and the target features are preprocessed and normalized. The target features include domain age, SSL certificate information, IP location similarity, domain Hamming distance, IP service and port information.

[0090] Based on the characteristics of the fraudulent behavior, the neighborhood radius and minimum number of points of the preset clustering algorithm are set, and the labeled data are clustered according to the preset clustering algorithm and the target features to generate clustering results;

[0091] Based on the clustering results, the characteristics and commonalities of different clusters are analyzed to identify potential target network entities involved in fraud.

[0092] This application uses the DBSCAN clustering algorithm for illustration. Several key features were selected based on the identification requirements of fraudulent activities, including domain age, SSL certificate information, IP location similarity, domain Hamming distance, IP service and port information, etc. These features reflect the typical characteristics and behavioral patterns of fraudulent network assets. For example, fraudulent websites typically have short domain lifespans; phishing websites and legitimate websites usually have significant differences in SSL certificates; fraudulent websites often use multiple IPs from the same geographical location; malicious domains typically have higher information entropy than normal domains; and fraudulent websites often open non-standard ports. Necessary preprocessing was performed on the selected features, such as numericalizing text information (hashing or word embedding) and transforming geographical location information to ensure data consistency and comparability. Simultaneously, numerical features were normalized to eliminate the influence of different units on the clustering results. Based on the characteristics of fraudulent network assets, such as domain age and registration information, an appropriate neighborhood radius (Eps) value was selected to define the neighborhood size of a point. This value needs to be determined based on experiments and data analysis to ensure effective clustering of fraudulent websites. Based on the complexity of the fraudulent network assets, a suitable minimum point count (MinPts) value is selected to define the minimum number of points required for a dense region. This value selection also needs to be based on experiments and data analysis to ensure that there are a sufficient number of points within the dense region to represent a genuine fraudulent network entity. By testing and setting reasonable Eps and MinPts parameters, any point whose average distance from all points within the dense region exceeds 0.7 is marked as noise, identifying noise points unrelated to the fraudulent network assets. Noise is further filtered using density-based Local Outlier Factor (LOF) and domain name reuse rate-based threshold filtering to handle points identified as noise. The DBSCAN clustering algorithm is used to perform cluster analysis on the marked data based on a preset neighborhood radius and minimum point count. The DBSCAN algorithm can identify clusters of arbitrary shapes based on density, making it well-suited for identifying complex fraud patterns. After cluster analysis, clustering results are generated, grouping similar network assets together to form different clusters. Based on the clustering results, the characteristics and commonalities of different clusters are analyzed. By comparing the similarities and differences between different clusters, potential target network entities with fraudulent behavior characteristics can be identified. Based on the analysis results, clusters representing potential target network entities involved in fraud were identified. These entities may share similar characteristics such as domain names, registration information, page content, or behavioral patterns. Dimensionality reduction techniques and visualization tools (such as PCA and t-SNE) were used to visualize the clustering results, providing a more intuitive understanding of the distribution and structure of fraudulent network assets. Simultaneously, tools such as SHAP (Shapley Additive Explanations) were used to interpret the clustering results, revealing the common characteristics and differences among fraudulent network assets.Based on cluster analysis and combined with actual needs, fraudulent websites can be tracked quickly and accurately. Case linkage can be automated and intelligently carried out based on cluster results, improving the efficiency and accuracy of combating fraud.

[0093] Examples of suspicious traffic scenarios: (1) Centralized release of new domain names, characterized by multiple different domain names pointing to the same server at a certain time, for example, a large-scale release of new domain names pointing to the same CNAME; (2) Geographical clustering, characterized by fraudulent activities originating from a specific geographical area or IP address range, for example, an IP address from a certain country or region initiates an active request communication; (3) Similar network assets, characterized by similar domain names, registration information, and page content; (4) Frequent access to specific resources, characterized by high-frequency access to specific resource files or pages; (5) Periodic requests, characterized by repeated requests or behaviors within a fixed time period.

[0094] Data characteristics of centralized new domain name releases include: domain name, target IP address or CNAME, and record addition time. Domain name, target IP / CNAME, and timestamp can be used as features for DBSCAN clustering to identify situations where newly released domain names point to the same server. Data characteristics of geographic location clusters include: source IP address and geographic location information. Geographic location information of IP addresses can be used for DBSCAN clustering to identify fraudulent activities concentrated in specific geographic areas. Data characteristics of similar online assets include: domain name, registration information, and numerical representation (hash or word embedding) of page content text information. These hash values ​​can be used as features for DBSCAN clustering to identify online assets with similar registration information or page content. Data characteristics of frequent access to specific resources include: access frequency and URL path (hash). Access frequency and URL path hash values ​​can be used for DBSCAN clustering to identify patterns of frequent access to specific resources. Data characteristics of periodic requests include: timestamp and access type. Timestamp and access type features can be used for DBSCAN clustering to identify patterns of periodic requests.

[0095] Different distance metrics are applied in different scenarios. For example, when clustering geographically related clusters, distance metrics related to IP geographical location are preferred. Cosine similarity is used to quantify the similarity between SSL certificates. Different weights are assigned to different features based on the importance of the fraudulent network assets. Based on actual test results, domain age, IP location, and SSL certificates are assigned relatively high feature weights, while other features are assigned relatively low weights.

[0096] By selecting features highly correlated with target fraudulent activities (such as domain age, SSL certificate information, IP location similarity, domain Hamming distance, IP service and port information, etc.), cluster analysis can focus on those features most likely to reveal fraudulent patterns, thereby improving identification accuracy. Cluster analysis is an automated data processing technique capable of rapidly processing large amounts of data and discovering hidden patterns and structures. Compared to manual review, this method significantly improves processing speed and efficiency, especially suitable for cybersecurity applications requiring rapid response. By adjusting the parameters of the preset clustering algorithm (such as neighborhood radius and minimum number of points), it can flexibly adapt to the characteristics of different types of fraudulent activities, making the method widely applicable. Furthermore, as new types of fraudulent methods emerge, the system's effectiveness and timeliness can be maintained by updating the feature set and algorithm parameters. Cluster analysis not only helps identify known fraudulent activities but also reveals potential fraud threats by discovering anomalous clusters or unclassified entities in the data. This is of great significance for preventing future fraudulent activities.

[0097] Optionally, the step of performing cluster analysis on the labeled data according to the preset clustering algorithm and the target features to generate clustering results includes:

[0098] The labeled data are grouped according to the neighborhood radius, the minimum number of points, and the target feature.

[0099] Traverse all data points in the marked data. When the number of data points within the radius of the target data point is greater than or equal to the minimum number of points, set the target data point as the core point.

[0100] Starting from the core point, adjacent data points are connected into a cluster based on density reachability, and clustering results are generated based on the clusters.

[0101] All data points are initially marked as "unvisited". This ensures that each data point is processed only once during clustering, avoiding redundant calculations. A randomly selected "unvisited" point from the marked data is chosen as the starting point. This step randomizes the start of the clustering process, which is helpful for handling datasets with multiple dense regions. For the selected point, all points within its neighborhood are calculated based on a preset neighborhood radius. This typically involves calculating the distance between the selected point and all other points, and identifying the set of points whose distance is less than or equal to the neighborhood radius. If the number of data points in the neighborhood of the selected point is greater than or equal to a preset minimum number of points (MinPts), the point is marked as a "core point" and considered sufficient to form the starting point of a new cluster. If the number of data points in the neighborhood is less than MinPts, the point is marked as a "noise point". These points may be located in sparse regions of the dataset or are too low in density compared to other clusters to form a separate cluster. However, it is worth noting that in some implementations, these noise points may be re-evaluated in subsequent iterations and may be added to other clusters, although in the original DBSCAN algorithm they are typically kept as noise. Starting from each core, adjacent data points are connected into clusters based on density reachability (i.e., reaching another core from one core through a series of density-connected cores). This process involves recursively checking the points in the neighborhood of each core and adding them to the corresponding cluster until no more points can be added. This process continues traversing the data points until all points have been visited and classified into their respective clusters (as cores, noise points, or part of a cluster). Based on this process, one or more clusters, along with a possible set of noise points, are ultimately generated. These clusters and noise points constitute the final result of the clustering analysis.

[0102] By utilizing neighborhood radius and minimum number of points as clustering criteria, this process can identify clusters of arbitrary shapes and sizes in the data. During clustering, points with fewer data points in their neighborhood than the minimum number of points are marked as noise or outliers. This approach helps reduce the impact of noise on the clustering results, making them more accurate and reliable. Simultaneously, since noise points are clearly distinguished, users can perform further analysis or processing as needed. By connecting adjacent data points into a cluster through density reachability, this embodiment ensures that data points within a cluster are density-connected. This continuity not only helps form compact and meaningful clusters but also aids in understanding the intrinsic relationships between data points. Because this embodiment clusters based on local density, it is robust to local variations or noise in the data. Furthermore, as the dataset size increases, this embodiment maintains good clustering performance and exhibits scalability. In addition to generating clusters, this embodiment also provides information on noise points. This information is crucial for further data analysis, anomaly detection, or data cleaning tasks.

[0103] S140. The metadata is aggregated and analyzed using situational awareness technology to identify and track target network entities from the potential target network entities, and a risk assessment is performed.

[0104] like Figure 2As shown, situational awareness and tracking includes observation-level metadata aggregation and intelligent parsing, understanding-level semantic analysis and correlation reasoning, prediction-level trend prediction and risk assessment, and a multi-dimensional interactive situational awareness dashboard. The main purpose of situational awareness is to observe elements in the environment within a certain time and space, understand the meaning of these elements, and predict their state in the near future. Situational awareness consists of three levels: observation, understanding, and prediction, and its output is directly fed into the decision-making and action cycle. Observation Layer: This layer involves the operator's sensory detection of significant information about the system they are operating and the environment in which it operates. For example, the operator needs to be able to see relevant displays or hear alarm signals. Basic situational awareness includes the perception of the status of elements such as various fraud types, fraudulent websites, fraudulent system nodes, fraudulent network asset clusters, relevant historical records, and relevant historical IP addresses. Data on websites, domains, IP addresses, traffic, and related information involved in fraud is collected through web crawlers and API interfaces. Understanding Layer: Situational awareness goes beyond simply observing the data displayed on the computer screen; it needs to understand the meaning or significance of this information in conjunction with the operator's objectives. In this process, the information gathered is gradually integrated to develop a comprehensive picture of the system, thereby forming a more complete overall understanding of the events unfolding. In terms of situational understanding, the focus is on the scenarios in which fraudulent activities are implemented, the detection characteristics of fraudulent activities, which fraudulent activities may be interconnected, and the correct prioritization of competing events. Machine learning algorithms, including decision trees and support vector machines (SVM), are used to classify and label the collected data. The prediction layer involves forward temporal inference of information to determine how it will affect the future state of the operating environment. This combines an understanding of the current situation with a mental model of the system, enabling prediction of possible next steps. Specifically, based on time series analysis and recurrent neural networks (RNNs), it predicts new fraudulent network deployment methods, predicts changes from legitimate assets to fraudulent network assets, and records changes in the clustering of the same assets over time.

[0105] Using data analysis tools and methods from situational awareness technology, preprocessed metadata is aggregated and analyzed. This includes, but is not limited to, correlation analysis (identifying relationships between different data sources), cluster analysis (grouping similar data points), and time series analysis (analyzing data trends over time). These analytical methods allow valuable information and patterns to be extracted from massive amounts of data. Based on the results of the aggregation analysis, key features related to target network entities are extracted. These features may include network behavior patterns, abnormal traffic characteristics, and user behavior characteristics. The extracted features are matched against a known target network entity feature database to identify potential target network entities. If a match is successful, the entity's activity trajectory and behavior patterns are further tracked. Identified target network entities are continuously monitored, and their latest activity data and metadata are collected to adjust analysis strategies and risk assessment results in a timely manner. Based on the target network entity's behavior patterns and characteristics, the potential threat type and severity are assessed. This includes assessing its attack intent, attack methods, and attack scope. System or network vulnerabilities that the target network entity may exploit are analyzed, and the likelihood of a successful attack is assessed. A risk assessment report is generated by combining the results of the threat assessment and vulnerability assessment. The report should include a detailed description of the target network entity, the type of threat, the vulnerability, the risk assessment level, and recommended countermeasures.

[0106] Optionally, the step of aggregating and analyzing the metadata using situational awareness technology to identify and track target network entities from the potential target network entities includes:

[0107] Data aggregation technology is used to aggregate preprocessed metadata in order to reveal potential fraud patterns and trends;

[0108] Natural language processing and semantic analysis techniques are used to understand the aggregated metadata in order to reveal the logical relationships between the data. Through association reasoning algorithms, the connection between different fraudulent behaviors and online assets is determined.

[0109] Using time series analysis and deep learning models, we predict the trends of fraudulent activities and the risks to online assets, and create a fraud trend prediction chart.

[0110] Data aggregation technology is used to summarize and integrate preprocessed metadata (such as log files, network traffic data, user behavior records, etc.) to extract valuable information from large amounts of scattered data. Specifically, the raw data undergoes cleaning, deduplication, and formatting to ensure quality and consistency. Data aggregation algorithms (such as hash tables and database aggregation queries) are used to merge similar or related data items, forming datasets that are easier to analyze. Aggregated data can reveal potential fraud patterns (such as frequent login failures and abnormal fund flows) and trends, providing a foundation for subsequent analysis. Natural language processing and semantic analysis techniques are used to deeply understand the aggregated metadata content, revealing the logical relationships between data and thus discovering potential clues to fraud. Specifically, text data (such as user comments and chat logs) is converted into a computer-understandable format, keywords and phrases are extracted, the meaning and context of the text data are analyzed, potential semantic patterns or anomalies are identified, and association reasoning algorithms (such as graph theory algorithms and clustering algorithms) are used to perform association analysis on data from different sources (such as user behavior, network traffic, and system logs) to determine the inherent connections between different fraudulent behaviors and online assets. By employing time series analysis and deep learning models, this study predicts trends in fraudulent activities and risks to online assets, providing a basis for developing effective defense strategies. Specifically, it analyzes time-varying patterns in historical data to identify periodic and seasonal characteristics of fraudulent activities. Deep learning algorithms (such as Recurrent Neural Networks (RNNs) and Long Short-Term Memory Networks (LSTMs)) are then used to train the model on a large dataset, learning complex patterns of fraudulent activities. Based on the trained model, future trends in fraudulent activities are predicted, and fraud trend prediction charts are generated to visually demonstrate changes in risk.

[0111] Data aggregation technology enables the integration of large amounts of scattered metadata into more meaningful and highly correlated datasets. This integration helps reveal patterns and trends of fraudulent behavior hidden behind complex data, significantly improving the accuracy of fraud identification. Combining natural language processing and semantic analysis techniques, the system can deeply understand the aggregated metadata content and uncover the logical relationships between data points. This not only helps identify direct clues to fraudulent activities but also, through association reasoning algorithms, discovers indirect but important connections between different fraudulent activities and online assets, constructing a more comprehensive network graph of fraudulent activities. Using time series analysis and deep learning models, future trends of fraudulent activities can be predicted. These models can learn patterns from historical data and infer future trends accordingly. Simultaneously, creating fraud trend prediction maps provides decision-makers with intuitive and visual references, helping them to formulate timely response strategies and reduce potential risks. By continuously monitoring and analyzing metadata in the network environment, the system can promptly detect and track target network entities, including potential fraudsters. This real-time monitoring and tracking capability significantly enhances network security and defense capabilities, protecting network assets from harm.

[0112] Optionally, the use of data aggregation technology to aggregate the preprocessed metadata to reveal potential fraud patterns and trends includes:

[0113] The types of fraudulent activities are used as keywords for aggregation. By statistically analyzing the distribution of each type in the overall fraudulent activities, the most frequent types of fraudulent activities can be identified.

[0114] Using the geographical location information of fraudulent activities as the aggregation keywords, a heat map or distribution map of fraudulent activities is drawn by counting the number of fraudulent activities in each geographical location.

[0115] Using the time information of fraudulent acts as the aggregation keyword, a timeline chart of fraudulent acts is drawn by statistically analyzing the number of fraudulent acts that occur in each time period.

[0116] By cross-combining the analysis results of a single dimension, a multi-dimensional cross-aggregated view is formed, where the single dimension includes type, location, and time.

[0117] This analysis uses the type of fraud as the aggregation keyword. This includes various fraud methods such as phishing, telephone fraud, and identity theft. By statistically analyzing the distribution of each type within the overall fraud activity, the most frequent types of fraud can be identified. This analysis helps determine current fraud hotspots and key targets for prevention. The statistical results are displayed in chart form, such as a fraud type distribution map, visually reflecting the proportion and distribution of each type of fraud. Another aggregation keyword is the geographical location of the fraud. This includes the IP address of the fraud perpetrator and the geographical location of the domain name resolution. By counting the number of frauds occurring at each geographical location, a heat map or distribution map of fraud can be created. This map display method can visually reveal the regional characteristics and concentrated areas of fraud. Geographical location-based aggregation results can guide relevant departments to strengthen prevention and crackdown efforts in specific areas. Finally, another aggregation keyword is the time information of the fraud. This includes the timestamp and duration of the fraud. By statistically analyzing the number of frauds occurring within each time period, a timeline of fraud can be created. This timeline display method can reveal the periodic, seasonal, or sudden characteristics of fraud activities. Based on time-series aggregation results, predictive models can be used to forecast future trends in fraudulent activities, allowing for proactive preventative measures. Cross-combining the analysis results from the aforementioned single dimensions (type, location, time) creates a multi-dimensional cross-aggregated view. This cross-aggregation reveals the inherent connections and mutual influences between different dimensions. By comprehensively analyzing the results of multi-dimensional cross-aggregation, a more comprehensive understanding of the complexity and diversity of fraudulent activities can be achieved. For example, the distribution of specific types of fraud across different geographical locations and time periods can be analyzed, or the evolutionary trends of different types of fraud within the same geographical location can be examined. The results of multi-dimensional cross-aggregation provide strong data support for developing more precise and effective anti-fraud strategies.

[0118] By using fraud type information as aggregation keywords, the system can accurately statistically analyze the distribution of various fraud types within the overall picture, quickly identifying the most frequent fraud types. This precise identification helps security teams prioritize and address high-risk fraud activities. Using geographic location information as aggregation keywords, combined with heatmaps or distribution maps, the system can visually reflect the geographical distribution characteristics of fraud activities. This map visualization technology allows security personnel to quickly locate hotspots of fraud activities, providing strong support for subsequent on-site investigations and crackdowns. Using time information as aggregation keywords, the system can statistically analyze the number of fraud activities within each time period and create timeline graphs. This time series analysis helps reveal the periodic, seasonal, or sudden characteristics of fraud activities, providing a temporal reference for developing targeted prevention measures. By cross-combining the analysis results of single dimensions (type, location, time) to form a multi-dimensional cross-aggregated view, the system can reveal the inherent connections and mutual influences between different dimensions. This comprehensive analysis helps security teams gain a more comprehensive understanding of the complexity and diversity of fraud activities, thereby developing more comprehensive and effective response strategies.

[0119] Optionally, the step of employing natural language processing and semantic analysis techniques to understand the aggregated metadata in order to reveal the logical relationships between the data, and determining the connection between different fraudulent behaviors and online assets through association reasoning algorithms, includes:

[0120] Entities are extracted from aggregated metadata using named entity recognition technology. These entities include the name of the fraud method, the type of victim, the IP address, and the domain name.

[0121] By using relation extraction technology, the relationships between the entities are identified and extracted. Combined with contextual information, semantic understanding is performed on the extracted entities and the relationships to reveal the underlying logical relationships.

[0122] Different data points are associated according to preset rules, and the associated data points are represented in the form of a network graph, where nodes represent entities and edges represent the relationships between entities.

[0123] By using association reasoning algorithms, path search and influence analysis are performed in the network graph to reveal the connections between different fraudulent activities and online assets.

[0124] Named Entity Recognition (NER) technology extracts key entities from aggregated metadata. These entities are fundamental to understanding the relationships between fraudulent activities and online assets, including but not limited to fraud names, victim types, IP addresses, and domain names. NER helps the system quickly locate and identify key information in text, providing foundational data for subsequent relation extraction and semantic understanding. Relation extraction technology identifies and extracts the relationships between these entities. This process involves not only simple entity-to-relation mapping but also requires in-depth semantic understanding of the extracted entities and relationships by incorporating contextual information. Relation extraction and semantic understanding can reveal complex logical relationships between entities, such as the association between a specific IP address and multiple fraudulent domain names, or the correspondence between a certain fraudulent method and a specific victim type. Based on preset rules or algorithms, different data points are associated, and these associated data points are represented in the form of a network graph. In the network graph, nodes represent entities (such as IP addresses, domain names, fraud methods, etc.), and edges represent relationships between entities (such as association, similarity, influence, etc.). Network graph representations visually illustrate the complex relationships between data, allowing operators to clearly see the inherent connections between different fraudulent activities and online assets. Path searching and influence analysis are performed within the network graph using association reasoning algorithms. These algorithms identify key nodes and paths in the network, revealing the potential connections and mutual influences between different fraudulent activities and online assets. Association reasoning algorithms not only help the system discover new fraud patterns and trends but also provide a scientific basis for developing targeted preventative measures. By updating the network graph and association reasoning results in real time, the system can maintain a keen awareness of fraudulent activities and a rapid response capability.

[0125] The following are the implementation details of specific levels of situational awareness:

[0126] (1) Observation-level metadata aggregation and intelligent analysis: This level adopts advanced data aggregation technology to deeply aggregate the collected metadata in order to reveal potential fraud patterns and trends. By utilizing machine learning and big data analysis, it analyzes the dynamic changes of various fraud network assets in real time, as well as their correlation with environmental factors, providing richer and more accurate data support for subsequent understanding and prediction levels.

[0127] Fraud Type Distribution Map: This map displays the distribution of different types of fraud, such as the geographical location of the fraudsters and the domain names used. Specifically, it collects information such as the geographical location, domain name, and type of fraud involved in fraud incidents, aggregates the data by type and geographical location, and uses a geographic map (such as a heat map) to display the distribution of different types of fraud.

[0128] Fraudulent Network Asset Clustering Graph: This graph uses maps or network diagrams to illustrate the clustering of fraudulent network assets, such as the relationships between IP addresses and domain names. Specifically, it collects IP addresses, domain names, fraudulent web pages, fraudulent applications, fraud types, and their relationships, and uses network diagrams (such as NetwrokX) to display the connections between IP addresses, domain names, fraudulent web pages, applications, and fraud types.

[0129] Historical Timeline: Displays the historical changes of assets in a fraudulent network, such as domain registration and usage. Specifically, it collects information on domain registration, changes, usage, and timestamps, arranges the data chronologically, and displays the historical record using a timeline chart.

[0130] Alerts and Notifications Dashboard: Displays real-time alerts and notifications related to fraudulent online assets, such as newly discovered fraudulent domains and abnormal traffic. Specifically, it collects real-time alert and notification data, such as abnormal traffic, suspicious domains, IP addresses, and types of fraud. The alert data is then filtered, categorized, and prioritized (e.g., by severity of fraud type), and displayed in real-time using an alert dashboard (such as Dash).

[0131] (2) Understanding Hierarchical Semantic Analysis and Association Reasoning: This level employs natural language processing and semantic analysis techniques to deeply understand the aggregated metadata, revealing the underlying meanings and logical relationships behind the data. Through association reasoning algorithms, it explores the intrinsic connections between different fraudulent behaviors and online assets, providing a scientific basis for identifying and preventing new types of fraud.

[0132] Fraud Behavior Analysis Chart: This chart uses graphs and diagrams to illustrate the characteristics and patterns of fraud behavior, such as common fraud methods and target victims. Specifically, it collects data on fraud behavior, such as methods, target victims, frequency, and amount of money defrauded. The data is aggregated according to fraud methods and target victim types, and then displayed using bar charts, pie charts, and other graphs to show the distribution of common fraud methods and target victims.

[0133] Association Analysis Graph: This graph displays the connections and similarities between different fraudulent activities, such as fraudulent domains using the same IP address or registration information. Specifically, it collects data containing association information, such as IP addresses, registration information, and domain names, and associates them based on the same IP address or registration information, using a network graph to illustrate the associations.

[0134] Event Priority Chart: This chart displays the priority of fraud events through lists or graphs, helping operators determine the order and strategy of response. Specifically, it collects data including fraud amount and fraud type that contains priority information, sorts them according to priority, and displays the priority ranking using lists or bar charts.

[0135] (3) Forecasting level trend prediction and risk assessment: This level uses time series analysis and deep learning models to accurately predict future fraud trends and network asset risks. Combined with a multi-dimensional risk assessment model, it assesses the risk level of various network assets in real time, providing forward-looking decision support for fraud prevention.

[0136] Fraud Trend Prediction Chart: This chart uses time series analysis to illustrate future trends of fraudulent online assets, such as predicting potential future fraudulent domain names. Specifically, it collects time series data on fraudulent domain names, including registration time and access frequency. The time series data is processed, including missing value imputation and smoothing. Time series models such as ARIMA and Prophet are used for trend prediction, and a time series trend chart is generated to display historical and predicted data. The periodicity, seasonality, and interval distribution of the data can be used for further in-depth analysis.

[0137] Risk Assessment Chart: By analyzing the characteristics and historical records of fraudulent network assets, this chart displays the risk levels of different assets, such as risk scores based on historical fraudulent behavior. Specifically, it collects the characteristics of each network asset (such as registration information, associated IPs, access frequency, etc.) and historical fraud records, calculates the risk score for each asset according to predefined scoring rules or models, and draws a risk score chart to display the risk levels of different network assets.

[0138] Change Correlation Graph: Records and visualizes the clustering changes of the same assets over time. If the changes are consistent, further correlation information is displayed between the two. Specifically, it collects the clustering results of the same assets at different points in time, records the clusters to which the assets belong at different points in time, and uses a change graph to show the clustering changes of the same assets. If the changes are consistent, it further displays detailed correlation information between the two.

[0139] (4) Multidimensional interactive situational awareness dashboard: integrates the analysis results of the three levels of observation, understanding and prediction, and provides operators with comprehensive and multidimensional situational awareness through an interactive visualization interface. Users can customize the view and layout according to their needs, adjust the analysis parameters in real time, and explore various information of fraudulent network assets in depth to achieve more flexible and accurate situational awareness.

[0140] Integrating visualizations across observation, understanding, and prediction: A comprehensive dashboard integrates visualizations from all three levels, providing holistic situational awareness. Operators can customize views as needed, such as selecting displayed charts and adjusting layouts. Interactive features are provided, such as clicking on a fraudulent domain to view detailed information or adjusting the time range to view historical records.

[0141] To create a comprehensive dashboard integrating observation, understanding, and prediction, and providing customizable views and interactive features, modern visualization frameworks such as Plotly Dash, Streamlit, or Bokeh can be used. Users can customize the view as needed by selecting different chart types via drop-down menus. Users can click on a category in the fraud behavior analysis chart to view detailed information, with the specific data clicked displayed in a window.

[0142] In addition to the DBSCAN clustering algorithm, other clustering algorithms can also be used in this application embodiment to achieve automated tracking and clustering of fraudulent network assets. For example, the K-Means clustering algorithm: by selecting an appropriate K value, fraudulent network assets can be divided into K clusters, and then the optimal cluster centers can be found through iterative optimization; hierarchical clustering: by gradually merging or splitting clusters, a hierarchical clustering tree can be formed, thereby better understanding the relationships between fraudulent network assets.

[0143] This application's embodiments can use deep learning methods, such as autoencoders or deep clustering networks, to learn complex patterns in fraudulent network assets and achieve more accurate clustering. Autoencoders: Low-dimensional representations of fraudulent network assets are learned by training an autoencoder, and then these representations are used for clustering. Deep clustering networks: Cluster assignments of fraudulent network assets are directly learned through end-to-end training.

[0144] This application's embodiments can use graph analysis methods to represent and analyze relationships between fraudulent network assets. For example, social network analysis: representing fraudulent network assets as nodes in a graph, representing the relationships between them as edges, and then using social network analysis methods to identify important nodes and communities; graph embedding: using graph embedding methods to learn low-dimensional representations of fraudulent network assets, and then using these representations for clustering.

[0145] The embodiments of this application can achieve more robust and flexible tracking and clustering of fraud network assets by combining multiple methods. For example, ensemble learning: by combining the results of multiple clustering algorithms, more robust and accurate clustering can be achieved; multi-view clustering: by integrating information from different views or sources, more comprehensive and accurate clustering can be achieved.

[0146] This embodiment also discloses a system for clustering fraudulent entities based on network traffic and metadata. Figure 3 This is a schematic diagram of the modules of the system for clustering fraudulent entities based on network traffic and metadata disclosed in an embodiment of this application, as shown below. Figure 3 As shown, the system includes a data acquisition module 301, a labeling module 302, a clustering module 303, and an execution module 304, wherein:

[0147] The acquisition module 301 is configured to acquire traffic logs and metadata from multiple data sources, including CDN, public cloud, private cloud and internal network of the target area;

[0148] The tagging module 302 is configured to tag the traffic in the traffic log using a preset tag propagation algorithm to distinguish between CDN traffic, public cloud traffic, private cloud traffic, Whois privacy protection traffic, malicious traffic, and business traffic.

[0149] Clustering module 303 is configured to perform cluster analysis on labeled data using a preset clustering algorithm to identify potential target network entities involved in fraud.

[0150] Execution module 304 is configured to perform aggregate analysis of the metadata using situational awareness technology to identify and track target network entities from the potential target network entities and to conduct risk assessment.

[0151] Optionally, the marking module 302 is configured to:

[0152] Initial labels are set for nodes of known categories based on configuration information. The known categories of nodes include CDN nodes, public cloud nodes, and Whois privacy-protected nodes.

[0153] The final label of each node is determined based on the connection relationships between nodes, including traffic or network connections recorded in the traffic log;

[0154] Traffic is labeled based on the node labels, traffic characteristics, and traffic behavior patterns of the nodes it passes through. The traffic characteristics include protocol type, port number, and request URL, and the traffic behavior patterns include access frequency and request time distribution.

[0155] Optionally, the clustering module 303 is configured to:

[0156] Based on the requirements for identifying fraudulent activities, target features are selected, and the target features are preprocessed and normalized. The target features include domain age, SSL certificate information, IP location similarity, domain Hamming distance, IP service and port information.

[0157] Based on the characteristics of the fraudulent behavior, the neighborhood radius and minimum number of points of the preset clustering algorithm are set, and the labeled data are clustered according to the preset clustering algorithm and the target features to generate clustering results;

[0158] Based on the clustering results, the characteristics and commonalities of different clusters are analyzed to identify potential target network entities involved in fraud.

[0159] Optionally, the clustering module 303 is configured to:

[0160] The labeled data are grouped according to the neighborhood radius, the minimum number of points, and the target feature.

[0161] Traverse all data points in the marked data. When the number of data points within the radius of the target data point is greater than or equal to the minimum number of points, set the target data point as the core point.

[0162] Starting from the core point, adjacent data points are connected into a cluster based on density reachability, and clustering results are generated based on the clusters.

[0163] Optionally, the execution module 304 is configured to:

[0164] Data aggregation technology is used to aggregate preprocessed metadata in order to reveal potential fraud patterns and trends;

[0165] Natural language processing and semantic analysis techniques are used to understand the aggregated metadata in order to reveal the logical relationships between the data. Through association reasoning algorithms, the connection between different fraudulent behaviors and online assets is determined.

[0166] Using time series analysis and deep learning models, we predict the trends of fraudulent activities and the risks to online assets, and create a fraud trend prediction chart.

[0167] Optionally, the execution module 304 is configured to:

[0168] The types of fraudulent activities are used as keywords for aggregation. By statistically analyzing the distribution of each type in the overall fraudulent activities, the most frequent types of fraudulent activities can be identified.

[0169] Using the geographical location information of fraudulent activities as the aggregation keywords, a heat map or distribution map of fraudulent activities is drawn by counting the number of fraudulent activities in each geographical location.

[0170] Using the time information of fraudulent acts as the aggregation keyword, a timeline chart of fraudulent acts is drawn by statistically analyzing the number of fraudulent acts that occur in each time period.

[0171] By cross-combining the analysis results of a single dimension, a multi-dimensional cross-aggregated view is formed, where the single dimension includes type, location, and time.

[0172] Optionally, the execution module 304 is configured to:

[0173] Entities are extracted from aggregated metadata using named entity recognition technology. These entities include the name of the fraud method, the type of victim, the IP address, and the domain name.

[0174] By using relation extraction technology, the relationships between the entities are identified and extracted. Combined with contextual information, semantic understanding is performed on the extracted entities and the relationships to reveal the underlying logical relationships.

[0175] Different data points are associated according to preset rules, and the associated data points are represented in the form of a network graph, where nodes represent entities and edges represent the relationships between entities.

[0176] By using association reasoning algorithms, path search and influence analysis are performed in the network graph to reveal the connections between different fraudulent activities and online assets.

[0177] It should be noted that the above embodiments of the apparatus are only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0178] This embodiment also discloses an electronic device, as shown in the reference. Figure 4 The electronic device may include: at least one processor 401, at least one communication bus 402, user interface 403, network interface 404, and at least one memory 405.

[0179] The communication bus 402 is used to enable communication between these components.

[0180] The user interface 403 may include a display screen and a camera. Optionally, the user interface 403 may also include a standard wired interface and a wireless interface.

[0181] The network interface 404 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).

[0182] The processor 401 may include one or more processing cores. The processor 401 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 405, and by calling data stored in memory 405. Optionally, the processor 401 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 401 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also be implemented as a separate chip without being integrated into the processor 401.

[0183] The memory 405 may include random access memory (RAM) or read-only memory. Optionally, the memory 405 may include a non-transitory computer-readable storage medium. The memory 405 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 405 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 405 may also be at least one storage device located remotely from the aforementioned processor 401. Figure 4 As shown, the memory 405, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for clustering fraudulent entities based on network traffic and metadata.

[0184] exist Figure 4In the electronic device shown, the user interface 403 is mainly used to provide an input interface for the user and to obtain the user input data; while the processor 401 can be used to call the application stored in the memory 405 for a method of clustering fraudulent entities based on network traffic and metadata. When executed by one or more processors 401, the electronic device executes one or more methods as described in the above embodiments.

[0185] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0186] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0187] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the shown or discussed mutual couplings or direct couplings or communication connections may be through some service interfaces; indirect couplings or communication connections between apparatuses or units may be electrical or other forms.

[0188] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0189] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0190] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory 405 and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory 405 includes various media capable of storing program code, such as a USB flash drive, external hard drive, magnetic disk, or optical disk.

[0191] The foregoing description is merely an exemplary embodiment of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Other embodiments of this disclosure will be readily apparent to those skilled in the art upon consideration of the disclosure in this specification. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.

Claims

1. A method for fraud entity clustering based on network traffic and metadata, the method comprising: The method is applied to a fraud-related network entity clustering and tracking platform, and comprises the following steps: Obtaining traffic logs and metadata from multiple data sources, including CDNs, public clouds, private clouds, and internal networks in target regions; Labeling the traffic in the traffic logs by a preset label propagation algorithm to distinguish CDN traffic, public cloud traffic, private cloud traffic, Whois privacy protection traffic, malicious traffic, and business traffic; Performing clustering analysis on the labeled data by a preset clustering algorithm to identify potential target network entities related to fraud, and selecting target features for clustering analysis according to the identification requirements of fraud behaviors, wherein the target features include domain name age, SSL certificate information, IP location similarity, domain name Hamming distance, IP service, and port information; Performing aggregated analysis on the metadata by a situational awareness technology to identify and track target network entities from the potential target network entities and perform risk assessment, specifically comprising the following steps: Using data aggregation technology to aggregate the preprocessed metadata to reveal potential fraud behavior patterns and trends; Using natural language processing and semantic analysis technology to understand the aggregated metadata to reveal the logical relationships between the data, and determining the relationships between different fraud behaviors and network assets by an association reasoning algorithm; Using time series analysis and deep learning models to predict the trends of fraud behaviors and the risks of network assets, and drawing fraud trend prediction graphs, The steps of using natural language processing and semantic analysis technology to understand the aggregated metadata to reveal the logical relationships between the data, and determining the relationships between different fraud behaviors and network assets by an association reasoning algorithm comprise the following steps: Extracting entities from the aggregated metadata by named entity recognition technology, wherein the entities include fraud means names, victim types, IP addresses, and domain names; Using relationship extraction technology to identify and extract the relationships between the entities, and combining context information to understand the semantics of the extracted entities and the relationships to reveal the logical relationships between the entities; According to a preset rule, associating different data points and representing the associated data points in the form of a network graph, wherein the nodes represent the entities and the edges represent the relationships between the entities; Performing path search and influence analysis in the network graph by an association reasoning algorithm to reveal the relationships between different fraud behaviors and network assets.

2. The method for cyber fraud detection entity clustering based on network traffic and metadata according to claim 1, wherein, The steps of labeling the traffic in the traffic logs by a preset label propagation algorithm comprise the following steps: Setting initial labels for nodes of known categories based on configuration information, wherein the nodes of known categories include CDN nodes, public cloud nodes, and Whois privacy protection class nodes; Determining the final labels of each node according to the connection relationships between the nodes, wherein the connection relationships include the traffic or network connections recorded in the traffic logs; Labeling the traffic according to the node labels, traffic features, and traffic behavior patterns of the nodes, wherein the traffic features include protocol types, port numbers, and request URLs, and the traffic behavior patterns include access frequency and request time distribution.

3. The method for cyber fraud detection entity clustering based on network traffic and metadata as claimed in claim 1 wherein, The steps of performing clustering analysis on the labeled data by a preset clustering algorithm to identify potential target network entities related to fraud comprise the following steps: Preprocessing and normalizing the target features; Setting a neighborhood radius and a minimum point number of the preset clustering algorithm according to characteristics of the fraud behavior, performing clustering analysis on the labeled data according to the preset clustering algorithm and the target features to generate a clustering result; Based on the clustering result, analyzing features and commonalities of different clusters to determine potential target network entities involved in fraud.

4. The method for cyber fraud detection entity clustering based on network traffic and metadata of claim 3, wherein, The clustering analysis on the labeled data according to the preset clustering algorithm and the target features to generate a clustering result includes: Grouping the labeled data according to the neighborhood radius, the minimum point number and the target features; Traversing all data points in the labeled data, and setting a target data point as a core point when a number of data points within the neighborhood radius of the target data point is greater than or equal to the minimum point number; Connecting adjacent data points into a cluster through density reachability from the core point, and generating a clustering result according to the cluster.

5. The method for cyber fraud detection entity clustering based on network traffic and metadata of claim 1, wherein, The data aggregation technology is used to aggregate the preprocessed metadata to reveal potential fraud behavior patterns and trends, including: Taking type information of fraud behavior as an aggregation key, and identifying a fraud behavior type with the highest occurrence rate by counting a distribution of each type in overall fraud behavior; Taking geographical location information of fraud behavior as an aggregation key, and drawing a fraud behavior heat map or distribution map by counting a number of fraud behaviors occurring at each geographical location; Taking time information of fraud behavior as an aggregation key, and drawing a fraud behavior timeline by counting a number of fraud behaviors occurring in each time period; Combining single-dimensional analysis results to form a multi-dimensional cross-aggregation view, the single dimension including type, location and time.

6. A system for fraud entity clustering based on network traffic and metadata, the system comprising: The system includes a collection module, a labeling module, a clustering module and an execution module, wherein: The collection module is configured to obtain traffic logs and metadata from multiple data sources including CDNs, public clouds, private clouds and internal networks of target regions; The labeling module is configured to label traffic in the traffic logs through a preset label propagation algorithm to distinguish CDN traffic, public cloud traffic, private cloud traffic, Whois privacy protection traffic, malicious traffic and business traffic; The clustering module is configured to perform clustering analysis on the labeled data through a preset clustering algorithm to identify potential target network entities involved in fraud, and select target features for clustering analysis according to fraud behavior identification requirements, the target features including domain name age, SSL certificate information, IP location similarity, domain name Hamming distance, IP service and port information; The execution module is configured to perform aggregation analysis on the metadata through a situation awareness technology to identify and track target network entities from the potential target network entities and perform risk assessment, specifically including: Using a data aggregation technology to aggregate the preprocessed metadata to reveal potential fraud behavior patterns and trends; The natural language processing and semantic analysis technology is used to understand the aggregated metadata, so as to reveal the logical relationship between the data, and the association reasoning algorithm is used to determine the relationship between different fraud behaviors and network assets; The time series analysis and deep learning model are used to predict the trend of fraud behaviors and the risk of network assets, and a fraud trend prediction graph is drawn, The execution module is further configured to extract entities from the aggregated metadata by using the named entity recognition technology, wherein the entities include fraud means names, victim types, IP addresses and domain names; The relationship extraction technology is used to identify and extract the relationship between the entities, and the semantic understanding is performed on the extracted entities and the relationship in combination with context information, so as to reveal the logical relationship between the entities; According to a preset rule, different data points are associated, and the associated data points are represented in the form of a network graph, wherein a node represents an entity, and an edge represents the relationship between entities; The association reasoning algorithm is used to perform path search and influence analysis in the network graph, so as to reveal the relationship between different fraud behaviors and network assets.

7. An electronic device, comprising: The electronic device includes a processor, a memory, a user interface and a network interface, the memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory, so that the electronic device performs the method according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores instructions, when the instructions are executed, the method according to any one of claims 1-5 is performed.

Citation Information

Patent Citations

  • Medical insurance cash-out fraud behavior identification method and system

    CN115330546A