Intelligence and Visualization of Attack Activities for Countering Cyber Attacks
Through a multi-attribute cluster identifier system, cluster analysis and malicious activity determination model are used to solve the problem that the existing technology is difficult to identify and prevent network attacks initiated by email, and effectively identify and mark malicious email activities.
Patent Information
- Application Number
- CN202080076915.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-03
- Filing Date
- 2020-11-03
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-11-03
AI Technical Summary
The prior art is difficult to effectively identify and prevent cyber attacks initiated by email, especially when attackers are constantly changing their strategies and using masquerading legitimate emails.
A multi-attribute cluster identifier system is used to determine models through cluster analysis and malicious activity, and an identifier with risk scores and attribute sets is generated to identify and mark malicious email activities.
It realizes effective identification and labeling of malicious email activities in the computing environment, reduces false positive rates, and improves rapid response capabilities to potential threats.
Smart Images

Figure CN114761953B_ABST
Abstract
Description
Background Art
[0001] Users typically rely on computing resources such as applications and services to perform various computing tasks. Computing resources can be exposed to different types of malicious activities that limit the operating capabilities of the computing resources. For example, cyberattacks remain an increasing avenue for malicious actors to attempt to expose, steal, damage, or gain unauthorized access to an individual's or company's assets. Specifically, email attacks remain an entry point for attackers to target enterprise customers of computing environments and can trigger damaging consequences.
[0002] Generally, email attacks do not target a single individual. Attackers typically target multiple individuals simultaneously via waves of email attacks. These attack campaigns usually last for a period of time to break through various email filters and affect potential victims. To generate an effective attack campaign, an attacker can first set up a sending infrastructure (e.g., a secure, authenticated or unauthenticated sending domain, sending IP, sending email script, etc.). Once the infrastructure is formed, email templates with specific payloads (e.g., URL hosting, infected documents attached to emails, etc.) can be created to deceive and lure users into interacting with the emails to deliver the payload. Additionally, post-delivery attack chains (e.g., capturing credentials for phishing attacks or other malware) can be embedded in the emails so that attackers can collect information from the payloads they deliver. Due to the increase in complex event detection, cybercriminals are constantly changing their strategies to effectively target the maximum number of individuals. Thus, it is important to develop cybersecurity systems to reduce or prevent cyberattacks against computing environments. Summary of the Invention
[0003] Aspects of the techniques described herein generally relate to systems, methods, computer storage media, etc. for providing multi-attribute cluster identifiers (i.e., enhanced multi-attribute-based fingerprints for clusters), where the multi-attribute cluster identifiers support identifying malicious activities (e.g., malicious email attack campaigns) in a computing environment (e.g., a distributed computing environment). Instances of activities with an attribute set (e.g., email messages) can be evaluated. The attribute set of an instance of an activity is analyzed to determine whether the instance of the activity is a malicious activity. Specifically, the attribute set of the instance of the activity is compared with multiple multi-attribute cluster identifiers of previous instances of the activity such that when the attribute set of the instance of the activity corresponds to an identified multi-attribute cluster identifier, the instance of the activity is determined to be a malicious activity. The risk score and the attribute set of the identified multi-attribute cluster identifier indicate the likelihood that the instance of the activity is a malicious activity.
[0004] For example, multiple multi-attribute cluster identifiers for an email attack campaign can be generated based on a set of attributes of previous emails in the email attack campaign. The malicious activity manager can then monitor the computing environment and analyze instances of the activity (e.g., emails) to support determining whether the instance of the activity is a malicious activity. When the set of attributes of an email corresponds to a multi-attribute cluster identifier having a risk score and a set of attributes, the email is identified as a malicious activity, where the risk score and the set of attributes indicate the likelihood that the instance of the activity is a malicious activity.
[0005] In operation, the malicious activity manager of the malicious activity management system supports generating multiple multi-attribute cluster identifiers having corresponding risk scores and sets of attributes. The generation of the multi-attribute cluster identifiers is based on a cluster analysis and a malicious activity determination model, and the malicious activity determination model supports determining that a set of attributes of an activity has a quantified risk (i.e., a risk score) indicating that the activity is a malicious activity. Generating multi-attribute cluster identifiers for an activity (e.g., an email attack campaign) can be based on: an initial clustering of instances of the activity based on a corresponding set of attributes (e.g., a set of attributes of emails in an email campaign).
[0006] The malicious activity manager then operates to monitor the computing environment, analyze instances of the activity (e.g., emails), and match the set of attributes of the instance with the identified multi-attribute cluster identifiers indicating that the activity is a malicious activity. In this way, the multi-attribute cluster identifiers corresponding to the type of malicious activity support effectively identifying instances of the activity being processed in the computing environment as malicious activities.
[0007] The malicious activity manager also supports providing a visualization of malicious activity operation data. The visualization is associated with different functions and features of the malicious activity manager (e.g., cluster analysis, risk score, set of attributes, evaluation results). The visualization is provided to allow for effective operation and control of the malicious activity manager from the perspective of the user interface, while the malicious activity manager provides feedback information when making decisions regarding malicious activity management.
[0008] The operations of the malicious activity manager ("malicious activity management operations") and additional functional components are executed as algorithms for performing specific functions of malicious activity management. In this way, the malicious activity management operations support restricting or preventing malicious activities in the computing environment, where malicious actors are constantly changing their strategies for performing malicious activities. The malicious activity management operations also support providing a visualization of malicious activity operation data that is helpful to the user, and the visualization of the activity operation data improves the efficiency of the computer in making decisions and user navigation of the graphical user interface for malicious activity management. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The techniques described herein are described in detail below with reference to the drawings, in which:
[0010] Figures 1A to 1D A block diagram showing an example malicious activity management environment for providing malicious activity management operations using a malicious activity manager, suitable for implementing aspects of the techniques described herein;
[0011] Figure 2 An example malicious activity management environment for providing malicious activity management operations according to aspects of the techniques described herein;
[0012] Figure 3 An example schematic diagram of malicious activity management processing data for providing malicious activity management operations according to aspects of the techniques described herein;
[0013] Figure 4 An example schematic diagram of malicious activity management processing data for providing malicious activity management operations according to aspects of the techniques described herein;
[0014] Figure 5 An example schematic diagram of malicious activity management processing data for providing malicious activity management operations according to aspects of the techniques described herein;
[0015] Figure 6 An example schematic diagram of malicious activity management processing data for providing malicious activity management operations according to aspects of the techniques described herein;
[0016] Figure 7 An example schematic diagram of malicious activity management processing data for providing malicious activity management operations according to aspects of the techniques described herein;
[0017] Figure 8 An example schematic diagram of malicious activity management processing data for providing malicious activity management operations according to aspects of the techniques described herein;
[0018] Figure 9 An example schematic diagram of malicious activity management processing data for providing malicious activity management operations according to aspects of the techniques described herein;
[0019] Figures 10A to 10C A schematic diagram of malicious activity management processing data for providing malicious activity management operations according to aspects of the techniques described herein;
[0020] Figures 11A to 11B A schematic diagram of malicious activity management processing data for providing malicious activity management operations according to aspects of the techniques described herein;
[0021] Figure 12A schematic diagram of malicious activity management processing data for providing malicious activity management operations according to aspects of the technology described herein;
[0022] Figure 13 Provide a first example method of providing malicious activity management operations according to aspects of the technology described herein;
[0023] Figure 14 Provide a second example method of providing malicious activity management operations according to aspects of the technology described herein;
[0024] Figure 15 Provide a third example method of providing malicious activity management operations according to aspects of the technology described herein;
[0025] Figure 16 Provide a block diagram of an exemplary distributed computing environment suitable for implementing aspects of the technology described herein; and
[0026] Figure 17 A block diagram of an exemplary computing environment suitable for implementing aspects of the technology described herein. Detailed Description
[0027] Overview of technical problems, technical solutions, and technical improvements
[0028] Email-initiated cyberattacks remain a preferred avenue for global cybercriminals. By sending several waves of emails at once, threat actors are better able to reach potential victims. The emails used in these attack campaigns are often highly customized to mimic legitimate emails and avoid sophisticated filtering systems used by security operations teams. By conducting attack campaigns over a period of time, threat actors can effectively measure their hit rates to ensure their chances of success. Additionally, attackers can continuously aggregate and update malicious activities by quickly adjusting their infrastructure, email templates, payloads, and post-delivery attack chains to ensure their chances of success. The rapid movement of threat actors poses a challenge to security operations teams responsible for thwarting these dangerous and evolving attack campaigns.
[0029] Generally speaking, the process of identifying, investigating, and responding to potential security incidents is a strict and complex process that requires human intervention at all stages of the process. In large enterprise organizations, it is difficult, time-consuming, and expensive to intelligently correlate potential threats based on indicators of compromise ("IOCs"). Security operations teams using traditional cybersecurity systems (e.g., cybersecurity management components) continuously monitor for IOCs by creating queries to analyze email activity and manually inspecting email for compromise patterns based on email content and business. However, since attackers are constantly developing and changing their attack campaigns (i.e., computer- or network-based system attacks), it is difficult to correlate individual potentially malicious emails with other threats in a large dataset.
[0030] For example, email phishing attacks have various characteristics that are manually analyzed by security operations teams and automatically analyzed by email filtering systems. Attackers can observe how their email attack campaigns interact with the target's threat detection systems and adjust various features of their email attack campaigns to defeat the filtering mechanisms. For example, cybercriminals can change the subject line, sending domain, URL, or any combination of other features to successfully deliver the email to the target inbox. Therefore, there is a need to automatically observe IOCs in large amounts of data to identify malicious email activity.
[0031] In addition, traditional cybersecurity systems analyze emails on a case-by-case basis by examining many features of the emails to determine the legitimacy of the emails. For example, a cybersecurity management component can analyze unique features (e.g., fingerprints) of emails with known malicious email characteristics (i.e., known malicious email fingerprints) as an indicator of potential threats. However, email fingerprints are generally misleading and unreliable for identifying potential threats because email fingerprints are generated based on the content of the emails. Thus, using email fingerprints to identify spam and malicious emails, malicious emails often bypass the filtering system or go undetected by the security operations team because malicious emails often have the same or similar fingerprints as legitimate emails. As such, there is a need to group emails to increase the speed of detecting potential attacks across large datasets. Thus, a comprehensive malicious activity management system that addresses the limitations of traditional systems will improve malicious activity management operations and malicious activity interfaces that support identifying malicious activity in a computing environment.
[0032] Embodiments of the present invention are directed to a simple and effective method, system, and computer storage medium for providing a multi-attribute cluster identifier (i.e., an enhanced multi-attribute-based fingerprint for a cluster) that supports identifying malicious activities (e.g., malicious email activities) in a computing environment (e.g., a distributed computing environment). An instance of an activity (e.g., an email message) having a set of attributes can be evaluated. The set of attributes of the instance of the activity is analyzed to determine whether the instance of the activity is a malicious activity. Specifically, the set of attributes of the instance of the activity is compared to a plurality of multi-attribute cluster identifiers of previous instances of the activity, such that when the set of attributes of the instance of the activity corresponds to an identified multi-attribute cluster identifier, it is determined that the instance of the activity is a malicious activity. The identified multi-attribute cluster identifier has a risk score and a set of attributes, where the risk score and the set of attributes indicate the likelihood that the instance of the activity is a malicious activity.
[0033] For example, a malicious activity management system can support clustering instances of activities (e.g., cyberattacks) based on attributes associated with the activities. Clusters of cyberattack instances can be analyzed based on visualization and interpretation of cluster segments (e.g., cyberattack segments) of the instances of the activities to generate multi-attribute cluster identifiers. For example, by clustering emails within a single cyberattack activity over a period of time (e.g., days, weeks, months, etc.). Malicious activity management operations can help determine the nature and impact of a cyberattack. Characteristics of the attack activity (such as IoCs) are identified across large datasets and the infrastructure used by cybercriminals for email sending and payload hosting is revealed. Emails are clustered into a unified entity (i.e., a cyberattack segment) to visualize attacker-specific signals and tenant policies (e.g., safe senders, technical reports, etc.). Allowing for quick categorization to determine the urgency of an email attack activity. Additionally, malicious activity management operations allow for bulk operations to be taken on all emails found in an attack activity, thus supporting effective remediation efforts. In this way, a malicious activity management system is able to effectively counteract rapidly changing email attack activities of attackers by clustering emails based on relevant attributes indicating malicious activities.
[0034] To identify which emails are part of the same attack campaign, malicious activity management operations support analyzing multiple attributes of emails to tie attack campaigns together. To build attack campaigns, malicious activity management operations support linking together various metadata values and applying filters at each step to reduce false positives. For example, clusters can be built by iteratively analyzing and filtering out certain attributes of the attacker's infrastructure (e.g., sending IP address, subject fingerprint, originating client IP, etc.) and including additional centroids based on a single centroid of an extended cluster. For each attack campaign discovered, malicious activity management operations will create an attack campaign guide. The attack campaign guide links to the metadata attributes (e.g., URL domain, IP address, etc.) that are part of the attack campaign. As a result, the attack campaign and its associated guide include a large number of emails with similar attributes indicating potential malicious activity.
[0035] Based on the metadata of individual email messages within an attack campaign, malicious activity management operations can compute email clusters and aggregate the statistics or attributes of these email clusters. Clusters are useful because certain aspects of an attack campaign will have attributes and infrastructure similar to legitimate emails. For example, an email could be sent from a legitimate provider (such as outlook.com), or a server could have been compromised and hosted a phishing website before being taken down. In these cases, it is expected that emails belonging to a cyber attack are included, rather than legitimate emails within the cluster.
[0036] Specifically, using machine learning techniques, malicious activity management operations can automatically cluster emails into attack campaigns by learning a large number of cyber attack attempts. Using fuzzy clustering techniques, emails can be grouped based on locality-sensitive hashing (LSH) fingerprints and then the clusters can be extended based on similar attributes in other emails to capture potential anomalous emails (note: LSH fingerprinting supports techniques for obtaining fingerprints from any type of document or content). Since LSH fingerprints are based on the text of an email message, or any other document or content (excluding URLs and HTML structures), the distance between two LSH fingerprints with similar content is very close. After generating clusters based on email fingerprints, multiple clustering steps can be performed to aggregate similar clusters, and smaller clusters and outliers can be filtered out to produce a final cluster that can be used to compute a risk score. Thus, clusters of potentially harmful email messages can be generated at any time for further analysis.
[0037] Additionally, malicious activity management operations support assigning a risk score to each cluster based on the behavior of filtering emails or the behavior of other users in the cluster towards the emails. For example, if most of the users who received emails during a specific attack campaign have never received emails from a specific sender, then this increases the suspicion that the attack campaign is malicious. As another example, if some of the users who received emails during a specific attack campaign move the emails to the "Junk" folder, then this is also a sign that the entire attack campaign is illegal. As another example, if malware detection software determines that the attachments in some of the email messages in an attack campaign are infected with potentially harmful code, it is very likely that all the attachments in the emails with a similar content template sent by the same sender during the attack campaign are also infected.
[0038] The risk score of a cluster can be calculated based on suspicion scores, anomaly scores, and impact scores generated by separate models. Each scoring model can operate at the tenant level, global level, and combined level. The trained suspicion score model uses historical data and aggregates information about various characteristics and attributes of the cluster (e.g., the ratio of the sender domain's first contact with the tenant or any other client globally). The trained anomaly model analyzes historical data, as well as whether the emails in the cluster are from new senders to the tenant, new linked hosts to the tenant, or from an old sender to the tenant and a new linked host to the tenant from a new location. For the trained anomaly model, historical data is defined at the tenant level and within a closed time period (e.g., six months). Three examples of historical data are the history of senders, the history of geographical locations, and the history of linked hosts. The characteristics of historical data include the age of the relationship, the number of days seen, the total number of messages (tenant divided by global), and the total number of error messages (tenant divided by global).
[0039] The impact score model can be divided into example impact scores, such as but not limited to high-value employee (HVE) scores, content classification scores, and website impact scores. The HVE model uses email communication patterns to build a graph of users with consistent email communication to identify which employees are the targets of a specific group of emails. The HVE model aims to identify those employees who, once compromised, will have a greater impact on the customer. For example, if compromised, the CEO can approve payments. In addition to other privileges and rights, administrators also have special access to sensitive data. The content classification model uses machine learning techniques to identify messages that request attributes such as credentials. The machine learning techniques are based on the hash subject and body tokens of the email message as well as structural features.
[0040] Finally, a website impact model, which can be implemented as part of the suspicion model, analyzes whether a website found in an email in a cluster is malicious, either alone or in any other suitable way, based on whether the website includes form entries, form entries that collect passwords, and whether the website resembles a proprietary website (e.g., Office 365). Other factors for the embodiments disclosed herein for determining the risk score of a cluster include: target metadata (e.g., if an attack campaign targets multiple users with administrative roles in IT management, members of the executive or finance teams, etc., the risk score will increase), attack campaign growth (if the total volume of an attack campaign increases over a period of time, the risk score increases because this means a larger total volume will occur in the future), user actions (e.g., if an attack campaign witnesses a large number of recipients clicking on hyperlinks or other items in an email, this means the email is crafted and appears credible, and thus is at higher risk), and payload type (e.g., a malware attack campaign attempting to infect Windows devices will not cause much harm if it attacks a mobile device with a different operating system), etc.
[0041] To further explain, an email message can belong to many different types of clusters. A fuzzy content cluster is an example, and a payload (website, malware, etc.) cluster is another example. Multi-attribute aggregation is also possible, and each multi-attribute aggregation can have its own associated suspicion score. Multiple suspicion scores can be combined into a single aggregated value. If desired, the aggregation can be the result of a statistical combination or the result of a trained model enhanced by heuristics. An example of the difference between a suspicion score and an anomaly score is: a suspicion score can be obtained as a result of a supervised training method (e.g., labeled data is accessible for both what we consider good or bad, and the system automatically learns it). An anomaly score can be, but is not necessarily, the result of an unsupervised statistical analysis or heuristics. Finally, if a compromise is successful, an impact score can assess the damage.
[0042] As a result of calculating an overall risk score, malicious activity management operations can take various actions based on the severity of the risk score. For example, an attack activity can be blocked such that future email messages that match the email message profile in the attack activity no longer reach their intended recipients. As another example, prioritized alerts can be triggered based on the risk score such that security operations processors and personnel can remediate the problem before victims (e.g., enter credentials into a malicious web form, install malware on a user device, etc.). As another example, affected users and devices can be automatically contained and conditional access can be automatically applied to data access. Additionally, malicious activity management operations can support automatic response to events. For example, according to an embodiment of the present invention, malicious emails can be automatically removed and affected or infected devices can be automatically remediated. As another example, malicious activity management operations can further support automatically clearing any policy or configuration issues. As yet another example, URL senders that may be determined to be or have been determined to be malicious (or at least potentially malicious) can be automatically blocked if they have not been blocked previously.
[0043] The risk score associated with a cluster also allows for the quick and efficient exploration of a particular email attack activity. For example, a user in a security operations team needs to prioritize emails with the highest associated risk. The risk score (e.g., high, medium, or low) will allow the user to quickly see which email clusters pose the greatest threat to the organization. Transmitting the risk score and relevant information about the attack activity to the UI provides the user with a powerful visualization tool. The user can expand a highly dangerous attack activity to view detailed information about the attack activity (e.g., IOCs) and take bulk actions on all emails in the attack activity. For example, alerts are sent to individuals who received emails in a high-risk score attack activity.
[0044] Advantageously, embodiments of the present disclosure can effectively identify, investigate, and respond to large-scale security incidents. By identifying email attack activities that have multiple common attributes, the scope of potential cyberattacks can be better understood, thus preventing future attacks. Clustering emails into a logical entity allows for large-scale actions to be taken across multiple tenants rather than at a granular level, which can be a time-consuming and highly manual process. The flexibility to visualize and explore attack activities at a high level and at a granular level enables stakeholders to make effective decisions and enables security teams to investigate attack activities to mitigate any threats and avoid future harm.
[0045] Overview of an example environment for providing malicious activity management operations using a security maintenance manager Aspects of the technical solution can be described by way of example and with reference to FIGS. 1 to Figure 12 and Figure 12associated with an exemplary solution environment (malicious activity management environment or system 100) suitable for implementing the embodiments of the solution. Generally, the solution environment includes a solution system suitable for providing malicious activity management operations based on a malicious activity manager. First, referring to Figure 1A , in addition to engines, managers, generators, selectors, or components (collectively referred to herein as "components") not shown, Figure 1A discloses a malicious activity management environment 100, a malicious activity management service 120, a malicious activity manager 110 having instance and property sets 112 of activities, cluster instructions 114, score instructions 116, a malicious activity model 118, a multi - attribute cluster identifier 120, analysis instructions 130, visualization instructions 140, malicious activity management operations 150, a malicious activity client 180 (including a malicious activity management interface 182 and malicious activity management operations 184).
[0046] The components of the malicious behavior management system 100 can communicate with each other through one or more networks (e.g., a public network or a virtual private network "VPN") as shown by network 190. Network 190 can include, but is not limited to, one or more local area networks (LANs) and / or wide area networks (WANs). The graphical editing client 180 can be a client computing device corresponding to the computing device described herein with reference to Figure 7 . The graphical editing system operations (e.g., graphical editing system operations 112 and graphical editing system operations 184) can be implemented by a processor executing instructions stored in a memory, as further described with reference to Figure 17 .
[0047] At a high level, the malicious activity manager 110 of the malicious activity management system supports generating different multi - attribute cluster identifiers with corresponding risk scores and property sets. Generating the multi - attribute cluster identifier is based on: cluster analysis and a malicious activity determination model that supports determining that an instance of an activity's property set has a quantified risk (i.e., a risk score) indicating that the activity is a malicious activity. Generating a multi - attribute cluster identifier for an activity (e.g., an email activity) can be based on initially clustering instances of the activity based on a corresponding property set (e.g., the property set of an email in an email activity).
[0048] The malicious activity manager 110 monitors the computing environment and analyzes instances of activities (e.g., emails), and based on matching the instance with a multi - attribute cluster identifier (i.e., a risk score and a property set) indicating that the activity is a malicious activity, identifies the instance of the activity as a malicious activity. Thus, the multi - attribute cluster identifier corresponding to the malicious activity type supports effectively identifying instances of activities being processed in the computing environment as malicious activities.
[0049] The Malicious Activity Manager 110 also supports providing visualization of malicious activity operation data. The visualization is associated with different functions and features of the Malicious Activity Manager (e.g., clustering analysis, risk scores, attribute sets, and evaluation results). The visualization is provided to allow for effective operation and control of the Malicious Activity Manager from the perspective of the user interface, while the Malicious Activity Manager feeds back information when making decisions regarding malicious activity management.
[0050] Various terms and phrases are used herein to describe embodiments of the present invention. Some of the terms and phrases used herein are described herein, but more details are included throughout the description.
[0051] An activity refers to a set of electronic messages that are delivered between computing devices. An activity includes a set of electronic messages stored as data in a machine. An instance of an activity can refer to an individual electronic message within the message set. An electronic message can have a static data structure that facilitates processing the electronic message based on machine instructions using machine operations. Each instance of an activity includes machine instructions embedded within the respective instances of the activity to support the execution of machine operations.
[0052] An example activity can be an email activity, and an instance of the activity can refer to an email message. An attack activity can be a network attack activity (e.g., an email network attack activity). An attack activity can be an umbrella term associated with operations and computational objects (i.e., activity objects) surrounding the management of instances of an activity.
[0053] A fingerprint refers to a short binary representation of a document. The short binary representation can be used for longer documents. The binary representation corresponds to a defined distance metric. Thus, similar documents will have a smaller distance. An example of a fingerprint is referred to as Locality-Sensitive Hashing (LSH). Refer Figure 1C , Figure 1C Shows three different fingerprints (i.e., Fingerprint A10C, Fingerprint B 20C, and Fingerprint C30C). Each fingerprint has a corresponding binary representation that is a distance metric that can be used to measure the similarity between different fingerprints.
[0054] A centroid can refer to a cluster of fingerprints. Specifically, a centroid is a cluster of fingerprints based on fingerprints from the body of an email message. They are created by finding all fingerprints that have N 2-bit common words. For example, if a centroid allows a maximum of 3 two-bit words to be different, then these two words will be in the same centroid, as shown below:
[0055] Fingerprint 1: 0100 0011 1100 1101 0011 1010 0000 1011
[0056] Fingerprint 2: 0110 1011 1100 1001 0011 1010 0000 1011
[0057] A centroid can be generated for a corresponding fingerprint. In particular, a centroid can be effectively generated to effectively cluster fingerprints into existing centroids. As shown, in Figure 1C a centroid 40C can be generated from fingerprint A10C, fingerprint B 20C, and fingerprint C 30C. The centroid value itself is a fingerprint that is as close as possible to multiple fingerprints. One technical implementation of generating a centroid can be based on a daily basis. Generating a centroid daily can result in the same fingerprint having different centroid values on different days. Thus, as an alternative to continuously calculating new centroids, a list of known centroids can be maintained. Then, it can be determined whether one or more fingerprints can be identified within a centroid. If one or more fingerprints are identified within an existing centroid, the fingerprint will be grouped into that centroid cluster. Fingerprints that are not within a centroid will be processed to create a new centroid.
[0058] Reference Figure 1D , Figure 1D shows a high-level representation of various operations and computational objects associated with an attack activity. In particular, Figure 1D it includes value proposition 10D, attack activity visualization 20D, risk score 30D, attack activity classifier 40D, clustering technique 50D, and signals, metadata, and reputation 60D. Attack activity objects can support implementing the technical solutions of the present disclosure, as discussed in more detail below. For example, an attack activity can include a value proposition (i.e., value 10D) that determines the protection, alerts, remediation, and reporting provided. Attack activities and attack activity objects can also provide insights into multiple aspects of a cyber attack experienced by a customer or administrator (i.e., attack activity visualization 20D). The ability to cluster a large number of individual emails (i.e., attack activity classifier 40D and clustering technique 50D), which often vary over time in various ways (email templates, weaponized URLs, sending infrastructure, etc.), is particularly valuable. Clustering can be based on email attributes (e.g., signals, metadata, and sender reputation 60D). Attack activity objects enable users to understand complex cyber attacks because they are summarized and presented in a visual graphical user interface (GUI). Unless otherwise specified, attack activities and attack activity objects can be used interchangeably herein.
[0059] A URL domain, or simply a domain, is an identifying string that defines an administrative autonomy, authorization, or control realm within the Internet. URL domain names are used in various network environments and for naming and addressing purposes in specific applications. The technical implementation of the present disclosure supports recording the domain of a URL provided to a customer. A URL domain can be associated with a sending IP address. An Internet Protocol (IP) address is a numerical label assigned to each device connected to a computer network that uses the Internet Protocol for communication. As used herein, the sending IP address among the IP addresses connected to a main server or a server group is responsible for managing emails sent from the URL domain of the main server or the server group.
[0060] Reference Figure 1B , the malicious activity management system 100, the malicious activity manager 10, and the malicious activity client 20 correspond to the components described in the reference Figure 1A In step 30, multiple instances of potential malicious activities (e.g., multiple email messages) are clustered based on the metadata and attributes associated with the instances (e.g., email messages) in the cluster. The malicious activity management system (e.g., the malicious activity manager 10) identifies attack activities through a series of steps. The malicious activity manager 10 identifies emails that are part of the same attack activity. Multiple attributes of the emails are used to link the attack activities together. The cluster can identify the centroid based on the identified centroid and by leveraging the metadata attributes associated with the email messages in the cluster. The centroid is calculated based on the fingerprint of the email messages and identifies the common metadata attributes among the email messages. The clustering can also be based on: applying a filter based on the common metadata attributes, where applying the filter based on the common metadata attributes identifies two or more subsets of the email messages and generates attack activity guidelines associated with each of the two or more subsets of the multiple email messages, where the attack activity guidelines are linked to the common metadata attributes corresponding to the subsets of the multiple email messages.
[0061] In step 32, a risk score is assigned to the cluster based on the historical information, attributes, and metadata associated with the messages in the cluster. The risk score can be based on a suspicion score, an anomaly score, and an impact score. Thus, in step 34, a suspicion score, an anomaly score, and an impact score are generated for the cluster segment; and in step 36, the suspicion score, the anomaly score, and the impact score are used to generate the risk score. The suspicion score is based on information about the tenant of the computing environment; the anomaly score is based on information about the information sender; and the impact score is based on the location of the individual targeted by the instances of the activity in the corresponding cluster, the content classification in the corresponding cluster, and the website information included in the previous instances of the activity.
[0062] In step 38, a multi-attribute customer identifier is generated, and the multi-attribute cluster identifier includes a risk score and an attribute set. In step 40, an instance of an activity is received in a computing environment. In step 50, an instance of an activity is received at a malicious activity manager. In step 52, the instance of the activity is processed based on a malicious activity model. Processing the instance of the activity is performed by using the malicious activity model. The malicious activity model includes Associated with a previous instance of the activity a plurality of multi-attribute cluster identifiers. The multi-attribute cluster identifier includes a corresponding risk score and an attribute set, wherein the risk score and the attribute set indicate the likelihood that the instance of the activity is a malicious activity. The malicious activity model is a machine learning model generated based on processing a plurality of email messages associated with a plurality of network attacks. The processing of the plurality of email messages is based on: clustering the plurality of email messages based on actions taken by recipients of the plurality of emails, attachments in the emails, and corresponding fingerprints of the plurality of emails. The malicious activity model is configurable to process the plurality of email messages at a tenant level, a global level, or a combined tenant-global level.
[0063] In step 54, it is determined whether the instance of the activity is a malicious activity. Determining whether the instance of the activity is a malicious activity is based on: comparing the attribute set of the instance of the activity with the plurality of multi-attribute cluster identifiers, Where the attributes of an instance of the activity Set matches the attribute set of at least one multi-attribute cluster identifier among multiple multi-attribute cluster identifiers. Multiple multi-attribute sets The risk score and attribute set of at least one multi-attribute cluster identifier among the cluster identifiers indicate the likelihood that an instance of the activity is a malicious activity Likelihood .
[0064] In step 56, a visualization of malicious activity operation data is generated and passed to the malicious activity client 20. The malicious activity operation data is selected from each of the following: cluster analysis, risk score, attribute set, evaluation results for generating corresponding cluster analysis visualizations, risk score visualizations, attribute set visualizations, and evaluation result visualizations. Based on the determination Activ The attribute set of an instance of the activity corresponds to the risk score and attribute set of at least one multi-attribute cluster identifier among multiple multi-attribute cluster identifiers indicating the likelihood that an instance of the activity is a malicious activity, and perform one or more Figure 2 remedial actions. One or more remedial actions are performed based on the severity of the risk score, where the risk score corresponds to one of the following: high, medium, or low. In step 42, a visualization of the malicious activity operation data with the instance of the activity is received and caused to be displayed.
[0065] Reference Figure 2 , Figure 2Shows different stages of malicious activity management operations that support identifying malicious activities in a computing environment. These stages include alert 210, analysis 220, investigation 230, impact assessment 240, containment 250, and response 260. In the alert 210 stage, a report can be triggered and queued in a Security Operations Center (SOC), or an email can be reported by an end user as a phishing email. Operationally, alerts combine dozens or hundreds of emails into a coherent entity. The nature and impact of the attack can be determined from the alert text itself. The attack activity entity can show the hacking infrastructure used based on email sending and payload hosting. Alerts can be visualized using visualization tools. In the alert stage, it can be determined whether an email is malicious. For example, are there any Indicators of Compromise (IOCs), and what does the latest intelligence indicate about the content of the email.
[0066] In the investigation stage 230, other emails that match the identified email are determined. The determination can be based on email attributes (e.g., sender, sending domain, fingerprint, URL, etc.). As mentioned earlier, all emails in an attack activity are intelligently clustered together over days and weeks. The attack activity entity indicates rich context data for matching emails to the cluster. In the impact assessment stage 240, the extent of the impact is determined. For example, which users received such an email, whether any users clicked on a malicious URL, and the users and their devices are investigated. The targeted users of the attack activity can be clear at a glance, including who clicked on what URL, at what time, whether the URL was blocked or allowed, etc. In the containment stage 250, actions are taken to contain the affected users and devices, and conditional access to data is applied by the users and devices. And, in the response stage 250, malicious emails can be removed, clean-up operations and policy configuration operations are performed, and any URLs and unblocked senders can be blocked. Visualization at this stage can support exploring the attack activity, performing bulk operations on all emails in the attack activity, identifying indicators of compromise, and identifying tenant policies.
[0067] Reference Figure 3 , Figure 3 Shows a flowchart and components that support providing malicious activity management operations, where the malicious activity management operations are used to identify malicious activities in a computing environment. Figure 3It includes the following phases: Email Online Protection (“EOP”) 302, Detection 306, and Model Training 308; and services: EOP Service 310, Cluster 320, History 330, External 340, Tenant 360, Global 370, Model Training 380, and Risk Score 390. EOP can refer to a hosted email security service that filters spam and removes viruses from emails. EOP can analyze message logs and extract phishing words associated with people (e.g., High-Value Employees HVE) associated with email messages. Based on the generated history 314 and using a fuzzy clustering algorithm 316, message data from EOP can be used to aggregate features on clusters of email messages 314. The generated history can be based on message data that includes specific tenant as well as global URLs, senders, sender and location combinations. As discussed in more detail herein, clustering features from the cluster, anomaly algorithm 318, IP, and recipient data 319 are processed by machine learning models (e.g., tenant models and global models) trained via Model Training 380 to generate a Risk Score 390.
[0068] Reference Figure 4 , Figure 4 An alternative view and components of a flowchart that supports malicious activity management operations for identifying malicious activity in a computing environment are provided. Figure 4 It includes the following phases: Cluster 402, Detection 404, Notification 406, and Investigation Remediation 408. After Cluster 402, as described herein, a Risk Score 418 can be generated based on a suspicion score, an anomaly score, and an impact score. The risk score can be used in the notification phase (i.e., Notification 406) for alerts (e.g., Alert 422) and automated orders (e.g., Automated Order 420). Notification 406 can also include filtering the risk score of corresponding emails based on anomaly indicators, suspicion indicators, and impact indicators in the investigation and remediation phase.
[0069] Turning to generating fingerprints to support clustering, first, traditional systems only use fingerprints based on email content. Just using fingerprints can produce many FPs / FNs (False Positives: Incorrect Positive Predictions / False Negatives: Incorrect Negative Predictions) because legitimate emails can have the same fingerprint as malicious emails when an attacker tries to impersonate a legitimate sender. Email messages (i.e., instances) of an email attack activity (i.e., activity) are expected to be used herein as example instances of the activity; however, this is not meant to be restrictive as the described embodiments can be applied to different types of activities and corresponding instances.
[0070] The malicious activity management system 100 can specifically utilize attack activity graph clustering techniques based on attack activity guidelines and metadata. Using the attack activity graph clustering technique, an attack activity guideline is generated for each detected activity. Then, the attack activity guideline will be linked to the metadata attributes that are part of the attack activity. Attack activities are generated based on linking metadata values together, and filters are applied at each step to reduce false positives.
[0071] The malicious activity management system 100 can also implement complete fuzzy clustering techniques. For example, LSH fingerprints can be generated based on the text of emails, excluding URLs and HTML structures. These fingerprints are similar to hash functions, but the distance between messages with similar content is close. Complete fuzzy clusters support running daily aggregations. The clustering operation on a given number of messages can be completed based on a defined time period (e.g., 300 million messages per day, and it can take approximately 2 hours to complete). In this way, for example, approximately 500,000 clusters can be output daily from 5 million possible clusters.
[0072] As Figure 5 shown, at step 510, a selected capacity of the instances of the activity (e.g., 10%) can be selected based on a suspicion model (e.g., selecting a set of suspicious instances of the activity). At step 520, fingerprints can be generated for the instances. At step 530, a first clustering step can be performed on a subset of the instances (e.g., 50 thousand batches). At step 540, smaller clusters (e.g., 50% of the total) are filtered. At step 560, a second clustering step can be performed on the instances (e.g., 50k batches). At step 570, a final clustering step is performed. At step 580, outlier clusters are filtered to generate cluster 590.
[0073] Refer to Figure 6 , Figure 6A flowchart for generating clusters based on centroids and URLs is shown. At step 610, a starting centroid is identified, which is BC16F415.983B9015.23D61D9B.56D2D102.201B6. As discussed, a centroid can refer to a cluster of fingerprints. Specifically, the centroid is a cluster of fingerprints based on those from the body of an email message. They are created by finding all fingerprints that have N two - bit common words. At step 620, a plurality of URLs associated with the centroid are identified. At step 630, a plurality of known URLs are filtered out from the URLs found in step 620. At step 640, the filtered set of URLs is identified. At step 650, the filtered URLs are used to identify a new centroid. At step 660, the new centroid is identified. At step 670, centroids (e.g., known centroids) are filtered out from the new centroids. At step 680, a set of instances (e.g., 37,466 messages) associated with the new centroid is identified.
[0074] Reference Figure 7 , Figure 7 illustrates generating attack activities based on identifying multiple common metadata attributes. For example, a first email 710 may include a first detected URL 712, and a second email 720 may include a second detected URL 722. It can be determined that the first URL 712 and the second URL 722 are the same (i.e., chinabic.net 730). The first email may have a first centroid 740, and the second email may have a second centroid 750. Several different metadata types can be identified and assigned priorities so as to generate attack activities using these metadata types (e.g., common attribute 770, metadata type 772, and priority 774). For example, metadata types include static centroids, URL domains, sending IPs, domains, senders, subject fingerprints, originating client IPs, etc. Each is associated with a corresponding priority. Attack activities (e.g., attack activity 760 with attack activity ID 762, centroid 764, and URL domain) associated with the first and second emails can be generated. Advantageously, if an email is sent to the MS CEO and the email is phishing / spam missed by the current system. For different users, the content of the email is slightly different. Thus, there is a similarity metric for emails sent by the same person, which leads to other attributes. Therefore, based on one message, embodiments of the present invention infer and are able to calculate that approximately 4500 messages in the attack activity are potentially malicious. To achieve this, in embodiments of the present invention, emails can be wisely selected using clustering techniques to generalize and capture what is abnormal.
[0075] Turning Figure 8 , Figure 8An attribute grid that can be used to cluster emails is shown. Specifically, malicious activity management operations include supporting the intelligent selection of attributes by tracking reputations using various services and attributes. The tracking of reputations can be for known entities. For each entity, the reputation is checked across multiple attributes, and clusters are expanded in that dimension. For example, if 4,500 emails are sent with a URL that has never been sent before, and these emails are from many different companies and are directed to high-value employees (HVE), this would be very suspicious.
[0076] Malicious activity management operations also support discerning malicious emails from legitimate emails. In some cases, there will be users reporting problems. In other cases, some analysts will view the received emails and conduct investigations. Embodiments of the present invention can use machine learning models to analyze all these characteristics of email clusters to determine whether an email is suspicious. For example, see how many supervisors an email is sent to. Extract features, historical data (user operations), and some other matters from the email. A machine learning model can be generated to distinguish whether the received cluster is malicious or legitimate. Then additional logic is used to determine whether the cluster is a mixed cluster, because a malicious message does not mean the entire message is bad. It is expected that various ways of grouping / clustering email messages according to attributes can be implemented. For example, by content similarity (e.g., fingerprinting using locality-sensitive hashing (LSH)).
[0077] Referring to additional clustering techniques, the cluster processes the operation of grouping / clustering emails in the following way: This way actually enables embodiments of the present invention to distinguish different emails, and also enables embodiments of the present invention to discern how similar emails are in different time periods. The cluster further considers how the maintenance of this information is carried out (e.g., how to maintain the cluster / grouping). And, given the expected cluster scale, various attributes or signals can be used to determine whether the cluster is malicious (e.g., spam, phishing, malware, etc.). A machine learning model can be used to determine that the cluster is malicious. And, visualization techniques can be implemented to relay, display, present, or otherwise demonstrate to the user various aspects of potential attacks in email attack activities (e.g., presenting an attack with thousands of emails from different places and going to different places). Because of the large scale of emails, the operations described here can also see activities at a global scale across many tenants and customers. In this way, embodiments of the present invention can establish attack activities due to the large amount of data accessible by embodiments of the present invention.
[0078] As mentioned above, the risk score can be based on several different scores. The scores can be used to determine the overall risk score of the associated email attack activity. Associated models can be used for each score calculation. One such score calculation is the suspicion score. It can be generated at the tenant and global levels. Features include engine information aggregation, such as how much was captured by which subsystem, model scores, and other subsystem results, historical data: for example, connection graph features with P2 senders and URL hosts. The training aspects of the suspicion model include: bad labels, such as clusters containing at least one phishing / malware manual rating (provider or analyst); good labels, such as clusters containing at least one good manual rating or one good label and no bad manual ratings. The model can be any conceivable / suitable model for generating a suspicion score based on email attributes (e.g., Figure 9 example features in). For example, an ensemble of trees consisting of 4 boosted trees of different lengths and 1 linear regression. The following are examples of features considered by the suspicion model.
[0079] Another score is the anomaly score. Anomaly features include new senders for the tenant, new link hosts for the tenant, old senders for the tenant from new locations, and new links for the tenant, etc. Any suitable model, algorithm, process, or method (such as a simple heuristic) can be used to calculate the anomaly score. The model can use historical data at the tenant and global levels to generate the anomaly score. The history can be defined as a closed time period (e.g., 6 months). The update frequency can be daily. The historical data is defined at the tenant level. Example types of historical data include but are not limited to: history of senders (e.g., the ideal address format would be P1+P2, even better; "original self-header + P1 address"), history of geographical locations (e.g., geographical grid calculated based on latitude / longitude, client IP, and connection IP address), history of link hosts. Features include relationship age, number of days seen, total number of messages (tenant / global), and total number of bad messages (tenant / global), etc.
[0080] An impact score can also be calculated. For example, an impact score can be calculated for a High-Value Employee (HVE). The impact score model can use email communication patterns to build a graph of users with consistent email communication. For example, given the CEO as a seed, the distance along the graph to other users is equivalent to the importance of the user. Another impact score that can be calculated is based on content classification. For example, embodiments of the present invention can use a machine learning model or any other suitable model that identifies messages requesting credentials based on / analyzes / uses hashed subject and body tags and structural features. In some embodiments, the model is trained on messages marked as brand imitation by the person being scored and on messages from corpus heuristics that are unlikely to be good messages for credential hunting. Yet another impact score that can be calculated by embodiments of the present invention is based on website impact signals. For example, embodiments of the present invention can use a model (machine learning or otherwise) to determine whether a website included in an email associated with an attack activity includes form entries, forms for collecting passwords, or whether the website brand looks like an own brand (e.g., Office365).
[0081] Malicious activity management operations also enable certain actions to be taken for attack activities based on the overall risk score or suspicion level of the attack activity. Based on the suspicion of the attack activity, actions can be taken. These actions include any of the following: the attack activity pattern is blocked so that future instances no longer reach the user population; alerts that can be triggered are prioritized so that security processes and personnel can quickly fix the problem before the victim surrenders their corporate credentials or installs malware on their device; through a good flow chart, the attack activity can be understood in detail to determine the set of Indicators of Compromise (IOCs) used in the attack. These IOCs can be checked / input into other security products; and an attack activity report can be automatically generated to describe the origin, timeline, intent, and appearance of the attack in plain English and nice charts. Then there is the degree of protection or the degree of protection failure for various aspects of the filtering solution. Then there are which users or devices were targeted and which of them became victims (see further description of the generation of attack activity reports below).
[0082] Reference Figures 10A to 10C and Figure 11AThrough FIGS. 11C, example dashboards with different types of visualizations are provided. Malicious activity management operations also support generating visualizations of malicious activity management operation data. The visualizations can be associated with generating attack activity reports, reports, summaries, and / or any other information that can be transmitted to any number of tenants, customers, or other users. These reports, reports, summaries, and / or any other information can be automatically generated to show what a high-level attack activity looks like, while also allowing a granular view of certain data / attributes associated with the attack activity. An "activity report" can be generated as a user-editable document that describes the attack report in plain English (localized, of course) along with data, charts, and graphs. The user-editable document is generated based on the attributes and information associated with the email activity. This is advantageous because reporting on malicious email messages or activities is a common activity in cybersecurity, where blog posts are typically used to detail various aspects of malicious attacks. Traditional security services provide customers with the attack flow that targets them.
[0083] In Figure 10A it, the dashboard 1010 includes a configurable interface for viewing malicious activity operation data (e.g., alert permissions, classification, data loss prevention, data governance, threat prevention). As shown, charts identifying the total number of emails delivered and the total number of emails clicked can be generated. The malicious activity operation data can be viewed based on different optional tags (e.g., attack activity, email, URL click, URL, and email source). Additional filters provided through the interface can include the attack activity name, attack activity ID, etc. As Figure 10B shown, the dashboard 1020 can show details of a specific attack activity. For example, the Phish.98D458AA attack activity can have an interface element that identifies the IP addresses associated with the attack activity and the number of emails allowed for the attack activity. Additional attributes of the attack activity can also be easily obtained from the attack activity visualization. In Figure 10C it, more granular details of the attack activity are shown in the dashboard 1030, including the attack activity type and name, affected data, and timeline data. As shown in the dashboard 1040, users at risk and related attributes can be easily identified, and the related attributes include recipient actions, sender IP, sender domain, and payload.
[0084] In Figure 11AIn it, the dashboard 1120 includes details of credential theft attack activities. Specifically, for the set of emails associated with the attack activities, the user actions taken (e.g., click, report spam, and delete unread, inbox, delete, and quarantine) can be shown. Attack activities associated with a specific user (such as HVE) can also be visualized to ensure that HVE is not part of a specific attack activity. In Figure 11B In it, the dashboard 1120 includes actions taken by the user (such as secure link block or secure link block override) that can also be shown. In Figure 12 In it, the dashboard 1210 includes the status and impact information of the attack activities (e.g., active status, start date, and impact - 4 inboxes and 1 click). In this way, visualization helps summarize malicious attack activity operation data to facilitate improved access to attack activity information and support additional operations based on the visualization.
[0085] The malicious activity management operation further supports the purge requests from customers, especially requests for them to project their research, knowledge, or other information to their teams and upper management when describing security incidents in their infrastructure. This enables users to understand, feel comfortable, and increase the use of the features and functions described here. Understanding examples of complex cyberattacks that have been blocked or remediated is the key to obtaining a meaningful return on investment (ROI).
[0086] Another aspect of the reporting function is the educational awareness aspect of real-world attacks. From a service perspective, it is necessary to extend attack activity intelligence and visualization beyond the portal and dashboard because these dashboards usually have restricted access. The security operations team and other personnel need to do a lot of manual work to obtain screenshots and collate information for internal communication, post-mortems, etc. The malicious activity management operation supports automatically or otherwise generating attack activity reports and generating a written summary of the cyberattack based on any email attack activity entity. The attack activity reports can be generated in different ways: 1) Generated by the user when viewing the attack activity details page. Behind the "export" button, the "write-up" button can trigger this function; 2) Generated by the user when viewing popular attack activities or security dashboards (context menu -> export write-up) or browsers or attack simulators; 3) Generated by the product team, attack team, pre-sales team, competitive team, which can request a record of the attack in a specific customer environment; 4) Scheduled attack activity reports can be generated and sent via email to one or more users in another user database.
[0087] Malicious activity management operations ensure that attack activity reports follow a consistent template in terms of the document sections and the type and granularity of the data included. This enables users of these reports to predict which sections to add, remove, or modify based on their specific audience for the final report. For example, if a CISO wishes to reach a broader employee population and educate them about a recent phishing incident, specific users who have already been victims of phishing or malware can be excluded, or certain mail flow vulnerabilities that enabled the attack (e.g., allow lists) can be excluded. The same CISO may wish to communicate to the mail management team the misconfigurations exploited by the attack by listing details of specific messages and message routing that evaded detection (e.g., timestamps and message IDs, domains, etc.). When communicating the same attack to other executives, the same CISO may choose to list the specific users affected, how they were affected, the exposure, and the specific measures taken by the internal security organization. The report templates generated by embodiments of the present invention can be considered parameterized text that will be populated with specific sections and data points. Sections of reports generated by embodiments of the present invention can include, but are not limited to: cover page, executive summary, nature of the threat and payload, timeline, propagation, mutation, victimology, impact within the enterprise, and follow-up actions.
[0088] Example method for providing malicious activity management based on malicious activity management operations
[0089] Reference Figure 13 、 Figure 14 and Figure 15 provide flowcharts showing methods for providing malicious activity management based on malicious activity management operations. These methods can be executed using the malicious activity management environment described herein. In an embodiment, one or more computer storage media having computer-executable instructions thereon can cause, when the instructions are executed by one or more processors, the one or more processors to execute the methods in the storage system.
[0090] Turning Figure 3 provides a flowchart showing method 1300 for providing malicious activity management operations. At step 1302, instances of activity in a computing environment are accessed. At step 1304, the instances of activity are processed based on a malicious activity model. The malicious activity model includes a plurality of multi-attribute cluster identifiers. The multi-attribute cluster identifiers include a risk score and a set of attributes that indicate the likelihood that the instances of activity are malicious activities. At step 1306, based on the processing of the instances of activity, it is determined that the instances of activity are malicious activities. At step 1308, a visualization of malicious activity operation data including the instances of activity is generated. The visualization identifies the instances of activity as malicious activities.
[0091] Turning Figure 4, a flowchart showing a method 1400 for providing malicious activity management operations is provided. At step 1402, a plurality of email messages are clustered into clusters based on metadata and attributes associated with the plurality of email messages in the cluster. At step 1404, a risk score is assigned to the cluster based on historical information, attributes, and metadata associated with the plurality of email messages in the cluster. At step 1406, the risk score of the cluster and corresponding information are transmitted.
[0092] Turn to Figure 5 , a flowchart showing a method 1500 for providing malicious activity management operations is provided. At step 1502, each of a suspicion score, an anomaly score, and an impact score is generated for a cluster segment associated with a type of activity. The type of activity is associated with a set of attributes. At step 1504, a risk score is generated based on the suspicion score, the anomaly score, and the impact score. At step 1506, a multi-attribute cluster identifier including the risk score and the set of attributes corresponding to the type of activity is generated. The risk score corresponds to one of the following: high, medium, or low. The risk score and the set of attributes indicate the likelihood that an instance of the activity is a malicious activity.
[0093] Example distributed computing environment
[0094] Now refer to Figure 16 , Figure 16 shows an example distributed computing environment 600 in which implementations of the present disclosure may be employed. Specifically, Figure 16 shows a high-level architecture of an example cloud computing platform 610 that may host a technology solution environment or portions thereof (e.g., a data trustee environment). It should be understood that such and other arrangements described herein are presented only as examples. For example, as described above, many of the elements described herein may be implemented as discrete or distributed components, or combined with other components, and implemented in any suitable combination and location. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, and function groupings) may be used in addition to or instead of those shown.
[0095] The data center can support a distributed computing environment 1600, which includes a cloud computing platform 1610, racks 1620, and nodes 1630 (e.g., computing devices, processing units, or blades) in the racks 1620. The technical solution environment can be implemented with a cloud computing platform 1610 that runs cloud services across different data centers and geographical regions. The cloud computing platform 1610 can implement a structure controller 1640 component for resource allocation, deployment, upgrade, and management for the provision and management of cloud services. Generally, the cloud computing platform 1610 stores data or runs service applications in a distributed manner. The cloud computing infrastructure 1610 in the data center can be configured to host and support the operation of endpoints of specific service applications. The cloud computing infrastructure 1610 can be a public cloud, a private cloud, or a dedicated cloud.
[0096] The node 1630 can be equipped with a host 1650 (e.g., an operating system or a runtime environment) that runs a defined software stack on the node 1630. The node 1630 can also be configured to perform specialized functions (e.g., a computing node or a storage node) within the cloud computing platform 1610. The node 1630 is assigned to run one or more parts of a tenant's service application. A tenant can refer to a customer who utilizes the resources of the cloud computing platform 1610. The service application components of the cloud computing platform 1610 that support a specific tenant can be referred to as tenant infrastructure or tenant. The terms service application, application, or service can be used interchangeably herein and generally refer to any software or part of software that runs on top of the data center or accesses the storage and computing device locations within the data center.
[0097] When the node 1630 supports more than one separate service application, the node 1630 can be partitioned into virtual machines (e.g., virtual machines 1652 and 1654). Physical machines can also run separate service applications simultaneously. The virtual machines or physical machines can be configured as personalized computing environments supported by resources 1660 (e.g., hardware resources and software resources) in the cloud computing platform 1610. It is expected that resources can be configured for specific service applications. In addition, each service application can be divided into functional parts such that each functional part can run on a separate virtual machine. In the cloud computing platform 1610, multiple servers can be used to run service applications and perform data storage operations in a cluster. In particular, the servers can perform data operations independently but present as a single device called a cluster. Each server in the cluster can be implemented as a node.
[0098] The client device 1680 can be linked to the service application in the cloud computing platform 1610. The client device 1680 can be any type of computing device, which can correspond to the reference Figure 17The described computing device 1700, e.g., client device 1680, may be configured to issue commands to cloud computing platform 1610. In an embodiment, client device 680 may communicate with a service application via a virtual Internet Protocol (IP) and a load balancer or other means that direct communication requests to a specified endpoint in cloud computing platform 1610. Components of cloud computing platform 610 may communicate with each other via a network (not shown), which may include, but is not limited to, one or more local area networks (LANs) and / or wide area networks (WANs).
[0099] Example operating environment
[0100] An overview of embodiments of the present invention has been briefly described. The following describes an example operating environment in which embodiments of the present invention may be implemented to provide a general background for various aspects of the present invention. Specifically, initially referring to Figure 17 , Figure 17 FIG. shows an example operating environment for implementing embodiments of the present invention and is generally designated as computing device 1700. Computing device 1700 is only one example of a suitable computing environment and is not intended to impose any limitation on the scope of use or functionality of the present invention. Nor should computing device 1700 be construed as having any dependency or requirement on any one or combination of the illustrated components.
[0101] The present invention may be described in the general context of computer code or machine-usable instructions, including computer-executable instructions, such as program modules, executed by a computer or other machine (e.g., a personal data assistant or other handheld device). Generally, program modules include routines, programs, objects, components, data structures, etc. Code that performs a particular task or implements a particular abstract data type. The present invention may be implemented in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. The present invention may also be implemented in a distributed computing environment where tasks are performed by remote processing devices linked through a communication network.
[0102] Referring to Figure 17 , computing device 1700 includes a bus 1710 that directly or indirectly couples the following devices: a memory 1712, one or more processors 1714, one or more presentation components 1716, an input / output port 1718, an input / output component 1720, and an illustrative power supply 1722. Bus 1710 represents one or more buses (e.g., an address bus, a data bus, or a combination thereof). For the sake of conceptual clarity, Figure 17The various blocks are shown by lines, and other arrangements of the described components and / or component functions are also contemplated. For example, a presentation component such as a display device can be considered an I / O component. Additionally, the processor has a memory. We recognize this as being of the nature of the art and reiterate that Figure 17 the illustration of is merely an illustration of an example computing device that can be used in conjunction with one or more embodiments of the present invention. No distinction is made between categories such as "workstation", "server", "laptop", "handheld device", etc., as all of these are considered to be within Figure 17 the scope of and are referred to as "computing devices".
[0103] Computing device 1700 generally includes various computer-readable media. Computer-readable media can be any available media that can be accessed by computing device 1700 and includes both volatile and nonvolatile media, removable and non-removable media. By way of example and not limitation, computer-readable media may include computer storage media and communication media.
[0104] Computer storage media includes volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD–ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by computing device 1700. Computer storage media does not itself include a signal.
[0105] Communication media typically embodies computer-readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media. The term "modulated data signal" means a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example and not limitation, communication media includes wired media such as a wired network or direct-wire connection, and wireless media such as acoustic, RF, infrared and other wireless media. Any of the foregoing combinations should also be included within the scope of computer-readable media.
[0106] The memory 1712 includes computer storage media in the form of volatile and / or non-volatile memory. The memory can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid state memory, hard disk drives, optical disk drives, etc. The computing device 1700 includes one or more processors that read data from various entities such as the memory 1712 or the I / O component 1720. The presentation component 1716 presents data indications to the user or other devices. Exemplary presentation components include display devices, speakers, printing components, vibration components, etc.
[0107] The I / O port 1718 allows the computing device 1700 to be logically coupled to other devices including the I / O component 1720, some of which may be built-in. Illustrative components include microphones, joysticks, gamepads, dish antennas, scanners, printers, wireless devices, etc.
[0108] Referring to the technical solution environment described herein, the embodiments described herein support the technical solutions described herein. The components of the technical solution environment may be integrated components including a hardware architecture and a software framework that support constrained computing and / or constrained query functions within a technical solution system. The hardware architecture refers to the physical components and their interrelationships, while the software framework refers to the software that provides functions that can be implemented with the hardware implemented on the device.
[0109] A software-based end-to-end system can run within system components to operate computer hardware to provide system functions. At a low level, a hardware processor executes instructions selected from the machine language (also referred to as machine code or native) instruction set of a given processor. The processor identifies the native instructions and executes the corresponding low-level functions, such as functions related to logic, control, and memory operations. Low-level software written in machine code can provide more complex functions to high-level software. As used herein, computer-executable instructions include any software, including low-level software written in machine code, high-level software such as application software, and any combination thereof. In this regard, system components can manage resources and provide services for system functions. Embodiments of the present invention contemplate any other variations and combinations thereof.
[0110] For example, a technical solution system may include an application programming interface (API) library that includes specifications of routines, data structures, object classes, and variables that can support the interaction between the hardware architecture of a device and the software framework of the technical solution system. These APIs include configuration specifications of the technical solution system such that different components therein can communicate with each other within the technical solution system as described herein.
[0111] The various components used herein have been identified. It should be understood that within the scope of this disclosure, any number of components and arrangements can be used to achieve the desired functionality. For example, for clarity of concept, the components in the embodiments shown in the figures are depicted by lines. Other arrangements of these and other components can also be implemented. For example, although some components are depicted as single components, many of the elements described herein can be implemented as discrete or distributed components, or combined with other components, and implemented in any suitable combination and location. Some elements can be omitted entirely. Additionally, the various functions described herein as being performed by one or more entities can be performed by hardware, firmware, and / or software, as described below. For example, the various functions can be implemented by a processor executing instructions stored in a memory. Thus, other arrangements and elements (e.g., machines, interfaces, functions, orders, and groupings of functions) can be used in addition to or instead of those shown.
[0112] The embodiments described in the following paragraphs can be combined with one or more of the specifically described alternatives. In particular, the claimed embodiments can incorporate references to more than one other embodiment. The claimed embodiments can specify further limitations of the claimed subject matter.
[0113] The subject matter of the embodiments of the present invention has been specifically described herein to meet statutory requirements. However, the description itself is not intended to limit the scope of this patent. On the contrary, the inventors have contemplated that the claimed subject matter can also be implemented in other ways, to combine other existing or future technologies, including different steps or combinations similar to the steps described herein. Additionally, although the terms "step" and / or "block" may be used herein to imply different elements of the methods employed, these terms should not be construed as implying any particular order among or between the various steps disclosed herein, unless the order of the individual steps is explicitly described.
[0114] For the purposes of this disclosure, the word "comprising" has the same broad meaning as the word "including", and the word "access" includes "receiving", "referencing", or "retrieving". Additionally, the word "communicating" has the same broad meaning as the words "receiving" or "transmitting", and is implemented via a software- or hardware-based bus, receiver, or transmitter using the communication media described herein. Further, words such as "a" and "an", unless otherwise indicated to the contrary, include both the plural and the singular. Thus, for example, in the presence of one or more features, the constraint of "a feature" is satisfied. Additionally, the term "or" includes conjunctive, disjunctive, and both (thus "a or b" includes "a or b", as well as "a and b").
[0115] For purposes of the foregoing detailed discussion, embodiments of the present invention have been described with reference to a distributed computing environment; however, the distributed computing environment described herein is merely exemplary. Components may be configured to perform novel aspects of the embodiments, where the term "configured to" may refer to "programmed to" perform a particular task or implement a particular abstract data type using code. Further, although embodiments of the present invention generally may relate to the technical solution environments and diagrams described herein, it should be understood that the described technology may be extended to other implementation environments.
[0116] Embodiments of the present invention have been described in connection with specific embodiments, which are illustrative in all respects and not restrictive. Alternative embodiments will become apparent to those of ordinary skill in the art to which the present invention pertains without departing from its scope.
[0117] As can be seen from the foregoing, the present invention is well-suited to achieve all of the above objects and goals, as well as other obvious and inherent advantages of the structure.
[0118] It should be understood that certain features and sub-combinations are useful and may be used without reference to other features or sub-combinations. This is contemplated by the claims and within the scope of the claims.
Claims
1. A computer-implemented method, the method comprising: Accessing an active instance in a computing environment, wherein the active instance comprises a set of attributes; Receiving, at a malicious activity model, the active instance, the malicious activity model comprising a plurality of multi-attribute cluster identifiers associated with previous instances of the activity, the multi-attribute cluster identifiers comprising corresponding risk scores and sets of attributes, the risk scores and the sets of attributes indicating the likelihood that the active instance of the activity is a malicious activity; Wherein the risk score is based on the following: a suspicion score, an anomaly score, and an impact score, each based on information from the cluster indicating a malicious activity; Wherein the suspicion score is based on information about a tenant of the computing environment; Wherein the anomaly score is based on information about the sender; Wherein the impact score is based on the location of one or more individuals targeted by the active instance of the activity in the corresponding cluster; Using the malicious activity model, determining that the active instance of the activity is a malicious activity based on comparing the set of attributes of the active instance of the activity with the plurality of multi-attribute cluster identifiers, wherein the set of attributes of the active instance of the activity matches the set of attributes of at least one of the plurality of multi-attribute cluster identifiers, and the risk score and the set of attributes of the at least one of the plurality of multi-attribute cluster identifiers indicate the likelihood that the active instance of the activity is a malicious activity; And Generating a visualization of malicious activity operation data comprising the active instance of the activity, wherein the visualization identifies the active instance of the activity as the malicious activity.
2. The method according to claim 1, wherein the active instance of the activity is an email message, and wherein generating the malicious activity model for a plurality of email messages comprises: Clustering the plurality of email messages into two or more clusters based on metadata and attributes associated with the plurality of email messages in the cluster; And Assigning a risk score to each of the two or more clusters based on historical information, attributes, and metadata associated with the plurality of email messages in the two or more clusters.
3. The method according to claim 2, wherein clustering the plurality of email messages into the two or more clusters is based on the following: Identifying centroids using metadata attributes of the associated plurality of email messages, wherein the centroids are calculated based on fingerprints of the plurality of email messages; Identifying common metadata attributes among the plurality of email messages; Applying a filter based on the common metadata attributes, wherein applying the filter based on the common metadata attributes identifies two or more subsets of the plurality of email messages; Generating an attack activity guide associated with each of the two or more subsets of the plurality of email messages, wherein the attack activity guide is linked to the common metadata attributes corresponding to the subset of the plurality of email messages.
4. The method according to claim 1, wherein the impact score is further based on the content classification in the corresponding cluster, or the website information included in the previous instance of the activity.
5. The method according to claim 1, wherein the multi-attribute cluster identifier is based on the suspicion score, the anomaly score, the impact score, and the risk score, wherein the suspicion score, the anomaly score, and the impact score are associated with a cluster segment having an activity type, and wherein the activity type is associated with a set of attributes of the activity type.
6. The method according to claim 1, wherein the malicious activity model is a malicious activity determination model, and the malicious activity determination model is used to compare the set of attributes of the instance of the activity to identify at least one of the multi-attribute cluster identifiers, and the at least one multi-attribute cluster identifier indicates the likelihood that the instance of the activity is a malicious activity.
7. The method according to claim 1, wherein the malicious activity model is a machine learning model generated based on processing multiple email messages associated with multiple network attacks, wherein the processing of the multiple email messages is based on clustering the multiple emails, and the clustering is based on: actions taken by the recipients of the multiple emails, attachments in the emails, and corresponding fingerprints of the multiple emails; and wherein the malicious activity model is configurable to process the multiple email messages at the tenant level, the global level, or a combined tenant-global level.
8. The method according to claim 1, wherein the malicious activity operation data is selected from each of the following: cluster analysis, risk score, set of attributes, evaluation results for generating corresponding visualizations of cluster analysis, risk score visualization, set of attributes visualization, and evaluation results visualization.
9. The method according to claim 1, wherein one or more remedial actions are performed based on determining that the set of attributes of the instance of the activity corresponds to a risk score and a set of attributes of at least one of the multi-attribute cluster identifiers, wherein the risk score and the set of attributes indicate the likelihood that the instance of the activity is a malicious activity, and wherein the one or more remedial actions are performed based on the severity of the risk score, and the risk score corresponds to one of the following: high, medium, or low.
10. A system, comprising: one or more computer processors; and a computer memory storing computer-usable instructions that, when used by the one or more computer processors, cause the one or more computer processors to perform operations, the operations including: Access an active instance in a computing environment, where the active instance includes a set of attributes; receive the instance of the activity at a malicious activity model, the malicious activity model including a plurality of multi-attribute cluster identifiers associated with previous instances of the activity, the multi-attribute cluster identifiers including corresponding risk scores and sets of attributes, the risk scores and the sets of attributes indicating the likelihood that the instance of the activity is a malicious activity; where the risk score is based on the following: a suspicion score, an anomaly score, and an impact score, each based on information indicating a malicious activity from the cluster, where the suspicion score is based on information about a tenant of the computing environment, where the anomaly score is based on information about the sender, where the impact score is based on the location of one or more individuals targeted by the instance of the activity in the corresponding cluster, Using the malicious activity model, based on comparing the set of attributes of the instance of the activity with the plurality of multi-attribute cluster identifiers, determine that the instance of the activity is a malicious activity, where the set of attributes of the instance of the activity matches the set of attributes of at least one multi-attribute cluster identifier among the plurality of multi-attribute cluster identifiers, and the risk score and the set of attributes of the at least one multi-attribute cluster identifier among the plurality of multi-attribute cluster identifiers indicate the likelihood that the instance of the activity is a malicious activity; and Generate a visualization of malicious activity operation data including the instance of the activity, where the visualization identifies the instance of the activity as the malicious activity.
11. The system according to claim 10, where the instance of the activity is an email message, and where generating the malicious activity model for a plurality of email messages includes: Cluster the plurality of email messages into two or more clusters based on metadata and attributes associated with the plurality of email messages in the cluster; And Assign a risk score to each of the two or more clusters based on historical information, attributes, and metadata associated with the plurality of email messages in the two or more clusters.
12. The system according to claim 11, where clustering the plurality of email messages into the two or more clusters is based on the following: Use metadata attributes of the associated plurality of email messages to identify a centroid, where the centroid is calculated based on the fingerprints of the plurality of email messages; Identify common metadata attributes among the plurality of email messages; Apply a filter based on the common metadata attributes, where applying the filter based on the common metadata attributes identifies two or more subsets of the plurality of email messages; Generate an attack activity guide associated with each of the two or more subsets of the plurality of email messages, where the attack activity guide is linked to the common metadata attributes corresponding to the subset of the plurality of email messages.
13. The system according to claim 10, wherein the impact score is further based on the content classification in the corresponding cluster, or website information included in the previous instance of the activity.
14. The system according to claim 10, wherein the multi-attribute cluster identifier is based on the suspicion score, the anomaly score, the impact score, and the risk score, the suspicion score, the anomaly score, and the impact score being associated with a cluster segment having an activity type, wherein the activity type is associated with a set of attributes of the activity type.
15. The system according to claim 10, wherein the malicious activity model is a malicious activity determination model, and the malicious activity determination model is used to compare the set of attributes of the instance of the activity to identify at least one of the multi-attribute cluster identifiers among the plurality of multi-attribute cluster identifiers, the at least one multi-attribute cluster identifier indicating the likelihood that the instance of the activity is a malicious activity.
16. The system according to claim 10, wherein the malicious activity model is a machine learning model generated based on processing a plurality of email messages associated with a plurality of network attacks, wherein the processing of the plurality of email messages is based on clustering the plurality of emails, the clustering being based on: actions taken by the recipients of the plurality of emails, attachments in the emails, and corresponding fingerprints of the plurality of emails, wherein the malicious activity model is configurable to process the plurality of email messages at the tenant level, the global level, or a combined tenant-global level.
17. The system according to claim 10, wherein the operation further comprises: Execute one or more remediation actions associated with the instance of the activity, wherein the one or more remediation actions are executed based on the severity of the risk score, and the risk score corresponds to one of the following: high, medium, or low.
18. One or more computer storage media having embodied thereon computer-executable instructions that, when executed by a computing system having a processor and a memory, cause the processor to: Access an instance of an activity in a computing environment, wherein the instance of the activity includes a set of attributes; Receive the instance of the activity at a malicious activity model, the malicious activity model including a plurality of multi-attribute cluster identifiers associated with previous instances of the activity, the multi-attribute cluster identifiers including corresponding risk scores and sets of attributes, the risk scores and the sets of attributes indicating the likelihood that the instance of the activity is a malicious activity; wherein the risk score is based on the following: a suspicion score, an anomaly score, and an impact score, each based on information from the cluster indicating a malicious activity, wherein the suspicion score is based on information about the tenant of the computing environment, wherein the anomaly score is based on information about the sender, wherein the impact score is based on the location of one or more individuals targeted by the instance of the activity in the corresponding cluster. Using the malicious activity model, based on comparing the set of attributes of the instance of the activity with the plurality of multi-attribute cluster identifiers, it is determined that the instance of the activity is a malicious activity, wherein the set of attributes of the instance of the activity matches the set of attributes of at least one multi-attribute cluster identifier among the plurality of multi-attribute cluster identifiers, and the risk score and the set of attributes of the at least one multi-attribute cluster identifier among the plurality of multi-attribute cluster identifiers indicate the likelihood that the instance of the activity is a malicious activity; And Generate a visualization of the malicious activity operation data including the instance of the activity, wherein the visualization identifies the instance of the activity as the malicious activity.
19. The medium according to claim 18, wherein the visualization of the malicious activity operation data supports visually exploring attack activities associated with the instance of the activity, performing batch operations on all emails in the attack activities, identifying indicators of compromise, and identifying tenant policies.
20. The medium according to claim 18, wherein the malicious activity operation data is selected from each of the following: clustering analysis, risk score, set of attributes, evaluation results for generating corresponding clustering analysis visualizations, risk score visualizations, set of attributes visualizations, and evaluation result visualizations.
21. The medium according to claim 18, wherein based on determining that the set of attributes of the instance of the activity corresponds to the risk score and the set of attributes of at least one multi-attribute cluster identifier among the plurality of multi-attribute cluster identifiers, one or more remedial actions are performed, wherein the risk score and the set of attributes indicate the likelihood that the instance of the activity is a malicious activity, and wherein the one or more remedial actions are performed based on the severity of the risk score, and wherein the risk score corresponds to one of the following: high, medium, or low.
22. The medium according to claim 18, wherein the multi-attribute cluster identifier is based on the suspicion score, the anomaly score, the impact score, and the risk score, wherein the suspicion score, the anomaly score, and the impact score are associated with a cluster segment having an activity type, and wherein the activity type is associated with a set of attributes of the activity type.
Citation Information
Patent Citations
External malware data item clustering and analysis
US20180270264A1