Service infrastructure and method for predicting and detecting potential anomalies therein
By generating and analyzing log entries in the service infrastructure, calculating the probability of exceptions using string tables and exception tables, marking suspicious domain names and performing corresponding actions, the problem of difficulty in automatically detecting potential exceptions in the prior art is solved, and rapid response and improved security are achieved.
Patent Information
- Application Number
- CN201911213000.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-11-30
- Filing Date
- 2019-12-02
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2039-12-02
AI Technical Summary
The prior art has difficulty automatically detecting potential abnormalities in service infrastructure, especially abusive signals of abuse and malware, resulting in the inability to take corrective actions in a timely manner.
By accessing the string table, a log entry related to the service infrastructure event is generated and a string searches for strings to determine the probability of exception. If the probability of anomaly exceeds a predetermined threshold, the domain name is marked as suspicious and perform corresponding actions, such as discarding data packets, sending alarms, or blocking data traffic.
It enables rapid detection and prediction of potential anomalies in service infrastructure, reduces losses caused by abuse and malware, and improves system security and reliability.
Smart Images

Figure CN111258796B_ABST
Abstract
Description
Technical Field
[0001] The present technology relates to systems and methods for Internet security. Specifically, the method allows a service infrastructure to predict and detect potential anomalies at the service infrastructure. Background Art
[0002] Data centers and cloud infrastructures integrate many servers to provide mutual hosting services to a large number of clients. A data center can include hundreds of thousands of servers and host millions of domains for its clients. Among these servers, File Transfer Protocol (FTP) servers are used to transfer data, software, and other information elements from clients for storage on the infrastructure. Generally, a domain is associated with an FTP server, although a given FTP server can serve more than one domain, depending on the capacity of the server. Although very useful, the File Transfer Protocol has known security issues.
[0003] Data centers and cloud infrastructures are, for example, often victims of malware that is misinstalled on domains hosted by these systems and causes spam, copyright infringement, or phishing. Although not the only entry point for malware into the data center, malicious third parties sometimes use FTP servers to install malware on hosted domains. A particular difficulty in preventing abuse of the infrastructure is related to the fact that data center operators are generally not allowed to inspect the content of client data received on the FTP servers. In addition to abuse by third parties that have no relationship with the infrastructure owner, it has been found that some clients of the data center are attempting to exploit FTP security flaws to access parts of the infrastructure that are beyond their legitimate content or space.
[0004] After malware is installed on the infrastructure, significant losses may occur. For example, when a third party installs malware in the domain of a legitimate client, the malware may cause anomalies, such as by sending spam. Such anomalies are sometimes detected by other instances outside the infrastructure. These other instances may internally blacklist the legitimate client, thereby causing the legitimate client to lose certain services. Even worse, clients typically purchase a range of consecutive IP addresses. Instances that may blacklist a legitimate client may actually block an entire IP address range.
[0005] In addition to abuse, other types of anomalies may also affect the operation of data centers and cloud infrastructures. Types of anomalies that may occur and be detected include, but are not limited to, operating system crashes, hardware component failures such as server or router failures, etc.
[0006] Parse the activity logs on infrastructure devices to attempt to identify past abuses and build a malware database. However, there is no solution that allows for the automatic detection of signals that are precursors to anomalies in the domains hosted on the infrastructure. In fact, it is considered practical to detect a few anomalies, especially a few abuses, at a fast enough rate to take corrective measures in a timely manner.
[0007] Another related problem is that blocking long-term information related to malware is not necessarily a viable solution. For example, a malicious party may obtain an IP address and use it to install malware. Then, the malicious party can release the IP address, which can later be obtained by a legitimate party. Permanently blocking the IP address may deprive the legitimate party of service.
[0008] Although the recent developments described above can provide benefits, there is still a need for improvement.
[0009] The subject matter discussed in the background section should not be regarded as prior art merely because it is mentioned in the background section. Similarly, problems mentioned in the background section or associated with the subject matter of the background section should not be regarded as having been previously recognized in the prior art. The subject matter in the background section only represents different approaches. Summary of the Invention
[0010] Based on the developers' recognition of the disadvantages associated with the prior art, embodiments of the present technology have been developed.
[0011] In particular, such disadvantages may include the difficulty of identifying the causes of anomalies (including abuses) that may affect the service infrastructure and the inability to take corrective measures in a timely manner due to the delay in detecting anomalies.
[0012] In one aspect, various implementations of the present technology provide a method for detecting potential anomalies at a service infrastructure, including:
[0013] Access a string table, where each corresponding entry in the string table defines a corresponding string and a corresponding anomaly probability for that string;
[0014] Generate a log entry in the database of the service infrastructure related to an event that occurs in the service infrastructure, the log entry including a string specifying one of a file name and an IP address, the log entry including a domain name hosted by the service infrastructure;
[0015] Search for the string in the string table; and
[0016] If the string is found in the string table and if the anomaly probability corresponding to the string exceeds a predetermined threshold, mark the domain name as suspicious.
[0017] In some implementations of the present technology, the method further includes: populating a domain table, where each given entry of the domain table includes: a given domain name, a given string associated with the given domain name in a time frame of interest, and a given association time corresponding to the latest association time in all log entries associating the given domain with the given string.
[0018] In some implementations of the present technology, populating the domain table includes: parsing a plurality of log entries in a database to extract a corresponding domain name, a corresponding string, and a corresponding association time from each log entry.
[0019] In some implementations of the present technology, each corresponding entry of the string table further defines: (i) a corresponding number of domains associated with the corresponding string and having an anomaly in the time frame of interest; and (ii) a corresponding number of domains associated with the corresponding string and not having an anomaly in the time frame of interest.
[0020] In some implementations of the present technology, the method further includes: at a service infrastructure, determining that an anomaly has occurred at a detection time related to an affected domain; accessing an anomaly table, where each entry of the anomaly table includes the name of a domain for which an anomaly has been detected in the time frame of interest and a corresponding anomaly time; if there is an anomaly table entry for the affected domain, updating the corresponding anomaly time in the anomaly table entry with the detection time; and if there is no anomaly table entry for the affected domain, then: creating a new anomaly table entry for the affected domain, the new anomaly table entry including the name of the affected domain and the detection time, extracting a list of strings associated with the affected domain from the domain table; for each string in the list of strings associated with the affected domain, incrementing the number of domains associated with the string and having an anomaly in the time frame of interest in the string table, and for each string in the list of strings associated with the affected domain, decrementing the number of domains associated with the string and not having an anomaly in the time frame of interest in the string table.
[0021] In some implementations of the present technology, the method further includes: periodically updating the time frame of interest such that it extends from a past predetermined duration until the current time; and after updating the time frame of interest: deleting entries in the domain table whose association time is earlier than the time frame of interest, deleting entries in the anomaly table whose anomaly time is earlier than the time frame of interest, and repopulating the string table according to the updated time frame of interest.
[0022] In some implementations of the present technology, repopulating the string table according to the updated time frame of interest includes: for each deleted entry in the domain table: if an entry in the exception table for the domain name present in the deleted entry in the domain table indicates that an exception was found for the domain name within the time frame of interest, decrement in the string table the number of domains associated with the string present in the deleted entry and for which an exception exists within the time frame of interest, and if there is no entry in the exception table for the domain name present in the deleted entry in the domain table, decrement in the string table the number of domains associated with the string present in the deleted entry and for which no exception exists within the time frame of interest; and scanning the domain table to find entries containing the domain names extracted from each deleted entry in the exception table; and for each found string of each found entry in the domain table: decrement in the string table the number of domains associated with the found string present in the deleted entry in the domain table and for which an exception exists within the time frame of interest, and increment in the string table the number of domains associated with the found string present in the deleted entry in the domain table and for which no exception exists within the time frame of interest.
[0023] In some implementations of the present technology, the method further includes: for each respective entry in the string table, calculating a respective anomaly probability based on the following four: (i) the respective number of domains associated with the respective string and for which no anomaly exists within the time frame of interest, (ii) the respective number of domains associated with the respective string and for which an anomaly exists within the time frame of interest, (iii) the number of domains listed in the exception table, and (iv) the total number of domains hosted by the service infrastructure.
[0024] In some implementations of the present technology, each respective anomaly probability is calculated by applying a Bayesian filter to the following four: (i) the respective number of domains associated with the respective string and for which no anomaly exists within the time frame of interest; (ii) the respective number of domains associated with the respective string and for which an anomaly exists within the time frame of interest; (iii) the number of domains listed in the exception table; and (iv) the total number of domains hosted by the service infrastructure.
[0025] In some implementations of the present technology, the method further includes calculating an aggregated probability of an anomaly of the service infrastructure related to one of the domain names in the domain table by: calculating the logarithm of each of the anomaly probabilities of the strings related to the one domain name in the domain name; and combining the logarithms of each of the anomaly probabilities of the strings related to the one domain name in the domain name.
[0026] In some implementations of the present technology, the method further includes calculating an aggregated probability of an anomaly related to one of the domain names in the domain table for the service infrastructure by: calculating a first aggregated probability of an anomaly on the one domain name in the domain name for the service infrastructure for a first subset of strings related to the one domain name in the domain name having the maximum anomaly probability; calculating a second aggregated probability of an anomaly on the one domain name in the domain name for the service infrastructure for a second subset of strings related to the one domain name in the domain name having the minimum anomaly probability; and combining the first aggregated probability and the second aggregated probability.
[0027] In some implementations of the present technology, if the aggregated probability of an anomaly related to a domain name present in the generated log entries for the service infrastructure is below a predetermined threshold, then the domain name is unconditionally marked as non-suspicious.
[0028] In some implementations of the present technology, the method further includes, after marking a domain name as suspicious, performing an action selected from the group including: discarding a received data packet related to the generation of a log entry, sending a message to a client related to the domain name, providing an alert on the service infrastructure, blocking all data traffic related to the domain name, and combinations thereof.
[0029] In other aspects, various implementations of the present technology provide a method for predicting potential anomalies at a service infrastructure, including:
[0030] Defining a string table, each corresponding entry of which defines: (i) a corresponding string, (ii) a corresponding number of domains hosted by the service infrastructure, associated with the corresponding string and having an anomaly during a time frame of interest, (iii) a corresponding number of domains hosted by the service infrastructure, associated with the corresponding string and not having an anomaly during the time frame of interest, and (iv) a corresponding anomaly probability of the string;
[0031] Generating a plurality of log entries in a database of the service infrastructure, each log entry related to an event occurring in the service infrastructure, each log entry associating a string specifying one of a file name and an IP address with a domain name hosted by the service infrastructure, and each log entry further recording a specific association time;
[0032] Parsing the plurality of log entries to populate a domain table, each given entry of which contains a given domain name, a given string associated with the given domain name during the time frame of interest, and a given association time corresponding to the latest association time in all log entries associating the given domain with the given string;
[0033] Detecting an anomaly occurring at a detection time related to an affected domain at the service infrastructure;
[0034] An access exception table, each entry of which includes the name of the domain where an anomaly has been detected in the time frame of interest and the corresponding anomaly time;
[0035] If there is an entry in the exception table for the affected domain, update the corresponding anomaly time in the exception table entry with the detection time; and
[0036] If there is no entry in the exception table for the affected domain, then:
[0037] Create a new entry in the exception table for the affected domain, the new entry including the name of the affected domain and the detection time,
[0038] Extract the list of strings associated with the affected domain from the domain table,
[0039] For each string in the list of strings associated with the affected domain, increment the number of domains associated with the string and having an anomaly in the time frame of interest, and
[0040] For each string in the list of strings associated with the affected domain, decrement the number of domains associated with the string and not having an anomaly in the time frame of interest.
[0041] In other aspects, various implementations of the present technology provide a service infrastructure, including:
[0042] A server configured to receive data packets and / or commands from a client;
[0043] A database configured to store multiple log entries, each corresponding log entry of the log including a corresponding string associated with a corresponding domain name;
[0044] A processor; and
[0045] A memory device including a non-transitory computer-readable medium having executable code stored thereon, the executable code including instructions for performing a method of detecting potential anomalies at the service infrastructure and / or a method of predicting potential anomalies at the service infrastructure when the executable code runs on the processor.
[0046] In the context of this specification, unless otherwise explicitly specified, a computer system may refer to but is not limited to "electronic device", "operating system", "system", "computer-based system", "controller unit", "monitoring device", "control device" and / or any combination thereof suitable for the current relevant task.
[0047] In the context of this specification, unless otherwise explicitly stated, the expressions "computer-readable medium" and "memory" are intended to include any medium of any nature and type, non-limiting examples of which include RAM, ROM, disks (CD-ROM, DVD, floppy disk, hard disk drive, etc.), USB keys, flash memory cards, solid state drives, and tape drives. Still in the context of this specification, "a" computer-readable medium and "the" computer-readable medium should not be construed as being the same computer-readable medium. Instead, at any appropriate time, "a" computer-readable medium and "the" computer-readable medium may also be construed as a first computer-readable medium and a second computer-readable medium, respectively.
[0048] In the context of this specification, unless otherwise explicitly specified, the words "first", "second", "third", etc. have been used as adjectives solely for the purpose of being able to distinguish from each other the nouns they modify, and not for the purpose of describing any particular relationship between these nouns.
[0049] Implementations of the present technology each have at least one of the above-mentioned objectives and / or aspects, but not necessarily all of them. It should be understood that some aspects of the present technology obtained by attempting to achieve the above-mentioned objectives may not meet that objective and / or may meet other objectives not explicitly mentioned herein.
[0050] Additional and / or alternative features, aspects, and advantages of implementations of the present technology will become apparent from the following description, the drawings, and the appended claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] To better understand the present technology and its other aspects and additional features, reference is made to the following description used in conjunction with the drawings, in which:
[0052] Figure 1 A service infrastructure according to an embodiment of the present technology is shown;
[0053] Figure 2 is a sequence diagram showing the operations of a method for detecting potential anomalies at a service infrastructure according to an embodiment of the present technology;
[0054] Figure 3a and Figure 3b is a sequence diagram showing the operations of a method for recording information about detected anomalies in a service infrastructure according to an embodiment of the present technology;
[0055] Figure 4a and Figure 4b is a sequence diagram showing the operations of a method for updating the time frame of interest of a service infrastructure according to an embodiment of the present technology;
[0056] Figure 5 is a sequence diagram showing the operations of a first method for calculating the aggregated probability of anomalies in a service infrastructure according to an embodiment of the present technology;
[0057] Figure 6 is a sequence diagram showing the operations of a second method for calculating the aggregated probability of anomalies in a service infrastructure according to an embodiment of the present technology; and
[0058] Figure 7 is a simplified illustration of an operator display outlining the anomaly probabilities of multiple managed domains according to an embodiment of the present technology.
[0059] It should also be noted that the drawings are not drawn to scale unless explicitly stated otherwise herein. Detailed Description
[0060] The example and conditional language mentioned herein are mainly intended to assist the reader in understanding the principles of the present technology and are not intended to limit the scope of the present technology to the specifically cited examples and conditions. It will be understood that those skilled in the art can design various arrangements that, although not explicitly described or shown herein, still implement the principles of the present technology and are included within the spirit and scope of the present technology.
[0061] In addition, for the sake of helping understanding, the following description may describe relatively simplified implementations of the present technology. As those skilled in the art will understand, various implementations of the present technology may have greater complexity.
[0062] In some cases, examples of modifications that are considered beneficial to the present technology may also be presented. This is done solely for the purpose of assisting understanding, again emphasizing, rather than to limit the scope of the present technology or to elaborate on the boundaries of the present technology. These modifications are not an exhaustive listing, and those skilled in the art can make other modifications that are still within the scope of the present technology. Additionally, in the absence of presenting examples of modifications, it should not be construed that no modifications can be made and / or that the described content is the only way to implement the element of the present technology.
[0063] Furthermore, all statements herein reciting the principles, aspects, and implementations of the present technology, as well as specific examples thereof, are intended to include their structural and functional equivalents, whether currently known or developed in the future. Thus, for example, those skilled in the art should recognize that any block diagram herein represents a conceptual view of an illustrative circuit embodying the principles of the present technology. Similarly, it will be understood that any flowchart, flow graph, state transition diagram, pseudocode, etc. represents various processes that can be substantially represented in a computer-readable medium and thus executed by a computer or processor, whether or not such a computer or processor is explicitly shown.
[0064] The functions of the various elements shown in the drawings can be provided by using dedicated hardware and hardware capable of executing software in association with appropriate software, including any functional block labeled "processor". When provided by a processor, the functions can be provided by a single dedicated processor, by a single shared processor, or by multiple individual processors (some of which may be shared). In some embodiments of the present technology, the processor can be a general-purpose processor such as a central processing unit (CPU) or a processor dedicated to a specific purpose such as a digital signal processor (DSP). Additionally, the explicit use of the term "processor" should not be construed as excluding hardware capable of executing software, and can implicitly and non-exclusively include an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a read only memory (ROM) for storing software, a random access memory (RAM), and a non-volatile storage device. Other conventional and / or custom hardware can also be included.
[0065] Software modules (or modules that are only implied to be software) can be represented herein as any combination of flowchart elements or other elements indicating the performance of process steps and / or text descriptions. These modules can be executed by hardware that is explicitly or implicitly shown. Additionally, it should be understood that a module can include, for example but not limited to, computer program logic, computer program instructions, software, a stack, firmware, a hardware circuit, or a combination thereof that provides the required capabilities.
[0066] By appropriately leveraging these basic principles, some non-limiting examples will now be considered to illustrate the various implementations of aspects of the present technology.
[0067] This technology solves some problems caused by the lack of previous solutions for automatically detecting signals before service infrastructure anomalies. A technology is introduced through which device logs (such as File Transfer Protocol (FTP) server logs) are stored in the service infrastructure and are regularly maintained to cover a time frame of interest, such as a time period of up to one (1) month until now. The logs can contain strings identifying, for example, file names or IP addresses. Analysis is performed to evaluate the probability of an anomaly. When a string related to the current event (e.g., receiving a data packet or command from a client or from a malicious party posing as a client) is found in the table, an anomaly is suspected, and the string is marked as related to an anomaly such as abuse of a domain hosted by the service infrastructure. Unless otherwise stated, in one aspect, when some of these data packets or commands can be related to recently occurring anomalies, this technology can rely on the similarity between the data packets or commands received at the service infrastructure. Thus, the technology includes a learning function that allows anomalies patterns to be associated over time. In particular, instead of searching for a specific cause of an anomaly, such as malware or a faulty router, this technology attempts to identify the user who has recently been associated with an anomaly, which is identified by the IP address or file name contained in the data packets or commands received from these users. Potentially legitimate but incorrect commands entered by the operator of the service infrastructure or external participants can also result in anomalies that this technology may detect. Although the examples provided herein mention data packets and commands, the anomaly probability of other types of events that result in log generation in the service infrastructure can be evaluated. The mention of data packets and commands is for illustrative purposes and does not imply a reduction in the generality of this disclosure.
[0068] It is believed that eventually a small number of abuses will be detected, or a small number of abuses will be detected quickly enough to take corrective measures in a timely manner. The learning function provided by this technology allows the scope of anomaly search to be extended and potential anomaly events to be detected quickly.
[0069] In one aspect of this technology, a Bayesian filter can be used to calculate the probability of an anomaly. Although false positive detections of anomalies may occur, the Bayesian filter takes into account the probability that a given string may occur legitimately.
[0070] When the domain or IP address of a legitimate user is used by a third party, automatically blocking the domain or IP address may result in blocking the legitimate user. In some regions, this may not be possible for legal reasons, and in any case, this may damage the relationship between the infrastructure owner and the client. Thus, in an implementation, information about a suspected anomaly may be provided on an operator interface instead of automatically blocking the domain or IP address, or information about a suspected anomaly may be provided on an operator interface before automatically blocking the domain or IP address. In any case, the present technology does include the possibility of blocking entities at the source of the anomaly as an optional feature.
[0071] Figure 1 FIG. 4 shows a service infrastructure 100 according to an implementation of the present technology. The service infrastructure 100 includes a number of servers 110 (only one is shown for simplicity of illustration) such as an FTP server that provides services to a client 120. The server 110 is the host of the domain of the client 120. The server 110 generates logs that are stored in a database 130. Other components of the service infrastructure 100 include a parser 140, a detector 150, a user interface 160, a signaling interface 170, and a blocker 180. The parser 140 may parse the logs stored in the database 130 and generate a result table stored in the database 130. The detector 150 detects anomalies, either internally notified by an external detection system 190. The detector uses the table stored by the parser 140 in the database 130 to calculate an anomaly probability stored in another table of the database 130. The detector 150 is also notified of data packets and commands received from the client 120 at the server 110 such as FTP commands. The detector 150 identifies domains that may be considered suspicious in view of suspicious information in the received data packets or commands. The user interface 160 allows communication between the detector 150 and an operator of the service infrastructure 100. The signaling interface 170 allows the detector 150 to communicate with the client 120 and with the external detection system 190.
[0072] More specifically, the database 130 stores, for each server 110, a log containing information about various events and activities of the server 110. Examples of such events and activities include commands received at the server 110, computations performed at the server 110, etc. Some log entries are generated when a data packet or command is received at the server 110 from the client 120, or when other events occur within the time frame of interest. For example, the log entries may be stored in plain text format. For each data packet or command received at the service infrastructure 100, an entry in the log may include: the date or date and time of the event, the domain hosted by the service infrastructure 100 that is the intended recipient of the command or data packet, a string (or simply "string") representing the file name included in the received data packet or command, and / or a string representing the IP address included in the received data packet or command.
[0073] The above-mentioned table includes a "domain table" generated by the parser 140 and stored in the database 130. The "domain table" provides information about the occurrences of strings related to file names or IP addresses in the following data packets or commands that are received at domains hosted by the service infrastructure 100 during the time frame of interest (e.g., in a time period extending from one month ago until the current time). For example, the domain table can be populated by parsing the logs of servers such as FTP servers that serve domains hosted by the service infrastructure 100. The domain table can alternatively be updated in real time as the service infrastructure 100 processes data traffic and commands. An example of the domain table is shown in Table I.
[0074] Table I
[0075] Domain Name String Date of the Last Occurrence of the String Domain 1 String_a August 26, 2018 Domain 1 String_b September 11, 2018 Domain 1 String_c September 22, 2018 Domain 2 String_c September 25, 2018 Domain 2 String_d September 20, 2018 Domain 3 String_f September 25, 2018 Domain 4 String_b September 14, 2018
[0076] Table I reflects the content of the domain table on September 25, 2018. As shown in Table I, the domain table contains information related to the occurrences of various strings found in the log. Each entry (each row) of the domain table includes three (3) relevant fields. The first field contains the domain name, and the second field contains the string, where the domain and the string are found in a common entry of the log. In the example shown, the third field contains a time value that can be represented as the "date of the last occurrence of the string" on that domain. The time value can be represented in other ways as the number of days since the string last occurred on that domain, in which case each entry is incremented regularly, for example, by one (1) day per day, so that the time since the last occurrence of the string is correctly reflected. If multiple entries in the log include a pair formed by the domain name and the string, the latest entry among these entries is used to populate the "date of the last occurrence of the string" field in the entry of the domain table. Although only a few rows are shown in Table I, in practical applications, the number of domains will be much larger, such as hundreds or thousands or millions of domains, and the number of strings will be on the order of, for example, millions or even billions of strings. In an embodiment, the domain name is used as the main search keyword to search for relevant information in the domain table.
[0077] Entries that are earlier than a predetermined number of days corresponding to the time frame of interest are deleted from the domain table. For example but not limited to, a predetermined number of days can be selected for deleting entries that are earlier than 30 days. Entries of strings that are not relevant to a given domain in the log are deleted from the domain table during the time frame limit that needs to maintain the total number of entries to be retained. It can be considered that hackers who generate anomalies often modify their IP addresses and file names in an attempt to bypass conventional Internet security solutions to set the duration of the time frame. It is envisioned to use longer or shorter intervals corresponding to longer or shorter time frames to delete entries from the domain table. Although the duration of the time frame can be expressed in days, the domain table can be checked at a higher rate (e.g., once every few minutes) to delete entries that are no longer valid, so as to spread the processing load related to the deletion operation over time.
[0078] Table II is an example of an "anomaly table" generated by detector 150 and stored in database 130.
[0079] Table II
[0080] Domain Name Date of the Last Exception Domain 1 September 15, 2018 Domain 2 September 25, 2018 Domain 4 August 26, 2018 Domain 5 September 3, 2018
[0081] Table II reflects the content of the exception table as of September 25, 2018. Each entry (each row) of the exception table contains a domain name field and a "date of last exception" field. In a variant, the time (i.e., hour) of the last exception on a given domain can also be inserted into the "date of last exception" field. In another variant, this field can indicate the number of days since an exception was detected for the given domain. In the example of Table II, the service infrastructure 100 serves five (5) domains named "Domain 1", "Domain 2", "Domain 3", "Domain 4", and "Domain 5". There is no entry for "Domain 3" in the exception table because there were no exceptions related to that domain during the time frame of interest. As a background operation, entries older than the time frame can be periodically (e.g., once a day) deleted from the exception table. As in the case of the domain table, the number of entries in the exception table will be much larger in a practical application, with the exception table having entries for hundreds, thousands, or millions of domains. Even in a large data center hosting millions of domains, the exception table is relatively small compared to other tables, and the search for domains that have recently been found to be exceptions is expected to be completed in a few milliseconds.
[0082] In the exception table, the number of entries represents the number of domains for which an exception has been detected during the time frame. The number of domains without exceptions can be inferred based on knowledge of the total number of domains hosted in the service infrastructure 100. As an alternative, the exception table can contain an entry for each domain hosted by the service infrastructure 100, in which case the "data of last exception" field for a given domain will be set to a Null value to indicate that no exception was associated with that given domain during the time frame.
[0083] Table III is a "string table" generated by the detector 150 and stored in the database 130.
[0084] Table III
[0085] String Existence of Domains without Exceptions Existence of Domains with Exceptions <![CDATA[p′(A|Str i )]]> String_a 106 58 0.935 String_b 58 22 0.905 String_c 45 17 0.893 String_d 1 4 0.882 String_e 50 15 0.871 String_f 1 0 0.000
[0086] Table III reflects the content of the string table as of September 25, 2018. Each entry (each row) of the string table includes: a first field providing a string name, a second field showing the number of domains hosted in the service infrastructure 100 in which the string has been detected as part of a data packet or command received and in which there were no exceptions during the time frame, a third field showing the number of domains in which the string has been detected as part of a data packet or command received and in which there were exceptions during the time frame, and the probability p′(A|Str i ) of an exception possibly occurring on the service infrastructure 100, assuming the string Str represents a file name or an IP address i that has been associated with at least one domain hosted on the service infrastructure 100 during the time frame of interest. The calculation of the value p′(A|Str i) In this way. The number of entries in the string table will be much larger in actual applications, with the string table having entries for up to billions of strings. The string table is indexed by strings rather than creation dates. Thus, searches for specific strings can be quickly performed with a moderate amount of processing work.
[0087] The content of the domain table can be updated by the parser 140 in the following manner. The parser 140 parses the log periodically (e.g., but not limited to, once a day, once an hour) or even continuously and effectively in real time. For each association between a domain name and a string (i.e., an IP address or a file name), the parser 140 updates the corresponding entry in the domain table or creates an entry if it did not previously exist in the domain table. The "date of the last occurrence of the string" field of a given entry is set to the actual date stored in the log; if more than one log entry is found for the same domain name and the same string, the date of the most recent log entry among these log entries is placed in the corresponding entry in the domain table.
[0088] When the parser 140 detects a new string that was not a previous part of the string table, the parser 140 creates a new entry in the domain table to associate the new string with the domain name related to that string, and also adds the association time to this new entry. The parser 140 also creates a new entry in the string table. In this entry, the "presence of non-exception domains" field is initially set to one (1), the "presence of exception domains" field is initially set to zero (0), and the probability p′(A|Str i ) is initially set to zero (0).
[0089] The parsing of the log can be performed at a fast rate to detect new strings that appear on the service infrastructure 100, and at a lower rate (e.g., once a day) to delete entries in the domain table that are no longer within the time frame of interest.
[0090] The same file name may result in multiple entries being generated. As a non - limiting example, the log may show that a 'PUT' FTP command is received to store an image in a folder of a domain. The FTP command may read as follows: PUT my_image.jpg / www / img / my_nicest_pics / my_image.jpg. This command requests that an image named'my_image' with the extension 'jpg' be placed in the folder named'my_nicest_pics' within the path ' / www / img' of the domain owned by the issuer of the PUT command. Each of 'www', 'img','my_nicest_pics','my_image.jpg', and'my_image' may constitute a different string for placing different entries in the domain table, and for the same domain hosting the folder / www / img / my_nicest_pics / , each of these entries in the domain table has the same value in the 'date of the last occurrence of the string' field. Some modifications may be made to the content of the PUT command in view of restricting the number of similar strings, for example, replacing uppercase letters with lowercase letters, replacing consecutive spaces with a single space, etc.
[0091] Continuing with this example, if the PUT command is received on September 25, 2018, all entries are filled with this date in the 'date of the last occurrence of the string' field. If a similar command requesting the'my_photograph.jpg' entry in the same folder is received the next day for the same domain, new entries are created in the domain table for'my_photograph.jpg' and'my_photograph' with the date of September 26, 2018. Now the entries for 'www', 'img','my_nicest_pics' have the date of September 26, 2018. The entries for'my_image.jpg' and'my_image' still have the date of September 25, 2018.
[0092] Entities of the service infrastructure 100 can detect anomalies in real - time. The entities of the service infrastructure 100 receive information about the anomaly from the following: received from an operator via the user interface 160, received from the client 120 where the anomaly is detected, received from a batch process implemented in the detector 150, or received via the signaling interface 170 from a detection system 190 that communicates with the detector 150. Regardless of how the anomaly is detected, the detector 150 updates the content of Table II and Table III to account for the new anomaly.
[0093] Referring to Tables I, II, and III, in a non-limiting example, an anomaly occurring on September 25, 2018 was detected, and the anomaly relates to Domain 2. Check the anomaly table to determine if there is an entry for Domain 2. If an entry exists, the "Date of Last Anomaly" field will be updated to September 25, 2018 without taking any other action. If the anomaly table does not have an entry for Domain 2, a new entry for Domain 2 is created in the anomaly table. Then, the following string table is updated: (i) search the domain table for all strings related to Domain 2 (in this example, Domain 2 is related to String_c and String_d), (ii) decrement the "Presence of Domain without Anomaly" for each of String_c and String_d in the string table, and (iii) increment the "Presence of Domain with Anomaly" for each of String_c and String_d in the string table. Although not required, the values p′(A|Str i ) of String_c and String_d can optionally be updated at this time.
[0094] In another non-limiting example, it was found that String_f was associated with Domain 3 on September 25, 2018. String_f is a new string that did not previously exist in the domain table and the string table, which means that String_f did not appear on the service infrastructure 100 during the time frame. A new entry for Domain 3 and String_f is created in the domain table. A new entry for String_f is also created in the string table. Assuming that no anomaly was indicated to occur on Domain 3 during the time frame, the "Presence of Domain without Anomaly" of String_f is set to one (1), and the "Presence of Domain with Anomaly" is set to zero (0). This reflects that String_f is only related to Domain 3 and no anomaly was detected in Domain 3 during the time frame. In the case where String_f has already been associated with Domain 3 in the domain table, the only action is to update the date of the entry to September 25, 2018.
[0095] In another non-limiting example, the domain table was scanned on September 26, 2018, and it was found that String_a was not associated with Domain 1 after August 26, 2018 or 31 days ago. In this example, the last occurrence of String_a associated with Domain 1 exceeded the predetermined number of days of 30 days corresponding to the time frame. The entries for Domain 1 and String_a were deleted from the domain table. An anomaly was indicated to have been recently detected on Domain 1 on September 15, 2018. String_a is no longer associated with Domain 1, and now String_a is associated with one less domain in which an anomaly was detected. Therefore, the "Presence of Domain with Anomaly" of String_a is decremented by one (1). In the case where no anomaly was detected on Domain 1 during the time frame, instead, the "Presence of Domain without Anomaly" of String_a will be decremented by one (1).
[0096] Continuing with the above non - restrictive example that occurred on September 26, 2018, scan the anomaly table and find that no anomalies were found on domain 4 since August 26, 2018 or 31 days ago, which in this example exceeds the predetermined number of 30 days corresponding to the time frame. Entries for domain 4 are deleted from the anomaly table. Scan the domain table and note that String_b was associated with domain 4 on September 14, 2018, which is shorter than the predetermined number of 30 days after the last anomaly on domain 4. In the string table, until now, the "presence of domains with anomalies" for String_b has reflected the presence of String_b on domain 4 during the time period when anomalies were found on domain 4. Now it is shown that the presence of String_b on domain 4 on September 14, 2018 has nothing to do with any anomalies after August 26, 2018. Assuming no anomalies are detected in domain 4 within the time frame, now String_b is related to one less domain where anomalies exist. Therefore, the "presence of domains with anomalies" for String_b decreases by one (1), and the "presence of domains without anomalies" increases by one (1). Repeat this process for all strings related to domain 4 in the domain table and for all domains in the anomaly table that may show no anomalies within the time frame.
[0097] Returning to the above example where String_f was associated with domain 3 on September 25, 2018, in the case where no other data packets or commands carrying String_f are received for any domain within the next 30 days, the entry connecting domain 4 and String_f will be deleted from the domain table on October 25, 2018. At the same time, the entry for String_f will be deleted from the string table because otherwise the "presence of domains without anomalies" will decrease to zero (0), and the "presence of domains with anomalies" will still be equal to zero (0). Generalizing this action, it can be shown that when there are no longer any entries for a given string in the domain table, the corresponding entry for that given string can be deleted from the string table.
[0098] As described below, a Bayesian filter can be used to perform the calculation of the probability p′(A|Str i ) of the string table.
[0099] First, assume that a string Str representing a file name or an IP address i has been associated with at least one domain hosted on the service infrastructure 100 within the time frame of interest. Then, the probability p(A|Str i ) of an anomaly possibly occurring on the service infrastructure 100 is calculated according to equation (1):
[0100]
[0101] Where:
[0102] p(A) is the number of domains on which an anomaly occurs in a time frame divided by the total number of domains;
[0103] is the number of domains on which no anomaly occurs in a time frame divided by the total number of domains;
[0104] p(Str i |A) is the probability that the string Str representing a file name or an IP address in a time frame i is associated with the domain of the service infrastructure 100 on the domain where an anomaly has occurred; and
[0105] is the probability that the string Str representing a file name or an IP address in a time frame i has been associated with the domain of the service infrastructure 100 on the domain where an anomaly has occurred.
[0106] In a non - limiting example, the anomaly table may show that the service infrastructure 100 hosts 1000 domains, and anomalies are detected for 50 of these domains during a time frame. Thus, p(A) equals 0.050, and equals 0.950. The given string Str can be found in the domain table i associated with 35 domains on which an anomaly occurs and 5 domains on which no anomaly occurs. In this case, p(Str i |A) equals 35 / 50 = 0.700, and equals 5 / (1000 - 50)=0.0052. Then, using equation (1), the probability p(A|Str i ) of the given string Str i equals 0.876, which means that if the given string Str i is received at the service infrastructure 100, there is an 87.6% chance of an anomaly occurring.
[0107] In a variant, instead of calculating the values p(A) and based on the content of the anomaly table, i fixed values can be assigned. For example, when the value p(Str |A) is very far from , setting p(A) equal to 0.100 and setting i equal to 0.900 tends to increase the value of the probability p(A|Str i ). Using these fixed values in the above example, the probability p(A|Str i ) will equal 0.937, which means that if the given string Str
[0108] Using equation (1) to calculate the probability p(A|Str i ) may be affected by the fact that some strings occur rarely. A string may appear only once in a time frame and may be related to an abnormal event on the domain to which it is associated. The string may or may not be related to the cause of the abnormality. For this string, p(Str i |A) is equal to 1.000, and is equal to 0.000. For p(A) and For any value of i ) becomes 1.000, or 100%. A weighted probability can be calculated using equation (2) to mitigate this effect:
[0109]
[0110] in:
[0111] n i Is a string Str with a file name and IP address associated with it i The number of different domains, n i is equal to the sum of the "existence of no anomalous domain" and "existence of anomalous domain" of the string; and
[0112] k is a constant weighting parameter value.
[0113] Without limitation, it has been found that when the log is a log of an FTP server, the weight parameter value k can be set to three (3) with good results. i The above example is associated with 35 domains on which the anomaly occurs and 5 domains on which the anomaly does not occur, n i Equal to 40 domains, probability p(A|Str i ) changes from 0.269 to 0.250 (if p(A) and using real values), or from 0.438 to 0.407 if p(A) and are set to fixed values 0.100 and 0.900 respectively). Except for the given string Str i The number of associated domains n i Except for very small cases, the effect of using a weighting parameter value k is not significant. Other weighting parameter values can be easily selected for other applications, for example by trial and error.
[0114] Each time an entry is modified in the string table, the probability p(A|Str i ) or weighted probability p′(A|Str i ) may be recalculated, or alternatively the entire string table may be recalculated at regular intervals, for example every few minutes.
[0115] In an embodiment, instead of deleting information older than the time frame from Tables I, II, and III, the same data can be moved to additional tables for m earlier time frames. The calculation of the probability of an anomaly being generated by a domain over a longer time period can be extended using Equation (3) or equivalently using Equation (4):
[0116] p″(A|Str i ) t = p′(A|Str i ) t + δp′(A|Str i ) t-1 + δ 2 p′(A|Str i ) t-2 + δ 3 p′(A|Str i ) t-3 +...(3)
[0117]
[0118] Where:
[0119] p″(A|Str i ) t is the time-weighted probability of an anomaly possibly occurring on the service infrastructure 100, assuming that the string Str representing the file name or IP address is associated with at least one domain hosted on the service infrastructure 100 in time frame t or one or more earlier time frames t - m; i δ is a decay factor greater than 0 and less than 1 used to multiply the weighted probability p′(A|Str
[0120] i i ) calculated in the earlier time frame t - m; and
[0121] H is an integer greater than or equal to 1 and specifying the m earlier time periods on which the weighted probability is calculated.
[0122] Thus, H represents the time range over which the adjusted probability of a domain generating an anomaly can be calculated. Using a small value for δ, δ m rapidly decreases, and for example, H can be restricted to a range between three (3) and five (5), such as three (3) to five (5) earlier time periods per month (1).
[0123] In a variant, the contents of Tables I, II, and III earlier than a time frame can be copied to equivalent Tables I', II', and III' for previous iterations, which are maintained in the service infrastructure 100 for a longer period such as one year and are reused to calculate the time-weighted probability p″(A|Str i ) t . The contents of these tables no longer change. In another variant, when performing the operation of removing information earlier than the time frame of interest from Tables I, II, and III, the current value of p(A|Str i ) or p′(A|Str i ) can be stored in the service infrastructure 100 together with the current date.
[0124] Equations (1), (2), (3), and (4) provide various ways to calculate the anomaly probability of a given string Str i representing a file name or IP address of a given domain. These probabilities can be aggregated in the following Equation (5) to provide the overall probability of anomalies for a given domain D:
[0125]
[0126] where:
[0127] I D is the total number of strings Str i that have been associated with the domain D within the time frame; and
[0128] p(A|D) is the aggregated probability of anomalies affecting the domain D within the time frame.
[0129] Although Equation (5) as shown combines the values of the weighted probability p′(A|Str i ), variants can combine the probability p(A|Str i ) or the time-weighted probability p″(A|Str i ) t .
[0130] The number of strings Str i that may be associated with the domain D can be large. Although there may often be anomaly attacks in the service infrastructure 100, for the vast majority of strings, the probability p′(A|Str i that any given string Str i) are expected to be small. Using floating - point arithmetic to calculate the aggregated probability p(A|D) of equation (5) may cause various products in this equation to become zero (0) because many intermediate calculation results are rounded. For this reason, in a variant, equations (6) and (7) can be used to calculate the aggregated probability p(A|D) of anomalies affecting domain D in a time frame:
[0131]
[0132] Where:
[0133]
[0134] In another variant, the computational complexity can be reduced and rounding of calculations can be avoided by calculating the aggregated probability p(A|D) for a small number of strings Str i with the lowest p′(A|Str i values for domain D (e.g., a subset containing 20 strings) and for another small number of strings Str i with the highest p′(A|Str i values for domain D (e.g., a subset containing 20 strings).
[0135] Figure 2 is a sequence diagram showing the operations of a method for detecting potential anomalies at a service infrastructure according to an embodiment of the present technology. Figure 2 Shows a sequence 200 including a plurality of operations that can be executed in a variable order, some operations can be executed simultaneously, and some operations are optional. Sequence 200 can start with operation 205, in which detector 150 provides / accesses a string table, where each corresponding entry in the string table defines a corresponding string and the corresponding anomaly probability of that string. Each corresponding entry in the string table can also define the number of corresponding domains in which there is an anomaly and the number of corresponding domains in which there is no anomaly in the time frame of interest, associated with the corresponding string. A log entry is generated at operation 210. This log entry is related to an event that occurs in service infrastructure 100. Examples of events that can cause a log entry to be generated include, but are not limited to, receiving a data packet or a command (e.g., an FTP command) at infrastructure 100. The log entry includes a string specifying one of the file names and IP addresses included in the data packet, the command, or any other event that causes the log entry to be generated. The log entry also includes the domain name hosted by service infrastructure 100, which is associated with the event that causes the log entry to be generated.
[0136] Database 130 can contain multiple log entries, each log entry defining the association between a specific string and a specific domain name, and each log entry also records the specific association time.
[0137] In one embodiment, there may be multiple strings in the log entry generated at operation 210, each string specifying a different one of the different file names or IP addresses included in the log entry. The following operations of sequence 200 may be repeated independently for each string present in the log entry.
[0138] Optionally, the domain name, string, and current generation time of the log entry at operation 210 can be used to update the domain table by creating a new entry or, if the corresponding entry was previously part of the domain table, refreshing the data of the last occurrence of the string. If there is more than one string in the log entry, the same number of entries can be updated or created in the domain table.
[0139] At operation 215, detector 150 accesses the string table and searches for a string in the string table (or, independently, for each string extracted from the data packet or command). Operation 220 determines whether the string is found in the log entry. If not found, sequence 200 ends at operation 225. If the string is found in the string table, detector 150 determines at operation 230 whether the anomaly probability corresponding to the string exceeds a predetermined threshold. The predetermined threshold can be a fixed threshold or a variable threshold. A typical but non-limiting example of a fixed threshold can be set to 0.8, that is, the probability that the string is related to an anomaly is set to 80%.
[0140] Different thresholds can be used for different types of anomalies, including, for example, abuse caused by malware or errors caused by failures of hardware components. The threshold may vary over time in view of the recently discovered abuse probability. For example, by using a very high threshold, false positive detections can be largely avoided. When a wider identification of potential anomalies is needed, a lower threshold can be used. The operator of service infrastructure 100 can use user interface 160 to modify the threshold.
[0141] In one embodiment, the threshold can be calculated based on the data accumulated from the most recent events occurring at the service infrastructure using equation (8):
[0142]
[0143] where,
[0144] I is the indicator function;
[0145] m is the number of entries in the string table.
[0146] Equation (8) calculates the minimum value x such that the 95% exception probability is less than x (the 95th percentile). If more exceptions are desired to be detected (at the risk of getting more false positives), this 95% ratio can be decreased. The ratio can be increased to decrease the number of false positives.
[0147] If the exception probability exceeds a predetermined threshold, then at operation 240 the domain name is marked as suspicious. If the threshold is not exceeded, then sequence 200 ends at operation 235. If the domain name has been marked as suspicious, detector 150 can optionally cause one or more of the following actions to be performed at operation 245: the received data packet related to generating a log entry at operation 210 can be discarded by the server 110 that has received the data packet, and as indicated by detector 150 via blocker 180, detector 150 can cause a message such as a text or an email to be sent to the client related to the domain name via signaling interface 170, detector 150 can cause an alert to be provided on user interface 160 of service infrastructure 100, and / or detector 150 can cause server 110 to block some or all data traffic related to the domain name. In any case, sequence 200 ends at operation 250.
[0148] Figure 3a and Figure 3b is a sequence diagram showing operations of a method of recording information about exceptions detected in a service infrastructure according to an embodiment of the present technology. Figure 3a and Figure 3b shows sequence 300 including a plurality of operations that can be performed in a variable order, some operations can be performed simultaneously, and some operations are optional. At operation 305, parser 140 can parse a plurality of log entries stored in database 130 to populate a domain table, as shown in Table I, in which each given entry includes a given domain name, a given string associated with the given domain name in a concerned time frame, and a given association time corresponding to the latest association time in all log entries associating the given domain with the given string. Considering that instead of or in addition to operation 305, the domain table can be populated in real time when a log entry is generated at operation 210 ( Figure 2 ), so this operation 305 is optional. Sequence 300 continues with operation 310, in which detector 150 provides / accesses an exception table, as shown in Table II, in which each entry includes the domain name for which an exception has been detected in a concerned time frame and the corresponding exception time. At operation 315, service infrastructure 100 determines that an exception related to the domain affected by the exception has occurred at the corresponding exception time. Detector 150 can make this determination autonomously. Optionally, detector 150 can rely on the operation at 230 ( Figure 2) The detection performed at [operation 210] determines that an anomaly has occurred in the domain name identified in the log entry generated at operation 210 when the anomaly probability corresponding to the string identified in the log entry generated at operation 210 exceeds a predetermined threshold. Alternatively, at sub-operation 317, the detector 150 can receive an anomaly indication for the domain affected by the anomaly from the detection system 190 via the signaling interface 170.
[0149] At operation 320, the detector 150 determines whether there is an anomaly table entry for the affected domain. If so, at operation 325, the detector 150 updates the corresponding anomaly time in the anomaly table entry with the detection time, and the sequence ends at operation 330.
[0150] If there is no anomaly table entry for the affected domain, then at operation 335, the detector 150 creates a new anomaly table entry for the affected domain. The new anomaly table entry includes the name of the affected domain and the detection time. Sequence 300 continues after operation 335 at Figure 3b At operation 340, the detector 150 extracts a list of strings associated with the affected domain from the domain table. At operation 345, for each string in the list of strings associated with the affected domain, the detector 150 increments the number of domains associated with that string and having an anomaly in the time frame of interest in the string table. At operation 350, for each string in the list of strings associated with the affected domain, the detector 150 decrements the number of domains associated with that string and not having an anomaly in the time frame of interest in the string table.
[0151] Then at operation 355, for each entry in the string table corresponding to one of the lists of strings associated with the affected domain, the detector 150 can calculate the corresponding anomaly probability based on: (i) the number of domains associated with that string and not having an anomaly in the time frame of interest; (ii) the number of domains associated with that string and having an anomaly in the time frame of interest; (iii) the number of domains listed in the anomaly table, and (iv) the total number of domains hosted by the service infrastructure 100. The detector 150 can use a Bayesian filter to calculate the anomaly probability. Then, sequence 300 ends at operation 360.
[0152] In one embodiment, the anomaly probability of the strings associated with the affected domain can be calculated at regular intervals, and the detector 150 performs an optional batch processing of the string table instead of following operations 345 and 350. In the same or another embodiment, the anomaly probability of the strings present in the log entry generated at operation 210 ( Figure 2 ) can also be calculated in real time by the detector 150 when processing data packets or commands.
[0153] Figure 4a and Figure 4b is a sequence diagram showing operations of a method for updating a service infrastructure 100's time frame of interest according to an embodiment of the present technology. Figure 4a and Figure 4b shows a sequence 400 including a plurality of operations that can be executed in a variable order, some operations can be executed simultaneously, and some operations are optional. The time frame of interest can be regarded as a sliding window spanning between the current time and a predetermined period before the current time, regardless of whether it is expressed in minutes, hours, days, or months. This sliding window is updated over time. In a non-limiting example, if the time frame of interest is expressed in days, its range can be updated once a day. At operation 410, the time frame of interest is updated so that it extends from a predetermined duration in the past to the current time, and this operation is periodically repeated at regular intervals. The database 130, the parser 140, the detector 150, or any other component of the service infrastructure 100 can update the time frame of interest. After the time frame of interest is updated, the database 130 or the parser 140 deletes entries in the domain table whose associated time is earlier than the time frame of interest at operation 420. The database 130 or the detector 150 deletes entries in the exception table whose exception time is earlier than the time frame of interest at operation 430.
[0154] After updating the time frame of interest, then at operation 440, the database 130 or the detector 150 repopulates the string table according to the updated time frame of interest. Operation 440 can include sub-operations 442, 444, and 446. At sub-operation 442, for each deleted entry in the domain table, if the entry in the exception table for the domain name that exists in the deleted entry in the domain table indicates that an exception was found for that domain name within the time frame of interest, then the number of domains associated with the string that exists in the deleted entry and for which an exception exists within the time frame of interest is decremented in the string table. At sub-operation 444, for each deleted entry in the domain table, if there is no entry in the exception table for the domain name that exists in the deleted entry in the domain table, then the number of domains associated with the string that exists in the deleted entry and for which no exception exists within the time frame of interest is decremented in the string table. Then at sub-operation 446, the domain table is scanned to find entries that contain the domain names extracted from each deleted entry in the exception table, and for each found string of each found entry in the domain table, the number of domains associated with the found string that exists in the deleted entry in the domain table and for which an exception exists within the time frame of interest is decremented in the string table, and the number of domains associated with the found string that exists in the deleted entry in the domain table and for which no exception exists within the time frame of interest is incremented in the string table.
[0155] Operation 450 can optionally follow operation 440. In operation 450, for each string associated with one or more deleted entries of the domain table, database 130 or detector 150 recalculates the anomaly probability based on the same parameters as they now represent after operations 410, 420, 430, and 440, the parameters including the number of domains associated with the string and without anomalies in the time frame of interest, the number of domains associated with the string and with anomalies in the time frame of interest, the number of domains listed in the anomaly table, and the total number of domains hosted by service infrastructure 100. However, in one embodiment, all anomaly probabilities in the string table can be calculated at regular intervals in an optional batch process of the string table.
[0156] Regardless of whether operation 450 follows operation 440, sequence 400 ends at operation 460.
[0157] Figure 5 is a sequence diagram showing operations of a first method for calculating an aggregated probability of anomalies of service infrastructure 100 according to an embodiment of the present technology. Figure 5 Shows sequence 500 including multiple operations that can be executed by detector 150 or by database 130 in a variable order, some operations can be executed simultaneously, and some operations are optional. In sequence 500 as a first variant for calculating the aggregated probability of anomalies of service infrastructure 100, operation 510 includes calculating the logarithm of each anomaly probability of a string related to one of the domain names. Then, operation 520 includes combining the logarithms of each anomaly probability of a string related to one of the domain names. Sequence 500 ends at operation 530.
[0158] Figure 6 is a sequence diagram showing operations of a second method for calculating an aggregated probability of anomalies of service infrastructure 100 according to an embodiment of the present technology. Figure 6 Shows sequence 600 including multiple operations that can be executed by detector 150 or by database 130 in a variable order, some operations can be executed simultaneously, and some operations are optional. In sequence 600 as a second variant for calculating the aggregated probability of anomalies of service infrastructure 100, operation 610 includes calculating a first aggregated probability of anomalies of service infrastructure 100 on one of the domain names for a first subset of strings related to one of the domain names having the highest anomaly probability. Operation 620 includes calculating a second aggregated probability of anomalies of service infrastructure 100 on one of the domain names for a second subset of strings related to one of the domain names having the lowest anomaly probability. The first aggregated probability and the second aggregated probability are combined at operation 630. Sequence 600 ends at operation 640.
[0159] In an embodiment including sequence 500 or 600, if the abnormal aggregation probability of the service infrastructure 100 relative to the domain name extracted from the data packet or command is lower than a predetermined threshold, then the domain name identified in the log entry generated at Figure 2 operation 210 can be unconditionally marked as non-suspicious by the detector 150. For example, the predetermined threshold for this operation can be set in the same or equivalent manner as expressed in operation 230 discussed above.
[0160] Each of the operations in sequences 200, 300, 400, 500, and 600 can be configured to be processed by one or more processors, and the one or more processors are coupled to a memory. The memory can include a non-transitory computer-readable medium having executable code stored thereon, and the executable code includes instructions for executing any one of sequences 200, 300, 400, 500, and 600 when the executable code runs on the one or more processors.
[0161] Figure 7 is a simplified illustration of an operator display summarizing the abnormal probabilities of multiple hosted domains according to an embodiment of the present technology. In Figure 1 the user interface 160 introduced in the description can provide a spreadsheet 700 on a display (not shown), and the spreadsheet 700 summarizes the contents of the domain table and the string table. As Figure 1 shown, the user interface 160 can access the information in the database 130 via the detector 150. In one embodiment, the user interface 160 can be directly connected to the database 130 and obtain information useful for preparing the spreadsheet 700 from the database 130.
[0162] The spreadsheet 700 includes a table 710 for each of the hosted domains of the service infrastructure 100, and the display can be arranged to provide a partial view of the spreadsheet 700. Although only four (4) domain names are shown in Figure 7 , more domain names can be displayed in an actual implementation. Each table 710 includes a domain name 712 and a ranking 714. In one embodiment, the ranking 714 allows the operator to view which domain name is most affected by the anomalies that occurred during the time frame of interest.
[0163] Each row 716 of each table 710 shows a string 718 that matches the domain name 712 in the domain table and a corresponding score 720 representing the probability p′(A|Str i ) of that string 718. In a given table 710, the rows 716 are arranged to list the strings 718 in descending order of score 720. The actual number of rows 716 displayed for each domain name may be greater than the number shown in Figure 7 .
[0164] In an example consistent with the content of Tables I, II, and III above Figure 7 , String_a has the highest score of 720 on the string table and exists in Domain 1. Thus, in a non-limiting example, Domain 1 is given the first rank (1). String_a is not associated with any other domain name in the domain table, so the second rank (2) is given to Domain 4, where String_b exists and has the second-highest score in the string table. Each row 716 also has the number 722 of domains with anomalies and the number 724 of domains without anomalies, and these values and the score 720 are extracted from the string table.
[0165] Returning to the example above, where the domain hosts the folder “ / www / img / my_nicest_pics / ”, different substrings may have been stored in different log entries. In Figure 7 the example shown, “String_b” actually equals “ / www / img / my_nicest_pics / ” (in this specific example, the string field 718 on the second row 716 of Table 710 for Domain 1 would actually show “ / www / img / my_nicest_pics / ”). The operator can select (e.g., click on) the string field 718 on the second row 716 of Table 710 for Domain 1, such that a window 730 listing multiple substrings related to String_b is displayed.
[0166] It is expected that other fields and information elements can be provided on the spreadsheet 700. For example, each Table 710 or a portion thereof can be displayed with a color code indicating the type of anomaly detected on a particular hosted domain (e.g., phishing, spam, etc.). Operator-selectable fields can be provided for simultaneously displaying multiple Tables 710, for displaying the Table 710 for a particular domain name, for selecting multiple displayed rows 716 in each Table 710, etc. The ranking 714 given to each domain name in each Table 710 can be calculated in other ways—e.g., based on the multiple scores 720 on the respective rows 716 in a given Table 710. Many modifications can be made to the presentation of the spreadsheet 700, and Figure 7 the illustration is not intended to be restrictive.
[0167] Although the above implementations are described and illustrated with reference to specific steps performed in a particular order, it will be understood that these steps can be combined, subdivided, or reordered without departing from the teachings of the present technology. At least some of the steps can be performed in parallel or serially. Thus, the order and grouping of the steps are not a limitation of the present technology.
[0168] It should be clearly understood that not all of the technical effects mentioned herein need to be present in every embodiment of the present technology.
[0169] Modifications and improvements to the above-described implementations of the present technology may become apparent to those skilled in the art. The foregoing description is intended to be exemplary and not restrictive. Thus, the scope of the present technology is intended to be limited only by the scope of the appended claims.
Claims
1. A method for detecting potential anomalies at a service infrastructure, comprising: accessing a string table, each corresponding entry of the string table defining a corresponding string and a corresponding anomaly probability for that string; generating, in a database of the service infrastructure, a log entry related to an event occurring in the service infrastructure, the log entry including a string specifying one of a file name and an IP address, the log entry including a domain name hosted by the service infrastructure; searching for the string in the string table; and if the string is found in the string table and if the anomaly probability corresponding to the string exceeds a predetermined threshold, then marking the domain name as suspicious, wherein the predetermined threshold is calculated based on: where: I is an indicator function; and m is the number of entries in the string table.
2. The method according to claim 1, further comprising: populating a domain table, each given entry of the domain table containing: a given domain name, a given string associated with the given domain name in a time frame of interest, and a given association time corresponding to the most recent association time in all log entries associating the given domain with the given string.
3. The method according to claim 2, wherein populating the domain table includes: parsing a plurality of log entries in the database to extract a corresponding domain name, a corresponding string, and a corresponding association time from each log entry.
4. The method according to claim 2, wherein each corresponding entry of the string table further defines: (i) a corresponding number of domains associated with the corresponding string and having an anomaly in the time frame of interest, and (ii) a corresponding number of domains associated with the corresponding string and not having an anomaly in the time frame of interest.
5. The method according to claim 4, further comprising: at the service infrastructure, determining that an anomaly has occurred at a detection time related to an affected domain; accessing an anomaly table, each entry of the anomaly table including the name of a domain for which an anomaly has been detected in the time frame of interest and a corresponding anomaly time; if there is an anomaly table entry for the affected domain, then updating the corresponding anomaly time in the anomaly table entry with the detection time; and if there is no anomaly table entry for the affected domain, then: creating a new anomaly table entry for the affected domain, the new anomaly table entry including the name of the affected domain and the detection time, extracting a list of strings associated with the affected domain from the domain table, for each string in the list of strings associated with the affected domain, incrementing, in the string table, the number of domains associated with that string and having an anomaly in the time frame of interest, and for each string in the list of strings associated with the affected domain, decrementing, in the string table, the number of domains associated with that string and not having an anomaly in the time frame of interest.
6. The method according to claim 5, wherein determining that an anomaly has occurred includes: receiving, at the service infrastructure, an anomaly indication for the affected domain.
7. The method according to claim 5, further comprising: periodically updating the time frame of interest such that the time frame of interest extends from a past predetermined duration until the current time; and after updating the time frame of interest: deleting entries in the domain table whose associated time is earlier than the time frame of interest, deleting entries in the exception table whose exception time is earlier than the time frame of interest, and refilling the string table according to the updated time frame of interest.
8. The method according to claim 7, wherein, refilling the string table according to the updated time frame of interest comprises: for each deleted entry in the domain table: if the entry in the exception table for the domain name present in the deleted entry in the domain table indicates that an exception was found for that domain name during the time frame of interest, decrementing the number of domains associated with the string present in the deleted entry and having an exception during the time frame of interest in the string table, if there is no entry in the exception table for the domain name present in the deleted entry in the domain table, decrementing the number of domains associated with the string present in the deleted entry and having no exception during the time frame of interest in the string table; scanning the domain table to find entries containing the domain names extracted from each deleted entry in the exception table; and for each found string of each found entry in the domain table: decrementing the number of domains associated with the found string present in the deleted entry in the domain table and having an exception during the time frame of interest in the string table, and incrementing the number of domains associated with the found string present in the deleted entry in the domain table and having no exception during the time frame of interest in the string table.
9. The method according to claim 5, further comprising: for each corresponding entry in the string table, calculating a corresponding anomaly probability based on the following four: (i) the corresponding number of domains associated with the corresponding string and having no anomaly during the time frame of interest, (ii) the corresponding number of domains associated with the corresponding string and having an anomaly during the time frame of interest, (iii) the number of domains listed in the exception table, and (iv) the total number of domains hosted by the service infrastructure.
10. The method according to claim 9, wherein, each corresponding anomaly probability is calculated by applying a Bayesian filter to the following four: (i) the corresponding number of domains associated with the corresponding string and having no anomaly during the time frame of interest, (ii) the corresponding number of domains associated with the corresponding string and having an anomaly during the time frame of interest, (iii) the number of domains listed in the exception table, and (iv) the total number of domains hosted by the service infrastructure.
11. The method according to claim 10, further comprising: calculating the aggregated probability of an anomaly of the service infrastructure related to one of the domain names in the domain table by the following means: Calculate the logarithm of each of the anomaly probabilities of the strings associated with the one domain name in the domain name; and Combine the logarithms of each of the anomaly probabilities of the strings associated with the one domain name in the domain name.
12. The method according to claim 10, further comprising: Calculate the aggregated probability of anomalies of the service infrastructure associated with one domain name in the domain table by: For a first subset of strings associated with the one domain name in the domain name having the maximum anomaly probability, calculate a first aggregated probability of anomalies of the service infrastructure on the one domain name in the domain name; For a second subset of strings associated with the one domain name in the domain name having the minimum anomaly probability, calculate a second aggregated probability of anomalies of the service infrastructure on the one domain name in the domain name; and Combine the first aggregated probability and the second aggregated probability.
13. The method according to claim 11, wherein If the aggregated probability of anomalies of the service infrastructure associated with the domain name present in the generated log entry is lower than a predetermined threshold, then unconditionally mark the domain name as non-suspicious.
14. The method according to claim 2, further comprising: After marking the domain name as suspicious, perform an action selected from the group consisting of: discarding the received data packet related to the generation of the log entry, sending a message to the client related to the domain name, providing an alert on the service infrastructure, and blocking all data traffic related to the domain name.
15. A service infrastructure, comprising: A server configured to receive data packets and / or commands from a client; A database configured to store a plurality of log entries, each corresponding log entry of the log including a corresponding string associated with a corresponding domain name; A processor; and A memory device including a non-transitory computer-readable medium, on which executable code is stored, the executable code including instructions for performing the method according to any one of claims 1 to 14 when the executable code runs on the processor.
16. The service infrastructure according to claim 15, wherein The processor implements a detector configured to detect an anomaly related to an affected domain occurring at a corresponding detection time.
17. The service infrastructure according to claim 16, further comprising a user interface operably connected to the detector, the detector configured to cause the user interface to issue a warning to an operator of the service infrastructure when the domain name is marked as suspicious.
18. The service infrastructure according to claim 16, further comprising a signaling interface operably connected to the detector, the detector configured to cause the signaling interface to send a message to a client related to the domain name when the domain name is marked as suspicious.
19. The service infrastructure according to claim 16, further comprising a blocker operatively connected to the detector, the detector being configured to cause the blocker to issue a command to the server to discard received data packets related to the generation of log entries when a domain name is marked as suspicious.
20. A method of recording information about detected anomalies at a service infrastructure, comprising: defining a string table, each respective entry of the string table defining: (i) a respective string, (ii) a respective number of domains hosted by the service infrastructure, associated with the respective string, and having an anomaly during a time frame of interest, (iii) a respective number of domains hosted by the service infrastructure, associated with the respective string, and not having an anomaly during the time frame of interest, and (iv) a respective anomaly probability of the string; generating a plurality of log entries in a database of the service infrastructure, each log entry being related to an event occurring in the service infrastructure, each log entry associating a string specifying one of a file name and an IP address with a domain name hosted by the service infrastructure, each log entry further recording a specific association time; parsing the plurality of log entries to populate a domain table, each given entry of the domain table containing a given domain name, a given string associated with the given domain name during the time frame of interest, and a given association time corresponding to the latest association time in all log entries associating the given domain with the given string; detecting an anomaly at the service infrastructure, the anomaly occurring at a detection time related to an affected domain; accessing an anomaly table, each entry of the anomaly table including the name of a domain for which an anomaly has been detected during the time frame of interest and a corresponding anomaly time; if there is an anomaly table entry for the affected domain, updating the corresponding anomaly time in the anomaly table entry with the detection time; and if there is no anomaly table entry for the affected domain, then: creating a new anomaly table entry for the affected domain, the new anomaly table entry including the name of the affected domain and the detection time, extracting a list of strings associated with the affected domain from the domain table, for each string in the list of strings associated with the affected domain, incrementing the number of domains associated with the string and having an anomaly during the time frame of interest in the string table, and for each string in the list of strings associated with the affected domain, decrementing the number of domains associated with the string and not having an anomaly during the time frame of interest in the string table.
21. A service infrastructure, comprising: a server configured to receive data packets and / or commands from a client; a database configured to store a plurality of log entries, each respective log entry of the logs including a respective string associated with a respective domain name; a processor; and A memory device comprising a non-transitory computer-readable medium having stored thereon executable code including instructions for performing the method according to claim 20 when the executable code is run on a processor.
Citation Information
Patent Citations
Methods and apparatus for detecting unwanted traffic in one or more packet networks utilizing string analysis
CN101529862A
Input error processing method and apparatus
CN105589570A