Digital availability surveillance method

A machine-learning model analyzes internet signals to detect ransomware attacks on healthcare networks, providing early warning and rapid response to minimize disruptions.

WO2026024760A1PCT designated stage Publication Date: 2026-01-29RGT UNIV OF CALIFORNIA +2
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/038721
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-22
Filing Date
2025-07-22
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Current systems lack the ability to rapidly identify healthcare facilities affected by ransomware attacks, which can cause significant disruptions with life-threatening consequences, and there is no standardized reporting of such incidents.

Method used

A machine-learning model that analyzes publicly available internet signals, such as email system responses and social media posts, to detect abnormal network behavior and identify potential ransomware attacks before they are publicly reported, using a handshake signal to assess network responsiveness and characterizing activities as normal or abnormal.

Benefits of technology

Enables early detection of ransomware attacks on healthcare networks, potentially days before public announcement, allowing for rapid response and minimizing disruption impacts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025038721_29012026_PF_FP_ABST
    Figure US2025038721_29012026_PF_FP_ABST
Patent Text Reader

Abstract

Digital surveillance methods and systems are described. One example method relates to detecting disruptions to a computer network associated with an organization. Another example method relates to detecting abnormal activities on a computer network. The disclosed embodiments include techniques to analyze a unique and varied collection of publicly available signals associated with specific organizations to allow abnormal behaviors of a digital network to be detected.
Need to check novelty before this filing date? Find Prior Art

Description

DIGITAL AVAILABILITY SURVEILLANCE METHODCROSS-REFERENCE TO RELATED APPLICATION

[0001] This patent document claims priority to and benefits of U.S. Provisional Application 63 / 674,158, entitled “DIGITAL AVAILABILITY SURVEILLANCE METHOD,” and filed on July 22, 2024. The entire content of the above noted patent application is incorporated by reference as part of the disclosure of this patent document.STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH

[0001] This invention was made with Government support under grant number SP4701- 23-C-0075 awarded by the Advanced Research Projects Agency for Health. The government has certain rights in the invention.TECHNICAL FIELD

[0002] The present patent document relates to systems, methods, and devices for detecting disruptions to a digital network before they are publicly reported.BRIEF DESCRIPTION OF THE DRAWINGS

[0003] FIG. 1 shows example data used to evaluate machine learning models in accordance with an embodiment of the disclosed technology.

[0004] FIG. 2 shows example results obtained in a study performed in accordance with an embodiment of the disclosed technology.

[0005] FIG. 3 shows a geospatial map obtained in accordance with an embodiment of the disclosed technology.

[0006] FIG. 4 shows example data obtained in a study performed in accordance with an embodiment of the disclosed technology.

[0007] FIG. 5 shows example data obtained in a study performed in accordance with an embodiment of the disclosed technology.

[0008] FIG. 6 shows rules used to categorize subnetworks according to an embodiment of the disclosed technology.

[0009] FIG. 7 shows example results obtained in a study performed in accordance with disclosed techniques.

[0010] FIG. 8 example results obtained in a study performed in accordance with disclosed techniques.

[0011] FIG. 9 shows example results obtained in a study performed in accordance with disclosed techniques.

[0012] FIG. 10 shows example results obtained in a study performed in accordance with disclosed techniques.

[0013] FIG. 11 shows example results obtained in a study performed in accordance with disclosed techniques.

[0014] FIG. 12 shows example results obtained in a study performed in accordance with disclosed techniques.

[0015] FIG. 13 shows a flowchart of a methodology based on the disclosed technology.

[0016] FIG. 14 show example results obtained in a study performed in accordance with disclosed techniques.

[0017] FIG. 15 shows example results obtained in a study performed in accordance with disclosed techniques.

[0018] FIG. 16 shows example results obtained in a study performed in accordance with disclosed techniques.

[0019] FIG. 17 shows example results obtained in a study performed in accordance with disclosed techniques.

[0020] FIG. 18 shows example results obtained in a study performed in accordance with disclosed techniques.

[0021] FIG. 19 shows example sources of ground truth data according to an embodiment of the disclosed technology.

[0022] FIG. 20 shows example data obtained in a study performed in accordance with disclosed techniques.

[0023] FIG. 21 shows example data obtained in a study performed in accordance with disclosed techniques.

[0024] FIG. 22 shows example data obtained in a study performed in accordance with disclosed techniques.

[0025] FIG. 23 shows example data obtained in a study performed in accordance with disclosed techniques.

[0026] FIG. 24 shows example data obtained in a study performed in accordance with disclosed techniques.

[0027] FIG. 25 shows example data obtained in a study performed in accordance with disclosed techniques.

[0028] FIG. 26 shows example data obtained in a study performed in accordance with disclosed techniques.

[0029] FIG. 27 shows example data obtained in a study performed in accordance with disclosed techniques.

[0030] FIG. 28 shows a screenshot of an example handshake signal time out according to an embodiment of the disclosed technology.

[0031] FIG. 29 shows an example LSTM architecture according to an embodiment of the disclosed technology.

[0032] FIG. 30 shows a flowchart of an example method according to an embodiment of the disclosed technology.

[0033] FIG. 31 shows a flowchart of an example method according to an embodiment of the disclosed technology.DETAILED DESCRIPTION

[0034] The disclosed technology relates to the development and use of techniques to regularly and proactively analyze a unique and varied collection of publicly available internet signals associated with specific institutions (e.g., healthcare institutions) to provide indicators of abnormal network behavior across an entire sector of a critical infrastructure.

[0035] Various embodiments disclosed herein are described in the context of a hospital or healthcare facility as non-limiting examples. The example embodiments are not exclusive to the healthcare domain and may be used to give insight over other infrastructures or organizations(e.g., schools, banks, etc.). Additionally, the disclosed embodiments may be implemented to detect various types of disruptions, attacks, or outages to a network.

[0036] In the current state of the art, no functionality exists to rapidly identify healthcare facilities affected by ransomware attacks. Deep technical expertise in the field of network scanning and Internet measurement has resulted in the development of methodologies and systems, such as those described herein, that can elucidate a host of IP addresses belonging to specific healthcare delivery organizations. Coupled with a frequently updated national database of hospitals, techniques disclosed herein permit monitoring of a substantial majority of a national healthcare infrastructure.

[0037] Ransomware attacks have already impacted numerous healthcare systems worldwide. When health IT systems are taken offline, it can cause a huge disruption to patients and providers, potentially with life-threatening consequences if emergency care and surgery cannot be provided. In the US, healthcare providers are not held to a strict reporting standard when these attacks occur, and they may delay announcing an attack for PR reasons. However, the impact to public health from a hospital being attacked is immediate, and many stakeholders (e.g., other hospitals, state / local / federal government) would be able to react more quickly to address these attacks if an automated detection system was in place. Besides hospitals, the disclosed embodiments may be used to detect cyber attacks on other critical systems (e.g., schools, banks, infrastructure, etc.).

[0038] Some example embodiments to be described herein involve a machine-learning model that takes multiple outputs (e.g., email system response, emergency dispatch, standard health record reporting, social media posts) to estimate whether a provider has been taken offline by a ransomware attack. These signals are publicly available and do not require permission from the provider in question. Among other features and benefits, the disclosed embodiments may be used to detect real-world cyberattacks using these signals. In some cases, such detection may occur hours or days before the attacks are announced publicly by the affected healthcare providers.

[0039] In one aspect, the disclosed embodiments include a method for detecting disruptions to a computer network associated with an organization. The method comprises: transmitting a handshake signal to a set of network addresses, wherein at least some network addresses in the set of network addresses are associated with a service or system of the organization; obtaining monitoring data associated with a respective network address in the setof network addresses; determining a responsiveness of the respective network address to the handshake signal based on the monitoring data; obtaining, based on the responsiveness, information related to an activity associated with the respective network address; characterizing the activity as normal or abnormal based on the information; and detecting a disruption to the network based on results of the characterizing, wherein the characterizing comprises obtaining a determination as to whether the activity is normal or abnormal using a machine learning (ML) algorithm trained by retrospective data from historical disruptions to the network.

[0040] In another aspect, the disclosed embodiments include a method for detecting abnormal activities on a computer network. The method comprises: obtaining a first set of network addresses, wherein each network address in the first set is connected to the computer network and associated with an organization; categorizing, based on predetermined rules, each network address in the first set according to a set of categories to obtain a second set of network addresses; determining subnetworks associated with at least some network addresses in the second set; assigning, based on the predetermined rules, each of the subnetworks to a category in the set of categories; determining a number of the subnetworks belonging to each category in the set of categories; obtaining a characterization of the computer network based on the number of the subnetworks belonging to each category; determining one or more correlated disruptions of the subnetworks; and determining an abnormal activity on the computer network based on the one or more correlated disruptions, wherein: at least some network addresses in the first set are associated with different organizations, at least some of the subnetworks are shared by the different organizations, and the characterization is based on the subnetworks that are shared by the different organizations.

[0041] In the description that follows, an example method for detecting ransomware attacks is disclosed in the context of a hospital as an example. The method can be implemented by software.

[0042] Hospitals have an increasing number of technologies and services that are network connected and thus interact with the Internet. From e-mail webservers to time-keeping systems for employees to APIs for electronic medical records, the "up-time" of a service can be assessed in accordance with the disclosed method by asking a particular IP address associated with these elements for a simple "handshake" - a brief connection between computers for the transmission of basic data. When this process proceeds uneventfully, empiric confirmation that the target service remains connected to a functional network is gained. As ransomware attacks are nearlyalways associated with subsequent loss of network functionality, the converse - a lack of response to what should be a routine request - is suggestive of abnormal network behavior. By combining a significant number of "signals" - i.e., the individual services or systems targeted in the presently disclosed technique - with a large number of addresses associated with hospitals (addresses that are not publicly available but can be enumerated from one initial starting address according to an embodiment of the disclosed technology), the digital infrastructure of a nation's healthcare system is able to be “scanned.” When abnormal responses are returned, a machine learning algorithm trained by retrospective data from historical ransomware attacks is utilized to determine a probability that such responses are associated with a ransomware attack induced network disruption as opposed to another cause (e.g., power outage, scheduled downtime). In some implementations, the machine learning model is trained to assign probabilities to aberrant data correlating to likelihood of ransomware attack.

[0043] Information gained from the example embodiments disclosed herein can be organized and visualized on a dashboard to provide near real-time insight into the network status of hospitals across the country, allowing for rapid detection and response to healthcare ransomware attacks.

[0044] Some disclosed embodiments rely on machine learning. The present patent document discloses various machine learning training approaches.

[0045] In one example machine learning training approach, a dataset is grouped by hospital name and date. Each group contains information about the IPs associated with that hospital, along with their last update timestamp. Additionally, a ransomware column is used to indicate whether a ransomware attack occurred on the hospital for that particular date. An IP is considered active on a given date if its last update was recorded on the previous day. For each (hospital, date) group, the percentage of active IPs is computed. Rows with missing data are removed, and duplicate entries are dropped to ensure data consistency. Before training the models, the percentage of active IPs feature was standardized using Standard Seal er from the s klearn library to normalize the values. The final dataset includes the scaled percentage of active IPs and the corresponding ransomware attack indicator as the target variable. The data can then be split into 75% training and 25% testing subsets.

[0046] In an example demonstration of disclosed techniques, the following list of machine learning models were evaluated and trained to predict ransomware attacks based on IP activity patterns:• Logistic Regression: A simple yet effective linear model that estimates the probability of a ransomware attack based on the weighted sum of input features. It serves as a strong baseline model.• Decision Tree: A non-linear model that splits data based on feature thresholds, creating a tree-like structure for classification. It is interpretable but can suffer from overfitting.• Random Forest: An ensemble learning method that builds multiple decision trees and combines their predictions to improve accuracy and reduce overfitting. It provides better generalization than a single decision tree.• XGBoost (Extreme Gradient Boosting): A powerful gradient boosting algorithm that optimizes tree-based learning through parallel computation and regularization techniques. It is highly effective in capturing complex patterns and is often used in competitive machine learning tasks.

[0047] These models were evaluated based on key classification metrics such as accuracy, precision, recall, and Fl -score to assess their effectiveness in predicting ransomware attacks. Example results of the model evaluations are shown in Table 1 below.

[0048] Table 1 : Example evaluation of machine learning models

[0049] FIG. 1 shows a graph of the positive rate vs. false positive rate of the four models (Logistic Regression, Decision Tree, Random Forest, XGBoost) as obtained in an example demonstration.

[0050] In the description that follows, example techniques to measure disruptions of networks during a technology outage are disclosed and an example study was performed in accordance with the disclosed techniques. A system capable of monitoring and detecting disruptions in a critical digital infrastructure is disclosed.

[0051] The example study was motivated by the question: what patient safety impacts were associated with a major national technology outage? In the study, to be described in further detail below, network disruptions coinciding with a faulty software update of July 19th, 2024 were measured at 759 US hospitals. Of the nearly 1100 specific internet-based services examined, 21.8% were characterized as corresponding with direct patient care functionality. Results of the study demonstrate that widespread technology failures affecting healthcare infrastructure may have commensurate negative impact to patient care systems.

[0052] A company which produces cybersecurity software (CrowdStrike, Austin, TX) makes "Falcon," a program intended to monitor and protect computers of large commercial enterprises from cybersecurity threats. On July 19th, 2024, a faulty update for the Falcon program was simultaneously distributed via the Internet to millions of personal computers (PCs) and servers running certain versions of the Windows (Microsoft, Redmond, Washington) operating system and the Crowdstrike Falcon software.

[0053] The update contained a programming error which caused computers that installed the update to reboot and crash. Installing the repair or "patch" for this update required direct manual access to each affected computer and could not be accomplished remotely through the Internet - a time intensive and laborious process resulting in downtimes of hours or even days for many organizations. The impact from the Crowdstrike Falcon update was immediate and global. Disruptions were sustained across dozens of industries in dozens of countries.

[0054] The cross-sectional study was conducted between the dates of July 5th, 2024, and August 3rd, 2024. Two weeks of data before (July 5th- July 18th), during (July 19th), and after (July 20th- August 3rd) the Crowdstrike outage was collected and analyzed. U.S. healthcare delivery organizations (HDOs) running the Epic (Epic Systems, Verona, WI) EHR and which had at least one publicly available Fast Healthcare Interoperability Resources (FHIR) internet endpoint were included.

[0055] Data for the study was collected as part of a pre-existing initiative to prospectively monitor hospital ransomware attacks.

[0056] Collected and catalogued were internet address ranges for HDOs and hospitals that use Epic as their EHR provider and which host external and public-facing internet services. A provider of historical internet address availability data (Censys, Ann Arbor, MI) was used to match hospital domains to IP ranges, and filtering using DomainName System (DNS) and X.509 Certificate information provided by Censys' daily snapshot of the internet address space wereused. Services running on the identified hosts were similarly enumerated. FHIR endpoints were identified by downloading additional data from the Lantern Project, a publicly available resource from MITRE (McLean, VA).

[0057] Hospital IP range scans were performed using an open-source tool designed to probe open network ports of internet-connected hosts (ZMap version 4.2.0, University of Michigan, Ann Arbor, Michigan). Address ranges were probed in rounds of 3 hours, and a positive scan (on any port) indicated an online host endpoint. Scans of FHIR endpoints were performed in rounds of 2.5 hours using a Python hypertext transfer protocol (HTTP) client. Positive scan results, i.e., endpoint was operational and communicating with outside internet traffic, as well as negative scans, i.e., system downtime, were recorded and stored on a secure internal server.

[0058] Downtime for the address range scans was defined as the period between when a deviation from the normative count of positive IP scans (by a factor greater than or equal to twice the standard deviation) occurs to when it recovers. Downtime for a hospital or HDO was defined as the maximum downtime for any network on which it hosted any service.

[0059] Downtime for the FHIR endpoint scans was defined as the time between a negative scan and the return of a positive scan. Endpoints that had no positive scan the entire duration of the study were excluded. Downtime instances were then collated into a dataset for subsequent analysis of potential clinical effects.

[0060] Once unresponsive network services were identified, further steps were taken to provide a clinical interpretation of the services related to each HDO or hospital to infer potential impacts on clinical operations. Four researchers independently analyzed the Crowdstrike outage dataset described above (n=l,098). The clinical relevance of the 1,098 affected services was investigated through (i) directly visiting site Uniform Resource Locators (URLs), (ii) 'Google Dorking', a technique harnessing search engines to identify specific text in website code, and (iii) Domain Name Service (DNS) evaluation through terminal queries. Where further confirmation was needed, Client for URL (curl) requests were used to check site availability, HTTP status codes, and headers to attempt to determine the clinical function of an IT service. Viewing the page source provided additional insights, such as metadata with information on the services behind a Login portal (e.g., remote staff access portals for patient records) and hidden redirects.

[0061] Through the combination of these methods, each affected service was assigned a manual label detailing its function within clinical operations (e.g., "Patient portal for healthcare record"). Based on the manual labels, each service was assigned one of four categories capturing its relevance to patient care: (1) Patient facing (e.g., radiological imaging systems), (2) Operationally Relevant (e.g., staff scheduling systems), (3) Research relevant (e.g., research databases for clinical trial operations) and (4) Not relevant or unknown (e.g., donation pages for academic institutions). For categories 1 to 3, the service had to affect patient experience and thus services such as research lab information webpages were excluded and placed in category 4. Services that could not be identified due to internal network or security restrictions, or were decommissioned and unavailable, were also placed in category 4.

[0062] FIG. 2 shows example results obtained in the study. The study showed that immediately following the Crowdstrike update on July 19, 2024, 759 health systems of X meeting inclusion criteria experienced endpoint disruption using a combination of the address space and FHIR scans. FIG. 2 demonstrates the overlap.

[0063] Further examination of affected domains across the hospitals yielded a total of 1,098 individual disrupted digital services that were investigated for clinical relevance. Figure 3 provides a more granular illustration of this downtime duration across the country, with the geographies of affected hospitals plotted across the United States and color-coded according to the length of recovery time. Specifically, Figure 3 shows a geospatial map of the locations of identified HDOs inferred to have service disruptions using address space scans. Data points on the map are color coded by the time the recovery following the CrowdStrike incident. Most hospital services recovered within 0-6 hours (n=321), with a smaller number of hospitals (n=43) experiencing outages of over 48 hours (Figure 3).

[0064] It was also observed that a set of 52 unique HDOs (comprising 299 individual hospitals) failed to respond to the FHIR scans immediately following the July 19thupdate. The change in FHIR outage detection occurring during the Crowdstrike event is presented in FIG. 4, which displays unresponsive HDO EPIC FHIR endpoints resulting from scans conducted two weeks before and two weeks after the Crowdstrike Outage. Specifically, FIG. 4 shows unresponsive HDO Epic FHIR endpoints prior (July 5th- July 18th, 2024), during (July 19th, 2024), and after the Crowdstrike outage (July 20th- August 3rd). The vertical red line marks the time point at which the Crowdstrike outage occurred.

[0065] The mean (SD) daily FHIR endpoint downtime was significantly increased (P <0.01) between the pre and post-periods with 6 events (3.72) prior, 128 during, and 11 (9.1) after the CrowdStrike incident. There was no significant difference in the means between the pre and post-periods (P = 0.065). The median observed downtime for all HDOs during the event was 5 hours (IQR, 2.5 - 8 hours), with a majority of HDOs recovering within 5 hours (31 / 52).

[0066] Clinical impacts were also assessed in the study. Table 2 demonstrates that the 1,098 affected services were found to be 21.8% Patient Facing, 15.4% Operationally Relevant, 5.3% Research Relevant and 57.5% Not relevant / unknown (Table 2). Patient Facing Services spanned imaging platforms, pre-hospital medicine health record systems, patient transfer portals, access to secure documentation and staff portals for viewing patient details (Table 2). In addition to staff portals, patient access platforms across diverse hospital systems were being impacted, which, when operating as usual, allow patients to schedule appointments, contact healthcare providers, access lab results and refill prescriptions. The "Not relevant" categories encompassed testing / staging environments for software that were pre-deployment, donation pages for institutions, information pages on educational services and educational resources for students (e.g., medical, nursing students) (Table 2).

[0067] Table 2: Evaluation of affected services at Crowdstrike impacted hospitals and their relevant clinical utility

[0068] The foregoing discloses example techniques which were implemented in a study to proactively monitor digital signals corresponding to HDOs using well-established Internet measurement techniques and report associated downtime occurring at IP addresses corresponding to 759 hospitals during the Crowdstrike outage of 2024. The disclosed techniques enabled disruptions in nearly 1100 unique services belonging to HDOs across the country to be measured. While a majority of the services (59%) were either associated with non-clinical elements or unable to be categorized, characterization of the remaining elements allows for the most granular description of this incident's impacts on healthcare infrastructure to date.

[0069] The study found an association between the 2024 Crowdstrike outage and many geographically diverse HDOs experiencing significant technical system downtime. Hospital recovery from the downtime was temporally varied. Prospective internet availability scanning of critical digital healthcare may serve as an early warning signal for adverse events such as ransomware attack, datacenter failure, or faulty software, and could serve an important public health function as healthcare continues expanding its dependence on digital technology.

[0070] Table 3 provides further examples of successfully identified clinical services from the Crowdstrike Outage dataset which were obtained in the study. Extracts of informative sections of source code, headers and webpage details are also shown in Table 3. Column 1 (Table 3) details the service from the Crowdstrike Outage Dataset that was investigated, and thecategory that was assigned following investigation (e.g., patient facing). Column 2 (Table 3) provides the technique for obtaining information on the service and the relevant obtained information and the final column provides the clinical translation of this service in the context of patient care. Identifiable information has been removed and replaced with "XX" to protect the identity of individual organizations.

[0071] Table 3: Example table of services affected by Crowdstrike Outage and the techniques used to translate the services to their clinical context.

[0072] Some disclosed techniques relate to monitoring service availability of networks. Example techniques to measure network disruptions through targeted probing are disclosed. An example study performed in accordance with the disclosed techniques is also described.

[0073] In the example study, active measurement techniques were used to monitor service availability across major U.S. hospitals over two years. The study was motivated by the increasingly critical role played by technology in health delivery (i.e., that network downtime can have direct, and adverse, health implications in practice). Exploring this question has required addressing three key measurement challenges that were explored in the study: establishing a mapping between physical hospitals and the particular IP address ranges that they depend upon, developing heuristics to distinguish outage anomalies from mere shifts in infrastructure deployment and, finally, validating this work using ground truth proxies to tie a subset of such measurements back to real -world events. In particular, publicly-reported hospital ransomware attacks were used as such a proxy (hospitals under attack will commonly disconnect affected services to prevent ongoing control and / or future contagion) and it was shown that 71% of applicable attacks manifest as anomalies in the collected data.

[0074] Active measurements can be used to infer service and network availability - if services on a network once responded, but now do not, this is a powerful indicator that some failure has occurred. Investigating this issue of characterizing hospital network service downtime presents several challenges: mapping hospitals to IP address lists, inferring hospital service outages, and validation against ransomware proxies, among others. Such challenges were a focus of the study. Another goal was to study targeted cyberattacks at a smaller granularity, affecting a specific sector of organizations (hospitals) scattered around the Internet. This goal involves two main challenges: determining the networks associated with a specific category of organizations, and detecting disruptions of those networks.

[0075] Unlike Internet-wide measurement, which seeks to cover as many of the networks connected to the Internet as possible, a focused measurement study requires identifying the networks associated with specific organizations.

[0076] To measure a hospital's network disruptions, the hospital's network infrastructure needed to be identified. While it is tempting to think of a hospital's network as a single entity, the reality is much more complicated. In fact, a hospital may operate its own network (e.g., announce its own Autonomous System), host some of its services in the cloud, have network connectivity entirely through a regional Internet Service Provider (ISP), or a mix of all three.

[0077] An example technique to identify a hospital's network includes identifying individual IP addresses operated by the hospital. Recognizing that the hospital may have varied level of control over these IP addresses, the IP addresses were categorized into three categories: Hospital, ISP, or Cloud. Lastly, these IP addresses were expanded into subnetworks, which further increases coverage.

[0078] The present patent document discloses various techniques to allow discovery of a hospital’s network.

[0079] In an example methodology to discover a hospital’ s network, the hospital's network is represented by a set of IP addresses. The method includes: (1) identifying a seed set of IP addresses, and (2) discovering additional IP addresses based on the seed set.

[0080] To identify a seed set of IP addresses associated with hospitals, the following steps were applied: a) identifying a seed set of domains; b) identifying additional domain names; and c) identifying IP addresses associated with each domain.

[0081] To identify a seed set of domains, a seed set of domains that are associated with hospitals was collected. A dataset compiled by the American Hospital Association (AHA) was used, the dataset including a list of hospital names and their associated domain names. For example, for the hospital Scripps Health, the AHA dataset contains the domain name scripps . org (FIG. 5). Domains that do not resolve were removed.

[0082] Specifically, FIG. 5 shows a flowchart of an example methodology which was applied to discovering network boundary of Scripps Health, a system operating four individual hospitals. The seed domain scripps . org, which was used to discover additional FQDNs (even those under different second-level domains) and IP addresses. The IP addresses eventually expand to two subnetworks, one hospital and one cloud.

[0083] Since the AHA dataset only contains the primary domain names associated with each hospital, identification of additional domain names that are associated with the same hospital is needed. For example, the AHA dataset only contains the domain name scripps . org, but Scripps Health operates a variety services under this domain, such as servicenow . scripps . org (FIG. 5).

[0084] To uncover such additional domain names, Censys's collection of X.509 certificates was utilized. Concretely, for each domain in the seed set, Censys's collection of X.509 certificates were used to to query Censys's certificate database. Censys returns all certificates that either have the seed domain as the Common Name (CN) or have a common name that is a subdomain of the seed domain. As an example, if the queried domain is scripps . org, Censys returns all certificates that have scripps . org as the CN, as well as certificates with CNs that are subdomains of the seed domain (e.g., servicenow . scripps . org). For all such certificates returned, both the CN and all domains listed in the Subject Alternative Names (SANs) were added to a domain list.

[0085] To discover IP addresses associated with the domains, Censys's host repository, which contains a list of IP addresses associated with each domain, was utilized. Censys aggregates data from multiple sources, including reverse DNS lookups, forward DNS lookups, protocol responses (e.g., HTTP responses), and the banner when available (via Zgrab).

[0086] While the seed set of IP addresses already provides good coverage of a hospital's network, the set only includes IP addresses that are searchable and discoverable using a domain name. However, other hosts can exist that are operated by the hospital, but do not have adiscoverable domain name associated with them (e.g., a host may still have port 443 and 80 open but does not return a domain name when Censys tries to scan it). To capture these hosts, the seed set of IP addresses was expanded into the subnetworks they belong to.

[0087] To expand IP addresses into subnetworks, a convenient approach is to assume a fixed subnetwork size (e.g., / 24) for every IP address associated with a hospital. This approach has two problems: (1) Real -world networks are not always perfectly aligned to a fixed size, and it is hard to determine the size of the subnetwork; and (2) not every IP address is worth expanding to a subnetwork (e.g., IP addresses operated by a cloud provider). To address the first issue, IP ownership and routability information was utilized to determine the prefix size of the subnetworks. To address the second issue, the IP addresses were categorized into four categories (Hospital, I SP, Cloud, and Others) and only the Hospital IP addresses were expanded into subnetworks.

[0088] To determine the prefix size, two sources were used to provide an informative threshold on the subnetwork prefix sizes. The IP WHOIS ownership entries reflect how subnetworks are allocated in practice. The BGP routability announcements reflect the routability of the subnetwork. When the two sources offer different prefix sizes, the most specific prefix size was selected. If the process yields a prefix size less specific than / 24, it was truncated to / 24 to make the active scanning more efficient. Furthermore, the same category of the IP address was generalized to the subnet, which should be the same for all IP addresses in the subnet per the most specific prefix rule.

[0089] Every month, this process was repeated to grow the seed domains and discover IP addresses to capture new domains and IP addresses, and in turn new subnets, that are associated with hospital organizations. Performing a sensitivity analysis on the frequency of refreshing these IP addresses indicated that a monthly frequency is sufficient to capture the evolution of hospital networks.

[0090] The IP addresses may be categorized to help facilitate analysis. For example, each IP address was assigned one of the following labels: Hospital, indicating that the IP is attributed to a hospital; Cloud, indicating the IP is shared by multiple tenants, and is hosted by a cloud service or data center (e.g., AWS or Epic Cloud); ISP, indicating it is hosted by an Internet Service Provider and may be shared with other customers of the ISP; or Others, which captures the remaining IPs. As will be explained in the description that follows, the first three categories capture the vast majority of IP addresses.

[0091] The category of an IP address was inferred based on its owner and connectivity providers. The owner information was retrieved from the IP WHOIS database, which reflects who the IP address is allocated to. The connectivity provider was determined using BGP origin announcements, which reflects the Autonomous System (AS) and hence organization the IP address is routed through. Then, the following four rules were applied to categorize the IP addresses:

[0092] Rule 1 : Categorize based on IP WHOIS and AS Organization Names. Regular expressions were used to match the name patterns for three types of organizations: Hospital, Cloud, and ISP. When IP WHOIS and AS organization names both map to the same category, this category was assigned to the network.

[0093] Rule 2: Categorize based on Most-Specific Announcement. In cases where the IP WHOIS and AS categories do not agree, the category from the most specific announcement was adopted as the representative category.

[0094] Rule 3 : Break Ties using BGP Announcement. When the prefix sizes w the same, use the BGP announcement to break ties, assuming that routing information is updated more frequently and is more central to the IP address's operation, thus more accurate.

[0095] Rule 4: Special Routing Relationship. IP addresses that belong to a hospital organization but are routed through an ISP were identified as Hospital IPs. This is because although the hospital organization may be assigned a less specific (e.g., / 16) block by the Regional Internet registry (RIR), they often announce the block as individual / 24 subnets through the ISP, a common practice for BGP configuration.

[0096] Anycast and CDN IP addresses were removed based on their AS numbers, because they are reverse proxies pointing to the hospital's online presence and not part of the hospital's online presence. Unroutable IP addresses were also removed.

[0097] Example results obtained in a study performed in accordance with the example methodology are disclosed herein.

[0098] To identify seed IP addresses, domains were collected from the AHA dataset, which contains domain information for 6,150 individual hospitals. After filtering out hospitals with no domains and domains that fail to resolve, 5,980 hospitals and 4,336 unique domains were identified. The list of domains was further expanded using Censys's collection of X.509certificates, resulting in a total of 34,774 FQDNs, among which 9,314 have different registered domains and the rest are subdomains of the seed domains.

[0099] In March 2025's snapshot, 34,774 FQDNs were identified using Censys' certificate transparency (CT) logs, 9,314 of which are under a different second-level domain than the seed domain. These FQDNs are associated with 61,570 unique IP addresses using Censys. These IP addresses are originated by 1,586 unique ASes and belong to 4,138 unique networks based on WHOIS registrations. Using the categorization and expansion methodologies disclosed above, these IP addresses were expanded into 19,989 subnetworks. The category breakdown of these subnetworks and the hospitals they provide coverage for is shown in Table 4.

[0100] Table 4: All networks that were identified and categorized; and the number of hospitals covered by each type of network and their average size (characterized by bed count)Hospital CoverageType # of # of # ofSubnets ASes WHOIS Hospital Description Count Beds ± stdevHospital 2,327 719 1,355 Has Hospital subnets 1,728 209 ± 258ISP 1,850 430 1,033 Has ISP but no Hospital subnets 984 137 ± 177Cloud 14,512 83 763 Only Cloud subnets 2,542 166 ± 200Others 1,300 550 1,022 Only uncategorized subnets 726 118 ± 154Total 19,989 1,586 4,138 All externally visible hospitals 5,980 170 ± 215

[0101] The self-reported bed counts from the AHA dataset were used to try to characterize the size of hospitals associated with these subnets. Unsurprisingly, hospitals with Hospital subnets (which are likely on-premise) are larger, with an average of 209 beds. While this allows for additional coverage on 984 more hospitals by including subnets categorized as ISP, these hospitals are much smaller at half of the number of beds on average. Notably, 2,542 hospitals are hosted exclusively on cloud services, which varies in size. In addition, of all 5,980 externally visible hospitals, 3,113 have only one associated subnet. These hospitals are on the slightly smaller end, with an average of 159 beds. Although 75% of hospitals have four or fewer subnets, the average of 50 subnets is biased by large extremes.

[0102] Hospital subnets were categorized using a categorization algorithm. A breakdown of the rules used to categorize Hospital subnets is shown in FIG. 6. The majority of subnets are categorized using Rule 1, where the IP WHOIS and the origin AS organization names match patterns of the same category. Notably, the routing relationship rule (Rule 4) captures 609subnets that are originated by an ISP AS but used exclusively by some hospital according to IP WHOIS registration data. These are hospital subnets that would have bee missed if only BGP routing information was relied upon to categorize them.

[0103] One of the most interesting findings from the mapping results is that many hospitals tend to share their subnets with other hospitals. These overlaps, or interdependencies across different hospital networks, complicate the attribution of network disruptions to individual hospitals. These overlaps are likely due to membership in a larger healthcare system that shares common network infrastructure. In other cases, hospitals with collaborative research projects also host services on each other's networks.

[0104] Network mappings were validated in two ways: using address-space mappings provided by two hospitals as ground truth and comparing with labels from the Autonomous System database ASdb.

[0105] As partial validation of the network identification and categorization, the addressspace mappings of two hospital systems were obtained: a large hospital system operating over 30 hospitals (with 500 beds) across six large geographic regions, and another small rural hospital with 160 beds in Southern California.

[0106] For the large hospital system, eight out of ten advertised / 24 blocks spanning four distinct DMZs were mapped. Of those ten prefixes, seven were originated by two ASNs under the hospital's control, five of which were mapped as Hospital. The discovery and mapping methodology was able to discover most of the network blocks used by the hospital. Moreover, three of the prefixes were originated by a data center's ASN, all of which were mapped as Cloud. One out of four sub- / 24s suballocated from the ISP was also mapped (the other three proved unmappable: two had no hosts on Censys and one lacked any CT log entries).

[0107] For the small rural hospital, one / 24 prefix spanning one DMZ as ISP was mapped. The prefix is owned by and originated from its ISP and, as such, further refinement of the mapping category was not possible.

[0108] ASdb was used as a coarse-grained dataset to compare the mapping labels with. To make the mapping comparable, the labels were limited to AS granularity using only AS organization names, and only look at the ASes registered in the U.S. according to WHOIS information. There are two categories in ASdb that were relevant to the study: the layer 1 category "Healthcare Services" and the layer 2 category "Hospitals and Medical Centers". Themappings labels were compared with the two ASdb categories using the latest January 2024 ASdb data. Results of the comparison are reported in Table 5.

[0109] Table 5: Categorization of Hospitals compared to that of ASDB. All ASes are limited to US only based on their WHOIS country.Classification Label Total In Both Only in ASDB Only in DisclosedMethodCount Example Count ExampleDisclosed Method (AS 577 - - - -Only)ASDB Layer 1 : Healthcare 2034 524 1510 AS29207 Takeda 53 AS55215 TrinityServices Pharmaceuticals HealthU.S.A., Inc.ASDB Layer 2: Hospitals 611 318 293 AS62635 259 AS32117 and Medical Ctrs VPNWholesaler.c Berkshire Health om Systems

[0110] Out of 577 ASes that were labeled as Hospital, 524 of them are also labeled as "Healthcare Services" in ASdb. Naturally, many ASes that belong to pharmaceutical companies, health insurance companies, biological research institutions, and other healthcare-related organizations were left out- they fall under the "Healthcare Services" label in ASdb, but are not hospitals. For the 55 ASes that are not labeled as "Healthcare Services" in ASdb, they were manually verified to all be hospitals. Based upon this overlap and the manual verification, it was confirmed that all ASes labeled as Hospital are indeed hospitals.[OHl] For comparison with the ASdb layer 2 category "Hospitals and Medical Centers", the labels only share 318 common ASes (55% of the Hospital ASes and 52% of ASdb's "Hospitals and Medical Centers" ASes). Upon closer scrutiny, many of the ASdb ASes are not hospitals at all, and it misses many hospitals caught using the disclosed techniques. The result is not surprising, as ASdb layer 2 categories are less accurate than the layer 1 categories.

[0112] In another aspect of the disclosed technology, example techniques to detect hospital disruptions are provided.

[0113] An organization's network might experience a disruption due to a variety of factors, some transient or permanent (e.g., scheduled maintenance, network reconfigurations), and some unplanned (e.g., power outages, hardware failures, cyber-attacks). To capture the impact ofunplanned disruptions, while minimizing the influence of smaller disruptions, the networks of an organization were aggregated and collectively monitored for correlated failures. To model such failures, an approach to finding correlated weather events among state-ASN pairs, as described elsewhere, was adapted by modifying it to the context of finer-grained networks like hospitals and other organizations.

[0114] In the description that follows, an example methodology for detecting dependent disruptions is disclosed, and two additional heuristics for distinguishing different types of failures are provided. In an example demonstration, the methodology was applied on two years of active probing data of hospital networks and identified disruptions were characterized.

[0115] A technique to detect dependent failures is presently disclosed .A single host failure F was defined as a host reachable at time t - 1, but unreachable on all the probed ports at time t. Consider the collective pool of all hosts belonging to an organization Horg. For each consecutive scan at time Z, the number of responding hosts Torg(f) and the number of failed hosts Forg(f) were counted.

[0116] For each Horg, for scan time Z, the average failure probability Pn (t,W) was calculated, using a window W (a set of W days prior to the current scan time Z) of historical data. The failure probability is the ratio of the number of disruptions Forgto the number of responding hosts Torg.

[0117] Binomial testing gives the probability of observing F failures in N trials, where N is the number of hosts in the organization pool Horg. If this probability is less than 0.01%, the disruption event was considered to be dependent. Thus, Equation 2 was used to solve for Fmin,Binom(0.0001 . (2)

[0118] For each day, if the number of failures F was greater than Fmin, the event was considered to be a dependent correlated failure event.

[0119] Techniques to separating transient, permanent and unplanned failures are presently disclosed.

[0120] While the binomial test hinges on the ability to filter statistically significant failures from others, it still has a few limitations. First, if the prior baseline of failure is very low (e.g.,Fmin = 1), then even a single failure can be statistically significant (e.g., a few hosts go offline) and are flagged as a dependent failure. Second, the binomial test only accounts for failures at a particular scan time, and does not measure the duration of the failure across multiple scans, i.e., the time until recovery.

[0121] To address these limitations, two intuitive heuristics are disclosed.

[0122] Magnitude of Failure. Z-score can be used to quantify the scale of the failure against a baseline (an unplanned disruption should have a higher magnitude of failure). Assuming N is the number of hosts in Horsat the scan prior to / , and Forgis the subset of those that failed, the Z-score for the disruption may be calculated as:

[0123] Where LI is the expected number of failures at time t under the assumption of the binomial distribution, and o is the standard deviation of the number of failures:F = N - PH(4)

[0124] Length of Failure and Extent of Recovery. For each dependent failure event, the recoverability of the event was tracked using the IP set from the last successful scan (t -1) prior to the failure event ( / ). For each future scan, the extent of recovery was calculated as the percentage of that IP set that is reachable.

[0125] Table 6 summarizes the obtained inferences based on a combination of these heuristics. Transient failures are short and recover fully, and usually fewer in numbers (hence low magnitude of failure, Z-score). Permanent failures are long, have a low rate of recovery, and a low magnitude of failure (Z-score). Unplanned failures in contrast are sudden, resulting in a large magnitude of host failures (high Z-score).

[0126] Table 6: Classification of disruption types by their statistical and recoverability characteristics.RecoveryZ-score - Failure TypeDuration ExtentLow Short High (-100%) TransientLow Long Low (-0%) PermanentHigh Varies Varies Unplanned

[0127] Stabilization Point. Unlike transient or permanent failures, unplanned failures (such as a cyberattack) may result in permanent changes in the network - hosts might be reconfigured, the network may be restructured, or newer firewall rules may be added. Thus, the recovery might not result in the same availability of hosts as before, changing the organizational footprint permanently.

[0128] To account for this, the stabilization point was calculated for each unplanned disruption event. Intuitively, this represents the beginning of a Ws sliding day period where the percentage recovery slows down or stops completely (no more than e change in recovered hosts). If in fact, a complete (100%) recovery is seen within the window, the stabilization point was set to the time of complete recovery.

[0129] However, this presents an additional challenge: A large unplanned event might result in a cluster of dependent disruptions (often on consecutive days), resulting in different stabilization points for the same unplanned event. To account for this, a simple clustering algorithm was applied to group failures that were within a window of Wc days of the start of each other, and the furthest stabilization point was considered as the point of stabilization for the entire event.

[0130] To evaluate the above methodology, historic probing data from Censys was used, and the extent of transient, permanent, and unplanned failures was then quantified for each hospital organization. Censys's Universal Internet Dataset was used, and contains a daily snapshot of the entire IPv4 address space and the services running on each host. Filtering was based on the network prefixes of U.S. hospitals, mapping all hospitals from the seed dataset to their pool of IP addresses using the methodology described in preceding sections. Filtering for transport layer liveness and aggregate ports across the same IP address on any given day was also implemented. Further, the daily granularity of the Censys dataset does not allow for detection of failures that are shorter than a day.

[0131] From Censys, 85.2 million daily probes were collected for 46,175 unique IP addresses representing 1,720 hospital organizations during this period. For each hospital organization, its set of responding IP addresses N were restricted to just the IP addresses identified as on-premises according to the analysis explained above (addresses labeled asHospital). Intuitively, these hosts were expected to most likely to be affected by a significant disruption across a hospital.

[0132] The choice of parameters for the binomial test, recovery, and stabilization heuristics are summarized in the description that follows. For these heuristics, parameter values were derived empirically based upon exploratory analysis of the data.

[0133] Dependent Failure Events. For each of the 1,720 organizations, it was tested whether each day can be a potential dependent disruption event. A window W of length 30 days was chosen for the binomial test. This resulted in 1,720 x 820 (Jan 2023 - Mar 2025) = 1,410,400 events as the upper bound of the number of dependent failure events.

[0134] Observed across all the organizations were 17,552 dependent failures (1.2%), and 1,570 (91%) of the organizations had at least one dependent failure event across the time period. Figure 7 shows the distribution of the number of total responding hosts, number of dependent failures, and the Fminvalues across all hospital organizations. Figure 7 further shows the median number of responding hosts per organization is 25, and the median average Fminrequired for a dependent failure is low even for larger aggregates.

[0135] Failure Heuristics. To classify transient failures, the shortest possible recovery duration was a day since Censys data has daily granularity. If an event recovers completely within one day, it was classified as a transient failure. For permanent failures, a low Z-score value of < 2 was used, and a long recovery window WR of 90 days over which to calculate recoverability. If the event exhibits low (< 5%) recovery within the window, it was classified as a permanent failure. A Z-score threshold of greater than 5 was used to capture sudden, drastic, unplanned failures.

[0136] Classifying the 17,522 dependent failures, 12,146 (70%) were labeled as transient, 2,596 (15%) labeled as permanent, and 2,745 (16%) labeled as unplanned. Observed were 689 (40%) organizations with at least one unplanned failure over the time period.

[0137] Recovery and Stabilization Heuristics. To calculate the recovery or stabilization point, a short sliding window Ws of 7 days was used. Events were clustered within a window of 3 days (Wc = 3) and the further recovery point for each group was calculated. Importantly, the stabilization threshold e was kept low at 5% for determining stabilization.

[0138] For the 2,745 unplanned failures, Figure 8 shows the distribution of the length of failure. A mean failure of 15 days was observed, as well as a median of 6 with a long tail offailures lasting up to the recovery window of 90 days. It was observed that 8 unplanned failures (0.3%) never stabilized to within the threshold of 5% recovery.

[0139] Common challenges with systems for large-scale detection of network failures are confirmation and attribution. The above description characterized a variety of disruptions among the hospitals in the study, but it is challenging to obtain ground truth about them: significant dependent failures in particular are very plausibly network issues at the organization, but they could also be due to substantial changes in infrastructure deployment, firewalls blocking active probes, etc.

[0140] The unfortunate occurrence of hospital ransomware, however, provides a serendipitous source of ground truth. Hospital ransomware events are generally made public, and hospitals frequently (although not always) shut down their networks when they are ransomed.

[0141] In the description that follows, publicly disclosed, historical ransomware events that have affected hospital organizations in the past two years were used as a ground-truth data set of major disruptions to provide partial validation of the disclosed network disruption measurement approach.

[0142] A combination of two popular ransomware data sets - Critical Ransomware Attacks (CIRA) and Comparitech - were used to gather historic ransomware events. Where possible, the date first reported using public announcements from the organization was refined. Only events affecting U.S. -based hospitals that provide direct patient care were considered, and those involving third-party vendors or other auxiliary healthcare services that were victims of ransomware attacks were filtered out.

[0143] A total of 53 hospitals systems that were affected by ransomware events from Jan 2023 to Mar 2025 were collected by combining the historical datasets. Categorizing their networks in accordance with disclosed techniques, observed were 22 (49%) hospitals with Cloud-only prefixes, 10 (22%) with ISP prefixes in addition, and 21 (36%) with Hospital prefixes.

[0144] For each of the 21 hospitals with Hospit l prefixes, the dependent failures were calculated for each organization, and 15 of them were observed to have dependent failures during the ransomware event. The remaining six hospitals did not have sufficient externally-visible events for the methodology to detect a dependent failure. Three (Cayuga Medical Center, TaylorRegional Hospital, and Warren General Hospital) of the remaining hospitals had a very small network footprint of four or fewer responding hosts, with at most two of those hosts failing during the ransomware event. The other three (Fairfield Memorial Hospital, Anna Jaques Hospital, and Palomar Health) did not have any visible network disruptions during the event.

[0145] Table 7 shows results for the 15 hospitals for which a dependent failure overlapping a known ransomware event at the hospital was detected. For each hospital the length of failure was calculated as previously discussed. The nearest dependent failure date, the failure categories, and the unattributed failure groups for each unplanned failure were also reported. For example, FIG. 9 shows a time-series graph of the number of responding hosts for the University Medical Center (UMC) in Lubbock, Texas. The ransomware event was first reported on Sep 26, 2024 and the first observed dependent failure event was on Sep 28, 2024, two days later. The stabilization date on Nov 9, 2024 was inferred, 42 days after the event was observed. Each dependent failure was categorized as transient, unplanned, or permanent.

[0146] Table 7: Matching past ransomware events on targeted organizations (hospital systems) with the dependent failure inference on Censys Hospital blocks. The total number of days with dependent failures was 819 days (Jan 1 '23 - Mar 31 '25). T - transient failures, P- permanent failure, U - unplanned failure, U* - unplanned failure groups (including the observed event).

[0147] The average number of host responses for the 15 hospitals in the ransomware set was large (70), and comparable with the average number of hosts across all hospitals in the dataset (FIG. 7). Such a large number of responding hosts supports more reliable detection. Further, the average length of failure for the hospitals was 20 days, indicating that unlike the mean length of unplanned failures across all hospitals, the ransomware set is notably higher. Such long durations highlight the utility of incorporating this feature into the detection method. Finally, the average number of days with dependent failures for the 15 hospitals ranged from 3 (0.4%) to 30 (3.7%) days, with a mean of 15 days (1.8%) - similar to the average number of dependent failures across the entire hospital dataset (FIG. 7).

[0148] The foregoing analysis focused at a more granular scale on the individual hospitals making up the U.S. health delivery ecosystem. But in doing so, the following broad scanning capabilities were leaned upon in multiple ways: first, the mapping of hospitals to IP addresses relies, in part, on the viability of scanning services at every IPv4 address and recording which hosts answer with strings associated with particular hospitals. Second, having identified per- hospital IP addresses, the availability of longitudinal data about active services at those addresses makes it possible to identify extended periods of downtime and correlate - post hoc - with external data sources. Together, these approaches enable mapping of the robustness of U.S. hospital IT infrastructure.

[0149] The coverage of networks larger than / 24 was also studied. The coverage metric was reported for each of the networks, calculated as the number of / 24 blocks scanned for divided by the total number of externally visible / 24 blocks in the network. All hosts seen by Censys in the network range were used as a proxy for all externally visible IP addresses and then grouped by / 24 prefixes.

[0150] The results are reported in FIG. 8 as a cumulative density function (CDF) of the coverage metric. For more than half of the networks, the entire network was able to be scanned. On average, 86% of the network was scanned.

[0151] FIG. 10 shows the CDF of coverage for networks larger than / 24. Another observation is that most large / 16 networks only have a few externally visible / 24 blocks (not shown in the FIG. 10). This confirms the hypothesis that scanning the entire network is not necessary, and capping them at / 24 is sufficient to cover most of the network while saving lots of time during scanning.

[0152] In the analysis, the mapping process described above was re-ran every month to capture the evolution of hospital network boundaries. Throughout January to May 2025, the following percentage of changes in the network boundaries of hospitals per month were observed: smaller than 4% of Hospital subnets, 3-9% of I SP subnets, and 25-45% of Cloud subnets. Unsurprisingly, the Cloud subnets are the most volatile, as their IP addresses may be easily rotated and reused, whereas the Hospital and I SP subnets are more stable. Since the analysis is based on Hospital subnets, it can be concluded that the boundaries of hospital networks are relatively stable (< 5% monthly change) over time and mapping them once a month is sufficient to capture these changes.

[0153] Overlaps and interdependencies between hospital networks are common, where one subnetwork are associated with services from multiple hospitals. FIG. 11 shows the distribution of 437 subnets that were associated with more than five hospitals. Some hosts as many as 130 hospitals. Unsurprisingly, most of these networks were categorized as Cloud. However, the number of Hospital and ISP networks are not negligible.

[0154] FIG. 12 shows a matrix of hosting relationships across Hospital and ISP networks hosting more than five networks, illustrating the dependencies between hospitals. The x-axis shows the networks, while the y-axis shows the hospitals that are hosting services in them. Bright spots in the matrix indicate that a hospital is hosting its services on a hospital network in the AS. The horizontal traces indicate that the hospital is hosting its services in multiple networks, whereas the vertical traces indicate that many hospitals are hosting their services on the same AS.

[0155] In another aspect, the present patent document discloses techniques relating to a "real-time" surveillance tool utilizing a machine learning model to rapidly identify HDO ransomware attacks. In an example embodiment, the foundation of the model are signals, which are defined in the ideal state as externally visible data that is highly correlated with a ransomware attack on a hospital, accessible without requiring hospitals to implement new functionality, and, ideally, longitudinally available, allowing for historical comparison against previous attacks. The following description discloses the development of the signal classification schema and initial signals which were interrogated, along with proof-of-concept demonstrations on past attacks.

[0156] Ransomware attacks affecting healthcare delivery organizations (HDO) pose substantial risk to patient safety and the financial and operational viability of our nationalhealthcare infrastructure. Current regulatory frameworks and status quo HDO behavior results in the delayed reporting of ransomware attacks on hospitals with attendant impacts on patient safety as well as the ability to rapidly provide resources to affected institutions and prepare unaffected hospitals within the "cyber blast radius" for the sequelae of increased patient diversion and workflow burdens. Currently, there is no proactive, ubiquitous, and timely mechanism to detect healthcare ransomware attacks without voluntary HDO reporting.

[0157] Example techniques related to a machine learning surveillance tool are disclosed. A case study analysis of using FHIR endpoint scanning to detect a recent attack on Ascension hospitals is also provided, demonstrating the promise of the disclosed techniques.

[0158] Table 8 lists several categories of signals that can be evaluated, in accordance with disclosed techniques, and their properties as per the classification scheme designed for each signal.

[0159] Table 8: Signals and associated Correlation, Access Restrictions, and HistoricalAvailability

[0160] The present patent document disclosed a measurement procedure which involved Censys and other passive measurement services (e.g., DomainTools) to test for host downtime in a hospitals' network address ranges, around the days of a ransomware attack.

[0161] To scale up such methodology, an example study was performed and involved using hospital data collected by American Health Association (AHA). A sample of all hospitals in California was provided and contained 421 hospitals. Using this sample, a large array of IP addresses hosting services within each hospital was built. This was achieved by focusing on discovering additional domains, and then focusing on discovering hosts associated with those domains. A seed domain is used to discover additional domains associated with that hospital- the website domain provided by the AHA dataset. Using Censys's Certificate data, a collection of X.509 certificates collected by Censys, observed as a part of a TLS handshake during its periodic scan of the internet. Those certificates that did not include the registered domain portion of the website domain in its Certificate Name (CN) were filtered out. For the certificates that remain, all other domains included in the certificate’s Subject Alternative Name (SAN) were appended to the domain list. This combination of SAN and CN filtering helps in discovering related domains given a registered domain for a hospital (FIG. 13). FIG. 13 shows an example methodology for finding hosts given website domain. Using this methodology for the 421 hospitals in the list, 620 total associated domains were obtained.

[0162] In some implementations, the disclosed methodology can involve searching for hosts (available on the internet) using a Censys's Host search. Specifically, all IPs with the domain included as a part of the banner message captured during Censys scans can be filtered. Using these two filtering techniques a combination of 185,501 host IP addresses for the 421 hospitals in California was obtained. FIG. 13 shows the Censys search queries used to discover related domains (using Censys Certs) and its associated IP addresses using Censys Host banner searches.

[0163] The disclosed techniques enable active monitoring of the availability of these IP addresses using active scanning tools like ZMap. Using such a scanning tool, a pipeline that scans these health related IP addresses, and provides a view of service availability can beachieved. To determine the ports to actively scan, the active ports returned by Censys Host search can be used.

[0164] Some disclosed embodiments rely on the Electronic Health Record (EHR)'s HL7® FHIR® Standard - a standard used to exchange healthcare information between Health Information Exchange (HIE) -to publicly monitor availability of a hospital's EHR. During a ransomware attack, a disruption in the use of such EHR systems by the hospitals is expected to be observed, either because the EPIC proxy or server experiences downtime or because the hospital network proactively removes their EHR offline. In one example embodiment, to observe such a disruption externally, the unauthenticated endpoints (e.g., / metadata) provided by the FHIR standard can be periodically queried. Using endpoints from the two large EHR providers (EPIC and Cerner), active monitoring on 2087 healthcare organizations (434 customers of EPIC and 1653 customers of Cerner) was achieved in an example demonstration.

[0165] During a ransomware attack on Ardent Health care networks, on November 23rd, the disclosed technology was used to confirm that an authentication request for metadata sent to the ardent healthcare endpoint timed out. FIG. 28 shows a graph noting Ardent handshake times out because of missing Server hello. Ardent's outage continued to be monitored using disclosed techniques and its restoration was observed on Jan 29th, 2024 (FIG. 14).

[0166] On Jan 31st, 2024, Ann & Robert H. Lurie Children's Hospital of Chicago in Chicago, Illinois, was ransomwared. FHIR polling of Lurie's EPIC API, in accordance with disclosed techniques, showed a downtime for the hospitals on the same day. Reconstructing the timeline between the polling logs (which were then at 6 hour intervals) and public disclosure of the outage (via Lurie Hospitals Twitter accounts), the polling captured the attack hours earlier. In fact, the first announcement of any outage (starting with a generic report of a disruption to telephone on official social media profiles) was 80 minutes after the latest polling window. Further, the public notification that MyChart was down was announced nearly 48 hours after the earliest estimate provided by the disclosed technique. It was noted that the API access was being restored (as noted in polling logs) on 28thFebruary at 20: 14:33UTC.

[0167] FIG. 15 shows a reconstructed timeline for Lurie Children Hospital EHR outage (Jan 31, 2024). Times in PST.

[0168] Services within a hospital system periodically "beacon" externally to obtain information. Some disclosed embodiments use NTP beacons (a protocol used by digital clocks to synchronize time) to observe a "gap" in periodic beaconing. In an example demonstrationusing three NTP servers hosted at the University of Maryland, Healthcare IP addresses in the global IP space that call an NTP service to set their internal clocks were studied. These NTP servers are geographically distributed throughout the continental US in high-availability cloud VPS providers (Newark, NJ, and Lenexa, KS, with lonos and Portland, OR with OVH).

[0169] Manually selecting 5 hospital system network prefixes (IP address range), periodic calls from all of them were observed. Some of these (e.g., Dignity Health) are large hospital systems (which operate many individual hospitals), and hence more periodic lookups are seen from those IPs. For Ardent Health and Tallahassee Memorial, which are smaller hospital systems, periodic beaconing is seen happening at a lesser frequency.

[0170] To determine the practicality of detecting HDO downtime over NTP beacons, variance within a prefix was characterized. To do this, a service that periodically reads a file of IPv4 prefixes was used. The inbound NTP requests were parsed and checked to see if the source address was covered by a prefix in the trie. Periodically, every hour, the address seen, the prefix they matched to, and its associated count were saved.

[0171] FIG. 16 shows NTP Lookups resulting from 5 hospital systems over a period of 5 days. FIG. 17 and FIG. 18 show a plot of the NTP request count observed over a span of 5 days. FIG. 17 shows the plot for all hospital systems, while FIG. 18 zooms in on one particular system near the bottom end of FIG. 17. Two important takeaways from both FIG. 17 and FIG. 18 are noted: 1) ehe beaconing is variable, with larger hospital systems beaconing information at a higher rate, and 2) even smaller hospital systems show almost continuous beaconing, and only short periods (few hours in a day) of no beaconing.

[0172] The disproportionate levels of NTP requests from large hospital systems possibly suggests the use of Network Address Translation (NAT). NAT allows multiple devices within a private network to access the internet through a single public address by translating a private IP address to a public address. It is suspected that some HDO have all their traffic externally going out of one point or gateway, i.e., their NAT.

[0173] Some disclosed embodiments rely on machine learning for weighing signals. The stated problem can be defined as a time series forecasting and modeling. It is a vital tool in decision making as it involves analyzing historical data patterns to predict future trends. Time series models capture underlying patterns, seasonality, and trends by applying statistical and machine learning techniques to provide valuable insights and forecasts.

[0174] Time series can be further divided into univariate and multivariate. Univariate time series analysis focuses on a single variable's historical data to understand its past behavior; this method is simple but may fail to capture complex relationships. Multivariate time series analysis deals with datasets containing multiple variables that are interrelated and can influence each other's behavior over time.

[0175] Some disclosed embodiments deal with multivariate time series forecasting using Vector Auto Regression (VAR). In a VAR algorithm, each variable is a linear function of the past values of itself and the past values of all the other variables. By estimating the coefficients, VAR models can capture the dynamic relationships between variables and then make predictions about their future values. Some disclosed embodiments use deep learning models like LSTM. It is well suited to handle sequential data which makes them ideal for time series analysis. LSTM can capture long term dependencies and retain information over a longer period.

[0176] Some disclosed embodiments rely on publicly available data. For example, the Lantern project, sponsored by the Office of the National Coordinator for Health Information Technology, has gathered availability data for FHIR endpoints since July of 2020 and with daily updates starting in July or 2023. This public dataset (available at github.com / onc-healthit / ) can be implemented in disclosed embodiments to provide longitudinal data. For example, this public dataset can provide an additional source of daily HDO availability data. Furthermore, it can provide, both historically and going forward, a further set of FHIR endpoints beyond those already identified according to an embodiment of the disclosed technology.

[0177] Part of the challenge in monitoring health delivery organization (HDO) endpoints is to identify what those endpoints are.

[0178] An example embodiment relates to a repeatable methodology for identifying such candidate endpoints and then feeding these to machine learning systems for further refinement. In an example methodology, a dataset from the American Hospital Association, which includes Web sites addresses for most of its members (including over 4300 distinct domain names), was used. From this set, a combination of Certificate Transparency logs (i.e., to identify TLS certificates for the registered domain) were used and other subject alternative names that might be included in the certificate were identified. Also identified were fully qualified domains, for those registered domains, as seen via passive DNS monitoring (e.g., such as provided by DomainTools) or as seen in the application banners seen during IP-level scanning (e.g., as found by Censys). This almost doubled the set of candidate sites. From all of these candidate fully-qualified domains (as well as any associated CNAMEs in DNS), both their DNS infrastructure and their IP addresses (i.e., "A records") can be identified. For each of these addresses, the enclosing network prefix was identified and, across all of these prefixes and across a range of port numbers widely used in health care, each was scanned with high frequency (every few hours).

[0179] An ongoing challenge for any external scanning effort is validating results: for example, when a FHIR endpoint at a hospital no longer responds, is it due to a network issue (perhaps entirely unrelated to the hospital network), a maintenance issue, or does it actually reflect a ransomware attack?

[0180] Some disclosed detection techniques are trained and validated using a ground truth data set. When training and validating a detection technique, it is extremely useful to have ground truth data about when ransomware attacks take place and when hospitals disable network services. In an example demonstration, news and social media data were used as sources of ground truth to validate the availability of FHIR network endpoints at three Ascension hospitals in Wisconsin, Illinois, and Michigan (Ascension was recently a victim of ransomware in May 2024). FIG. 19 shows example sources of ground truth data. Specifically, FIG. 19 shows tweets from a hospital employee and news agencies reporting on the attack which can be used to compare and correlate with FHIR measurements.

[0181] FIG. 20 and FIG. 21 show, respectively, timelines of the availability of the FHIR endpoints at the three hospitals together with the tweets and news reports. Each hospital has a time series curve for its FHIR endpoint roughly every 2-3 hours. If the scanner could reach the FHIR endpoint, the curve has a point at the bottom. If it could not, it has a point at the top. The tweets and news report are shown in FIG. 20 and FIG. 21 as dashed vertical lines. FIG. 20 shows data for a week that spans the start of the ransomware attack. FIG. 21 shows the same data but a much longer time span of nearly a month.

[0182] These results show both the promise of scanning FHIR endpoints. The FHIR scanning clearly reveals and reflects the downtime of Ascension hospital services at Wisconsin and Providence (Michigan), and for extended periods of time (from May 8 to May 25 for Providence, and from May 8-23 and ongoing again from May 27). The FHIR scanning also detects FHIR downtimes before reports on Twitter or in the news, as highlighted in FIG. 20, further demonstrating its value as an early detection mechanism. The FHIR scanning also showsthat Ascension Illinois was not impacted in the same way as Ascension Wisconsin and Providence (Michigan).

[0183] In some disclosed embodiments, scanning results are visualized. In one example embodiment, scanning results are visualized on a dashboard that shows live and historical status of hospital networks for the country, including which hospitals have been attacked by ransomware, how long their networks and services have been impacted by the ransomware, the signals indicating network status, and any additional data (e.g., news, social media, etc.) correlated with attacks.

[0184] Some disclosed embodiments utilize databases (e.g., Censys, Shodan) or other crawling services to actively measure parameters (e.g., host connectivity) across the internet's address space. In an example embodiment, Censys is used to capture a lack of IT service availability during a ransomware attack by checking for host downtime in a hospitals' network address ranges around the days of the attack. In one demonstration, to get a comprehensive list of ransomware attacks, the Temple University Dataset - a manually curated list of ransomware attacks from news articles, and social media accounts — was examined. Entries that only belonged to the healthcare and public health sectors were filtered out and only those attacks dating after 2021 were considered (from when Censys data available is available). This resulted in 72 different ransomware attacks on hospitals from 2021-2023. Several historical attacks and two recent ransomware attacks were selected for further proof-of-concept analysis. Each hospital with its ransomware date and news articles is listed in Table 9 below.

[0185] Table 9: Historical Ransomware Reference for Proof-of-Concept Analysis.

[0186] To infer which hosts were active or inactive during historical ransomware attacks, historical services and the associated IP addresses they resolved to during the ransomware period were located. DomainTools Passive DNS (pDNS) repository was leveraged to discover such historical services (Domain Tools, Seattle, Washington). Domain Name System (or DNS) is the protocol used to successfully resolve a human-parsable domain name to an internet IP address that locates the host. pDNS is a collection of successful DNS resolutions from volunteering name servers around the globe, to further the scientific community's understanding of internet connectivity. Using this pDNS repository, the most common services announced for each hospital domain, roughly associated with the period of the ransomware attack, were filtered out. The IP address for each service was tracked, and Censys's host database was queried to search for their availability surrounding and including those periods.

[0187] A typical Censys host response includes several fields: the services available on that host, and information specific to each service (e.g., encoding information, certificate names, etc.), the time when each service was observed, other metadata on that host (e.g., location, operating system, etc.), and lastly the last updated timestamp for that host entry. If a host is down (such as during a ransomware attack), Censys does not complete its scanning for that host and instead logs no data in its database. Further, the last updated time for that entry would be the last successful scan of that host.

[0188] FIG. 22 shows example results obtained using a disclosed methodology as applied to data collected from Atlantic General Hospital's ransomware attack. These two services were manually selected since both hosts continuously beacon to Censys daily and were hosted within the Atlantic General network. Censys shows a drop-off coinciding with the time the ransomware attack was reported (information derived from the Temple dataset) and failed to update its database. The data also captures the disparity in bringing different services back online (e.g., Kronos came back up a day quicker than the exchange traffic) - a fact most likely associated with the complex manual process required to restart the services that went down. The processwas repeated for the other hospitals in the study, and similar correlations between network drop- offs from services and the ransomware attack were observed (FIGS. 23-25).

[0189] In Oakbend Medical Center's ransomware attack, the autodiscover service (a service that resolves to the local exchange server) came back up much quicker than the portal gateway. In Tallahassee Memorial Hospital's case, it was observed that one service (Optilink) that went down around the same time as the ransomware attack came back up nearly 22 days later. Lastly, with Scripps Health, a number of services were seen to go down at roughly the same time, and a wide disparity in services coming back up was observed (e.g., services such as Electronic Health Records (Haiku) come back up much later).

[0190] In some disclosed embodiments, scanning for downtime is automated. In one example demonstration, a disclosed technique involving automated scanning was applied to a ransomware attack on 9thNovember, 2023 that occurred in Tri-City Medical Centre in Oceanside, California. The approach involved first identifying the two address spaces Tri-City uses - 205.167.50.0 / 23 and 12.157.153.0 / 24 for their network services. Around the periods corresponding to the ransomware attack, the availability of every Censys host within these address spaces was graphed (Figure 26). FIG. 26 shows that around five services (corresponding to Windows IIS server, exchange server, etc.) were seen to go down around a similar time as the onset of the ransomware. FIG. 26 also shows that these services were restored two weeks later on November 21st(which corresponds to news reports that Tri-City had restored most of their systems).

[0191] It is important, however, to assess the false positive nature of these signals - in other words, how often do they go down due to non-ransomware related causes? In one example study, data for the different services identified as affected in the Scripps Ransomware attack was extracted, only for the full year of 2021 instead of the four week period of the attack in May. Interestingly, two perturbations are observed (FIG. 27). The first is attributed to the ransomware attack in May, 2021. The second, however, occurs later in the year (Dec, 2021) and does not coincide with a known ransomware attack. As is the case with any large-scale measurement, the most likely answer for this is attributed to a failure / error in Censys's scanning - though planned network downtime, utility maintenance, or a number of other causes could certainly be plausible. Combining multiple signals across additional categories (some described below) can enable differentiation from such 'false positives.'

[0192] Services within a hospital system periodically "beacon" externally to obtain information. Some disclosed techniques enable visibility into their behavior including where they are coming from and the ability to observe their inactivity during certain periods such as ransomware attacks.

[0193] Network-time protocol (NTP) is one example of a beacon. Some example systems based on the disclosed technology use an NTP server that allows visibility of IP addresses in the global IP space that call an NTP service to set their internal clocks. Such calls happen when the NTP server is elected in the load-balancing rotation. If there is a drop-off for long periods of time without NTP probes from hospital networks, the system can be used to infer a disruption in service from such systems.

[0194] During the process of providing care to patients, a hospital records, processes, and reports vast amounts of healthcare data to a variety of other HDOs (Health information exchanges), insurance entities (billing claims), and public health agencies (vaccinations, deaths, births, etc.).

[0195] It can be assumed that the operational impact of an ongoing ransomware attack, information regularly reported is very likely to cease or be delayed when workflows shift from digital to analog. Embodiments of the disclosed technology can be implemented to find such gaps in information which coincide with ransomware periods. In accordance with one example embodiment, if the number of vaccinations delivered within the last 24 hours is seen to go down drastically, it may be correlated with a disruption in delivering health care (such a ransomware attack).

[0196] Some disclosed embodiments include data sets assembled to train machine learning algorithms to develop predictive models. Such predictive models can, e.g., surveil national health infrastructure to identify successful ransomware attacks at the earliest possible moment, allowing for rapid assistance and regional load balancing.

[0197] Disclosed techniques may be implemented to investigate diverse types of signals. For example, both traditionally associated healthcare signals such as health information exchanges, insurance claim submissions, and FHIR API availability, as well as novel data streams like social media posts, regional EMS dispatch, and external phone / fax system availability, among others. An example of a promising signal candidate is the electronic health information exchange (HIE). Most modern HDOs currently share medical records with otherorganizations as required under the 2009 HITECH act, wherein a requirement of EHR Use is participation in HIE.

[0198] Some disclosed embodiments use deep learning approaches to model time series data to predict the likelihood of attacks based on abnormal signal behavior using long short-term memory (LSTM) recurrent neural networks. In some implementations, models are trained with retrospective institutional data and subsequently trained on simulated data if necessary. Input features can be represented as a time series of validated signals. Continuous features can be normalized prior to modeling.

[0199] In some disclosed embodiments, the LSTM model is trained on a temporal sequence of signals, which returns a hidden vector for each state that subsequently passes into a fully connected layer prior to the final prediction. FIG. 29 shows an example architecture of an LSTM based on the disclosed technology. The LSTM model can be built on retrospective data and / or optimized with prospective data.

[0200] One example function of the model is to uncover which signal features contribute to a given outcome. To aid in model interpretation, some disclosed embodiments use the SHapley Additive exPlanations (SHAP model). Once given SHAP values, interpretability is improved because features are concrete and assigned importance. Features can then be validated based on scientific rationale and further analysis.

[0201] Model performance can be evaluated, for example, via area under the receiver operating characteristics curve (AUC), Fl -score, sensitivity, specificity, precision, recall, and accuracy. For example, metrics can be averaged by 10-fold cross-validation, in which data can be split into 10 folds (in which 9 folds will serve as the training set and 1-fold as validation set) and calculated repeatedly until each fold has served as a validation set.

[0202] Some disclosed embodiments use trained machine learning predictive models. Existing, simulated, or prospectively collected data streams derived from identified signals can be leveraged in a deep learning approach to train models, including long short-term memory (LTSM) recurrent neural networks, with a primary predicted outcome aimed to identify likelihood of active disruption (e.g., ransomware attack) affecting an organization.

[0203] The identified signals, in addition to the machine learning predictive model, can be leveraged to build a real-time analytic and data visualization dashboard. This dashboard canallow real-time visualization of a map (e.g., map of the United States) and display abnormal activity in established patterns to provide a potential early-warning of a ransomware attack.

[0204] Among other features and benefits, the disclosed embodiments can provide significant advantages to major federal government agencies who are increasingly focused on obtaining better visibility into the occurrence of these critical safety events. In addition, large cloud providers are likely to have interest in the disclosed technology in addition to companies oriented around predictive analytics.

[0205] FIG. 30 shows a flowchart of an example method 3000 for detecting disruptions to a computer network associated with an organization according to an embodiment of the disclosed technology. At step 3010, the method 3000 comprises transmitting a handshake signal to a set of network addresses. In some implementations of the method 3000, at least some network addresses in the set of network addresses are associated with a service or system of the organization. At step 3020, the method 3000 comprises obtaining monitoring data associated with a respective network address in the set of network addresses. At step 3030, the method 3000 comprises determining a responsiveness of the respective network address to the handshake signal based on the monitoring data. At step 3040, the method 3000 comprises obtaining, based on the responsiveness, information related to an activity associated with the respective network address. At step 3050, the method 3000 comprises characterizing the activity as normal or abnormal based on the information. At step 3060, the method 3000 comprises detecting a disruption to the network based on results of the characterizing. In some implementations of the method 3000, the characterizing comprises obtaining a determination as to whether the activity is normal or abnormal using a machine learning (ML) algorithm trained by retrospective data from historical disruptions to the network.

[0206] FIG. 31 shows a flowchart of an example method 3100 for detecting abnormal activities on a computer network according to an embodiment of the disclosed technology. At step 3110, the method 3100 comprises obtaining a first set of network addresses. In some implementations of the method 3100, each network address in the first set is connected to the computer network and associated with an organization. At step 3120, the method 3100 comprises categorizing, based on predetermined rules, each network address in the first set according to a set of categories to obtain a second set of network addresses. At step 3130, the method 3100 comprises determining subnetworks associated with at least some network addresses in the second set. At step 3140, the method 3100 comprises assigning, based on the predeterminedrules, each of the subnetworks to a category in the set of categories. At step 3150, the method 3100 comprises determining a number of the subnetworks belonging to each category in the set of categories. At step 3160, the method 3100 comprises obtaining a characterization of the computer network based on the number of the subnetworks belonging to each category. At step 3170, the method 3100 comprises determining one or more correlated disruptions of the subnetworks. At step 3180, the method 3100 comprises determining an abnormal activity on the computer network based on the one or more correlated disruptions. In some implementations of the method 3100, at least some network addresses in the first set are associated with different organizations. In some implementations of the method 3100, at least some of the subnetworks are shared by the different organizations. In some implementations of the method 3100, the characterization is based on the subnetworks that are shared by the different organizations.

[0207] Various implementations of features of the disclosed technology can be made based on the above disclosure, including the examples listed below.

[0208] Example 1. A method for detecting disruptions to a computer network associated with an organization, comprising: transmitting a handshake signal to a set of network addresses, herein at least some network addresses in the set of network addresses are associated with a service or system of the organization; obtaining monitoring data associated with a respective network address in the set of network addresses; determining a responsiveness of the respective network address to the handshake signal based on the monitoring data; obtaining, based on the responsiveness, information related to an activity associated with the respective network address; characterizing the activity as normal or abnormal based on the information; and detecting a disruption to the network based on results of the characterizing, wherein the characterizing comprises obtaining a determination as to whether the activity is normal or abnormal using a machine learning (ML) algorithm trained by retrospective data from historical disruptions to the network.

[0209] Example 2. The method of example 1, comprising: in response to the determination indicating that the activity is abnormal, using the ML algorithm to determine a probability that the activity is associated with the disruption.

[0210] Example 3. The method of example 1, comprising: determining one or more correlations between the activity associated with the respective network address and additional activities associated with additional network addresses in the set of network addresses, wherein the detecting is based on the one or more correlations.

[0211] Example 4. The method of example 1, wherein the set of network addresses are determined by: determining a seed set of domains associated with the organization using a first dataset, determining a set of domain names associated with the organization using information obtained from a publicly available database, adding the set of domain names to the seed set, obtaining, from a repository of aggregated data, a list of network addresses associated with each domain in the seed set, and mapping each domain in the seed set to network addresses in the list.

[0212] Example 5. The method of example 1, comprising: classifying the disruption according to a set of types, wherein the set of types allows the disruption to be classified as transient, permanent, or unplanned.

[0213] Example 6. The method of example 5, wherein the set of types are classified based on a time duration associated with downtime instances of the network.

[0214] Example 7. The method of example 1, comprising: performing a first scanning of the set of network addresses to identify host endpoints of the network that are internet- connected, wherein the first scanning comprises probing open ports of the network at one or more predetermined times; performing a second scanning of other endpoints of the network to determine an operational status of the other endpoints at one or more additional predetermined times; and identifying downtime instances of the network based on results of the first scanning and the second scanning.

[0215] Example 8. The method of example 7, wherein at least some of the other endpoints are identified using a publicly available resource or the ML algorithm.

[0216] Example 9. The method of example 7, comprising: collating the downtime instances into a dataset; and analyzing the dataset to determine effects to services or systems of the organization.

[0217] Example 10. The method of example 7, wherein the disruption is classified based on a statistical significance associated with the downtime instances.

[0218] Example 11. The method of example 1, comprising: reporting the disruption prior to the disruption being detected by the organization or other entity.

[0219] Example 12. The method of example 1, wherein the monitoring data is associated with at least one of an Application Programming Interface (API) or a Network Time Protocol (NTP).

[0220] Example 13. A method for detecting abnormal activities on a computer network, comprising: obtaining a first set of network addresses, wherein each network address in the first set is connected to the computer network and associated with an organization; categorizing, based on predetermined rules, each network address in the first set according to a set of categories to obtain a second set of network addresses; determining subnetworks associated with at least some network addresses in the second set; assigning, based on the predetermined rules, each of the subnetworks to a category in the set of categories; determining a number of the subnetworks belonging to each category in the set of categories; obtaining a characterization of the computer network based on the number of the subnetworks belonging to each category; determining one or more correlated disruptions of the subnetworks; and determining an abnormal activity on the computer network based on the one or more correlated disruptions, wherein: at least some network addresses in the first set are associated with different organizations, at least some of the subnetworks are shared by the different organizations, and the characterization is based on the subnetworks that are shared by the different organizations.

[0221] Example 14. The method of example 13, wherein obtaining the first set comprises: determining a seed set of domains associated with the organization using a first dataset, determining a set of domain names associated with the organization using information obtained from a publicly available database, adding the set of domain names to the seed set, obtaining a list of network addresses associated with each domain in the seed set, and mapping each domain in the seed set to network addresses in the list.

[0222] Example 15. The method of example 13, wherein obtaining the first set comprises: using subnetting techniques to determine additional network addresses associated with the organization that are not discoverable by a domain name; and adding the additional network addresses to the first set.

[0223] Example 16. The method of example 13, wherein obtaining the second set comprises: determining an ownership or network connectivity property associated with each network address in the first set, wherein the categorizing of each network address in the first set is based on the ownership or network connectivity property.

[0224] Example 17. The method of example 16, wherein determining the subnetworks comprises determining a prefix size of the subnetworks based on the ownership or network connectivity property.

[0225] Example 18. The method of example 13, wherein the predetermined rules include rules based on at least one of pattern recognition, a prefix size of the subnetworks, or routing information.

[0226] Example 19. The method of example 13, comprising: obtaining a map describing a region of coverage provided by the computer network, wherein the abnormal activity is visualized on the map within the region of coverage.

[0227] Example 20. The method of example 19, wherein the map is provided to a device for display.

[0228] Example 21. A method for detecting ransomware attacks, comprising: detecting whether a data network is functioning abnormally by monitoring publicly accessible signals associated with the data network; determining, upon detecting that the data network is functioning abnormally, a probability that the data network is functioning abnormally due to a ransomware attack; and reporting results of the detecting such that individuals are informed of functionality of the data network.

[0229] Example 22. The method of example 21 , wherein the data network is associated with a healthcare facility.

[0230] Example 23. The method of example 22, wherein the reporting occurs prior to the healthcare facility announcing a finding that the data network is functioning abnormally.

[0231] Example 24. The method of example 21, wherein the probability is determined by a machine-learning algorithm trained by retrospective data from historical ransomware attacks.

[0232] Example 25. The method of example 24, wherein the machine-learning algorithm is trained to assign probabilities to aberrant data correlating to likelihood of the ransomware attack.

[0233] Example 26. The method of example 21, wherein the signals are based on one or more of e-mail system responses, emergency dispatches, standard health record reporting, and social media posts.

[0234] Example 27. The method of example 1, wherein the detecting whether the data network is functioning abnormally comprises submitting a request to the data network.

[0235] Example 28. The method of example 27, wherein a lack of response by the data network to the request is indicative that the data network is functioning abnormally.

[0236] Example 29. The method of example 1, wherein the data network comprises a computer network connected to the Internet.

[0237] Example 30. The method of example 1, wherein the data network comprises a telephone network.

[0238] From the foregoing, it will be appreciated that specific embodiments of the invention have been described herein for purposes of illustration, but that various modifications may be made without deviating from the scope of the invention. Accordingly, the invention is not limited except as by the appended claims.

[0239] Implementations of the subject matter and the functional operations described in this patent document can be implemented in various systems, digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Implementations of the subject matter described in this specification can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a tangible and non-transitory computer readable medium for execution by, or to control the operation of, data processing apparatus. The computer readable medium can be a machine- readable storage device, a machine-readable storage substrate, a memory device, a composition of matter effecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing unit" or "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.

[0240] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts storedin a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.

[0241] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).

[0242] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a processor for performing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0243] While this patent document contains many specifics, these should not be construed as limitations on the scope of any invention or of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this patent document in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one ormore features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.

[0244] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. Moreover, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.

[0245] Only a few implementations and examples are described and other implementations, enhancements and variations can be made based on what is described and illustrated in this patent document.

Claims

1. A method for detecting disruptions to a computer network associated with an organization, comprising: transmitting a handshake signal to a set of network addresses, wherein at least some network addresses in the set of network addresses are associated with a service or system of the organization; obtaining monitoring data associated with a respective network address in the set of network addresses; determining a responsiveness of the respective network address to the handshake signal based on the monitoring data; obtaining, based on the responsiveness, information related to an activity associated with the respective network address; characterizing the activity as normal or abnormal based on the information; and detecting a disruption to the network based on results of the characterizing, wherein the characterizing comprises obtaining a determination as to whether the activity is normal or abnormal using a machine learning (ML) algorithm trained by retrospective data from historical disruptions to the network.

2. The method of claim 1, comprising: in response to the determination indicating that the activity is abnormal, using the ML algorithm to determine a probability that the activity is associated with the disruption.

3. The method of claim 1, comprising: determining one or more correlations between the activity associated with the respective network address and additional activities associated with additional network addresses in the set of network addresses, wherein the detecting is based on the one or more correlations.

4. The method of claim 1, wherein the set of network addresses are determined by: determining a seed set of domains associated with the organization using a first dataset, determining a set of domain names associated with the organization using information obtained from a publicly available database,adding the set of domain names t :, obtaining, from a repository of aggregated data, a list of network addresses associated with each domain in the seed set, and mapping each domain in the seed set to network addresses in the list.

5. The method of claim 1, comprising: classifying the disruption according to a set of types, wherein the set of types allows the disruption to be classified as transient, permanent, or unplanned.

6. The method of claim 5, wherein the set of types are classified based on a time duration associated with downtime instances of the network.

7. The method of claim 1, comprising: performing a first scanning of the set of network addresses to identify host endpoints of the network that are internet-connected, wherein the first scanning comprises probing open ports of the network at one or more predetermined times; performing a second scanning of other endpoints of the network to determine an operational status of the other endpoints at one or more additional predetermined times; and identifying downtime instances of the network based on results of the first scanning and the second scanning.

8. The method of claim 7, wherein at least some of the other endpoints are identified using a publicly available resource or the ML algorithm.

9. The method of claim 7, comprising: collating the downtime instances into a dataset; and analyzing the dataset to determine effects to services or systems of the organization.

10. The method of claim 7, wherein the disruption is classified based on a statistical significance associated with the downtime instances.

11. The method of claim 1, comprising: reporting the disruption prior to the disruption being detected by the organization or other entity.

12. The method of claim 1, wherein the monitoring data is associated with at least one of an Application Programming Interface (API) or a Network Time Protocol (NTP).

13. A method for detecting abnormal activities on a computer network, comprising: obtaining a first set of network addresses, wherein each network address in the first set is connected to the computer network and associated with an organization; categorizing, based on predetermined rules, each network address in the first set according to a set of categories to obtain a second set of network addresses; determining subnetworks associated with at least some network addresses in the second set; assigning, based on the predetermined rules, each of the subnetworks to a category in the set of categories; determining a number of the subnetworks belonging to each category in the set of categories; obtaining a characterization of the computer network based on the number of the subnetworks belonging to each category; determining one or more correlated disruptions of the subnetworks; and determining an abnormal activity on the computer network based on the one or more correlated disruptions, wherein: at least some network addresses in the first set are associated with different organizations, at least some of the subnetworks are shared by the different organizations, and the characterization is based on the subnetworks that are shared by the different organizations.

14. The method of claim 13, wherein obtaining the first set comprises: determining a seed set of domains associated with the organization using a first dataset, determining a set of domain names associated with the organization using information obtained from a publicly available database, adding the set of domain names to the seed set, obtaining a list of network addresses associated with each domain in the seed set, and mapping each domain in the seed set to network addresses in the list.

15. The method of claim 13, wherein obtaining the first set comprises: using subnetting techniques to determine additional network addresses associated with the organization that are not discoverable by a domain name; and adding the additional network addresses to the first set.

16. The method of claim 13, wherein obtaining the second set comprises: determining an ownership or network connectivity property associated with each network address in the first set, wherein the categorizing of each network address in the first set is based on the ownership or network connectivity property.

17. The method of claim 16, wherein determining the subnetworks comprises determining a prefix size of the subnetworks based on the ownership or network connectivity property.

18. The method of claim 13, wherein the predetermined rules include rules based on at least one of pattern recognition, a prefix size of the subnetworks, or routing information.

19. The method of claim 13, comprising: obtaining a map describing a region of coverage provided by the computer network, wherein the abnormal activity is visualized on the map within the region of coverage.

20. The method of claim 19, wherein the map is provided to a device for display.

Citation Information

Patent Citations

  • Detecting connectivity disruptions by observing traffic flow patterns

    US11539728B1

  • Malicious port scan detection using source profiles

    US20200244684A1

  • Systems and methods for event assignment of dynamically changing islands

    US20240022483A1