Method, computing system and computer readable storage medium

By using artificial intelligence analysis and automated processes, clean snapshot files after a ransomware attack can be identified and recovered, solving the problems of inefficiency and errors in traditional ransomware recovery methods and achieving efficient and accurate data recovery.

CN121658285APending Publication Date: 2026-03-13COHESITY INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-25
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Traditional ransomware recovery methods rely on manual intervention, which is time-consuming and prone to errors, leading to data loss and operational interruption. Existing snapshot recovery methods cannot effectively identify damaged files.

Method used

Artificial intelligence technology is used to analyze file metadata, content and behavior patterns, automatically identify clean snapshot files, extract clean files from multiple snapshots, and restore them using a cleanroom environment, reducing human intervention and errors.

Benefits of technology

It improves the accuracy and efficiency of ransomware recovery, reduces data loss and recovery time, and protects recovered data from further damage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121658285A_ABST
    Figure CN121658285A_ABST
Patent Text Reader

Abstract

The invention relates to a method, a computing system, and a computer readable storage medium. Techniques for recovery of impaired snapshots are described. An exemplary method includes identifying, by a data platform implemented by a computing system, a baseline snapshot from a plurality of snapshots of protected data, where the baseline snapshot includes one or more files, each file not presenting an indication of impairment; for each file in an abnormal snapshot of the plurality of snapshots, identifying, by the data platform, a clean version of the file from one or more intermediate snapshots between the abnormal snapshot of the plurality of snapshots and the baseline snapshot; and storing, by the data platform, a clean snapshot including a respective clean version of a respective file identified for the file in the exception snapshot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a data platform for computing systems. Background Technology

[0002] Ransomware and other malware attacks pose a significant threat to organizations by encrypting or compromising critical data. Traditional recovery methods are manual, time-consuming, error-prone, and require specialized expertise. This results in substantial data loss, operational disruption, and financial impact.

[0003] Snapshots are typically used for quickly restoring to a previous state or for creating consistent backups without disrupting system operation, and therefore can be used for malware recovery. A snapshot can be taken of any object or collection of objects stored on a computing system's storage and / or disk, and the snapshot can be saved as one or more files. Examples of snapshots include file system snapshots, which are point-in-time copies or representations of the entire file system or a specific subset thereof. Snapshots capture the state of files and directories at a specific point in time, providing a snapshot of how the file system data appears at that particular point. File system data may include file system objects (e.g., files, directories, etc.), metadata, or both. Snapshots can also be taken of application workloads such as virtual machines, groups of one or more containers, or bare-metal processes. For example, a virtual machine snapshot captures the state of a virtual machine at a specific point in time and typically involves saving the virtual machine's virtual disk, storage state, configuration data, and virtual machine snapshot metadata as multiple distinct files, usually with corresponding file types and formats.

[0004] The primary problem with traditional malware recovery methods, especially ransomware, is their inefficiency and risk. Traditional recovery processes rely heavily on human intervention, making them time-consuming and error-prone. Restoring an entire snapshot is a sluggish method that often includes compromised data. Summary of the Invention

[0005] This article describes a technique for generating clean snapshots that include files that do not present anomaly or damaged indicators. This may involve identifying clean files within a damaged snapshot. This technique may include using artificial intelligence (AI) to identify clean files. Snapshots that represent the state of an object relatively late in time will better represent the current state of the object being protected, as anticipated updates are captured in subsequent snapshots. Ransomware and other malware often incrementally infect files, i.e., by infecting various files over time rather than all files simultaneously, to avoid triggering alerts. Therefore, malware that is infecting files included in stored snapshots often also incrementally infects files from different snapshots. Thus, it is common for different snapshots to have different groups of infected files. However, current snapshot recovery methods require that the snapshot does not present anomaly or damaged indicators (IOCs) in any of the snapshot files in the snapshot file to be considered a candidate snapshot for recovery.

[0006] The data platform based on the described technology can analyze file metadata, content, and behavioral patterns to distinguish clean snapshot files from infected snapshot files, thus enabling a fine-grained snapshot recovery approach, rather than relying on isolated security features such as ransomware detection, data classification, or support from security platforms and various Data Security Posture Management (DSPM) providers. Instead of relying on a single snapshot, the data platform can examine several snapshots to increase the probability of finding clean files and such clean files in more recent snapshots. Identified clean files located in different snapshots can be used for recovery and can be recovered in a secure, isolated environment (i.e., a cleanroom) to prevent further contamination. This contributes to the integrity of the recovered data.

[0007] In some examples, data platforms can significantly reduce recovery time by automating the process compared to manual methods. Furthermore, the use of AI can improve the accuracy of identifying clean files and reduce data loss compared to existing manual methods and sluggish approaches that rely on identifying completely uninfected snapshots for recovery. In some examples, AI-based chatbots can provide user-friendly interactions and collect feedback for improvement.

[0008] The technology disclosed herein can provide one or more technical advantages for realizing one or more practical applications. As described above, automation can accelerate the recovery process and reduce business downtime. AI-driven and personalized identification of clean files in different snapshots can reduce data loss by applying preferences for clean objects identified in more recent snapshots. Automation can achieve efficiency savings by reducing the need for professional personnel. Recovery in a cleanroom environment can protect recovered data from further damage.

[0009] While the techniques described in this disclosure may be described in relation to the snapshot functionality of a data platform, similar techniques can be applied to backup or archiving functions or other data protection features provided by a data platform. In some examples, the techniques described herein can be used to provide security responses for applications or other workloads, including those related to or unrelated to snapshots, backups, or archiving.

[0010] In one example, this disclosure describes a method comprising: identifying anomalous snapshots from a plurality of snapshots of protected data by a data platform implemented via a computing system, wherein the anomalous snapshots include one or more damaged files, the one or more damaged files including an indication of damage; identifying a baseline snapshot by the data platform from the plurality of snapshots, wherein the baseline snapshot does not include any files presenting an indication of damage; for each damaged file in the anomalous snapshots, identifying a clean version of the damaged file by the data platform from the baseline snapshots or from one or more intermediate snapshots between the anomalous snapshot and the baseline snapshot, the clean version not including an indication of damage; and storing by the data platform a clean snapshot including a corresponding clean version of each corresponding damaged file in the anomalous snapshots.

[0011] In another example, this disclosure describes a computing system comprising: a processing component configured to: identify anomalous snapshots from a plurality of snapshots of protected data, wherein the anomalous snapshots include one or more damaged files, the one or more damaged files including an indication of damage; identify a baseline snapshot from the plurality of snapshots, wherein the baseline snapshots do not include any files presenting an indication of damage; for each damaged file in the anomalous snapshots, identify a clean version of the damaged file from the baseline snapshots or from one or more intermediate snapshots between the anomalous snapshot and the baseline snapshot, the clean version not including an indication of damage; and store a clean snapshot including a corresponding clean version of each corresponding damaged file in the anomalous snapshots.

[0012] In yet another example, this disclosure describes a non-transitory computer-readable medium including instructions that, when executed, cause processing circuitry of a computing system to: identify anomalous snapshots from a plurality of snapshots of protected data, wherein the anomalous snapshots include one or more damaged files, the one or more damaged files including an indication of damage; identify baseline snapshots from the plurality of snapshots, wherein the baseline snapshots do not include any files presenting an indication of damage; for each damaged file in the anomalous snapshots, identify a clean version of the damaged file from the baseline snapshots or from one or more intermediate snapshots between the anomalous snapshot and the baseline snapshot, the clean version not including an indication of damage; and store a clean snapshot including a corresponding clean version of each corresponding damaged file in the anomalous snapshots.

[0013] Details of one or more embodiments of the present invention are set forth in the accompanying drawings and detailed description below. Other features, objects, and advantages of the invention will become apparent from the description and drawings, and from the claims. Attached Figure Description

[0014] Figures 1A to 1B This is a block diagram illustrating an exemplary system configured to support malware recovery according to one or more aspects of the technology described in this disclosure.

[0015] Figure 2 This is a block diagram illustrating an exemplary system configured to support malware recovery according to the technology of this disclosure.

[0016] Figure 3 This is a block diagram illustrating an example of multiple snapshots that can be used to construct a clean snapshot according to the technology of this disclosure.

[0017] Figure 4 This is a flowchart illustrating exemplary operations of a data protection manager according to the technology of this disclosure when performing various aspects of the construction of a clean snapshot.

[0018] Figure 5 This is a use case diagram illustrating a cleanroom configuration using an AI interface according to the technology disclosed herein.

[0019] Figures 6A to 6C This is a flowchart illustrating an exemplary technique for continuously improving a VM criticality machine learning model through a feedback loop, according to the technology disclosed herein.

[0020] Figure 7 This is a flowchart illustrating a malware recovery operation mode according to the technology disclosed herein.

[0021] Throughout the text and accompanying figures, similar figure labels indicate similar elements. Detailed Implementation

[0022] Figures 1A to 1B This is a block diagram illustrating one or more examples of an exemplary system configured to support malware recovery according to the techniques described in this disclosure. Figure 1A In the example, system 100 includes application system 102. Application system 102 represents a collection of hardware devices, software components, and / or data storage that can be used to implement one or more applications or services provided via network 113 to one or more mobile devices 108 and one or more client devices 109. Application system 102 may include one or more physical or virtual computing devices that execute workloads 174 for the applications or services. Workloads 174 may include one or more virtual machines, groups of one or more containers (e.g., Kubernetes pods), bare-metal processes, and / or other types of workloads.

[0023] exist Figure 1A In the example, application system 102 includes application servers 170A-170M (collectively referred to as "application server 170") connected via a network to database server 172 that implements the database. Other examples of application system 102 may include one or more load balancers, web servers, network devices (such as switches or gateways), or other devices for implementing one or more applications or services and delivering one or more applications or services to mobile device 108 and client device 109. Application system 102 may include one or more file servers. One or more file servers may implement the main file system of application system 102. (In such instances, file system 153 may be a secondary file system that provides backup, archiving, and / or other services to the main file system. References to file systems herein may include a main file system or a secondary file system, such as the main file system of application system 102 or file system 153 operating as a main file system or a secondary file system.)

[0024] Application system 102 may be deployed on-premises and / or located in one or more data centers, each of which is part of a public cloud, private cloud, or hybrid cloud. These applications or services may be distributed applications. These applications or services may support enterprise software, financial software, office or other productivity software, data analytics software, end-user relationship management, web services, educational software, database software, multimedia software, information technology, healthcare software, or other types of applications or services. These applications or services may be provided as a service (-aaS) as Software as a Service (SaaS), Platform as a Service (PaaS), Infrastructure as a Service (IaaS), Data Storage as a Service (dSaaS), or other types of services.

[0025] In some examples, application system 102 may represent an enterprise system that includes one or more workstations in the form of desktop computers, laptops, mobile devices, enterprise servers, network devices, and other hardware to support enterprise applications. Enterprise applications may include enterprise software, financial software, office or other productivity software, data analytics software, end-user relationship management, web services, educational software, database software, multimedia software, information technology, healthcare software, or other types of applications. Enterprise applications may be delivered as services from external cloud service providers or other providers, executed natively on application system 102, or both.

[0026] exist Figure 1A In this example, system 100 includes a data platform 150 that provides a file system 153 and archiving functionality to application system 102 using storage system 105 and a separate storage system 115. Data platform 150 implements a distributed file system 153 and storage architecture to facilitate application system 102's access to file system data and facilitate data transfer between storage system 105 and application system 102 via network 111. In the case of a distributed file system, data platform 150 enables devices of application system 102 to access file system data via network 111 using communication protocols as if such file system data were stored locally (e.g., on the hard drive of application system 102's device). Exemplary communication protocols for accessing files and objects include Server Message Blocks (SMB), Network File System (NFS), or Amazon® Simple Storage Service (S3®). File system 153 can be the primary or secondary file system of application system 102.

[0027] File system manager 152 represents a collection of hardware devices and software components that implement file system 153 for data platform 150. Examples of file system functions provided by file system manager 152 include storage space management, which includes deduplication, file naming, directory management, metadata management, partitioning, and access control. File system manager 152 executes communication protocols to facilitate application system 102's access to files and objects stored on storage system 105 via network 111.

[0028] Data platform 150 includes a storage system 105 having one or more storage devices 180A–180N (collectively, “storage devices 180”). Storage device 180 may represent one or more physical or virtual computing and / or storage devices that include storage media or otherwise have access to storage media. Such storage media may include one or more of the forms of flash drives, solid-state drives (SSDs), hard disk drives (HDDs), electrically programmable memory (EPROM) or electrically erasable programmable memory (EEPROM) and / or other types of storage media used to support data platform 150. Different storage devices in storage device 180 may have combinations of different types of storage media. Each of storage devices 180 may include system memory. Each of storage devices 180 may be a storage server, a network attached storage (NAS) device, or disk storage that may represent a computer device. Storage system 105 may be a redundant array of independent disks (RAID) system. In some examples, one or more storage devices in storage device 180 are both computing devices and storage devices, executing software for data platform 150, such as file system manager 152 and data protection manager 154 in the example of system 100. In some examples, a separate computing device (not shown) executes software for data platform 150, such as file system manager 152 and data protection manager 154 in the example of system 100. Each storage device in storage device 180 may be considered and referred to as a "storage node" or simply a "node". Storage device 180 may represent a virtual machine running on a supported hypervisor, a cloud virtual machine, a physical rack server, or a computing model installed in a convergence platform.

[0029] In various examples, data platform 150 can run on a physical system, virtually, or cloud-natively. For example, data platform 150 can be deployed as a physical cluster, a virtual cluster, or a cloud-based cluster running in a private cloud, hybrid private / public cloud, or public cloud deployed by a cloud service provider. In some examples of system 100, multiple instances of data platform 150 can be deployed, and file system 153 can be replicated among the various instances. In some cases, data platform 150 can be a computing cluster representing a single management domain. The number of scalable storage devices 180 can meet performance requirements.

[0030] Data platform 150 can implement and provide multiple storage domains to one or more tenants, or isolate workloads 174 that require different data policies. A storage domain can be a data policy domain, which determines the policies for deduplication, compression, encryption, tiering, and other operations performed on objects stored using the storage domain. In this way, data platform 150 provides users with the flexibility to select a global data policy or a workload-specific data policy. Data platform 150 may support partitioning.

[0031] A view can be a protocol export residing within a storage domain. A view can inherit the data policy of its storage domain, but additional data policies can be specified for the view. Views can be exported via SMB, NFS, S3, and / or another communication protocol. Policies determining the data processing and storage performed by the data platform 150 can be assigned at the view level. Protection policies can specify backup frequency and retention policies, which may include data lockout periods. Snapshots 142 or archives created according to the protection policy inherit the data lockout and retention periods specified by the protection policy.

[0032] Each of Network 113 and Network 111 may be the Internet, or may include or represent any public or private communications network or other network. For example, Network 113 may be cellular, Wi-Fi®, ZigBee®, Bluetooth®, Near Field Communication (NFC), satellite, enterprise, service provider, and / or other types of networks that support the transfer of data between computing systems, servers, computing devices, and / or storage devices. One or more devices of such types may use any suitable communication technology to transmit and receive data, commands, control signals, and / or other information across Network 113 or Network 111. Each of Network 111 or Network 113 may include one or more network hubs, network switches, network routers, satellite dish, or any other network equipment. Such network devices or components may be operatively coupled to each other, thereby providing information exchange between computers, devices, or other components (e.g., between one or more client devices or systems and one or more computer / server / storage devices or systems). Figures 1A to 1B Each of the devices or systems shown may be operatively coupled to network 111 and / or network 113 using one or more network links. The links coupling such devices or systems to network 111 and / or network 113 may be Ethernet, Asynchronous Transfer Mode (ATM) or other types of network connections, and such connections may be wireless and / or wired connections. Figures 1A to 1B One or more of the devices or systems shown or otherwise located on network 111 and / or network 113 may be located locally and / or remotely relative to one or more other devices or systems shown.

[0033] Application system 102 can generate objects and other data using file system 152 provided by data platform 150. File system manager 152 can create, manage, and store the objects and other data in storage system 105. For this purpose, application system 102 may be alternatively referred to as the "source system," and file system 153 used by application system 102 may be alternatively referred to as the "source file system." Application system 102 may communicate directly with storage system 105 via network 111 to transfer objects for some purposes, and may also communicate indirectly with file system manager 152 via network 111 to obtain objects or metadata from storage system 105 for some purposes. File system manager 152 generates metadata and stores it in storage system 105. The collection of data stored in storage system 105 and used to implement file system 153 is referred to herein as file system data. File system data may include the aforementioned metadata and objects. Metadata may include file system objects, tables, trees, or other data structures; metadata generated to support deduplication; or metadata used to support snapshots. For example, such as... Figure 1A As shown in the example, storage system 105 can store metadata of file system 153 in a tree data structure. The stored objects may include files, databases, applications, workloads 174, system images, directory information, or other types of objects used by application system 102. Objects of different types and objects of the same type can be deduplicated relative to each other.

[0034] Data platform 150 includes a data protection manager 154 that provides one or more data protection functions for application system 102, such as backups or snapshots of file system data of file system 153, workload 174, operating system, database of database server 172, or databases of servers 170, 172. Hereinafter, this disclosure will refer to snapshot 142, but the technique can be applied to other of the aforementioned data protection functions. In the example of system 100, data protection manager 154 stores one or more snapshots 142 of application system 102 data to storage system 115 via network 111. Application system 102 data may be, for example, file system data stored on storage system 105 or data local to application system 102, workload 174, operating system, database of database server 172, or databases of servers 170, 172, or other operations, configurations, or other data related to application system 102.

[0035] Storage system 115 includes one or more storage devices 140A-140X (collectively referred to as "storage devices 140"). Storage device 140 may represent one or more physical or virtual computing and / or storage devices that include storage media or otherwise have access to storage media. Such storage media may include one or more of the forms of flash drives, solid-state drives (SSDs), hard disk drives (HDDs), optical discs, electrically programmable memory (EPROM), or electrically erasable programmable memory (EEPROM) and / or other types of storage media. Different storage devices of storage device 140 may have different mixtures of storage media. Each of storage devices 140 may include system memory. Each of storage devices 140 may be a storage server, a network attached storage (NAS) device, or disk storage that may represent a computing device. Storage system 115 may include a redundant array of independent disks (RAID) system. Storage system 115 may be capable of storing a much larger amount of data than storage system 105. Storage device 140 may also be configured for long-term storage of information, making it more suitable for archiving purposes.

[0036] In some examples, storage systems 105 and / or 115 may be storage systems deployed and managed by a cloud storage provider and referred to as “cloud storage systems.” Example cloud storage providers include, for example, Amazon Web Services (AWS™) of Amazon Inc., Azure® of Microsoft Inc., Dropbox™ of Dropbox Inc., Oracle Cloud™ of Oracle Inc., and Google Cloud Platform (GCP) of Google Inc. In some examples, storage system 115 is located alongside storage system 105 in a data center, on-premises, or in a private, public, or hybrid private / public cloud. Storage system 115 may be considered a “backup” or “secondary” storage system to the primary storage system 105. Storage system 115 may be referred to as an “external target” of snapshot 142. When deployed and managed by a cloud storage provider, storage system 115 may be referred to as “cloud storage.” Storage system 115 may include one or more interfaces for managing the transfer of data between storage systems 105 and 115 and / or between application system 102 and storage system 115. Data platform 150 supporting application system 102 relies on primary storage system 105 to support latency-sensitive applications. However, because storage system 105 is typically more difficult to scale or more expensive to scale, data platform 150 may use secondary storage system 115 to support use cases such as backup, snapshots, and archiving. Typically, each snapshot in snapshot 142 is a copy of application system 102 data created by data protection manager 154 to support rapid recovery, typically due to loss or corruption of some data in file system 153 or application system 102. File system archives (“archives”) are copies of file system 153 used to support long-term retention and review. A “copy” may include data needed to restore or view the state of file system 153 or other application system 102 data at the time of snapshot, backup, or archive.

[0037] Data Protection Manager 154 can back up or snapshot data at any time according to Policy 158, which specifies, for example, periodicity and timing (daily, weekly, etc.), the data to be backed up, retention period, storage location, access control, etc. An initial snapshot of the file system data can correspond to the state of the data at an initial time (the time the initial snapshot was created). According to Policy 158, the initial snapshot may include all or less of the data. For example, the initial backup / snapshot may include all objects of file system 153 or one or more selected objects of file system 153, some of workloads 174, all of workloads 174, or a portion thereof.

[0038] One or more subsequent incremental backups / snapshots of file system 153 may correspond to the corresponding state of data at the respective subsequent creation time, i.e., after the creation time corresponding to the initial backup / snapshot. Subsequent backups / snapshots may correspond to incremental backups of one or more objects, workloads 174, or other data related to application system 102 of file system 153. Some of the file system data of file system 153 stored on storage system 105 at the initial creation time of the backup / snapshot may also be stored on storage system 105 at subsequent backup / snapshot creation times. Subsequent incremental backups / snapshots may include data not previously stored on storage system 115. Data included in subsequent backups / snapshots may be deduplicated by data protection manager 154 against data included in one or more previous backups / snapshots (including the initial backup / snapshot) to reduce the amount of storage used. (References to "time" in this disclosure may refer to a date and / or time. Time may be associated with a date. For example, multiple backups / snapshots may occur at different times on the same date.)

[0039] In system 100, data protection manager 154 stores a snapshot of the data of application system 102 as snapshot 142 in storage system 115. (In some examples, data protection manager 154 also or alternatively stores a backup in storage system 115.) Data protection manager 154 can use any snapshot in snapshot 142 to subsequently restore the data of application system 102 to its state at the time the snapshot was created. As described above, data protection manager 154 can perform deduplication on data included in subsequent snapshots by comparing it to data included in one or more previous snapshots. For example, deduplication can be performed on a second object of file system 153 included in a second snapshot by comparing it to a first object of file system 153 included in an earlier first snapshot. Similarly, deduplication can be performed on a first workload 174 included in a second snapshot by comparing it to an earlier version of the first workload 174 included in an earlier first snapshot.

[0040] Data Protection Manager 154 can apply deduplication as part of the write process of a snapshot in snapshot 142, which writes (i.e., stores) data into storage system 115. Deduplication can be implemented in various ways. For example, the method can be fixed-length or variable-length, the block size used for the file system can be fixed or variable, and the deduplication domain can be applied globally or on a workload basis. Fixed-length deduplication involves defining data streams at fixed intervals. Variable-length deduplication involves defining data streams at variable intervals to improve the ability to match data, regardless of the file system block size method being used. This algorithm is more complex than fixed-length deduplication algorithms but is more efficient in most cases and typically produces less metadata. Variable-length deduplication can include variable-length, sliding-window deduplication. The length of any deduplication operation (whether fixed-length or variable-length) determines the size of the block being deduplicated.

[0041] End users or applications can access (e.g., read or write) data stored in storage system 115. Applications can execute on application system 102, data platform 150, or other systems. End users or applications may delete some of the data due to malicious attacks (e.g., viruses, ransomware, etc.), rogue or malicious administrators, and / or human error. User credentials may be compromised, therefore, data stored in storage system 115 may be vulnerable to ransomware attacks. To reduce the possibility of accidental or malicious data deletion or corruption, data locks with data lock periods can be applied to snapshot applications.

[0042] Data Protection Manager 154 may apply security services in the form of security services 165 (“Service 165”), which analyze file system 153, snapshot 142, application system 102, etc., to identify security vulnerabilities, including one or more of ransomware attacks, malware attacks, unauthorized data access, and the presence of malicious code. Services 165 may each be implemented using one or more microservices, workloads, or other executable instances. Services 165 may each provide dedicated security analysis capabilities that allow end users to perform keyword searches in an attempt to summarize security vulnerabilities identified by the corresponding services within Services 165. This forces end users to input keyword searches specific to the underlying services within Services 165 via interface 160 of Data Platform 150, which may require end users to have a specific understanding of each service within Services 165. As a result, end users may frequently have to perform multiple keyword searches and manually summarize security vulnerabilities, which can be frustrating and may lead end users to contact support personnel at Data Platform 150. End users may waste computing resources locally (e.g., at application system 102) and at data platform 150 trying to better understand security vulnerabilities affecting file system 153 or application system 102 (e.g., especially in the form of ransomware, which may lock files stored in snapshot 142 and prevent successful recovery of snapshot 142).

[0043] Workload 174 may include one or more virtual machines (VMs) that may be vulnerable to ransomware or other malicious attacks. According to various examples of the techniques described in this disclosure, data protection manager 154 can create a clean, functional VM by restoring files from multiple snapshots 142. Data protection manager 154 may use a secure environment, namely a cleanroom within cleanroom 167, to restore the files. This process can be designed to counteract ransomware or other malicious attacks that may have compromised one or more of the workload 147, file system 153, or snapshots 142.

[0044] Data Protection Manager 154 can initiate the process by identifying a snapshot from snapshot 142 that clearly shows no signs of damage. This clean baseline snapshot 141 can be used as the basis for reconstruction. In some examples, instead of restoring the entire snapshot, Data Protection Manager 154 can carefully select files from subsequent snapshots 142 to ensure they are clean and free from malicious alterations. This fine-grained approach reduces the risk of reintroducing threats.

[0045] Using the recovered clean files 162, the data protection manager 154 can construct a new, deployable virtual machine (VM). This VM should be identical to the original VM before the attack, minus the malicious components. By selectively restoring clean files 162 up to the current snapshot's clean state, the data protection manager 154 can significantly reduce the risk of recovering compromised data. For example, restoring only the necessary clean files 162 saves time and resources compared to restoring the entire snapshot 142. The data protection manager 154 can be designed to precisely reconstruct the VM's pre-attack state, such as the state of virtual disks, storage, and / or configuration data. By eliminating malicious components, the reconstructed VM may be less vulnerable to future attacks.

[0046] Snapshot 142 comprises numerous snapshots from many different points in time. As used herein, each snapshot S (which may be one or more of snapshots in snapshot 142) comprises one or more files containing data that stores the protected data. Snapshots may be fully or partially hydrated, where “hydration” is the process of transforming a partially hydrated snapshot (typically a thin, spatially valid representation of the data) into a fully usable dataset. This typically involves “filling” the snapshot with the actual data represented by it, making it a complete and fully accessible copy of the data at a given point in time.

[0047] For VM snapshots, each snapshot may include a separate file to save the virtual machine's virtual disk, storage state, configuration data, and virtual machine snapshot metadata. As used herein, the term "snapshot S[N]" for snapshot recovery operations refers to the most recent snapshot in which anomalies or irregularities have been identified. Snapshot S[N] can be the starting point for the recovery process of protected data. As used herein, the term "file F[N]" refers to the complete set of files contained within snapshot S[N]. File F[N] may include all files, including both clean files and affected files.

[0048] As used herein, the term “affected files M[N]” refers to a subset of files in F[N] that have been compromised by ransomware or other malicious activity. Compromised files are those that show signs of tampering, such as, but not limited to, modified content, changed file extensions, or the presence of compromised indicators (IOCs). As used herein, the term “clean files”162 refers to files within F[N] that have not been affected by malicious activity and remain intact. Clean files162 can be the primary target for recovery.

[0049] As used herein, the term "baseline snapshot" refers to a reference point, i.e., a snapshot taken when the file system 153 is known to be clean and free from any ransomware or malicious influence. Baseline snapshot 141 can be used as a benchmark for comparison and recovery. Clean snapshot 164 can represent an undamaged version of the recovered VM.

[0050] The aforementioned terminology establishes a framework for understanding the malware recovery process. By identifying the affected files and a clean baseline 141, the data protection manager 154 can focus on restoring clean files 162 to reconstruct a healthy system state to a clean snapshot 164. The disclosed technique can be more precise than restoring the entire snapshot 142 because it avoids the reintroduction of compromised data. The following description... Figures 3 to 4 An exemplary process is shown for identifying clean files in a set of snapshots in order to reconstruct clean snapshots for the purpose of recovering protected data.

[0051] Based on various examples of the techniques described in this disclosure, data platform 150 may support the execution of an AI “bot” that can rely on one or more machine learning (ML) models 163 (“ML model 163”) (e.g., decision tree, clustering, linear regression, Naive Bayes, k-nearest neighbor (kNN)). ML model 163 may be trained relative to various knowledge bases 166, including a general security knowledge base, a security environment knowledge base (e.g., access control data, encryption mechanisms, data loss prevention mechanisms), a VM criticality knowledge base, a data platform security-specific knowledge base (e.g., documents about security services provided by the data platform), an account-specific security knowledge base (e.g., logs and / or other data reflecting security vulnerabilities of a specific account associated with an end user), and other security proximity knowledge bases. The security knowledge base may include the identification of user or other actions and security vulnerabilities at the network, computing, or other electronic system, which, when used to train ML model 163, allows ML model 163 to streamline ransomware recovery. As described herein, the bot may be implemented in data platform 150 in the form of interface 226 and may be referred to as interface 226.

[0052] According to various examples of the techniques described in this disclosure, data platform 150 can support the execution of a robot that streamlines the ransomware recovery process by automating the configuration of cleanroom 167. Traditionally, users need to manually switch between different recovery environments or contexts. The disclosed robotics technology can eliminate this manual intervention, enhancing user experience and efficiency. Cleanroom 167 refers to one or more cleanroom environments. Cleanrooms are typically used as secure, isolated environments for analyzing, identifying, and recovering from malware infections without the risk of further malware spread or causing additional damage to the system. Cleanroom 167 can be implemented using a separate, disconnected network (air-gap system) or virtual machine not connected to data platform 150.

[0053] Data Protection Manager 154 may include an ML model 163 in the form of a Large Language Model (LLM), which may reference one or more knowledge bases 166 in various ways to obtain configuration data 169 (general, specific, and / or cleanroom-specific). This configuration data may form the basis for natural language messages, summaries, explanations, or descriptions used to configure cleanroom 167 for a user, as well as natural language responses to natural language queries input by the end user. LLM 163 may be executed by data platform 150 or on a third-party platform. In some examples, Data Protection Manager 154 may apply an LLM (which is an example of ML model 163 and may be referred to as "LLM 163") to interact with the user, such as prompting the user with information or confirming actions (e.g., configuration responses). The prompt is an "actionable prompt" because Data Protection Manager 154 may perform an action in response to confirmation from the user (e.g., user input approving the deletion of cleanroom 167). In some examples, Data Protection Manager 154 may receive user input (e.g., approval, permission, or confirmation) before performing any action to ensure that no action is taken without user approval.

[0054] In some examples, interface 226 can automatically analyze a specific ransomware attack, the affected data types, and the desired recovery target. Based on this information, interface 226 can automatically configure the environment of cleanroom 167.

[0055] Users can interact with interface 226 using natural language (e.g., speech-to-text, text chat messages, etc.) to input queries, commands, and other information. Interface 226 can use one or more ML models 163 to process these queries, commands, and other information to derive a configuration. Based on this configuration, data protection manager 154 can retrieve security and / or configuration data 169 from one or more of a general security knowledge base, a general configuration knowledge base, a data platform-specific knowledge base, an account-specific configuration knowledge base, or other security-proximity knowledge bases (shown as knowledge base 166). Data protection manager 154 can invoke LLM 163 to provide the exported configuration, monitored actions, security analysis output from service 165, security data retrieved from various knowledge bases 166, or various subsets thereof.

[0056] LLM 163 can formulate natural language responses based on such input. For example, LLM 163 can provide a user-friendly interface that allows users to monitor the recovery process and make adjustments as needed, without delving into complex configuration settings. LLM 163 may include one or more suggested actions (e.g., configuration settings) for the user to confirm or describe one or more actions that interface 226 has taken to set up the cleanroom to create cleanroom documents 162. LLM 163 can formulate natural language responses based on such input. Interface 226, then executed by data platform 150, can output natural language responses from LLM 163. In some examples, interface 160 may provide one or more APIs 161, and other systems may make API calls (e.g., requests) to interface 226 to allow users to interact with data platform 150 using natural language.

[0057] Data Protection Manager 154 may process snapshot 142, such as through a data security ML model of ML model 163 (which is an example of ML model 163 and may be referred to as "data security model 163"), to detect security vulnerabilities (or potential security vulnerabilities), compromises (or potential compromises), or both, in snapshot 142, workload 174, components or data of application system 102, data or file collections of file system 153. Data security model 163 can be trained to detect security vulnerabilities or compromises with respect to various knowledge bases 166, including general security knowledge bases, data platform security-specific knowledge bases (e.g., documents about security services provided by the data platform), account-specific security knowledge bases (e.g., logs and / or other data reflecting security vulnerabilities of specific accounts associated with end users), and other security-proximity knowledge bases. For example, in response to receiving an indication of a security vulnerability, data security model 163 may determine whether a snapshot in snapshot 142 has been compromised, as will be further described below.

[0058] In some examples, the Data Protection Manager 154 may employ a VM criticality ML model of ML model 163 (which is an example of ML model 163 and may be referred to as "VM Criterion Model 163"), which can be trained to provide accurate VM criticality assessments. Accurate VM criticality assessments can help determine the prioritization of recovery processes and resource allocation. Identifying critical VMs can help the Data Protection Manager 154 focus security efforts on high-value assets.

[0059] As mentioned above, complex user interfaces and convolutional business processes often lead to user frustration and reduced productivity. Users frequently rely on extensive documentation and tutorials to navigate systems effectively, which can be time-consuming and inefficient. By introducing a bot (interface 226), organizations can simplify user interactions and automate task setup. For example, interface 226 can provide users with a more natural and conversational way to interact with data platform 150.

[0060] Interface 226 can handle the complex task of switching between different recovery contexts, ensuring that the cleanroom 167 environment is isolated and secure. Interface 226 can optimize resource allocation within cleanroom 167 based on recovery needs, such as, but not limited to, storage, computing power, and network connectivity. Interface 226 can save time and reduce errors associated with manual configuration.

[0061] A simplified process makes ransomware recovery easier for users with varying levels of technical expertise. Automating the process reduces the risk of security breaches due to human error. Interface 226 ensures that resources are allocated efficiently for the recovery process.

[0062] For example, in the case where ransomware is detected on one of workloads 174, interface 226 can break down complex tasks into simpler steps during detection, making the process easier for the user. Interface 226 can provide real-time assistance, eliminating the need for users to constantly refer to manuals. A more streamlined and intuitive user experience leads to increased satisfaction and productivity. Interface 226 can simplify the complex task of configuring cleanroom 167. Interaction with interface 226 should be intuitive and easy to understand, even for users without technical expertise.

[0063] This dynamic, interactive process not only ensures that users can interact with interface 226 using simple natural language commands to create, modify, or delete cleanrooms 167, thereby enhancing safety measures and configurations without human intervention, but also represents a significant advancement in safety management because it can automatically set the necessary infrastructure, resources, and safety parameters based on user input. In some examples, interface 226 can implement chatbots or virtual assistants.

[0064] Figure 1B System 190 is Figure 1A A variant of system 100, in which data platform 150 stores snapshot 142 to a snapshot storage system 105 residing locally, or in other words, locally on data platform 150. In some examples of system 190, storage system 105 allows users or applications to create, modify, or delete clean snapshots 164 via file system manager 152. In system 190, Figure 1BStorage system 105 is a local storage system used by data protection manager 154 for initial storage and cumulative snapshot 142. Data protection manager 154 can store tree data at storage system 105, including nodes having references (e.g., pointers) to one or more clean files 162.

[0065] Figure 2 This is a block diagram illustrating an exemplary system configured to support malware recovery according to the technology of this disclosure. Figure 2 System 200 can be described as Figure 1A System 100 or Figure 1B This document describes an exemplary or alternative implementation of system 190 (where object 142 is written to local snapshot storage system 115). Figure 1A and Figure 1B Contextual description Figure 2 One or more aspects of.

[0066] exist Figure 2 In the example, system 200 includes network 111, data platform 150 implemented through computing system 202, and storage system 115. Figure 2 In this context, network 111, data platform 150, and storage system 115 can correspond to... Figure 1A The network 111, data platform 150, and storage system 115 are described. Although only one snapshot storage system 115 is depicted, the data platform 150 can use multiple instances of the snapshot storage system 115 to apply the techniques according to this disclosure. Different instances of the storage system 115 may be deployed by different cloud storage providers, the same cloud storage provider, by an enterprise, or by other entities.

[0067] The computing system 202 can be implemented as any suitable computing system, such as one or more server computers, workstations, mainframes, appliances, cloud computing systems, and / or other computing systems capable of performing the operations and / or functions described in one or more examples according to this disclosure. In some examples, the computing system 202 represents a cloud computing system, server farm, and / or server cluster (or a portion thereof) that provides services to other devices or systems. In other examples, the computing system 202 may represent or be implemented through one or more virtualized computing instances (e.g., virtual machines, containers) of a cloud computing system, server farm, data center, and / or server cluster.

[0068] exist Figure 2In the example, computing system 202 may include one or more communication units 215, one or more input devices 217, one or more output devices 218, and one or more storage devices of local storage system 105. Local storage system 105 may include interface module 226, file system manager 152, ML model 163 and policy 158, and data protection manager 154 and service 165. Local storage system 105 may also include knowledge base 166 and interface 160 and API 161. One or more of the devices, modules, storage areas or other components of computing system 202 may be interconnected to enable inter-component communication (physical, communicative and / or operational). In some examples, such connectivity may be provided via communication channels (e.g., communication channel 212), which may represent one or more of system buses, network connections, inter-process communication data structures, or any other method for conveying data.

[0069] The computing system 202 includes processing components. Figure 2 In the example, the processing component includes one or more processors 213, which are configured to implement functions associated with or related to the computing system 202. Figure 2 The functionality associated with one or more modules shown and described below and / or the execution of instructions associated with the above. One or more processors 213 may be part of a processing circuitry system, and / or may include a processing circuitry system that performs operations according to one or more examples of this disclosure. Examples of processors 213 include microprocessors, application processors, display controllers, auxiliary processors, one or more sensor hubs, and any other hardware configured to function as a processor, processing unit, or processing device. The computing system 202 may use one or more processors 213 to perform operations according to one or more examples of this disclosure using software, hardware, firmware, or a mixture of hardware, software, and firmware residing in and / or executing at the computing system 202.

[0070] One or more communication units 215 of computing system 202 can communicate with devices outside computing system 202 by transmitting and / or receiving data, and in some respects can operate as both an input device and an output device. In some examples, communication unit 215 can communicate with other devices via a network. In other examples, communication unit 215 can transmit and / or receive radio signals on a radio network (such as a cellular radio network). In other examples, communication unit 215 of computing system 202 can transmit and / or receive satellite signals on a satellite network. Examples of communication unit 215 include network interface cards (e.g., Ethernet cards), optical transceivers, radio frequency transceivers, GPS receivers, or any other type of device capable of transmitting and / or receiving information. Other examples of communication unit 215 may include devices capable of communicating via Bluetooth®, GPS, NFC, ZigBee® and cellular networks (e.g., 3G, 4G, 5G) and Wi-Fi® radios in mobile devices, as well as Universal Serial Bus (USB) controllers. Such communication may comply with, implement, or follow appropriate protocols, including Transmission Control Protocol / Internet Protocol (TCP / IP), Ethernet, Bluetooth®, NFC, or other technologies or protocols.

[0071] One or more input devices 217 may represent any input device of computing system 202 not otherwise described separately herein. Input devices 217 may generate, receive, and / or process input. For example, one or more input devices 217 may generate or receive input from a network, a user input device, or any other type of device for detecting input from a person or machine.

[0072] One or more output devices 218 may represent any output device of computing system 202 not otherwise described separately herein. Output devices 218 may generate, present, and / or process output. For example, one or more output devices 218 may generate, present, and / or process output in any form. Output devices 218 may include one or more USB interfaces, video and / or audio output interfaces, or any other type of device capable of generating tactile, audio, visual, video, electrical, or other outputs. Some devices may function as both input and output devices. For example, a communication device may transmit data to and receive data from other systems or devices via a network.

[0073] One or more storage devices of the local storage system 105 within the computing system 202 may store information for processing during operation of the computing system 202, such as random access memory (RAM), flash memory, solid-state drive (SSD), hard disk drive (HDD), etc. The storage devices may store program instructions and / or data associated with one or more modules among the modules described according to one or more examples of this disclosure. One or more processors 213 and one or more storage devices may provide an operating environment or platform for such modules, which may be implemented as software, but in some examples may include any combination of hardware, firmware, and software. One or more processors 213 may execute instructions, and one or more storage devices of the storage system 105 may store instructions and / or data of one or more modules. The combination of processors 213 and the local storage system 105 may retrieve, store, and / or execute instructions and / or data of one or more applications, modules, or software. The processor 213 and / or the storage devices of the local storage system 105 may also be operatively coupled to one or more other software and / or hardware components, including but not limited to the computing system 202 and / or one or more devices or systems shown as connected to the computing system 202.

[0074] File system manager 152 can perform functions related to providing file system 153, as described above. Figure 1A The file system manager 152 can generate and manage file system metadata for constructing file system data for file system 153, and store the file system metadata and file system data in local storage system 105. The file system metadata may include one or more trees that describe objects within file system 153 and the file system 153 hierarchy, and can be used to write to or retrieve objects within file system 153. The file system manager 152 can interact with and / or cooperate with one or more modules of computing system 202, including interface module 226 and data protection manager 154.

[0075] Data Protection Manager 154 can perform functions related to ransomware recovery by implementing a recovery process, as described above. Figure 1A The above describes the operations related to ML model 163 (such as VM criticality model 163 and LLM 163 as described above), interface 226, interface 160, and service 165. Data protection manager 154 enables storage system 105 to store, retrieve, and update knowledge base 166. For example, data protection manager 154 enables storage system 105 to store, retrieve, and update knowledge base 166 during the training and inference of ML model 163.

[0076] Data Protection Manager 154 can generate one or more snapshots 142 and store file system data as tree data within snapshot 142 in snapshot storage system 115. Data Protection Manager 154 can generate and manage tree data for creating, viewing, retrieving, or restoring any snapshot in snapshot 142. Data Protection Manager 154 can generate and manage file system metadata for creating, viewing, retrieving, or restoring objects such as VMs for any snapshot in snapshot 142. In some examples, VMs can be restored in a secure and isolated environment (cleanroom 167).

[0077] Local storage system 105 can store one or more cleanrooms 167. Cleanroom 167 can include multiple verified clean files 164, which can be used to generate a clean snapshot 164 by combining the verified clean files 167. In some examples, data protection manager 154 can automatically analyze a specific ransomware attack, the affected data types, and the desired recovery target. Based on this information and based on user interaction, data protection manager 154 can automatically configure the environment of cleanroom 167. If multiple cleanrooms exist, data protection manager 154 can present a list of options for the user to choose from. If no cleanroom is configured, data protection manager 154 can prompt the user to register a new cleanroom before initiating recovery.

[0078] Local storage system 105 may include configuration data 169, which may describe the requirements for the corresponding cleanroom 167 on storage system 115, as well as other metadata about snapshots, such as checksums, encrypted data, compressed data, etc. Figure 2 In this context, data protection manager 154 causes file system metadata to be stored in local storage system 105. In some examples, data protection manager 154 causes some or all of the file system metadata to be stored in snapshot storage system 115. Data protection manager 154, optionally or in conjunction with file system manager 152, can use the file system metadata to restore any snapshot in snapshot 142 to a clean snapshot 164 in cleanroom 167 implemented by data platform 150, which can then be presented to other systems by file system manager 152.

[0079] Interface module 226 can execute an interface through which other systems or devices can determine the operation of file system manager 152 or data protection manager 154. Another system or device can communicate via the interface module 226 to specify one or more policies 158.

[0080] This can be achieved by modifying system 200. Figure 1B Example of system 190. In the modified system 200, snapshot 142 can be stored in local snapshot storage system 115.

[0081] The interface module 240 of the snapshot storage system 115 can execute an interface through which users can create, modify, or delete one or more cleanrooms 167 for restoring clean snapshots 164. The interface module 240 can execute and present an API. The interface presented by the interface module 240 can be gRPC, HTTP, RESTful, command line, graphical user interface, Web interface, or other interfaces.

[0082] Figure 3 This is a block diagram illustrating examples of multiple snapshots that can be used to construct clean snapshots according to various techniques of this disclosure. (The following is in...) Figures 1A to 1B In the context of Figure 3 All aspects. Figure 3 Multiple read-write (RW) snapshots were depicted, some of which contained malicious files and were therefore considered infected. For example... Figure 3 As shown in the example, Data Protection Manager 154 can identify in Figures 1A to 1B and Figure 2 The last (i.e., most recent) clean snapshot (S[NK]) 302 is represented as baseline snapshot 141. As described above, this step may involve a look-back analysis of the snapshot from anomalous snapshot 310 to pinpoint the clean starting point (baseline snapshot 302). Thus, in some examples, anomalous snapshot S[N] 310 may represent a snapshot suspected of being compromised. Anomalous snapshot S[N] 310 may be the most recent snapshot suspected of being compromised. Data protection manager 154 can continue the process by examining previous snapshots S[N-1] 308. In some examples, data protection manager 154 may continue the examination sequentially with older snapshots 306, 304 until baseline snapshot 302 is found to be free of anomalies, impairment indicators (IOCs), or any other signs of compromise. This baseline snapshot 302 may be designated as S[NK], where K represents the number of snapshots between anomalous snapshot 310 and baseline snapshot 302.

[0083] exist Figure 3 In the example, once baseline snapshot 302 is determined, data protection manager 154 can shift its focus to reconstructing a clean version of the protected data. In some examples, data protection manager 154 can check if each file 320 present in anomalous snapshot 310 also exists in baseline snapshot 302. If a file exists in both snapshots, data protection manager 154 can further examine all intermediate snapshots 304-308 to ensure that the file remains unchanged and free of anomalies or IOCs. Files appearing in snapshots after baseline snapshot 302 but before anomalous snapshot 310 can be carefully examined in each snapshot up to anomalous snapshot 310 to verify their integrity and legitimacy.

[0084] exist Figure 3 In the example, the combined steps performed by the data protection manager 154 can be designed to restore a snapshot of data, such as a VM, to a clean state. The data protection manager 154 can identify a point in time when the snapshot is known to be clean (undamaged) (baseline snapshot 302). The data protection manager 154 can verify the integrity of the file from that clean point to the present.

[0085] In some examples, utilizing a verified list of clean files from baseline snapshot 302 to anomalous snapshot 310, the final step performed by data protection manager 154 could be to create a clean snapshot excluding any corrupted files. For example, data protection manager 154 could combine all files that have passed the verification process from baseline snapshot 302 to anomalous snapshot 310 into a final clean snapshot 164. In operation, data protection manager 154 can assume these files are trustworthy and have not been maliciously tampered with. For example, when data protection manager 154 identifies any file 330 in anomalous snapshot 310 as affected or corrupted, data protection manager 154 can explicitly exclude these affected files from the final clean snapshot 164. In this way, data protection manager 154 can ensure that the reconstructed object contains no malicious elements.

[0086] Data protection manager 154 can preferentially use instances of files from more recent snapshots that have been determined to be clean. Therefore, a more recent version of the file can be used to reconstruct a clean snapshot 164.

[0087] Still referencing Figure 3 The following example illustrates the reconstruction of a clean snapshot 164 using files with file extensions. In this example, the anomalous snapshot 310 can be S[5], and the baseline snapshot 302 can be S[2]. Files can have the following extensions: .txt (text document), .doc (MICROSOFT WORD document), .pdf (Portable Document Format), .jpg (JPEG image), .png (Portable Web Graphics), .xlsx (MICROSOFT EXCEL spreadsheet), .rar (compressed archive).

[0088] In the following example, baseline snapshot 302 (S[2]) contains the following files: a.txt, b.doc, c.pdf, and d.jpg. All files in baseline snapshot 302 are determined by data protection manager 154 to be clean at that point.

[0089] Snapshot S[3] contains all files from baseline snapshot 302 plus e.png. Although file e.png is new, data protection manager 154 can consider the file clean because the file does not show any anomalies or IOCs. Snapshot S[4] contains all files from S[3] plus f.xlsx. In this example, file f.xlsx may be new, but data protection manager 154 can consider the file clean. Anomaly snapshot 310 S[5] contains all files from S[4] plus g.rar.

[0090] The file g.rar may be new, and Data Protection Manager 154 can identify it as corrupted. Data Protection Manager 154 can now determine that files c.pdf, e.png, and f.xlsx are corrupted, even if they were clean in previous snapshots. This could indicate a potential data breach or modification.

[0091] Data Protection Manager 154 can individually assess the integrity of each file 320 in the anomalous snapshot 310. In other words, each file 320 can be assessed to determine whether the file 320 is corrupted. Files that are indeed considered clean from the anomalous snapshot 310 itself or from previous snapshots 304-308 can be included in the final clean snapshot 164. In this example, the final clean snapshot 164 may include the following clean files {a.txt, b.doc, c.pdf, d.jpg, e.png, f.xlsx}. a.txt and b.doc are taken from snapshot S[5], g.rar is excluded from the clean snapshot, and c.pdf, d.jpg, e.png, and f.xlsx are taken from snapshot S[4] which is closer than S[3]. Files identified as corrupt, such as newly introduced files with malicious content or files that have been modified, as well as clean versions of files for which they do not exist, can be excluded, as in the example above with g.rar. The disclosed technique involves moving backward through snapshots 302-310 to verify file integrity, ensuring that the final clean snapshot 164 is free of contamination. By inspecting each file individually, the data protection manager 154 can reduce the risk of including corrupted data.

[0092] The disclosed techniques can be optimized through the automation and prioritization of critical documents. Backward analysis can enhance the confidence in the integrity of the final clean snapshot 164.

[0093] Figure 4 This is a flowchart illustrating exemplary operation of a data protection manager according to the technology of this disclosure when performing various examples of a clean snapshot construction. For example, such as Figure 4 As shown in the example, the data security model of ML model 163 can determine whether the current snapshot contains an anomaly (402).

[0094] If an anomaly is detected (decision box 402, is a branch), the data security model 163 may send information about the detected anomaly to the data protection manager 154. The data protection manager 154 may locate the most recent clean snapshot (404) (e.g., baseline snapshot 302). As described above, the data protection manager 154 may identify file differences between the clean snapshot and the anomalous snapshot (406). As used herein, the term "nominal snapshot" refers to a snapshot containing suspicious or corrupted files. The data protection manager 154 may remove files that only exist in the baseline snapshot 141 (408). In some examples, the data protection manager 154 may add clean files from the anomalous snapshot to the clean snapshot 164 in the cleanroom 167 (410). As used herein, the term "clean snapshot" refers to a snapshot without anomalies or malicious content. For example, as described above, the data protection manager 154 may inspect clean files in a previous snapshot (412). When the data protection manager 154 finds clean files (decision box 414, is a branch), the data protection manager 154 may add them to the clean snapshot 164 (410). As used in this article, the term "clean file" refers to a file that has been determined to be free of anomalies.

[0095] Data Protection Manager 154 can continue process 420 until all retrievable clean files are identified (decision box 416) and added to clean snapshot 164. As described above, clean snapshot 164 can be reconstructed in the cleanroom 167 environment.

[0096] Currently, configuring a cleanroom presents challenges for users due to its complex user interface, convolutional business processes, and extensive documentation. Navigating a complex system can be time-consuming and error-prone. Understanding the procedures and requirements can be overwhelming. Users typically need to refer to manuals or guides to complete the task.

[0097] According to the technology of the present invention, interface 226 can improve the process by automating configuration, simplifying interaction, eliminating context switching, and providing guidance. Interface 226 can essentially handle all technical aspects of setting up cleanroom 167. Thus, users can interact with interface 226 using natural language, providing instructions and preferences. Interface 226 can manage multiple tasks and configurations without user intervention. Interface 226 can provide suggestions and recommendations based on user needs. Tasks can be completed faster with less effort. The process can become more intuitive and user-friendly. Automation reduces human error. Users can focus on core tasks rather than configuration.

[0098] Figure 5 This is a use case diagram illustrating cleanroom configurations using AI interfaces according to various technologies disclosed herein.

[0099] As described above, the AI ​​robot can be implemented in the data platform 150 in the form of interface 226, and can be referred to as interface 226. When user 502 attempts to access information about the cleanroom or initiate cleanroom restoration, and the cleanroom is not currently registered, data protection manager 154 can first prompt 506 regarding cleanroom details. Data protection manager 154 can request 506 cleanroom configuration information from user 502. Cleanroom configuration information may include, but is not limited to: the address of data platform 150, the password for accessing data platform 150, the desired name of cleanroom 167, etc.

[0100] User 502 can provide the details requested in 508. Data Protection Manager 154 can use the provided credentials to make the necessary API calls 514 to service 165 to register cleanroom 167. If user 502 initially requested cleanroom details, Data Protection Manager 154 can provide newly registered cleanroom information 526. In some examples, if user 502 initiates a recovery request, Data Protection Manager 154 continues the recovery process using the registered cleanroom 167. The disclosed technology is designed to be simple and intuitive for users.

[0101] Data Protection Manager 154 may request only the necessary details for registering Cleanroom 167. Data Protection Manager 154 may interact with the environment of Interface 160 and Service 165 to complete the registration. In some examples, the actions of Data Protection Manager 154 may depend on an initial query 504 from user 502.

[0102] When user 502 interacts with data protection manager 154, multiple cleanrooms 167 may have been registered. Data protection manager 154 should be able to handle these situations effectively. If only one cleanroom 167 is registered, data protection manager 154 can directly provide the details of cleanroom 167.

[0103] If multiple cleanrooms 167 exist, the data protection manager 154 can present a list of available cleanrooms 167 and allow the user 502 to select a cleanroom for details. If only one cleanroom 167 is registered, the recovery process can begin automatically using that cleanroom 167. If no cleanroom 167 is configured, the data protection manager 154 can prompt the user to register a new cleanroom before initiating recovery.

[0104] When multiple options are available, user 502 may have the option to select a specific cleanroom 167. Data Protection Manager 154 may provide user 502 with clear instructions and options. The option to register a new cleanroom 167 is always available to user 502.

[0105] In some examples, user 502 can request the deletion of a specific cleanroom 167 by providing the name of that specific cleanroom to data protection manager 154.

[0106] Data Protection Manager 154 can confirm user 502's request to delete the specified cleanroom 167. This step may be important to prevent accidental deletion. Data Protection Manager 154 can initiate a deletion process, which may involve removing cleanroom 167 from the records of data platform 150 and potentially deleting associated data (depending on the specific implementation). Data Protection Manager 154 can notify user 502 that cleanroom 167 has been successfully deleted.

[0107] See again Figure 5 This flowchart outlines the disclosed technologies for registering a new cleanroom. Interface 226 can be the interface through which user 502 interacts with data protection manager 154. LLM (Large Language Model) 163 can handle user interactions, process requests, and generate responses. Interface 160 can be responsible for cleanroom management. Cleanroom service 528 can be a specific service or module within service 165 used to handle the operation of cleanroom 167.

[0108] The disclosed technology for registering a new cleanroom 167 using interface 226 and a language model can begin with a user request 504 to register a new cleanroom 167. By sending request 504, user 502 can express their desire to register a new cleanroom 167 through interface 226. LLM 163 can analyze query 504 and select an appropriate workflow 505 to register the cleanroom 167. In some examples, LLM 163 can prompt user 502 506 to provide registration details, such as the cleanroom name, hostname, and password.

[0109] The collected user input can be processed and prepared for the next step. LLM 163 can determine, based on the provided details, which APIs might be necessary for registering cleanroom 167. In some examples, LLM 163 can handle user interaction, workflow selection, and API determination. Interface 226 can process the response 510 from LLM 163 containing the selected workflow and can determine necessary actions, including but not limited to identifying the required API calls and associating user input with API parameters. Interface 226 can pass the extracted information to interface 160 via API call 514. Interface 160 can handle API interactions. Interface 160 can make the necessary API calls to the relevant service (e.g., cleanroom service 528) to create cleanroom 167. Cleanroom service 528 can process the registration request, create a new cleanroom record, and return a response 516. Interface 160 can wait for the API response 516, which may include the status of the registration process. Figure 5As shown, API response 516 can be fed back to LLM 163 via interface 226. LLM 163 can receive a response for parsing from interface 226, can parse the received response 520, and can provide message 526 to user 502, notifying them of successful or failed registration. Therefore, in some examples, LLM 163 can handle user interaction, determine the required action, and generate a final response 526 to user 502.

[0110] Interface 160 manages the technical aspects of interaction with service 165 and / or external systems via API. While LLM 163 can focus on natural language understanding and user interaction, Interface 160 handles technical execution. This architecture allows for easier integration of different APIs and systems. Interface 160 can handle multiple API calls simultaneously, thus improving performance.

[0111] Based on various technologies, Data Protection Manager 154 can be configured to determine the criticality of a virtual machine (VM) based on collected data points, such as, but not limited to, file extensions, backup frequency, security labels, and data read / write patterns. During data collection, Data Protection Manager 154 can collect relevant data points from the VM. Data Protection Manager 154 can derive insights from the collected data to assess VM criticality. Data Protection Manager 154 can create machine learning models to improve predictive accuracy. Data Protection Manager 154 can incorporate user feedback to improve the model. Data Protection Manager 154 can collect data points from the VM, including but not limited to file extensions, backup frequency, data read / write volumes, and security labels. Data Protection Manager 154 can analyze the collected data to identify patterns and correlations. Data Protection Manager 154 can use simple rules or heuristics to determine the initial VM criticality level. For example, frequent backups, high data activity, and sensitive data may indicate high criticality. Conversely, infrequent backups, low data activity, and no sensitive data may indicate low criticality. In other words, if VM backups are performed frequently and large amounts of data are read and written, or if a security label indicates that the VM contains highly sensitive data, the data protection manager 154 can infer that the VM is highly critical due to its usage patterns and the nature of its file content. The inference provided by the data protection manager 154 can highlight frequent backups and highly sensitive data. On the other hand, if VM backups are infrequent and a security label indicates the absence of sensitive data, the data protection manager 154 can infer that the VM has low usage and contains less important documents. The insights provided by the data protection manager 154 (e.g., LLM 163) can highlight low criticality, detailing the lack of sensitive data. According to the technology of the present invention, the data protection manager 154 can use supervised learning techniques to create a machine learning model. The data protection manager 154 can train the model on collected data using VM criticality as the target variable. The data protection manager 154 can employ a feedback mechanism to allow the model to learn from user corrections and improve accuracy over time. The data protection manager 154 can incorporate end-user feedback to improve the model and adapt to different usage patterns. Data Protection Manager 154 can use feedback to identify areas where the model is inaccurate and make necessary adjustments. Accurate VM criticality assessment helps prioritize backup, disaster recovery planning, and resource allocation. Identifying critical VMs helps Data Protection Manager 154 focus security efforts on high-value assets. By understanding VM importance, organizations can optimize resource utilization and improve efficiency.

[0112] Using various technologies, the Data Protection Manager 154 can collect end-user feedback on the VM criticality level (low, medium, high). The Data Protection Manager 154 can store this feedback in a knowledge base 166 for analysis and model improvement. The Data Protection Manager 154 can utilize existing data points (file extensions, backup frequency, data read / write) and predicted criticality levels generated by the VM criticality model 163. This combined dataset can be used as training data for the improved model. The Data Protection Manager 154 can use the existing supervised VM criticality model 163 to predict the criticality of the VM or data source based on the collected data points. The Data Protection Manager 154 can pass VM details, predicted criticality, and raw data points to the LLM 163 to generate insights. The LLM 163 can analyze the provided information and generate human-readable insights into the VM's criticality. The Data Protection Manager 154 can present the predicted criticality, generated insights, and raw data points to user 502. The Data Protection Manager 154 can use user 502's feedback to correct the predictions of the VM criticality model 163 and improve the model's accuracy over time. Incorporating customer feedback enhances the ability of the VM Criterion Model 163 to accurately assess VM criticality. LLM 163 provides valuable insights into the reasoning behind the VM Criterion Model 163, increasing transparency. Feedback loops ensure that the VM Criterion Model 163 can adapt to changing conditions and end-user requirements.

[0113] Figures 6A to 6C This is a flowchart illustrating exemplary techniques for continuously improving a VM criticality machine learning model through feedback loops, according to various techniques of this disclosure. A knowledge base 166 (e.g., a VM criticality knowledge base) may store the current dataset and the updated dataset. Machine learning service 602 may be a service provided by service 165, responsible for training, deploying, and using the VM criticality model 163 for prediction. Client 502 may represent a user or application interacting with data platform 150. In some examples, client 502 may represent data protection manager 154. In some examples, during or before the data recovery described above (e.g., ransomware recovery), data protection manager 154 may send data to machine learning service 602 for prediction. For example, as... Figure 6AAs shown, client 502 can provide feedback 604 regarding the accuracy of machine learning predictions. Client 502 can send the prediction feedback data 604 to machine learning service 602. Machine learning service 602 can update dataset 606 with new information. Machine learning service 602 can use the updated dataset to periodically retrain VM keyability model 163. Machine learning service 602 can save the new dataset to knowledge base 166. Machine learning service 602 can use the trained VM keyability model 163 to make predictions on new data. Data protection manager 154 can continuously improve VM keyability model 163 by combining new data and feedback. Knowledge base 166 can efficiently store and manage datasets. In some examples, machine learning service 602 can handle model training, deployment, and prediction.

[0114] During the data collection phase, the data protection manager 154 can collect relevant data points about the VM, such as, but not limited to, VM snapshot metadata (changed files, written bytes, etc.), file extensions, copy frequency, data read / write patterns, and existing labels. The VM criticality model 163 can be used as a baseline for assessing VM importance. In some examples, the data protection manager 154 can allow user 502 to provide feedback on the predictions of the VM criticality model 163 through interface 226. The data protection manager 154 can update the training and validation datasets with new feedback data via machine learning service 602, such as... Figure 6A As shown.

[0115] like Figure 6B As shown, the data protection manager 154 can also employ machine learning service 602 to periodically retrain the VM criticality model 163 with updated datasets. The data can be used as the basis for the VM criticality model 163 and can be continuously enriched with user feedback. The VM criticality model 163 can predict VM criticality, such as... Figure 6C As shown, and can be improved through feedback loops. Interface 226 allows users to provide feedback on the predictions of the VM criticality model 163. Feedback loops can ensure continuous improvement of the VM criticality model 163 by incorporating user input.

[0116] Now for reference Figure 6CIn some examples, the data protection manager 154 can collect relevant VM snapshot metadata (changed files, written bytes, etc.) and send the collected data 610 to a machine learning (ML) service 602 for analysis. In some examples, the ML service 602 can obtain the latest trained VM keyness model 163 from a knowledge base 166. The ML service 602 can apply the VM keyness model 163 to the provided metadata to generate keyness predictions. The prediction results can be sent back 614 to the client 502 for further action. The data protection manager 154 can continuously collect user feedback on the predictions of the VM keyness model 163. This feedback can be used to update the training dataset and retrain the VM keyness model 163. This iterative process ensures that the VM keyness model 163 adapts to changing data patterns and can improve accuracy over time. In some examples, the data protection manager 154 can employ LLM 163 to provide detailed insights based on the predicted keyness levels. LLM 163 analyzes the VM snapshot metadata 610 and keyness ratings to generate human-readable interpretations. User 502 can receive the predicted criticality level and generated insights. Data Protection Manager 154 can collect VM snapshot metadata 610 and receive prediction results 610 and insights. ML service 602 can manage model training, deployment, and prediction generation. In some examples, knowledge base 166 can store VM criticality model 163, training data, and feedback data. LLM 163 can provide valuable insights. In some examples, Data Protection Manager 154 can utilize actionable information provided by ML service 602 to manage data recovery processes (e.g., ransomware recovery).

[0117] Figure 7 This is a flowchart illustrating a malware recovery operation mode according to the technology disclosed herein. Figures 1A to 1B and Figures 2 to 3 Contextual description Figure 7 Some aspects. Data platform 150 (such as through data protection manager 154) can identify a baseline snapshot (702) from multiple snapshots of the protected data. The baseline snapshot may include one or more files, each of which does not present an indication of damage. As described above in conjunction with... Figure 3The step described herein may involve a backward analysis of the snapshot from the anomalous snapshot 310 to pinpoint a clean starting point (baseline snapshot 302). For each damaged file in the anomalous snapshot among multiple snapshots, the data protection manager 154 may identify a clean version of the damaged file from the baseline snapshot or one or more intermediate snapshots between the anomalous snapshot and the baseline snapshot (704). As used herein, the term "nominal snapshot" refers to a snapshot containing a suspected or damaged file. Files from the anomalous snapshot 310 itself or from previous snapshots 304-308 that are indeed considered clean may be included in the final clean snapshot 164. The data protection manager 154 may store a clean snapshot that includes a corresponding clean version of each of the corresponding damaged files in the anomalous snapshot (706). The disclosed technique can be more accurate than restoring the entire snapshot 142 because it avoids the reintroduction of damaged data.

[0118] While the techniques described in this disclosure are primarily concerned with backup or snapshot functions performed by the data protection manager of a data platform, similar techniques may be additionally or alternatively applied to archiving, copying, or cloning functions performed by the data platform. In such cases, snapshot 142 is an archive, a copy, or a clone, respectively.

[0119] For the processes, devices, and other examples or illustrations described herein included in any flowchart, certain operations, actions, steps, or events included in any technology described herein may be performed in a different order, and may be added, combined, or omitted entirely (e.g., not all described actions or events are necessary for the practice of the technology). Furthermore, in some examples, operations, actions, steps, or events may be performed concurrently, for example, through multithreading, interrupt handling, or multiple processors, rather than sequentially. Additionally, some operations, actions, steps, or events may be performed automatically even if not explicitly identified as automatically executed. Moreover, some operations, actions, steps, or events described as automatically executed may alternatively not be automatically executed; rather, in some examples, such operations, actions, steps, or events may be performed in response to input or another event.

[0120] The detailed descriptions set forth herein in conjunction with the accompanying drawings are intended to describe various configurations and are not intended to represent only configurations in which the concepts described herein can be practiced. Specific details are included for the purpose of providing a comprehensive understanding of the various concepts. However, it will be apparent to those skilled in the art that these concepts can be practiced without these specific details. In some instances, well-known structures and components are shown in block diagram form to avoid obscuring such concepts.

[0121] According to one or more aspects of this disclosure, the term "or" may be interpreted as "and / or" unless otherwise specified in the context. Additionally, while phrases such as "one or more" or "at least one" may be used in some instances, they may not be used in others; those instances where such language is not used may be interpreted as having this implied meaning unless otherwise specified in the context.

[0122] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored as one or more instructions or code on and / or transmitted via a computer-readable medium and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium (such as a data storage medium) or a communication medium (including any medium that facilitates the transfer of a computer program (e.g., according to a communication protocol) from one place to another). In this way, a computer-readable medium may generally correspond to (1) a tangible computer-readable storage medium that is non-transitory, or (2) a communication medium such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. Computer program products may include computer-readable media.

[0123] For example, and not as a limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store program code in the form of desired instructions or data structures and is accessible by a computer. Furthermore, any connection is appropriately referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology (such as infrared, radio, and microwave), then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology (such as infrared, radio, and microwave) are included in the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather refer to non-transient tangible storage media. As used, disks and optical discs include compact discs (CDs), laser discs, optical discs, digital universal discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically, while optical discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0124] The instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuit systems. Therefore, the terms "processor" or "processing circuit system" as used herein can each refer to any of the foregoing structures or any other structure suitable for implementing the described techniques. Additionally, in some examples, the described functionality may be provided within dedicated hardware and / or software modules. Moreover, these techniques can be implemented entirely within one or more circuit or logic elements.

[0125] As used herein, a processing component may include the processing circuitry system as described above. In some examples, the processing component may include at least one processor and at least one memory having computer code including a set of instructions that, when executed by the at least one processor, cause the at least one processor to perform any of the functions described herein. In some examples, the processing component may receive computer code including the set of instructions from at least one memory coupled to the processing component.

[0126] The techniques disclosed herein can be implemented in a wide variety of devices or apparatuses, including wireless handsets, mobile or non-mobile computing devices, wearable or non-wearable computing devices, integrated circuits (ICs) or IC sets (e.g., chip sets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed techniques, but they do not necessarily need to be implemented by different hardware units. Rather, as described above, various units may be combined within hardware units or provided by a collection of interoperable hardware units (including one or more processors as described above) combined with suitable software and / or firmware.

[0127] Additional details

[0128] This disclosure implements at least the following exemplary process. This process can be performed by a data platform implemented through a computing system. The process begins by identifying anomalous snapshots containing corrupted files. This may involve analyzing multiple snapshots to detect snapshots containing one or more indications of corruption in files. A baseline snapshot is identified. The baseline snapshot serves as a reference point and contains a version of the file that does not contain any indications of corruption.

[0129] For each damaged file in the anomalous snapshot, the process identifies a clean version from the baseline snapshot or from one or more intermediate snapshots. One or more intermediate snapshots exist between the anomalous snapshot and the baseline snapshot within a set of snapshots. The process then involves storing a clean snapshot with clean versions of the corresponding damaged files. This clean snapshot includes the corresponding clean versions of the files identified as damaged in the anomalous snapshot, effectively creating a recovered version of the data without any damaged elements.

[0130] This process allows for the systematic identification and recovery of damaged files within a snapshot, providing a mechanism to maintain data integrity and security in a protected data environment. It enables rapid identification and recovery of clean file versions without rolling back the entire system to a baseline snapshot. This method also allows for targeted recovery of damaged files while maintaining the integrity of the entire system. By focusing on specific damaged files, this process can be completed much faster than traditional full system recovery methods.

[0131] By selectively restoring only clean versions of the damaged files, this method preserves any legitimate changes made to files unaffected between snapshots. This targeted approach retains valuable data and user modifications that occurred after the baseline snapshot was created. Therefore, it strikes a balance between data recovery and the preservation of the most recent valid changes.

[0132] The ability to identify clean versions from intermediate snapshots provides a range of recovery points for data recovery. This flexibility allows for more precise selection of file versions, potentially enabling file recovery from points closer to the damaged snapshot. It offers finer granularity in selecting the most appropriate clean version for each file.

[0133] The aspects of this disclosure include the following embodiments.

[0134] Example 1: A method comprising: identifying baseline snapshots from a plurality of snapshots of protected data by a data platform implemented through a computing system, wherein the baseline snapshots include one or more files, each of which does not show an indication of corruption; for each file in an aberrant snapshot among the plurality of snapshots, identifying a clean version of the file by the data platform from one or more intermediate snapshots between the aberrant snapshot and the baseline snapshot; and storing clean snapshots by the data platform, the clean snapshots including a corresponding clean version of the corresponding file identified for the file in the aberrant snapshot.

[0135] Example 2. The method as described in Example 1, wherein the abnormal snapshot is compromised by malware, the method further comprising: using the data platform and one or more machine learning models to analyze one or more of the following: the malware or the data type affected by the malware.

[0136] Example 3. The method as described in Example 1, wherein the abnormal snapshot includes the most recent snapshot that has been determined to be corrupted.

[0137] Example 4. The method as described in any one of Examples 1 to 3, wherein identifying the clean version of the file includes identifying the clean version of the file in a secure environment.

[0138] Example 5. The method as described in Example 4, further comprising: obtaining registration information related to the security environment using natural language by one or more machine learning models of the data platform; and initiating the registration process of the security environment by the one or more machine learning models.

[0139] Example 6. The method as described in any one of Examples 1 to 5, further comprising: training the one or more machine learning models by the data platform using a dataset that includes at least a security environment knowledge base.

[0140] Example 7. The method as described in any one of Examples 1 to 6, further comprising: in response to receiving a deletion request from a user, deleting the security environment by the one or more machine learning models of the data platform.

[0141] Example 8. The method as described in any one of Examples 1 to 7, further comprising: restoring at least a portion of the protected data by the data platform based on the clean snapshot.

[0142] Example 9. The method of any one of Examples 1 to 8, wherein the protected data includes a first application workload, the method further comprising: predicting the criticality of the first application workload by one or more machine learning models of the data platform; obtaining user feedback from the data platform indicating the accuracy of the criticality prediction; and providing the user feedback to the one or more machine learning models by the data platform to generate a revised one or more machine learning models.

[0143] Example 10. The method as described in Example 9, wherein the protected data includes a second application workload, the method further comprising: predicting the criticality of the second application workload by the revised one or more machine learning models of the data platform, wherein the revised one or more machine learning models incorporate user feedback indicating the accuracy of the criticality prediction of the first application workload into the prediction of the criticality of the second application workload.

[0144] Example 11. The method as described in any one of Examples 1 to 10, wherein identifying the clean file for each file in the anomalous snapshot comprises: iterating through the one or more intermediate snapshots, and verifying the integrity of the corresponding file when a corresponding file for the file in the anomalous snapshot exists in one of the intermediate snapshots.

[0145] Example 12. The method as described in any one of Examples 1 to 11, wherein the baseline snapshot does not include any file that presents an indication of damage.

[0146] Example 13. A computing system comprising: a memory storing instructions; and processing circuitry executing the instructions to: identify a baseline snapshot from a plurality of snapshots of protected data, wherein the baseline snapshot includes one or more files, each file not exhibiting an indication of corruption; for each file in an aberrant snapshot among the plurality of snapshots, identify a clean version of the file from one or more intermediate snapshots between the aberrant snapshot and the baseline snapshot; and store a clean snapshot, the clean snapshot including a corresponding clean version of the corresponding file identified for the file in the aberrant snapshot.

[0147] Example 14. The computing system as described in Example 13, wherein the anomalous snapshot is compromised by malware, and the processing circuitry further executes the instructions to: use one or more machine learning models to analyze one or more of the following: the malware or the data type affected by the malware.

[0148] Example 15. The computing system as described in Example 13, wherein the abnormal snapshot includes the most recent snapshot that has been determined to be corrupted.

[0149] Example 16. A computing system as described in any one of Examples 13 to 15, wherein identifying the clean version of the file includes identifying the clean version of the file in a secure environment.

[0150] Example 17. The computing system as described in Example 16, wherein the processing circuitry further executes the instructions to: obtain registration information related to the security environment using natural language by one or more machine learning models; and initiate a registration process for the security environment by the one or more machine learning models.

[0151] Example 18. A computing system as described in any one of Examples 13 to 17, wherein the processing circuitry further executes the instructions to: train the one or more machine learning models with a dataset that includes at least a security environment knowledge base.

[0152] Example 19. A computing system as described in any one of Examples 13 to 18, wherein the processing circuitry further executes the instructions to: delete the security environment by the one or more machine learning models in response to receiving a deletion request from a user.

[0153] Example 20. A non-transitory computer-readable medium including instructions that, when executed, cause processing circuitry of a computing system to: identify a baseline snapshot from a plurality of snapshots of protected data, wherein the baseline snapshot includes one or more files, each of which does not exhibit an indication of corruption; for each file in an aberrant snapshot among the plurality of snapshots, identify a clean version of the file from one or more intermediate snapshots between the aberrant snapshot and the baseline snapshot; and store a clean snapshot, the clean snapshot including a corresponding clean version of the corresponding file identified for the file in the aberrant snapshot.

[0154] Various examples of this disclosure have been described. Any combination of the described systems, operations, or functions is contemplated.

Claims

1. A method, the method comprising: A data platform implemented through a computing system identifies anomalous snapshots from multiple snapshots of protected data, wherein the anomalous snapshots include one or more damaged files, and the one or more damaged files include an indication of damage; The data platform identifies a baseline snapshot from the plurality of snapshots, wherein the baseline snapshot does not include any files that show signs of damage; For each damaged file in the anomalous snapshot, the data platform identifies a clean version of the damaged file from the baseline snapshot or from one or more intermediate snapshots between the anomalous snapshot and the baseline snapshot, the clean version not including an indication of damage; and The data platform stores a clean snapshot, which includes a clean version of each of the corresponding damaged files in the abnormal snapshot.

2. The method as described in claim 1, wherein, The abnormal snapshot was compromised by malware, and the method further includes: The data platform uses one or more machine learning models to analyze one or more of the following: the malware or data types affected by the malware.

3. The method as described in claim 1, wherein, The abnormal snapshots include the most recent snapshots that have been identified as corrupted.

4. The method of claim 1, wherein, Identifying the clean version of the damaged file includes identifying the clean version of the damaged file in a secure environment.

5. The method of claim 4, further comprising: Registration information related to the security environment is obtained using natural language from one or more machine learning models of the data platform; as well as The registration process for the secure environment is initiated by one or more machine learning models.

6. The method of claim 5, further comprising: The data platform trains the one or more machine learning models using a dataset that includes at least a security environment knowledge base.

7. The method of claim 5, further comprising: In response to receiving a deletion request from a user, the security environment is deleted by one or more machine learning models of the data platform.

8. The method of claim 1, further comprising: At least a portion of the protected data is recovered by the data platform based on the clean snapshot.

9. The method of claim 1, wherein, For each damaged file in the abnormal snapshot, identifying clean files includes: Iterate through the one or more intermediate snapshots, and when the corresponding file for the damaged file in the abnormal snapshot exists in one of the intermediate snapshots, verify the integrity of the corresponding file.

10. A computing system, the computing system comprising: Processing component, the processing component being configured to: Identify anomalous snapshots from multiple snapshots of protected data, wherein the anomalous snapshots include one or more corrupted files, and the one or more corrupted files include an indication of corruption; Identify a baseline snapshot from the plurality of snapshots, wherein the baseline snapshot does not include any files that present an indication of damage; For each damaged file in the anomalous snapshot, a clean version of the damaged file is identified from the baseline snapshot or from one or more intermediate snapshots between the anomalous snapshot and the baseline snapshot, the clean version not including an indication of damage; and Store a clean snapshot containing a clean version of each of the corresponding damaged files in the abnormal snapshot.

11. The computing system of claim 10, wherein, The abnormal snapshot was compromised by malware, and the processing component is further configured to: Use one or more machine learning models to analyze one or more of the following: the malware or the data types affected by the malware.

12. The computing system of claim 10, wherein, The abnormal snapshots include the most recent snapshots that have been identified as corrupted.

13. The computing system according to any one of claims 10 to 12, wherein, Identifying the clean version of the damaged file includes identifying the clean version of the damaged file in a secure environment.

14. The computing system of claim 13, wherein, The processing component is further configured to: Registration information related to the security environment is obtained using natural language by one or more machine learning models; and The registration process for the secure environment is initiated by one or more machine learning models.

15. A computer-readable storage medium comprising instructions that, when executed, cause one or more processors of a computing system to perform the method as described in any one of claims 1 to 9.