Data distribution and security in multi-layer storage infrastructure

A multi-tier storage infrastructure with ensemble learning and encryption techniques addresses data vulnerability issues by securing file data across layers, ensuring secure and recoverable data access.

JP7841827B2Active Publication Date: 2026-04-07INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-05-13
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Conventional software-as-a-service technologies leave remote data vulnerable to security attacks and privacy breaches, especially in the event of an attack on a data center or network infrastructure.

Method used

Implementing a multi-tier storage infrastructure with a cloud, fog, and local computing layers, using ensemble learning models for data distribution and encryption, and applying hash transformations with cyclic error correction codes to protect and recover file data across these layers.

Benefits of technology

Enhances data security by requiring unauthorized access to all layers to intercept usable file data, and facilitates authorized data recovery even if data is lost in one layer, making it difficult for unauthorized users to access complete data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007841827000001
    Figure 0007841827000001
  • Figure 0007841827000002
    Figure 0007841827000002
  • Figure 0007841827000003
    Figure 0007841827000003
Patent Text Reader

Abstract

Techniques for data distribution and security in a multi-tier storage infrastructure are described. A related computer-implemented method includes receiving file data associated with a user for storage in a managed services domain, applying an ensemble learning model to devise a data distribution technique for the file data based on context information associated with the user, and encrypting the file data. The method further includes partitioning the file data for storage among a cloud computing tier, a fog computing tier, and a local computing tier by performing a hash transformation and applying at least one cyclic error correcting code based on the data distribution technique. In one embodiment, the method further includes receiving a data access request associated with the file data, authenticating the data access request, and recovering the file data via decryption.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The various embodiments described herein generally relate to data distribution and security in multi-layer storage infrastructures, and more specifically, the various embodiments describe techniques for distributing and protecting file data in a storage infrastructure having multiple layers, including a fog computing layer. [Overview of the Initiative]

[0002] Various embodiments described herein provide data distribution and security techniques. Various embodiments further provide data recovery techniques after data distribution. A relevant computer implementation method includes receiving user-related file data for storage in a managed service domain; applying an ensemble learning model to devise a data distribution technique for the file data based on contextual information related to the user; and encrypting the file data. The method further includes, based on the data distribution technique, distributing the file data for storage across a cloud computing layer, a fog computing layer, and a local computing layer by performing a hash transform and applying at least one cyclic error correction code. In one embodiment, the method further includes receiving a data access request related to the file data; authenticating the data access request; and recovering the file data via decryption.

[0003] One or more additional embodiments relate to a computer program product including a computer-readable storage medium implementing program instructions. According to such embodiments, the program instructions are executable by a computing device, which may be made to perform one or more steps of the computer implementation method described above, or to implement one or more embodiments related to the computer implementation method described above, or both. One or more further embodiments relate to a system comprising at least one processor and memory storing an application program. When the application program is executed on the at least one processor, it performs one or more steps of the computer implementation method described above, or to implement one or more embodiments related to the computer implementation method described above, or both.

[0004] To understand in detail how the above embodiments are achieved, a more specific description of the embodiments of the present invention outlined above can be obtained by referring to the accompanying drawings.

[0005] However, the accompanying drawings only illustrate typical embodiments of the present invention and should not be construed as limiting the scope of the invention. Other equally effective embodiments of the present invention may be possible. [Brief explanation of the drawing]

[0006] [Figure 1] This figure shows a cloud computing environment according to one or more embodiments. [Figure 2] This figure shows an abstraction model layer provided by a cloud computing environment according to one or more embodiments. [Figure 3] This figure shows a managed service domain related to a cloud computing environment according to one or more embodiments. [Figure 4]This figure shows a multi-tier storage infrastructure related to a managed service domain according to one or more embodiments. [Figure 5] This figure shows a method for processing file data according to one or more embodiments. [Figure 6] This figure shows a method of applying an ensemble learning model to devise a data distribution method for file data based on user context information, according to one or more embodiments. [Figure 7] This figure shows a method for deriving topic context data by applying natural language processing (NLP) to user context information, according to one or more embodiments. [Figure 8] This figure shows a method for creating multiple encoded feature vectors according to one or more embodiments. [Figure 9] This figure shows a method for deriving multiple file data portions by applying NLP to file data based on derived topic context data, according to one or more embodiments. [Figure 10] This figure shows a method for splitting file data according to one or more embodiments. [Figure 11] This figure shows a method for authenticating data access requests according to one or more embodiments. [Modes for carrying out the invention]

[0007] The various embodiments described herein relate to data processing technologies within a multi-tier storage infrastructure incorporating a cloud computing layer, a fog computing layer, and a local computing layer. The various embodiments provide technologies for distributing and protecting file data within the multi-tier storage infrastructure. The various embodiments further provide technologies for restoring distributed and protected file data upon receipt of an authenticated request. The various embodiments further provide data backup technologies incorporating data bins. In relation to the various embodiments, the cloud computing layer of the multi-tier storage infrastructure includes a managed service domain associated with the cloud computing environment. The cloud computing environment is a virtualized environment where one or more computing functions are available as a service. A cloud server system configured to implement the data distribution and security technologies related to the various embodiments described herein can utilize the artificial intelligence capabilities of machine learning knowledge models, specifically ensemble learning models, and knowledge base information associated with such models.

[0008] Various embodiments may offer advantages over conventional technologies. Conventional software-as-a-service technologies, for example, may leave remote data vulnerable to security attacks, privacy breaches, or both, in the event of an attack on a data center, network infrastructure, or both. Various embodiments improve computer technology by facilitating the distribution of file data across multiple layers of a multi-tier storage infrastructure, making it more difficult for unauthorized users to access the entire file data. According to various embodiments, data is partitioned across storage layers, requiring access to all layers to restore all file data across layers. Therefore, unauthorized users cannot intercept usable versions of file data without accessing the entire file data across all layers. Furthermore, various embodiments facilitate the inclusion of data bins in the local storage layers of a multi-tier storage infrastructure, allowing authorized users to recover complete data if file data is lost in one or more storage layers. Some embodiments may not include all of these advantages, and such advantages are not necessarily required in all embodiments.

[0009] Reference will now be made in detail to various embodiments of the present invention. It should be understood, however, that the present invention is not limited to the particular embodiments described. Rather, the present invention is contemplated to be implemented and practiced in any combination of the following features and elements, whether related to different embodiments or not. Furthermore, each embodiment may achieve other possible solutions or advantages over the prior art or both, but the present invention is not limited by whether a particular advantage is achieved by a given embodiment. Accordingly, the following aspects, features, embodiments, and advantages are merely illustrative and are not to be considered elements or limitations of the appended claims, except as explicitly recited in one or more of the claims. Similarly, when the term "the present invention" is used, it should not be construed as a generalization of the inventive subject matter disclosed herein and is not to be considered an element or limitation of the appended claims, except as explicitly recited in one or more of the claims.

[0010] The present invention can be a system, method, or computer program product, or a combination thereof, integrated at any possible level of technical detail. The computer program product may include a computer-readable storage medium storing computer-readable program instructions for causing a processor to execute aspects of the present invention.

[0011] A computer-readable storage medium can be a tangible device that holds and stores instructions for use by an instruction execution device. As an example, a computer-readable storage medium may be an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or a suitable combination thereof. As a more specific example of a computer-readable storage medium, there are a portable computer diskette, a hard disk, a RAM, a ROM, an erasable programmable ROM (EPROM or flash memory), a static random access memory (SRAM), a CD-ROM, a DVD, a memory stick, a floppy disk, a punched card, or a mechanically encoded device that records instructions in a raised structure in a groove, and a suitable combination thereof. A computer-readable storage medium as used herein should not be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.

[0012] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing devices / processing devices. Alternatively, they can be downloaded via a network (such as the Internet, a LAN, a WAN, or a wireless network, or a combination thereof) to an external computer or an external storage device. The network can include a copper transmission cable, an optical transmission fiber, a wireless transmission, a router, a firewall, a switch, a gateway computer, or an edge server, or a combination thereof. A network adapter card or network interface within each computing device / processing device receives the computer-readable program instructions from the network and transfers them for storage in a computer-readable storage medium in each computing device / processing device.

[0013] The computer-readable program instructions for performing the operations of the present invention may be assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk and C++, and conventional procedural programming languages ​​such as the C programming language and similar programming languages. The computer-readable program instructions can be executed as a standalone software package, either entirely on the user's computer or partially on the user's computer. Alternatively, they can be executed partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including LANs and WANs, or it may be connected to an external computer (for example, via the Internet using an Internet Service Provider). In some embodiments, electronic circuits, including, for example, programmable logic circuits, field-programmable gate arrays (FPGAs), and programmable logic arrays (PLAs), can execute computer-readable program instructions by utilizing state information of computer-readable program instructions in order to customize the electronic circuits for the purpose of performing aspects of the present invention.

[0014] Aspects of the present invention are described herein with reference to flowcharts or block diagrams, or both, of methods, apparatus (systems), and computer program products according to embodiments of the present invention. Each block in a flowchart or block diagram, or both, and combinations of blocks in a flowchart or block diagram, or both, are executable by computer-readable program instructions.

[0015] These computer-readable program instructions can be provided to the processor of a general-purpose computer, a dedicated computer, or other programmable data processing device for the production of a machine. This creates a means for these instructions, executed via the processor of such a computer or other programmable data processing device, to perform functions / operations identified in one or more blocks in a flowchart or block diagram, or both. These computer-readable program instructions can further be stored in a computer-readable storage medium that can be instructed to function in a particular manner for a computer, a programmable data processing device, or other device, or a combination thereof. Thus, the computer-readable storage medium containing the instructions constitutes a product containing instructions for performing functions / operations identified in one or more blocks in a flowchart or block diagram, or both.

[0016] Alternatively, a computer execution process may be generated by loading computer-readable program instructions into a computer, another programmable device, or other device, and having a series of operational steps executed on that computer, other programmable device, or other device. This ensures that the instructions executed on the computer, other programmable device, or other device perform functions / operations identified by one or more blocks in a flowchart, block diagram, or both.

[0017] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions containing one or more executable instructions for performing a specific logical function. In some other implementations, the functions shown within a block may be executed in an order different from the order shown in each diagram. For example, depending on the functions involved, two consecutively shown blocks may actually be achieved as a single process, executed simultaneously or nearly simultaneously, executed in a manner that partially or entirely overlaps in time, or the blocks may be executed in reverse order. Each block in a block diagram or flowchart or both, and combinations of multiple blocks in a block diagram or flowchart or both, are executable by a dedicated hardware-based system that performs a specific function or operation, or executes a combination of dedicated hardware and computer instructions.

[0018] In specific embodiments, techniques relating to data distribution and security in multi-layer storage infrastructure are described. However, the techniques described herein may be adapted to various purposes in addition to those specifically described herein. Therefore, the description of specific embodiments is illustrative and not limiting to the invention.

[0019] The various embodiments described herein can be provided to end users through cloud computing infrastructure. While this disclosure includes a detailed description of cloud computing, it should be understood that the implementation forms of the teachings described herein are not limited to cloud computing environments. Rather, the various embodiments described herein can be implemented in combination with any other type of computing environment that is currently known or will be developed in the future.

[0020] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be quickly provisioned and released with minimal administrative effort or interaction with service providers. Thus, cloud computing allows users to access virtual computing resources in the cloud (e.g., storage, data, applications, and even full virtualized computing systems) regardless of the underlying physical systems (or the location of those systems) used to provide the computing resources. This cloud model may include at least five characteristics, at least three service models, and at least four deployment models.

[0021] The characteristics are as follows:

[0022] On-demand self-service: Cloud consumers can unilaterally prepare computing power, such as server time and network storage, automatically as needed, without requiring human interaction with service providers.

[0023] Broad network access: Computing power is available over the network and accessible through standard mechanisms. This facilitates use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, personal digital assistants (PDAs)).

[0024] Resource pooling: A provider's computing resources are pooled and delivered to multiple consumers using a multi-tenant model. Various physical and virtual resources are dynamically allocated and reallocated as needed. Generally, consumers have a sense of location independence because they do not manage or know the exact location of the resources provided. However, consumers may be able to identify the location at a higher level of abstraction (e.g., country, state, data center).

[0025] Rapid Elasticity: Computing power can be prepared quickly and flexibly, allowing it to scale out automatically and immediately, and to be quickly released and scale in immediately. To consumers, the computing power available for preparation often appears unlimited and can be purchased in any quantity at any time.

[0026] Service Measurement: Cloud systems leverage metering capabilities at a level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, active user accounts) to automatically control and optimize resource usage. Resource usage can be monitored, controlled, and reported, providing transparency to both service providers and consumers.

[0027] The service model is as follows:

[0028] Software as a Service (SaaS): The functionality offered to consumers is the ability to use the provider's applications running on a cloud infrastructure. These applications can be accessed from various client devices via thin client interfaces such as web browsers (e.g., webmail). Consumers do not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application functions, except for configuring a limited number of user-specific applications.

[0029] Platform as a Service (PaaS): The functionality offered to consumers is the ability to deploy applications they have created or acquired to cloud infrastructure using programming languages ​​and tools supported by the provider. Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, and storage, but they can control the deployed applications and, in some cases, the configuration of their hosting environment.

[0030] Infrastructure as a Service (IaaS): The functionality provided to consumers is the provision of processors, storage, networking, and other basic computing resources that enable consumers to deploy and run any software, including operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but they can control the operating system, storage, and deployed applications, and in some cases, partially control certain network components (e.g., host firewalls).

[0031] The deployment model is as follows:

[0032] Private Cloud: This cloud infrastructure is operated exclusively for a specific organization. This cloud infrastructure can be managed by that organization or a third party and can reside on-premises or off-premises.

[0033] Community Cloud: This cloud infrastructure is shared by multiple organizations to support a specific community with common interests (e.g., mission, security requirements, policies, and compliance). This cloud infrastructure can be managed by the organization or a third party and can reside on-premises or off-premises.

[0034] Public Cloud: This cloud infrastructure is provided to a large number of people or large industry groups and is owned by organizations that sell cloud services.

[0035] Hybrid Cloud: This cloud infrastructure combines two or more cloud models (private, community, or public). While maintaining the unique entities of each model, they are bound together by standards or individual technologies to achieve data and application portability (e.g., cloud bursting for load balancing across clouds).

[0036] Cloud computing environments are service-oriented environments that emphasize statelessness, low coupling, modularity, and semantic interoperability. At the core of cloud computing is the infrastructure, which includes a network of interconnected nodes.

[0037] Figure 1 shows an exemplary cloud computing environment 50 according to one or more embodiments. As shown, the cloud computing environment 50 may include one or more cloud computing nodes 10. Local computer devices used by cloud consumers (e.g., personal digital assistants or mobile phones 54A, desktop computers 54B, laptop computers 54C, or automotive computer systems 54N, or a combination thereof) can communicate with these nodes. The nodes 10 can communicate with each other. The nodes 10 can be grouped physically or virtually (not shown) in one or more networks, such as the private, community, public, or hybrid clouds or a combination thereof. This allows the cloud computing environment 50 to provide infrastructure, platforms, or software as a service, or a combination thereof, without requiring cloud consumers to maintain resources on their local computer devices. Note that the types of computer devices 54A-N shown in Figure 1 are merely examples, and it should be understood that the computing nodes 10 and the cloud computing environment 50 can communicate with any type of electronic device via any type of network or network addressable connection (e.g., using a web browser) or both.

[0038] Figure 2 shows a set of functional abstraction layers provided by a cloud computing environment 50 according to one or more embodiments. The components, layers, and functions shown in Figure 2 are illustrative and are not limited to the various embodiments described herein. Various layers and corresponding functions are provided as illustrated. Specifically, the hardware and software layer 60 includes hardware components and software components. Examples of hardware components may include a mainframe 61, a reduced instruction set computer (RISC) architecture-based server 62, server 63, blade server 64, storage 65, and a network and network components 66. In some embodiments, the software components may include network application server software 67 and database software 68. The virtualization layer 70 provides an abstraction layer. From this layer, virtual entities such as a virtual server 71, virtual storage 72, a virtual network 73 including a virtual private network, a virtual application and operating system 74, and a virtual client 75 can be provided.

[0039] As an example, the management layer 80 can provide the following functions: Resource preparation 81 can enable the dynamic procurement of computing resources and other resources used to perform tasks within the cloud computing environment 50. Metering and pricing 82 can enable cost tracking as resources are used within the cloud computing environment 50, and billing or invoicing for the consumption of these resources. As an example, these resources may include licenses for application software. Security enables not only protection of data and other resources, but also identification and verification of cloud consumers and tasks. User portal 83 can provide consumers and system administrators with access to the cloud computing environment. Service level management 84 can enable the allocation and management of cloud computing resources to ensure that requested service levels are met. Service Level Agreement (SLA) planning and execution 85 can enable the pre-arrangement and procurement of cloud computing resources that are expected to be needed in the future in accordance with the SLA.

[0040] The workload layer 90 provides examples of available functions for the cloud computing environment 50. Examples of workloads and functions that can be provided from this layer include mapping and navigation 91, software development and lifecycle management 92, virtual classroom education delivery 93, data analysis processing 94, transaction processing 95, and data distribution and recovery 96. Data distribution and recovery 96 may enable the distribution of file data across storage layers and the recovery of such file data in response to authenticated requests, according to various embodiments described herein.

[0041] Figure 3 shows a managed services domain 300 within a cloud computing environment 50. Data distribution and recovery 96 and other workload / function-related functions may be performed in the managed services domain 300. The managed services domain 300 includes a cloud server system 310. In one embodiment, the cloud server system 310 includes a cloud server application 315. The cloud server application 315 is configured to enable, facilitate, or perform both data distribution and recovery according to the various embodiments described herein. The cloud server application 315 represents a single application or multiple applications. The cloud server application 315 includes an ensemble learning model 320 or is operably coupled with the ensemble learning model 320. Furthermore, in one embodiment, the managed services domain 300 includes a cloud computing interface 330, a database system 340, a directory server system 350, and multiple application server clusters 3601-360. n This includes one or more of the following. The cloud computing interface 330 enables communication between the cloud server system 310 and one or more client systems that interact with the managed service domain 300. As detailed with reference to Figure 4, the cloud computing interface 330 enables communication with the fog server system, local machines, or other components or combinations that interact with the managed service domain 300. Furthermore, in one embodiment, the cloud server system 310 includes a database system 340, a directory server system 350, or application server clusters 3601-360. n It is configured to communicate with or a combination of these. Furthermore, application server clusters 3601-360 nThe application servers within the domain may be configured to communicate with each other, with servers in other domains, with other components that interact with the managed services domain 300, or a combination of these.

[0042] The database system 340, in any configuration, coordinates and manages the knowledge base of the ensemble learning model 320. The knowledge base associated with the ensemble learning model 320, in any configuration, includes all repositories, ontologities, files, or documents, or a combination thereof, related to the model and its associated data processing. The database system 340 may include one or more database servers, which can coordinate, manage, or both, various elements of the knowledge base. The database system 340 includes the ensemble learning model 320 and multiple application server clusters 3601-360. n , and the relationships between knowledge bases are stored. In one embodiment, the database system 340 is managed via a database management system (DBMS), or optionally a relational database management system (RDBMS). In a further embodiment, the database system 340 includes one or more databases, some or all of which may be relational databases. In a further embodiment, the database system 340 includes one or more ontology trees or other ontology structures.

[0043] Directory server system 350 is application server cluster 3601-360 nThis facilitates client authentication within managed service domain 300, including client authentication for one or more of the following: For authentication for one or more of the multiple applications within managed service domain 300, the client can provide environment details and authentication information related to such applications. Application server clusters 3601-360 n It stores elements of various applications and provides application services to one or more client systems. The cloud server application 315 and other data collection components of the managed service domain 300 are configured to provide appropriate notifications regarding the collection of any user context data. One or more aspects of the managed service domain 300 are further configured to provide users with the option to opt in or opt out of such user context data collection at any time. In any aspect, one or more elements of the managed service domain 300 are further configured to send at least one notification to affected users based on user specifications (e.g., periodically or whenever such user context data collection occurs).

[0044] FIG. 4 is a diagram showing a multi-layer storage infrastructure 400 including a cloud computing layer 403 (i.e., cloud layer), a fog computing layer 413 (i.e., fog layer), and a local computing layer 423 (i.e., local layer). The cloud layer 403 of the multi-layer storage infrastructure 400 includes, is operatively coupled to, communicatively coupled to, or takes a combined form with the cloud server system 310 and other elements of the managed service domain 300. The cloud server system 310 stores cloud layer data 405. The cloud layer data 405 includes, as an optional aspect, file data distributed according to various embodiments described herein. The cloud server system 310 stores or facilitates the storage of the cloud layer data within one or more storage components or database components or both related to the managed service domain 300 (e.g., within one or more components described with respect to FIG. 3, or within one or more other hardware-based storage or database components virtualized within the cloud layer 403 as an optional aspect). The fog layer 413 of the multi-layer storage infrastructure 400 includes fog server systems 4101 to 41 n 0. The fog server systems 4101 to 41 n 0 each store fog server data 4151 to 415 n within one or more hardware-based storage or database components respectively associated with the fog server systems 4101 to 41 n 0. Such hardware-based storage or database components are virtualized within the fog layer 413 as an optional aspect. As an alternative configuration, the fog layer 413 includes a single fog server system corresponding to one or a combination of elements 4101 to 41 n 0, and this system stores elements 4151 to 415 nIt stores fog layer data corresponding to one or a combination of the following. In relation to the various embodiments described herein, fog computing within the fog layer, specifically fog layer 413 as described in Figure 4, refers to distributed computing that includes one or more devices or systems located around the cloud computing environment, situated between cloud resources and local resources. Fog computing associated with such a fog layer can function as decentralized intermediary between remote cloud computing associated with the cloud layer, specifically cloud layer 403 as described in Figure 4, and local computing devices or systems located in the local layer, specifically local layer 423 as described in Figure 4. Fog computing enables the aggregation of data associated with multiple devices into a storage node with regional connectivity.

[0045] The local tier 423 of the multi-layer storage infrastructure 400 is local machines 4201-420 n Includes local machines 4201-420. n Each of these represents or includes at least one device or system, or both, that incorporates at least one hardware component. Local machines 4201-420 n These are, for example, local machines 4201-420 n Within one or more hardware-based storage or database components associated with each, local tier data 4251-425 n It will remember this. Furthermore, local machines 4201~420 n These are local machines 4201-420, respectively. n Within one or more hardware-based storage or database components associated with each of them, data bins 4351-435 n Includes: Fog server systems 4101-410 nEach of these may, in any manner, be one or more local machines 4201-420 n It provides regional connectivity, data aggregation functionality, or both. As an alternative configuration, the local layer 423 contains elements 4201-420 n It includes a single local machine corresponding to one or a combination of the following, and this local machine is element 4251-425 n It stores local layer data corresponding to one of the elements, and elements 4351-435 n It contains (or is operablely coupled with) a data bin corresponding to one or a combination of the following.

[0046] As an alternative configuration, the multi-layer storage infrastructure 400 includes additional intermediate layers in addition to the three layers described above, for example, an intermediate layer between the cloud layer 403 and the fog layer 413, or an intermediate layer between the fog layer 413 and the local layer 423, or both. Such intermediate layers, in any manner, buffer data for transmission between the cloud layer 403, the fog layer 413, or the local layer 423, or a combination thereof. In addition to or instead of this, such intermediate layers incorporate the characteristics of adjacent layers. For example, the intermediate layer between the cloud layer 403 and the fog layer 413, in any manner, incorporates certain elements of both the cloud layer 403 and the fog layer 413. In another example, the intermediate layer between the fog layer 413 and the local layer 423 incorporates certain elements of both the fog layer 413 and the local layer 423.

[0047] In one embodiment, the managed service domain 300 is a hybrid cloud interface that is communicably coupled to the fog layer 413 of the multi-layer storage infrastructure 400 via at least one network connection. According to such an embodiment, the managed service domain 300 further, in any manner via at least one network connection and the fog layer 413, connects to at least one local machine associated with a user in the local layer 423 of the multi-layer storage infrastructure 400 (e.g., local machines 4201-420).n It is communicatively coupled to at least one of the following. In such embodiments, the fog layer 413 is, in any manner, a component of a virtual private cloud in a hybrid cloud environment, where an on-demand pool of configurable resources is allocated to one or more designated users. In an additional embodiment, at least one local machine is communicatively coupled to a managed service domain 300 (specifically, one or more of its components) via an application programming interface (API) that facilitates cloud connectivity. In a further embodiment, local machines 4201-420 n At least one of these communicates with one or more components of the managed service domain 300 via a data exchange format such as JavaScript Object Notation (JSON) or Extensible Markup Language (XML). In a further embodiment, local machines 4201-420 n At least one of these is an edge computing device. According to such further embodiments, if the file data processed according to various embodiments relates to data collected from sensors or distributed computing locations in relation to an Internet of Things (IoT) infrastructure, for example, such local machines can interact with the cloud layer 403 or the fog layer 413 or both via the edge computing device. Local machines 4201-420 n By including at least one of these as an edge computing device, a hybrid fog-edge computing implementation becomes possible within the multi-tier storage infrastructure 400, providing the local system advantages of edge computing and the interoperability and aggregation capabilities of fog computing.

[0048] Figure 5 shows a method 500 for processing file data. Specifically, method 500 relates to encrypting file data and distributing it across storage layers of a multi-tier storage infrastructure (e.g., multi-tier storage infrastructure 400), and further relates to decrypting and restoring the file data in response to receiving an authenticated data access request. In one embodiment, one or more steps related to method 500 are performed in an environment where computing power is provided as a service (e.g., a cloud computing environment 50). According to such an embodiment, one or more steps related to method 500 are performed in a managed service domain (e.g., a managed service domain 300) within the environment. In any embodiment, this environment is a hybrid cloud environment. In a further embodiment, one or more steps related to method 500 are performed in one or more other environments, such as a client-server network environment or a peer-to-peer network environment. A centralized cloud server system within a managed service domain (e.g., a cloud server system 310 within a managed service domain 300) can facilitate processing by method 500 and other methods further described herein. More specifically, a cloud server application within a cloud server system (e.g., cloud server application 315) can perform or facilitate the performance of one or more steps of Method 500 and other methods described herein. Data processing techniques facilitated or performed through a cloud server system within a managed service domain can be associated with data distribution and recovery workloads within the workload layer of the functional abstraction layers provided by the environment (e.g., data distribution and recovery 96 within workload layer 90 of cloud computing environment 50).In a further embodiment, the cloud server application performs one or more steps of Method 500 and other methods described herein via one or more programming instructions encoded via a high-level programming language (e.g., Python, C, C++, or C# or a combination thereof).

[0049] Method 500 first involves step 505, in which the cloud server application receives user-related file data for storage in the managed service domain. In any embodiment, the received file data is data that the user or related entity intends to store in the multi-tier storage infrastructure for privacy, accessibility, security, or a combination thereof. In one embodiment, the received file data is data from a single file. Alternatively, the received file data is data from multiple files. In one embodiment, the cloud server application receives the file data in the managed service domain via a cloud computing interface (e.g., cloud computing interface 330) in any embodiment. According to such an embodiment, the cloud server application receives the received file data within the cloud server system and subsequently processes it. In a further embodiment, the cloud server application receives the file data via a user-related client interface in step 505. According to such further embodiments, and other embodiments described herein, the client interface is a user system or device in the multi-tier storage infrastructure (e.g., local machines 4201-420 in the multi-tier storage infrastructure 400). nA user interface in the form of a graphical user interface (GUI), command-line interface (CLI), or both, which is installed on one of the above, presented via at least one client application that can otherwise access it, or is operable / communicatively coupled to the multi-tier storage infrastructure.

[0050] In step 510, the cloud server application applies an ensemble learning model (e.g., ensemble learning model 320) to devise a data distribution technique for file data based on contextual information relevant to the user. In relation to the various embodiments described herein, the ensemble learning model identifies topics or groups of topics or both related to the file data and incorporates and coordinates multiple artificial intelligence techniques, including machine learning techniques and, if necessary, deep learning techniques, to classify the file data based on such topic identification. The multiple artificial intelligence techniques may include one or more natural language processing (NLP) models or one or more multi-class classification techniques or both. In relation to the various embodiments described herein, the data distribution technique divides the file data across storage layers of a multi-tier storage infrastructure and further distributes redundant data across storage layers to maintain dependencies necessary for restoring or transferring the file data or both. The cloud server application devises a data distribution technique based on the topic identification and classification obtained via the ensemble learning model. The data distribution technique assigns portions or sub-portions of the file data to their respective storage layers. As further described herein, each storage tier, in any embodiment, correlates with data buckets formed by a cloud server application partitioning file data. In relation to the various embodiments described herein, a data bucket is a data structure configured to store distributed data. A data bucket, in any embodiment, is or contains a container or buffer.

[0051] During the application of the ensemble learning model in step 510, the cloud server application performs topic analysis on the received file data, taking into account contextual information related to the user. In one embodiment, the contextual information related to the user is based on past or currently observed user patterns of application access and data access, or both, as well as various attributes related to the user. In such an embodiment, the cloud server application, in any manner, applies the ensemble learning model to devise a data distribution technique based at least partially on one or more of several user contextual factors. These user contextual factors, in any manner, include the user's data usage frequency, the user's application usage frequency, the user's system configuration, the user's file storage patterns, the user's attributes (e.g., the user's occupation, the organization associated with the user, the user's personal or business contacts), or the user's data content related to applications or files used in the past, recently, or currently (e.g., the user's data file type and / or size), or a combination thereof. The method for applying the ensemble learning model in step 510 is illustrated with reference to Figure 6.

[0052] In step 515, the cloud server application encrypts the file data. The cloud server application encrypts the file data in preparation for executing the data distribution technique devised in step 510. In one embodiment, the cloud server application devises an encryption strategy that incorporates one or more encryption techniques based on user context information, for example, based on security or privacy settings or both associated with such information. In addition to or instead of this, the cloud server application devises an encryption strategy based on the data type or data size associated with the file data or a portion of the file data or both. In an additional embodiment, the cloud server application completes encryption according to step 515 before further processing of the file data to avoid security risks associated with data interception, for example, man-in-the-middle (MitM) attacks. In such an additional embodiment, the cloud server application encrypts the file data before splitting the file data according to the devised data distribution technique. In addition to or instead of this, the cloud server application encrypts the file data before performing a hash transformation. In a further embodiment, the cloud server application applies a first iteration of encryption to the entire file data, and then applies one or more subsequent iterations of encryption to each portion of the file data distributed across layers of the multi-layer storage infrastructure as determined via an ensemble learning model. In a further embodiment, one or more encryption techniques related to the encryption strategy include symmetric key algorithms such as TDES (Triple Data Encryption Algorithm) and AES (Advanced Encryption Standard algorithm). According to such a further embodiment, the cloud server application encrypts the file data by applying TDES in any manner. In addition to or instead of this, the cloud server application encrypts the file data by applying AES in any manner.In a further embodiment, one or more encryption techniques include an asymmetric key algorithm such as RSA (Rivest-Shamir-Adleman). In a further embodiment, the cloud server application applies multiple encryption algorithms to all or specific parts of the file data to enhance security or privacy or both. According to such a further embodiment, the cloud server application, in any manner, applies a relatively strong encryption algorithm or a larger number of encryption algorithms or both to file data or parts of file data that are marked as sensitive data rather than non-sensitive data, or that are associated with sensitive data.

[0053] In step 520, the cloud server application divides the file data encrypted in step 515 based on the data distribution technique devised in step 510 and distributes and stores it among the storage layers of the multi-layer storage infrastructure. In accordance with step 520, the cloud server application performs the data distribution technique. In one embodiment, the cloud server application divides the file data among the cloud computing layer, fog computing layer, and local computing layer of the multi-layer storage infrastructure (e.g., cloud layer 403, fog layer 413, and local layer 423 of the multi-layer storage infrastructure 400). In an alternative embodiment, the cloud server application divides the file data among one or more intermediate layers in addition to the cloud layer, fog layer, and local layer. Such one or more intermediate layers may include an intermediate layer between the cloud layer and the fog layer that incorporates specific elements of both the cloud layer and the fog layer in any manner. In addition to or instead of this, such one or more intermediate layers may include an intermediate layer between the fog layer and the local layer that incorporates specific elements of both the fog layer and the local layer in any manner. By partitioning file data across storage layers, the cloud server application completes the distribution of each portion of the file data using the devised data distribution technology.

[0054] When splitting file data according to step 520, the cloud server application performs a hash transformation and applies at least one cyclic error correcting code. The cloud server application performs the hash transformation by applying one or more cryptographic hash functions. In one embodiment, the cloud server application performs the hash transformation by applying the MD5 message digest algorithm (MD5). In an additional embodiment, the cloud server application performs the hash transformation by applying the Secure Hash Algorithm 2 (SHA-2). In a further embodiment, the cloud server application performs the hash transformation by applying multiple cryptographic hash functions (e.g., a joint application of MD5 and SHA-2). The cloud server application applies one or more of these cryptographic hash functions to generate a hash value in bit array format. In a further embodiment, the cloud server application adds further security when generating the hash value by randomizing such a hash value by adding a random string, i.e., a salt, to the file data string to be hashed before generating the hash value. In a further embodiment, the cloud server application performs hash transformation by separately hashing file data that is divided across each storage layer of the multi-tier storage infrastructure. In this case, a different hash value is generated for each string portion of the file data distributed across the storage layers. Specifically, the cloud server application optionally generates a first set of hash values ​​for the file data distributed across the cloud computing layer, a second set of hash values ​​for the file data distributed across the fog computing layer, and a third set of hash values ​​for the file data distributed across the local computing layer.According to such further embodiments, the cloud server application optionally applies different cryptographic hash functions to generate different sets of hash values ​​depending on the different security requirements of each computing layer. In further embodiments, the cloud server application integrates the hash transformation and error correction code into a single algorithm or procedure. In further embodiments, the cloud server application applies at least one cyclic error correction code by applying at least one error correction code from the BCH (Bose-Chaudhuri-Hocquenghem) error correction code class, i.e., by applying at least one BCH error correction code. According to such further embodiments, the cloud server application optionally applies a Reed-Solomon subset of BCH error correction codes. A method relating to the partitioning of file data according to step 520 is described with reference to Figure 10.

[0055] In one embodiment, the cloud server applies an ensemble learning model, encrypts file data, or partitions file data across storage layers, or a combination thereof, according to steps 510-520, by executing a bash script containing a set of tasks and subtasks related to the method steps. More specifically, the cloud server application encrypts file data by executing such a bash script based on the output of the ensemble learning model. In relation to various embodiments, the bash script is a file containing a set of commands to facilitate the execution of a command-line program. According to such embodiments, the bash script facilitates the execution of automation techniques that combine encryption and file data partitioning with all elements of the ensemble learning process. Thus, the cloud server application executes a bash script to automate the process flow in any manner, in relation to one or more of the various embodiments. In an additional embodiment, the cloud server application further executes hardware-based or virtualized data bins associated with the user's local machine (e.g., local machines 4201-420 in the multi-tier storage infrastructure 400). n Data bins 4351-435 are associated with one of these. n One of these stores the entire file data in an unpartitioned state. According to such an embodiment, if a cloud server application determines that partitioned file data has been lost between one or more layers of a multi-layer storage infrastructure, it can recover the file data via the data bin.

[0056] In step 525, the cloud server application receives a data access request related to file data. In one embodiment, the cloud server application receives the data access request via a client interface associated with the user. The client interface is, in any embodiment, the same interface on which the file data was received in step 505, or a different client interface. In step 530, the cloud server application authenticates the data access request. In any embodiment, the cloud server application uses a directory server system associated with the managed service domain (e.g., directory server system 350) for authentication. In one embodiment, the cloud server application authenticates the data access request by applying one or more cryptographic hash functions (in any embodiment, in a manner similar to how one or more cryptographic hash functions were applied when splitting the file data in step 520). In any embodiment, if authentication of the data access request fails, the cloud server application repeats step 530 or proceeds to the end of method 500. The method relating to the authentication of the data access request according to step 530 is described with reference to Figure 11.

[0057] In step 535, the cloud server application restores the file data through decryption. In one embodiment, the cloud server application restores the file data, which has been partitioned between storage layers of a multi-tier storage infrastructure, in the same format in which it was originally constructed. Alternatively, the cloud server application restores the file data in a format modified according to the user's preference (e.g., a compressed format, an encrypted format, or both). In one embodiment, through the use of a key, the cloud server application decrypts the file data by applying one or more decryption techniques corresponding to one or more encryption techniques applied in step 515. In a further embodiment, if the cloud server application determines that the file data has been lost between one or more storage layers (e.g., due to corruption of data elements in one or more of the storage layers), it restores the entire file data by accessing a data bin in which the file data is stored in an unpartitioned form. In a further embodiment, the cloud server application performs one or more of steps 525-535 in a procedure separate from steps 505-520.

[0058] Figure 6 shows a method 600 for applying an ensemble learning model. Method 600 provides a substep in step 510 of Method 500 according to one or more embodiments. Method 600 first derives topical context data in step 605 by having a cloud server application apply NLP to contextual information related to the user. According to step 605, the cloud server application applies NLP to user context information to derive topical context data. The cloud server application applies one or more NLP models / algorithms to user context information to derive topical context data. The derived topical context data includes topics or groups of topics related to user context information, or both. Such topics / groups of topics may, in any embodiment, include topics related to one or more of several user context factors, for example, topics related to the user's application activity, topics related to the user's data activity, topics related to the user's resource usage or user memory patterns, or topics related to user attributes (such as the user's occupation, the user's affiliated organizations, or the user's contacts), or a combination thereof. In any configuration, the derived topic context data may further include topical metadata that describes the association between a user and a topic / group of topics, or between topics / groups of topics, or both. Specifically, the topical metadata may describe relationship information between each topic / group of topics and user data, user applications, or user attributes or combinations thereof, such as date / time information or access information for user data related to security-related topics.

[0059] The cloud server application applies the unsupervised learning capabilities of an ensemble learning model to user context aspects identified within user context information via NLP in order to identify topic patterns within the user context information. In one embodiment, the user context aspects include data elements obtained by parsing or other means from one or more of a plurality of user context factors. In one embodiment, the cloud server application applies NLP to the user context information by incorporating at least one natural language understanding (NLU) technique, in any manner. In addition to or instead of this, the cloud server application applies NLP to the user context information by incorporating at least one automatic speech recognition (ASR) technique. In a further embodiment, the cloud server application parses text from the text elements of the user context information to identify the user context aspects to which NLP is applied. In addition to or instead of this, the cloud server application identifies user context aspects for NLP analysis by applying audiovisual processing to the audiovisual elements of the user context information. When the cloud server application identifies audio related to a user (e.g., speech utterances of the user or a related contact), it may optionally apply speech recognition (e.g., speech-to-text conversion) to derive text-based elements from the audio, and then apply NLP to those text-based elements. In addition to, or instead of, when the cloud server application identifies visual images related to a user (e.g., still images, videos, or both of the user's activities or the activities of a related contact), it may optionally apply video recognition (e.g., video-to-text conversion) to derive text-based elements from the visual images, and then apply NLP to those text-based elements.The cloud server application, in any embodiment, applies one or more other forms of audiovisual processing to audio or visual images or both within the user context information to identify further user context elements. In a further embodiment, the cloud server application derives user context elements for NLP analysis by analyzing monitoring data collected in relation to the user (e.g., monitoring data collected as a result of collecting Internet of Things (IoT) sensor data from multiple monitoring sensors that are attached to the user, attached to devices related to the user, installed in an environment related to the user, or a combination thereof). The cloud server application, in any embodiment, applies NLP to text elements derived from sensor data or sensor metadata or both.

[0060] In one embodiment, the cloud server application quantifies the topical relevance between derived topic context data by assigning a user context score to each topic, each group of topics, or both. The user context score indicates the relative importance of each topic, each group of topics, or both to the user in terms of file storage access. The cloud server application, in any embodiment, prioritizes storing data associated with topics, each group of topics, or both that have relatively high user context scores in order to speed up user access. Conversely, the cloud server application, in any embodiment, allows for more flexible storage of such data, for example, by allowing storage of data associated with topics, each group of topics, or both that have relatively low user context scores in locations where storage costs are relatively low but access latency is relatively high. According to such embodiments, the cloud server application, in any embodiment, assigns each user context score on a predetermined numerical scale, for example, an integer scale from 0 to 100.

[0061] As further described herein, in any embodiment, the cloud server application encodes feature vectors based on topic patterns identified within user context information and derives topic context data based on processing of the encoded feature vectors, for example, through the application of at least one clustering algorithm. In relation to various embodiments, the feature vector is an n-dimensional representation of data points that describes each feature of the data points in numerical form. In relation to various embodiments, the data points are entities represented in the data, for example, individuals, organizations, or application elements. In relation to various embodiments, the features are measurable properties or characteristics associated with the data points. A method relating to the derivation of topic context data according to step 605 is described with reference to Figure 7.

[0062] In step 610, the cloud server application derives multiple file data portions by applying NLP to the file data based on the topic context data derived in step 605. According to step 610, in order to derive multiple file data portions, the cloud server application applies NLP to the file data based on the derived topic context data. Using the derived topic context data, which includes topics / groups of topics related to user context information, as input, the cloud server application applies the unsupervised learning capabilities of an ensemble learning model to the file data aspects via NLP to identify topic patterns between the file data. In one embodiment, the cloud server application applies NLP to the file data by incorporating at least one NLU technique in any manner. In addition to or instead of this, the cloud server application applies NLP to the file data by incorporating at least one ASR technique. In a further embodiment, the cloud server application parses text from the text elements of the file data to identify the file data elements to which NLP is applied. In addition to or instead of the above, the cloud server application identifies file data elements for NLP analysis by applying audiovisual processing to the audiovisual elements within the file data. When it identifies audio (e.g., speech or other sounds) within the file data, the cloud server application optionally applies speech recognition (e.g., speech-to-text) to identify text-based elements from the audio, and then applies NLP to those text-based elements. When it identifies visual images (e.g., still images or videos or both) within the file data, the cloud server application optionally applies video recognition (e.g., video-to-text) to identify text-based elements from the visual images, and then applies NLP to those text-based elements.The cloud server application may, in any manner, apply one or more other forms of audiovisual processing to audio or visual images or both within the file data in order to identify further file data elements.

[0063] As will be further described herein, based on topic patterns identified among file data, the cloud server application may, in any manner, create multiple file data portions and assign data points within the file data to the multiple file data portions, for example, through the application of at least one clustering algorithm. A method relating to deriving multiple file data portions according to step 610 will be illustrated with reference to Figure 9.

[0064] In step 615, the cloud server application applies at least one multi-class classification technique to the multiple file data portions derived in step 610. Following step 615, the cloud server application associates the derived file data portions, or their respective sub-parts, or both, with one or more classes from the multiple classes. Based on the topic labeling determined via NLP applied in steps 605-610, the cloud server application applies the supervised learning capabilities of an ensemble learning model to associate the derived file data portions, or their respective sub-parts, with one or more classes from the multiple classes. In one embodiment, the cloud server application applies a random forest classifier (RFC) decision tree ensemble. In addition to or instead of this, the cloud server application applies a support vector machine (SVM) model (an SVM ensemble model in any embodiment). The cloud server application uses the derived file data portions as input to achieve supervised learning. Based on processing at each decision tree node related to at least one classification technique, the cloud server application generates output containing information that identifies associations between each file data portion and one or more of several classes. In relevant embodiments, the output includes, for each file data portion or sub-part, a numerical output (e.g., integer, binary, or one-hot encoded vector output, or a combination thereof) that corresponds to or otherwise refers to one or more specific classes of several classes. In additional relevant embodiments, the output includes data relating to the distribution of file data portions based on several classes, such as statistical data or metadata or both relating to the relationships between several classes and each file data portion or sub-part.

[0065] The cloud server application, in accordance with step 615, associates a file data portion, or each of its sub-parts, or both, with one or more classes based on one or more specific elements of user context information, for example, based on one or more of several user context factors. In one embodiment, the cloud server application classifies the file data portion by incorporating several user context elements. These user context elements include one or more elements relating to the frequency of data use, the relevance of data, the data type, the data size, the complexity of data, the sensitivity of data, or the priority of data related to a particular user task, or a combination thereof. In a further embodiment, the cloud server application classifies the file data portion based on its relevance to each data type. Specifically, the cloud server application may, in any manner, classify file data portions having the most frequently used data types (or data related to the most frequently used applications) into one or more classes associated with the local tier, file data portions relating to less frequently used data types into one or more classes associated with the fog tier, and file data portions relating to the least frequently used data types into one or more classes associated with the cloud tier. In addition to or instead of this, the cloud server application classifies the file data portion based on the relevance of the data. Specifically, the cloud server application may, in any manner, classify the file data portion containing data deemed most relevant to user activity into one or more classes associated with the local layer, the file data portion containing less relevant data into one or more classes associated with the fog layer, and the file data portion containing the least relevant data into one or more classes associated with the cloud layer. In addition to or instead of this, the cloud server application may classify the file data portion based on the relevance of the application.Specifically, the cloud server application may, in any manner, classify the portion of file data related to the applications most frequently used during user activity into one or more classes associated with the local tier, the portion of file data related to applications that are not used relatively frequently during user activity into one or more classes associated with the fog tier, and the portion of file data related to the least frequently used applications into one or more classes associated with the cloud tier. In addition to or instead of this, the cloud server application classifies the file data portion based on file size. Specifically, the cloud server application may, in any manner, classify the portion of file data related to larger user files into one or more classes associated with the cloud tier, the portion of file data related to more compact user files into one or more classes associated with the fog tier, and the portion of file data related to the most compact user files into one or more classes associated with the local tier.

[0066] In an additional embodiment, the cloud server application classifies file data portions or subports that are less relevant, less compact, less frequently used, or a combination of these into one or more classes associated with the cloud tier; file data portions or subports that are more relevant, more compact, more frequently used, or a combination of these into one or more classes associated with the fog tier; and the most relevant, most compact, or most frequently used file data portions or subports into one or more classes associated with the local tier. In addition to or instead of this, the cloud server application classifies file data portions that are more complex, more sensitive, or both into one or more classes associated with the cloud tier for the purpose of accessibility. In addition to or instead of this, the cloud server application classifies file data portions that have the highest relative priority in relation to user tasks into classes associated with all storage tiers, so that users can access such file data portions in any data storage scenario. In addition to or instead of this, the cloud server application classifies any portion of file data or its sub-portions that are necessary for remote data processing, related to the availability of remote data access, or both, into one or more classes associated with the cloud tier, or one or more classes associated with the fog tier, or both. In addition to or instead of this, the cloud server application classifies any portion of file data or its sub-portions related to local resources into one or more classes associated with the local tier.

[0067] In a further embodiment, the cloud server application classifies each file data portion, or a sub-part thereof, or both, into one or more classes according to step 615, at least in part, based on the respective user context score assigned to each topic or group of topics in the derived topic context data. According to such a further embodiment, the cloud server application optionally assigns each file data portion or sub-part associated with a topic or group of topics or both having a relatively high user context score to one or more classes that correlate to a relatively high storage priority in order to facilitate relatively fast user access. Conversely, the cloud server application optionally assigns each file data portion or sub-part associated with a topic or group of topics or both having a relatively low user context score to one or more classes that correlate to a more flexible storage configuration.

[0068] The cloud server application classifies each file data portion into one or more classes of a plurality of classes according to step 615, based on one or more classification techniques. In one embodiment, one or more classification techniques classify with respect to storage layers related to a multi-tier storage infrastructure, e.g., cloud layer, fog layer, and local layer. According to the first classification technique, each of the plurality of classes is associated with only one of the storage layers. Specifically, the first class is associated with the cloud layer, the second class is associated with the fog layer, and the third class is associated with the local layer. According to this first classification technique, the cloud server application associates a file data portion whose entirety is distributed across a particular layer of storage layers with a single class associated with that particular layer (i.e., a cloud layer class, a fog layer class, or a local layer class). For example, the cloud server application may associate a file data portion whose entirety is distributed across the cloud layer with a cloud layer class. Conversely, the cloud server application may associate a file data portion distributed across multiple layers with multiple classes of a plurality of classes, based on layer distribution. In this case, the cloud server application may subdivide the file data portion into its respective sub-parts based on the associated class. For example, the cloud server application can associate the file data portion distributed between the cloud tier and the fog tier with both the cloud tier class and the fog tier class. In this case, the cloud server application can identify the cloud tier sub-part of the file data portion based on its association with the cloud tier class, and further identify the fog tier sub-part of the file data portion based on its association with the fog tier class.

[0069] According to the second classification technique, each of several classes is associated with one or a combination of storage layers in any manner. Specifically, the first class is associated with the cloud layer, the second class with the fog layer, the third class with the local layer, the fourth class is associated with both the cloud and fog layers in any manner, the fifth class is associated with both the cloud and local layers in any manner, the sixth class is associated with both the fog and local layers in any manner, and the seventh class is associated with all three layers in any manner. According to the second classification technique, a cloud server application, similar to the first classification technique, associates a portion of file data distributed entirely across a specific storage layer with a single class associated with that specific layer (i.e., a cloud layer class, a fog layer class, or a local layer class). Conversely, a cloud server application associates a portion of file data distributed across multiple layers with a class that reflects each of those layers. For example, a cloud server application can associate the file data portion, which is distributed between the cloud layer and the fog layer, with classes associated with both the cloud layer and the fog layer.

[0070] In one embodiment, the cloud server application associates one or more numerical class values ​​with each of a plurality of file data portions, each of its sub-parts, or both. According to such an embodiment, the cloud server application optionally creates an encoded feature vector containing numerical data that associates one or more numerical class values ​​with each of the plurality of file data portions, each of its sub-parts, or both. In an alternative embodiment, according to one or more classification techniques, the cloud server application performs classification with respect to an intermediate layer related to the multi-layer storage infrastructure, in addition to the cloud layer, fog layer, and local layer. In such an alternative embodiment, one or more classification techniques include each class associated with an intermediate layer between the cloud layer and the fog layer, or an intermediate layer between the fog layer and the local layer, or a combination thereof.

[0071] In step 620, the cloud server application determines a distribution plan for multiple file data portions based on the application of at least one multi-class classification technique. The distribution plan determined in step 620 is the execution of the data distribution technique devised in step 510. In one embodiment, the cloud server application determines the distribution plan based on the classification of file data portions via at least one multi-class classification technique. Based on the computational intelligence described above, the cloud server application calculates the distribution ratio of file data stored in each layer of the multi-layer storage infrastructure (e.g., cloud layer, fog layer, and local layer). In an optional embodiment, the cloud server application calculates the distribution ratio based on the classification of file data portions into multiple classes. In an additional embodiment, the cloud server application includes in the distribution plan a procedure for distributing the file data to the respective buckets in the cloud layer, fog layer, and local layer. In such an additional embodiment, the cloud server application optionally specifies a distribution order that specifies the order in which the file data is distributed to the respective buckets. According to one variation of the distribution order, the cloud server application may distribute file data first to the buckets in the cloud tier, then to the buckets in the fog tier, and then to the buckets in the local tier. According to another variation of the distribution order, the cloud server application may distribute file data first to the buckets in the cloud tier, then to the buckets in the local tier, and then to the buckets in the fog tier. According to yet another variation of the distribution order, the cloud server application may distribute file data first to the buckets in the fog tier, then to the buckets in the cloud tier, and then to the buckets in the local tier. According to yet another variation of the distribution order, the cloud server application may distribute file data first to the buckets in the fog tier, then to the buckets in the local tier, and then to the buckets in the cloud tier.According to yet another variation of the distribution order, the cloud server application may distribute the file data first to buckets in the local layer, then to buckets in the cloud layer, and then to buckets in the fog layer. According to yet another variation of the distribution order, the cloud server application may distribute the file data first to buckets in the local layer, then to buckets in the fog layer, and then to buckets in the cloud layer. If the multi-layer storage infrastructure includes an intermediate layer, the distribution order, in any embodiment, includes the buckets in the intermediate layer in addition to the buckets in the cloud layer, the buckets in the fog layer, and the buckets in the local layer. In a further embodiment, the cloud server application appends ensemble learning metadata to the file data based on the NLP and multi-class classification applied in steps 605-615. According to such a further embodiment, the cloud server application, in any embodiment, uses the ensemble learning metadata to determine the distribution ratio between layers of the multi-layer storage infrastructure. The ensemble learning metadata, in any embodiment, incorporates information relating to the portion of the file data derived in step 610. Information relating to each file data portion, which may be incorporated into the ensemble learning metadata in any manner, may include the size of the file data portion, the data type associated with the file data portion, the application associated with the file data portion, the topic classification associated with the file data portion, or the class or a combination thereof associated with the file data portion.

[0072] In step 625, the cloud server application trains an ensemble learning model based on the NLP and multi-class classification applied in steps 605-615. In accordance with step 625, the cloud server application updates the ensemble learning model based on elements obtained as a result of applying the model, or determined, or both. In one embodiment, the cloud server application stores one or more topic context data elements based on the topic context data derived in step 605 for future analysis of the above user, or a user having similar data usage characteristics and / or application usage characteristics, or both. According to such an embodiment, the cloud server application optionally analyzes one or more topic context data elements and, based on such analysis, identifies and stores topic context data patterns (e.g., topic patterns based on one or more of several user context factors). In any embodiment, the cloud server application utilizes the stored topic context data patterns to facilitate future topic context data processing. In one embodiment, the cloud server application stores one or more file data elements based on the multiple file data portions derived in step 610 to facilitate the future organization of file data portions related to the above-mentioned user, or to a user having similar data usage characteristics or application usage characteristics, or both. In such additional embodiments, the cloud server application optionally analyzes one or more file data elements and, based on such analysis, identifies and stores file data patterns (e.g., topic patterns between file data or other file data patterns). The cloud server application optionally utilizes the stored file data patterns to facilitate future file data processing. In further embodiments, the cloud server application stores one or more classification elements determined in step 615 with respect to the multiple file data portions derived.In such further embodiments, the cloud server application optionally analyzes one or more classification elements and, based on such analysis, identifies and stores classification patterns that associate one or more classes with each of a plurality of file data parts, subparts thereof, or both. Specifically, the cloud server application optionally stores encoded feature vectors that associate numerical class values ​​with each of a plurality of file data parts, subparts thereof, or both. The cloud server application optionally uses the stored classification patterns to facilitate future multi-class classification processing. In further embodiments, the cloud server application trains an ensemble learning model based on the ensemble learning metadata created, or based on patterns determined in relation to method 600, or both, and optionally stores such ensemble learning metadata, patterns, or both in a knowledge base related to the ensemble learning model.

[0073] In summary, applying an ensemble learning model according to Method 600 includes: deriving topic context data by applying NLP to user-related context information; deriving multiple file data portions by applying NLP to file data based on the topic context data; applying at least one multi-class classification technique to the multiple file data portions; determining a distributed plan for the multiple file data portions based on the application of the at least one multi-class classification technique; and training an ensemble learning model based on the applied NLP and the applied at least one multi-class classification technique.

[0074] Figure 7 shows a method 700 for deriving topic context data. Method 700 provides a substep in step 605 of Method 600 according to one or more embodiments. Method 700 first creates multiple coded feature vectors in step 705 by applying at least one NLP model to raw data related to a user. According to step 705, the cloud server application applies at least one NLP model to raw data related to a user in order to create multiple coded feature vectors. The cloud server application creates multiple coded feature vectors based on topic patterns identified in the raw data. Raw data related to a user may optionally include unstructured data or structured data from user context information that has not been NLP processed. In one embodiment, in connection with applying at least one NLP model in step 705, the cloud server application applies a named entity recognition (NER) model to the user context elements. Such an NER model processes user context information by identifying and classifying entities in the raw text of the raw data based on topics. The application of the NER model yields at least one encoded topic assignment vector. In an additional embodiment, in connection with applying at least one NLP model, the cloud server application applies latent Dirichlet allocation (LDA) to user context elements to identify topics and relationships within the text. According to such an additional embodiment, the cloud server application optionally applies LDA modeling to the raw text of the raw data to topic classify data points represented within the raw text.The application of the LDA model yields at least one encoded topic assignment vector. In a further embodiment, in connection with the application of at least one NLP model, the cloud server application applies Bidirectional Encoder Representations from Transformers (BERT) to user context elements. When applying BERT, the cloud server application applies transformer-based NLP utilizing neural network attention techniques to the raw text of the raw data. By utilizing neural network attention, the cloud server application enhances focus on important elements of the raw text input in order to perform sentence embedding with respect to the raw text input. The application of the BERT model yields at least one encoded real number vector based on the sentence embedding. In connection with various embodiments, sentence embedding encompasses a set of techniques for mapping sentences to real number vectors.

[0075] In a further embodiment, relating to the application of at least one NLP model, the cloud server application applies a recurrent neural network (RNN) model to user context elements to establish machine learning (deep learning) based connections between consecutive data points related to user context information. The application of the RNN model causes at least one encoded feature vector to reflect a combination of topic and time series. According to such a further embodiment, the cloud server application optionally utilizes a long short-term memory recurrent neural network (LSTM-RNN) architecture configured to store time-series pattern data with respect to text elements related to raw data or more general user context information. The cloud server application optionally applies an LSTM-RNN model for the purpose of storing time-series pattern data. In relation to the LSTM-RNN model, the cloud server application stores time-dependent usage characteristics for each user context element that is available as input to one or more of the at least one NLP models. Based on usage characteristics of user context elements over time, the cloud server application, in any manner, uses LSTM-RNN data to determine patterns between user context elements. The patterns determined between user context elements, in any manner, reflect the user usage over time of one or more previously used (e.g., most recently used, or most recently used within a given period) applications, data types, or system resources.Specifically, using LSTM-RNN modeling, the cloud server application, in any embodiment, derives at least one timestamped pattern relating to one or more user context elements, and thus, based on the timestamp, identifies captured data usage patterns, captured application usage patterns, or captured system resource usage patterns, or a combination thereof. The application of the LSTM model yields encoded time-series pattern vectors, and the application of the LSTM-RNN modeling further yields encoded feature vector information that reflects both topic and time series. In a further embodiment, the cloud server application applies NLP in conjunction with a gated recurrent unit (GRU) architecture to user context elements in the raw data or more general user context information. The application of GRUs yields encoded time-series pattern vectors based on fewer parameters than LSTM. In any embodiment, relating to the application of at least one NLP model, the cloud server application combines one or more features from the one or more models described above.

[0076] In a further embodiment, the cloud server application assigns weight values ​​to each of the multiple encoded feature vectors to specify the relative importance of the feature vectors created by applying each of the NLP models from at least one NLP model in step 705. For example, the cloud server application may assign relatively high weight values ​​to the encoded feature vectors associated with NEP and relatively low weight values ​​to the encoded feature vectors associated with LDA to indicate that the application of the NEP model is relatively more important than the application of the LDA model in a particular scenario. In a further embodiment, the cloud server application creates multiple encoded feature vectors by deriving data representations related to the application type or data type accessed by the user, the frequency of application or data access by the user, the amount of computing resources used related to application or data access by the user, or memory patterns or combinations thereof related to application or data access by the user. A method for creating multiple encoded feature vectors according to step 705 is illustrated with reference to Figure 8.

[0077] In step 710, the cloud server application obtains a numerical topical output by applying at least one clustering algorithm to the multiple encoded feature vectors created in step 705. According to step 710, the cloud server application applies at least one clustering algorithm to the multiple encoded feature vectors in order to obtain a numerical topical output. The multiple encoded feature vectors, or their elements, created in step 705 are inputs to at least one clustering algorithm. In one embodiment, the cloud server application concatenates one or more of the multiple encoded feature vectors before applying at least one clustering algorithm. In an additional embodiment, the cloud server application associates each feature of one of the multiple encoded feature vectors with its respective topic based on the numerical topical output obtained for each feature as a result of applying at least one clustering algorithm. For example, a cloud server application can associate each feature that yields a numerical topic output value "1" with a security-related topic, each feature that yields a numerical topic output value "2" with a client engagement-related topic, and each feature that yields a numerical topic output value "3" with a quantum computing-related topic. According to such additional embodiments, the cloud server application associates data points from the raw data related to the user with one or more respective topics based on the topic associations of data point features represented in an encoded feature vector representing data points, such as the one created in step 705.For example, a cloud server application can associate a data point from the raw data with security-related topics and client engagement-related topics based on the fact that the data point has features related to security-related topics and client engagement-related topics (this is determined by the respective numerical topic output values ​​obtained for such features). According to such an additional embodiment, the cloud server application derives a numerical topic output value for a data point from the raw data in accordance with step 710 by concatenating or otherwise processing the respective numerical topic output values ​​for each feature of the data point, in any manner. The numerical topic output value for such a data point is represented in any manner as a feature vector, i.e., a clustering-derived supplement to the encoded feature vector representing the data point created in step 705. In a further embodiment, the cloud server application encodes each value of the numerical topic output in any manner in binary format or one-hot encoding format. In a further embodiment, at least one clustering algorithm applied in connection with step 710 includes a k-means clustering algorithm. In addition to or instead of this, at least one clustering algorithm includes an expectation-maximization algorithm. In addition to or instead of this, at least one clustering algorithm includes a hierarchical clustering algorithm, such as agglomerative hierarchical clustering. In addition to or instead of this, at least one clustering algorithm includes a mean-shift clustering algorithm. Following the coding and clustering in steps 705-710, the cloud server application associates features from user context information with numerical topic information.Cloud server applications can then use this kind of numerical topic information to process and classify file data received from users.

[0078] In summary, deriving topic context data according to Method 700 includes creating multiple coded feature vectors by applying at least one NLP model to user-related raw data, and obtaining a numerical topic output by applying at least one clustering algorithm to these multiple coded feature vectors.

[0079] Figure 8 shows a method 800 for creating multiple encoded feature vectors. Method 800 provides a substep in step 705 of Method 700 according to one or more embodiments. Method 800 first derives, in step 805, a cloud server application derives at least one data representation based on the application type or data type accessed by the user. In one embodiment, the encoded vector representation includes a set of encoded feature vectors. The set of encoded feature vectors includes one or more encoded feature vectors. In any embodiment, the cloud server application derives the set of encoded feature vectors by deriving a binary vector (e.g., a bit array) based on the application access history via one-hot coding. According to such embodiments, one-hot coding enables the representation of categorical variables in binary vector form. In addition to or instead of this, the cloud server application derives the set of encoded feature vectors via word embedding. In relation to various embodiments, word embedding refers to assigning numerical values ​​to text words based on context such that the numerical values ​​of relatively related words are closer to each other than the numerical values ​​of relatively unrelated words. In a related embodiment, a cloud server application applies word embedding to derive a set of coded features representing relatively related words that are closer to each other in a real-valued vector. In a related additional embodiment, a cloud server application applies word embedding to derive a set of coded features representing relatively related data types that are closer to each other in a real-valued vector, relating to user context information. In a further embodiment, a cloud server application associates data points with data types in attribute-value pair format.

[0080] In step 810, the cloud server application derives at least one data representation relating to the frequency of user access to the application or data. In one embodiment, the cloud server application encodes user access to a particular application or a particular data point in attribute-value pair format. Such attribute-value pair format can associate a particular application / data point with a metadata color value or vector value that describes the frequency of application / data usage, the date / time of application / data usage, or other elements or combinations thereof of application / data usage. In a further embodiment, the cloud server application encodes user access to a particular data type in attribute-value pair format. Such attribute-value pair format can associate a particular data type with a metadata vector value that describes the frequency of data type usage, the date / time of data type usage, or other elements or combinations thereof of data type usage.

[0081] In step 815, the cloud server application derives at least one data representation relating to computing resource usage related to user application access or data access, or to memory patterns related to user application access or data access, or both. In one embodiment, the cloud server application collects information relating to historical memory patterns related to user data access. For example, the cloud server may record memory patterns that reflect the memory of frequently used user application data across all storage layers. In any embodiment, the cloud server application stores such resource usage information or memory pattern information, or both, in a standardized encoding format, such as binary vector format or attribute-value pair format or both. The cloud server application incorporates the data representations derived in steps 805-815, or elements thereof, into one or more of a plurality of encoded feature vectors. According to various embodiments, the cloud server application, in any embodiment, performs a subset of steps 805-815, or in any embodiment, performs steps 805-815 in any order, or both.

[0082] In summary, creating multiple coded feature vectors according to Method 800 includes: deriving at least one data representation based on the application type or data type accessed by the user; deriving at least one data representation related to the frequency of the user's application access or data access; and deriving at least one data representation related to the computing resource usage or storage patterns associated with the user's application access or data access.

[0083] Figure 9 shows a method 900 for deriving multiple file data portions. Method 900 provides a substep in step 610 of Method 600 according to one or more embodiments. Method 900 first identifies topic patterns between data points in file data by a cloud server application in step 905 by applying at least one NLP model considering the topic context data derived according to step 605. According to step 905, the cloud server application applies at least one NLP model considering the topic context data to identify topic patterns. In one embodiment, the cloud server application derives topic information related to the file data by considering the topic context data. According to such an embodiment, the cloud server application identifies file data portions by associating data points in the file data with topics or groups of topics in the topic context data, in any manner. Specifically, based on the topic context data, the cloud server application applies at least one NLP model to the file data to determine topic patterns between data points in the file data. In a relevant embodiment, the cloud server application associates each feature of the data points in the file data with its respective topic based on the identified topic patterns. For example, a cloud server application may associate a feature with a security-related topic if it has identified a pattern between the features of a data point in a file data file and the corresponding features in topic context data related to that security-related topic. According to such a related embodiment, the cloud server application may, in any manner, associate data points in a file data file with a topic based on the topic association of the features of such data points.For example, based on identified topic patterns, if it is determined that the characteristics of the data points in each file data are associated with security-related topics and client engagement-related topics, the cloud server application can associate the data points in each file data with security-related topics and client engagement-related topics.

[0084] In one embodiment, in connection with applying at least one NLP model in step 905, the cloud server application applies an NER model to the file data elements of the file data. Such an NER model extracts information from the file data by identifying and classifying entities based on topics, taking topic context data into consideration. Based on topic identification via the NER model, the cloud server application, in any embodiment, identifies all data points or subsets of data points in the topic context data related to the file data. In a further embodiment, in connection with applying at least one NLP model, the cloud server application applies LDA to identify topics and relationships between texts in the file data. According to such a further embodiment, the cloud server application, in any embodiment, applies LDA modeling to the file data elements, taking topic context data into consideration, in order to topic classify the data points represented in the file data. In a further embodiment, in connection with applying at least one NLP model, the cloud server application applies BERT to the file data elements, taking topic context data into consideration. When applying BERT, the cloud server application applies transformer-based NLP that utilizes neural network attention techniques to the texts in the file data. By utilizing neural network attention, cloud server applications enhance focus on key elements of raw text input in order to perform sentence embedding with respect to raw text input.

[0085] In a further embodiment, in connection with applying at least one NLP model, the cloud server application applies an RNN model to establish machine learning-based connections between consecutive data points related to file data, taking topic context data into consideration. According to such a further embodiment, the cloud server application optionally utilizes an LSTM-RNN architecture configured to store time-series pattern data with respect to file data elements. The cloud server application optionally applies an LSTM-RNN modeling for the purpose of storing time-series pattern data. In connection with the LSTM-RNN modeling, the cloud server application stores time-dependent usage characteristics for each file data element that is available as input to one or more of the at least one NLP models. Based on the time-dependent usage characteristics of the file data elements, the cloud server application optionally uses the LSTM-RNN data to predict patterns between file data elements (e.g., relationships determined over time between file data patterns and user usage patterns). Specifically, using the LSTM-RNN modeling, the cloud server application optionally derives at least one timestamped pattern with respect to one or more of the file data elements. A cloud server application may, in any manner, identify a relationship between one or more file data elements and one or more previously used applications, data points, or system resources, and thus associate such one or more file data elements with a recorded usage pattern over time with respect to one or more previously used applications, data points, or system resources. Thus, in connection with deriving file data portions, a cloud server application may associate such one or more file data elements with one or more file data portions based at least in part on such usage patterns.For example, depending on whether a file data element is associated with an application previously used by the user (e.g., most recently used or used within a given time), the cloud server application may associate such file data element with one or more portions of file data that reflect compatibility or other associations with such previously used applications. In a further embodiment, the cloud server application applies NLP to the file data element in conjunction with the GRU architecture, taking topic context data into consideration. In any manner, the cloud server application combines one or more features of the models described above when applying NLP techniques. Based on applying at least one NLP model taking topic context data into consideration, the cloud server application associates each data point in the file data with one or more identified topic patterns.

[0086] In step 910, the cloud server application assigns data points in the file data to multiple file data parts by applying at least one clustering algorithm based on the topic patterns identified in step 905. According to step 910, the cloud server application applies at least one clustering algorithm based on the identified topic patterns to assign data points in the file data to multiple file data parts. Based on the identified topic patterns associated with the data points in the file data (in any embodiment, represented in the form of coded feature vectors), the cloud server application applies at least one clustering algorithm to assign the data points to multiple file data parts based on topic pattern correlations between the data points. In any embodiment, the cloud server application organizes each file data part such that each data point in each file data part has a specific topic pattern correlation determined via at least one clustering algorithm. In one embodiment, the at least one clustering algorithm applied in connection with step 910 includes a k-means clustering algorithm. In addition to or instead of this, the at least one clustering algorithm includes an expectation maximization algorithm. In addition to or instead of this, at least one clustering algorithm includes a hierarchical clustering algorithm, e.g., condensed hierarchical clustering. In addition to or instead of this, at least one clustering algorithm includes a mean-shift clustering algorithm. In a further embodiment, the cloud server application assigns a numerical value (e.g., an integer value, a binary value, or a one-hot encoded value) to each of the multiple file data portions. In any embodiment, the cloud server application assigns a numerical value to each of the multiple file data portions based on the respective numerical values ​​associated with the corresponding elements in the topic context data derived in step 605.For example, based on the example described above with respect to step 710, the cloud server application may assign the number "1" to a portion of the file data containing data points and characteristics related to security, the number "2" to a portion of the file data containing data points and characteristics related to client engagement, and the number "3" to a portion of the file data containing data points and characteristics related to quantum computing. Depending on whether each portion of the file data is determined to contain data points and characteristics related to multiple topics, the cloud server application may, in any manner, assign multiple numbers corresponding to multiple topics to such a portion of the file data, or assign a single number corresponding to the most relevant topic among the multiple topics to such a portion of the file data.

[0087] In summary, deriving multiple file data portions according to Method 900 includes identifying topic patterns between data points in the file data by applying at least one NLP model considering topic context data, and assigning those data points in the file data to multiple file data portions by applying at least one clustering algorithm based on the identified topic patterns.

[0088] Figure 10 shows a method 1000 for partitioning file data. Method 1000 provides a substep in step 520 of Method 500 according to one or more embodiments. Method 1000 first involves a cloud server application storing separate portions of file data between the cloud computing layer, fog computing layer, and local computing layer of a multi-tier storage infrastructure in step 1005. In relation to various embodiments, storing separate portions means storing non-redundant parts. In one embodiment, the cloud server application stores the partitioned data between storage layers in buckets in each storage layer. According to such an embodiment, the cloud server application stores cloud data in cloud server buckets in the cloud layer, fog data in fog server buckets in the fog layer, and local data in local machine buckets in the local layer. In any embodiment, the cloud server application indexes each bucket based on a bucket key value so that portions or sub-portions of file data having the same bucket key value are stored in a single bucket. The bucket key value of each bucket corresponds to, in any manner, a class value assigned to a portion or sub-portion of the file data, or is otherwise related. The cloud server application, in any manner, specifies or retrieves access control data for each bucket. Bucket access control may also be specified for the bucket via its respective access control list (ACL). In a further embodiment, the cloud server bucket is the primary bucket, i.e., where the cloud server application begins partitioning the file data according to the data distribution technique devised in step 510.According to such further embodiments, the cloud server application performs partitions on the file data located in the primary bucket in order to store separate portions of the file data in buckets corresponding to other storage tiers. The cloud server application may, in any manner, perform each partition sequentially or simultaneously. In an alternative embodiment, in step 1005, the cloud server application stores separate portions of the file data in one or more intermediate tiers of the multi-tier storage infrastructure, in addition to the cloud tier, fog tier, and local tier. According to such an alternative embodiment, the cloud server application may, in any manner, store the data classified to be stored in an intermediate tier in an intermediate bucket associated with that intermediate tier.

[0089] In step 1010, the cloud server application stores redundant dependencies between the cloud computing layer, fog computing layer, and local computing layer of the multi-tier storage infrastructure. The redundant dependencies stored between layers include libraries related to the execution of file data or parts / subparts thereof. In addition to or instead of this, the redundant dependencies include application packages related to file data or parts / subparts thereof. In addition to or instead of this, the redundant dependencies include information linking each file data part or subpart to facilitate, for example, the recovery or transfer of file data. In an alternative embodiment, in step 1010, in addition to the cloud computing layer, fog computing layer, and local computing layer, the cloud server application stores redundant dependencies between one or more intermediate layers of the multi-tier storage infrastructure.

[0090] In summary, splitting file data according to Method 1000 includes storing separate portions of the file data between the cloud computing layer, the fog computing layer, and the local computing layer, and storing redundant dependencies between each of the cloud computing layer, the fog computing layer, and the local computing layer.

[0091] Figure 11 shows a method 1100 for authenticating a data access request. Method 1100 provides a substep in step 530 of Method 500 according to one or more embodiments. Method 1100 first involves a cloud server application parsing an identity and password from a data access request in step 1105. In step 1110, the cloud server application fetches a stored hash value pre-associated with the user based on the identity information parsed in step 1105. In one embodiment, the cloud server application retrieves the stored hash value by coordinating with a directory server system associated with the managed service domain. In step 1115, the cloud server application hashes the password parsed in step 1105 by applying one or more cryptographic hash functions corresponding to the stored hash value to the parsed password. In one embodiment, the cloud server application communicates with the directory server system to facilitate the application of one or more cryptographic hash functions. In step 1120, the cloud server application determines whether the result of the password hash matches the stored hash value. If the cloud server application determines that the password hash result matches the stored hash value, it completes the authentication of the data access request in step 1125. Conversely, if the cloud server application determines that the password hash result does not match the stored hash value, it proceeds to the end of method 1100 without completing the authentication of the data access request.

[0092] To summarize, authenticating a data access request according to Method 1100 includes: parsing the identification information and password from the data access request; fetching a stored hash value pre-associated with the user based on the parsed identification information; hashing the parsed password by applying one or more cryptographic hash functions corresponding to the stored hash value to the parsed password; and completing the authentication of the data access request if it is determined that the result of the password hash matches the stored hash value.

[0093] While various embodiments of the present invention have been described as examples, they are not intended to be exhaustive or limit the scope to the disclosed embodiments. Any kind of modification and equivalent configuration of the described embodiments should be included within the scope of protection of the present invention. Accordingly, the scope of the present invention should be most broadly described in relation to modes for carrying out the invention, according to the claims that follow, and should encompass all possible equivalent modifications and equivalent configurations. As will be apparent to those skilled in the art, many modifications and variations are possible without departing from the scope of each described embodiment. The terms used herein have been selected to best describe the principles, practical applications, or technical improvements to the art found in the market of each embodiment, or to enable other those skilled in the art to understand each embodiment described herein.

Claims

1. To store user-related file data in the managed service domain, Applying an ensemble learning model to identify topics, groups of topics, or both related to the file data based on contextual information associated with the user, and classifying the file data based on the identification results, Encrypting the classified file data, The file data is divided to be stored across the cloud computing layer, the fog computing layer, and the local computing layer by performing a hash transformation and applying at least one cyclic error correction code. Computer implementation methods, including those mentioned above.

2. Receiving data access requests related to the aforementioned file data, Authenticating the aforementioned data access request, The process of recovering the aforementioned file data through decryption, The computer implementation method according to claim 1, further comprising:

3. Applying the aforementioned ensemble learning model means The computer implementation method according to claim 1, comprising deriving topic context data by applying natural language processing (NLP) to the context information related to the user.

4. Applying the aforementioned ensemble learning model means The computer implementation method according to claim 3, further comprising deriving a plurality of file data portions by applying NLP to the file data based on the topic context data.

5. Applying the aforementioned ensemble learning model means The computer implementation method according to claim 4, further comprising applying at least one multi-class classification technique to the plurality of file data portions.

6. Applying the aforementioned ensemble learning model means The computer implementation method according to claim 5, further comprising determining a distribution plan for the plurality of file data portions based on the application of at least one multi-class classification technique.

7. The deriving of the aforementioned topic context data is: By applying at least one NLP model to the raw data associated with the user, multiple coded feature vectors are created. Numerical topic output is obtained by applying at least one clustering algorithm to the plurality of encoded feature vectors. The computer implementation method according to claim 3, including the method described in claim 3.

8. Creating the aforementioned multiple coded feature vectors is, The computer implementation method according to claim 7, comprising deriving at least one data representation based on the application type or data type accessed by the user.

9. Creating the aforementioned multiple coded feature vectors is, To derive at least one data representation related to the frequency of application access or data access by the user, To derive at least one data representation related to the usage of computing resources or storage patterns related to application access or data access by the user, The computer implementation method according to claim 7, including the method described in claim 7.

10. The computer implementation method according to claim 7, wherein each feature of the plurality of encoded feature vectors is associated with a topic based on the numerical topic output obtained for each feature.

11. The process of deriving the aforementioned multiple file data portions is as follows: By applying at least one NLP model considering the aforementioned topic context data, the topic patterns between data points in the file data are identified. Assigning the data points in the file data to the plurality of file data portions by applying at least one clustering algorithm based on the identified topic pattern, The computer implementation method according to claim 4, including the method described in claim 4.

12. Splitting the aforementioned file data is The computer implementation method according to claim 1, comprising storing separate portions of the file data between the cloud computing layer, the fog computing layer, and the local computing layer.

13. Splitting the aforementioned file data is The computer implementation method according to claim 1, further comprising storing redundant dependencies between the cloud computing layer, the fog computing layer, and the local computing layer, respectively.

14. Authenticating the aforementioned data access request is Analyzing the identification information and password from the aforementioned data access request, Based on the analyzed identification information, the stored hash value pre-associated with the user is fetched, The process involves hashing the analyzed password by applying one or more cryptographic hash functions corresponding to the stored hash value to the analyzed password. In response to determining that the hash result of the password matches the stored hash value, the authentication of the data access request is completed. The computer implementation method according to claim 2, including the method described in claim 2.

15. A computer program including program instructions, wherein the program instructions are executable by a computing device, and the computing device, To store user-related file data in the managed service domain, Applying an ensemble learning model to identify topics, groups of topics, or both related to the file data based on contextual information associated with the user, and classifying the file data based on the identification results, Encrypting the classified file data, The file data is divided to be stored across the cloud computing layer, the fog computing layer, and the local computing layer by performing a hash transformation and applying at least one cyclic error correction code. A computer program that executes something.

16. The program instructions are sent to the computing device: Receiving data access requests related to the aforementioned file data, Authenticating the aforementioned data access request, The process of recovering the aforementioned file data through decryption, The computer program according to claim 15, which further performs the following.

17. Applying the aforementioned ensemble learning model means The computer program according to claim 15, comprising deriving topic context data by applying NLP to the context information related to the user.

18. At least one processor, A system comprising memory storing an application program, wherein the application program, when executed on at least one processor, To store user-related file data in the managed service domain, Applying an ensemble learning model to identify topics, groups of topics, or both related to the file data based on contextual information associated with the user, and classifying the file data based on the identification results, Encrypting the classified file data, The file data is divided to be stored across the cloud computing layer, the fog computing layer, and the local computing layer by performing a hash transformation and applying at least one cyclic error correction code. A system that performs actions including those mentioned above.

19. The aforementioned operation is, Receiving data access requests related to the aforementioned file data, Authenticating the aforementioned data access request, The process of recovering the aforementioned file data through decryption, The system according to claim 18, further comprising:

20. Applying the aforementioned ensemble learning model means The system according to claim 18, comprising deriving topic context data by applying NLP to the context information related to the user.

Citation Information

Patent Citations

  • Method for enhancing data migration security in cloud storage

    CN112764677A

  • Distributed data archive system

    JP2006012192A

  • Information processing apparatus, information processing method and distributed processing system

    JP2020034982A

  • Information processing device, information processing method and information processing program, and terminal

    JP2020123006A

  • Matrix-based Error Correction and Erasure Code Methods and Apparatus and Applications Thereof

    US20100218037A1