Hash-based allocation of applications to virtualized computing environments
By organizing virtualized computing environments into a ring structure based on application properties and caching data, the method addresses long loading times and uneven resource allocation, enhancing user experience and efficiency in cloud computing platforms.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2026-03-12
AI Technical Summary
In cloud computing environments, the process of loading software applications is time-consuming, leading to degraded user experience due to long wait times before applications are ready for use, and existing methods of assigning virtualized computing environments either result in increased power consumption or uneven wear leveling.
A method is introduced where virtualized computing environments are organized into a ring structure based on application properties, with more popular applications allocated more resources and data cached in these environments, allowing instances to be launched quickly from local caches.
This approach reduces loading times, balances cache sizes and bandwidth usage, and evenly distributes applications, thereby improving user experience and reducing power consumption and wear leveling.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
TECHNICAL AREA
[0001] At least one embodiment relates to the allocation of applications to virtualized computing environments. For example, at least one embodiment relates to processors or computing systems used to provide and enable hash-based allocation of applications to virtualized computing environments based on application properties, according to various novel techniques described herein. BACKGROUND
[0002] In a cloud computing environment, a user can access and stream software applications (such as games) on their local client device via an application hosting platform. During a startup sequence, the software application may load multiple files, including assets, textures, graphics, user-associated data, and so on. Loading all the files associated with the software application can be time-consuming. An application hosting platform can load the software application upon a user request. However, such techniques can result in users having to wait a considerable amount of time before the software application is ready to stream. This can degrade the overall user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0003] Various embodiments according to the present disclosure are described with reference to the drawings. The drawings show: Fig. 1 A block diagram of an exemplary cloud computing environment architecture, according to at least one embodiment. Fig. 2 a block representation of a system that enables hash-based application-to-server mapping in a cloud environment, according to at least one embodiment. Fig. 3 an exemplary ring-based allocation of applications to virtualized computing environments using application properties, according to at least one embodiment. Fig. 4 An exemplary ring-based allocation of applications to virtualized computing environments using the availability of a virtual computing environment, according to at least one embodiment. Fig. 5 An exemplary allocation of applications to virtualized computing environments using multiple rings, according to at least one embodiment. Fig. 6 A flowchart of an exemplary procedure for allocating applications to virtualized computing environments, according to at least one embodiment. Fig. 7A an inference and / or training logic, according to at least one embodiment. Fig. 7B an inference and / or training logic, according to at least one embodiment. Fig. 8 an exemplary data center system, according to at least one embodiment. Fig. 9 a computer system, according to at least one embodiment. Fig. 10 a computer system, according to at least one embodiment. Fig. 11 at least parts of a graphics processor, according to at least one embodiment; Fig. 12 at least parts of a graphics processor, according to at least one embodiment; Fig. 13 An exemplary data flow diagram for an advanced computing pipeline, according to at least one embodiment. Fig. 14 A system diagram for an exemplary system for training, adapting, instantiating and deploying machine learning models in an advanced computing pipeline, according to at least one embodiment. Fig. 15A and Fig. 15B a data flow diagram for a process for training a machine learning model and a client-server architecture for improving annotation tools with pre-trained annotation models, according to at least one embodiment. DETAILED DESCRIPTION
[0004] Applications can be hosted in a cloud computing environment (e.g., a cloud-based gaming platform) by an application hosting platform. An instance of an application hosted by the application hosting platform can be delivered to a client device via a virtualized computing environment (e.g., a virtual machine (VM) or a container). A user can access the application (e.g., a game application, a collaborative content creation application, a streaming application, a multimedia application, an entertainment application, an interactive application (which may contain one or more of these other types of applications), or any other type of application) on a client device (e.g., a computer, a game console, a mobile phone, or a smartphone) through the application hosting platform. For example, a user can access the application hosting platform (e.g.,(log in) and select a game application for access through their client device. The application hosting platform can communicate with the game application via a number of application programming interfaces (APIs) and can begin loading an instance of the game application into the virtualized computing environment for access by the user device. For example, the application hosting platform can launch the game application from a command line within the virtualized computing environment when the user selects the game application. To launch the game application within the virtualized computing environment, the application hosting platform can begin by loading the game application files into the virtualized computing environment for the user.
[0005] An instance of the game application can be allocated to the virtualized computing environment in response to the user selecting the game application on the client device connected to the application hosting platform. The application hosting platform may require a significant amount of time to download the game application instance data from a storage server to the virtualized computing environment. For example, it may take several seconds or several minutes for the application hosting platform to download generic and user data associated with the game application to the virtualized computing environment. During this time, the user may not be able to access the game application and may see a loading screen while the virtualized computing environment prepares for user interaction with the game application.After a few seconds or several minutes, the virtualized computing environment may be ready, and the user can access the game application and the main menu. However, the overall user experience interacting with the game application instance can be negatively impacted by the long wait time between selecting the game application and gaining access to it.
[0006] In some examples, a game application instance can be randomly assigned to a virtualized computing environment (e.g., a VM) from a pool of available virtualized computing environments (e.g., available VMs). However, randomly assigning the game application instance to a virtualized computing environment from the pool of available virtualized computing environments can require all available virtualized computing environments to remain online for extended periods. This can lead to increased power consumption in the cloud computing platform. In some other examples, the game application instance can be assigned to a specific virtualized computing environment. For example, each instance of the game application can be assigned to the same virtualized computing environment every time the game application is selected.Assigning every instance of a game application to the same virtualized compute environment can lead to increased wear leveling within those environments. For example, virtualized compute environments running instances of more popular games may experience faster wear than those running instances of less popular games. In some cases, different applications may load the CPU, GPU, or other memory in different ways. Therefore, having a single application running in a virtualized compute environment can lead to increased wear and tear within that environment.
[0007] Aspects of the present disclosure address the aforementioned and other deficiencies by providing methods and systems that allocate virtualized computing environments to applications (e.g., game applications, a collaborative content creation application, a streaming application, a multimedia application, an entertainment application, an interactive application (which may contain one or more of these other types of applications), or any other type of application) and store application data in respective caches of the virtualized computing environments. For example, multiple game applications provided by an application hosting platform can be identified, and for each game application, one or more virtualized computing environments (e.g.,Virtual machines or containers are allocated, and application data for the game application can be stored in the caches of one or more virtualized computing environments. In some examples, multiple virtualized computing environments can be associated with a unique identifier for a particular game application. When the game application is selected by a user associated with a client device connected to the application hosting platform, an instance of the game application can run in a virtualized computing environment associated with the unique identifier and use application data stored in the virtualized computing environment's cache.
[0008] In some examples, a deployment manager of a cloud computing platform might organize some or all of the virtualized compute environments into a ring. Furthermore, the deployment manager might identify a number of game applications that can run on the cloud computing platform and organize these applications within the ring. The deployment manager can organize the applications based on one or more properties of the applications. For example, more popular applications might be assigned a larger segment of the ring than less popular applications. Therefore, more popular applications might be allocated a greater number of virtualized compute environments than less popular applications.Furthermore, the deployment manager can store application data associated with game applications in respective caches of the virtualized computing environments, based on the locations of the game applications and the virtualized computing environments within the ring. When the user selects a game application, an instance of the game application can be executed in one of the virtualized computing environments, again based on the locations of the game applications and the virtualized computing environments within the ring.
[0009] In one example, the application hosting platform can host four game applications, each with a unique identifier. Additionally, the cloud computing platform can provide twenty virtualized computing environments capable of running instances of the four game applications. These virtualized computing environments can generally be evenly distributed along the ring. The deployment manager can organize the game applications within the ring based on one or more properties of the game applications, such as their popularity. For example, the deployment manager might allocate thirty percent of the ring to a first game application, forty percent to a second, ten percent to a third, and twenty percent to a fourth.Furthermore, the deployment manager can assign virtualized computing environments to the game applications based on their respective locations and the virtualized computing environments within the ring. For example, the deployment manager can assign six virtualized computing environments to the first application, eight to the second, two to the third, and four to the fourth. Additionally, the deployment manager can store data associated with the game applications in the respective caches of the virtualized computing environments.The deployment manager can, for example, store data associated with the first game application in the caches of the corresponding six virtualized computing environments, data associated with the second game application in the caches of the corresponding eight virtualized computing environments, data associated with the third game application in the caches of the corresponding two virtualized computing environments, and data associated with the fourth game application in the caches of the corresponding four virtualized computing environments. When a user selects one of the game applications, an instance of the game application can be run in the corresponding virtualized computing environment using the data stored in that environment's cache.
[0010] In some examples, the deployment manager can enable or disable virtualized computing environments based on one or more parameters, such as a time parameter. For instance, the deployment manager might determine that game applications run more frequently during a certain time of day (such as in the evening) than during a different time of day (such as in the morning). The deployment manager can enable most or all of the virtualized computing environments during the first time of day to allow game applications to run faster. Additionally or alternatively, the deployment manager can disable some of the virtualized computing environments during the second time of day to reduce power consumption.In some examples, the deployment manager can reassign a virtualized compute environment from a first game application to a second game application based on one or more conditions. For example, based on the detection that a second game application is becoming more popular than the first, the deployment manager can reassign a unique identifier for the second game application from one virtualized compute environment to another. Additionally, the deployment manager can delete application data associated with the second game application from the cache of the virtualized compute environment and store application data associated with the second game application in a cache of the other virtualized compute environment. In some examples, each virtualized compute environment's cache can store data for a single game application.In some other examples, the cache of at least one virtualized computing environment can store data for multiple game applications. This can allow the virtualized computing environment to efficiently host each of the multiple game applications.
[0011] Some benefits of this disclosure include reduced loading times for game applications on a cloud computing platform. For example, when a user launches a game application, an instance of the game application can be run in a virtualized computing environment using data stored in a cache within that environment. Faster loading of game applications improves the user experience. Other benefits of this disclosure include improved balancing of cache sizes and bandwidth usage.For example, virtualized computing environments with larger caches and lower bandwidth can store more game application data in their caches, while those with smaller caches and higher bandwidth can store less data. This allows more data to be accessed from the cloud computing platform's data store while a game application is running. Further advantages of this disclosure include allocating virtualized computing environments to game applications based on application popularity and enabling and disabling them based on one or more parameters, such as the time of day. This can reduce power consumption on the cloud computing platform.Some advantages of the present disclosure further include the fact that it enables gaming applications to be distributed more evenly among virtualized computing environments, thereby reducing the wear leveling of the virtualized computing environments.
[0012] The invention is defined by the claims. To illustrate the invention, aspects and embodiments that may or may not be within the scope of protection of the claims are described here.
[0013] Devices, systems, and techniques for allocating application hosting platforms in a virtualized computing environment are disclosed. A method may include allocating a first set of virtualized computing environments to a first application based on one or more properties of the first application, and allocating a second set of virtualized computing environments to a second application of the plurality of applications based on one or more properties of the second application, wherein the second set of virtualized computing environments differs from the first set of virtualized computing environments.The procedure may include causing, in response to a request to run the first application, an instance of the first application to be executed in a virtualized computing environment of the first set of virtualized computing environments, using data stored in a cache of the virtualized computing environment.
[0014] The revelation extends to all novel aspects or features described and / or illustrated here.
[0015] Further features of the disclosure are characterized by the independent and dependent claims.
[0016] Any feature of one aspect of the disclosure can be applied in any suitable combination to other aspects of the disclosure. In particular, procedural aspects can be applied to apparatus or system aspects, and vice versa.
[0017] Furthermore, features implemented in hardware can be implemented in software and vice versa. Any reference to software and hardware features herein should be interpreted accordingly.
[0018] Each system or device feature described herein can also be provided as a process feature, and vice versa. System and / or device aspects that are functionally described (including means plus functional features) can alternatively be expressed in terms of their corresponding structure, such as a suitably programmed processor and allocated working memory.
[0019] It is also understood that certain combinations of the various features described and defined in each aspect of the revelation can be implemented and / or provided and / or used independently of one another.
[0020] The disclosure also provides computer programs and computer program products comprising software code designed to perform one of the methods described herein when executed on a data processing device and / or to embody one of the device and system features described herein, including one or all component steps of a method.
[0021] The disclosure also provides a computer or computing system (including networked or distributed systems) with an operating system that supports a computer program for carrying out the procedures described herein and / or for embodying the device or system features described herein.
[0022] The revelation also provides a computer-readable medium on which one or more of the aforementioned computer programs are stored.
[0023] The revelation also provides a signal that transmits one or more of the aforementioned computer programs.
[0024] The disclosure extends to methods and / or devices and / or systems as described herein with reference to the accompanying drawings.
[0025] Aspects and embodiments of the disclosure will now be described purely by way of example with reference to the attached drawings.
[0026] Fig. Figure 1 illustrates a block diagram of an exemplary system architecture 100, according to at least one embodiment. The system architecture 100 (here also referred to as the "system") comprises an application hosting platform 102, an application developer platform 104, a server machine 106, client devices 108A-N (collectively and individually referred to as client device(s) 108) and a data storage device 112, each of which is connected to a network 120. In implementations, the network 120 can include a public network (e.g., the Internet), a private network (e.g., a Local Area Network (LAN) or Wide Area Network (WAN)), a wired network (e.g., an Ethernet network), a wireless network (e.g., an 802.11 network or a Wi-Fi network), a cellular network (e.g., a Long Term Evolution (LTE) network), routers, hubs, switches, server computers, and / or a combination thereof.
[0027] In some implementations, Datastore 112 is persistent storage capable of holding content items as well as data structures for tagging, organizing, and indexing those content items. Datastore 112 can be hosted by one or more storage devices, such as main memory, magnetic or optical storage disks, tapes or hard disks, NAS, SAN, etc. In some implementations, Datastore 112 may be a network-connected file server, while in other implementations, Datastore 112 may be a different type of persistent storage, such as an object-oriented database, a relational database, etc., which can be hosted by Platform 102 or by one or more different machines connected to Platform 102 via Network 120.
[0028] The Application Hosting Platform 102 can be configured to host files of one or more applications (e.g., Application 130A, Application 130B, etc.) provided by an application developer (e.g., via the Application Developer Platform 104). The Application Developer Platform 104 can be used by an application developer (e.g., a user, a company, an organization, etc.). For example, an application developer might be a video game developer creating a video game (represented by an Application 130) that users can interact with on client devices 108. The Application Hosting Platform 102 can provide users with access to an Application 130 (or an instance of an Application 130) provided by the Application Developer Platform 104 through a respective client device 108A-N.For example, the application hosting platform 102 can enable users to use, upload, download, and / or search for applications 130. In at least one embodiment, the application hosting platform 102 can include a website (e.g., one or more web pages) or a client application or component that can be used to provide users with access to applications 130. In at least one embodiment, each application 130 can consist of generic data 132 (e.g., data consisting exclusively of user data) and user data 134, e.g., generic data 132A and user data 134A of application 130A, generic data 132B and user data 134B of application 130B, and so on.In at least one embodiment, the application hosting platform 102 can be an example of a cloud-hosted gaming service platform, a cloud-hosted collaborative content creation platform for heterogeneous content creation applications, a video streaming hosting platform, a test platform for simulated or enhanced content, a machine learning training platform, a machine learning deployment platform, or a video conferencing hosting platform. In at least one embodiment, the application 130 (or instance of the application 130) can be an example of a gaming application, a video conferencing application, a content creation application, a cloud-hosted application, a collaborative content creation application, a cloud-hosted collaborative content creation application, a video streaming application, a machine learning application, or a simulation application.
[0029] In at least one embodiment, the server 106 can host a virtualized computing environment 160 that runs an instance of the application 130. The server 106 can, for example, be a computer system containing one or more physical devices (e.g., a processing device (e.g., a GPU), memory, one or more I / O devices, etc.) and a hypervisor and / or host operating system that manages one or more virtualized computing environments 160. A virtualized computing environment 160 can, for example, correspond to a virtual machine running a guest operating system and one or more guest applications, including an instance of the application 130, or to a container running an application, such as an instance of the application 130. One or more servers 106 can be deployed, and each server 106 can host one or more virtualized computing environments 160.In at least one embodiment, each Server 106 can correspond to Computer System 900 and / or Computer System 1000, which with respect to . Fig. 9 and Fig. 10 are described.
[0030] The application hosting platform 102 can include a deployment manager 140 that assigns applications 130 (e.g., game applications) to specific sets of virtualized computing environments 160. In at least one embodiment, the deployment manager 140 assigns applications 130 to specific sets of virtualized computing environments 160 by organizing some or all of the virtualized computing environments 160 into a ring and organizing applications 130 within the ring based on properties of the applications 130. For example, applications 130 that are more popular can be assigned to a larger segment of the ring than applications 130 that are less popular. Therefore, a larger number of virtualized computing environments 160 can be allocated to the more popular applications 130 than to the less popular applications 130.
[0031] The application hosting platform 102 can include an application load manager 150, which enables the loading of application data (generic data 132 and user data 134) of an application 130 into each cache 162 of a set of virtualized computing environments 160, based on the location of the application 130 and the set of virtualized computing environments 160 in the ring. When the user selects the application 130, an instance of the application 130 can be executed on one of the virtualized computing environments 160, based on the location of the application 130 and the virtualized computing environment 160 in the ring, and using the data stored in the cache 162 of the virtualized computing environment 160. This reduces the time between selecting the application 130 and executing an instance of the application 130, as explained in more detail herein.At least some of the generic data 132 and user data 134 can be displayed via a user interface (UI) on each client device accessing an instance of a respective application 130. The set of generic data 132 and user data 134 for the application 130 can be defined by an application developer via the application developer platform 104. A user of the respective client device 108 can interact with the instance of the application 130 by interacting with an instance of the application 130 loaded with the generic data 132 and user data 134 via a GUI provided by the application hosting platform 102. The GUI of the application hosting platform 102 can be rendered by a client component of the application hosting platform 102 or by a web browser hosted by the client device 108.
[0032] The virtualized computing environment 160 can be instantiated to facilitate the execution of applications 130 for access by client devices 108 and can be deconstructed in response to an event or condition (e.g., in response to a request from a user of the client device 108). For example, when a specific event or condition is detected, the application load manager 150 can transmit a deconstruction request to the server 106, which can cause a hypervisor to deconstruct the virtualized computing environment 160.
[0033] A user of a given client device 108 can interact with the application 130 (e.g., via the application hosting platform GUI), which is loaded with generic data 132 and user data 134, to navigate through the application via the respective client device 108. In an illustrative example, applications 130A and 130B could be video game applications (e.g., game applications) developed by a video game developer. In another illustrative example, applications 130A and 130B could be content creation or asset creation applications on a cloud-hosted collaborative content creation platform. Generic data 132 and user data 134 of a given video game application 130 can be presented on an application hosting platform GUI on a client device 108 for use by a user of the client device 108.In at least one embodiment, the generic data 132 can include assets, textures, graphics of the application 130, background graphics, title screens, main menu screens, memory allocation of the game application 130 on the client device 108, shaders, etc. In at least one embodiment, the user data 134 can include data that the application 130 loads specifically for a particular client device 108 that the game application 130 has selected, e.g., for a specific user. For example, the user data 134 can include specific user settings (e.g., resolution or shader settings), user-defined content (e.g., user customizations to a character created by the user in a game), user-defined key bindings, etc. Accordingly, a user can customize the experience of running the game application 130 on the respective client device 108.
[0034] The client devices 108 can include, but are not limited to, televisions, smartphones, mobile phones, personal digital assistants (PDAs), portable media players, netbooks, laptop computers, e-readers, tablet computers, desktop computers, set-top boxes, game consoles, and the like. As discussed above, each client device 108 can include a client component of the application hosting platform 102 (or a web browser) that provides a GUI enabling a user of the client device 108 to request the execution of the application 130. The GUI can provide a rendered version of the generic data 132 and user data 134 for presentation during the runtime of the application 130 and allow the user to provide input during the runtime of the application 130.
[0035] In at least one embodiment, the server machine 106 can be separate from a server computer that supports the application hosting platform 102. In other embodiments, the server machine 106 can be part of the application hosting platform 102. In at least one embodiment, one or more server machines 106, application hosting platforms 102, application developer platforms 104, and data storage 112 can be part of a cloud environment that client devices 108A-N can access via the network 120.
[0036] Fig. Figure 2 illustrates a block diagram of a cloud environment 200 according to at least one embodiment. In at least one embodiment, the cloud environment 200 includes an application hosting platform 102, a login platform 250, and one or more servers 106, each providing one or more virtual computing environments 160. The client device 108 is connected to the cloud environment via the network 120, as described in Figure 2. Fig. 1 described. In at least one embodiment, the login platform 250 can include an account manager 205, an account linking component 215, and an identity manager 210. The application hosting platform 102 can include a deployment manager 140 and an application load manager 150, which includes a scheduler 225 and a software application service 230. The application hosting platform 102 can also include a game application engine 240 and an application file repository 260.
[0037] In at least one embodiment, the account linking component 215 can link an account of a user of the client device 108 with an account that is linked to the application hosting platform 102. In at least one embodiment, the account linking component 215 can also determine whether an account of the user of the client device 108, which is linked to an application (e.g., the application 130, as referred to in Fig. 1) is assigned, linked, or whether the user is "logged in". In at least one embodiment, if the user's account associated with application 130 is not linked, the account linking component 215 can facilitate linking the account. In at least one embodiment, if the user's account associated with application 130 is linked, the account linking component 215 can store the respective credentials or access and refresh tokens (e.g., tokens used to enable user access to application 130). In at least one embodiment, the account linking component 215 can link or maintain credentials for many different users for a number of applications 130 hosted by the application hosting platform 130.
[0038] In at least one embodiment, the account manager 205 can manage an account associated with the application 130 for a user of the client device 108. For example, the account manager 205 can provide a UI that allows the user of the client device 108 to log in to an account associated with the application 130.
[0039] In at least one embodiment, the login platform 250 can contain a unique account manager 205 for each unique application 130 hosted on the application hosting platform 130. In at least one embodiment, the identity manager 210 can validate (e.g., check) to ensure that the user's login credentials are associated with an account of the application 130. That is, the identity manager 210 can allow the user to access their account associated with the application 130.
[0040] The application hosting platform 102 can include a deployment manager 140 that can organize some or all of the virtualized computing environments 160 into a ring of virtualized computing environments, identify a number of applications 130, and organize the applications 130 in the ring based on one or more properties of the applications 130. Properties of the applications 130 can include, for example, application popularity, application size, application rating, etc. Based on the ring structure of the virtualized computing environments 160 and the applications 130, the deployment manager 140 can assign multiple sets of virtualized computing environments 160 to multiple applications 130, with each set of virtualized computing environments 160 being assigned to a specific application 130 (e.g.,(A first set of virtualized computing environments 160 is assigned to a first application 130, a second set of virtualized computing environments 160 is assigned to a second application 130, a third set of virtualized computing environments 160 is assigned to a third application 130, and so on). In at least one embodiment, the deployment manager 140 temporarily rotates the sets of virtualized computing environments 160 on the ring based on the expected wear characteristics of the respective applications 130, such that the wear characteristics of the virtualized computing environments 160 are evenly distributed among the virtualized computing environments 160. In at least one embodiment, the expected wear characteristics of each application 130 can be determined by the utilization of resources (GPU, CPU, hard disk, etc.).) is measured by the respective application 130 over time during multiple cloud sessions and relevant resource usage parameters, such as the percentage of GPU / CPU usage, the maximum GPU / CPU usage, read / write bytes, etc., are monitored, which can indicate the heavy use of GPU / CPU or hard drive by the application.
[0041] In at least one embodiment, the deployment manager 140 stores allocation data (the mapping of a unique identifier of a specific application 130 to a corresponding set of one or more virtualized computing environments 160) in the allocation data store 270, which may be, for example, a database, a table, a file, etc. In at least one embodiment, the deployment manager 140 can change the allocation for the application 130 if the properties of the application change. For example, if the application 130 becomes more popular than another application, one or more virtualized computing environments from the set previously allocated to the other application can be reassigned to the application 130. In at least one embodiment, the deployment manager 140 can allocate multiple applications to the same virtualized computing environment (e.g., for applications that are less popular).
[0042] The application hosting platform 102 may further include an application file repository 260, which stores files of various applications 130 (including, for example, generic data and user data of applications 130) registered with the application hosting platform 102, and an application load manager 150, which can enable the loading of an instance of the application 130 into the virtualized computing environment 160. In at least one embodiment, the application load manager 150 or the deployment manager 140 loads application data of each application 130 into caches 162 of corresponding sets of virtualized computing environments 160 assigned to the application 130 (for example, based on allocation data in the allocation data store 270).When the user selects an application 130, an instance of the application 130 can be run in a virtualized computing environment 160, which is selected from the allocated set of virtualized computing environments 160 based on the locations of applications 130 and virtualized computing environments 160 in the ring and using the data stored in the cache 162 of that virtualized computing environment 160. In at least one embodiment, the selected virtualized computing environment 160 is an available virtualized computing environment in the ring from the allocated set of virtualized computing environments 160. For example, the selected virtualized computing environment 160 is the first available virtualized computing environment in the ring from the allocated set of virtualized computing environments 160.In at least one embodiment, if the assigned set of virtualized computing environments 160 does not contain an available virtual computing environment, an available virtual computing environment is selected from another set of virtual computing environments 160 in the ring for the instance of the application 130.
[0043] In at least one embodiment, the application hosting platform 102 can receive a request (e.g., a user request) to terminate the application 130 and can stop the execution of the instance of the application 130, causing the virtualized computing environment 160 to become available for subsequent allocations in the respective set of virtualized computing environments 160. In at least one embodiment, the application hosting platform 102 can activate and / or deactivate virtualized computing environments 160 based on one or more parameters (e.g., time of day, day of the week, or month, etc.).
[0044] In at least one embodiment, the application hosting platform 102 can cause the instance of the application 130 in the virtualized computing environment 160 to stream application content to the client device 108, e.g., to the GUI of the application hosting platform (AHP), as described in reference to Fig. 1 described. In one embodiment, the application 130 can be an example of a game application (e.g., a video game) or a platform or application for collaborative content creation. In such examples, the virtualized computing environment 160 can store or load an instance of the game application. The instance of the game application can stream game content to the client device 108 when a user of the client device initiates a session.
[0045] In at least one embodiment, the application load manager 150 of the application hosting platform 102 can include a scheduler 225. In one embodiment, the scheduler 225 can be configured to initiate or start an instance of a game application 130 for a user of the client device 108. In one embodiment, the scheduler 225 can issue a request to allocate (e.g., create) a virtualized computing environment 160 to the server 106. For example, the scheduler 225 can initiate the creation of the virtualized computing environment 160 by sending a request to a virtual system manager (which manages the creation, deconstruction, etc., of virtualized computing environments 160 across server 106) or directly to a hypervisor or host operating system of the server 106.
[0046] In at least one embodiment, the application 130 is integrated into an API plug-in 235 of the application hosting platform (AHP) (e.g., a designated software development kit, SDK) which may be configured to communicate with the application hosting platform 102 via a predefined set of API commands to start and / or load an instance of the game application 130 in the virtualized computing environment 160. In at least one embodiment, the application hosting platform 102 may have a unique AHP API plug-in for each software application 130 that it hosts. Alternatively, a common AHP API plug-in may be operable with all software applications that are hosted / registered with the application hosting platform 102.
[0047] Fig. Figure 3 illustrates an exemplary ring-based allocation of 300 applications to virtualized computing environments using application properties, according to at least one embodiment.
[0048] A deployment manager of a cloud computing platform can identify one or more virtualized computing environments capable of hosting one or more applications. These virtualized computing environments can be, for example, virtual machines or containers. Furthermore, the cloud computing platform's deployment manager can identify one or more applications (for example, game applications) that are served by an application hosting platform. The deployment manager can also organize the applications into a ring of applications based on one or more application properties. For example, more popular applications can be assigned to a larger segment of the ring than less popular applications. Therefore, more popular applications can be allocated a greater number of virtualized computing environments than less popular applications.
[0049] As in Fig. As shown in Figure 3, the deployment manager identifies four applications. For example, the deployment manager might identify a first application as 302, a second application as 304, a third application as 306, and a fourth application as 308. Each application can be assigned a corresponding identifier, such as a Content Management Server identifier (ID) (cmsId). For example, the first application might be assigned cmsId A 310 (302), the second application might be assigned cmsId B 312 (304), the third application might be assigned cmsId C 314 (306), and the fourth application might be assigned cmsId D 316 (308).
[0050] As further in Fig. As shown in Figure 3, a cloud computing platform provides eighteen virtualized computing environments capable of running instances of the four applications. These virtualized computing environments are shown as virtualized computing environments 318 through 352. The deployment manager can organize the applications into a ring of applications based on one or more of the applications' properties. In some aspects, the deployment manager can organize the applications in the ring based on their respective popularity. For example, based on application popularity, the deployment manager can allocate thirty percent of the ring to the first application 302, twenty percent to the second application 304, ten percent to the third application 306, and forty percent to the fourth application, a game application 308.
[0051] The deployment manager can assign virtualized computing environments to one or more applications based on the respective locations of the virtualized computing environments and the applications within the application ring. For example, the deployment manager can assign six virtualized computing environments (318 to 328) to the first application (302), three virtualized computing environments (330 to 334) to the second application (304), two virtualized computing environments (336 and 338) to the third application (306), and seven virtualized computing environments (340 to 352) to the fourth application.
[0052] The deployment manager can store application-related data in the respective caches of the associated virtualized computing environments. For example, the deployment manager can store data associated with the first application, 302, in the caches of the corresponding six virtualized computing environments (318 to 328); data associated with the second application, 304, in the caches of the corresponding three virtualized computing environments (330 to 334); data associated with the third application, 306, in the caches of the corresponding two virtualized computing environments (336 and 338); and data associated with the fourth application, 308, in the caches of the corresponding seven virtualized computing environments (340 to 352).When a user selects one of the applications, an instance of the application can be run in a corresponding virtualized computing environment using the data stored in the cache of the virtualized computing environment.
[0053] While Fig. Figure 3 shows four applications and eighteen virtualized computing environments organized in a ring structure; it is understood that this is provided for illustrative purposes only. For example, any number of applications and virtualized computing environments can be identified, any number of virtualized computing environments can be assigned to any number of applications, and any type of structure or list can be used to organize and assign the virtualized computing environments and corresponding applications.
[0054] Fig. Figure 4 illustrates an exemplary ring-based allocation 400 of applications to virtualized computing environments using virtualized computing environment availability, according to at least one embodiment. Virtualized computing environments 418 to 452 (each corresponding to virtualized computing environments 318 to 352 in Example 300) are capable of hosting applications 402 to 408 (each corresponding to applications 302 to 308 in Example 300).
[0055] In some aspects, when a user request to run the first application 402 is received, an available virtualized computing environment can be selected from the set assigned to the first application 402 to run the instance of the first application 402. For example, the instance of the first application 402 can run in the first available virtualized computing environment from the set of virtualized computing environments assigned to the first application 402. As shown, virtualized computing environments 418 and 420 are unavailable (they may be in use, for example), while virtualized computing environments 422 through 428 are available. Therefore, the instance of the first application 402 can run in virtualized computing environment 422.
[0056] In some aspects, when a user request to run a second application (404) is received, an available virtualized computing environment can be selected from the set assigned to the second application (404) to run the instance of the second application (404). For example, the instance of the second application (404) can run in the first available virtualized computing environment from the set of virtualized computing environments assigned to the first application (404). As shown, virtualized computing environment 430 is unavailable (it may be in use, for example), while virtualized computing environments 432 and 434 are available. Therefore, the instance of the second application (404) can run in virtualized computing environment 432.
[0057] In some aspects, when a user request to run a third application 406 is received, an available virtualized computing environment from the set assigned to third application 406 can be selected to run the instance of third application 406. For example, the instance of third application 406 can run in the first available virtualized computing environment from the set of virtualized computing environments assigned to third application 406. As shown, both virtualized computing environments 436 and 438 are available. Therefore, the instance of third application 406 can run in virtualized computing environment 436.
[0058] In some aspects, when a user request to run a fourth application 408 is received, an available virtualized computing environment can be selected from the set assigned to fourth application 408 to run the instance of fourth application 408. For example, the instance of fourth application 408 can run in the first available virtualized computing environment from the set of virtualized computing environments assigned to fourth application 408. As shown, virtualized computing environments 440 through 444 are unavailable (they may be in use, for example), while virtualized computing environments 446 through 452 are available. Therefore, the instance of fourth application 408 can run in virtualized computing environment 446.
[0059] In some aspects, the virtualized computing environments can be enabled and / or disabled based on one or more parameters, such as a time parameter. For example, it can be determined that applications run more frequently during a certain time of day (such as in the evening) than during a different time of day (such as in the morning). Therefore, most or all of the virtualized computing environments can be enabled during the first time of day to allow applications to run faster, and some of the virtualized computing environments can be disabled during the second time of day to reduce power consumption.For example, during the first time of day, each of the virtualized computing environments 418 to 428 can be activated for the first application 402, while during the second time of day, the virtualized computing environments 418 and 420 can be deactivated for the first application 402 in order to reduce power consumption.
[0060] In some aspects, the deployment manager can reassign a virtualized compute environment from one application to another based on one or more conditions. For example, based on the detection that the second application (404) is becoming more popular than the first (402), the deployment manager can reassign the unique identifier of the second application (404) from one virtualized compute environment to another. Additionally, application data associated with the first application (404) can be deleted from the cache of virtualized compute environment (428), and application data associated with the second application (404) can be stored in a cache of virtualized compute environment (428). In some aspects, each virtualized compute environment's cache can store data for a single application.In other aspects, the cache of at least one virtualized computing environment can store data for multiple applications. This can enable the virtualized computing environment to efficiently host each of the multiple applications.
[0061] Fig. Figure 5 illustrates an exemplary allocation of 500 applications to virtualized computing environments using multiple rings, according to at least one embodiment.
[0062] In some aspects, the deployment manager can organize applications and virtualized computing environments into multiple rings. For example, the deployment manager can organize applications and virtualized computing environments into multiple overlapping rings. This can enable the storage of data for multiple applications in caches of the virtualized computing environments. In a two-ring structure, for example, the cache of each virtualized computing environment can store data for two applications.
[0063] As in Fig. As shown in Figure 5, the deployment manager identifies four applications. For example, the deployment manager might identify a first application as 502, a second application as 504, a third application as 506, and a fourth application as 508. Each application can be assigned a corresponding identifier. For example, the first application might be assigned cmsId A 510 (502), the second application might be assigned cmsId B 512 (504), the third application might be assigned cmsId C 514 (506), and the fourth application might be assigned cmsId D 516 (508).
[0064] As further in Fig. As shown in Figure 5, a cloud computing platform can provide nineteen virtualized computing environments capable of running instances of the four applications. These virtualized computing environments are shown as virtualized computing environments 518 through 554. The deployment manager can organize the applications in the ring based on one or more application properties. For example, the deployment manager can organize the applications in the ring based on their popularity.
[0065] As shown in Example 500, the first application is assigned to virtualized computing environments 518 to 536 (502), the second application is assigned to virtualized computing environments 538 to 554 (504), the third application is assigned to virtualized computing environments 518 to 526 (506), and the fourth application is assigned to virtualized computing environments 528 to 554 (508). From the perspective of the virtualized computing environments, virtualized computing environment 518 stores data for applications 502 and 506, virtualized computing environment 520 stores data for applications 502 and 506, virtualized computing environment 522 stores data for applications 502 and 506, virtualized computing environment 524 stores data for applications 502 and 506, virtualized computing environment 526 stores data for applications 502 and 506, and virtualized computing environment 528 stores data for applications 502 and 508.The virtualized computing environment 530 stores data for applications 502 and 508, the virtualized computing environment 532 stores data for applications 502 and 508, the virtualized computing environment 534 stores data for applications 502 and 508, the virtualized computing environment 536 stores data for applications 502 and 508, the virtualized computing environment 538 stores data for applications 504 and 508, the virtualized computing environment 540 stores data for applications 504 and 508, the virtualized computing environment 542 stores data for applications 504 and 508, the virtualized computing environment 544 stores data for applications 504 and 508, the virtualized computing environment 546 stores data for applications 504 and 508, the virtualized computing environment 548 stores data for applications 504 and 508, the virtualized computing environment 550 data for applications 504 and 508,The virtualized computing environment 552 stores data for applications 504 and 508; the virtualized computing environment 554 stores data for applications 504 and 508.
[0066] In some aspects, when a user request to run First Application 502 is received, an available virtualized computing environment can be selected from the set assigned to First Application 502 to run the instance of First Application 502. For example, the instance of First Application 502 can run in any of the several available virtualized computing environments assigned to First Application 502. As shown, virtualized computing environments 518 through 522 and virtualized computing environments 528 through 524 are unavailable (they may be in use, for example), while virtualized computing environments 524, 526, and 536 are available. Therefore, the instance of First Application 502 can run in virtualized computing environment 524.
[0067] In some aspects, when a user request to run a second application 504 is received, an available virtualized computing environment can be selected from the set assigned to the second application 504 to run the instance of the second application 504. For example, the instance of the second application 504 can run in the first available virtualized computing environment among the several virtualized computing environments assigned to the second application 504. As shown, virtualized computing environments 538 and 540 are unavailable (they may be in use, for example), while virtualized computing environments 542 through 554 are available. Therefore, the instance of the second application 504 can run in virtualized computing environment 542.
[0068] In some aspects, when a user request to run a third application 506 is received, an available virtualized computing environment can be selected from the set assigned to the third application 506 to run the instance of the third application 506. For example, the instance of the third application 506 can run in the first available virtualized computing environment among the several virtualized computing environments assigned to the third application 506. As shown, virtualized computing environments 518 through 522 are unavailable (they may be disabled, for example), and virtualized computing environment 524 is assigned to the first application 502, while virtualized computing environment 526 is available. Therefore, the instance of the third application 506 can run in virtualized computing environment 526.
[0069] In some aspects, when a user request to run a fourth application 508 is received, an available virtualized computing environment can be selected from the set assigned to the fourth application 508 to run the fourth application 508 instance. For example, the fourth application 508 instance can run in the first available virtualized computing environment among the several assigned to the fourth application 508. As shown, virtualized computing environments 528 through 534 and virtualized computing environments 538 through 540 are unavailable (they may be disabled, for example), and virtualized computing environment 542 is assigned to the second application 504, while virtualized computing environments 536 and 544 are available. Therefore, the fourth application 508 instance can run in virtualized computing environment 536.
[0070] Fig. Figure 6 illustrates a flowchart of an exemplary method 600 for allocating application hosting platforms in a virtualized computing environment, according to at least one embodiment. Method 600 can be performed using one or more processing units or processors (e.g., CPUs, GPUs, accelerators, physics processing units (PPUs), data processing units (DPUs), etc.) that can contain (or communicate with) one or more storage devices. According to some aspects of the disclosure, Method 600 can be performed using a processing device. According to some aspects of the disclosure, Method 600 can be performed using processing units of the application hosting platform 102 from Fig. 1. According to some aspects of the disclosure, processing units performing the procedure 600 can execute instructions stored on a non-volatile, computer-readable storage medium. According to some aspects of the disclosure, the procedure 600 can be performed using multiple processing threads (e.g., CPU threads and / or GPU threads), with individual threads performing one or more individual functions, routines, subroutines, or operations of the procedure. According to some aspects of the disclosure, the processing of threads implementing the procedure 600 can be synchronized (e.g., using semaphores, critical sections, and / or other thread synchronization mechanisms). Alternatively, the processing of threads implementing the procedure 600 can be performed asynchronously with respect to each other.Several operations of procedure 600 can be performed in a different order compared to the one in . Fig. The operation of Procedure 600 can be performed in the sequence shown in 6. Some operations of Procedure 600 can be performed simultaneously with other operations. According to some aspects of the disclosure, one or more of the operations shown in the disclosure can be performed simultaneously with other operations. Fig. The 6 operations shown are not always performed.
[0071] With reference to Fig. In block 605, processing units executing procedure 600 can identify a multitude of virtualized computing environments for running a multitude of applications provided by an application hosting platform. In some aspects, the multitude of virtualized computing environments is organized into a ring of virtualized computing environments.
[0072] In Block 610, processing units of a first application can assign a first set of virtualized computing environments to the multitude of applications based on one or more properties of the first application. In some aspects, the one or more properties of the first application include the size of the first application, the popularity of the first application, and / or a ranking of the first application. In some aspects, assigning the first set of virtualized computing environments to the first application involves assigning a unique identifier of the first application to the first set of virtualized computing environments.
[0073] In block 615, processing units can store data associated with the first application in respective caches of the first set of virtualized computing environments.
[0074] In Block 620, processing units of a second application can assign a second set of virtualized computing environments to the multitude of applications based on one or more properties of the second application. The second set of virtualized computing environments can differ from the first set. In some aspects, the one or more properties of the second application include the size of the second application, its popularity, and / or its ranking. In some aspects, assigning the second set of virtualized computing environments to the second application involves assigning a unique identifier to the second set of virtualized computing environments.
[0075] In block 625, processing units can store data assigned to the second application in respective caches of the second set of virtualized computing environments.
[0076] In block 630, processing units, in response to a request to run the first application, can cause an instance of the first application to run in a virtualized compute environment of the first set of virtualized compute environments, using data stored in a virtualized compute environment cache. In some aspects, virtualized compute environment assignments for the first and second applications are stored in an assignment data store and then used to select an available virtualized compute environment from the assigned virtualized compute environments in the ring. In some aspects, the selected virtualized compute environment is the first available virtualized compute environment in the ring associated with the first set of virtualized compute environments.
[0077] In some aspects, processing units can receive a request to terminate the first application and can reassign the virtualized computing environment as an available virtualized computing environment in the first set of virtualized computing environments.
[0078] In some aspects, processing units can activate or deactivate a specific virtualized computing environment from the first set of virtualized computing environments, based on one or more parameters. In some aspects, one or more parameters include a time-of-day parameter.
[0079] In some aspects, processing units can switch a virtualized computing environment from the first set of virtualized computing environments to the second set of virtualized computing environments based on one or more properties of the first application or one or more properties of the second application.
[0080] In some aspects, processing units can determine the first set of virtualized computing environments and the second set of virtualized computing environments based on one or more properties of the first application and one or more properties of the second application. In some aspects, the first set of virtualized computing environments and the second set of virtualized computing environments are temporarily rotated on the ring for reasons that, without limitation, include the expected wear characteristics of the first and second applications.
[0081] In some aspects, each virtualized computing environment stores application data for a single application in its cache. In other aspects, at least one virtualized computing environment stores application data for two or more applications in its cache.
[0082] In some aspects, processing units can receive a request to run one of two or more applications and to run an instance of the application in at least one virtualized computing environment that stores the application data for the two or more applications in the cache of the virtualized computing environment.
[0083] Fig. Figure 7A illustrates inference and training logic 715, which is used to perform inference and / or training operations associated with one or more embodiments. Details relating to inference and / or training logic 715 are given below in conjunction with Fig. 7A and / or 7B provided.
[0084] In at least one embodiment, the inference and / or training logic 715 may, without limitation, include a code and / or data store 701 for storing forward and / or output weights and / or input / output data and / or other parameters to configure neurons or layers of a neural network that is trained and / or used for inference in aspects by one or more embodiments. In at least one embodiment, the training logic 715 may include or be coupled to the code and / or data store 701 to store graph code or other software for controlling the timing and / or sequence in which weight and / or other parameter information for configuring the logic is loaded, including integer and / or floating-point units (collectively, arithmetic logic units, or ALUs).In at least one embodiment, a code, such as a graph code, loads weight or other parameter information into processor ALUs based on a neural network architecture to which the code corresponds. In at least one embodiment, the code and / or data store 701 stores weight parameters and / or input / output data of each layer of a neural network that was trained or used in conjunction with one or more embodiments during the forward propagation of input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments. In at least one embodiment, a portion of the code and / or data store 701 may be contained in another on-chip or off-chip data store, including an L1, L2, or L3 cache or system memory of a processor.
[0085] In at least one embodiment, a section of the code and / or data memory 701 can be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or the code and / or data memory 701 can be a cache memory, dynamic randomly addressable memory (DRAM), static randomly addressable memory (SRAM), non-volatile memory (e.g., flash memory), or other memory.In at least one embodiment, the choice of whether the code and / or the code and / or data storage 701 is, for example, internal or external to a processor or comprises DRAM, SRAM, Flash or another type of memory, may depend on on-chip versus off-chip available memory, latency requirements of the training and / or inference functions performed, the batch size of the data used in the inference and / or training of a neural network or a combination of these factors.
[0086] In at least one embodiment, the inference and / or training logic 715 may, without limitation, include a code and / or data store 705 for storing backward and / or output weights and / or input / output data corresponding to neurons or layers of a neural network that is trained and / or used for inference in aspects of one or more embodiments. In at least one embodiment, the code and / or data store 705 stores weight parameters and / or input / output data of each layer of a neural network that was trained or used in conjunction with one or more embodiments during the backward propagation of input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments.In at least one embodiment, the training logic 715 can include or be coupled to the code and / or data memory 705 to store graph code or other software for controlling the timing and / or sequence in which weight and / or other parameter information for configuring the logic is loaded, including integer and / or floating-point units (collectively, arithmetic logic units, or ALUs). In at least one embodiment, a code, such as a graph code, loads weight or other parameter information into processor ALUs based on a neural network architecture to which that code corresponds.In at least one embodiment, a portion of the code and / or data memory 705 may be contained in another on-chip or off-chip data memory, including an L1, L2, or L3 cache or system memory of a processor. In at least one embodiment, a portion of the code and / or data memory 705 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data memory 705 may be a cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other memory.In at least one embodiment, the choice of whether the code and / or data storage 705 is, for example, internal or external to a processor, or comprises DRAM, SRAM, Flash or another type of memory, may depend on on-chip versus off-chip available memory, latency requirements of the training and / or inference functions performed, the batch size of the data used in the inference and / or training of a neural network, or a combination of these factors.
[0087] In at least one embodiment, the code and / or data memory 701 and the code and / or data memory 705 can be separate memory structures. In at least one embodiment, the code and / or data memory 701 and the code and / or data memory 705 can be the same memory structure. In at least one embodiment, the code and / or data memory 701 and the code and / or data memory 705 can be partly the same memory structure and partly separate memory structures. In at least one embodiment, a portion of the code and / or data memory 701 or the code and data memory 705 can be contained in another on-chip or off-chip data memory, including an L1, L2, or L3 cache or system memory of a processor.
[0088] In at least one embodiment, the inference and / or training logic 715 may, without limitation, include one or more arithmetic logic units (“ALUs”) 710, including integer and / or floating-point units, to perform logical and / or mathematical operations that are at least partially based on or indicated by a training and / or inference code (e.g., graph code), wherein a result thereof may generate activations (e.g., output values of layers or neurons within a neural network) that are stored in an activation memory 720, which are functions of input / output and / or weight parameter data that are stored in the code and / or data memory 701 and / or code and / or data memory 705.In at least one embodiment, activations stored in the activation memory 720 are generated according to linear algebraic and / or matrix-based mathematics, which are performed by ALU(s) 710 in response to the execution of instructions or other code, wherein weight values stored in the code and / or data memory 705 and / or code and / or data memory 701 are used as operands together with other values, such as bias values, gradient information, momentum values or other parameters or hyperparameters, some or all of which may be stored in the code and / or data memory 705 or code and / or data memory 701 or in another memory on or off the chip.
[0089] In at least one embodiment, the ALU(s) 710 are included in one or more processors or other hardware logic devices or circuits, while in another embodiment, the ALU(s) 710 may be external to a processor or other hardware logic device or circuit that uses them (e.g., a coprocessor). In at least one embodiment, the ALU(s) 710 may be included within the execution units of a processor or otherwise within a series of ALUs that can be accessed by the execution units of a processor, either within the same processor or distributed across different processors of different types (e.g., central processing units, graphics processing units, fixed functional units, etc.).In at least one embodiment, the code and / or data memory 701, the code and / or data memory 705, and the activation memory 720 can be located on a single processor or other hardware logic device or circuit, while in another embodiment they can be located on different processors or other hardware logic devices or circuits, or on a combination of identical and different processors or other hardware logic devices or circuits. In at least one embodiment, a portion of the activation memory 720 can be contained in another on-chip or off-chip data memory, including an L1, L2, or L3 cache or system memory of a processor.Furthermore, the inference and / or training code can be stored with other code that a processor or other hardware logic or circuitry can access and be retrieved and / or processed using a processor's retrieval, decoding, scheduling, execution, elimination, and / or other logic circuitry.
[0090] In at least one embodiment, the activation memory 720 can be a cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or another type of memory. In at least one embodiment, the activation memory 720 can be located wholly or partially inside or external to one or more processors or other logic circuits. In at least one embodiment, the choice of whether the activation memory 720 is, for example, internal to or external to a processor, or whether it comprises DRAM, SRAM, flash memory, or another type of memory, can depend on on-chip versus off-chip available memory, latency requirements of the training and / or inference functions performed, the batch size of the data used in the inference and / or training of a neural network, or a combination of these factors. In at least one embodiment, the inference and / or training logic 715, which is implemented in Fig. Figure 7A illustrates that it can be used in conjunction with an application-specific integrated circuit (“ASIC”), such as a Google TensorFlow® processing unit, a Graphcore™ inference processing unit (IPU), or an Intel Corp. Nervana® processor (e.g., “Lake Crest”). In at least one embodiment, inference and training logic 715, which is illustrated in Figure 7A, ... Fig. 7A illustrates how it can be used in conjunction with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware, or other hardware such as data processing unit (“DPU”) hardware or field programmable gate arrays (“FPGAs”).
[0091] Fig. Figure 7B illustrates the inference and / or training logic 715 according to at least one or more embodiments. In at least one embodiment, the inference and / or training logic 715 can, without limitation, include hardware logic in which computing resources are dedicated or otherwise used exclusively in connection with weight values or other information corresponding to one or more layers of neurons within a neural network. In at least one embodiment, the inference and / or training logic 715, which is described in Fig. Figure 7B illustrates that it can be used in conjunction with an application-specific integrated circuit (ASIC), such as a Google TensorFlow® processing unit, a Graphcore™ inference processing unit (IPU), or an Intel Corp. Nervana® processor (e.g., “Lake Crest”). In at least one embodiment, inference and training logic 715, which is illustrated in Figure 7B, ... Fig. Figure 7B illustrates the inference and / or training logic 715, which can be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware, such as data processing unit (DPU) hardware or field programmable gate arrays (FPGAs). In at least one embodiment, the inference and / or training logic 715 includes, without limitation, a code and / or data memory 701 and a code and / or data memory 705, which can be used to store a code (e.g., graph code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. In at least one embodiment, which is illustrated in Figure 7B, the inference and / or training logic 715 includes, without limitation, a code and / or data memory 701 and a code and / or data memory 705, which can be used to store a code (e.g., graph code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. Fig. As illustrated in Figure 7B, each of the code and / or data memory 701 and code and / or data memory 705 is assigned to a dedicated computing resource, such as computing hardware 702 and computing hardware 706, respectively. In at least one embodiment, each of the computing hardware 702 and computing hardware 706 comprises one or more ALUs that perform mathematical functions, such as linear algebraic functions, only on information stored in the code and / or data memory 701 and the code and / or data memory 705, respectively, the result of which is stored in the activation memory 720.
[0092] In at least one embodiment, each of the code and / or data storage 701 and 705 and the corresponding computing hardware 702 and 706, respectively, corresponds to different layers of a neural network, such that a resulting activation from a "memory / computing pair 701 / 702" of the code and / or data storage 701 and the computing hardware 702 is provided as an input to a "memory / computing pair 705 / 706" of the code and / or data storage 705 and the computing hardware 706, to reflect a conceptual organization of a neural network. In at least one embodiment, each of the memory / computing pairs 701 / 702 and 705 / 706 can correspond to more than one neural network layer. In at least one embodiment, additional memory / computing pairs (not shown) may be included after or in parallel to the memory / computing pairs 701 / 702 and 705 / 706 in the inference and / or training logic 715.
[0093] Fig. Figure 8 illustrates an exemplary data center 800 in which at least one embodiment can be used. In at least one embodiment, the data center 800 includes an infrastructure layer 810, a framework layer 820, a software layer 830, and an application layer 840.
[0094] In at least one embodiment, as in Fig. Figure 8 shows that the infrastructure layer 810 of the data center can include a resource orchestrator 812, clustered compute resources 814, and node compute resources (“node RRs”) 816(1)-816(N), where “N” is any positive integer. In at least one embodiment, the node RRs 816(1)-816(N) can include, but are not limited to, any number of central processing units (CPUs) or other processors (including accelerators, field programmable gate arrays (FPGAs), data processing units, graphics processing units, etc.), memory devices (e.g., dynamic solid-state storage), storage devices (e.g., solid-state or disk drives), network input / output devices (“NW I / O”), network switches, virtual machines (“VMs”), power supply modules, and / or cooling modules, etc.In at least one embodiment, one or more Node RRs among Node RRs 816(1)-816(N) can be a server that has one or more of the above-mentioned computing resources.
[0095] In at least one embodiment, the grouped compute resources 814 can include separate groupings of node RRs located in one or more racks (not shown) or in many racks in data centers at different geographic locations (also not shown). Separate groupings of node RRs within grouped compute resources 814 can include grouped computing, networking, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, multiple node RRs, including the CPUs or processors, can be grouped within one or more racks to provide compute resources to support one or more workloads.In at least one embodiment, the one or more racks can also include any number of power supply modules, cooling modules and network switches in any combination.
[0096] In at least one embodiment, the resource orchestrator 812 can configure or otherwise control one or more node RRs 816(1)-816(N) and / or grouped computing resources 814. In at least one embodiment, the resource orchestrator 812 can include a software design infrastructure (“SDI”) management entity for the data center 800. In at least one embodiment, the resource orchestrator can include hardware, software, or a combination thereof.
[0097] In at least one embodiment, as in Fig. As shown in Figure 8, the framework layer 820 includes a job scheduler 822, a configuration manager 824, a resource manager 826, and a distributed file system 828. In at least one embodiment, the framework layer 820 can include a framework that supports the software 832 of software layer 830 and / or one or more applications 842 of application layer 840. In at least one embodiment, the software 832 or the one or more applications 842 can each include web-based service software or applications such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 820 can be a type of free and open-source software web application framework, such as Apache Spark™ (hereinafter "Spark"), which can utilize a distributed file system 828 for processing large amounts of data (e.g., "Big Data") without being restricted to it.In at least one embodiment, the job scheduler 822 can include a Spark driver to facilitate the scheduling of workloads supported by different layers of the data center 800. In at least one embodiment, the configuration manager 824 can be able to configure different layers, such as the software layer 830 and the framework layer 820, including Spark and the distributed file system 828, to support the processing of large amounts of data. In at least one embodiment, the resource manager 826 can be able to manage clustered or grouped compute resources allocated or assigned to support the distributed file system 828 and the job scheduler 822. In at least one embodiment, the clustered or grouped compute resources can include the grouped compute resource 814 on the infrastructure layer 810 of the data center.In at least one embodiment, the resource manager 826 can coordinate with the resource orchestrator 812 to manage these allocated or assigned computing resources.
[0098] In at least one embodiment, the software contained in software layer 830 may include software 832 that is used by at least sections of the node RRs 816(1)-816(N), the grouped compute resources 814, and / or the distributed file system 828 of framework layer 820. The one or more types of software may include, among others, web search software, email virus scanning software, database software, and streaming video content software.
[0099] In at least one embodiment, the applications 842 contained in the application layer 840 may include one or more types of applications used by at least sections of the node RRs 816(1)-816(N), the grouped compute resources 814, and / or the distributed file system 828 of the framework layer 820. One or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computations, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.
[0100] In at least one embodiment, the configuration manager 824, the resource manager 826, and / or the resource orchestrator 812 can implement any number and type of self-modifying actions based on any set and type of data acquired in any technically feasible way. In at least one embodiment, self-modifying actions can relieve a data center operator of the data center 800 of potentially making poor configuration decisions and potentially avoiding underutilized and / or poorly functioning sections of a data center.
[0101] In at least one embodiment, the Data Center 800 may contain tools, services, software, or other resources for training one or more machine learning models or for predicting or inferring information using one or more machine learning models according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model may be trained by calculating weight parameters according to a neural network architecture, using the software and computing resources described above with reference to the Data Center 800.In at least one embodiment, trained machine learning models corresponding to one or more neural networks can be used to infer or predict information using the resources described above with reference to the Computing Center 800 by using weight parameters calculated by one or more training techniques such as those described herein.
[0102] In at least one embodiment, the data center can use CPUs, application-specific integrated circuits (ASICs), GPUs, DPUs, FPGAs, or other hardware to perform training and / or inference using the resources described above. Furthermore, one or more of the software and / or hardware resources described above can be configured as a service to allow users to train or infer information, such as image capture, speech capture, or other artificial intelligence services.
[0103] The inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 715 are provided below in conjunction with Fig. 7A and / or 7B are provided. In at least one embodiment, the inference and / or training logic 715 can be located in the system of Fig. 8 for inference or prediction operations that are based at least partially on weight parameters calculated using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.
[0104] Such components can be used to generate synthetic data that mimics error cases in a network training process, which can help improve network performance while limiting the amount of synthetic data to avoid overfitting.
[0105] Fig. Figure 9 is a block diagram illustrating an exemplary computer system 900, which may be a system of interconnected devices and components, a system-on-a-chip (SoC), or any other combination thereof, formed with a processor, which may include execution units for carrying out an instruction, according to at least one embodiment. In at least one embodiment, the computer system 900 may, without limitation, include a component, such as a processor 902, to employ execution units, including logic for carrying out algorithms on process data, according to the present disclosure, as in the embodiment described herein.In at least one embodiment, the Computer System 900 may include processors such as the PENTIUM® processor family, Xeon™, Itanium®, XScale™ and / or StrongARM™, Intel® Core™ or Intel® Nervana™ microprocessors available from Intel Corporation in Santa Clara, California, although other systems (including PCs with other microprocessors, engineering workstations, set-top boxes, and the like) may also be used. In at least one embodiment, the Computer System 900 may run a version of a WINDOWS operating system available from Microsoft Corporation, Redmond, Washington, although other operating systems (for example, UNIX and Linux), embedded software, and / or graphical user interfaces may also be used.
[0106] Embodiments can be used in other devices, such as handheld devices and embedded applications. Some examples of handheld devices include mobile phones, Internet Protocol devices, digital cameras, personal digital assistants (“PDAs”), and handheld PCs. In at least one embodiment, embedded applications can include a microcontroller, a digital signal processor (“DSP”), a system-on-a-chip, network computers (“NetPCs”), set-top boxes, network hubs, wide area network (“WAN”) switches, edge devices, Internet of Things (“IoT”) devices, or any other system capable of executing one or more instructions, according to at least one embodiment.
[0107] In at least one embodiment, the computer system 900 can include, without limitation, the processor 902, which can include, without limitation, one or more execution units 908 for performing machine learning model training and / or inference according to the techniques described herein. In at least one embodiment, the computer system 900 is a single-processor desktop or server system, but in another embodiment, the computer system 900 can be a multiprocessor system.In at least one embodiment, the processor 902 can, without restriction, include a Complex Instruction Computer (CISC) microprocessor, a Reduced Instruction Set Computing (RISC) microprocessor, a Very Long Instruction Word (VLIW) microprocessor, a processor implementing a combination of instruction sets, or any other processing device, such as a digital signal processor. In at least one embodiment, the processor 902 can be coupled to a processor bus 910, which can transmit data signals between the processor 902 and other components in the computer system 900.
[0108] In at least one embodiment, a processor 902 can, without limitation, include an internal Level 1 (“L1”) cache memory (“cache”) 904. In at least one embodiment, a processor 902 can have a single internal cache or multiple levels of an internal cache. In at least one embodiment, the cache memory can be located outside of a processor 902. Other embodiments can also include a combination of both internal and external caches, depending on specific implementations and requirements. In at least one embodiment, a register bank 906 can store different types of data in different registers, including, without limitation, integer registers, floating-point registers, status registers, and instruction pointer registers.
[0109] In at least one embodiment, an execution unit 908, which contains unrestricted logic for performing integer and floating-point operations, is also located in a processor 902. In at least one embodiment, a processor 902 can also contain a microcode ("ucode") read-only memory ("ROM") that stores microcode for specific macro instructions. In at least one embodiment, an execution unit 908 can contain logic for executing a packaged instruction set 909. In at least one embodiment, by including a packaged instruction set 909 in an instruction set of a general-purpose processor 902, together with associated switching technology for executing instructions, operations used by many multimedia applications can be performed using packaged data in a general-purpose processor 902.In one or more embodiments, many multimedia applications can be accelerated and run more efficiently by using the full width of a processor data bus to perform operations on packed data, thereby eliminating the need to transfer smaller units of data over the processor data bus to perform one or more operations on a single data element each.
[0110] In at least one embodiment, an execution unit 908 can also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, the computer system 900 can include a working memory 920 without restriction. In at least one embodiment, the working memory 920 can be implemented as a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, or another storage device. In at least one embodiment, the working memory 920 can store one or more instructions 919 and / or data 921, represented by data signals that can be executed by the processor 902.
[0111] In at least one embodiment, the system logic chip can be coupled to the processor bus 910 and the main memory 920. In at least one embodiment, the system logic chip can include, without restriction, a memory controller hub (MCH) 916, and the processor 902 can communicate with the MCH 916 via the processor bus 910. In at least one embodiment, the MCH 916 can provide a high-bandwidth memory path 918 to the main memory 920 for instruction and data storage and for storing graphics instructions, data, and textures. In at least one embodiment, the MCH 916 can route data signals between the processor 902, the main memory 920, and other components in the computer system 900, and for bridging data signals between the processor bus 910, the main memory 920, and a system I / O 922.In at least one embodiment, the system logic chip can provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH 916 can be coupled to the main memory 920 via a high-bandwidth main memory path 918, and the graphics / video card 912 can be coupled to the MCH 916 via an Accelerated Graphics Port (“AGP”) interconnect 914.
[0112] In at least one embodiment, the computer system 900 can use the system I / O 922, which is a proprietary hub interface bus, to couple the MCH 916 to the I / O controller hub (“ICH”) 930. In at least one embodiment, an ICH 930 can provide direct connections to some I / O devices via a local I / O bus. In at least one embodiment, a local I / O bus can, without limitation, include a high-speed I / O bus for connecting peripheral devices to a memory 920, a chipset, and a processor 902.Examples may include, without limitation, an audio controller 929, a firmware hub (“flash BIOS”) 928, a wireless transceiver 926, a data storage device 924, a legacy I / O controller 923 containing user input and keyboard interfaces 925, a serial expansion port 927, such as a Universal Serial Bus (“USB”), and a network controller 934, which in at least one embodiment may include a data processing unit. The data storage device 924 may include a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.
[0113] In at least one embodiment, Fig. 9 represents a system that includes interconnected hardware devices or “chips”, whereas in other embodiments Fig. Figure 9 may represent an exemplary system-on-a-chip (“SoC”). In at least one embodiment, devices can be interconnected using proprietary interconnects, standardized interconnects (e.g., PCIe), or a combination thereof. In at least one embodiment, one or more components of the Computer System 900 are interconnected using Compute Express Link (CXL) interconnects.
[0114] The inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 715 are provided below in conjunction with Fig. 7A and / or 7B are provided. In at least one embodiment, the inference and / or training logic 715 can be located in the system of Fig. 9 for inference or prediction operations that are based at least partially on weight parameters calculated using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.
[0115] Such components can be used to generate synthetic data that mimics error cases in a network training process, which can help improve network performance while limiting the amount of synthetic data to avoid overfitting.
[0116] Fig. Figure 10 is a block diagram illustrating a System 1000 for using a Processor 1010, according to at least one embodiment. In at least one embodiment, the System 1000 can be, for example, and without limitation, a notebook, a tower server, a rack server, a blade server, a laptop, a desktop, a tablet, a mobile device, a telephone, an embedded computer, an edge device, an IoT device, or any other suitable electronic device.
[0117] In at least one embodiment, the system 1000 can include, without restriction, the processor 1010, which is communicatively coupled to any suitable number or type of components, peripherals, modules, or devices. In at least one embodiment, the processor 1010 is coupled using a bus or interface, such as a 1°C bus, a system management / bus (“SMBus”), a low-pin-count (LPC) bus, a serial peripheral interface (“SPI”), a high-definition audio (“HDA”) bus, a serial-advanced technology attachment (“SATA”) bus, a universal serial bus (“USB”) (versions 1, 2, 3), or a universal asynchronous receiver / transmitter (“UART”) bus. In at least one embodiment, Fig. 10 represents a system that includes intermediate hardware devices or “chips”, whereas in other embodiments Fig. 10 can represent an exemplary system-on-a-chip (“SoC”). In at least one embodiment, the components described in Fig. The 10 illustrated devices can be connected using proprietary connections, standardized connections (e.g., PCIe®), or a combination thereof. In at least one embodiment, one or more components of Fig. 10 interconnected using Compute Express Link (CXL) connections.
[0118] In at least one embodiment, Fig. 10 a display 1024, a touchscreen 1025, a touchpad 1030, a near field communication (NFC) unit 1045, a sensor hub 1040, a thermal sensor 1046, an Express chipset (EC) 1035, a Trusted Platform Module (TPM) 1038, a BIOS / firmware / flash memory (BIOS, FW Flash) 1022, a DSP 1060, a drive 1020, such as a solid-state drive (SSD) or a hard disk drive (HDD), a wireless local area network (WLAN) unit 1050, a Bluetooth unit 1052, a wireless wide area network (WWAN) unit 1056, a global positioning system (GPS) 1055, a camera (USB 3.0 camera) 1054, such as a The device may include a USB 3.0 camera and / or a Low-Power Double Data-Rate (LPDDR) memory unit (LPDDR3), for example implemented in an LPDDR3 standard. These components can each be implemented in any suitable manner.
[0119] In at least one embodiment, other components can be communicatively coupled to the processor 1010 via the components described above. In at least one embodiment, an accelerometer 1041, an ambient light sensor (ALS) 1042, a compass 1043, and a gyroscope 1044 can be communicatively coupled to the sensor hub 1040. In at least one embodiment, a thermal sensor 1039, a fan 1037, a keyboard 1036, and a touchpad 1030 can be communicatively coupled to the EC 1035. In at least one embodiment, a loudspeaker 1063, an audio unit 1064, and a microphone (“Mic”) 1065 can be communicatively coupled to an audio unit (“audio codec and Class-D amplifier”) 1062, which in turn can be communicatively coupled to the DSP 1060.In at least one embodiment, the audio unit 1064 can, for example, and without limitation, include an audio encoder / decoder (“codec”) and a Class-D amplifier. In at least one embodiment, a SIM card (“SIM”) 1057 can be communicatively coupled with the WWAN unit 1056. In at least one embodiment, components such as the WLAN unit 1050 and Bluetooth unit 1052, as well as the WWAN unit 1056, can be implemented in a next-generation form factor (“NGFF”).
[0120] The inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 715 are provided below in conjunction with Fig. 7A and / or 7B are provided. In at least one embodiment, the inference and / or training logic 715 can be located in the system of Fig. 10 for inference or prediction operations that are based at least partially on weight parameters calculated using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.
[0121] Such components can be used to generate synthetic data that mimics error cases in a network training process, which can help improve network performance while limiting the amount of synthetic data to avoid overfitting.
[0122] Fig. Figure 11 is a block representation of a processing system according to at least one embodiment. In at least one embodiment, the system 1100 includes one or more processors 1102 and one or more graphics processors 1108 and can be a single-processor desktop system, a multi-processor workstation system, or a server system with a large number of processors 1102 or processor cores 1107. In at least one embodiment, the system 1100 is a processing platform integrated into a system-on-a-chip (SoC) circuit for use in mobile, handheld, edge, or embedded devices.
[0123] In at least one embodiment, the system 1100 can include or be integrated into a server-based gaming platform, a gaming console comprising a gaming and media console, a mobile gaming console, a handheld gaming console, or an online gaming console. In at least one embodiment, the system 1100 is a mobile phone, a smartphone, a tablet computer, or a mobile internet device. In at least one embodiment, the processing system 1100 can also include, be coupled to, or be integrated into a wearable device, such as a wearable smartwatch, smart glasses, augmented reality, or virtual reality device.In at least one embodiment, the processing system 1100 is a television set or a set-top box device with one or more processors 1102 and a graphical interface generated by one or more graphics processors 1108.
[0124] In at least one embodiment, the one or more processors 1102 each include one or more processor cores 1107 for processing instructions that, when executed, perform operations for system and user software. In at least one embodiment, the one and the multiple processor cores 1107 are each configured to process a specific instruction set 1109. In at least one embodiment, the instruction set 1109 can enable complex instruction set computing (CISC), reduced instruction set computing (RISC), or very-long instruction word (VLIW) computing. In at least one embodiment, the one or the multiple processor cores 1107 can each process a different instruction set 1109, which may include instructions for enabling the emulation of other instruction sets.In at least one embodiment, the one or more processor cores 1107 may also include other processing devices, such as a digital signal processor (DSP).
[0125] In at least one embodiment, the processor 1102 includes a cache memory 1104. In at least one embodiment, the processor 1102 can have a single internal cache or multiple levels of an internal cache. In at least one embodiment, the cache memory is shared by different components of the processor 1102. In at least one embodiment, the processor 1102 also uses an external cache (e.g., a Level 3 (L3) cache or Last-Level Cache (LLC)) (not shown), which can be shared by the processor cores 1107 using known cache coherence techniques. In at least one embodiment, a register bank 1106 is additionally included in the processor 1102, which can contain different types of registers for storing different types of data (e.g., integer registers, floating-point registers, status registers, and an instruction pointer register).In at least one embodiment, the register bank 1106 may include general-purpose registers or other registers.
[0126] In at least one embodiment, one or more processors 1102 are coupled to one or more interface buses 1110 for transmitting communication signals, such as address, data, or control signals, between the processor 1102 and other components in the system 1100. In at least one embodiment, the interface bus 1110 can be a processor bus, such as a version of a Direct Media Interface (DMI) bus. In at least one embodiment, the interface bus 1110 is not limited to a DMI bus and can include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), memory buses, or other types of interface buses. In at least one embodiment, the processor(s) 1102 can include an integrated memory controller 1116 and a platform controller hub 1130.In at least one embodiment, the main memory controller 1116 enables communication between a main memory device and other components of the system 1100, while the platform controller hub (PCH) 1130 provides connections to I / O devices via a local I / O bus.
[0127] In at least one embodiment, a memory device 1120 can be a dynamic random-access memory (DRAM) device, a static random-access memory (SRAM) device, a flash memory device, a phase-change memory device, or another memory device that has suitable performance to serve as process memory. In at least one embodiment, the memory device 1120 can be operated as system memory for the system 1100 to store data 1122 and instructions 1121 for use when one or more processors 1102 are executing an application or process.In at least one embodiment, the memory controller 1116 is also coupled to an optional external graphics processor 1112, which can communicate with one or more graphics processors 1108 in the one or more processors 1102 to perform graphics and media operations. In at least one embodiment, a display device 1111 can be connected to the processor(s) 1102. In at least one embodiment, the display device 1111 can include one or more internal displays, such as in a mobile electronic device or a laptop device, or external displays connected via a display interface (e.g., DisplayPort, etc.).In at least one embodiment, the display device 1111 may include a head-mounted display (HMD) such as a stereoscopic display device for use in virtual reality (VR) applications or augmented reality (AR) applications.
[0128] In at least one embodiment, the platform controller hub 1130 enables peripheral devices to be connected to the memory device 1120 and the processor 1102 via a high-speed I / O bus. In at least one embodiment, the I / O peripheral devices include, without limitation, an audio controller 1146, a network controller 1134, a firmware interface 1128, a wireless transceiver 1126, touch sensors 1125, and a data storage device 1124 (e.g., a hard disk drive, flash memory, etc.). In at least one embodiment, the data storage device 1124 can be connected via a storage interface (e.g., SATA) or via a peripheral bus, such as a peripheral component interconnect bus (e.g., PCI, PCI Express). In at least one embodiment, the touch sensors 1125 can include touchscreen sensors, pressure sensors, or fingerprint sensors.In at least one embodiment, the wireless transceiver 1126 can be a WiFi transceiver, a Bluetooth transceiver, or a mobile network transceiver such as a 3G, 4G, or Long-Term Evolution (LTE) transceiver. In at least one embodiment, the firmware interface 1128 enables communication with the system firmware and can, for example, be a Unified Extensible Firmware Interface (UEFI). In at least one embodiment, the network controller 1134 enables a network connection to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to the interface bus 1110. In at least one embodiment, the audio controller 1146 is a multi-channel high-definition audio controller. In at least one embodiment, the system 1100 includes an optional legacy I / O controller 1140 for coupling older devices (e.g.,Personal System 2 devices (PS / 2 devices) with the system. In at least one embodiment, the platform controller hub 1130 can also be connected to one or more connection input devices of one or more Universal Serial Bus (USB) controllers 1142, such as combinations of keyboard and mouse 1143, a camera 1144, or other USB input devices.
[0129] In at least one embodiment, an instance of the memory controller 1116 and platform controller hub 1130 can be integrated into a discrete external graphics processor, such as an external graphics processor. In at least one embodiment, the platform controller hub 1130 and / or the memory controller 1116 can be external to one or more processor(s) 1102. For example, in at least one embodiment, the system 1100 can include an external memory controller 1116 and platform controller hub 1130, which can be configured as a memory controller hub and peripheral controller hub within a system chipset that communicates with the processor(s) 1102.
[0130] The inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 715 are provided below in conjunction with Fig. 7A and / or 7B provided. In at least one embodiment, parts or all of the inference and / or training logic 715 can be incorporated into the graphics processor 1108. For example, in at least one embodiment, training and / or inference techniques described herein can use one or more ALUs embodied in a graphics processor. Furthermore, in at least one embodiment, inference and / or training operations described herein can be performed using logic that differs from that in Fig. 7A or Fig. The logic illustrated in Figure 7B differs. In at least one embodiment, weight parameters can be stored in on-chip or off-chip memory and / or registers (shown or not shown), thereby configuring graphics processor ALUs to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0131] Such components can be used to generate synthetic data that mimics error cases in a network training process, which can help improve network performance while limiting the amount of synthetic data to avoid overfitting.
[0132] Fig. Figure 12 is a block representation of a processor 1200 comprising one or more processor cores 1202A-1202N, an integrated memory controller 1213, and an integrated graphics processor 1208, according to at least one embodiment. In at least one embodiment, the processor 1200 may include additional cores up to and including the additional core 1202N, represented by dashed boxes. In at least one embodiment, the one or more processor cores 1202A-1202N each include one or more internal cache units 1204A-1204N. In at least one embodiment, each processor core also has access to one or more shared cache units 1206.
[0133] In at least one embodiment, the one or more internal cache units 1204A-1204N and the one or more shared cache units 1206 constitute a cache memory hierarchy within the processor 1200. In at least one embodiment, the cache memory units 1204A-1204N can include at least one level of an instruction and data cache within each processor core and one or more levels of a shared mid-level cache, such as Level 2 (L2), Level 3 (L3), Level 4 (L4), or other levels of a cache, wherein a highest level of a cache in front of external memory is classified as LLC. In at least one embodiment, cache coherence logic maintains coherence between different cache units 1206 and 1204A-1204N.
[0134] In at least one embodiment, the processor 1200 can also include a set of one or more bus controller units 1216 and a system agent core 1210. In at least one embodiment, the one or more bus controller units 1216 manage a set of peripheral buses, such as one or more PCI or PCI Express buses. In at least one embodiment, the system agent core 1210 provides management functionality for various processor components. In at least one embodiment, the system agent core 1210 includes one or more integrated memory controllers 1213 for managing access to various external memory devices (not shown).
[0135] In at least one embodiment, one or more of the processor cores 1202A-1202N include support for simultaneous multi-threading. In at least one embodiment, the system agent core 1210 includes components for coordinating and processing the cores 1202A-1202N during multi-threading. In at least one embodiment, the system agent core 1210 may additionally include a power control unit (PCU) that includes logic and components for regulating one or more power states of the one or more processor cores 1202A-1202N and the graphics processor 1208.
[0136] In at least one embodiment, the processor 1200 additionally includes the graphics processor 1208 for performing graphics processing operations. In at least one embodiment, the graphics processor 1208 is coupled with one or more shared cache units 1206 and the system agent core 1210, which includes one or more integrated memory controllers 1213. In at least one embodiment, the system agent core 1210 also includes a display controller 1211 for driving the graphics processor output to one or more coupled displays. In at least one embodiment, the display controller 1211 can also be a separate module coupled to the graphics processor 1208 via at least one intermediate connection, or it can be integrated into the graphics processor 1208.
[0137] In at least one embodiment, a ring-based interconnect unit 1212 is used to couple internal components of the processor 1200. In at least one embodiment, an alternative interconnect unit can be used, such as a point-to-point interconnect, a switched interconnect, or other techniques. In at least one embodiment, the graphics processor 1208 is coupled to the ring connection 1212 via an I / O link 1213.
[0138] In at least one embodiment, the I / O link 1213 represents at least one of several types of I / O intermediaries, including an on-packet I / O intermediary that enables communication between different processor components and an embedded high-performance memory module 1218, such as an eDRAM module. In at least one embodiment, each of the processor cores 1202A-1202N and the graphics processor 1208 use embedded memory modules 1218 as a shared last-level cache.
[0139] In at least one embodiment, the processor cores 1202A-1202N are homogeneous cores executing a common instruction set architecture. In at least one embodiment, the processor cores 1202A-1202N are heterogeneous with respect to the instruction set architecture (ISA), wherein one or more of the processor cores 1202A-1202N execute a common instruction set, while one or more of the other cores of the processor cores 1202A-1202N execute a subset of a common instruction set or a different instruction set. In at least one embodiment, the processor cores 1202A-1202N are heterogeneous with respect to the microarchitecture, wherein one or more cores exhibiting relatively higher power consumption are coupled with one or more high-performance cores exhibiting lower power consumption.In at least one embodiment, the 1200 processor can be implemented on one or more chips or as an integrated SoC circuit.
[0140] The inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 715 are provided below in conjunction with Fig. 7A and / or 7B provided. In at least one embodiment, parts or all of the inference and / or training logic 715 may be implemented in the processor 1200. For example, in at least one embodiment, the training and / or inference techniques described herein may utilize one or more of the ALUs, graphics core(s) 1202A-1202N, or other components embodied in the graphics processor 1208. Fig. 12. Furthermore, in at least one embodiment, the inference and / or training operations described herein can be performed using logic that differs from that described in Fig. 7A or Fig. The logic illustrated in Figure 7B differs from that of other embodiments. In at least one embodiment, weight parameters can be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALUs of the 1200 graphics processor to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0141] Such components can be used to generate synthetic data that mimics error cases in a network training process, which can help improve network performance while limiting the amount of synthetic data to avoid overfitting.
[0142] Fig. Figure 13 is an exemplary data flow diagram for a process 1300 of generating and deploying an image processing and inference pipeline, according to at least one embodiment. In at least one embodiment, the process 1300 can be provided for use with imaging devices, processing devices, and / or other types of devices in one or more facilities 1302. The process 1300 can be executed in a training system 1304 and / or a deployment system 1306. In at least one embodiment, the training system 1304 can be used to perform training, deployment, and implementation of machine learning models (e.g., neural networks, object detection algorithms, computer vision algorithms, etc.) for use in the deployment system 1306.In at least one embodiment, the deployment system 1306 can be configured to offload processing and computing resources to a distributed computing environment in order to reduce infrastructure requirements at the facility 1302. In at least one embodiment, one or more applications in a pipeline can use or call services (e.g., inference, visualization, computation, AI, etc.) of the deployment system 1306 during application execution.
[0143] In at least one embodiment, some applications used in advanced processing and inference pipelines can use machine learning models or other AI to perform one or more processing steps. In at least one embodiment, machine learning models can be trained at the facility 1302 using data 1308 (such as imaging data) generated at the facility 1302 (and stored on one or more Picture Archiving and Communication System (PACS) servers at the facility 1302), they can be trained using imaging or sequencing data 1308 from another facility or facilities, or a combination thereof.In at least one embodiment, the training system 1304 can be used to provide applications, services and / or other resources for generating functional, deployable machine learning models for the deployment system 1306.
[0144] In at least one embodiment, a model registry 1324 can be backed up by object storage, thereby supporting versioning and object metadata. In at least one embodiment, the object storage can be accessed, for example, via cloud storage (e.g., Cloud 1426 from Fig. 14) compatible application programming interface (API) from within a cloud platform. In at least one embodiment, machine learning models within the model registry 1324 can be uploaded, listed, modified, or deleted by developers or partners of a system interacting with an API. In at least one embodiment, an API can provide access to procedures that allow users with appropriate permissions to associate models with applications, so that models can be executed as part of the execution of containerized instantiations of applications.
[0145] In at least one embodiment, a training pipeline 1404 ( Fig. 14) include a scenario in which the facility 1302 trains its own machine learning model or has an existing machine learning model that needs to be optimized or updated. In at least one embodiment, imaging data 1308 generated by an imaging device(s), sequencing devices, and / or other device types can be received. In at least one embodiment, after imaging data 1308 has been received, AI-assisted annotation 1310 can be used to assist in generating annotations according to the imaging data 1308, which are to be used as ground-truth data for a machine learning model. In at least one embodiment, AI-assisted annotation 1310 can include one or more machine learning models (e.g., convolutional neural networks (CNNs)) that can be trained to generate annotations according to certain types of imaging data 1308 (e.g.,to generate AI-supported annotations 1310 (from certain devices). In at least one embodiment, AI-supported annotations 1310 can then be used directly or adapted or fine-tuned using an annotation tool to generate ground-truth data. In at least one embodiment, the AI-supported annotations 1310, labeled clinical data 1312, or a combination thereof can be used as ground-truth data to train a machine learning model. In at least one embodiment, a trained machine learning model can be designated as an output model 1316 and used by the deployment system 1306 as described herein.
[0146] In at least one embodiment, the training pipeline 1404 ( Fig. 14) include a scenario in which the facility 1302 requires a machine learning model for use in performing one or more processing tasks for one or more applications in the deployment system 1306, but the facility 1302 may not currently have such a machine learning model (or may not have a model that is optimized, efficient, or effective for such purposes). In at least one embodiment, an existing machine learning model can be selected from the model registry 1324. In at least one embodiment, the model registry 1324 may contain machine learning models trained to perform a range of different inference tasks on imaging data. In at least one embodiment, machine learning models in the model registry 1324 may be trained on imaging data from facilities that are different from the facility 1302 (e.g.,(remote facilities). In at least one embodiment, machine learning models may be trained on imaging data from one location, two locations, or any number of locations. In at least one embodiment, when training on imaging data from a specific location, the training may take place at that location, or at least in a manner that protects the confidentiality of imaging data or restricts the transfer of imaging data from the user's own premises. In at least one embodiment, after a model has been trained—or partially trained—at a location, a machine learning model may be added to model registry 1324.In at least one embodiment, a machine learning model can then be retrained or updated on any number of other facilities, and a retrained or updated model can be made available in the model registry 1324. In at least one embodiment, a machine learning model can then be selected from the model registry 1324—and designated as the output model 1316—and can then be used in the deployment system 1306 to perform one or more processing tasks for one or more applications of a deployment system.
[0147] In at least one embodiment, the training pipeline 1404 ( Fig. 14) include a scenario in which the facility 1302 requires a machine learning model for use in performing one or more processing tasks for one or more applications in the deployment system 1306, but the facility 1302 may not currently have such a machine learning model (or may not have a model that is optimized, efficient, or effective for such purposes). In at least one embodiment, a machine learning model selected from the model registry 1324 may not be fine-tuned or optimized for imaging data 1308 generated at the facility 1302 due to differences in populations, robustness of training data used to train a machine learning model, diversity of anomalies in training data, and / or other problems with training data.In at least one embodiment, AI-assisted annotation 1310 can be used to support the generation of annotations corresponding to the imaging data 1308, which are to be used as ground-truth data for retraining or updating a machine learning model. In at least one embodiment, labeled data 1312 can be used as ground-truth data for training a machine learning model. In at least one embodiment, retraining or updating a machine learning model can be referred to as model training 1314. In at least one embodiment, the model training 1314—e.g., AI-assisted annotations 1310, labeled clinical data 1312, or a combination thereof—can be used as ground-truth data for retraining or updating a machine learning model.In at least one embodiment, a trained machine learning model can be designated as an output model 1316 and used by the deployment system 1306 as described herein.
[0148] In at least one embodiment, the deployment system 1306 can include software 1318, services 1320, hardware 1322, and / or other components, features, and functionality. In at least one embodiment, the deployment system 1306 can include a software "stack" such that software 1318 can be built on top of services 1320 and use services 1320 to perform some or all of the processing tasks, and services 1320 and software 1318 can be built on top of hardware 1322 and use hardware 1322 to perform the processing, storage, and / or computing tasks of the deployment system 1306. In at least one embodiment, the software 1318 can include any number of distinct containers, each container capable of executing an instantiation of an application.In at least one embodiment, each application can perform one or more processing tasks in an advanced processing and inference pipeline (e.g., inference, object detection, feature detection, segmentation, image enhancement, calibration, etc.). In at least one embodiment, an extended processing and inference pipeline can be defined based on selections of various containers that are desired or required for processing imaging data 1308, in addition to containers that receive and configure imaging data for use by each container and / or for use by the facility 1302 after processing by a pipeline (e.g., to convert outputs back into a usable data type). In at least one embodiment, a combination of containers within the software 1318 (e.g.,which constitute a pipeline) are referred to as a virtual instrument (as described in more detail herein), and a virtual instrument can use services 1320 and hardware 1322 to perform some or all of the processing tasks of applications instantiated in containers.
[0149] In at least one embodiment, a data processing pipeline can receive input data (e.g., imaging data 1308) in a specific format in response to an inference request (e.g., a request from a user of a deployment system 1306). In at least one embodiment, input data can represent one or more images, video, and / or other data representations generated by one or more imaging devices. In at least one embodiment, data can undergo preprocessing as part of the data processing pipeline to prepare the data for processing by one or more applications.In at least one embodiment, post-processing can be performed on an output from one or more inference tasks or other processing tasks of a pipeline to prepare output data for a subsequent application and / or to prepare output data for transmission and / or use by a user (e.g., in response to an inference request). In at least one embodiment, inference tasks can be performed by one or more machine learning models, such as trained or deployed neural networks, which may include output models 1316 of the training system 1304.
[0150] In at least one embodiment, tasks of the data processing pipeline can be encapsulated in one or more containers, each representing a discrete, fully functional instantiation of an application and virtualized computing environment capable of referencing machine learning models. In at least one embodiment, containers or applications can be published in a private area (e.g., a restricted access area) from a container registry (described in more detail herein), and trained or deployed models can be stored in the model registry 1324 and assigned to one or more applications. In at least one embodiment, images from applications (e.g.,Container images) are available in a container registry, and after it has been selected by a user from a container registry for use in a pipeline, an image can be used to generate a container for instantiating an application for use by a user system.
[0151] In at least one embodiment, developers (e.g., software developers, clinicians, physicians, etc.) can develop, publish, and store applications (e.g., as containers) for performing image processing and / or inference on supplied data. In at least one embodiment, the development, publication, and / or storage can be performed using a software development kit (SDK) associated with a system (e.g., to ensure that a developed application and / or container is consistent with or compatible with a system). In at least one embodiment, an application being developed can be deployed locally (e.g., at a first facility, to data from a first facility) as a system (e.g., System 1400 of) using an SDK that can support at least some of the services 1320. Fig. 14) be tested. In at least one embodiment, because DICOM objects can contain between one and hundreds of images or other data types, and due to variations in the data, a developer may be responsible for managing (e.g., defining constructs, incorporating preprocessing into an application, etc.) the extraction and preparation of incoming data. In at least one embodiment, after being validated by System 1400 (e.g., for precision), an application may be available in a container registry for selection and / or implementation by a user to perform one or more data processing tasks at a user's facility (e.g., a second facility).
[0152] In at least one embodiment, developers can then make applications or containers available over a network for access and use by users of a system (e.g., System 1400 from Fig. 14) share. In at least one embodiment, completed and validated applications or containers can be stored in a container registry, and associated machine learning models can be stored in the model registry 1324. In at least one embodiment, a requesting entity—providing an inference or image processing request—can search a container registry and / or model registry 1324 for an application, container, dataset, machine learning model, etc., select a desired combination of elements to include in a data processing pipeline, and send an image processing request.In at least one embodiment, a request may include input data (and in some examples, associated patient data) necessary to execute a request, and / or it may include a selection of one or more applications and / or machine learning models to be executed when processing a request. In at least one embodiment, a request may be forwarded to one or more components of the deployment system 1306 (e.g., a cloud) to perform processing in the data processing pipeline. In at least one embodiment, the processing by the deployment system 1306 may involve referencing selected elements (e.g., applications, containers, models, etc.) from a container registry and / or a model registry 1324. In at least one embodiment, after results have been generated by a pipeline, results may be sent to a user for reference (e.g.,(for viewing in a viewing application suite running on a local workstation or on-site terminal).
[0153] In at least one embodiment, services 1320 can be used to support the processing or execution of applications or containers in pipelines. In at least one embodiment, services 1320 can include computing services, artificial intelligence services (AI services), visualization services, and / or other service types. In at least one embodiment, services 1320 can provide functionality common to one or more applications in software 1318, so that functionality can be abstracted to a service that can be called or used by applications. In at least one embodiment, functionality provided by services 1320 can run dynamically and more efficiently, while also scaling well by allowing applications to process data in parallel (e.g., using a parallel computing platform 1430). Fig. 14)) to process. In at least one embodiment, the service 1320 can be shared between and by different applications, instead of each application sharing the same functionality offered by the service 1320 requiring its own instance of the service 1320. In at least one embodiment, services can include an inference server or inference engine that can be used as non-restrictive examples for performing acquisition or segmentation tasks. In at least one embodiment, a model training service can be included that can provide machine learning model training and / or retraining capabilities. In at least one embodiment, a data augmentation service can also be included that can provide GPU-accelerated data (e.g., DICOM, RIS, CIS, REST-compatible, RPC, raw, etc.) extraction, resizing, scaling, and / or other augmentation.In at least one embodiment, a visualization service can be used that can add image rendering effects—such as ray tracing, rasterization, denoising, sharpening, etc.—to enhance realism in two-dimensional (2D) and / or three-dimensional (3D) models. In at least one embodiment, virtual instrument services can be included that provide beamforming, segmentation, inference, imaging, and / or support for other applications within pipelines of virtual instruments.
[0154] In at least one embodiment, where a service 1320 includes an AI service (e.g., an inference service), one or more machine learning models can be executed by calling an inference service (e.g., an inference server) (e.g., as an API call) to execute one or more machine learning models, or to process them, as part of application execution. In at least one embodiment, if another application includes one or more machine learning models for segmentation tasks, an application can call an inference service to execute machine learning models to perform processing operations associated with one or more segmentation tasks.In at least one embodiment, software 1318, which implements an advanced processing and inference pipeline that includes a segmentation application and an anomaly detection application, can be streamlined because each application can call the same inference service to perform one or more inference tasks.
[0155] In at least one embodiment, the hardware 1322 can include GPUs, CPUs, DPUs, graphics cards, an AI / deep learning system (e.g., an AI supercomputer such as NVIDIA's DGX), a cloud platform, or a combination thereof. In at least one embodiment, different types of hardware 1322 can be used to provide efficient, purpose-built support for software 1318 and services 1320 in the deployment system 1306. In at least one embodiment, the use of GPU processing can be implemented for local processing (e.g., in the facility 1302), within an AI / deep learning system, in a cloud system, and / or in other processing components of the deployment system 1306 to improve the efficiency, precision, and effectiveness of image processing and generation.In at least one embodiment, software 1318 and / or services 1320 can be optimized for GPU processing with respect to deep learning, machine learning, and / or high-performance computing, as non-limiting examples. In at least one embodiment, at least a portion of the computing environment of the deployment system 1306 and / or the training system 1304 can be run in a data center, in one or more supercomputers or high-performance computing systems, using GPU-optimized software (e.g., a hardware and software combination of NVIDIA's DGX system). In at least one embodiment, hardware 1322 can include any number of GPUs that can be called upon to perform parallel data processing as described herein. In at least one embodiment, a cloud platform can further include GPU processing for GPU-optimized execution of deep learning tasks, machine learning tasks, or other computing tasks.In at least one embodiment, the cloud platform can further include DPU processing to transfer data received over a network and / or through a network controller or other network interface directly to one or more GPUs (e.g., their memory). In at least one embodiment, the cloud platform (e.g., NVIDIA's NGC) can run as a hardware abstraction and scaling platform using one or more AI / deep learning supercomputers and / or GPU-optimized software (e.g., as provided on NVIDIA's DGX systems). In at least one embodiment, the cloud platform can integrate an application container clustering or orchestration system (e.g., Kubernetes) across multiple GPUs to enable seamless scaling and load balancing.
[0156] Fig. Figure 14 is a system representation for an exemplary system 1400 for generating and deploying an imaging deployment pipeline, according to at least one embodiment. In at least one embodiment, the system 1400 can be used to perform the process 1300 of Fig. 13 and / or other processes that include advanced processing and inference pipelines. In at least one embodiment, the system 1400 may include a training system 1304 and a deployment system 1306. In at least one embodiment, the training system 1304 and the deployment system 1306 may be implemented using the software 1318, services 1320, and / or hardware 1322 as described herein.
[0157] In at least one embodiment, System 1400 (e.g., Training System 1304 and / or Deployment System 1306) can be implemented in a cloud computing environment (e.g., using Cloud 1426). In at least one embodiment, System 1400 can be implemented locally with respect to a healthcare facility or as a combination of cloud and local computing resources. In at least one embodiment, access to APIs in Cloud 1426 can be restricted to authorized users by established security measures or protocols. In at least one embodiment, a security protocol can include web tokens that can be signed by an authentication service (e.g., AuthN, AuthZ, Gluecon, etc.) and carry the corresponding authorization.In at least one embodiment, APIs of virtual instruments (described herein), or other instantiations of System 1400, can be restricted to a set of public IPs that have been audited or authorized for interaction.
[0158] In at least one embodiment, various components of the System 1400 can communicate between themselves and with each other using a range of different network types, including, without limitation, local area networks (LANs) and / or wide area networks (WANs) via wired and / or wireless communication protocols. In at least one embodiment, communication between devices and components of the System 1400 (e.g., for transmitting inference requests, receiving results of inference requests, etc.) can be carried out via one or more data buses, wireless data protocols (WiFi), wired data protocols (e.g., Ethernet), etc.
[0159] In at least one embodiment, the training system 1304 can be combined with training pipelines 1404, similar to those described herein with respect to Fig. 13. In at least one embodiment, training pipelines 1404 can be used when one or more machine learning models in deployment pipelines 1410 are to be used by the deployment system 1306 to train or retrain one or more (e.g., pre-trained) models and / or to implement one or more of the pre-trained models 1406 (e.g., without the need for retraining or updating). In at least one embodiment, one or more output models 1316 can be generated as a result of the training pipelines 1404. In at least one embodiment, training pipelines 1404 can include a number of processing steps, such as conversion or adaptation of imaging data (or other input data). In at least one embodiment, different training pipelines 1404 can be used for different machine learning models used by the deployment system 1306.In at least one embodiment, the training pipeline 1404 can be similar to a first one with respect to . Fig. The example described in point 13, used for a first machine learning model, the training pipeline 1404 can be similar to a second one in terms of Fig. The example described in section 13 can be used for a second machine learning model, and the training pipeline 1404 can be similar to a third in terms of Fig. The example described in section 13 can be used for a third machine learning model. In at least one embodiment, any combination of tasks within a training system 1304 can be used, depending on what is required for each machine learning model. In at least one embodiment, one or more machine learning models can already be trained and ready for use, so that machine learning models may not require any processing by the training system 1304 and can be implemented by the deployment system 1306.
[0160] In at least one embodiment, the one or more output models 1316 and / or pretrained models 1406 can include any type of machine learning model, depending on the implementation or embodiment. In at least one embodiment, machine learning models used by the System 1400 may, without restriction, include one or more machine learning models employing linear regression, logistic regression, decision trees, support vector machines (SVMs), the naive Bayes classifier, k-nearest neighbors (Knn), K-means clustering, random forest, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., auto-encoders, convolutional, recurrent, perceptrons, long / short-term / memory (LSTM), Hopfield, Boltzmann, deep-belief, unfolding, generative adversarial, liquid-state machine, etc.), and / or other types of machine learning models.
[0161] In at least one embodiment, the training pipelines 1404 AI-supported annotation, as detailed herein with regard to at least Fig. described in 16B, include. In at least one embodiment, labeled data 1312 (e.g., traditional annotation) can be generated by a number of techniques. In at least one embodiment, labels or other annotations can be generated within a drawing program (e.g., an annotation program), a computer-aided design (CAD) program, a labeling program, another type of program suitable for generating annotations or labels for ground truth, and / or, in some examples, drawn by hand. In at least one embodiment, ground truth data can be produced as follows: synthetically (e.g., generated from computer models or renderings), real (e.g., conceived and produced from real-world data), machine-automated (e.g., using feature analysis and learning to extract features from data and then generate labels), or human-annotated (e.g.,The labeler, or annotation expert, defines the location of the labels), and / or a combination thereof. In at least one embodiment, for each instance of imaging data 1308 (or a type of data used by machine learning models), there can be corresponding ground-truth data generated by the training system 1304. In at least one embodiment, AI-assisted annotation can be performed as part of the deployment pipelines 1410; either in addition to or instead of AI-assisted annotation included in training pipelines 1404. In at least one embodiment, the system 1400 can include a multi-layer platform that includes a software layer (e.g., software 1318) for diagnostic applications (or other application types) capable of performing one or more medical imaging and diagnostic functions. In at least one embodiment, the system 1400 (e.g.,(via encrypted links) the system may be communicatively coupled to PACS server networks of one or more institutions. In at least one embodiment, the System 1400 may be configured to access or reference data from PACS servers to perform operations such as training machine learning models, deploying machine learning models, image processing, inference, and / or other operations.
[0162] In at least one embodiment, a software layer can be implemented as a secure, encrypted, and / or authenticated API through which applications or containers can be called (e.g., invoked) from an external environment(s) (e.g., the facility 1302). In at least one embodiment, applications can then invoke or execute one or more services 1320 to perform computational, AI, or visualization tasks assigned to the respective applications, and software 1318 and / or services 1320 can utilize hardware 1322 to perform processing tasks in an effective and efficient manner.
[0163] In at least one embodiment, the deployment system 1306 can execute the deployment pipelines 1410. In at least one embodiment, the deployment pipelines 1410 can include any number of applications that can be applied sequentially, non-sequentially, or otherwise to imaging data (and / or other data types) generated by imaging devices, sequencing devices, genomics devices, etc. – including AI-assisted annotation, as described above. In at least one embodiment, a deployment pipeline 1410, as described herein, can be designated for an individual device as a virtual instrument for a device (e.g., a virtual ultrasound instrument, a virtual CT scan instrument, a virtual sequencing instrument, etc.).In at least one embodiment, there can be more than one input pipeline 1410 for a single device, depending on the desired information from the data generated by the device. In at least one embodiment, if anomaly detection from an MRI machine is desired, there can be a first input pipeline 1410, and if image enhancement for the output of an MRI machine is desired, there can be a second input pipeline 1410.
[0164] In at least one embodiment, an image generation application may include a processing task that involves the use of a machine learning model. In at least one embodiment, a user may choose to use their own machine learning model or select a machine learning model from a model registry. In at least one embodiment, a user may implement their own machine learning model or select a machine learning model for inclusion in an application to perform a processing task. In at least one embodiment, applications may be selectable and customizable, and by defining constructs of applications, the deployment and implementation of applications for a particular user are presented as a more seamless user experience.In at least one embodiment, by utilizing other features of the system 1400 - such as services 1320 and hardware 1322 - deployment pipelines 1410 can be even more user-friendly, provide simpler integration and produce more accurate, efficient and timely results.
[0165] In at least one embodiment, the deployment system 1306 may include a user interface 1414 (e.g., a graphical user interface, a web interface, etc.) that can be used to select applications for inclusion in the deployment pipeline(s) 1410, to arrange applications, to modify or change applications or parameters or constructs thereof, to use and interact with the deployment pipeline(s) 1410 during setup and / or deployment, and / or to otherwise interact with the deployment system 1306. In at least one embodiment, the user interface 1414 (or another user interface), although not illustrated in relation to the training system 1304, may be used to select models for use in the deployment system 1306, to select models for training or retraining in the training system 1304, and / or to otherwise interact with the training system 1304.
[0166] In at least one embodiment, a pipeline manager 1412 can be used, in addition to an application orchestration system 1428, to manage the interaction between applications or containers of deployment pipelines 1410 and services 1320 and / or hardware 1322. In at least one embodiment, the pipeline manager 1412 can be configured to allow application-to-application, application-to-service 1320, and / or application-to-service or hardware 1322 interactions. In at least one embodiment, the pipeline manager 1412, although illustrated as included in the software 1318, which is not to be interpreted as restrictive, and in some examples (e.g., as in Fig. (12 illustrated) in services 1320. In at least one embodiment, the application orchestration system 1428 (e.g., Kubernetes, Docker, etc.) can include a container orchestration system that can group applications into containers as logical units for coordination, management, scaling, and deployment. In at least one embodiment, each application can be run in a self-contained environment (e.g., at a kernel level) by assigning applications from a deployment pipeline(s) 1410 (e.g., a reconstruction application, a segmentation application, etc.) to individual containers to increase speed and efficiency.
[0167] In at least one embodiment, each application and / or container (or each image thereof) can be developed, modified, and deployed individually (e.g., a first user or developer can develop, modify, and deploy a first application, and a second user or developer can develop, modify, and deploy a second application separately from the first user or developer). This allows the focus and attention to be concentrated on a task of a single application and / or container without being hindered by tasks of one or more other applications or containers. In at least one embodiment, communication and cooperation between different containers and applications can be supported by the pipeline manager 1412 and the application orchestration system 1428.In at least one embodiment, the application orchestration system 1428 and / or the pipeline manager 1412 can, as long as an expected input and / or output from each container or application is known to a system (e.g., based on constructs of applications or containers), enable communication among and between, and the sharing of resources among and between, each of the applications or containers. In at least one embodiment, since one or more of the applications or containers in a deployment pipeline(s) 1410 can share the same services and resources, the application orchestration system 1428 can orchestrate, balance, and determine the sharing of services or resources between and among different applications or containers.In at least one embodiment, a scheduler can be used to track resource requests from applications or containers, the current or planned use of these resources, and resource availability. In at least one embodiment, a scheduler can therefore allocate resources to different applications and distribute resources between and among applications with respect to the requirements and availability of a system. In some examples, a scheduler (and / or another component of the application orchestration system 1428) can determine resource availability and distribution based on constraints imposed on a system (e.g., user constraints), such as quality of service (QoS), urgency of data output needs (e.g., to determine whether to perform real-time or delayed processing), and so on.
[0168] In at least one embodiment, services 1320, which are used and shared by applications or containers in the deployment system 1306, can include compute services 1416, AI services 1418, visualization services 1420, and / or other service types. In at least one embodiment, applications can call (e.g., execute) one or more services 1320 to perform processing operations for an application. In at least one embodiment, the compute services 1416 can be used by applications to perform supercomputing or other high-performance computing (HPC) tasks. In at least one embodiment, one or more compute services 1416 can be used to perform parallel processing (e.g., using a parallel computing platform 1430) to process data across one or more applications and / or one or more tasks of a single application, essentially simultaneously.In at least one embodiment, the parallel computing platform 1430 (e.g., NVIDIA's CUDA) can enable general-purpose computing on GPUs (GPGPU) (e.g., GPUs 1422). In at least one embodiment, a software layer of the parallel computing platform 1430 can provide access to virtual instruction sets and parallel computing elements of GPUs for executing computing kernels. In at least one embodiment, the parallel computing platform 1430 can include main memory, and in at least one embodiment, main memory can be shared between and among multiple containers and / or between and among different processing tasks within a single container.In at least one embodiment, interprocess communication (IPC) calls can be generated for multiple containers and / or for multiple processes within a container to use the same data from a shared segment of the main memory of the Parallel Computing Platform 1430 (e.g., when multiple different stages of an application or multiple applications process the same information). In at least one embodiment, instead of copying data and moving it to different locations in main memory (e.g., a read / write operation), the same data can be used at the same location in main memory for any number of processing tasks (e.g., at the same time, at different times, etc.).In at least one embodiment, when data is used to generate new data as a result of processing, this information can be stored in a new data location and shared between different applications. In at least one embodiment, the data location and a location of updated or modified data can be part of a definition of how a payload is understood within containers.
[0169] In at least one embodiment, the AI services 1418 can be used to perform inference services for executing a machine learning model(s) assigned to an application (e.g., tasked with performing one or more processing tasks of an application). In at least one embodiment, the AI services 1418 can utilize the AI system 1424 to execute a machine learning model (e.g., neural networks such as CNNs) for segmentation, reconstruction, object detection, feature detection, classification, and / or other inference tasks. In at least one embodiment, applications of the one or more deployment pipelines 1410 can use one or more output models 1316 from the training system 1304 and / or other application models to perform inference on imaging data.In at least one embodiment, two or more examples of inference using the application orchestration system 1428 (e.g., a scheduler) may be available. In at least one embodiment, a first category may include a high-priority / low-latency path that can achieve higher quality-of-service agreements, such as for performing inference on urgent requests in emergencies or for a radiologist during diagnosis. In at least one embodiment, a second category may include a standard-priority path that can be used for requests that may not be urgent or when the analysis can be performed at a later time. In at least one embodiment, the application orchestration system 1428 may distribute resources (e.g., services 1320 and / or hardware 1322) based on priority paths for different inference tasks of the AI services 1418.
[0170] In at least one embodiment, the shared memory can be connected to the AI services 1418 within the system 1400. In at least one embodiment, shared memory can be operated as a cache (or other type of storage device) and used to process inference requests from applications. In at least one embodiment, when an inference request is sent, a request can be received by a set of API instances of the deployment system 1306, and one or more instances can be selected (e.g., for best fit, for load balancing, etc.) to process a request.In at least one embodiment, to process a request, a request can be entered into a database; a machine learning model can be found from the model register 1324 if it is not already in a cache; a validation step can ensure that the appropriate machine learning model is loaded into a cache (e.g., shared memory); and / or a copy of a model can be saved in a cache. In at least one embodiment, a scheduler (e.g., a pipeline manager 1412) can be used to start an application referenced in a request if no application is already running or if there are not enough instances of an application. In at least one embodiment, an inference server can be started if no inference server has already been started to execute a model. Any number of inference servers can be started per model.In at least one embodiment, models can be cached in a pull model where inference servers are clustered, if load balancing is advantageous. In at least one embodiment, inference servers can be statically loaded on corresponding distributed servers.
[0171] In at least one embodiment, inference can be performed using an inference server running in a container. In at least one embodiment, an instance of an inference server can be associated with a model (and optionally with multiple versions of a model). In at least one embodiment, if an instance of an inference server does not exist when a request to perform inference on a model is received, a new instance can be loaded. In at least one embodiment, a model can be submitted to an inference server when it is started, so that the same container can be used to serve different models, as long as the inference server is running as a different instance.
[0172] In at least one embodiment, during application execution, an inference request for a given application can be received, and a container (which, for example, hosts an instance of an inference server) can be loaded (if it is not already) and a startup process can be invoked. In at least one embodiment, preprocessing logic in a container can load incoming data, decode it, and / or perform other additional preprocessing on it (for example, using one or more CPUs and / or one or more GPUs and / or one or more DPUs). In at least one embodiment, after data has been prepared for inference, a container can perform inference on data as needed. In at least one embodiment, this can involve a single inference call on an image (for example, a hand X-ray) or it can require inference on hundreds of images (for example, a breast CT scan).In at least one embodiment, an application can summarize results before completion, which can include, without limitation, generating a single confidence score, pixel-level segmentation, voxel-level segmentation, generating a visualization, or generating text to summarize findings. In at least one embodiment, different models or applications can be assigned different priorities. For example, some models can have a real-time priority (TAT < 1 min), while others can have a lower priority (e.g., TAT < 12 min). In at least one embodiment, model execution times can be measured by a requesting institution or entity and can include partner network traversal time as well as execution on an inference service.
[0173] In at least one embodiment, the transmission of requests between Services 1320 and inference applications can be hidden behind a software development kit (SDK), and robust transport can be provided through a queue. In at least one embodiment, a request is placed in a queue via an API for a single application / tenant ID combination, and an SDK will retrieve a request from the queue and pass it to an application. In at least one embodiment, a queue name can be provided in an environment from which an SDK will retrieve it. In at least one embodiment, asynchronous communication through a queue can be useful because it allows each instance of an application to pick up work as it becomes available.The results can be transferred back through a queue to ensure no data is lost. In at least one embodiment, queues can also provide the ability to segment work, with highest-priority work going to a queue with the most associated instances of an application, while lowest-priority work going to a queue with a single associated instance, which processes tasks in the order they were received. In at least one embodiment, an application can run on a GPU-accelerated instance generated in Cloud 1426, and an inference service can perform inference on a GPU.
[0174] In at least one embodiment, the visualization services 1420 can be used to generate visualizations for viewing outputs from applications and / or the one or more deployment pipelines 1410. In at least one embodiment, GPUs 1422 can be used by the visualization services 1420 to generate visualizations. In at least one embodiment, rendering effects, such as ray tracing, can be implemented by the visualization services 1420 to generate higher-quality visualizations. In at least one embodiment, visualizations can include, without limitation, 2D image rendering, 3D volume rendering, 3D volume reconstruction, 2D tomography layers, virtual reality displays, augmented reality displays, etc. In at least one embodiment, virtualized environments can be used to provide a virtual interactive display or environment (e.g.,to generate a virtual environment for interaction by users of a system (e.g., doctors, nurses, radiologists, etc.). In at least one embodiment, the visualization services 1420 may include an internal visualizer, cinematography, and / or other rendering or image processing capabilities or functionality (e.g., ray tracing, rasterization, internal optics, etc.).
[0175] In at least one embodiment, the hardware 1322 may include GPUs 1422, AI system 1424, cloud 1426, and / or other hardware used to run the training system 1304 and / or deployment system 1306. In at least one embodiment, GPUs 1422 (e.g., NVIDIA TESLA and / or QUADRO GPUs) may include any number of GPUs used to perform processing tasks by the compute services 1416, AI services 1418, visualization services 1420, other services, and / or any features or functionality of software 1318. For example, with regard to AI services, 1418 GPUs 1422 can be used to perform preprocessing on imaging data (or other data types used by machine learning models), postprocessing on outputs of machine learning models, and / or to perform inference (e.g., to run machine learning models).In at least one embodiment, the Cloud 1426, the AI system 1424, and / or other components of the System 1400 can utilize GPUs 1422. In at least one embodiment, the Cloud 1426 can include a GPU-optimized platform for deep learning tasks. In at least one embodiment, the AI system 1424 can utilize GPUs, and the Cloud 1426—or at least a part that performs deep learning or inference—can be executed using one or more AI systems 1424. As such, although the Hardware 1322 is illustrated as discrete components, this is not to be construed as restrictive, and any components of the Hardware 1322 can be combined with or utilized by other components of the Hardware 1322.
[0176] In at least one embodiment, the AI system 1424 may include a dedicated computing system (e.g., a supercomputer or an HPC) configured for inference, deep learning, machine learning, and / or other artificial intelligence tasks. In at least one embodiment, the AI system 1424 (e.g., NVIDIA's DGX) may include GPU-optimized software (e.g., a software stack) that can be run using a variety of GPUs 1422, in addition to DPUs, CPUs, RAM, storage, and / or other components, features, or functionality. In at least one embodiment, one or more AI systems 1424 may be deployed in the cloud 1426 (e.g., in a data center) to perform some or all of the AI-based processing tasks of the system 1400.
[0177] In at least one embodiment, the Cloud 1426 can include a GPU-accelerated infrastructure (e.g., NVIDIA's NGC) that can provide a GPU-optimized platform for performing processing tasks of the System 1400. In at least one embodiment, the Cloud 1426 can include one or more AI systems 1424 for performing one or more AI-based tasks of the System 1400 (e.g., as a hardware abstraction and scaling platform). In at least one embodiment, the Cloud 1426 can be integrated into an application orchestration system 1428 utilizing multiple GPUs to enable seamless scaling and load balancing between and among applications and services 1320. In at least one embodiment, the cloud 1426 can have the task of performing at least some services 1320 of the system 1400, including computing services 1416, AI services 1418 and / or visualization services 1420, as described herein.In at least one embodiment, the Cloud 1426 can perform small and large batch inference (e.g., running NVIDIA's TENSOR RT), provide an accelerated parallel computing API and platform 1430 (e.g., NVIDIA's CUDA), run an application orchestration system 1428 (e.g., KUBERNETES), provide a graphics rendering API and platform (e.g., for ray tracing, 2D graphics, 3D graphics, and / or other rendering techniques to produce higher-quality cinematography), and / or provide other functionality for the System 1400.
[0178] Fig. Figure 15A illustrates a data flow diagram for a process 1500 for training, retraining, or updating a machine learning model according to at least one embodiment. In at least one embodiment, the process 1500 can be performed using, as a non-limiting example, the system 1400 of Fig. 14. In at least one embodiment, the process 1500 can utilize services 1320 and / or hardware 1322 of system 1400, as described herein. In at least one embodiment, refined models 1512 generated by the process 1500 can be executed by the deployment system 1306 for one or more containerized applications in deployment pipelines 1410.
[0179] In at least one embodiment, the model training 1314 can be the retraining or updating of an initial model 1504 (e.g., a pre-trained model) using new training data (e.g., new input data, such as a customer record 1506, and / or new ground truth data associated with the input data). In at least one embodiment, to retrain or update the initial model 1504, one or more output or loss layers of the initial model 1504 can be reset or deleted and / or replaced by one or more updated or new output or loss layers. In at least one embodiment, an initial model 1504 can be used with previously fine-tuned parameters (e.g.,Weights and / or biases) that remain from previous training, so that the training or retraining 1314 will not take as long or require as much processing as training a model from scratch. In at least one embodiment, during model training 1314, by readjusting or replacing output or loss layers of the initial model 1504, parameters for a new dataset can be updated and retuned based on loss calculations that affect the precision of the output or loss layers when generating predictions on a new customer dataset 1506 (e.g., image data 1308 from ). Fig. 13) are assigned.
[0180] In at least one embodiment, pre-trained models 1406 can be stored in a data store or registry (e.g., model registry 1324 of Fig. 13) are stored. In at least one embodiment, pre-trained models 1406 may have been trained at least partially at one or more facilities that are different from a facility performing the process 1500. In at least one embodiment, to protect the privacy and rights of patients, subjects, or clients at different facilities, pre-trained models 1406 may have been trained on-site using customer or patient data generated on-site. In at least one embodiment, pre-trained models 1406 may be trained using the Cloud 1426 and / or other hardware 1322; however, confidential, privacy-protected patient data may not be transferred to, used by, or accessible by any component of the Cloud 1426 (or other hardware outside the user's premises).In at least one embodiment, if the pre-trained Model 1406 is trained using patient data from more than one institution, it may have been trained individually for each institution before being trained on patient or customer data from another institution. In at least one embodiment, such as when customer or patient data has been cleared from privacy concerns (e.g., by waiver, for experimental use, etc.) or when customer or patient data is included in a public dataset, customer or patient data from any number of institutions may be used to train the pre-trained Model 1406 on-premises and / or off-premises, such as in a data center or other cloud computing infrastructure.
[0181] In at least one embodiment, when selecting applications for use in the deployment pipelines 1410, a user can also select machine learning models to be used for specific applications. In at least one embodiment, a user may not have a model to use, so a user may select a pre-trained model 1406 for use with an application. In at least one embodiment, the pre-trained model 1406 may not be optimized for generating accurate results on a customer dataset 1506 of a user's facility (e.g., based on patient diversity, demographics, types of medical imaging devices used, etc.).In at least one embodiment, before the pre-trained model 1406 is inserted into the deployment pipeline 1410 for use with one or more applications, the pre-trained model 1406 can be updated, retrained and / or fine-tuned for use at a particular facility.
[0182] In at least one embodiment, a user can select the pre-trained model 1406 that needs to be updated, retrained, and / or fine-tuned, and the pre-trained model 1406 can be designated as the initial model 1504 for a training system 1304 within the process 1500. In at least one embodiment, the customer data set 1506 (e.g., imaging data, genomic data, sequencing data, or other data types generated by devices at a facility) can be used to perform model training 1314 (which may, without limitation, include transfer learning) on the initial model 1504 to generate the refined model 1512. In at least one embodiment, ground-truth data corresponding to the customer data set 1506 can be generated by the training system 1304.In at least one embodiment, ground truth data can be at least partially generated by clinicians, scientists, physicians, medical professionals at an institution (e.g. as labeled clinical data 1312 from . Fig. 13) will be generated.
[0183] In at least one embodiment, the AI-powered annotation 1310 can be used in some examples to generate ground-truth data. In at least one embodiment, the AI-powered annotation 1310 (e.g., implemented using an AI-powered annotation SDK) can utilize machine learning models (e.g., neural networks) to generate suggested or predicted ground-truth data for a customer dataset. In at least one embodiment, the user can use annotation tools 1510 within a user interface (a graphical user interface (GUI)) on the computing device 1508.
[0184] In at least one embodiment, the user 1510 can interact with a GUI via the computing device 1508 to edit or fine-tune (auto-)annotations. In at least one embodiment, a polygon editing feature can be used to move the vertices of a polygon to more accurate or fine-tuned locations.
[0185] In at least one embodiment, once the customer dataset 1506 has associated ground-truth data, ground-truth data (e.g., from AI-assisted annotation, manual labeling, etc.) can be used during model training 1314 to generate a refined model 1512. In at least one embodiment, the customer dataset 1506 can be applied to the initial model 1504 any number of times, and ground-truth data can be used to update parameters of the initial model 1504 until an acceptable level of precision is achieved for the refined model 1512. In at least one embodiment, after the refined model 1512 has been generated, the refined model 1512 can be deployed within one or more deployment pipelines 1410 at a facility to perform one or more processing tasks related to medical imaging data.
[0186] In at least one embodiment, the refined model 1512 can be uploaded to the pretrained models 1406 in the model registry 1324 for selection by another facility. In at least one embodiment, this process can be performed at any number of facilities, allowing the refined model 1512 to be further refined on new datasets any number of times to generate a more universal model.
[0187] Fig. Figure 15B is an exemplary illustration of a client-server architecture 1532 for improving annotation tools with pre-trained annotation models, according to at least one embodiment. In at least one embodiment, the AI-supported annotation tools 1536 can be instantiated based on a client-server architecture 1532. In at least one embodiment, the AI-supported annotation tools 1536 can assist radiologists in imaging applications, for example, in identifying organs and anomalies. In at least one embodiment, imaging applications can include software tools that help the user 1510, as a non-restrictive example, to identify a few extreme points on a particular organ of interest in raw images 1534 (e.g., in a 3D MRI or CT scan) and automatically receive annotated results for all 2D sections of a particular organ.In at least one embodiment, results can be stored in a data working memory as training data 1538 and used (for example, and without limitation) as ground truth data for training. In at least one embodiment, a deep learning model can, for example, when the computing device 1508 sends extreme points for AI-assisted annotation 1310, receive this data as input and return inference results for a segmented organ or segmented anomaly. In at least one embodiment, pre-instantiated annotation tools, such as the AI-assisted annotation tool 1536B, can be used. Fig.15B, by making API calls (e.g., API call 1544) to a server, such as an annotation assistant server 1540, which may contain a set of pre-trained models 1542 stored, for example, in an annotation model registry. In at least one embodiment, an annotation model registry may store pre-trained models 1542 (e.g., machine learning models, such as deep learning models) that are pre-trained to perform AI-assisted annotation on a specific organ or anomaly. These models may further be updated using training pipelines 1404. In at least one embodiment, pre-installed annotation tools may be improved over time as new labeled clinical data 1312 are added.
[0188] Such components can be used to generate synthetic data that mimics error cases in a network training process, which can help improve network performance while limiting the amount of synthetic data to avoid overfitting.
[0189] Other variations are within the spirit of the present disclosure. Although disclosed techniques are susceptible to various modifications and alternative constructions, certain illustrated embodiments thereof are shown in the drawings and have been described in detail above. It is understood, however, that the disclosure is not intended to be limited to any particular disclosed form, but rather that it is intended to cover all modifications, alternative constructions, and equivalents that fall within the spirit and scope of protection of the disclosure, as defined in the attached claims.
[0190] The use of the terms "a," "an," and "one," and "the," "a," and similar references in connection with the description of disclosed embodiments (particularly in connection with the following claims) is to be interpreted as covering both the singular and the plural, unless otherwise specified herein or clearly contrary to the context, and not as defining a term. The terms "comprising," "having," "including," and "containing" are to be interpreted as open terms (meaning "including but not limited to"), unless otherwise specified. "Connected," when not modified and referring to physical connections, is to be interpreted as partially or completely contained in, attached to, or connected to one another, even if something is interposed.The enumerations of value ranges contained herein are intended merely as a shorthand method for referring individually to each value falling within the range, unless otherwise stated herein, and each individual value is included in the patent description as if it were listed individually herein. In at least one embodiment, the use of the term "set" (e.g., "a set of elements") or "subset," unless otherwise noted or in contradiction to the context, is to be understood as a non-empty collection comprising one or more members. Furthermore, unless otherwise noted or in contradiction to the context, the term "subset" of a corresponding set does not necessarily denote a proper subset of the corresponding set, but the subset and the corresponding set may be the same.
[0191] Subjunctive language, such as sentences of the form "at least one of A, B, and C" or "at least one of A, B, and C," unless expressly stated otherwise or otherwise clearly in conflict with the context, is understood otherwise with the appropriate context, as it is generally used to indicate that an element, term, etc., can be either A, B, or C, or a non-empty subset of the set of A, B, and C. For example, in an illustrative example of a set with three members, the subjunctive sentences "at least one of A, B, and C" and "at least one of A, B, and C" refer to one of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Thus, such subjunctive language is not generally intended to imply that certain embodiments require the presence of at least one of A, at least one of B, and at least one of C.Furthermore, unless otherwise stated or inconsistent with the context, the term "plurality" indicates the plural (e.g., "a plurality of elements" indicates multiple elements). In at least one embodiment, the number of elements in a plurality is at least two, but may be more if either explicitly stated or indicated by the context. Additionally, unless otherwise stated or clearly indicated by the context, the phrase "based on" means "at least partially based on" and not "exclusively based on".
[0192] Operations of the processes described herein may be performed in any suitable order unless otherwise specified herein or otherwise clearly contrary to the context. In at least one embodiment, a process such as the processes described herein (or variations and / or combinations thereof) is performed under the control of one or more computer systems configured with executable instructions and is implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed collectively on one or more processors, by hardware, or combinations thereof. In at least one embodiment, the code is stored on a computer-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions that can be executed by one or more processors.In at least one embodiment, a computer-readable storage medium is a non-volatile computer-readable storage medium that excludes transient signals (e.g., a propagating transient electrical or electromagnetic transmission) but includes a non-volatile data storage circuit (e.g., buffers, cache, and queues) within transceivers of transient signals. In at least one embodiment, code (e.g., executable code or source code) is stored on a set of one or more non-volatile computer-readable storage media containing executable instructions (or other working memory for storing executable instructions) which, when executed by one or more processors of a computer system (i.e., as a result of their execution), cause the computer system to perform operations described herein.In at least one embodiment, a set of non-volatile, computer-readable storage media comprises several non-volatile, computer-readable storage media, wherein one or more of the individual non-volatile storage media within the set lack all code, while several non-volatile, computer-readable storage media collectively store all code. In at least one embodiment, executable instructions are executed such that different instructions are executed by different processors.
[0193] Accordingly, computer systems in at least one embodiment are configured to implement one or more services that individually or jointly perform operations of the methods described herein, and such computer systems are configured with appropriate hardware and / or software that enables the execution of operations. Furthermore, a computer system implementing at least one embodiment of the present disclosure is a single device, and in another embodiment, a distributed computer system comprising several devices that operate differently, such that the distributed computer system performs the operations described herein and such that a single device does not perform all operations.
[0194] The use of any and all examples or exemplary formulations (e.g., "such as") provided herein is intended solely to better illustrate embodiments of the disclosure and does not constitute a limitation of the scope of protection of the disclosure unless otherwise claimed. No expression in the patent description should be interpreted as indicating that an unclaimed element is essential to the practice of the disclosure.
[0195] All references, including publications, patent applications and patents cited herein, are hereby incorporated by reference to the same extent as if each reference were individually and specifically indicated for inclusion by reference and set forth herein in their entirety.
[0196] In the description and claims, the terms "coupled" and "connected," along with their derivatives, may be used. It is understood that these terms are not synonymous. Rather, in certain examples, "connected" or "coupled" may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. "Coupled," however, can also mean that two or more elements are not in direct contact with each other but nevertheless cooperate or interact.
[0197] Unless expressly stated otherwise, terms such as "processing", "calculating", "calculating", "determining" or the like are understood to refer throughout the patent description to actions and / or processes of a computer or computer system or similar electronic computing device that manipulate and / or transform data represented as physical or electronic quantities in registers and / or working memories of the computer system into other data represented in a similar manner as physical quantities in working memories, registers or other such information storage, transmission or display devices of the computer system.
[0198] Similarly, the term "processor" can refer to any device or part of a device that processes electronic data from registers and / or main memory and converts that electronic data into other electronic data, which may be stored in registers and / or main memory. A "computing platform" may have one or more processors. In the sense used herein, "software" procedures may include, for example, software and / or hardware units that perform work over time, such as tasks, threads, and intelligent agents. Furthermore, each procedure may refer to multiple procedures for executing instructions sequentially or in parallel, continuously or intermittently.In at least one embodiment, the terms “system” and “method” are used interchangeably herein, insofar as the system may embody one or more methods and methods may be regarded as a system.
[0199] This document may refer to the acquisition, capture, reception, or input of analog or digital data into a subsystem, computer system, or computer-implemented machine. In at least one embodiment, the process of acquiring, capturing, receiving, or inputting analog and digital data can be achieved in a variety of ways, such as receiving data as a parameter of a function call or an application programming interface call. In at least one embodiment, processes for acquiring, capturing, receiving, or inputting analog or digital data can be achieved by transmitting data over a serial or parallel interface.In at least one embodiment, processes for obtaining, capturing, receiving, or inputting analog or digital data can be performed by transmitting data over a computer network from a providing entity to a receiving entity. In at least one embodiment, reference can also be made to providing, outputting, transmitting, sending, or presenting analog or digital data. In various examples, processes for providing, outputting, transmitting, sending, or presenting analog or digital data can be achieved by transmitting data as input or output parameters of a function call, a parameter of an application programming interface, or an interprocess communication mechanism.
[0200] Although the descriptions presented here are exemplary embodiments of the described techniques, other architectures may be used to implement the described functionality and are intended to be within the scope of protection of this disclosure. Furthermore, although specific divisions of responsibilities may be defined above for descriptive purposes, various functions and responsibilities may be distributed and divided in different ways depending on the circumstances.
[0201] Furthermore, it is understood that, although the subject matter has been described in language specific to structural features and / or methodological actions, the subject matter defined in the attached patent claims is not necessarily limited to the features or actions described above. Rather, specific features and actions are disclosed as exemplary forms of implementing the claims.
[0202] The disclosure of this application also contains the following numbered clauses: Clause 1 Method, comprising: identifying a plurality of virtualized computing environments for running a plurality of applications; assigning a first set of virtualized computing environments from the plurality of virtualized computing environments to a first application from the plurality of applications based on one or more properties of the first application, wherein data associated with the first application are stored in respective caches of the first set of virtualized computing environments;Assigning a second set of virtualized computing environments from the plurality of virtualized computing environments to a second application from the plurality of applications based on one or more properties of the second application, wherein the second set of virtualized computing environments differs from the first set of virtualized computing environments, and wherein data associated with the second application is stored in respective caches of the second set of virtualized computing environments; and causing, in response to a request to execute the first application, an instance of the first application to be executed in a virtualized computing environment from the first set of virtualized computing environments using data stored in a cache of the virtualized computing environment. Clause 2 Procedure according to Clause 1, wherein one or more features of the first application include at least one of a size of the first application, a popularity of the first application or a classification of the first application, and wherein one or more features of the second application include at least one of a size of the second application, a popularity of the second application or a classification of the second application. Clause 3 Procedure according to one of the preceding clauses, wherein the plurality of virtualized computing environments is organized in a ring of virtualized computing environments and wherein the virtualized computing environment is an available virtualized computing environment selected from the ring of virtualized computing environments. Clause 4 Procedure according to Clause 3, wherein the available virtualized computing environment in the ring of virtualized computing environments is a first available virtualized computing environment selected from the first set of virtualized computing environments in the ring of virtualized computing environments. Clause 5 Procedure according to one of Clauses 3 or 4, wherein, if the first set of virtualized computing environments does not contain an available virtualized computing environment, selecting a virtualized computing environment for a second instance of the first application from at least a part of the second set of virtualized computing environments in the ring of virtualized computing environments, wherein at least the part of the second set of virtualized computing environments in the ring of virtualized computing environments is also assigned to the first application. Clause 6 Procedure according to any of the preceding clauses, further comprising: Receiving a request to terminate the first application and reassigning the virtualized computing environment as an available virtualized computing environment in the first set of virtualized computing environments. Clause 7 Procedure according to any of the preceding clauses, further comprising activating or deactivating a specific virtualized computing environment of the first set of virtualized computing environments based on one or more parameters. Clause 8 Procedure according to Clause 7, wherein one or more parameters include a time-of-day parameter or a day-of-the-week parameter. Clause 9 Procedure according to any of the preceding clauses, further comprising switching a particular virtualized computing environment from the first set of virtualized computing environments to the second set of virtualized computing environments based on at least one of the one or more properties of the first application or the one or more properties of the second application. Clause 10 Procedure according to any of the preceding clauses, further comprising determining the first set of virtualized computing environments and the second set of virtualized computing environments from the plurality of virtualized computing environments based on one or more properties of the first application or the one or more properties of the second application. Clause 11 Procedure according to Clause 10, wherein the first set of virtualized computing environments and the second set of virtualized computing environments are temporarily rotated on a ring of virtualized computing environments based on at least one of the expected wear characteristics of the first application or the expected wear characteristics of the second application. Clause 12 Procedure according to any of the preceding clauses, wherein the assignment of the first set of virtualized computing environments to the first application comprises the assignment of a first unique identifier of the first application to the first set of virtualized computing environments, and wherein the assignment of the second set of virtualized computing environments to the second application comprises the assignment of a second unique identifier of the second application to the second set of virtualized computing environments. Clause 13 Procedure according to one of the preceding clauses, wherein at least one specific virtualized computing environment of the plurality of virtualized computing environments stores application data for two or more applications of the plurality of applications in the cache of the specific virtualized computing environment. Clause 14 Procedure according to Clause 13, further comprising: Causing, in response to a request to execute a particular application of the two or more applications, an instance of the particular application to be executed in at least one virtualized computing environment which stores the application data for the two or more applications in the cache of the particular virtualized computing environment. Clause 15 System, comprising: a main memory and at least one processor coupled to the main memory for performing operations, comprising: designating a first set of virtualized computing environments from a plurality of virtualized computing environments to execute a first application from the plurality of applications based on one or more properties of the first application; designating a second set of virtualized computing environments from the plurality of virtualized computing environments to execute a second application from the plurality of applications based on one or more properties of the second application, wherein the second set of virtualized computing environments is different from the first set of virtualized computing environments;and cause, in response to a request to run the first application, an instance of the first application to be executed in a virtualized computing environment of the first set of virtualized computing environments, using data stored in a corresponding cache of the virtualized computing environment. Clause 16 System according to Clause 15, wherein one or more features of the first application include at least one of a size of the first application, a popularity of the first application or a rating of the first application, and wherein one or more features of the second application include at least one of a size of the second application, a popularity of the second application or a rating of the second application. Clause 17 System according to one of Clauses 15 or 16, wherein the plurality of virtualized computing environments is organized in a ring of virtualized computing environments and wherein the virtualized computing environment is an available virtualized computing environment selected from the ring of virtualized computing environments. Clause 18 System according to Clause 17, wherein the available virtualized computing environment in the ring of virtualized computing environments is a first available virtualized computing environment selected from the first set of virtualized computing environments in the ring of virtualized computing environments. Clause 19 Non-volatile, computer-readable medium on which instructions are stored, wherein the instructions, when executed by a processor, cause the processor to: assign one or more sets of virtualized computing environments from a plurality of virtualized computing environments to one or more applications from a plurality of applications based on one or more properties of the one or more applications; and cause, in response to a request to execute a particular application from the plurality of applications, an instance of the particular application to be executed in a particular virtualized computing environment of a set of virtualized computing environments assigned to the particular application, using data stored in a corresponding cache of the particular virtualized computing environment. Clause 20 Non-volatile computer-readable medium according to Clause 19, wherein the plurality of virtualized computing environments is organized in a ring of virtualized computing environments and wherein the specific virtualized computing environment is an available virtualized computing environment selected from the set of virtualized computing environments in the ring of virtualized computing environments assigned to the specific application.
[0203] It is understood that the aspects and embodiments described above are purely exemplary and that modifications of details may be made within the scope of protection of the claims.
[0204] Each device, each method and each feature disclosed in the description, and (where applicable) the claims and drawings, may be provided independently or in any suitable combination.
[0205] Reference numerals appearing in the claims are for illustrative purposes only and do not restrict the scope of protection of the claims.
Claims
[1] Procedure, encompassing: Identifying a variety of virtualized computing environments for running a variety of applications; Assigning a first set of virtualized computing environments from the plurality of virtualized computing environments to a first application from the plurality of applications based on one or more properties of the first application, wherein data associated with the first application are stored in respective caches of the first set of virtualized computing environments; Assigning a second set of virtualized computing environments from the plurality of virtualized computing environments to a second application from the plurality of applications based on one or more properties of the second application, wherein the second set of virtualized computing environments differs from the first set of virtualized computing environments, and wherein data associated with the second application is stored in respective caches of the second set of virtualized computing environments; and In response to a request to run the first application, cause an instance of the first application to run in a virtualized computing environment of the first set of virtualized computing environments, using data stored in a cache of the virtualized computing environment. [2] Method according to claim 1, wherein one or more features of the first application include at least one of a size of the first application, a popularity of the first application or a classification of the first application, and wherein one or more features of the second application include at least one of a size of the second application, a popularity of the second application or a classification of the second application. [3] Method according to one of the preceding claims, wherein the plurality of virtualized computing environments is organized in a ring of virtualized computing environments and wherein the virtualized computing environment is an available virtualized computing environment selected from the ring of virtualized computing environments. [4] Method according to claim 3, wherein the available virtualized computing environment in the ring of virtualized computing environments is a first available virtualized computing environment selected from the first set of virtualized computing environments in the ring of virtualized computing environments. [5] Method according to one of claims 3 or 4, wherein, if the first set of virtualized computing environments does not contain an available virtualized computing environment, selecting a virtualized computing environment for a second instance of the first application from at least a part of the second set of virtualized computing environments in the ring of virtualized computing environments, wherein at least the part of the second set of virtualized computing environments in the ring of virtualized computing environments is also assigned to the first application. [6] Method according to any one of the preceding claims, further comprising: Receiving a request to terminate the first application and Reassigning the virtualized computing environment as an available virtualized computing environment in the first set of virtualized computing environments. [7] Method according to any of the preceding claims, further comprising activating a specific virtualized computing environment of the first set of virtualized computing environments or deactivating a specific virtualized computing environment of the first set of virtualized computing environments based on one or more parameters. [8] Method according to claim 7, wherein one or more parameters comprise a time-of-day parameter or a day-of-the-week parameter. [9] Method according to any of the preceding claims, further comprising switching a specific virtualized computing environment from the first set of virtualized computing environments to the second set of virtualized computing environments based on at least one of the one or more properties of the first application or the one or more properties of the second application. [10] Method according to any of the preceding claims, further comprising determining the first set of virtualized computing environments and the second set of virtualized computing environments from the plurality of virtualized computing environments based on one or more properties of the first application or one or more properties of the second application. [11] Method according to claim 10, wherein the first set of virtualized computing environments and the second set of virtualized computing environments are temporarily rotated on a ring of virtualized computing environments based on at least one of the expected wear characteristics of the first application or the expected wear characteristics of the second application. [12] Method according to any of the preceding claims, wherein assigning the first set of virtualized computing environments to the first application comprises assigning a first unique identifier of the first application to the first set of virtualized computing environments, and wherein assigning the second set of virtualized computing environments to the second application comprises assigning a second unique identifier of the second application to the second set of virtualized computing environments. [13] Method according to one of the preceding claims, wherein at least one specific virtualized computing environment of the plurality of virtualized computing environments stores application data for two or more applications of the plurality of applications in the cache of the specific virtualized computing environment. [14] The method of claim 13, further comprising: To cause, in response to a request to run a specific application of one of two or more applications, an instance of the specific application to be executed in at least one virtualized computing environment that stores the application data for the two or more applications in the cache of the specific virtualized computing environment. [15] System, encompassing: a memory and comprising at least one processor coupled with the main memory for performing operations: Designing a first set of virtualized computing environments from a multitude of virtualized computing environments to run a first application from the multitude of applications based on one or more properties of the first application; Designing a second set of virtualized computing environments from the plurality of virtualized computing environments to run a second application from the plurality of applications based on one or more properties of the second application, wherein the second set of virtualized computing environments differs from the first set of virtualized computing environments; and In response to a request to run the first application, cause an instance of the first application to be executed in a virtualized computing environment of the first set of virtualized computing environments, using data stored in a corresponding cache of the virtualized computing environment. [16] System according to claim 15, wherein one or more features of the first application include at least one of a size of the first application, a popularity of the first application or a classification of the first application, and wherein one or more features of the second application include at least one of a size of the second application, a popularity of the second application or a classification of the second application. [17] System according to one of claims 15 or 16, wherein the plurality of virtualized computing environments is organized in a ring of virtualized computing environments and wherein the virtualized computing environment is an available virtualized computing environment selected from the ring of virtualized computing environments. [18] System according to claim 17, wherein the available virtualized computing environment in the ring of virtualized computing environments is a first available virtualized computing environment selected from the first set of virtualized computing environments in the ring of virtualized computing environments. [19] Non-volatile, computer-readable medium on which instructions are stored, wherein the instructions, when executed by a processor, cause the processor to do the following: Assigning one or more sets of virtualized computing environments from a variety of virtualized computing environments to one or more applications from a variety of applications based on one or more properties of the one or more applications; and To cause, in response to a request to run a specific application from the multitude of applications, an instance of the specific application to be executed in a specific virtualized computing environment of a set of virtualized computing environments assigned to the specific application, using data stored in a corresponding cache of the specific virtualized computing environment. [20] Non-volatile computer-readable medium according to claim 19, wherein the plurality of virtualized computing environments is organized in a ring of virtualized computing environments and wherein the specific virtualized computing environment is an available virtualized computing environment selected from the set of virtualized computing environments in the ring of virtualized computing environments assigned to the specific application.