Input Encoding Using Associative Learning
Federated learning with input encoding techniques facilitates secure collaboration among multiple entities by sharing model parameters and updates, addressing data privacy concerns and improving model training efficiency.
Patent Information
- Application Number
- JP2023560054
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-04-26
- Filing Date
- 2022-04-24
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2042-04-24
AI Technical Summary
Distributed machine learning environments face challenges in combining data due to privacy and confidentiality concerns, making it difficult to train neural networks across multiple entities without sharing sensitive data.
Implementing federated learning with input encoding techniques, where encoders are identified and distributed among local entities to train machine learning models, ensuring data privacy by sharing only model parameters and updates.
Enables secure and efficient collaboration among multiple client machines by allowing them to train their own models using shared encoders, enhancing model performance without exposing sensitive data.
Smart Images

Figure 0007805076000007 
Figure 0007805076000008 
Figure 0007805076000009
Abstract
Description
[Technical Field]
[0001] The present embodiments relate to training machine learning models through a collaborative process. More particularly, the embodiments relate to techniques for applying input encoding methods in federated learning. [Background technology]
[0002] Artificial intelligence (AI) is a field of computer science that focuses on computers and computer behavior as they relate to humans. AI refers to intelligence when a machine can make informed decisions that maximize the likelihood of success in a particular topic. More specifically, AI can learn from data sets to solve problems and provide appropriate recommendations. For example, in the field of artificial intelligence computer systems, natural language systems (such as the IBM Watson® artificial intelligence computer system or other natural language question-answering systems) process natural language based on knowledge acquired by the system. To process natural language, the system may be trained using data obtained from a database or corpus of knowledge, but the results obtained may be incorrect or inaccurate for various reasons.
[0003] Machine learning (ML), a subset of artificial intelligence (AI), uses algorithms to learn from data and make predictions based on this data. ML is the application of AI through the creation of models, including neural networks, that can exhibit learning behavior by performing tasks they were not explicitly programmed to perform. ML workloads require large data sets, fast parallel access to the data, and algorithms to train to support the learning. Training AI models relies on large, high-quality data sets. In an enterprise, this data can be distributed across various entities and data centers, making it difficult to combine and analyze. Collecting data into a single repository for training is often impossible or impractical.
[0004] Federated learning is a distributed machine learning process in which different parties (e.g., clients) collaborate to jointly train a machine learning model without the need to share the training data with other parties. Each client has access only to its own data; therefore, the data is not independently and uniformly distributed across clients. Summary of the Invention
[0005] Embodiments include systems, computer program products, and methods for secured federated machine learning.
[0006] In one aspect, a system for use with an artificial intelligence (AI) platform for training machine learning models through a collaborative process is provided. A processing unit is operatively coupled to the memory and in communication with the AI platform, which includes tools in the form of an enrollment manager, an evaluator, and a director. The enrollment manager is operative to arrange first and second participating entities in a collaborative relationship. The first participating entity trains a first machine learning model using a first encoder on a first training data set, and the second participating entity trains a second machine learning model using a second encoder on a second training data set. The evaluator, operatively coupled to the enrollment manager, is operative to measure performance of the first and second machine learning models and selectively identify at least one of the first and second machine learning models based on the measured performance. The director, operatively coupled to the evaluator, is operative to share the encoder of the selectively identified machine learning model with each of the participating entities. The shared encoder is configured to be applied by the first and second participating entities to train first and second machine learning models, respectively. The director is further configured to merge the first and second trained machine learning models to form a single shared model configured to be distributed to the participating entities.
[0007] In another aspect, a computer program product for training machine learning models through a collaborative process is provided. The computer program product includes a computer-readable storage medium having program code embodied thereon, the program code being executable by a processor to arrange first and second participating entities in a collaborative relationship. The first participating entity trains a first machine learning model using a first encoder on a first training data set, and the second participating entity trains a second machine learning model using a second encoder on a second training data set. The program code is operative to measure performance of the first and second machine learning models and selectively identify at least one of the first and second machine learning models based on the measured performance. The program code is further operative to share the encoder of the selectively identified machine learning model with each of the participating entities. The shared encoder is configured to be applied by the first and second participating entities to train the first and second machine learning models, respectively. The program code is further configured to merge the first and second trained machine learning models to form a single shared model configured to be distributed to participating entities.
[0008] In yet another aspect, a method for training machine learning models through a collaborative process is provided. First and second participating entities are arranged in a collaborative relationship. The first participating entity trains a first machine learning model using a first encoder on a first training data set, and the second participating entity trains a second machine learning model using a second encoder on a second training data set. Performance of the first and second machine learning models is measured, and at least one of the first and second machine learning models is selectively identified based on the measured performance. An encoder for the selectively identified machine learning model is shared with each of the participating entities. The shared encoder is configured to be applied by the first and second participating entities to train the first and second machine learning models, respectively. The first and second trained machine learning models are then merged to form a single shared model configured to be distributed to the participating entities.
[0009] These and other features and advantages will become apparent from the following detailed description of the presently preferred embodiments taken in conjunction with the accompanying drawings.
[0010] The drawings referenced in this specification form a part of this specification. Features shown in the drawings are intended to be illustrative of only some embodiments and not all embodiments, unless expressly stated otherwise. [Brief explanation of the drawings]
[0011] [Figure 1] FIG. 1 is a system diagram illustrating a computer system configured for federated learning according to an embodiment. [Figure 2] FIG. 1 is a block diagram illustrating participating entities in a star protocol arrangement for supporting collaboration within a federated learning environment. [Figure 3]FIG. 2 is a block diagram illustrating the artificial intelligence platform and tools shown and described in FIG. 1 and their associated application program interfaces. [Figure 4A] 1 is a flowchart illustrating an embodiment of a method that includes sharing a model and corresponding encoder with an operably coupled entity. [Figure 4B] 1 is a flowchart illustrating an embodiment of a method that includes sharing a model and corresponding encoder with an operably coupled entity. [Figure 5] 10 is a flowchart illustrating an embodiment of a method that includes selecting two or more encoders based on performance evaluation from test data. [Figure 6] 1 is a flowchart illustrating an embodiment of a method that includes creating a union of encoders for use in federated learning. [Figure 7] 1 is a flow chart illustrating an embodiment of a method that includes performing performance evaluation in a distributed manner. [Figure 8] FIG. 8 is a block diagram illustrating an example computer system / server of a cloud-based support system for implementing the systems and processes described above with respect to FIGS. 1-7. [Figure 9] FIG. 1 is a block diagram illustrating a cloud computing environment. [Figure 10] FIG. 1 is a block diagram illustrating a series of functional abstraction model layers provided by a cloud computing environment. DETAILED DESCRIPTION OF THE INVENTION
[0012] It will be readily understood that the components of the present embodiments, as generally described and illustrated in the Figures herein, could be arranged and designed in a wide variety of different configurations. Thus, the following detailed description of the present apparatus, system, method, and computer program product embodiments as illustrated in the Figures is not intended to limit the scope of the claimed embodiments, but is merely representative of selected embodiments.
[0013] Throughout this specification, references to "selected embodiments," "one embodiment," or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment. Thus, the appearances of the phrases "selected embodiments," "in one embodiment," or "in an embodiment" in various places throughout this specification are not necessarily referring to the same embodiment.
[0014] The illustrative embodiments will be best understood by reference to the drawings, in which like parts are designated by like numerals throughout. The following description is intended to be merely exemplary and merely sets forth certain selected embodiments of devices, systems, and processes consistent with embodiments claimed herein.
[0015] A neural network is a model of how the human brain processes information. The basic unit of a neural network is called a neuron, which is typically organized into layers. Neural networks function by simulating a large number of interconnected processing units that resemble abstract versions of neurons. A neural network typically has three parts: an input layer with units representing input fields, one or more hidden layers, and an output layer with one or more units representing target fields. These units are connected with various connection strengths, or weights. Input data is presented to the first layer, and values are propagated from each neuron to all neurons in the next layer. Finally, the output layer provides results. Complex deep learning neural networks are designed to emulate how the human brain functions, allowing computers to be trained to support poorly defined abstractions and problems. Neural networks and deep learning are often used in image recognition, speech, and computer vision applications.
[0016] Neural networks consist of interconnected layers with corresponding algorithms and adjustable weights. The weights are adjusted using optimization algorithms such as stochastic gradient descent. More specifically, gradient descent is an optimization algorithm used to minimize a function by iteratively moving in the direction of steepest descent, defined by its negative gradient. In ML, gradient descent is used to update the parameters of neural networks and corresponding neural models. This is straightforward when training on a single physical machine or across computers within a single entity. However, when multiple entities are involved, sharing data may be impossible due to communication limitations or legal reasons (such as regulations like HIPAA). It is understood in the art that sharing insights from data can lead to the construction of desirable or improved neural models. However, sharing data can lead to other issues, such as breaches of confidentiality and privacy, as other participating entities may reverse engineer (e.g., reconstruct) the data from shared insights.
[0017] As shown and described herein, systems, computer program products, and methods are provided for supporting and enabling collaboration among multiple client machines via a central server for training machine learning models. In the field of federated learning, it is understood that each client machine has access only to its own data, and the server receives only model parameters or updates to model parameters from the client machines, thereby ensuring the privacy of the client machine's data. This embodiment is directed to federated learning, in which one or more encoders are identified and distributed. Generally, an encoder is a device or process that transforms data from one format to another. In the context of machine learning, encoded input features are encoded as input features.
number
number
number
number
number
number
[0018] Embodiments are directed to the application of input encoding techniques in conjunction with federated learning, and more particularly, to the evaluation and selection of one or more encoders using a trust metric. The selected encoder is shared or otherwise distributed with operatively coupled local entities (e.g., client machines) to enable the local entities to train their own local entity models (also referred to herein as base models), which utilize the selected encoder and local data. Thus, the federated learning process herein limits distribution among members of a federated learning environment to encoders and corresponding models, and the federated learning process allows local entities to benefit from the performance evaluations of members of the federated learning system, the selection of corresponding encoders, and the model updates.
[0019] Referring to FIG. 1, a schematic diagram (100) illustrating a computer system configured for federated learning is provided. As shown in the figure, a server (110), also referred to herein as a federated learning server, is provided in communication with multiple computing devices (180), (182), (184), (186), (188), and (190) via a network connection (105). The server (110) is comprised of a processing unit (112) in communication with a memory (116) via a bus (114). The server (110) is shown with an artificial intelligence (AI) platform (150) for supporting federated learning to selectively identify encoders within a distributed, federated, private, and secure environment. The server (110) communicates with one or more of the computing devices (180), (182), (184), (186), (188), and (190) via the network connection (105). More specifically, computing devices 180, 182, 184, 186, 188, and 190 communicate with each other and with other devices or components via one or more wired and / or wireless data communication links, each of which may include one or more wires, routers, switches, transmitters, receivers, etc. In this network configuration, server 110 and network connections 105 enable communication detection, recognition, and resolution. Other embodiments of server 110 may be used with components, systems, subsystems, and / or devices other than those shown herein.
[0020] The AI platform (150) is shown herein configured to receive input (102) from a variety of sources. For example, the AI platform (150) may receive input from a network (105) and utilize a data source (160), also referred to herein as a corpus or knowledge base, to evaluate the performance of a machine learning model. As shown, the data source (160) may comprise a library (162) or, in one embodiment, multiple libraries, where the library (162) comprises a dataset A (164 A ), dataset B (164 B ), dataset C (164 C ), and datasets D (164 D ), in one embodiment, the library 162 may include a reduced number of datasets or an increased number of datasets. Similarly, in one embodiment, the libraries in the data source 160 may be organized by common subject or theme, although this is not a requirement.
[0021] The AI platform (150) includes tools for supporting and enabling machine learning collaboration. Various computing devices (180), (182), (184), (186), (188), and (190) in communication with the network (105) may function as local entities containing local machine learning models. The AI platform (150) functions as a platform for enabling and supporting collaboration corresponding to the local machine learning models without sharing insights or data. As shown and described herein, the tools in support of federated learning are local to the AI platform (150). More specifically, as described in detail herein, the AI platform (150) and corresponding tools enable sharing of local encoders and local machine learning models without sharing data. Response outputs (132) in the form of encoders and corresponding neural models having desired performance or accuracy are obtained and shared with entities encompassing and participating in the collaboration. In one embodiment, the AI platform (150) communicates the response output (132) via the network (105) to a server (110) or to members of a collaborative topology such as those shown and described in FIG. 2 that are operatively coupled to one or more of the computing devices (180), (182), (184), (186), (188), and (190).
[0022] In various embodiments, the network 105 may include local network connections and remote connections, allowing the AI platform 150 to operate in environments of any size, including local and global (e.g., the Internet). The AI platform 150 serves as a back-end system to support collaboration. In this manner, any process adds data to the AI platform 150, which also includes an input interface for receiving requests and responding accordingly.
[0023] The AI platform (150) is shown herein with several tools for supporting collaboration of neural models, including a registration manager (152), an evaluator (154), and a director (156). The registration manager (152) functions to register participating entities in a collaboration, including arranging the registered entities within a topology and establishing communication direction and communication protocols between the entities within the topology. For example, in an embodiment, the registered entities are arranged in a ring topology. However, the communication protocol may vary. Examples of protocols include, but are not limited to, a linear direction protocol, a broadcast protocol, and an all-reduce protocol. As further shown and described herein, the entities that are registered members of the collaboration are referred to herein as local entities and are each operably coupled to the server (110) via a network connection (105). Each of the entities includes a local machine learning model and a corresponding local encoder, and each local machine learning model is trained using a local training dataset. For example, the computing entity (190) may run a local machine learning model (192 A ), the corresponding encoder (192 B ), and the local model training dataset (192 C ) In response, the enrollment manager (152) establishes a configuration of the participating entities within the federated learning environment to support collaboration between the entities.
[0024] The evaluator (154), shown operatively coupled to the enrollment manager (152) herein, functions to address the performance of the participating entities' local machine learning models. The evaluator (154) measures the performance of each of the enrolled local models against one or more test datasets and selectively identifies at least one of the enrolled local models based on the performance measurements. Each enrolled local model includes a corresponding encoder E(·). The director (156) shares the encoder associated with the selectively identified local model with the participating entities. In an embodiment, the director (156) shares the local model corresponding to the selected or identified encoder with the participating entities. The participating entities receive the identified encoder and / or model and apply their local data to train their local models based on the shared encoder. After a local model is trained for an identified or selected encoder, the corresponding local entity shares the trained local model with the AI platform (150), which is then subject to merging by the director (156) to create a single model that is communicated to members of the collaborative topology as a single shared model as the response output (132). Thus, sharing among participating entities is limited to the encoder and trained model parameters or model parameter updates, and in embodiments, the associated machine learning model, thereby supporting the benefits of learning while precluding data sharing among participating entities.
[0025] The embodiments shown and described herein may be extended to accommodate multiple machine learning models and corresponding encoders and data sets local to a participating entity. Referring to FIG. 2, a block diagram (200) is provided illustrating participating entities in a star protocol arrangement (also referred to herein as a hub and spoke protocol arrangement) for supporting collaboration. As shown in the figure, a central entity (210) is operatively coupled to local entities (220), (230), (240), and (260). The central entity (210) is comprised of the tools of the AI platform shown in FIG. 1, including a registration manager (252), an evaluator (254), and a director (256). The local entity (220) is herein referred to as a model 0,0 (222 0,0 ) and model 0,1 (222 0,1 ) along with the two machine learning models. 0,0 (222 0,0 ) is the encoder 0,0 (224 0,0 ) and datasets 0,0 (226 0,0 ) and the model 0,1 (222 0,1 ) is the encoder 0,1 (224 0,1 ) and datasets 0,1 (226 0,1 ) The remaining local entities (230), (240), and (260) are shown with one local model and associated encoder and dataset. The number of models shown for the local entities is for illustrative purposes and should not be considered limiting. As an example, local entity 1 (230) is shown with model 1,0 (232 1,0 ), encoder 1,0 (234 1,0 ), and datasets 1,0 (236 1,0), and local entity 2 (240) is shown with the model 2,0 (242 2,0 ), encoder 2,0 (244 2,0 ), and datasets 2,0 (246 2,0 ), and local entity 3 (260) is shown with the model 3,0 (262 3,0 ), encoder 3,0 (264 3,0 ), and datasets 3,0 (266 3,0 ) are shown. While only one of the local entities is shown with multiple models, in the illustrated embodiment, one or more additional local entities may comprise multiple models. Similarly, in an embodiment, the number of multiple local models may be greater than the number of local models shown. In the embodiment shown and described in FIG. 2, the evaluator (254) functions to measure the performance of the local models. Two machine learning models (e.g., model 0,0 (222 0,0 ) and model 0,1 (222 0,1 In the case of a local entity that includes multiple local models, such as local entity 0 (220) shown with ( ), the evaluator (254) measures the performance of each of the learned models for selective discrimination. The process of model performance evaluation is shown and explained in Figures 4A-7. Thus, a participating entity may be composed of or equipped with multiple local models to participate in federated learning.
[0026] Each entity separately shares its local performance evaluation data, or its encoders and / or models, with the central entity (210) via local communication channels (228), (238), (248), and (268), respectively, as shown and described in Figures 4A and 4B. The evaluator (254) may selectively identify a single model and corresponding encoder to be shared in federated learning, or in an exemplary embodiment, the evaluator (254) may selectively identify two or more local models from multiple received local models. In the case of identification of multiple models, the evaluator creates a union of the encoders for the director (256) to share with participating entities. Thus, the server (210), or in the embodiment shown in Figure 2, the central entity (210), evaluates the local entity models against a common test dataset.
[0027] The embodiments shown and described in Figures 1 and 2 are directed to a server (110) or central entity (210) that performs performance evaluation on a common data set. In an exemplary embodiment, performance evaluation may be performed local to participating entities (also referred to herein as distributed performance evaluation). With reference to Figure 2, the central entity (210) may support distributed performance evaluation via an evaluator (254), which is configured to share with each participating entity an encoder of a local entity model identified by the performance evaluation. After performance evaluation and identification or selection of one or more encoders, the central entity (210) distributes the identified or selected encoders, and in embodiments, the corresponding models, to each of its operably coupled entities via respective communication channels. In an embodiment, the central entity (210) receives the encoders and corresponding local entity models from each participating entity, and the evaluator shares the received encoders and corresponding local entity models with each participating entity. Local entity 0 (220) is in the dataset 0,0 (226 0,0 ) using a local dataset such as the received local entity model (e.g., model 1,0 (232 1,0 ), model 2,0 (242 2,0 ), and models 3,0 (262 3,0 A local distributed performance evaluation is performed for each local entity to evaluate each of the received local entity models. For example, local entity 1 (230) is the first entity in the dataset 1,0 (236 1,0 ) and evaluate the performance of the received local model using a local dataset such as 2,0 (246 1,0) and evaluate the performance of the received local model using a local dataset such as 3,0 (266 3,0 ) and evaluate the performance of the received local model using a local dataset. 0,0 (226 0,0 ), dataset 1,0 (236 1,0 ), dataset 2,0 (246 2,0 ), and datasets 3,0 (266 3,0 The local datasets are different, such that the local entities (110) and the received local entity models (210) are different datasets. The performance evaluation data from each local entity and each received local entity model is shared with the evaluator (154), (254) for selective identification of one or more of the models. Thus, as shown herein, the performance evaluation may be hierarchical, with one level of evaluation being performed at a local level involving different datasets, and a second level of evaluation being performed at a central level by the server (110) or central entity (210).
[0028] The registration manager (152) is responsible for establishing the topology and communication protocol. In one embodiment, the registration manager (152) establishes a fully connected topology, also known as a mesh topology, and a corresponding broadcast protocol, in which each participating entity directly transmits (e.g., broadcasts) its encoders to all other participating entities in the topology via the topology. The evaluator (154) further supports and enables selective model and encoder identification, and may employ central performance evaluation by the evaluator (154) or (254) or distributed performance evaluation by each participating local entity. The goal of performance evaluation is for each participating entity to receive and benefit from one or more selectively identified encoders of other participating entities. In a star topology, each participating member entity can communicate directly with the AI platform (150), so the evaluator (154) is configured to evaluate the performance of the model. In embodiments, one or more of the local entities may be registered but not actively participate with the AI platform (150), in which case the evaluator (154) may limit the sharing of selectively identified encoders with contributing entities or, depending on the performance evaluation broadcast protocol, require identified non-contributing entities to broadcast their local models to the server (110) or to broadcast their local models to each of the participating members of the topology. Thus, as shown and described herein, the star topology employs a broadcast protocol, and in one embodiment, validation of entity participation, to support federated machine learning.
[0029] In some example embodiments, the server (110) may be an IBM Watson® system from International Business Machines Corporation (Armonk, New York), which has been extended using mechanisms of the example embodiments described below. The IBM Watson® system shown and described herein includes tools for implementing federated machine learning based on performance evaluation and associated sharing protocols. These tools enable selective identification and sharing of encoders without sharing the underlying data, thereby allowing the data to remain confidential or private.
[0030] The Enrollment Manager (152), Evaluator (154), and Director (156), hereinafter collectively referred to as AI Tools or AI Platform Tools, are shown embodied or integrated within the AI Platform (150) of the Server (110). The AI Tools may be implemented within a separate computing system (e.g., 190) connected to the Server (110) via the Network (105). Whenever embodied, the AI Tools function to support and enable federated machine learning in an iterative manner, including the sharing of encoders and local models between participating entities, without sharing or disclosing the underlying data.
[0031] The types of information processing systems that can utilize the AI platform (150) range from small handheld devices, such as handheld computers / cell phones (180), to large mainframe systems, such as mainframe computers (182). Examples of handheld computers (180) include personal digital assistants (PDAs) and personal entertainment devices, such as MP4 players, portable televisions, and compact disc players. Other examples of information processing systems include pen or tablet computers (184), laptop or notebook computers (186), personal computer systems (188), and servers (190). As shown, various information processing systems can be networked together using a computer network (105). Types of computer networks (105) that may be used to interconnect various information handling systems include local area networks (LANs), wireless local area networks (WLANs), the Internet, the public switched telephone network (PSTN), other wireless networks, and any other network topology that may be used to interconnect information handling systems. Many information handling systems include a non-volatile data store, such as a hard drive or non-volatile memory, or both. Some information handling systems may use separate non-volatile data stores (e.g., a server (190) may use a non-volatile data store (190)). A ), and the mainframe computer (182) uses a non-volatile data store (182 A )). Non-volatile data store (182 A) can be a component external to the various information handling systems or can be internal to one of the information handling systems.
[0032] The information processing system employed to support the AI platform (150) may take a variety of forms, some of which are illustrated in FIG. 1. For example, the information processing system may take the form of a desktop, server, portable, laptop, notebook, or other form factor computer or data processing system. In addition, the information processing system may adopt other form factors, such as a personal digital assistant (PDA), gaming device, ATM machine, mobile phone device, communications device, or other device including a processor and memory. In addition, the information processing system need not necessarily embody a north bridge / south bridge controller architecture, as it is understood that other architectures may be employed.
[0033] An application program interface (API) is understood in the art as intermediate software between two or more applications. With respect to the AI platform (150) shown and described in FIG. 1, one or more APIs may be utilized to support one or more of the tools (152)-(156) and their associated functionality. Referring to FIG. 3, a block diagram (300) illustrating the tools (152)-(156) and their associated APIs is provided. As shown, multiple tools are incorporated within the AI platform (305), including a registration manager (352) associated with API0 (312), an evaluator (354) associated with API1 (322), and a director (356) associated with API2 (332). Each of the APIs may be implemented in one or more languages and interface specifications. API 0 (312) provides functional support for registering participating entities in a collaboration, establishing communication protocols between entities in the topology, and, in embodiments, establishing the direction of communication between the entities; API 1 (322) provides functional support for addressing the performance of participating entities' local machine learning models; and API 2 (332) provides functional support for selectively sharing encoders associated with identified local models with participating entities. As shown in the figure, each of APIs (312), (322), and (332) is operatively coupled to an API orchestrator (360) (otherwise referred to as an orchestration layer), which is understood in the art to function as an abstraction layer for transparently connecting separate APIs. In one embodiment, the functionality of separate APIs may be combined or combined. As such, the configuration of APIs shown herein should not be considered limiting. Accordingly, the functionality of the tools may be embodied or supported by each of those APIs, as shown herein.
[0034] 4A and 4B, a flowchart (400) is provided illustrating the communication and sharing of a model and corresponding encoder with operatively coupled entities. As shown, multiple entities (e.g., client machines) may share a variable X Total In an exemplary embodiment, there are a minimum of two entities communicating with an operably coupled entity (e.g., a server). Each client machine includes at least one machine learning model trained using a corresponding encoder and a corresponding data set. In an exemplary embodiment, one or more of the client machines may include two or more machine learning models and corresponding encoders. Each of the models is trained on data provided by the entity, including encoded and unencoded data. If an entity trains two or more models, the models differ in their choice of encoding. The encodings meet certain requirements in that they must be uniquely decodable. Each model associated with each of the client machines and trained using a corresponding encoder and data set is identified (404). An entity count variable is initialized (406), and the entity X An evaluation is performed to identify whether the entity contains more than one model (408). After a negative response to this evaluation, X This is followed by passing the model of the entity to the server (410). X This is followed by an internal performance evaluation of the model associated with the entity 412. In an exemplary embodiment, this internal evaluation involves testing the entity against a common test data set (e.g., an internal or private data set). X, identifying the best performing model among these models (416), and passing the identified best performing model, and optionally an associated encoding, to a server (418).
[0035] After either step (410) or (418), an entity count variable is incremented (420), and an evaluation is performed to determine whether each of the identified entities participating in the federated learning environment has passed a model to the server (422). A negative response to this evaluation is followed by a return to step (408), while a positive response to this evaluation ends the first aspect of model performance evaluation, which is directed to selective internal performance evaluation. In embodiments, entities that do not pass a model to the server or otherwise share a model are considered entities that do not participate in the federated learning process.
[0036] As shown and described in FIG. 1, the server performs a performance evaluation on each of the models received from each entity, and this evaluation is performed using a common data set. This allows for uniform performance evaluation because each model is evaluated against the same data set. It is understood that participating entities may opt out of federated learning. As such, the number of entities operatively coupled to the server may not match the number of models received by the server. The variable Y Total is assigned to the number of models received by the server (424), and the variable Y Total It is determined whether ≠ 1 exceeds the integer ≠ 1 (426). A negative response is an indication that only one entity is participating in federated learning, and the process is terminated. After a positive response to the determination in step ≠ 1, performance data, also referred to herein as test data, is provided or otherwise identified, and the server uses the performance data to map each of the models (e.g., model Y) (428). Various protocols for federated learning exist. In the embodiment shown and described herein, the server identifies the model with the best-performing output from the test data and distributes the corresponding encoder and, optionally, the corresponding model to each of the entities participating in the federated learning (430). Each of the participating entities communicating with the server then trains their internal entity models against the received encoder and shares their trained models with the federated learning server (432). The federated learning server merges the trained models to create a single shared model, which is then communicated to each of the participating entities in the collaborative topology (434). Thus, in the embodiment shown and described, the federated learning process evaluates each participating entity's model against selected common test data.
[0037] Referring to FIG. 5, a flowchart (500) is provided illustrating a process for selecting two or more encoders based on performance evaluation from test data. The process illustrated herein continues from step (422) of FIG. 4A. Application of the performance data results in performance output data. As shown and described herein, a threshold setting is employed to evaluate the output from the performance evaluation. In an exemplary embodiment, the threshold is configurable. The variable Y Total is assigned to the number of models received by the server (502). Each of the models is evaluated by the server using performance data, also referred to herein as test data (504). In an exemplary embodiment, the performance data for each of the models is the same so that the comparison is uniform or relatively uniform. A value corresponding to the model's performance, also referred to herein as output, is used to evaluate the performance of the models. Y Output data, called Z, is obtained from each of the models along with the evaluated test data (506). The output data is evaluated against a threshold (508) to identify models whose performance meets or exceeds the threshold, and the variable Z Totalis assigned a quantity of models whose performance meets or exceeds the threshold 510. Thus, the initial encoder selection process focuses on evaluating the performance of the models against a defined test data set.
[0038] After the first phase of evaluation is completed, whether at least one model has been identified (e.g., variable Z Total A determination is made (512) as to whether at least two models have been identified (e.g., whether variable Z is greater than 0). A negative response to the determination in step (512) is followed by either resetting the threshold or selecting a different test data set (514), or in embodiments, both resetting the threshold and selecting a different test data set, followed by returning to step (504). Alternatively, a positive response to the determination in step (512) is followed by a determination of whether at least two models have been identified (e.g., whether variable Z is greater than 0). Total An evaluation follows to determine whether Z is greater than 1 (516). A negative response to this determination indicates that only one model has been identified as meeting at least the performance threshold, and the server distributes the corresponding encoder, and optionally the corresponding model, to each of the entities participating in the federated learning (518). After a positive response to the determination in step (516), the set of models (e.g., Z Total ) is followed by a subsequent evaluation to see if the best performing model exceeds a selected performance threshold for the corresponding configuration (520). In an exemplary embodiment, the configuration for federated learning is configurable, and therefore need not be a static configuration. After a positive response to the determination in step (520), the set of models Z Total This is followed by identifying (522) a single model from the set Z, and returning to step (518). As shown herein, after a negative response to the determination in step (520), TotalThis is followed by distributing the corresponding encoder for each of the models represented by {circumflex over (x)}, and optionally the corresponding model, to each of the entities participating in the federated learning (524). Accordingly, a threshold may be utilized to identify the encoders and corresponding models for use in the federated learning.
[0039] Referring to Figure 6, a flowchart (600) is provided illustrating a process for creating a union of encoders for use in federated learning. The union of encoders is directed to identifying two or more models that have performance data that meets or exceeds a threshold. As shown and described in Figure 5, the variable Z Total represents a model with evaluated performance data that meets or exceeds a threshold (602). In the embodiment shown and described in this figure, the variable Z Total represents at least two models. Each model includes a corresponding encoder. The server can Total ) and create a concatenation of the encodings received from the variable X (604). Total is assigned (606) to the quantity of entities operably coupled to the server and participating in federated learning, which may or may not equal the number of models in the example embodiment. Total In an embodiment, the encoder's transmission or communication within the federated learning environment includes the corresponding model. Each of the entities participating in the federated learning (e.g., X=1 to X) then communicates with the server. Total ) train their internal entity models on the received combined data of the encoders (610). After local training using each internal data, the internal entity models are communicated or otherwise shared with the federated learning server (612), which merges the internal entity models into a single model (614). After merging the models, the federated learning server distributes the merged model to each of the participating entities (e.g., X = 1 to XTotal ) (616). Thus, in the embodiment shown and described, the federated learning process evaluates the model of each participating entity on selected common test data.
[0040] As shown in Figures 4A, 4B, 5, and 6, the model is evaluated on a test data set at selected locations. Referring to Figure 7, a flowchart (700) is provided illustrating a process for performing evaluation in a distributed manner. The variable X Total is operatively coupled to a server or central entity and assigned to the number of entities participating in the federated learning environment (702). It is understood that each participating entity may include a single model, or in embodiments, multiple models. Total Entity X ) for each variable Y Total represents the number of local models for each entity (704). Each entity participating in federated learning performs a local evaluation of the performance of the model, which is referred to herein as a test set. X We use a local test data set, denoted as Total Entity X ) receives models from each of the other participating entities (706) and applies a local test data set to each of the received models, thereby generating output in the form of model performance data (708). The aspect of receiving models from other entities in the federated learning environment is limited to the model itself, including the encoder; no data is moved. Performance data for each model evaluated by each entity is communicated to a server or central entity (710). Performance evaluations are performed by the server or central entity using the model performance data results. Thus, multiple performance evaluations are performed local to each entity using test data sets local to each entity.
[0041] Similar to the process shown and described in FIGS. 4A-6, a threshold setting is employed by the central entity upon receipt of each piece of model performance data to evaluate the output from the performance (712). Additionally, a variable k is assigned or represents the number of models employed in communicating the corresponding encoder. Next, it is determined whether k is set to the integer 1 (714). A positive response to this determination is followed by the central entity or server identifying or selecting the model with the best performance data (716). The central entity or server then communicates or otherwise distributes the encoder associated with the identified model; in an exemplary embodiment, the associated model is attached to the encoder, and the central entity or server communicates the encoder and / or model to each of the entities in the federated learning environment (718). The local data, along with the received encoder, is then used to train a local model (726). After a negative response to the determination in step (714), the central entity selects the top k models and associated encoders and creates a union of the top k models and associated encoders (720). The union of the models and encoders is distributed from the central entity to each of the local entities participating in or being members of the federated learning (722). This distribution does not include data (e.g., no data is moved). Each of the entities participating in the federated learning, communicating with the server, then trains each of their internal entity models with the received union of encoders and local data (724). As shown in the figure, after the local models are trained using the selected encoders or union of encoders as shown in steps (724) and (726), the local models are returned to the federated learning server (e.g., server), which aggregates the models to form a single model, which is then returned to the local entities (728).In an exemplary embodiment, the local entity may resume training the single model formed, thereby forming an updated model, and pass or otherwise communicate the updated model to the federated learning server, which may continue the process of aggregating the models. In an embodiment, the federated learning process may continue until a performance metric is achieved. Thus, in the embodiment shown and described, the federated learning process employs distributed performance evaluations across different test data sets to form a single aggregated model that is managed by the federated learning server and communicated to participating entities.
[0042] The AI platform with a federated learning environment as shown and described in Figures 1-7 is directed to a structure or configuration of local entities and servers as shown in Figure 2. However, the configuration shown herein should not be considered limiting. In example embodiments, the configuration may take different forms, including different distribution paths, such as a hierarchical structure including a hierarchy of local entities around a central entity.
[0043] Aspects of the functional tools 152-156 and their associated functionality may be embodied in a computer system / server at a single location, or in one embodiment, configured within a cloud-based system sharing computational resources. Referring to Figure 8, a block diagram 800 is provided illustrating an example of a computer system / server 802 (hereinafter referred to as a host 802) in communication with a cloud-based support system to implement the processes described above with respect to Figures 1-7. The host 802 is operational with numerous other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, or configurations, or combinations thereof, that may be suitable for use with the host (802) include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and file systems that include any of these systems, devices, and the like (e.g., distributed storage environments and distributed cloud computing environments).
[0044] The host (802) may be described in the general context of instructions executable by a computer system, such as program modules being executed by the computer system. Typically, program modules may include routines, programs, objects, components, logic, data structures, etc. that perform particular tasks or implement particular abstract data types. The host (802) may be executed in a distributed cloud computing environment (810) where tasks are performed by remote processing devices linked through a communications network. In a distributed cloud computing environment, program modules may be located in both local and remote computer system storage media, including memory storage devices.
[0045] As shown in FIG. 8, the host (802) is depicted in the form of a general-purpose computing device. The components of the host (802) may include, but are not limited to, one or more processors or processing units (804) (e.g., hardware processors), system memory (806), and a bus (808) that couples various system components, including the system memory (806), to the processor (804). The bus (808) represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures. Examples of such architectures include, but are not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnects (PCI) bus. The host (802) typically includes a variety of computer system-readable media. Such media can be any available media that is accessible by the host (802) and includes both volatile and non-volatile media, removable and non-removable media.
[0046] The memory (806) may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) (830) and / or cache memory (832). By way of example only, a storage system (834) may be provided for reading from and writing to non-removable, non-volatile magnetic media (not shown, typically referred to as a "hard drive"). Although not shown, a magnetic disk drive may be provided for reading from and writing to removable, non-volatile magnetic disks (e.g., "floppy disks"), and an optical disk drive may be provided for reading from and writing to removable, non-volatile optical disks, such as CD-ROMs, DVD-ROMs, or other optical media. In such examples, each may be connected to the bus (808) by one or more data media interfaces.
[0047] For example, a program / utility (840) including a set of (at least one) program modules (842) may be stored in memory (806), including, but not limited to, an operating system, one or more application programs, other program modules, and program data. Each of the operating system, one or more application programs, other program modules, and program data, or a combination thereof, may include an implementation of a network environment. The program modules (842) typically perform the functions and / or methods of an embodiment for applying an input encoding method in associative learning. For example, the set of program modules (842) may include tools (152) through (156) described in FIG. 1.
[0048] The host (802) may communicate with one or more external devices (814), such as a keyboard, a pointing device, a display (824), one or more devices that allow a user to interact with the host (802), or any device (e.g., a network card, a modem, etc.) that allows the host (802) to communicate with one or more other computing devices, or a combination thereof. Such communication may occur through an input / output (I / O) interface (822). Additionally, the host (802) may communicate with one or more networks, such as a local area network (LAN), a general wide area network (WAN), or a public network (e.g., the Internet), or a combination thereof, through a network adapter (820). As shown, the network adapter (820) communicates with the other components of the host (802) via a bus (808). In one embodiment, multiple nodes of a distributed file system (not shown) communicate with a host (802) through an I / O interface (822) or through a network adapter (820). Although not shown, it should be understood that other hardware and / or software components may be used with the host (802), including, but not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archive storage systems.
[0049] In this document, the terms "computer program medium," "computer usable medium," and "computer readable medium" are used generally to refer to media such as main memory (806), including RAM (830), cache (832), and storage systems (834), such as removable storage drives and hard disks installed in hard disk drives.
[0050] Computer programs (also called computer control logic) are stored in the memory (806). The computer programs may be received via a communications interface, such as a network adapter (820). When executed, such computer programs enable the computer system to perform the features of the present embodiments described herein. In particular, when executed, the computer programs enable the processing unit (804) to perform the functions of the computer system. Thus, such computer programs represent the controller of the computer system.
[0051] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. A non-exhaustive list of more specific examples of computer-readable storage media includes portable floppy disks, hard disks, dynamic or static random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), magnetic storage devices, portable compact disc read-only memory (CD-ROM), digital versatile disks (DVDs), memory sticks, floppy disks, mechanically encoded devices such as punch cards or ridge structures in grooves on which instructions are recorded, and any suitable combination thereof. As used herein, computer-readable storage media should not be construed as being ephemeral signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses passing through fiber optic cable), or electrical signals transmitted over wires.
[0052] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or storage device over a network (e.g., the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof) that may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage on a computer-readable storage medium within each computing / processing device.
[0053] The computer-readable program instructions for carrying out the operations of the present embodiments may be source or object code written in any combination of one or more programming languages, including assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Java, Smalltalk, C++, and traditional procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer as a standalone software package, partially on the user's computer and on a remote computer, or entirely on a remote computer or server or cluster of servers. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, to perform aspects of the embodiments, electronic circuitry including, for example, programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), may execute computer-readable program instructions to personalize the electronic circuitry by utilizing state information of the computer-readable program instructions.
[0054] In one embodiment, the host (802) is a node in a cloud computing environment. As known in the art, cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) and for rapidly provisioning and releasing these resources with minimal administrative effort or interaction with a service provider. This cloud model may include at least five characteristics, at least three service models, and at least four deployment models. Examples of such characteristics are:
[0055] On-demand self-service: Cloud customers can automatically provision server time, network storage, and other computing power as needed, without any unilateral or human interaction with the service provider.
[0056] Wide network access: Cloud capabilities are available over the network and can be accessed using standard mechanisms, facilitating usage by heterogeneous thin- or thick-client platforms (e.g., mobile phones, laptops, and PDAs).
[0057] Resource Pool: The provider's computing resources are pooled and offered to multiple consumers using a multi-tenant model. Various physical and virtual resources are dynamically allocated and reallocated according to demand. There is a sense of location independence; consumers typically have no control or knowledge regarding the exact location of the resources offered, although higher layers of abstraction may allow for the location to be specified (e.g., country, state, or data center).
[0058] Rapid Elasticity: Cloud capacity can be quickly and elastically provisioned, in some cases automatically, to scale out quickly, and quickly released to scale in quickly. Capacity available for provisioning often appears to consumers as unlimited, available for purchase in any quantity at any time.
[0059] Metered Services: Cloud systems leverage metering capabilities to automatically control and optimize resource usage at whatever layer of abstraction is appropriate for the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both providers and consumers of the services used.
[0060] The service model is as follows:
[0061] SaaS (Software as a Service): The consumer is provided with the ability to use the provider's applications running on a cloud infrastructure. Those applications can be accessed from a variety of client devices through thin-client interfaces such as web browsers (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or individual application features, except for the possibility of setting limited user-specific application configuration settings.
[0062] PaaS (Platform as a Service): The ability offered to a consumer is to deploy applications they create or acquire, written using programming languages and tools supported by the provider, onto a cloud infrastructure. The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, or storage, but does have control over the deployed applications and, in some cases, the configuration of the application hosting environment.
[0063] Infrastructure as a Service (IaaS): The capability provided to a customer is the provisioning of processing, storage, network, and other basic computing resources, upon which the customer can deploy and run any software, which may include operating systems and applications. The customer does not manage or control the underlying cloud infrastructure, but rather controls the operating systems, storage, deployed applications, and in some cases, limited control over selected network components (e.g., host firewalls).
[0064] The deployment model is as follows:
[0065] Private Cloud: This cloud infrastructure is operated solely for the organization, can be managed by the organization or a third party, and can reside on-premise or off-premise.
[0066] Community Cloud: This cloud infrastructure is shared by multiple organizations to support a specific community with shared interests (e.g., mission, security requirements, policy, and compliance considerations). It can be managed by these organizations or a third party and can reside on-premises or off-premises.
[0067] Public cloud: This cloud infrastructure is available for use by the general public or large industry organizations and is owned by an organization that sells cloud services.
[0068] Hybrid cloud: This cloud infrastructure is a combination of two or more clouds (private, community, or public) that remain distinct but are joined together by standardized or proprietary technologies that allow for data and application portability (e.g., cloud bursting to balance load between clouds).
[0069] A cloud computing environment is a service-oriented environment that emphasizes statelessness, low coupling, modularity, and semantic interoperability. At the heart of cloud computing is an infrastructure that consists of a network of interconnected nodes.
[0070] Referring now to FIG. 9, an exemplary cloud computing network (900) is shown. As shown, the cloud computing network (900) includes a cloud computing environment (950) that includes one or more cloud computing nodes (910) with which local computing devices used by cloud subscribers may communicate. Examples of these local computing devices include, but are not limited to, a personal digital assistant (PDA) or mobile phone (954A), a desktop computer (954B), a laptop computer (954C), or an automotive computer system (954N), or combinations thereof. Individual nodes within the nodes (910) may further communicate with each other. The nodes (910) may be physically or virtually grouped together in one or more networks (not shown), such as a private cloud, community cloud, public cloud, or hybrid cloud, as previously described herein, or combinations thereof. This allows the cloud computing environment (900) to provide an infrastructure, platform, and / or SaaS that eliminates the need for cloud users to maintain resources on local computing devices. The types of computing devices (954A-954N) shown in Figure 9 are intended as examples only, and it is understood that the cloud computing environment (950) can communicate with any type of computer-controlled device via any type of network and / or network-addressable connection (e.g., a connection using a web browser).
[0071] Referring now to Figure 10, there is shown a set of functional abstraction layers (1000) provided by the cloud computing network of Figure 9. It should be understood that the components, layers, and functions shown in Figure 10 are intended to be illustrative only and are not limiting to the embodiments. As shown, a hardware and software layer (1010), a virtualization layer (1020), a management layer (1030), and a workload layer (1040), and corresponding functions are provided.
[0072] The hardware and software layer (1010) includes hardware and software components. Examples of hardware components include mainframes (e.g., IBM® zSeries® systems), RISC (Reduced Instruction Set Computer) architecture-based servers (e.g., IBM pSeries® systems), IBM xSeries® systems, IBM BladeCenter® systems, storage devices, networks, and network components. Examples of software components include network application server software (e.g., IBM WebSphere® application server software) and database software (e.g., IBM DB2® database software). (IBM, zSeries, pSeries, xSeries, BladeCenter, WebSphere, and DB2 are trademarks of International Business Machines Corporation, registered in many jurisdictions worldwide.)
[0073] The virtualization layer (1020) comprises an abstraction layer that can provide virtual entities such as virtual servers, virtual storage, virtual networks including virtual private networks, virtual applications and operating systems, and virtual clients.
[0074] In one example, the management layer (1030) may provide functionality for resource provisioning, metering and pricing, a user portal, service layer management, and SLA planning and execution. Resource provisioning dynamically procures computing and other resources used to perform tasks within the cloud computing environment. Metering and pricing tracks costs as resources are utilized within the cloud computing environment and sends bills or invoices for the utilization of those resources. In one example, these resources may include application software licenses. Security verifies the identity of cloud users and tasks and protects data and other resources. The user portal provides users and system administrators with access to the cloud computing environment. Service layer management allocates and manages cloud computing resources to meet required service layers. Service level agreement (SLA) planning and execution proactively prepares and procures cloud computing resources in accordance with SLAs in anticipation of future demand.
[0075] The Workload Layer (1040) illustrates examples of functionality available in a cloud computing environment. Examples of workloads and functionality that may be provided from this layer include, but are not limited to, mapping and navigation, software development and lifecycle management, virtual classroom education delivery, data analytics processing, transaction processing, and federated machine learning.
[0076] While particular embodiments of the present invention have been shown and described, it will be apparent to those skilled in the art, based on the contents of this specification, that changes and modifications may be made without departing from the embodiments and their broader aspects. Accordingly, the appended claims encompass within their scope all such changes and modifications that are within the true scope of the embodiments. It should be further understood that the embodiments are defined solely by the appended claims. Where a specific number of introduced claim elements is intended, such intention will be expressly recited in the claims; it will be understood by those skilled in the art that, in the absence of such recitation, no such limitation exists. As an aid to understanding with respect to non-limiting examples, the appended claims below include the use of the introductory phrases "at least one" and "one or more" to introduce claim elements. However, the use of such phrases should not be construed as meaning that the introduction of a claim element by the indefinite article "a" or "an" limits any particular claim containing such introduced claim element to embodiments containing only one such element, even if the same claim also contains the introductory phrases "one or more" or "at least one" and an indefinite article such as "a" or "an," and the same applies to the use of definite articles in the claims.
[0077] The present embodiments may be systems, methods, or computer program products, or combinations thereof. Additionally, selected aspects of the present embodiments may take the form of entirely hardware embodiments, entirely software embodiments (including firmware, resident software, microcode, etc.), or embodiments combining software or hardware aspects or both, all of which may be referred to generally herein as "circuits," "modules," or "systems." Furthermore, aspects of the present embodiments may take the form of a computer program product embodied in a computer-readable storage medium containing computer-readable program instructions for causing a processor to execute aspects of the present embodiments. Thus embodied, the disclosed systems, methods, or computer program products, or combinations thereof, function to improve the functionality and operation of artificial intelligence platforms to apply input encoding techniques in federated learning and enable training of models using a common encoder.
[0078] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. A non-exhaustive list of more specific examples of computer-readable storage media includes portable floppy disks, hard disks, dynamic or static random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), magnetic storage devices, portable compact disc read-only memory (CD-ROM), digital versatile disks (DVDs), memory sticks, floppy disks, mechanically encoded devices such as punch cards or ridge structures in grooves on which instructions are recorded, and any suitable combination thereof. As used herein, computer-readable storage media should not be construed as being ephemeral signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses passing through fiber optic cable), or electrical signals transmitted over wires.
[0079] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or storage device over a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network). This network may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage on a computer-readable storage medium within each computing / processing device.
[0080] The computer-readable program instructions for carrying out the operations of the present embodiments may be source or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or object code written in one or more programming languages, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer as a standalone software package, partially on the user's computer and on a remote computer, or entirely on a remote computer or server or cluster of servers. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, to perform aspects of the present embodiments, electronic circuitry including, for example, programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), may execute computer-readable program instructions to personalize the electronic circuitry by utilizing state information of the computer-readable program instructions.
[0081] Aspects of the present embodiments are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to the present embodiments. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0082] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to create a machine, where the instructions, executed by the processor of the computer or other programmable data processing apparatus, create means for performing the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions may be stored on a computer-readable storage medium and capable of directing a computer, programmable data processing apparatus, or other device, or combination thereof, to function in a particular manner, such that the computer-readable storage medium on which the instructions are stored comprises an article of manufacture containing instructions that implement aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0083] Computer-readable program instructions may be loaded into a computer, other programmable data processing apparatus, or other device such that the instructions, which execute on the computer, other programmable apparatus, or other device, perform the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams, thereby causing a series of operable steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process.
[0084] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, comprising one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions shown in the blocks may occur out of the order shown in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or in the reverse order, depending on the functionality involved. It is also noted that each block in the block diagrams and / or flowchart diagrams, and combinations of blocks included in the block diagrams and / or flowchart diagrams, may be implemented by a special-purpose hardware-based system that performs the specified function(s) or operation(s), or executes a combination of special-purpose hardware and computer instructions.
[0085] Although specific embodiments have been described herein for purposes of illustration, it will be understood that various modifications may be made without departing from the scope of the embodiments. Accordingly, the scope of protection for the embodiments is limited only by the appended claims and their equivalents.
Claims
1. a processing unit operably coupled to the memory; an artificial intelligence (AI) platform in communication with the processing unit, wherein the AI platform applies input encoding in federated learning, and wherein the AI platform: an enrollment manager configured to arrange at least first and second participating entities in a cooperative relationship, wherein the first participating entity trains a first machine learning model using a first encoder on a first training data set and the second participating entity trains a second machine learning model using a second encoder on a second training data set; operatively coupled to the registration manager; measuring the performance of the first and second machine learning models; Selectively identifying at least one of the first and second machine learning models based on the measured performance. an evaluator configured to perform a director operatively coupled to the evaluator and configured to share encoders of the selectively identified machine learning models with each of the first and second participating entities, the shared encoders configured to be applied by the first and second participating entities to train the first and second machine learning models, respectively, and the first and second machine learning models trained using the shared encoders configured to be merged to form a single machine learning model; A system comprising:
2. 10. The system of claim 1, wherein at least the first participating entity trains at least two machine learning models, each of the at least two machine learning models including a separate encoder and a separate training data set, and further comprising the evaluator configured to measure performance of the at least two machine learning models on first test data, and based on their measured performance, selectively identify and share one of the at least two machine learning models with an entity operably coupled to the first and second participating entities.
3. 2. The system of claim 1, further comprising the evaluator configured to test combinations of the first and second machine learning models, wherein the selective identification includes combinations of the first and second encoders.
4. 4. The system of claim 3, further comprising the evaluator configured to create a combination of the first and second encoders and share the combination with the first and second participating entities.
5. The evaluator is configured to share the first and second machine learning models between the first and second participating entities; the first participating entity evaluates the first and second machine learning models using first test data and measures first performance of the first and second machine learning models based on the first test data; the second participating entity evaluates the first and second machine learning models using second test data different from the first test data, and measures second performance of the first and second machine learning models based on the second test data; the first and second participating entities share measured first and second performance data with the evaluator; The system of claim 1 , wherein the evaluator selectively identifies one or more of the first and second machine learning models.
6. 10. The system of claim 1, further comprising the director configured to merge the first and second trained machine learning models to form the single machine learning model.
7. 1. A computer program product for applying input encoding in associative learning, the computer program product comprising: arranging at least first and second participating entities in a cooperative relationship to train machine learning models, wherein the first participating entity trains a first machine learning model using a first encoder on a first training data set and the second participating entity trains a second machine learning model using a second encoder on a second training data set; evaluating the first and second machine learning models by an evaluation entity operably coupled to the first and second participating entities; measuring the performance of the first and second machine learning models; Selectively identifying at least one of the first and second machine learning models based on the measured performance; and said evaluating comprising: sharing encoders of the selectively identified machine learning models with each of the first and second participating entities, wherein the shared encoders are configured to be applied by the first and second participating entities to train the first and second machine learning models, respectively, and the first and second machine learning models trained using the shared encoders are configured to be merged to form a single machine learning model; A computer program for executing the above.
8. 8. The computer program product of claim 7, wherein at least the first participating entity trains at least two machine learning models, each of the at least two machine learning models including a separate encoder and a separate training data set, and wherein the computer program causes the processor to evaluate performance of the at least two machine learning models on first test data and selectively identify and share with the evaluation entity one of the at least two machine learning models based on their measured performance.
9. the processor, sharing the first and second machine learning models between the first and second participating entities; evaluating, by the first participating entity, the first and second machine learning models using first test data and measuring first performance of the first and second machine learning models based on the first test data; evaluating, by the second participating entity, the first and second machine learning models using second test data different from the first test data, and measuring second performances of the first and second machine learning models based on the second test data; Sharing the measured first performance and second performance data by the first and second participating entities with the evaluation entity; selectively identifying, by the evaluation entity, one or more of the first and second machine learning models; 8. The computer program product of claim 7, further comprising:
10. A method for computer information processing, comprising: arranging at least first and second participating entities in a cooperative relationship to train machine learning models, wherein the first participating entity trains a first machine learning model using a first encoder on a first training data set and the second participating entity trains a second machine learning model using a second encoder on a second training data set; evaluating the first and second machine learning models by an evaluation entity operably coupled to the first and second participating entities; measuring the performance of the first and second machine learning models; selectively identifying at least one of the evaluated first and second machine learning models based on the measured performance; and said evaluating comprising: sharing encoders of the selectively identified machine learning models with each of the first and second participating entities, wherein the shared encoders are configured to be applied by the first and second participating entities to train the first and second machine learning models, respectively, and the first and second machine learning models trained using the shared encoders are configured to be merged to form a single machine learning model; A method comprising:
11. 11. The method of claim 10, wherein at least the first participating entity trains at least two machine learning models, each of the at least two machine learning models including a separate encoder and a separate training data set, the method further comprising the first participating entity evaluating performance of the at least two machine learning models on first test data and selectively identifying and sharing with the evaluation entity one of the at least two machine learning models based on their measured performance.
12. The method comprises: sharing the first and second machine learning models between the first and second participating entities; the first participating entity evaluating the first and second machine learning models using first test data and measuring first performance of the first and second machine learning models based on the first test data; the second participating entity evaluating the first and second machine learning models using second test data different from the first test data and measuring second performance of the first and second machine learning models based on the second test data; the first and second participating entities sharing measured first and second performance data with the evaluation entity; the evaluation entity selectively identifying one or more of the first and second machine learning models; The method of claim 10 further comprising:
Citation Information
Patent Citations
Distributed machine learning system, apparatus, and method
JP2019526851A
Information processing device, information processing method, and program
JP2020067954A
Medical information processing apparatus, medical information processing system, and medical information processing method
JP2021056995A
Moderator for federated learning
WO2021070189A1