Client-based retraining and evaluation of artificial intelligence models
By performing decentralized retraining and performance evaluation of artificial intelligence models at multiple clients, the AI model performance deterioration caused by noisy data and poor training data sets is solved, and more efficient and generalized AI model training is achieved.
Patent Information
- Application Number
- CN202411615570.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-10-09
- Filing Date
- 2024-11-13
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art faces the problem of noise input-output data and poor training datasets when decentralized retraining of artificial intelligence models at multiple clients, resulting in the performance of retrained AI models that may deteriorate.
Provides a framework that allows acquisition and retraining of AI models at multiple clients and evaluates the performance of AI models through an iterative and decentralized approach. The framework includes deploying AI models to multiple clients, obtaining the training data set of clients, retraining the AI model, aggregating the client's weights, forming an AI model in the third training state, and evaluating it.
By distributed retraining and cross-evaluating the performance of the AI model at multiple clients, the negative impact of noisy data on model performance is reduced, the generalization ability and overall performance of the AI model are improved, and the privacy data of the client is protected.
Smart Images

Figure CN120012958A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of priority from European patent application No. 23209774.1 filed on November 14, 2023, the contents of which are incorporated by reference. Technical Field
[0003] Various examples of the present disclosure generally relate to retraining an artificial intelligence model at multiple clients and evaluating the retrained artificial intelligence model at multiple clients. Background Art
[0004] An artificial intelligence (AI) model (sometimes also called a machine learning model) is a function that is trained using pairs of input and output data that form a training data set. Various types of AI models are known, including support vector machines and deep neural networks. The parameters of an AI model are set in optimization to best reproduce the output data based on the input data.
[0005] The training of AI models is usually computationally expensive. In addition, the accuracy of AI models depends on the training dataset. For example, it is desirable that the training dataset fully samples the input space observed in actual deployment scenarios. Otherwise, the predictions made by the AI model may be inaccurate.
[0006] In order to address such aspects, a known technology is that, first, the AI model is pre-trained at a central authority based on an initial training data set. The initial training data set is usually collected by experts. Individual input-output data is carefully selected (curate) by experts. Then, the pre-trained AI model can be deployed to multiple clients. The client is a device or node that is not directly affected by the central authority. That is, they can use the pre-trained AI model independently. The client can perform reasoning based on the pre-trained AI model. Performing reasoning means collecting input data and making predictions of the AI model without available ground truth.
[0007] Scenarios are known in which the user then confirms or abandons a prediction made by the AI model, thereby generating baseline truth output data. Therefore, paired input-output data can be included in another training data set generated at each client. Based on such a training data set determined at the client, it is possible to retrain the AI model.
[0008] Such retraining of AI models can be performed in a decentralized manner at the client without involving a central authority. That is, different clients collect different training data sets, and retrain the AI model locally to obtain the corresponding updated training state of the AI model. This has the advantage that the training data set does not have to be provided to a central authority. Privacy and data security can thus be ensured.
[0009] However, such decentralized retraining techniques for AI models face certain constraints and drawbacks. In real-world environments, the input-output data residing at participating clients is inherently noisy. Some clients may generate poor training data sets. Due to this, the performance of the retrained AI model may potentially deteriorate. Summary of the invention
[0010] The present disclosure provides a framework for the deployment of artificial intelligence models from a central server. The present framework can be used to retrain artificial intelligence models at multiple clients and evaluate the retrained artificial intelligence models at multiple clients. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] A more complete appreciation of the present disclosure and its many attendant aspects will be readily obtained as it becomes better understood by reference to the following detailed description when considered in conjunction with the accompanying drawings.
[0012] Figure 1 A system of a central server implementing a central authority for multiple clients of the system according to various examples is schematically illustrated.
[0013] Figure 2 A device according to various examples is schematically illustrated, wherein the device may implement a central authority or a client.
[0014] Figure 3 Schematically illustrates functional aspects of a client according to various examples.
[0015] Figure 4 Schematically illustrates functional aspects of a grouping of multiple clients according to various examples.
[0016] Figure 5 is a flow chart of a method according to various examples.
[0017] Figure 6 is a flow chart of a method according to various examples.
[0018] Figure 7 is a flow chart of a method according to various examples.
[0019] Figure 8 is a flow chart of a method according to various examples. DETAILED DESCRIPTION
[0020] Regardless of grammatical term usage, individuals with male, female, or other gender identities are included in the term.
[0021] Some examples of the present disclosure generally provide multiple circuits or other electrical devices. All references to circuits and other electrical devices and the functions each of them provides are not intended to be limited to the things only included in the illustrations and descriptions herein. Although specific marks can be assigned to the various circuits or other electrical devices disclosed, such marks are not intended to limit the scope of operation for circuits and other electrical devices. Such circuits and other electrical devices can be combined and / or separated in any way based on the desired specific type of electrical implementation. It is to be recognized that any circuit or other electrical device disclosed herein may include any number of microcontrollers, graphics processor units (GPUs), integrated circuits, memory devices (e.g., flash memory, random access memory (RAM), read-only memory (ROM), electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM) or other suitable variants thereof), and software that acts together to perform (multiple) operations disclosed herein. In addition, any one or more of the electrical devices may be configured to execute a program code, which is contained in a non-transitory computer-readable medium programmed to perform any number of functions as disclosed.
[0022] In the following, embodiments of the present invention will be described in detail with reference to the accompanying drawings. It is to be understood that the following description of the embodiments is not to be considered in a restrictive sense. The scope of the present invention is not intended to be limited by the embodiments described below or by the accompanying drawings, which are considered to be illustrative only.
[0023] Various disclosed aspects relate to acquisition of a training data set for retraining an AI model at multiple clients. Various aspects relate to retraining of an AI model at multiple clients. Various disclosed aspects relate to an iterative and decentralized method for determining a client that generates a noisy / poor quality training data set. Various aspects relate to evaluation of performance of an AI model in a training state after retraining. Various aspects relate to cross-evaluation of performance at multiple clients.
[0024] The present disclosure provides a computer-implemented method. The method includes deploying an artificial intelligence model. The artificial intelligence model is deployed in a first training state. The artificial intelligence model is deployed to multiple clients. The method also includes, at each of the multiple clients: obtaining a corresponding training data set. The method further includes, at each of the multiple clients: retraining the artificial intelligence model; performing this retraining to obtain an artificial intelligence model in a corresponding second training state. The method also includes associating the multiple clients with multiple groups. The method also includes, for each of the multiple groups: aggregating the weights of the artificial intelligence model in the second training state associated with the clients in the corresponding group to obtain an artificial intelligence model in a corresponding third training state. The method further includes evaluating the artificial intelligence model in each of the third training states.
[0025] In addition, the present disclosure provides a computer-implemented method for use in a client, the method comprising: obtaining an artificial intelligence model in a first training state from a central authority (e.g., a central server); performing reasoning based on the artificial intelligence model in the first training state; obtaining a training data set based on the execution of the reasoning; retraining the artificial intelligence model based on the training data set to obtain an artificial intelligence model in a second training state; providing the artificial intelligence model in the second training state to the central authority and at least one of one or more additional clients; establishing an artificial intelligence model in multiple third training states based on information obtained from the central authority or at least one of the one or more additional clients; and evaluating the artificial intelligence model in each of the multiple third training states based on benchmarking relative to the performance of the artificial intelligence model in the third training state and based on the training data set.
[0026] In addition, the present disclosure provides a computer-implemented method for use in a central authority, the method comprising: deploying an artificial intelligence model in a first training state to multiple clients; obtaining an artificial intelligence model in a corresponding second training state from each of the multiple clients; associating the multiple clients with multiple groups; for each of the multiple groups: aggregating the weights of the artificial intelligence models in the second training state associated with the clients within the corresponding group to obtain an artificial intelligence model in a corresponding third training state; and providing the artificial intelligence model in the third training state to multiple clients for evaluation.
[0027] It is understood that the features mentioned above and those yet to be explained below can be used not only in the respective combination indicated but also in other combinations or alone, without departing from the scope of the present invention.
[0028] Various examples of the present disclosure generally relate to predictions using AI models. As a general rule, various types of AI models can be amenable to the techniques disclosed herein, e.g., support vector machines, deep neural networks, random forest algorithms, to name just a few examples.
[0029] Predictions using AI models can be used in a variety of use cases. For example, an AI model can operate based on medical data, such as medical imaging data. The AI model can estimate one or more properties of the medical imaging data. For example, it would be possible to perform image reconstruction based on magnetic resonance imaging (MRI) data. For example, it would be possible to detect one or more features in the medical imaging data, such as lesions in an MRI image, fractures in a computed tomography (CT) image, to give just a few examples. Anatomical structures can be segmented in an image. Treatment planning based on an AI model is another use case.
[0030] Various examples of the present disclosure generally relate to a training process for training an AI model. An iterative training process is disclosed. The AI model initially exists in an initial training state and is retrained to obtain an updated training state. This can be repeated multiple times.
[0031] For example, training a deep neural network (as an example of an AI model) includes using a training data set consisting of paired input-output data. Each pair provides an example from which the deep neural network can learn by adjusting its internal parameters to map the input to the expected output (called the benchmark truth). The goal during training is to optimize the parameters of the network to minimize the difference between the network's prediction and the benchmark truth. This optimization is usually implemented using a gradient descent algorithm. For each input-output pair, the network's current prediction is compared with the benchmark truth to calculate the error. Backpropagation is then used to calculate the gradient of this error with respect to each network parameter. These gradients indicate how much each parameter should be adjusted to reduce the total error. The gradient descent algorithm uses these gradients to update the parameters of the network in the direction of reducing the error. This process is iteratively performed for the entire training data set until the network's prediction is closely consistent with the benchmark truth for most input-output pairs.
[0032] According to various examples, retraining and in particular computationally expensive optimization are performed at multiple clients based on training data sets acquired locally at the multiple clients. This may be referred to as distributed retraining. There is no requirement for a central authority to evaluate the training data set before performing retraining. This protects privacy because the training data set does not need to be exposed to a central authority or other clients. By implementing retraining in a distributed manner at multiple clients, the training data set may be retained at each of the multiple clients. There is no requirement to share the training data set between the multiple clients or with a central authority. Thus, privacy is maintained. Figure 1 A corresponding system is disclosed in .
[0033] Figure 1 The system 90 is schematically illustrated. The system 90 includes a central server 91, which implements a central authority of the system 90. For example, the central server 91 can be maintained and operated by a developer of an AI model. The developer of the AI model can also be a manufacturer of certain medical imaging equipment operating at the sites of multiple clients 92, 93, 94.
[0034] As a general rule, the client may be implemented differently depending on the specific use case. For example, for the processing of medical image data, the client may be implemented by a computer connected to a radiology device (e.g., a magnetic resonance tomography (MRT) device or a CT device) at a hospital.
[0035] like Figure 1 The clients 92, 93, 94 depicted in the figure are connected to the central server 91 via the Internet 95. The central server 91 can provide certain data or information to each of the clients 92, 93, 94. For example, the central server 91 can deploy an AI model in an initial training state to each of the clients 92, 93, 94; thereby, facilitating reasoning tasks at each of the clients 92, 93, 94. The clients 92, 93, 94 can also provide data to the central server, such as a retrained AI model (in a further training state). Typically, the input data used for reasoning by the AI model is retained locally at each of the clients 92, 93, 94, that is, not shared with the central server 91 or other clients.
[0036] Figure 2 The device 80 is schematically illustrated. The device 80 may implement Figure 1 The device 80 includes a central server 91 or any of the clients 92, 93, 94 of the system 90. The device 80 includes a processor 97 coupled to a memory 99. The processor 97 can load program code from the memory 99 and execute the program code. Based on the execution of the program code, the processor 97 can perform one or more tasks as disclosed herein, for example: deploying an AI model to multiple clients; obtaining an AI model from a central authority; communicating with one or more clients and / or a central authority; controlling a graphical user interface output to a user via a human-machine interface 96; obtaining medical imaging data via a communication interface 98; processing medical imaging data based on an AI model; retraining the AI model to obtain an AI model in an updated training state; communicating the weights of the AI model via the communication interface 98; and the like.
[0037] Figure 3Retraining the AI model at each of the clients 92, 93, 94 is schematically illustrated. The respective clients 92, 93, 94 store data defining the AI model in the first training state 81. The respective clients 92, 93, 94 also store a training data set 82. Based on the training data set 82, retraining of the AI model in the first training state 81 may be performed to obtain data defining the AI model in the second training state 83.
[0038] The training data set 82 may be acquired locally. That is, the training data set may include paired input-output data acquired at one or more sensors or machines or processes executed locally at the respective clients 92, 93, 94 (e.g., at one or more machines connected to the respective clients 92, 93, 94 via a local area network (rather than via a public network such as the Internet 95)). The training data set 82 may be acquired at one or more machines under the control of the same operator who also operates the respective clients 92, 93, 94. Typically, the training data set 82 is of limited size, for example, if compared to a training data set used in the initial training of an AI model at a central authority such as a central server 91 (e.g., acquired using a dedicated measurement campaign).
[0039] Various techniques are based on the discovery that retraining on a small training data set generated on a single client can produce a new corresponding training state of the AI model that does not generalize well on other clients. In the case of data poisoning (i.e., incorrect or poor labeling of data occurs at the client), an AI model trained on such client(s) may result in poor generalization or suboptimal performance on other clients.
[0040] For the above reasons, it is desirable to evaluate the performance of the AI model in the second training state 83.
[0041] In one example, such evaluation may include a cross-evaluation process at multiple clients.
[0042] In a reference implementation of such a cross-evaluation process, if there are N clients that retrain an AI model in a first training state, then there are N second training states of the available AI models after retraining. This would mean that there are N instances of AI models that need to be evaluated and they need to be communicated to all available clients. From the perspective of a client, it should not only train the model on its local data, but it also needs to evaluate N training states (its own training state and the training states from the remaining N-1 clients) on its training dataset. This reference implementation of the cross-evaluation process is not only time-consuming but also may not be feasible because the required resource requirements are too huge.
[0043] Therefore, according to the present disclosure, for the purpose of evaluating a new training state of an AI model, the concept of grouping multiple clients into groups is used. This is shown in Figure 4 FIG.
[0044] Figure 4 Schematically illustrates aspects regarding the grouping of multiple clients (clients 92-1, 92-2, 92-3 here as part of group 70). Figure 4 FIG. illustrates a scenario in which AI models in multiple second training states 83-1, 83-2, 83-3 (as obtained at multiple clients 92-1, 92-2, 92-3) are combined to obtain an AI model in a third training state 85.
[0045] More generally, multiple clients may be associated with multiple groups. For example, all N clients are divided into M groups and M << N. These groups may be non-overlapping, or in other words, mutually exclusive of each other. In the context of the example provided above, the N clients are broken down into groups (group 1, group 2, group 3, …, group M-1, group M) and each group consists of L clients, where L = N / M.
[0046] Then, as shown in Figure 4 FIG., the weights of the AI models in the second training state (i.e., as retrained at each client) are aggregated within a group to form an AI model in a third training state. In other words, the weights of L individual retrained AI models are fused to form a single training state.
[0047] As a general rule, in various disclosed scenarios, such combination / fusion of weights can use various aggregation strategies, e.g., averaging of weights or weighted average of weights. The number of data points used to train the corresponding training state being fused can be used as a weighting scheme. That is, the larger the underlying training dataset, the greater the weight.
[0048] By combining multiple training states, the number of training states to be evaluated can be reduced from N to M, ie, Model 1, Model 2, Model 3, ... Model M.
[0049] It is then possible to evaluate an instance of the AI model that is in the third training state.
[0050] For example, each third training state may be evaluated at a central authority. To this end, the third training state may be communicated to a central server implementing the central authority. A benchmark data set may be available at the central authority and used for said evaluation.
[0051] In other examples, a cross-evaluation process at multiple clients can be used. For the cross-evaluation process, the AI model in each of the third training states is transferred to the clients that form the remaining other groups. For example, the third training state obtained for group 1 is transferred to at least one client of each of group 2, group 3, ..., group M, and is evaluated there, that is, cross-validation is achieved. For example, the AI model in the third training state can be benchmarked relative to the AI model in the first training state, as previously deployed by a central authority. Such a benchmark test can include performing reasoning based on input data of a corresponding training data set obtained at the client (or another data set of paired input-output suitable for verification), and comparing the deviation from the benchmark fact output data of the training data set. For larger deviations, poor performance is determined.
[0052] Then, two situations can be imagined: First, the AI model in the third training state may perform better than the AI model in the first training state. That is, the AI model in the third training state generalizes well on other clients and can therefore be regarded as a replacement or update for the current version of the AI model. A new deployment can be triggered. Second, the AI model in the third training state may also perform worse than the AI model in the first training state. The AI model in the third training state does not generalize well on other clients.
[0053] In such a scenario, it may be possible to break down a specific client group associated with a poorly performing third training state into smaller subgroups. For example, group 1 may be broken down into B smaller subgroups (subgroup 1_1, subgroup 1_2, subgroup 1_3, subgroup 1_4, …, subgroup 1_B), and B << M. That is, in response to evaluating that the AI model in the corresponding third training state does not meet a predefined benchmark, those clients previously associated with the corresponding group are re-associated with a plurality of newly formed (sub)groups that replace the initial group. Each of the newly formed groups consists of P clients, where P = M / B. Then, the previously described steps can be repeatedly performed for the newly formed subgroups. That is, from each of the newly formed subgroups, the corresponding second training states of the AI models associated with the clients in that subgroup are combined to form the corresponding third training state. Then the cross-evaluation process may be re-executed. That is, then the third training state is transferred to the clients of the remaining groups, and their performance is compared with the AI models in the first training state based on the corresponding local training datasets. If the performance of the newly determined third training state is not inferior to the performance of the first training state, they are retained; if the performance is inferior, the above-mentioned process is repeated.
[0054] The technique of starting with relatively large groups and breaking down those larger groups associated with poorly performing third training states into smaller groups can be labeled as an "iterative cross-evaluation process". Clients are initially associated with relatively large groups; this initially minimizes the computational effort of performing the cross-evaluation process. However, if a poorly performing third training state of the AI model (associated with a relatively large group) is detected, breaking down this group into smaller groups in the process can be repeated. This helps to identify those clients that compromise data quality (e.g., provide a poor-quality training dataset).
[0055] Finally, all the third training states that pass the evaluation can be combined. That is, the weights of all the third training states in the cross-evaluation process can be combined, for example, by averaging or by determining a weighted average. Similar techniques as described above for combining the weights of the AI models in the second training state to obtain the AI models in the corresponding third training state can also be applied. This results in a fourth training state, which can then be redeployed, for example, by a central authority.
[0056] Since all the components of the AI model in the fourth training have been successfully evaluated, it provides a carefully selected performance. In addition to achieving faster evaluation, this method also helps to identify abnormal clients where there is a possibility of data poisoning and avoid their contribution when proceeding to create an updated version of the algorithm. This information can also be considered to inform the participating clients of potential data poisoning or even remove the client from the training steps.
[0057] In a scenario where there are N clients, then the number of evaluations that need to be performed using conventional evaluation techniques is equal to N2. However, in the proposed method, the N clients are divided into M subgroups to produce M AI models that need to be evaluated. Let's say, one of the M models in the third training state performs poorly, so these clients are further decomposed to form B groups with a corresponding number of third training states that also need to be evaluated at (M-1) subgroups. Therefore, the number of evaluations required in this scenario is equal to [M2+(M-1)x B]. It should be noted here that B< <M<<N。
[0058] Figure 5 is a flow chart of a method according to various examples. Figure 5 The methods usually involve client-based retraining of AI models. For example, Figure 5 The method may be performed by a system including a central server and a plurality of clients, for example, Figure 1 The system 90 is executed. Distributed processing at multiple devices is possible. For example, processors at multiple devices can load program code from corresponding memories and execute the program code to participate in Figure 5 Method (reference Figure 2 :Device 80).
[0059] At box 3005, the AI model in a first training state is deployed.
[0060] Deploying AI models can include Figure 1 ) For example, to multiple clients via the Internet (refer to Figure 1 : Clients 92, 93, 94) transmit one or more messages. Such messages may include the weights of the AI model in the first training state.
[0061] Deploying the AI model may include receiving one or more such messages at each of a plurality of clients.
[0062] At block 3010, inference may then be performed based on the AI model in the first training state deployed in block 3005. This means that input data is provided to the AI model in the first training state, and output data is obtained from the AI model in the first training state after corresponding calculations. When performing inference at 3010, no ground truth is required to be available.
[0063] Thus, grouping box 3091 corresponds to the reasoning stage of using the AI model to perform reasoning.
[0064] Then, retraining begins at group box 3092.
[0065] Retraining includes, at box 3015, obtaining a training data set, wherein a corresponding training data set is obtained at each client. Therefore, each training data set is client-specific. Different clients have different training data sets. Each training data set includes a set of paired input data-output data. The output data is a reference fact associated with the nominal prediction of the AI model. There are various options available for obtaining such a training data set, and generally obtaining the training data set at box 3015 can be based on box 3010, that is, based on the execution of reasoning. In particular, it is often possible that, although the reference fact is not available when performing reasoning, based on user interaction, the output data is then positively verified to correspond to the nominal output or is changed by the user to constitute the reference fact. Give a specific example: it is possible that, at box 3010, the AI model is used to segment a certain structure in medical imaging data, such as a lesion. Then, this segmentation prediction provided by the AI model in the first training state can be output to the user via a graphical user interface. The user can or only confirm that the segmentation is correct, thereby generating a reference fact. Alternatively, the user can change the segmentation, such as by locally adapting the segmentation map, etc., to generate a reference fact. Over time, a training data set of sufficient size is obtained at box 3015 for performing retraining of the AI model.
[0066] Finally, it is then possible to retrain the AI model at block 3020 to obtain a second training state. This is performed at each of the multiple clients based on the local training data set. Thus, once block 3020 is completed, the number of second training states corresponds to the number of clients.
[0067] According to an example, certain clients may be pre-identified as unreliable, for example by a central authority or based on a local process implemented at the client. For example, certain clients may be identified as anomalous clients with a possibility of data poisoning due to noisy measurement data, etc. This may then generally suppress the influence of the AI model in the respective second training state on further processing, more specifically, on subsequently deployed training states of the AI model. There are different options for achieving such suppression of the influence of the respective client on the formation of further training states to be deployed, and in Figure 5 An option is illustrated in FIG. Here, at box 3025, for all clients marked as outliers, the corresponding second training state of the AI model is discarded. This can be done, for example, by a central authority. It would also be possible to pre-configure such clients so that they do not need to retrain the AI model to even obtain the second training state.
[0068] At box 3035, all clients are associated with multiple groups. Each client can be associated with exactly a single group. The groups may not overlap. It will also be possible that the groups overlap. The group is a logical collection of multiple clients. The groups are formed for the purpose of evaluating the AI model.
[0069] Thus, at block 3040, the weights of the AI models in the second training state associated with any given group of clients are aggregated separately to obtain AI models in the corresponding third training state. Thus, at block 3040, the number of third training states corresponding to the number of groups is obtained.
[0070] It is then possible to evaluate the AI model in each of the third training states, for example, using a cross-evaluation process at multiple clients, at box 3045. Thus, as will be appreciated from above, boxes 3025, 3035, 3040, and 3045 correspond to the evaluation / validation marked by grouping box 3039.
[0071] When the AI models in each of the third training states are evaluated at box 3045, the weights of the AI models in those third training states associated with the positive results of the evaluation are aggregated at box 3050 to obtain the AI model in the fourth training state.
[0072] Then, block 3005 may be re-executed to deploy the AI model in the fourth training state. To this end, the AI model in the fourth training state may be communicated to a central authority, or the aggregation of block 3050 may be implemented directly at the central authority.
[0073] As will be appreciated from the above, the computational complexity of the evaluation at block 3045 may be reduced by combining the plurality of respective second training states of the AI model into a corresponding third training state of the AI model at block 3040. The grouping at block 3035 may take into account a plurality of criteria, such as the total count of the groups, the number of clients per group, and / or the size of the training data set at each of the plurality of clients.
[0074] For example, it is desirable that the number of pairs of input data-output data forming the basis of each third training state is comparable across the groups. Thus, the groups may be formed such that the sum of the training data sets or sizes for all clients in each group is approximately stable across all groups. Furthermore, the number of groups may be said to exceed a certain minimum threshold, such as four groups or ten groups, etc. This may also be based on the number of participating clients.
[0075] Figure 6 is a flow chart of a method according to various examples. Figure 6 About Figure 5A specific example implementation of the evaluation process of the grouping block 3093 of FIG. The example implementation includes aspects related to conditional iterative regrouping of clients depending on the results of the evaluation process.
[0076] For example, Figure 6 The method may be performed by a system including a central server and a plurality of clients, for example, Figure 1 The system 90 is executed. Distributed processing at multiple devices is possible. For example, processors at multiple devices can load program code from corresponding memories and execute the program code to participate in Figure 6 Method (reference Figure 2 :Device 80).
[0077] At block 3104, an initial grouping is performed. This initial grouping may be based on one or more of the following decision criteria: the total count of groups; the number of clients in each group; the number of pairs of input data-output data in the training data set (i.e., the size of the training data set). This has been discussed above in conjunction with Figure 5 Explained.
[0078] At box 3015, a third training state of the AI model is obtained for each group according to the initial grouping of box 3104. For example, if there are M groups, there are M third training states.
[0079] Then, at block 3110, a current group is selected from among all current groups. In particular, the current group is selected from a series of all groups that have not been previously evaluated.
[0080] At optional box 3115, it can be determined whether the currently selected group of the current iteration of box 3110 is trusted. In a real-world environment, the trust associated with some clients will be higher than the rest of the clients. The trust level can be assigned by a central authority. This means that when compared to other clients, trustworthy clients generate cleaner data or less noisy data. If the model to be evaluated is generated by a group consisting of trustworthy clients, the corresponding evaluation can be skipped because this group will not produce poor / noisy data. According to the example, as indicated by the "yes" path exit box 3115, the third training state of the AI model that evaluates the trust group (i.e., a group that includes only trusted clients or a majority of trusted clients) can be skipped / bypassed.
[0081] At block 3120, the third training state of the AI model of the currently selected set in the current iteration of block 3110 is evaluated relative to the current baseline. Typically, the current baseline is an AI model previously deployed by a central authority, i.e., an AI model in a first training state in the terms used above.
[0082] Block 3120 may include a cross-evaluation process. To this end, the AI model in the third training state associated with the currently selected group may be provided to one or more clients of one or more remaining groups and evaluated locally relative to the current benchmark based on their corresponding local training data sets. In some examples, the corresponding third training state may be provided to all other clients of all remaining groups for evaluation relative to their local training data sets.
[0083] At block 3125, it is determined whether the corresponding third training state of the AI model of the currently selected group (current iteration of block 3110) meets the predefined benchmark, that is, the result of judgment block 3120. For example, it can be determined whether the corresponding third training state of the AI model performs better than the first training state of the AI model for all remaining clients.
[0084] If the predefined benchmark is met, the currently estimated third training state is added to the combination queue block 3145. Also, the corresponding group is marked as evaluated so that it will not be processed in any further iteration of block 3110. Subsequently, at block 3150, it is determined whether another group that has not yet been evaluated is to be processed at another iteration of block 3110.
[0085] However, if at box 3125, the performance of the corresponding third training state of the AI model does not meet the predefined benchmark, box 3130 is executed. At box 3130, it is determined whether the currently selected group (the current iteration of box 3110) includes more than the lower threshold (e.g., configured by the central authority). For example, it can be checked whether the currently selected group includes more than a single client. In the affirmative case, box 3135 is executed. If the check at box 3110 does not pass, the current group is discarded, that is, marked as evaluated, and because it has a negative evaluation result, it is not added to the combination queue. Otherwise, it is determined that the current group can be further decomposed into smaller groups, so that at box 3135, in response to the evaluation that the AI in the current third training state does not meet the predefined benchmark, the client previously associated with the current group is associated with a plurality of newly formed groups that replace the corresponding group. In addition, the weights of the AI model in the second training state for all clients in each of the newly formed groups are combined to obtain a plurality of new third training states. Then, box 3110 is repeated.
[0086] Figure 7 is a flow chart of a method according to various examples. Figure 7 The method may be performed by one of the plurality of clients, for example, Figure 1 For example, the processor may load the program code from the corresponding memory and execute the program code to execute Figure 7 Method (reference Figure 2 : processor 97 and memory 99).
[0087] Figure 7 Implementation of the method Figure 5 and optional Figure 6 Client-side processing of the method.
[0088] At block 3205, an AI model in a first training state is obtained. That is, the client is deployed with an AI model in a first training state. The AI model in the first training state can be obtained from a central authority. Figure 5 Box 3005 discusses corresponding techniques for deploying AI models.
[0089] At block 3210, the AI model in the first training state as obtained in block 3205 is used to perform reasoning; Figure 5 This is explained in box 3010 in .
[0090] At block 3215, a training data set is obtained, for example, based on the execution of the inference at block 3210. Figure 5 Box 3020 discusses the corresponding techniques.
[0091] At block 3220, retraining is performed to produce a second training state of the AI model. Figure 5 : Frame 3020 and combined Figure 3 The corresponding techniques are discussed.
[0092] At block 3225, the AI model (i.e., its weights) in the second training state is provided to, for example, a central authority or one or more additional clients. The training data set may be retained locally at the client to maintain privacy.
[0093] Then, at block 3230, the AI model is established in a plurality of third training states based on information obtained from at least one of the central authority or one or more additional clients. For example, data defining the plurality of third training states may be obtained from the central authority or other clients.
[0094] As previously explained in conjunction with block 3040, the third training state is obtained by combining the weights of the plurality of second training states. This combination may be performed at a central authority or other client. It will also be possible to perform Figure 7In this case, block 3230 may include obtaining AI models in multiple second training states from multiple additional clients or a central authority, and associating the client and multiple additional clients with multiple groups, and then aggregating the weights of the AI models in the second training state associated with the clients in each group to obtain the corresponding third training state of the AI model. Here, grouping information obtained from the central authority may be considered, for example, indicating a mapping of clients to groups or one or more grouping criteria, such as the size of the training data set, etc.
[0095] At block 3235 , the third training state is evaluated. The corresponding techniques have been previously discussed in conjunction with block 3045 .
[0096] Figure 8 is a flow chart of a method according to various examples. Figure 8 Involved with Figure 5 method and optionally with Figure 6 The server-side processing associated with the method, ie, the logic implemented at the central authority.
[0097] For example, a processor at the server may load program code from a corresponding memory and execute the program code to perform Figure 8 Method (reference Figure 2 : processor 97 and memory 99).
[0098] At block 3305, the AI model in the first training state is deployed to multiple clients. Figure 5 Box 3005 discusses the corresponding technology.
[0099] At block 3310, an AI model in a corresponding second training state is obtained from each of the plurality of clients. Block 3310 is combined with the above Figure 7 The boxes 3225 discussed are interrelated.
[0100] At block 3315, multiple clients are associated with multiple groups. Figure 5 Box 3035 of the method discusses corresponding techniques.
[0101] At block 3320, for each of the plurality of groups, the weights of the AI models in the second training state associated with the clients in the corresponding group are aggregated, thereby obtaining AI models in a plurality of third training states.
[0102] At block 3325, multiple third training states are evaluated. This can be implemented either at the server by benchmarking against a local training dataset and comparing to the performance of the AI model in the first training state deployed, for example, in block 3305; or can be offloaded to multiple clients by providing the AI model in the third training state to the multiple clients.
[0103] At block 3330, based on the evaluation result of block 3325, the weights of the AI models in the second training state and / or the third training state are aggregated to obtain an AI model in a fourth training state. Figure 6 (Particularly the performance check at box 3125 and subsequent boxes) discuss corresponding techniques.
[0104] Then, at box 3335, the AI model in the fourth training state may be redeployed.
[0105] Although Figure 8 A scenario is illustrated in which the grouping of clients is performed at a central authority at box 3315 and the weights of the AI model in the second training state are aggregated at the central authority at box 3320, but in conjunction with the previous Figure 7 : In some scenarios discussed in block 3230, it will also be possible to perform such aggregation and / or grouping at the client. To this end, the auxiliary grouping information may be provided to multiple clients by a central authority.
[0106] In summary, techniques for retraining and evaluating AI models have been disclosed. The disclosed evaluation process is faster than conventional evaluation techniques because it requires fewer operations. Distributed evaluation of real-world training datasets is possible. Furthermore, the evaluation mechanism helps isolate clients that would degrade the performance of the overall algorithm if taken into account.
[0107] To further summarize, at least the following examples have been disclosed.
[0108] Example 1. A computer-implemented method comprising:
[0109] - deploying (3005) the artificial intelligence model in the first training state (81) to a plurality of clients (80, 92, 92-1, 92-2, 92-3, 93, 94);
[0110] - At each of the plurality of clients (80, 92, 92-1, 92-2, 92-3, 93, 94): obtaining (3015) a corresponding training data set (82);
[0111] - at each of the plurality of clients (80, 92, 92-1, 92-2, 92-3, 93, 94): retraining (3020) the artificial intelligence model to obtain an artificial intelligence model in a corresponding second training state (83, 83-1, 83-2, 83-3);
[0112] - associating (3035) a plurality of clients (80, 92, 92-1, 92-2, 92-3, 93, 94) with a plurality of groups (70);
[0113] - for each of the plurality of groups (70): aggregating (3040) weights of artificial intelligence models in a second training state (83, 83-1, 83-2, 83-3) associated with clients (80, 92, 92-1, 92-2, 92-3, 93, 94) within the respective group (70) to obtain an artificial intelligence model in a respective third training state (85); and
[0114] -Evaluating (3045) the artificial intelligence model in each of the third training states (85).
[0115] Example 2. The computer-implemented method of Example 1,
[0116] Wherein, at multiple clients (80, 92, 92-1, 92-2, 92-3, 93, 94), a cross-evaluation process is used to evaluate the artificial intelligence model in each of the third training states (85).
[0117] Example 3. The computer-implemented method of Example 2,
[0118] Therein, the cross-evaluation process includes a benchmark relative to the performance of the artificial intelligence model in a first training state (81), which is based on a corresponding training data set at a corresponding client (80, 92, 92-1, 92-2, 92-3, 93, 94).
[0119] Example 4. The computer-implemented method of Example 2 or Example 3,
[0120] The cross-evaluation process includes providing an artificial intelligence model in a corresponding third training state (85) associated with a given one of the multiple groups (70) to at least one other group (70) of the multiple groups (70), and benchmarking the artificial intelligence model in a corresponding third training state (85) associated with a given one of the multiple groups at each of the at least one other group.
[0121] Example 5. The computer-implemented method of any of the preceding examples, further comprising:
[0122] -For each of the multiple groups: in response to the evaluation that the artificial intelligence model in the corresponding third training state does not meet the predefined benchmark, reassociating the clients previously associated with the corresponding group with multiple newly formed groups that replace the corresponding group, and repeating the aggregation of weights and evaluating the newly formed groups.
[0123] Example 6. The computer-implemented method of any of the preceding examples, further comprising:
[0124] -At each of the plurality of clients (80, 92, 92-1, 92-2, 92-3, 93, 94): in response to identifying at least one of the corresponding training data set or the artificial intelligence model in the corresponding second training state (83, 83-1, 83-2, 83-3) as an anomaly, suppressing the influence of the artificial intelligence model in the corresponding second training state (83, 83-1, 83-2, 83-3) on the aggregation of weights. \
[0125] Example 7. The computer-implemented method of any of the preceding examples, further comprising:
[0126] - Bypassing (3115) said evaluation of at least one group (70) comprising clients (80, 92, 92-1, 92-2, 92-3, 93, 94) marked as trustworthy.
[0127] Example 8. The computer-implemented method of any of the preceding examples,
[0128] Therein, groups (70) are non-overlapping.
[0129] Example 9. The computer-implemented method of any of the preceding examples,
[0130] Wherein, clients (80, 92, 92-1, 92-2, 92-3, 93, 94) are associated with a group (70) depending on one or more predefined criteria, the one or more predefined criteria being selected from: the total count of the groups (70); the number of clients in each group (70); the size of the training data set at each of the multiple clients (80, 92, 92-1, 92-2, 92-3, 93, 94).
[0131] Example 10. The computer-implemented method of any of the preceding examples, further comprising:
[0132] -When performing the evaluation on each of the artificial intelligence models in the third training state (85): aggregating (3050) the weights of the artificial intelligence models in those third training states (85) associated with the positive results of the evaluation to obtain the artificial intelligence model in the fourth training state, and deploying (3005) the artificial intelligence model in the fourth training state.
[0133] Example 11. The computer-implemented method of any of the preceding examples,
[0134] Among them, the artificial intelligence model performs one or more medical imaging processing tasks.
[0135] Example 12. A computer-implemented method for use in a client, comprising:
[0136] - obtaining (3005) an artificial intelligence model in a first training state from a central authority;
[0137] - performing (3010) reasoning based on the artificial intelligence model in the first training state;
[0138] - based on said execution of reasoning, obtaining (3015) a training data set;
[0139] - retraining (3020) the artificial intelligence model based on the training data set to obtain an artificial intelligence model in a second training state;
[0140] - providing the artificial intelligence model in the second training state to at least one of the central authority and the one or more further clients;
[0141] - establishing an artificial intelligence model in a plurality of third training states based on information obtained from at least one of the central authority or one or more additional clients;
[0142] -Evaluating the artificial intelligence model in each of the plurality of third training states based on a benchmark relative to the performance of the artificial intelligence model in the third training state and based on the training data set.
[0143] Example 13. The computer-implemented method of Example 12,
[0144] Wherein, the establishment of the artificial intelligence model in multiple third training states includes:
[0145] -Obtaining artificial intelligence models in multiple third training states from a central authority.
[0146] Example 14. The computer-implemented method of Example 12,
[0147] Wherein, the establishment of the artificial intelligence model in multiple third training states includes:
[0148] - obtaining artificial intelligence models in a plurality of second training states from a plurality of further clients,
[0149] - associating the client and a plurality of further clients with a plurality of groups,
[0150] -For each of the plurality of groups: aggregating weights of artificial intelligence models in a second training state associated with clients within the corresponding group to obtain an artificial intelligence model in a corresponding third training state.
[0151] Example 15. The computer-implemented method of Example 14,
[0152] Therein, the client and a plurality of further clients are associated with a plurality of groups based on grouping information obtained from a central authority.
[0153] Example 16. A computer-implemented method for use in a central authority, the method comprising:
[0154] - deploying the artificial intelligence model in the first training state to multiple clients;
[0155] - obtaining, from each of the plurality of clients, an artificial intelligence model in a corresponding second training state;
[0156] -Associate multiple clients with multiple groups;
[0157] - for each of the plurality of groups: aggregating weights of artificial intelligence models in a second training state associated with clients within the corresponding group to obtain an artificial intelligence model in a corresponding third training state; and
[0158] -Provide the artificial intelligence model in the third training state to multiple clients for evaluation.
[0159] Example 17. The computer-implemented method of Example 16, further comprising:
[0160] - Obtain evaluation results from multiple clients,
[0161] - aggregating the weights of the artificial intelligence model in the second training state and / or the third training state based on the evaluation result to obtain the artificial intelligence model in the fourth training state, and
[0162] -Deploy the artificial intelligence model in the fourth training state to multiple clients.
[0163] Example 18. A computer-implemented method for use in a central authority, the method comprising:
[0164] - deploying the artificial intelligence model in the first training state to multiple clients,
[0165] -Providing grouping information to multiple clients, the grouping information indicating association of the multiple clients with the multiple groups, for aggregation of weights of artificial intelligence models in corresponding second training states during a cross-evaluation process at the multiple clients.
[0166] Example 19. A system comprising a central authority and a plurality of clients, wherein the system is configured to perform the method of Example 1.
[0167] Example 20. A client comprising a processor and a memory, the processor being configured to load a program code from the memory and execute the program code, wherein execution of the program code causes the processor to perform the method of Example 12.
[0168] Example 21. A central server configured to implement a central authority for multiple clients, the central server comprising a processor and a memory, the processor configured to load a program code from the memory and execute the program code, wherein execution of the program code causes the processor to perform the method of Example 16 or Example 18.
[0169] Although the invention has been shown and described with respect to certain preferred embodiments, equivalents and modifications will occur to others skilled in the art upon the reading and understanding of the specification. The present invention includes all such equivalents and modifications, and is limited only by the scope of the appended claims.
Claims
1. A computer-implemented method comprising: deploying the artificial intelligence model in a first training state to a plurality of clients; At each of the plurality of clients, obtaining a corresponding training data set; At each of the plurality of clients, retraining the artificial intelligence model to obtain an artificial intelligence model in a second training state; associating the plurality of clients with a plurality of groups; for each of the plurality of groups, aggregating weights of the artificial intelligence models in the second training state associated with the clients in the corresponding group to obtain an artificial intelligence model in a third training state; as well as The artificial intelligence model is evaluated in each of the third training states.
2. The computer-implemented method of claim 1, in, The artificial intelligence model in each of the third training states is evaluated using a cross-evaluation process at the plurality of clients.
3. The computer-implemented method of claim 2, in, The cross-evaluation process includes a benchmark relative to the performance of the artificial intelligence model in the first training state, the benchmark being based on a corresponding training data set at a corresponding client.
4. The computer-implemented method of claim 2, in, The cross-evaluation process includes providing an artificial intelligence model in a corresponding third training state associated with a given one of the multiple groups to at least one other group of the multiple groups, and benchmarking the artificial intelligence model in the corresponding third training state associated with a given one of the multiple groups at each of the at least one other group.
5. The computer-implemented method of claim 1 , further comprising: For each of the plurality of groups, in response to the evaluation that the artificial intelligence model in the corresponding third training state does not meet the predefined benchmark, the clients previously associated with the corresponding group are reassociated with a plurality of newly formed groups that replace the corresponding group, and the aggregation of weights and evaluation of the newly formed groups are repeated.
6. The computer-implemented method of claim 1 , further comprising: At each of the plurality of clients, in response to identifying at least one of the corresponding training data set or the artificial intelligence model in the corresponding second training state as an anomaly, the influence of the artificial intelligence model in the corresponding second training state on the aggregation of weights is suppressed.
7. The computer-implemented method of claim 1 , further comprising: The evaluation is bypassed for at least one group including a client marked as trustworthy.
8. The computer-implemented method of claim 1, wherein: The groups are non-overlapping.
9. The computer-implemented method of claim 1, wherein: The client is associated with the group depending on one or more predefined criteria, the one or more predefined criteria comprising: the total count of the group; the number of clients per group; or A size of a training data set at each of the plurality of clients.
10. The computer-implemented method of claim 1, further comprising: When performing the evaluation on the artificial intelligence models in each of the third training states, the weights of the artificial intelligence models in those third training states associated with the positive results of the evaluation are aggregated to obtain the artificial intelligence model in the fourth training state.
11. The computer-implemented method of claim 10, further comprising deploying the artificial intelligence model in the fourth training state.
12. The computer-implemented method of claim 1, in, The artificial intelligence model performs one or more medical imaging processing tasks.
13. A computer-implemented method for use in a client, comprising: Obtaining an artificial intelligence model in a first training state from a central authority; performing reasoning based on the artificial intelligence model in the first training state; acquiring a training data set based on said performing of inference; Retrain the artificial intelligence model based on the training data set to obtain an artificial intelligence model in a second training state; providing the artificial intelligence model in the second training state to at least one of the central authority and one or more additional clients; establishing an artificial intelligence model in a plurality of third training states based on information obtained from at least one of the central authority or the one or more additional clients; as well as The artificial intelligence model in each of the plurality of third training states is evaluated based on a benchmark relative to the performance of the artificial intelligence model in the third training state and based on the training data set.
14. The computer-implemented method of claim 13, in, The establishing of the artificial intelligence model in the plurality of third training states includes obtaining the artificial intelligence model in the plurality of third training states from the central authority.
15. The computer-implemented method of claim 13, wherein: The establishing of the artificial intelligence model in the plurality of third training states includes: obtaining, from the plurality of additional clients, artificial intelligence models in a plurality of second training states; associating the client and the plurality of additional clients with a plurality of groups; and For each of the plurality of groups, weights of the artificial intelligence models in the second training state associated with the clients in the corresponding group are aggregated to obtain an artificial intelligence model in a corresponding third training state.
16. The computer-implemented method of claim 15, in, The client and the plurality of additional clients are associated with the plurality of groups based on grouping information obtained from the central authority.
17. A device comprising: Non-transitory computer readable media for storing program code; as well as One or more processor units in communication with the non-transitory computer readable medium, the one or more processor units operating with the program code to perform operations including: Obtaining an artificial intelligence model in a first training state from a central authority; performing reasoning based on the artificial intelligence model in the first training state; acquiring a training data set based on said performing of inference; Retrain the artificial intelligence model based on the training data set to obtain an artificial intelligence model in a second training state; providing the artificial intelligence model in the second training state to at least one of the central authority and one or more additional clients; establishing an artificial intelligence model in a plurality of third training states based on information obtained from at least one of the central authority or the one or more additional clients; as well as The artificial intelligence model in each of the plurality of third training states is evaluated based on a benchmark relative to the performance of the artificial intelligence model in the third training state and based on the training data set.
18. The apparatus according to claim 17, wherein: The establishing of the artificial intelligence model in the plurality of third training states includes obtaining the artificial intelligence model in the plurality of third training states from the central authority.
19. The apparatus according to claim 17, wherein: The establishing of the artificial intelligence model in the plurality of third training states includes: obtaining, from the plurality of additional clients, artificial intelligence models in a plurality of second training states; associating the client and the plurality of additional clients with a plurality of groups; and For each of the plurality of groups, weights of the artificial intelligence models in the second training state associated with the clients in the corresponding group are aggregated to obtain an artificial intelligence model in a corresponding third training state.
20. The device according to claim 19, in, The client and the plurality of additional clients are associated with the plurality of groups based on grouping information obtained from the central authority.