A federated incremental learning method, electronic device, and application product based on subspace aggregation
Patent Information
- Application Number
- CN202610911538.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-24
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2046-06-24
AI Technical Summary
[0004]在上述场景下,单纯采用传统联邦学习方法难以处理任务随时间连续到达的问题,单纯采用集中式增量学习方法又难以满足原始数据不能跨机构共享的要求
Smart Images

Figure CN122452819B_ABST
Abstract
Description
Technical Field
[0001] This disclosure belongs to the field of artificial intelligence algorithm technology, and specifically relates to a federated incremental learning method, electronic device and program product based on subspace aggregation. Background Technology
[0002] Federated Learning (FL) is a distributed learning technique that enables collaborative modeling across multiple clients. Its basic idea is to achieve joint training of a global model by exchanging model parameters, gradient information, or other intermediate representations, while each client retains its original data. Federated Learning is particularly suitable for applications with high requirements for data privacy and security, such as healthcare, finance, and industrial inspection. Especially in the field of medical artificial intelligence, there are often strict data compliance boundaries between different institutions. Traditional centralized training methods require the aggregation of raw data, resulting in high implementation costs, significant privacy risks, and difficulty in long-term stable deployment in multi-center scenarios.
[0003] Incremental learning, or continuous learning, is a technique that enables models to continuously learn new knowledge as new tasks are introduced. Its core principle is to allow the model to adapt to new tasks while maintaining the performance of previously learned tasks, avoiding catastrophic forgetting—the phenomenon where learning new tasks significantly degrades the performance of older tasks. In real-world applications, data distribution is typically not static but evolves over time. For example, in medical settings, new disease categories may emerge, imaging equipment may change, acquisition protocols may be adjusted, and data distribution may differ between different medical institutions.
[0004] In the aforementioned scenarios, traditional federated learning methods struggle to handle the issue of tasks arriving sequentially over time, while centralized incremental learning methods fail to meet the requirement that raw data cannot be shared across institutions. Therefore, Federated Incremental Learning (FIL), which balances data privacy protection with continuous task learning, is gradually becoming a technological trend. Summary of the Invention
[0005] One aspect of this disclosure is a federated incremental learning method based on subspace aggregation, suitable for joint training of incremental tasks that occur over time by multiple clients without sharing the original data, comprising the following steps:
[0006] The server initializes the global model and distributes it to each client.
[0007] Each client performs local training of the current incremental task based on local data and calculates the gradient of the model parameters;
[0008] Each client constructs a local principal gradient subspace based on the gradient, which compactly represents the main update direction of the current task in the parameter space in a low-rank form;
[0009] During subsequent incremental task training, each client performs orthogonal projection constraints on the current gradient based on the aggregated subspace issued by the server, in order to suppress parameter updates along historical key directions;
[0010] After completing the training for the current task, each client removes the existing historical direction components from the gradient information of the current task, extracts the newly added orthogonal directions, and incrementally updates the local principal gradient subspace. The updated local principal gradient subspace is then uploaded to the server.
[0011] The server receives the local main gradient subspace uploaded by each client, generates an aggregated subspace according to a preset aggregation strategy, and distributes the aggregated subspace to each client.
[0012] The server aggregates the local model parameters uploaded by each client to obtain new global model parameters.
[0013] Furthermore, the construction of the local principal gradient subspace includes: concatenating multiple sampled gradient vectors into a gradient matrix, performing singular value decomposition on the gradient matrix, and selecting the principal direction as the basis of the local principal gradient subspace according to the energy preservation threshold.
[0014] The orthogonal projection constraint includes: removing the projection component of the current gradient on the aggregated subspace to obtain the constrained gradient, and updating the local model parameters based on the constrained gradient.
[0015] The incremental update of the local principal gradient subspace includes: removing the projection components of the gradient matrix of the current task onto the historical subspace basis to perform deduplication; performing singular value decomposition on the deduplicated matrix and extracting new directions according to the cumulative energy threshold; concatenating the new directions with the original historical subspace basis and performing orthogonal normalization to obtain the updated local principal gradient subspace.
[0016] Furthermore, the preset aggregation strategy includes unified global aggregation and / or client-oriented adaptive aggregation.
[0017] The unified global aggregation includes: the server merging the local principal gradient subspaces uploaded by all clients to obtain a unified aggregation matrix, performing singular value decomposition on the unified aggregation matrix and selecting the principal direction according to the server-side energy threshold to generate a global aggregation subspace, and uniformly distributing the global aggregation subspace to all clients.
[0018] The client-oriented adaptive aggregation includes: the server calculating the geometric similarity between the local principal gradient subspaces of each client; selecting a set of neighboring clients whose local principal gradient subspaces are similar to those of each client; constructing an aggregation matrix based on the set of neighboring clients and the client's own local principal gradient subspace; performing singular value decomposition on the aggregation matrix and selecting the principal direction according to the server-side energy threshold to generate an aggregation subspace oriented towards that client; and then separately distributing the aggregation subspace oriented towards that client to the corresponding client.
[0019] Furthermore, the geometric similarity is measured using the Grassmann distance metric.
[0020] The model parameters are aggregated by weighting the number of samples from each client.
[0021] The client's local data exhibits non-independent identical distribution characteristics, and / or the incremental task is a medical image classification task.
[0022] One aspect of this disclosure is a federated incremental learning system based on subspace aggregation, comprising a server and multiple clients.
[0023] The server includes:
[0024] The global model management module is used to initialize the global model and distribute it to each client, as well as to receive local model parameters uploaded by each client and aggregate them to obtain new global model parameters.
[0025] The subspace aggregation module is used to receive the local main gradient subspace uploaded by each client, generate an aggregated subspace according to a preset aggregation strategy, and distribute the aggregated subspace to each client.
[0026] The client includes:
[0027] The local training module is used to perform local training of the current incremental task based on local data and to calculate the gradient of the model parameters;
[0028] The local subspace construction module is used to construct a local principal gradient subspace based on the gradient, which compactly represents the main update direction of the current task in the parameter space in a low-rank form;
[0029] The projection constraint module is used to perform orthogonal projection constraints on the current gradient based on the aggregated subspace issued by the server during subsequent incremental task training, so as to suppress parameter updates along historical key directions.
[0030] The subspace update module is used to remove existing historical direction components from the gradient information of the current task after completing the training of the current task, extract the newly added orthogonal directions and incrementally update the local principal gradient subspace, and upload the updated local principal gradient subspace to the server.
[0031] In one aspect of this disclosure, an electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, the processor running the computer program to implement the described subspace aggregation-based federated incremental learning method.
[0032] In one aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the described federated incremental learning method based on subspace aggregation.
[0033] In one aspect of this disclosure, a computer program product includes a computer program that is executed by a processor to implement the described subspace aggregation-based federated incremental learning method. Attached Figure Description
[0034] The above and other objects, features, and advantages of exemplary embodiments of the present invention will become readily apparent from the following detailed description taken in conjunction with the accompanying drawings. Several embodiments of the invention are illustrated in the drawings by way of example and not limitation, wherein:
[0035] Figure 1 A general flowchart of a federated incremental learning method according to one embodiment of this disclosure.
[0036] Figure 2. Schematic diagram of client-side local principal gradient subspace construction and update according to one of the embodiments of this disclosure.
[0037] Figure 3 is a schematic diagram of server-side dual-mode subspace aggregation according to one of the embodiments of this disclosure. Detailed Implementation
[0038] According to existing solutions, in long-term multi-center operation scenarios, clients not only face problems such as inconsistent task order, imbalanced class distribution, and unbalanced sample size, but are also affected by differences in device model, imaging protocol, sampling noise, and annotation standards. These factors lead to significant non-independent identically distributed (Non-IID) characteristics among clients, causing inconsistent or even conflicting model update directions in the same round of federated aggregation, thereby increasing the risk of historical knowledge loss and underfitting to new tasks. Several solutions have been proposed to address this problem, including:
[0039] 1. Experience replay-based methods
[0040] This method preserves historical samples, category prototypes, or generates pseudo-samples, reusing old task information during subsequent task training to mitigate catastrophic forgetting. Typical techniques include memory buffer-based replay, prototype-based replay, and generative replay. The drawback of this method is:
[0041] (1) It is necessary to store old task-related samples or data representations for a long time, which poses a risk of data leakage in privacy-sensitive scenarios;
[0042] (2) Even if only a small number of historical samples or compressed feature representations are preserved, it will result in high compliance review costs and obstacles to cross-agency deployment;
[0043] (3) The capacity management and sampling strategy of the memory bank have a significant impact on performance and require additional optimization;
[0044] (4) In a federal scenario, uploading samples or prototypes to a server further amplifies privacy risks.
[0045] 2. Methods based on parameter regularization
[0046] This method imposes additional constraints on important parameters during model optimization to preserve knowledge from previous tasks as much as possible. Examples include elastic constraint methods based on parameter importance weights (such as Elastic Weight Consolidation) and soft constraint methods based on knowledge distillation. The drawback of this method is:
[0047] (1) As tasks continue to arrive, regularization constraints are constantly added, and the model’s ability to adapt to new tasks often gradually decreases, i.e., it lacks plasticity.
[0048] (2) When there are significant differences between the old and new tasks or when the data distribution changes significantly, relying solely on parameter regularization is often insufficient to effectively combat forgetting;
[0049] (3) It is difficult to simultaneously maintain historical knowledge and learn new tasks.
[0050] 3. Parameter Isolation-Based Methods
[0051] This method allocates relatively independent parameter spaces or network structures for different tasks, such as task-specific sub-networks, dynamically scalable networks, and masking mechanisms. The drawback of this method is:
[0052] (1) The model parameter size continues to increase with the number of tasks, resulting in higher storage, computing and maintenance costs;
[0053] (2) It is difficult to meet the requirements of model compactness and running efficiency in long-term continuous learning scenarios;
[0054] (3) In federated scenarios with a large number of clients and a continuous increase in tasks, the pressure of communication and synchronization increases significantly.
[0055] 4. Gradient projection-based method
[0056] This method, represented by Gradient Projection Memory (GPM), or subspace constraint methods, typically mitigates catastrophic forgetting by extracting the principal direction subspace of historical tasks in the parameter space and imposing orthogonal projection constraints on the current gradient during subsequent task training. The core idea of this class of methods is to represent important update directions of historical tasks in a low-rank subspace and preserve existing task knowledge by restricting parameter updates in those directions during subsequent tasks. The drawback of this method is:
[0057] (1) This type of method is mainly aimed at single-machine continuous learning scenarios, and does not consider the non-independent and identically distributed characteristics between different clients in a federated environment, making it difficult to directly apply to multi-client collaborative training scenarios.
[0058] (2) This type of method usually only uses the local subspace information of a single model in historical tasks, lacks a mechanism for unified or targeted aggregation of knowledge from multiple clients on the server side, and is difficult to make full use of cross-client shared knowledge;
[0059] (3) Such methods usually do not establish a closed-loop collaborative mechanism between client-side local subspace updates, server-side aggregated subspace generation and distribution, and subsequent task projection constraints. Therefore, it is difficult to simultaneously maintain historical knowledge and adapt to new tasks in heterogeneous federated incremental learning scenarios.
[0060] In this disclosure, "multi-center" can refer to multiple independent data holding institutions or medical centers (such as different hospitals, different medical institutions, and different data centers). "Multiple clients" can refer to distributed nodes participating in collaborative training within a federated system. In a multi-center healthcare scenario, each center typically corresponds to one client within the federated architecture.
[0061] In summary, the existing technology has the following shortcomings:
[0062] First, experience-based replay methods require the retention of historical samples, category prototypes, or pseudo-samples, which poses a risk of privacy leakage and increases the storage and management burden, making them particularly difficult to apply to highly compliant scenarios such as healthcare.
[0063] Secondly, parameter regularization-based methods weaken the model's ability to adapt to new tasks and are difficult to effectively resist forgetting when there are large differences in tasks.
[0064] Third, methods based on parameter isolation, structural expansion, or task-specific modules can cause the model parameter size to continuously expand as the number of tasks increases, making it difficult to meet the efficiency requirements in long-term operation.
[0065] Fourth, while gradient projection-based methods can mitigate catastrophic forgetting in single-machine continuous learning to some extent, they typically lack systematic design for client heterogeneity, server-side subspace aggregation, and targeted knowledge sharing in federated scenarios, making them difficult to directly apply to heterogeneous federated incremental learning scenarios.
[0066] Therefore, in view of the aforementioned deficiencies in the prior art, the technical problems to be solved by this disclosure include at least the following:
[0067] 1. How to effectively preserve historical task knowledge in a federated scenario and suppress catastrophic forgetting during incremental learning, without sharing original historical samples or relying on experience replay?
[0068] 2. In a heterogeneous federated environment where client data exhibits non-independent identically distributed (Non-IID) characteristics, how can we avoid over-averaging and irrelevant interference caused by simply aggregating all client knowledge in a unified manner, and improve the relevance and effectiveness of the aggregation results for subsequent task training of each client?
[0069] 3. In long-term operation scenarios with a continuously increasing number of tasks, how to avoid the expansion of storage, communication and computing overhead caused by the continuous expansion of model structure or the continuous accumulation of sample library, so as to make the representation of historical knowledge compact enough and continuously incrementally updated.
[0070] 4. How to establish a closed-loop collaborative mechanism between the local continuous learning process on the client side and the cross-client knowledge sharing process on the server side, so as to form an adjustable collaborative relationship between the maintenance of historical knowledge (stability) and the adaptation to new tasks (plasticity).
[0071] 5. How to provide a flexibly configurable general framework that can achieve a balance between knowledge retention capability, system overhead, and federated training efficiency through parameterized adjustment under different heterogeneity intensities, communication resource conditions, and continuous learning task lengths.
[0072] According to one or more embodiments, a federated incremental learning method based on subspace aggregation establishes a unified training framework that coordinates client-side knowledge preservation and server-side subspace aggregation guidance for federated incremental learning scenarios. Specifically, the method includes:
[0073] On the client side, a compact representation of historical task knowledge is constructed by a local principal gradient subspace, and the local principal gradient subspace is uploaded to the server.
[0074] The server receives the local main gradient subspace uploaded by each client and generates an aggregated subspace through a preset aggregation strategy.
[0075] In subsequent task training, the client implements orthogonal projection constraints based on the aggregated subspace issued by the server, so as to maintain existing task knowledge without retaining the original historical samples and reduce irrelevant interference caused by data heterogeneity.
[0076] The preset aggregation strategy includes at least one of the following: First, uniformly aggregate the local principal gradient subspaces uploaded by all clients to generate a global aggregated subspace, and then distribute the global aggregated subspace to all clients. Second, for each client, calculate a set of nearest neighbor clients similar to its local principal gradient subspace, and generate an aggregated subspace oriented towards that client based on the set of nearest neighbor clients and the client's own local principal gradient subspace, and then distribute the aggregated subspace oriented towards that client to the corresponding client.
[0077] Through the above method, this embodiment of the disclosure achieves the synergy of preserving historical knowledge, adapting to new tasks, and sharing knowledge across clients without transmitting the original training samples.
[0078] The symbol parameters involved in this disclosure are defined as follows:
[0079]
[0080] According to one or more embodiments, a federated incremental learning system based on subspace aggregation is disclosed. This system is suitable for multiple clients to jointly train on new tasks that continuously arrive over time without sharing the original data, thereby mitigating the catastrophic forgetting problem in incremental learning and improving the adaptability to new tasks under heterogeneous data conditions. Embodiments of this disclosure consider both knowledge retention during continuous task learning on the client side and heterogeneity control during cross-client knowledge sharing on the server side.
[0081] Assume the federated learning system includes a server and There are 1 client, and the client set is denoted as . ,in, Indicates the first One client. Assume the entire incremental learning process includes... A series of consecutive tasks, denoted as [task sequence]. ,in Indicates the task index. (Number) The client in the first The local datasets for each task are denoted as follows: ,in Indicates the first One sample, Indicates the corresponding label, This indicates that the client is in the... The number of samples on each task. Let... Indicates the index of the neural network layer. This represents the number of gradient samples used to construct the local principal gradient subspace.
[0082] The steps of federated incremental learning in this publicly disclosed system include: within a complete task cycle, the client side sequentially performs local training, gradient sampling, and construction or update of the local principal gradient subspace; the server side performs subspace aggregation and global model update. The subspace aggregation includes two implementation methods: unified global aggregation and client-oriented adaptive aggregation. These steps are executed cyclically in task order until all incremental task training is completed.
[0083] The server-side aggregation step can be implemented in two ways: The first is unified global aggregation, where the server receives the local principal gradient subspaces uploaded by all clients after each task, performs unified aggregation on the local principal gradient subspaces to generate a global aggregated subspace, and then distributes the global aggregated subspace to all clients; The second is client-oriented adaptive aggregation, where the server receives the local principal gradient subspaces uploaded by all clients after each task, calculates the set of client subspaces closest to its local principal gradient subspace for each client, and uses the client's own local principal gradient subspace and the set of closest client subspaces together for aggregation to generate an aggregated subspace for that client, and then distributes the aggregated subspace for that client to the corresponding client.
[0084] Specifically, a federated incremental learning method based on subspace aggregation includes the following steps:
[0085] S1, the server initializes the global model and distributes it to each client.
[0086] Server initializes global model parameters In the When a task begins, the server sends the current global model parameters to each client participating in the training, and each client uses these parameters as the local model initialization parameters.
[0087] The purpose of step S1 is to provide a consistent initial model state for all participating clients within the same task cycle, so that subsequent local gradient sampling, local subspace construction, and server aggregation are all based on the same global parameter starting point, thereby improving the comparability and aggregability of information uploaded by different clients.
[0088] S2, the client performs local training and calculates gradients.
[0089] For the Any client in the task During local training, targeting the network's first... Layer parameters The gradient is calculated based on the current mini-batch of samples to characterize the update direction of this layer on the current task:
[0090] (1)
[0091] in, Represents the loss function. Indicates the client In the The current small batch of data on each task Indicates the client's local model number. Layer parameters. This gradient calculation provides a foundation for subsequent construction of the local principal gradient subspace and orthogonal projection constraints.
[0092] In practice, the gradient samples used to construct the gradient matrix can come from multiple mini-batch samples in the current task training set, or from representative samples from several iterations at the end of training. By combining multiple gradient vectors, the main changing trends of the current task in the parameter space of this layer can be represented more stably.
[0093] S3, construct the client-side local principal gradient subspace.
[0094] After each task training is completed, each client uses the data sampled during the training process of that task. The gradient vectors are used to construct the gradient matrix, which is used to extract the principal gradient directions for the task. For the client... The The layer has the following gradient matrix:
[0095] (2)
[0096] when At this time, singular value decomposition is performed on the gradient matrix to extract the main direction of gradient change for that layer:
[0097] (3)
[0098] in , and These are the left singular vector matrix, singular value matrix, and right singular vector matrix of the gradient matrix, respectively.
[0099] Select the smallest integer , so that:
[0100] (4)
[0101] in, Indicates the first The energy of the layer maintains a threshold. Denotes the Frobenius norm. Represents a matrix of The order is truncated approximation. Then, the preceding... The left singular vectors are used as the basis of the local principal gradient subspace for the first task, denoted as . This leads to the client. In the Local principal gradient subspace of the layer This process allows historical task knowledge to be compressed and represented as a low-rank local principal gradient subspace.
[0102] By setting an energy retention threshold, a balance can be struck between the ability to retain historical knowledge and the dimensionality of the subspace. A higher threshold results in more retained main directions and more comprehensive knowledge preservation; a lower threshold results in a more compact subspace dimension and lower communication and storage costs. Therefore, this invention allows for flexible adjustment of knowledge representation accuracy and system overhead based on specific application scenarios.
[0103] S4, orthogonal projection constraints are executed in subsequent tasks.
[0104] when At the same time, in order to maintain knowledge of historical tasks, the training is performed locally on the client side. Before each task, first receive and read the corresponding layer aggregated subspace basis matrix sent by the server, and denot it uniformly as follows. For the current gradient The gradient after constraint is obtained by removing the projected component of the gradient in the aggregate subspace distributed by the server through orthogonal projection constraints.
[0105] (5)
[0106] Then, the local model parameters are updated using the constrained gradient:
[0107] (6)
[0108] in, This represents the learning rate. Through the orthogonal projection constraint described above, parameter updates along the historical critical gradient direction can be suppressed during the training of new tasks, thereby reducing the disruption of existing task knowledge.
[0109] In this step, the client does not completely freeze parameter updates, but only suppresses the projection component of the current gradient in the direction of the aggregated subspace sent by the server, thus preserving the update degrees of freedom in the other directions for the new task. Since the aggregated subspace is obtained by the server based on the local principal gradient subspaces uploaded by each client in the historical task stages, this constraint can both introduce the effect of preserving historical knowledge and not completely restrict the learning of the new task.
[0110] S5, update the client's main gradient subspace.
[0111] Complete the first step on the client. After training on each task, the gradient matrix for the current task is reconstructed. To avoid repeatedly adding directions that already exist in the historical subspace to the new local principal gradient subspace, the gradient matrix must first be deduplicated, i.e., its projection components onto the basis matrix of the historical subspace must be removed.
[0112] (7)
[0113] Historical projection components are denoted as For the deduplicated matrix Perform singular value decomposition:
[0114] (8)
[0115] in, , and These are the left singular vector matrix, singular value matrix, and right singular vector matrix of the deduplicated matrix, respectively. Select the smallest integer... , so that:
[0116] (9)
[0117] in, Indicates the first The cumulative energy threshold of the layer, Represents the residual gradient matrix after deduplication. of The order truncation approximation is then applied. Subsequently, the newly extracted orthogonal directions are concatenated with the original history subspace basis matrix and orthogonally normalized to obtain the updated history subspace basis matrix:
[0118] (10)
[0119] in, This indicates that an orthogonal normalization operation is performed on the input matrix. This step enables incremental expansion of the local principal gradient subspace and avoids redundant storage of existing directions.
[0120] After updating the local subspace, the client can send the updated subspace basis matrix, corresponding singular values, and necessary layer index information to the server. Compared to directly uploading a large number of historical samples or complete training trajectories, this knowledge representation is more compact, which helps reduce communication overhead and long-term storage burden in federated scenarios.
[0121] S6, the server performs subspace aggregation.
[0122] After receiving the local principal gradient subspace information uploaded by each client, the server can generate an aggregated subspace according to a preset aggregation strategy. The aggregation strategy includes two implementation methods: unified global aggregation and client-oriented adaptive aggregation. Specifically, it includes...
[0123] S601, the server performs unified global aggregation.
[0124] In the unified global aggregation implementation, the server receives all local principal gradient subspace information uploaded by all clients after each task is completed, and performs unified aggregation on the local principal gradient subspace. For the first... Let the basis matrices of the local principal gradient subspaces uploaded by all clients be respectively... The corresponding singular value matrices are respectively The server constructs a unified aggregation matrix based on the local principal gradient subspace of all client layers:
[0125] (11)
[0126] Subsequently, the server performs singular value decomposition on the unified aggregation matrix:
[0127] (12)
[0128] in, , and These are the left singular vector matrix, singular value matrix, and right singular vector matrix of the unified aggregation matrix, respectively.
[0129] Further based on server-side energy thresholds Select the main direction to form the first The layer global aggregation subspace, denoted as This information is then distributed to all clients, ensuring that each client performs orthogonal projection constraints based on this subspace during the next training task. In this invention, the historical aggregated subspace basis matrix used for orthogonal projection constraints in step S4 is uniformly denoted as... When step S6 adopts a unified global aggregation method, each client receives the same aggregation subspace, therefore... .
[0130] In this implementation, the server uniformly aggregates the local principal gradient subspaces uploaded by all clients, enabling the generated global aggregated subspace to comprehensively represent the shared key gradient directions of multiple clients in historical tasks. This global aggregated subspace is then distributed to all clients as consistent constraint information. Therefore, this implementation is more conducive to maintaining historical knowledge globally, mitigating global forgetting, and improving the overall stability of the federated incremental learning process. This implementation is suitable for application scenarios where client distribution differences are relatively small, task structures are relatively consistent, or where global consistency constraints are emphasized.
[0131] S602, the server performs client-oriented adaptive subspace aggregation.
[0132] After completing training for the current task, each client sends its updated local principal gradient subspace information to the server. The server then calculates the subspace similarity between different clients. For any two clients... and The The Grassmann distance of the local principal gradient subspace of a layer is defined as:
[0133] (13)
[0134] in, This represents the set of principal angles between two subspaces. The server uses the distance metric described above to measure the gradient geometric similarity between clients. For each client... The server selects the one with the smallest distance. Each client is considered as its nearest neighbor set, denoted as [a_n] . ,in This indicates the number of neighboring clients.
[0135] Subsequently, instead of performing a unified aggregation of all client-side local subspaces, the server performs subspace aggregation separately for each client. For each client... The Layers, servers are based on their nearest neighbor set Construct the corresponding aggregation matrix:
[0136] (14)
[0137] Then, singular value decomposition is performed on the aggregate matrix:
[0138] (15)
[0139] in, , and For the client respectively The constructed aggregate matrix includes the left singular vector matrix, singular value matrix, and right singular vector matrix.
[0140] Further based on server-side energy thresholds Select the main direction to form the client-oriented first... Layered aggregate subspace, denoted as The data is then distributed to the corresponding clients, enabling each client to perform orthogonal projection constraints based on its received subspace during the next training task. In this invention, the historical aggregated subspace basis matrix used for orthogonal projection constraints in step S4 is uniformly denoted as... When step S6 adopts a client-oriented adaptive aggregation method, different clients receive their respective corresponding aggregation subspaces, therefore there are .
[0141] In this implementation, the server does not uniformly aggregate all client local principal gradient subspaces. Instead, it selects only client subspaces that are more similar to the gradient geometry of the target client to participate in the aggregation. This makes the generated aggregated subspace more closely match the current data distribution characteristics and knowledge needs of the target client. Therefore, this implementation reduces interference from irrelevant client information, which is more conducive to the target client's learning and adaptation to subsequent new tasks, and improves the plasticity in the continuous learning process. This implementation is particularly suitable for application scenarios with strong client heterogeneity and significant differences in task evolution paths among different clients.
[0142] S7, the server performs model parameter aggregation.
[0143] After completing local training for the current communication round, the server also performs weighted aggregation of the model parameters uploaded by each client to obtain new global model parameters. This represents the communication round. The weighted aggregation formula for the model parameters is:
[0144] (16)
[0145] in, Indicates the client In the The local model parameters uploaded in the round.
[0146] Through the steps S1-S7 above, this disclosure introduces task-related local principal gradient subspace representations under the federated learning framework, and combines orthogonal projection constraints, local incremental updates and server-side adaptive subspace aggregation mechanisms to achieve the preservation of historical task knowledge and the effective absorption of new task knowledge, thereby improving the performance of federated incremental learning in scenarios with heterogeneous data and continuous task arrival.
[0147] Figure 1 shows the overall flowchart of the federated incremental learning method of this embodiment, illustrating the subspace construction, updating, aggregation, and client-server interaction processes, corresponding to steps S1-S7.
[0148] In A1 (corresponding to S1), the server initializes the global model parameters and distributes them to each client, providing a consistent local training starting point for all participating nodes.
[0149] In A2 (corresponding to S2), each client performs training based on local private data, calculates the gradient of each network layer for the current task, and represents the update direction of the task in the parameter space.
[0150] In A3 (corresponding to S3), the client combines multiple sampled gradients into a gradient matrix, extracts the main update direction by decomposition, constructs a low-rank local principal gradient subspace, and compresses historical task knowledge into a compact directional representation.
[0151] A4 (corresponding to S4) involves the client reading the aggregated subspace issued by the server before training for subsequent new tasks, performing orthogonal projection constraints on the current local gradient, and removing update components along historical key directions to suppress the destruction of existing task knowledge.
[0152] In A5 (corresponding to S5), after the client completes the training of the current task, it removes the duplicate directions that already exist in the historical subspace in the gradient matrix, extracts the newly added orthogonal directions and splices them with the historical subspace to update, forming an expanded local principal gradient subspace, and then uploads it to the server.
[0153] In A6 (corresponding to S6), after receiving the subspaces from each client, the server executes one of two aggregation methods and distributes the data. Unified global aggregation merges all client subspaces into a consistent global aggregation subspace, which is then distributed uniformly to all clients. Client-oriented adaptive aggregation, on the other hand, selects nearby neighbor subspaces for targeted aggregation based on subspace geometric similarity, generating a more specific dedicated aggregation subspace, which is then distributed individually.
[0154] A7 (corresponding to S7) is a server that aggregates the local model parameters uploaded by each client by weighted by sample size to generate new global model parameters for use in the next round of communication or the next task cycle.
[0155] Steps A1 to A7 are executed cyclically in task order, forming a closed loop from local training, subspace extraction, server aggregation, constraint feedback, and continuous learning. This achieves cross-client historical knowledge preservation and adaptation to new tasks without sharing the original data.
[0156] Figure 2 shows a schematic diagram of the construction and update of the local principal gradient subspace on the client side, illustrating the complete process from gradient sampling, gradient matrix construction, SVD decomposition, deduplication to subspace splicing and orthogonal normalization, corresponding to steps S3 and S5 respectively.
[0157] In B1 (corresponding to S3), the client samples multiple gradient vectors during task training and concatenates these gradients into a gradient matrix to stably represent the main changing trend of the current task in the parameter space.
[0158] B2 (corresponding to S3) decomposes the gradient matrix and selects the main directions according to the preset energy preservation threshold to form the local principal gradient subspace of the task, realizing a low-rank compact representation of historical task knowledge.
[0159] In B3 (corresponding to S5), in subsequent new tasks, the client reconstructs the gradient matrix and removes the directional components that already exist in the historical subspace to avoid duplicate storage of existing knowledge directions.
[0160] B4 (corresponding to S5) decomposes the deduplicated residual matrix and extracts the newly added orthogonal directions according to the cumulative energy threshold to represent the new knowledge content brought about by the current task.
[0161] B5 (corresponding to S5) concatenates the new direction with the original historical subspace basis, and after orthogonal normalization, forms the updated local principal gradient subspace, completing the incremental expansion of the subspace, and uploading the updated subspace information to the server.
[0162] Figure 3 shows a schematic diagram of server-side dual-mode subspace aggregation, illustrating the processing flow of unified global aggregation and client-oriented adaptive aggregation, corresponding to steps S601 and S602 respectively.
[0163] C1, the server receives all local principal gradient subspace information uploaded by the clients.
[0164] In C2, under the unified global aggregation mode, the server merges the subspaces of all clients into a single global aggregation subspace and distributes it uniformly to all clients, enabling each client to perform orthogonal projection based on consistent historical knowledge constraints during subsequent training.
[0165] In C3, under the client-oriented adaptive aggregation mode, the server calculates the geometric similarity between the subspaces of each client and selects a set of neighboring clients with similar structures for each client.
[0166] C4: The server performs targeted aggregation based on each client's own subspace and its neighbor set, generating a dedicated aggregation subspace for that client, and distributes it separately to the corresponding client, so that different clients can learn based on targeted constraints that fit their own data distribution in subsequent training.
[0167] The embodiments disclosed herein have at least the following beneficial effects:
[0168] (1) Privacy-friendly, zero original sample sharing
[0169] This disclosure employs a local principal gradient subspace to perform low-rank compressed representation of historical task knowledge on the client side, and uploads the local principal gradient subspace to the server. The server generates an aggregated subspace based on the local principal gradient subspaces uploaded by each client, and performs constraints based on the aggregated subspace during subsequent task training. The knowledge carrier exchanged between the client and the server consists only of the subspace basis matrix and singular values, without involving the original samples or prototypes. Since the subspace basis matrix is a low-rank directional representation obtained by truncating the gradient matrix using SVD, it does not contain original image pixel or label information. Therefore, it is not necessary to transmit original training samples during cross-institutional transmission, and the server does not need to centrally store historical data. Compared with experience replay methods, this disclosure is more conducive to meeting the compliance requirements of long-term collaborative learning in privacy-sensitive scenarios such as healthcare.
[0170] (2) Effectively resists catastrophic forgetting and has high model stability
[0171] This disclosure combines the unified global aggregation method in step S601 with the orthogonal projection constraint in step S4. In step S601, the server performs unified aggregation on the local principal gradient subspaces uploaded by all clients to generate a global aggregated subspace. This global aggregated subspace comprehensively represents the key gradient directions shared by multiple clients during the historical task phase. Subsequently, in step S4, each client applies orthogonal projection constraints to its current gradient based on the global aggregated subspace issued by the server, thereby suppressing excessive updates of parameters along the global historical key directions. Since this constraint acts on the historical knowledge structure shared by all clients, it is more conducive to mitigating catastrophic forgetting from a global perspective and improving the overall stability of the federated incremental learning process. Compared with methods that rely solely on local constraints of a single client, this method more fully preserves the shared historical knowledge across clients and is more suitable for application scenarios that emphasize global consistency and the preservation of global historical knowledge.
[0172] (3) It takes into account both unified aggregation and targeted aggregation, and adapts to multiple scenarios.
[0173] This disclosure provides a dual-mode subspace aggregation strategy in steps S601 and S602. When the differences in client data distribution are relatively small, or when the application scenario emphasizes global consistency constraints, the unified global aggregation method in step S601 can be adopted. This method aggregates all client subspaces to form a consistent global aggregated subspace, which is more conducive to improving the ability to retain historical knowledge and overall stability. When the client heterogeneity is strong, or when the application scenario emphasizes the target client's adaptability to subsequent new tasks, the client-oriented adaptive aggregation method in step S602 can be adopted. This method only aggregates client subspaces that are closer to the target client, which can reduce the interference from irrelevant client information and make the generated aggregated subspace more in line with the data distribution characteristics and knowledge needs of the target client, thereby being more conducive to improving the plasticity in the learning process of new tasks.
[0174] Therefore, this disclosure does not adopt a single fixed aggregation method, but rather enables the system to flexibly adjust between stability and flexibility by configuring two modes: unified global aggregation and client-oriented adaptive aggregation, thereby adapting to federated incremental learning scenarios with different heterogeneity intensities and different task evolution conditions.
[0175] (4) Coordination of stability and plasticity
[0176] This disclosure presents a closed-loop training framework consisting of continuous updating of the local principal gradient subspace on the client side, generation and distribution of the aggregated subspace on the server side, and step S4, which executes orthogonal projection constraints based on the server-distributed aggregated subspace. After the client completes the current task, it extracts and uploads the local principal gradient subspace. The server then aggregates and distributes the aggregated subspace based on the local principal gradient subspaces uploaded by each client. In subsequent task training, the client executes orthogonal projection constraints based on the aggregated subspace, thus forming a closed-loop collaborative mechanism of "local knowledge extraction—server aggregation—constraint feedback—next task learning." In this closed loop, if the unified global aggregation method in step S601 is adopted, the aggregated subspace distributed by the server contains more comprehensive globally shared historical knowledge, thus improving overall stability. If the client-oriented adaptive aggregation method in step S602 is adopted, the aggregated subspace distributed by the server better fits the gradient geometry of the target client, thus improving adaptability and flexibility for subsequent new tasks. Therefore, this disclosure does not maintain historical knowledge by simply increasing the strength of constraints, nor does it rely solely on weak constraints to pursue the ability to learn new tasks. Instead, it establishes an adjustable and coordinated relationship between maintaining historical knowledge and adapting to new tasks through different choices of server-side aggregation modes.
[0177] (5) Low communication and storage overhead, easy to expand to multi-task scenarios
[0178] This disclosure uses the low-rank subspace of the gradient matrix as the knowledge carrier. Compared to directly retaining a large number of historical samples, a complete gradient history, or a large-scale additional network structure, the storage size of the low-rank subspace basis matrix is controlled by the energy preservation threshold and the cumulative energy threshold, and is typically much smaller than the number of parameters in the original model; therefore, it has the advantages of compact representation, low storage overhead, low communication burden, and easy extension to multi-task scenarios.
[0179] (6) Applicability and Engineering Value
[0180] The subspace construction, projection, update, and aggregation mechanism proposed in this disclosure is independent of specific task types. Therefore, it is applicable not only to medical image classification tasks, but also to medical image segmentation, multi-center collaborative diagnostic model updates, continuous updates of industrial quality inspection models, and other deep learning scenarios that need to simultaneously consider data privacy, continuous task evolution, and client heterogeneity. It has good versatility and engineering application value.
[0181] According to one or more embodiments, a federated incremental learning method based on subspace aggregation considers the implementation steps of adopting a unified global aggregation method in a scenario with heterogeneous distribution of medical image categories, including performing the aforementioned steps S1-S7, wherein step S6 adopts the unified global aggregation method described in step S601.
[0182] Here, the MedMNIST medical image classification dataset is used to construct a federated incremental learning task. To verify the applicability under heterogeneous class distributions, two data partitioning methods, distribution-based and quantity-based, can be used respectively.
[0183] In the distribution-based heterogeneous category distribution scenario, the PathMNIST dataset is used. PathMNIST is derived from colorectal cancer histopathological image data, containing 107,180 images covering nine tissue types. During federated partitioning, a Dirichlet distribution is used to construct the category distribution differences between different clients. In this embodiment, there are 10 clients, with the heterogeneity parameter set to 0.3. During incremental task partitioning, the categories are constructed into four consecutive tasks in the order [2,2,2,3].
[0184] In a heterogeneous scenario with quantity-based category distribution, the OrganAMNIST dataset was used. OrganAMNIST is derived from abdominal CT images and contains 58,850 images covering 11 organ categories. During federated partitioning, each client holds only a subset of categories to simulate the inconsistent number of visible categories across different medical institutions; during incremental task partitioning, the categories are constructed into four consecutive tasks in the order [3,3,3,2].
[0185] In one embodiment, ResNet-18 is used as the local backbone network for each client, and the output dimension of the classification head corresponds to the current cumulative number of classes. The optimizer uses SGD, with momentum set to 0.9 and an initial learning rate of 0.01. The local batch size for each client is set to 64, with 3 epochs of local training per communication round, and the number of communication rounds between the server and client is set to 20. The gradient sampling number used to construct the local principal gradient subspace is specified. The value is set to 256, meaning that at the end of training for each task, a gradient matrix is constructed by sampling a total of 256 gradient vectors from multiple mini-batches during the current task's training process. This is the layer energy preservation threshold during single-task subspace extraction. Set to 0.95, the cumulative energy threshold during incremental updates of the local principal gradient subspace. Set to 0.97. During unified global aggregation, the server concatenates the local principal gradient subspaces uploaded by each client column-wise, performs singular value decomposition, and then applies the energy threshold from the server side. A principal direction of 0.99 is selected to form a global aggregate subspace.
[0186] During training, the server initializes and distributes the global model according to step S1; each client completes local training, gradient sampling, local principal gradient subspace construction, orthogonal projection constraints, and subspace updates according to steps S2-S5; the server performs unified global aggregation of the local principal gradient subspaces uploaded by all clients according to step S601, and completes model parameter aggregation according to step S7. The above process is repeated in the order of incremental tasks until all tasks are trained.
[0187] This embodiment is used to verify that, in cases where the distribution of medical image categories is uneven or different clients hold inconsistent numbers of categories, this embodiment of the disclosure supports the preservation of historical knowledge during the federated incremental learning process by forming consistent historical knowledge constraints through unified global aggregation without sharing the original image data.
[0188] Table 1. Experimental results of the unified global aggregation method in a heterogeneous medical image category distribution scenario in MedMNIST.
[0189]
[0190] Table 1 shows the experimental results of the unified global aggregation method in the heterogeneous category distribution scenario of MedMNIST. As can be seen from Table 1, in the heterogeneous category distribution scenario of PathMNIST, the FedSubmerg-G method (unified global aggregation method) corresponding to the present disclosure embodiment achieves an ACC of 87.25 and a BWT of -12.39, which is better than the listed comparison methods. In the quantity-based heterogeneous category distribution scenario of OrganAMNIST, FedSubmerg-G achieves an ACC of 73.65 and a BWT of -8.51, which is also better than the listed comparison methods. The above results show that when there is an imbalance in the category ratio or an inconsistent number of visible categories among clients, unified global subspace aggregation can form a more effective historical knowledge constraint, improving the overall classification performance while reducing forgetting during incremental learning.
[0191] Therefore, this embodiment demonstrates that in scenarios with heterogeneous distribution of medical image categories, without uploading original image data or historical samples, it is possible to maintain historical task knowledge and continuously learn new tasks through the construction of local principal gradient subspaces, unified global subspace aggregation, and orthogonal projection constraints.
[0192] According to one or more embodiments, a federated incremental learning method based on subspace aggregation is implemented in a client-oriented adaptive aggregation method in a multi-center medical image feature distribution heterogeneous scenario, and is executed according to the aforementioned steps S1-S7, wherein step S6 adopts the client-oriented adaptive aggregation method described in step S602.
[0193] In one embodiment, a heterogeneous scenario of real feature distribution is constructed using skin cancer identification data. Specifically, four skin cancer image data sources are selected: HAM, D7P, BCN20000, and DMF. These four data sources come from different real medical institutions or data centers. Among them, HAM contains 9,041 cases, D7P contains 1,926 cases, BCN20000 contains 12,148 cases, and DMF contains 1,065 cases. Due to differences in acquisition equipment, image color distribution, imaging quality, lesion morphology, case composition, and sample size among different medical institutions, the above data sources naturally form a heterogeneous real feature distribution in a multi-center medical image scenario. The incremental task is divided into [2,2,2].
[0194] ResNet-18 is used as the local backbone network for each client, with the classification head output dimension corresponding to the current cumulative number of classes. The optimizer uses SGD with a learning rate of 0.01. The local batch size for each client is set to 32, with 5 epochs of local training per communication round, and 20 communication rounds between the server and client. The gradient sampling number used to construct the local principal gradient subspace is specified. Set to 256. Layer energy preservation threshold during single-task subspace extraction. Set to 0.95, the cumulative energy threshold during incremental updates of the local principal gradient subspace. The value is set to 0.97. During the client-oriented adaptive subspace aggregation process, the server uses Grassmann distance to measure the geometric similarity between the local principal gradient subspaces of different clients, and selects the two clients with the smallest distance as nearest neighbors for each client; that is, the number of nearest neighbors is set to 2. Subsequently, the server constructs an aggregation matrix based on the target client's own local principal gradient subspace and the corresponding subspaces of its nearest neighbors, and then adjusts the matrix according to the server-side energy threshold. Select the main direction as 0.98 to generate the aggregate subspace oriented towards this client.
[0195] During training, the server initializes and distributes the global model according to step S1; each client completes local training, gradient sampling, local principal gradient subspace construction, orthogonal projection constraints, and subspace updates according to steps S2-S5; the server calculates the Grassmann distance between the local principal gradient subspaces of different clients according to step S602, and selects the nearest neighbor client for each client based on the distance result, generating an aggregated subspace oriented towards that client; subsequently, the server completes model parameter aggregation according to step S7. The above process is repeated in the order of incremental tasks until all tasks are trained.
[0196] To further illustrate the technical effects of this embodiment, the client-oriented adaptive subspace aggregation implementation of this invention is compared with various federated learning and federated incremental learning methods. Evaluation metrics include ACC (Average Accuracy) and BWT (Backward Transfer). Specifically defined as... , in, Indicates completion of the first After training on the nth task, the model was on the th... Accuracy on the test set of each task Indicates completion of all tasks. After the first task, on the... Accuracy on the test set of each task This indicates that the model has just completed the first... After the first task, on the... Accuracy on the test set for each task. ACC measures the model's average recognition performance over consecutive tasks, while BWT measures the impact of training on new tasks on previously learned tasks. The closer the BWT value is to 0, the lower the degree of forgetting.
[0197] The engineering parameters suggested in this disclosure are as follows:
[0198] Table 2: Experimental results of client-oriented adaptive aggregation method in heterogeneous scenario of skin cancer medical image feature distribution.
[0199]
[0200] Table 2 shows the experimental results of the client-oriented adaptive aggregation method in a heterogeneous scenario of skin cancer medical image feature distribution. As can be seen from Table 2, in a heterogeneous scenario of skin cancer medical image feature distribution consisting of four real medical institutions or data centers, the FedSubmerg-A method (client-oriented adaptive aggregation method) of this invention achieves an ACC of 81.50 and a BWT of -10.30. Compared with conventional federated learning methods such as FedAvg, FedProx, and Scaffold, this method improves both average accuracy and forgetting suppression; compared with federated incremental learning methods such as Fed-EWC, Fed-GPM, and Fed-DER, this method also achieves a higher ACC while maintaining a smaller degree of forgetting. These results indicate that in real multi-center medical image data, there is often heterogeneity in feature distribution among different medical institutions caused by differences in acquisition conditions, imaging quality, case composition, and sample size. Simply performing uniform aggregation on all client subspaces may introduce constraint directions that do not match the data distribution of some clients. This invention measures the geometric similarity between local principal gradient subspaces of clients using Grassmann distance and generates corresponding aggregated subspaces for different clients. This reduces interference from irrelevant client information and improves the relevance of the aggregated results to the training of subsequent incremental tasks for each client.
[0201] Therefore, in the embodiments of this disclosure, in real multi-center medical image feature distribution heterogeneous scenarios, more targeted cross-client knowledge sharing can be achieved through an adaptive subspace aggregation mechanism of "similarity selection - targeted aggregation - targeted distribution", and the performance of federated incremental learning can be improved while keeping the data from leaving the local machine.
[0202] It should be understood that in the embodiments of the present invention, the term "and / or" is merely a description of the relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, the character " / " in this document generally indicates that the preceding and following associated objects have an "or" relationship.
[0203] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0204] In the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the couplings or direct couplings or communication connections shown or discussed may be indirect couplings or communication connections through some interfaces, apparatuses, or units, or they may be electrical, mechanical, or other forms of connection.
[0205] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0206] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A federated incremental learning method based on subspace aggregation, characterized in that, Includes the following steps: The server initializes the global model and distributes it to each client. Each client performs local training of the current incremental task based on local data and calculates the gradient of the model parameters. The incremental task is a medical image classification task. Each client constructs a local principal gradient subspace based on the gradient, which compactly represents the update direction of the current task in the parameter space in a low-rank form. During subsequent incremental task training, each client performs orthogonal projection constraints on the current gradient based on the aggregated subspace issued by the server, in order to suppress parameter updates along historical directions; After completing the training of the current incremental task, each client removes the existing historical direction components from the gradient information of the current incremental task, extracts the newly added orthogonal directions, and incrementally updates the local principal gradient subspace. The updated local principal gradient subspace is then uploaded to the server. The server receives the local main gradient subspace uploaded by each client, generates an aggregated subspace according to a preset aggregation strategy, and distributes the aggregated subspace to each client. The server aggregates the local model parameters uploaded by each client to obtain new global model parameters; The construction of the local principal gradient subspace includes: concatenating multiple sampled gradient vectors into a gradient matrix, performing singular value decomposition on the gradient matrix, and selecting the principal direction as the basis of the local principal gradient subspace according to the energy preservation threshold; The preset aggregation strategy includes unified global aggregation, which includes: the server merging the local principal gradient subspaces uploaded by all clients to obtain a unified aggregation matrix, performing singular value decomposition on the unified aggregation matrix and selecting the principal direction according to the server-side energy threshold to generate a global aggregation subspace, and uniformly distributing the global aggregation subspace to all clients. Since the subspace basis matrix is a low-rank directional representation obtained by truncating the gradient matrix using SVD, it does not contain the original image pixel or label information. Therefore, there is no need to transmit the original training samples during cross-institutional transmission, and the server does not need to centrally store historical data.
2. The method as described in claim 1, characterized in that, The orthogonal projection constraints include, Remove the projected component of the current gradient onto the aggregated subspace to obtain the constrained gradient, and update the local model parameters based on the constrained gradient.
3. The method as described in claim 1, characterized in that, The incremental update of the local principal gradient subspace includes The gradient matrix of the current task is deduplicated by removing its projection components on the historical subspace basis. The deduplicated matrix is then subjected to singular value decomposition and new directions are extracted according to the cumulative energy threshold. The new directions are then concatenated with the original historical subspace basis and orthogonally normalized to obtain the updated local principal gradient subspace.
4. The method as described in claim 1, characterized in that, The preset aggregation strategy includes client-oriented adaptive aggregation.
5. The method as described in claim 4, characterized in that, The client-oriented adaptive aggregation includes: the server calculating the geometric similarity between the local principal gradient subspaces of each client; selecting a set of neighboring clients whose local principal gradient subspaces are similar to those of each client; constructing an aggregation matrix based on the set of neighboring clients and the client's own local principal gradient subspace; performing singular value decomposition on the aggregation matrix and selecting the principal direction according to the server-side energy threshold to generate an aggregation subspace oriented towards that client; and then separately distributing the aggregation subspace oriented towards that client to the corresponding client.
6. The method as described in claim 1, characterized in that, The client's local data exhibits a non-independent, identically distributed characteristic.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the method as described in any one of claims 1 to 6.
8. A computer program product, comprising a computer program, characterized in that, The computer program is executed by a processor to implement the method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Dynamic cross-language cooperation method for multilingual large model federal learning
CN121902919A
Federal learning optimization method and system oriented to heterogeneous environment, and storage medium
CN122047394A