A method and system for federated learning aggregation of quantum weights

CN122886731APending Publication Date: 2026-10-09SHENZHEN POLYTECHNIC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611213394.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-11
Publication Date
2026-10-09

AI Technical Summary

Technical Problem

Weighted方法(wpQFL)通过固定插值系数和将本地模型与全局平均进行线性混合但系数不可自适应调整,无法随数据异质性程度变化而自动调节;Clustered方法(mdQFL)通过-means将参数向量划分为个簇后以簇代表均值作为聚合结果,但簇数需预先指定且在的小规模联邦中退化为简单的离群剔除操作;Density方法(qFedInf)对参数向量拟合高斯混合模型(GMM)后以似然度进行加权,但需预先指定混合分量数且在高维小样本条件下(如)似然面估计质量差

Benefits of technology

(一)实现了参数无关的自适应聚合。本发明采用数据驱动的方式,以每轮客户端参数向量间成对距离的中位数作为核带宽,带宽随参数分布离散程度自动调节,不引入任何人工预设超参数。由于带宽完全由本轮参数向量的统计特性决定,避免了现有技术中因依赖数据异质性未知程度而需反复调参的缺陷,降低了量子联邦学习系统的部署门槛和运维成本。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122886731A_ABST
    Figure CN122886731A_ABST
Patent Text Reader

Abstract

The application discloses a kind of nuclear weighting quantum federal learning aggregation method and system, comprising: obtaining the variational quantum circuit parameter vector uploaded by each client;The pair distance between any two clients is calculated;With the median of all pair distances as bandwidth;According to the pair distance and bandwidth, the nuclear affinity is calculated, and the affinity matrix is obtained;According to the affinity matrix, the nuclear affinity of each client relative to other clients is summed to obtain the centrality score;The aggregation weight is determined according to the centrality score;The parameter vector of each client is weighted and aggregated according to the aggregation weight, and the global model parameter vector is obtained and broadcasted.The application realizes parameter-independent adaptive aggregation, which can effectively suppress client drift and improve aggregation accuracy in the non-independent and identically distributed data scenario without relying on artificial hyperparameter configuration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of quantum computing and artificial intelligence technology, and in particular relates to a kernel-weighted quantum federated learning aggregation method and system. Background Technology

[0002] Quantum Federated Learning (QFL) is a machine learning paradigm that collaboratively trains variational quantum circuits (VQCs) on distributed quantum computing nodes. In a typical QFL system, a parameter server and... Each client connects via a communication network. Holding private local datasets The system goal is to train a shared VQC classifier. By solving the minimization problem Collaborative training is completed while ensuring that the original data is not leaked by each client. VQC includes data encoding and feature mapping. and parameterized proposed circuit Two phases. Training is conducted in synchronous communication rounds. Process: In each round, the server will display the current global model. Broadcast to all clients, each client executes. Upload local parameters after local stochastic gradient descent. The server uses aggregate functions Generate the next round of global model .

[0003] The aggregation mechanism commonly used in QFL systems is the Federated Average (FedAvg), defined as... The unweighted arithmetic mean of the client parameter vectors is taken: This mechanism is optimal under the condition of independent and identically distributed (IID) client data, because all local models converge to similar regions in the parameter space. However, when the client data is non-independent and identically distributed (non-IID), each client's local model drifts towards its own local optimum during local training, resulting in the phenomenon of "client drift." Under this condition, FedAvg still allocates equal resources to each client. The voting weights cause anomalous parameters from individual divergent clients to participate in the global model construction with the same weight, pulling the global model away from the direction of group consensus. The root of this structural flaw lies in FedAvg's aggregation mechanism, which completely ignores the geometric distribution relationship between client parameter vectors and cannot distinguish between consensus clients converging into the shared region and outlier clients deviating into low-density regions.

[0004] To address the client drift problem in FedAvg with non-IID data, existing technologies have proposed several alternative aggregation schemes, but these schemes all have structural flaws. The first type of scheme suppresses the influence of outlier clients by introducing adjustable hyperparameters. The Weighted method (wpQFL) uses fixed interpolation coefficients... and Linearly mix the local model with the global average. However, the coefficients cannot be adaptively adjusted and cannot automatically adjust with changes in the degree of data heterogeneity; the Clustered method (mdQFL) uses... -means divides the parameter vector into After each cluster, the cluster mean is used as the aggregation result, but the number of clusters... It needs to be specified in advance and in In small-scale federations, it degenerates into a simple outlier removal operation; the Density method (qFedInf) applies a Gaussian mixture model (GMM) to the parameter vectors and then weights them by likelihood, but the number of mixture components must be specified in advance. And under high-dimensional small sample conditions (such as...) The likelihood surface estimation quality is poor. A common drawback of the above solutions is that the optimal values ​​of the hyperparameters depend on the unknown degree of data heterogeneity, which cannot be predetermined during actual deployment, leading to unstable performance or incurring high hyperparameter search costs.

[0005] The second type of approach circumvents the hyperparameter problem by altering the system topology, but introduces other structural costs. The Personalized method (wpQFL) employs a hard threshold decision rule to calculate the distance from each client to the global model. Distance to its own historical parameters ,like If the global model is retained, the midpoint is taken; otherwise, the midpoint is used. This method retains an independent local model for each client instead of outputting a single global model, increasing server-side complexity. The first approach incurs high memory overhead and is unsuitable for standard QFL deployments requiring a single global model. The Chain approach (ccQFL) arranges clients in a serial training chain to eliminate a central server. Each client receives the model from the previous client, trains it locally, and then passes it to the next client. The parameters of the client at the end of the chain constitute the global model. However, this introduces order dependency and error cascading propagation problems. If the data distribution of the client at the beginning of the chain is extremely skewed, its error will propagate along the chain and amplify at each level. The common drawback of these two approaches is that, in order to eliminate hyperparameters, they sacrifice the simplicity of the standard QFL star topology and the uniformity of the global model.

[0006] In summary, existing QFL polymerization techniques generally suffer from the following structural defects: either they rely on hyperparameters that need to be preset manually (…). The optimal value of hyperparameters is unknown in actual deployment; alternatively, the system topology can be changed to circumvent hyperparameters, but this introduces order dependencies, memory overhead, or loss of global model consistency. The root cause of these shortcomings is that existing aggregation methods fail to utilize the geometric distribution information carried by the client parameter vectors in each round of communication to drive adaptive weighting, instead relying on externally preset configurations or topologies to indirectly address client drift. Summary of the Invention

[0007] To address the aforementioned technical problems, this invention proposes a kernel-weighted quantum federated learning aggregation method and system to resolve the issues present in the prior art.

[0008] To achieve the above objectives, this invention provides a kernel-weighted quantum federated learning aggregation method, applied to a quantum federated learning system deployed on a parameter server, comprising: Obtain the local model parameter vector uploaded by each client participating in federated learning. The local model parameter vector is a vector composed of trainable parameters in the variable quantum circuit. Calculate the pairwise distance between any two clients based on the local model parameter vectors provided by each client; The bandwidth used for kernel density estimation in the parameter group composed of each client is determined based on the pairwise distance; Based on the pairwise distance and the bandwidth, calculate the kernel affinity between any two clients to obtain the affinity matrix; Based on the affinity matrix, the core affinity of each client relative to all other clients is summed to obtain the centrality score of each client; The aggregation weight of each client is determined based on the centrality score of each client. The local model parameter vectors of each client are weighted and aggregated according to the aggregation weight to obtain a global model parameter vector, which is then broadcast through the network communication interface of the parameter server.

[0009] Optionally, the pairwise distance is the squared Euclidean distance between the local model parameter vectors of each client. The pairwise distances form a symmetric distance matrix, the diagonal elements of the distance matrix are zero, and the distance matrix is ​​stored in the memory of the parameter server in a form that stores only the upper triangular elements or only the lower triangular elements.

[0010] Optionally, the process of determining the bandwidth used for kernel density estimation in the parameter group consisting of each client based on the pairwise distance includes: The median of all pairwise distances is used as the square of the Gaussian kernel bandwidth; When the total number of clients participating in federated learning is less than 3 or the square of the bandwidth is less than a preset stable threshold, the parameter groups formed by the clients are indistinguishable in the parameter space, and the aggregate weights of the clients are equal.

[0011] Optionally, the process of calculating the kernel affinity between any two clients includes: The nuclear affinity between any two clients is obtained by dividing the negative pairwise distance by the square of twice the bandwidth as the exponent of the natural constant exponential function. The kernel function is any one of the Gaussian kernel function, Laplace kernel function, or Cauchy kernel function.

[0012] Optionally, the centrality score of a client can be obtained by summing the kernel affinity of the client relative to all other clients participating in federated learning.

[0013] Optionally, the aggregation weight of each client is determined based on the centrality score of each client, including: Using the sum of the centrality scores of all clients as a normalization factor, the centrality score of each client is divided by the normalization factor to obtain the normalized aggregate weight of that client. When the normalization factor is lower than a preset stable threshold, the aggregation weights of each client are equal.

[0014] Optionally, the local model parameter vectors of each client are weighted and aggregated according to the aggregation weights, including: The global model parameter vector is obtained by multiplying the local model parameter vector of each client with the aggregate weight of that client and summing the results over all clients.

[0015] This invention also provides a kernel-weighted quantum federated learning aggregation system, applied to a quantum federated learning system deployed on a parameter server, comprising: The pairwise distance calculation module is used to obtain the local model parameter vectors uploaded by each client and calculate the pairwise distance between the local model parameter vectors of any two clients; the local model parameter vector is a vector composed of trainable parameters in the variable quantum circuit. A bandwidth selection module is used to determine the bandwidth based on the paired distance; The nuclear affinity calculation module is used to calculate the nuclear affinity based on the pairwise distance and the bandwidth, and output the affinity matrix; The centrality calculation module is used to sum the kernel affinity of each client relative to other clients based on the affinity matrix, and output the centrality score; The weight normalization module is used to determine the aggregate weight based on the centrality scores of each client. The weighted aggregation module is used to perform weighted aggregation on the local model parameter vectors of each client according to the aggregation weight, output a global model parameter vector, and broadcast the global model parameter vector through the network communication interface of the parameter server.

[0016] Optionally, the pairwise distance calculation module includes a vector subtraction operator and a square accumulator, which calculate the squared Euclidean distance using the vector subtraction operator and the square accumulator; the bandwidth selection module includes a sorting network composed of a hardware comparator and a multiplexer, which calculates the median of all pairwise distances using the sorting network, and uses the median as the square of the Gaussian kernel bandwidth; the kernel affinity calculation module includes an exponential function calculation unit, which calculates the Gaussian kernel affinity based on the squared Euclidean distance and the square of the bandwidth using the exponential function calculation unit.

[0017] Optionally, the output of the weighted aggregation module is connected to the network communication interface, and the global model parameter vector is broadcast to each client through the network communication interface; after receiving the global model parameter vector, each client writes each element of the global model parameter vector into the register of the variable quantum circuit controller to set the rotation angle of each rotating door in the variable quantum circuit.

[0018] Compared with the prior art, the present invention has the following advantages and technical effects: (i) Parameter-independent adaptive aggregation is achieved. This invention adopts a data-driven approach, using the median of the pairwise distances between client parameter vectors in each round as the kernel bandwidth. The bandwidth automatically adjusts according to the dispersion of parameter distribution, without introducing any manually preset hyperparameters. Since the bandwidth is entirely determined by the statistical characteristics of the parameter vectors in this round, it avoids the shortcomings of existing technologies that require repeated parameter tuning due to the unknown degree of data heterogeneity, thus reducing the deployment threshold and operation and maintenance costs of the quantum federated learning system.

[0019] (ii) Performance remains unaffected under independent and identically distributed operating conditions. When the data from each client is independently and identically distributed, the local model converges to similar regions in the parameter space, the pairwise distance between the parameter vectors of each client approaches zero, the kernel affinity approaches 1, and the centrality is approximately equal. This invention automatically degenerates into the standard federated average. This ensures that under normal operating conditions with uniform data distribution, the aggregation accuracy remains consistent with the federated average, eliminating performance concerns when deploying this invention.

[0020] (III) Suppressing Client Drift in Non-Independent and Identically Distributed Data Scenarios. This invention calculates the centrality score of each client in the parameter population through kernel density estimation. Consensus clients located in high-density regions accumulate high centrality due to their proximity to multiple clients, thus gaining greater aggregation weights. Outlier clients deviating to low-density regions have centrality approaching zero due to their isolation from all other clients, and their contributions are automatically attenuated. Since the aggregation result is dominated by the majority of consensus clients, this invention addresses the problem of model drift towards outlier clients caused by equal-weight aggregation in federated averages at the root of the geometric distribution of the parameter space, thereby improving model accuracy in non-independent and identically distributed data scenarios.

[0021] (iv) The aggregation advantage increases with the scale of the federation. The increase in the number of clients enriches the samples in the parameter space, makes the distinction between high-density consensus regions and low-density outlier regions clearer, and improves the discriminative power of centrality scores. As a result, in larger-scale federated systems, it is possible to more accurately identify and suppress outlier clients, while more fully aggregating the effective information of consensus clients.

[0022] (v) Low computational and storage overhead, allowing direct deployment in existing systems. The computational complexity of this invention is only related to the square of the number of clients and the parameter dimension. It requires fewer distance calculations and kernel evaluations per round, and does not increase the number of trainable parameters for the variable quantum circuit, nor does it change the communication topology or number of rounds in quantum federated learning. Therefore, it can serve as a plug-and-play replacement for federated averaging modules, requiring no modification to the physical structure of the quantum circuit, no increase in qubit resources, and no adjustment to the client's local training protocol for deployment. Attached Figure Description

[0023] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a schematic diagram illustrating the principle of the kernel weighted aggregation (KWA) mechanism in an embodiment of the present invention; Figure 2 This is a data distribution diagram of each client under the Dir(0.5) non-independent and identically distributed condition according to an embodiment of the present invention; Figure 3 This is a comparison of the convergence curves of KWA and FedAvg under the IID and Dir(0.5) conditions in this embodiment of the invention. Figure 4 This is a diagram showing the evolution trajectory of the KWA aggregation weights in T=10 rounds of communication according to an embodiment of the present invention. Figure 5 This is a comparison of the convergence curves of KWA and FedAvg under different federation sizes (N=4, 8, 16) in this embodiment of the invention. Figure 6 This is a flowchart of a method according to an embodiment of the present invention. Detailed Implementation

[0024] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0025] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0026] Example 1 like Figure 6 As shown, this embodiment provides a kernel-weighted quantum federated learning aggregation method, including: In each round of communication executed on the parameter server side, the received data is processed... Client parameter vector The following six steps are performed to obtain the weighted aggregated global model parameters. Indicates the first The local model parameter vector uploaded by each client is a 3D real vector ( (This refers to the total number of trainable parameters in a variable quantum circuit). For client number, The total number of clients participating in federal training.

[0027] The following detailed explanation, with reference to the accompanying diagrams and specific steps, will be provided. Figure 1 This is a schematic diagram illustrating the principle of the KWA mechanism of this invention in a two-dimensional parameter space, showing... The spatial distribution of client parameter vectors, pairwise affinity, and the aggregation weight allocation results derived from centrality.

[0028] S1. Pair distance calculation.

[0029] The server calculates the pairwise squared Euclidean distance between all client parameter vectors: This step quantizes the relative positional relationships of the clients in the parameter space. The distance matrix, stored in server memory as a floating-point number, provides the geometric basis for subsequent density estimation. The distance matrix is ​​symmetric with zero diagonal elements; only the upper triangular distance needs to be calculated. Each element. For , For a typical QFL configuration, this step only requires Second-rate 3D vector distance calculation. Wherein, Indicates the client With the client The squared Euclidean distance between the parameter vectors, and The first The and the first Uploaded by each client dimensional parameter vector, This represents the total number of clients participating in this round of aggregation.

[0030] S2. Bandwidth selection.

[0031] The Gaussian kernel bandwidth is determined using the median heuristic: Compared to the mean, the median is naturally robust to extreme outliers. Even if there exists an outlier parameter vector that is extremely far from all other clients, the median still reflects the typical distance scale between most clients. This bandwidth is entirely data-driven and does not introduce any adjustable hyperparameters. In training epochs with a compact parameter distribution... Automatic scaling down to enhance kernel resolution in rounds with discrete parameter distributions. Automatic scaling avoids premature commitment to a single region of the parameter space, achieving round-level adaptive adjustment. This occurs when all client parameters are equal within machine precision. Or, the paired information is insufficient to support a meaningful density estimate. When this happens, it automatically degenerates into a uniform weight and directly returns the FedAvg result. Among these, This is the square of the Gaussian kernel bandwidth, controlling the smoothness of the kernel function. The larger the core, the flatter it is and the more uniform the weights tend to be. The smaller the kernel, the sharper the kernel and the more sensitive the weights are to distance. This represents the median operation, taking the median value after sorting all pairs of squared distances; ( ) is the set of squared Euclidean distances between all different client pairs calculated in step S1.

[0032] S3. Nuclear affinity calculation.

[0033] For each different client pair Calculate Gaussian kernel affinity: When the distance between the parameter vectors of two clients is much smaller than the bandwidth ( ), affinity is close to This indicates that the two converge to similar parameter regions and are considered "neighbors"; when the distance is much greater than the bandwidth ( The affinity index decays to near its maximum. This indicates that the two are isolated from each other in the parameter space. Affinity matrix Stored in server memory as a floating-point matrix, with diagonal elements Not used. Experiments have shown that alternative kernel functions, such as the Laplace kernel, are preferred. Cauchy nucleus And the difference in classification accuracy of Gaussian kernel is less than This indicates that the method is insensitive to the form of the kernel function, and the Gaussian kernel is chosen as the default due to its smoothness and natural correspondence with the squared Euclidean distance. Among these, For the client For the client The nuclear affinity, with a value range of The larger the value, the closer the two are in the parameter space; For the natural constant An exponential function with base 0; The squared Euclidean distance between the two clients calculated in step S1; This represents the squared bandwidth determined by the median heuristic in step S2. Affinity matrix. for OK A square matrix of columns, where the first column is... Line number Column elements are Diagonal elements are not used.

[0034] S4. Centrality calculation.

[0035] Define client The centrality of is the sum of its affinities to all other clients: Models located in high-density regions of the parameter space and adjacent to multiple other clients will accumulate high centrality scores; outliers isolated in low-density tails will acquire near-zero centrality because their affinity with all other clients is close to zero. Centrality Essentially a client In the present The nonparametric kernel density estimates of the parameter population formed by the parameter vectors do not require assumptions about the specific form of the parameter distribution. For the client Central score ( ), equal to client For all other clients ( nuclear affinity sum; This represents the total number of clients. Higher centrality indicates a denser parameter space region where the client resides and is closer to other clients.

[0036] S5. Weight normalization.

[0037] Normalize the centrality scores into aggregate weights: The weights are non-negative and their sum is . If the sum of all centralities is below the numerical stability threshold... This indicates that the parameters of all clients are extremely dispersed, and the affinity between them is negligible, thus degenerating into uniform weights. This normalization gives clients in high-density regions greater voting power, while the contributions of low-density outliers are automatically attenuated, with the degree of attenuation proportional to the degree of isolation. Completely isolated clients have weights approaching zero. For the client The normalized aggregate weights satisfy and ; For the client calculated in step S4 Central fractions; denominator For all The sum of the centrality of each client.

[0038] S6. Weighted aggregation.

[0039] The global model is a weighted average of all local parameter vectors: An alternative to the unweighted average in standard FedAvg aggregation. Among them, For KWA aggregation output The global model parameter vector will be broadcast to all clients as the initial model for the next communication round; The first calculated in step S5 Normalized aggregate weights for each client; For the first Uploaded by each client 3D local parameter vector; This represents the total number of clients. This aggregation result... The global model, used for the next communication round, is broadcast to all clients via the server's network communication interface, completing one round of QFL training. Under IID data conditions, the parameters of each client converge to similar regions, with pairwise distances... affinity centrality Weight KWA automatically reverts to standard FedAvg without any conditional checks or manual switching.

[0040] This invention also provides a kernel-weighted aggregation system deployed on a quantum federated learning parameter server, comprising the following six sequentially connected modules: (1) Paired distance calculation module: receiving Uploaded by each client A dimensional parameter vector is used to calculate all pairwise squared Euclidean distances using a vector subtraction operator and a square accumulator, and the output is... The symmetric distance matrix is ​​stored in server memory as a double-precision floating-point matrix. This module can utilize the matrix multiplication unit of the GPU for batch parallel distance calculations.

[0041] (2) Bandwidth selection module: Reads the upper triangular elements of the distance matrix The median of a scalar value is calculated using a sorting network consisting of a hardware comparator and a multiplexer, and the output is the scalar bandwidth. It stores data in double-precision floating-point numbers. This module does not contain any configurable registers or parameter tables; bandwidth is entirely determined by the statistical characteristics of the input data.

[0042] (3) Nuclear affinity calculation module: Reads the distance matrix and bandwidth This is achieved through exponential function calculation units or lookup tables, calculating each pair Gaussian nuclear affinity Output The affinity matrix has its diagonal elements set to zero.

[0043] (4) Centrality calculation module: sum the affinity matrix row by row and skip the diagonal elements, and generate the matrix using a floating-point adder tree. dimensional centrality vector .

[0044] (5) Weight normalization module: Summing all elements of the centrality vector yields Calculation via floating-point divider array ,like Then output Output Power Vector Non-negative and sum to .

[0045] (6) Weighted aggregation module: using weight vectors right indivual Weighted summation of dimensional parameter vectors Output the aggregated result A global parameter vector is broadcast to all clients via the server's network communication interface (TCP / IP or gRPC protocol).

[0046] In actual deployment, the above six modules run as computer programs on the central processing unit (CPU) or graphics processing unit (GPU) of the quantum federated learning parameter server. Data transmission between modules is completed through the server's memory bus. Each client receives global parameters. Then, the parameter values ​​are written into the registers of the variable quantum line controller via the data bus of the client computer, which is used to set the parameters of each component in the line. and The rotation angle of the revolving door. This aggregation system, as a plug-and-play alternative to FedAvg, requires no modification to the physical structure of the quantum circuitry, no increase in the number of qubits, and no change to the client-side local training protocol, local optimizer, learning rate, or local steps. If everything remains unchanged, deployment can be completed.

[0047] The computational complexity of KWA is Floating-point operations per round: Step 1 requires Calculate all paired distances. Step 2 requires... Sorting (can be optimized to [specific value] when using a comparator network) Step 3 requires Calculate affinity, steps 4-6 each require... or .for , The typical configuration, totaling Sub-distance calculation and Secondary nuclear assessment; for , Larger configuration, total Sub-distance calculation and Secondary core evaluation. The computational workload mentioned above is much smaller than... The overhead of each VQC execution and parameter shift gradient calculation, involving tens to hundreds of quantum gate operations and multiple quantum state measurements per VQC execution, does not constitute a system performance bottleneck. KWA introduces... Additional memory is used to store the parameter vector and Used for temporary affinity matrices, without increasing the number of trainable parameters in VQC. When The system automatically degenerates to the standard FedAvg.

[0048] Figure 1 Taking N=4 clients as an example, each dot represents the position of the parameter vector uploaded by a client in the parameter space. The lines connecting the dots represent pairwise kernel affinity. The thicker the line, the stronger the affinity and the closer the two clients are in the parameter space. Each dot is labeled with its aggregation weight, obtained through centrality normalization. The isolated client located in the lower left corner (with a weight close to...) Because they are far from other clients, the weights are automatically reduced, and the three clients in the upper left corner that are close to each other share the majority of the voting power. This diagram intuitively illustrates the core principle of KWA: regions with higher density in the parameter space receive greater aggregation weights, and outliers are automatically suppressed.

[0049] like Figure 2 As shown, the Dirichlet distribution under conditions The data distribution across clients is illustrated in the figure, which uses five binary classification tasks of MNIST handwritten digits (0 vs 9, 1 vs 8, 2 vs 7, 3 vs 6, 4 vs 5) as examples to show the distribution of the number of samples of each class held by each client in each task. "The filled bars with diagonal lines represent the number of samples in the first category," "The diagonally filled bars represent the number of samples in the second category, with the numbers on the bars indicating the specific sample size. As can be seen from the graph, in..." Within each partition, the proportion of samples from the two classes differs significantly across clients. Some clients are dominated by samples from one class. This label imbalance is the root cause of the drift of the local model in different directions.

[0050] like Figure 3 As shown, KWA and FedAvg are in Comparison of convergence curves under different conditions. The horizontal axis in the graph represents the number of communication rounds. The vertical axis represents the test classification accuracy. The solid line represents KWA, the dashed line represents FedAvg, and the shaded area represents the results of five random seed tests. Standard deviation interval. Subplot (a) shows the 0 vs 9 digit pair classification task, and subplot (b) shows the 2 vs 7 digit pair classification task. As can be seen from the figures, the KWA curve separates upward from the FedAvg curve in the later stages of training, ultimately achieving higher terminal accuracy on both tasks, demonstrating the effect of centrality weighting in suppressing drift of non-IID clients.

[0051] like Figure 4 As shown, KWA aggregate weights in The evolutionary trajectory in round-robin communication. The horizontal axis represents the number of communication rounds, and the vertical axis represents the aggregate weight assigned to each client. ( The horizontal dashed line marks the uniform weighted baseline. The four curves with different line types correspond to the four clients. The weight is in the first position. The round is determined by the initial random parameters, and the previous rounds are... to After adjustments, the distribution tends to stabilize, and the final distribution range is approximately... to The figure shows that KWA's centrality-weighted mechanism can converge to a stable weight allocation scheme within a limited number of communication rounds, without manual intervention or parameter configuration.

[0052] like Figure 5 As shown, different federal sizes ( KWA and FedAvg are under the following conditions. Comparison of convergence curves under different conditions. The three subplots in the figure correspond to... , and The federation size for each client is shown on the horizontal axis, representing the number of communication rounds, and the vertical axis represents the test classification accuracy. Solid lines (circular markers) represent KWA, dashed lines (square markers) represent FedAvg, and the shaded area represents the results of five random seed tests. Standard deviation interval. The two curves basically overlap. KWA began to show a slight advantage; Significant differences were observed over time—FedAvg entered a plateau period after the second round of communication, while KWA continued to improve until the tenth round, validating the scalability characteristic of centrality weighting, which benefits from increased centrality weighting as the federation size expands.

[0053] This invention achieves truly parameter-independent operation, reducing the deployment threshold and maintenance costs of QFL systems. In existing technologies, Weighted requires configuration... and Two interpolation coefficients; for Clustered, the number of clusters must be specified. Density requires specifying the number of Gaussian mixture components. The optimal values ​​of these hyperparameters depend on the unknown severity of data heterogeneity, and in actual deployment, they need to be determined through repeated trials or costly hyperparameter searches. This invention employs a data-driven median bandwidth heuristic, where bandwidth... The accuracy is entirely determined by the parameter vector received in each round, without introducing any hyperparameters that need to be manually set. Bandwidth ablation experiments show that replacing the median with the mean results in a smaller change in average precision than... Replacing the Gaussian kernel with a Laplace or Cauchy kernel results in a smaller change in accuracy than... The proof method is insensitive to specific modeling choices and requires no precise calibration. In a typical configuration, KWA runs completely automatically, requiring no manual intervention or parameter configuration steps.

[0054] This invention automatically restores the data to standard FedAvg under IID conditions, ensuring lossless performance under normal operating conditions. When all client data is IID, the local model, after local training, converges to a similar region in the parameter space, the pairwise distance between the parameter vectors of each client approaches zero, and the kernel affinity approaches zero. With approximately equal centrality, the KWA aggregate weights degenerate into uniform weights. As shown in Table 1, the average classification accuracy of KWA under the IID condition is... With FedAvg The difference is only The two are statistically indistinguishable, verifying the IID limit reduction property of the theoretical analysis. Figure 3 A comparison of the convergence curves of KWA and FedAvg under the Dir(0.5) condition is presented. Figure 3 (a) in the text represents tasks 0 vs 9. Figure 3 In (b) of the diagram, which compares tasks 2 and 7, we can see that the KWA curves separate from FedAvg and maintain higher terminal accuracy. In contrast, the Personalized method's accuracy decreases to [a lower value] under the same conditions. The Weighted method is reduced to This indicates that existing alternatives suffer from significant performance losses even under normal operating conditions.

[0055] This invention adaptively suppresses outlier clients under non-IID conditions, improving the accuracy and robustness of the aggregation model. Figure 2 The data distribution of each client under the Dir(0.5) partition is shown, and the imbalance of this label is the source of client drift. Figure 3 (a) and (b) illustrate the convergence separation process of KWA and FedAvg under the non-IID condition. As shown in Table 2, under the Dirichlet distribution... In simulated non-IID scenarios, KWA achieved the highest average accuracy among all seven methods. Exceeding FedAvg ( Among existing technologies, all five alternative methods rank below FedAvg: Weighted Chain Clustered Density Personalized The gap between FedAvg and third-ranked Weighted is... It is approximately the gap between KWA and FedAvg. The performance of the Chain, Clustered, and Density methods decreased by [number] times under non-IID conditions compared to their IID counterparts. , and KWA and FedAvg showed the smallest non-IID performance degradation, respectively. and . Figure 4 The evolution trajectory of the KWA aggregation weights in T=10 rounds of communication is shown. After the first 2-3 rounds of adjustment, the weights tend to stabilize, with a final distribution range of approximately 0.220 to 0.291, deviating from the uniform baseline by approximately 0.25. This indicates that centrality weighting produces a moderate but automatic weight redistribution under the Dir(0.5) condition. These results demonstrate that the advantages claimed by existing alternatives in the original study do not hold under a controlled uniform benchmark, and KWA is the only aggregation method that maintains optimal or near-optimal performance under both IID and non-IID conditions.

[0056] The aggregation advantage of this invention increases with the scale of the federation, adapting to future large-scale QFL deployments. As shown in Table 3, the number of clients... from Increase to and In large-scale experiments, such as Figure 5 As shown, KWA's accuracy advantage over FedAvg stems from... ( Growth to ( ), ultimately reaching ( ), advantage growth approximately The accuracy of FedAvg on the 2v7 digit pair classification task increased by several times. time steep drop time (decline) KWA remained stable at ( ).exist In this scenario, FedAvg plateaus after the second round of communication and stops improving, while KWA continues to improve until the tenth round. This trend is consistent with the design principle of centrality-weighted algorithms: as the number of clients in the federation increases, the parameter space density estimation benefits from richer samples, and the boundary between consensus and outlier regions becomes sharper. FedAvg accumulates the negative impact of client drift due to its continuous use of equal weights, while KWA automatically eliminates outlier interference through centrality-weighted algorithms.

[0057] This invention has extremely low computational and storage overhead and does not constitute a system performance bottleneck. The computational complexity of KWA per round is... ,exist , With a larger configuration, only Sub-distance calculation and Secondary core evaluation, with the single-precision floating-point arithmetic capabilities (tens of GFLOPS) of modern CPUs, can be completed in microseconds. Compared to each round... Each VQC execution and parameter shift gradient calculation involves tens to hundreds of quantum gate operations and multiple quantum state measurements, taking milliseconds to seconds on a simulator. KWA's computational overhead is negligible. Additional memory overhead is... It does not increase the number of trainable parameters in VQC. As a direct replacement for FedAvg, KWA does not change the communication topology of the QFL system, does not increase the number of communication rounds, and does not affect the client's local training protocol. It can be deployed as an aggregation layer plugin in the standard QFL software stack.

[0058] Table 1 The bolded values ​​in Table 1 indicate the optimal values ​​for that column. The accuracy difference between KWA and FedAvg under the IID condition is +0.0011, and the paired t-test value is p=0.178. The two are statistically indistinguishable, verifying the theoretical property that KWA can be reduced to FedAvg under the IID limit.

[0059] Table 2 The bolded values ​​in Table 2 indicate the best values ​​for that column. KWA achieved the highest average accuracy of 0.7120. The difference between FedAvg and Weighted, which ranked third, was 0.0391. Chain, Clustered, and Density showed a decrease in performance of 0.143, 0.165, and 0.177 from their IID values, respectively, while KWA showed the smallest decrease (0.0761).

[0060] Table 3 In Table 3, the Δ row represents the accuracy advantage of KWA over FedAvg. The average advantage of KWA increases from +0.0005 when N=4 to +0.023 when N=16. On the 2v7 task with N=16, FedAvg drops to 0.739 while KWA remains at 0.796 (Δ=+0.057).

[0061] This invention provides a kernel-weighted aggregation (KWA) method that automatically identifies consensus and outlier regions in the parameter space through nonparametric kernel density estimation. It assigns higher aggregation weights to clients in high-density regions and automatically decays the weights of isolated clients without any manual hyperparameter configuration. It also provides a data-driven bandwidth selection mechanism, using the median of the pairwise squared distances between client parameter vectors as the kernel bandwidth, automatically adjusting the bandwidth according to the dispersion of parameter distribution in each round. Finally, it provides an aggregation system that implements the above method, seamlessly replacing the FedAvg aggregation module in existing QFL systems, improving model accuracy in non-IID data scenarios without increasing quantum bit resources, changing the VQC structure, or introducing additional communication overhead.

[0062] This invention is applied to distributed quantum machine learning scenarios on noisy medium-scale quantum devices (NISQ), and solves the problems of model drift and aggregation quality degradation caused by non-independent identically distributed (non-IID) client data in quantum federated learning. It constructs the centrality weights of the parameter vectors of each client through nonparametric kernel density estimation, and realizes adaptive weighted aggregation of heterogeneous clients.

[0063] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A kernel-weighted quantum federated learning aggregation method, applied to a quantum federated learning system deployed on a parameter server, characterized in that, Includes the following steps: Obtain the local model parameter vector uploaded by each client participating in federated learning. The local model parameter vector is a vector composed of trainable parameters in the variable quantum circuit. Calculate the pairwise distance between any two clients based on the local model parameter vectors provided by each client; The bandwidth used for kernel density estimation in the parameter group composed of each client is determined based on the pairwise distance; Based on the pairwise distance and the bandwidth, calculate the kernel affinity between any two clients to obtain the affinity matrix; Based on the affinity matrix, the core affinity of each client relative to all other clients is summed to obtain the centrality score of each client; The aggregation weight of each client is determined based on the centrality score of each client. The local model parameter vectors of each client are weighted and aggregated according to the aggregation weight to obtain a global model parameter vector, which is then broadcast through the network communication interface of the parameter server.

2. The kernel-weighted quantum federated learning aggregation method according to claim 1, characterized in that, The pairwise distance is the squared Euclidean distance between the local model parameter vectors of each client. The pairwise distances form a symmetric distance matrix, the diagonal elements of the distance matrix are zero, and the distance matrix is ​​stored in the memory of the parameter server in a form that stores only the upper triangular elements or only the lower triangular elements.

3. The kernel-weighted quantum federated learning aggregation method according to claim 1, characterized in that, The process of determining the bandwidth used for kernel density estimation in the parameter group composed of each client based on the pairwise distance includes: The median of all pairwise distances is used as the square of the Gaussian kernel bandwidth; When the total number of clients participating in federated learning is less than 3 or the square of the bandwidth is less than a preset stable threshold, the parameter groups formed by the clients are indistinguishable in the parameter space, and the aggregate weights of the clients are equal.

4. The kernel-weighted quantum federated learning aggregation method and system according to claim 3, characterized in that, The process of calculating the kernel affinity between any two clients includes: The nuclear affinity between any two clients is obtained by dividing the negative pairwise distance by the square of twice the bandwidth as the exponent of the natural constant exponential function. The kernel function is any one of the Gaussian kernel function, Laplace kernel function, or Cauchy kernel function.

5. The kernel-weighted quantum federated learning aggregation method according to claim 1, characterized in that, The centrality score of a client is obtained by summing its kernel affinity relative to all other clients participating in federated learning.

6. The kernel-weighted quantum federated learning aggregation method according to claim 1, characterized in that, Based on the centrality scores of each client, determine the aggregation weight for each client, including: Using the sum of the centrality scores of all clients as a normalization factor, the centrality score of each client is divided by the normalization factor to obtain the normalized aggregate weight of that client. When the normalization factor is lower than a preset stable threshold, the aggregation weights of each client are equal.

7. The kernel-weighted quantum federated learning aggregation method according to claim 1, characterized in that, The local model parameter vectors of each client are weighted and aggregated according to the aggregation weights, including: The global model parameter vector is obtained by multiplying the local model parameter vector of each client with the aggregate weight of that client and summing the results over all clients.

8. A kernel-weighted quantum federated learning aggregation system, applied to a quantum federated learning system deployed on a parameter server, characterized in that, include: The pairwise distance calculation module is used to obtain the local model parameter vectors uploaded by each client and calculate the pairwise distance between the local model parameter vectors of any two clients. The local model parameter vector is a vector composed of trainable parameters in the variable quantum circuit. A bandwidth selection module is used to determine the bandwidth based on the paired distance; The nuclear affinity calculation module is used to calculate the nuclear affinity based on the pairwise distance and the bandwidth, and output the affinity matrix; The centrality calculation module is used to sum the kernel affinity of each client relative to other clients based on the affinity matrix, and output the centrality score; The weight normalization module is used to determine the aggregate weight based on the centrality scores of each client. The weighted aggregation module is used to perform weighted aggregation on the local model parameter vectors of each client according to the aggregation weight, output a global model parameter vector, and broadcast the global model parameter vector through the network communication interface of the parameter server.

9. The kernel-weighted quantum federated learning aggregation system according to claim 8, characterized in that, The pairwise distance calculation module includes a vector subtraction operator and a square accumulator. The pairwise distance calculation module calculates the squared Euclidean distance using the vector subtraction operator and the square accumulator. The bandwidth selection module includes a sorting network composed of a hardware comparator and a multiplexer. The bandwidth selection module calculates the median of all pairwise distances using the sorting network, and uses the median as the square of the Gaussian kernel bandwidth. The kernel affinity calculation module includes an exponential function calculation unit. The kernel affinity calculation module calculates the Gaussian kernel affinity based on the squared Euclidean distance and the square of the bandwidth using the exponential function calculation unit.

10. The kernel-weighted quantum federated learning aggregation system according to claim 8, characterized in that, The output of the weighted aggregation module is connected to the network communication interface, and the global model parameter vector is broadcast to each client through the network communication interface. After receiving the global model parameter vector, each client writes each element of the global model parameter vector into the register of the variable quantum circuit controller to set the rotation angle of each rotating door in the variable quantum circuit.