Prediction method based on spatiotemporal quantum federated learning for edge scenarios

By constructing a spatiotemporal graph and quantum state encoding for edge device clusters, the problem of non-independent and identically distributed data in traditional federated learning is solved, achieving efficient model aggregation and prediction, which is suitable for edge computing scenarios.

CN121503589BActive Publication Date: 2026-03-27CHINA TOWER CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-08
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Traditional federated learning methods assume that edge device data is independent and identically distributed, resulting in poor model generalization performance and unstable global model convergence. Existing quantum federated solutions rely on real quantum hardware, have high deployment thresholds, and are difficult to apply in conventional edge computing architectures.

Method used

By constructing a dual graph structure based on geographical distance and temporal similarity, spatiotemporal data of edge device clusters are obtained, client model parameters are encoded as quantum states, target clients are filtered using spatiotemporal correlation matrix and parameter difference matrix, dynamic aggregation operation is performed to generate a global model, and prediction is made when convergence conditions are met.

Benefits of technology

It accurately captures the spatiotemporal dynamic patterns of edge data, improves model generalization performance, reduces communication overhead, handles spatiotemporal heterogeneity between devices, and improves model convergence speed and generalization performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121503589B_ABST
    Figure CN121503589B_ABST
Patent Text Reader

Abstract

The application discloses a prediction method for edge scene based on space-time quantum federated learning, relates to the technical field of federated learning, and utilizes space-time data to construct a space-time graph of an edge device cluster and calculate a space-time correlation matrix, which can accurately capture the space-time dynamic mode of edge data and break through the limitation of traditional methods that ignore space-time heterogeneity; local model parameters of each client are encoded into quantum states, and the model parameter difference degree between any two clients in the edge device cluster is calculated, which further reduces communication overhead and greatly reduces network resource consumption of the edge device; based on a double pruning strategy, the space-time correlation matrix and the parameter difference matrix are used to screen a target client set from multiple clients and perform a dynamic aggregation operation to obtain a global model for scene prediction, which exhibits significant advantages in the edge computing architecture, can effectively handle the space-time heterogeneity between devices, reduce communication overhead, and improve model convergence speed and generalization performance.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of federated learning, and in particular to a prediction method based on spatiotemporal quantum federated learning for edge scenarios. BACKGROUND

[0002] With the rapid development of edge computing technology, perception and decision tasks in the fields of intelligent driving, industrial Internet of Things, smart city, etc. increasingly rely on AI models deployed on terminal devices. In order to utilize widely distributed edge data while protecting data privacy, federated learning as a distributed machine learning paradigm is widely applied in model collaborative training in edge scenarios.

[0003] In related technologies, traditional federated learning algorithms such as FedAvg, FedProx, etc. mainly aggregate model updates of each client by direct averaging or weighted averaging to generate a global model. However, the applicant realizes that the traditional method generally assumes that the data generated by edge devices is independent and identically distributed, which makes it difficult for the trained model to capture real spatiotemporal dynamic patterns, and the generalization performance is severely limited. Secondly, the non-uniform data distribution of the traditional method leads to performance fluctuations of the global model in the spatiotemporal dimension, thereby affecting the convergence stability and final accuracy of the global model. In addition, some existing federated learning schemes that introduce quantum computing need to rely on real quantum hardware, which has a high practical threshold and is difficult to deploy in conventional edge computing architectures. SUMMARY

[0004] Therefore, the present application provides a prediction method based on spatiotemporal quantum federated learning for edge scenarios, which mainly aims to solve the problems that the traditional federated learning has poor generalization performance and unstable global model convergence due to the assumption of independent and identically distributed data, and the existing quantum federated scheme relies on real quantum hardware, which has a high deployment threshold in conventional edge computing architectures.

[0005] According to a first aspect of the present application, a prediction method based on spatiotemporal quantum federated learning for edge scenarios is provided, which comprises:

[0006] obtaining spatiotemporal data of an edge device cluster and local model parameters of each client in the edge device cluster;

[0007] constructing a dual graph structure based on geographical distance and temporal similarity using the spatiotemporal data to obtain a spatiotemporal graph of the edge device cluster, and calculating a spatiotemporal correlation matrix of the edge device cluster;

[0008] encoding the local model parameters of each client into quantum states, and calculating the model parameter difference degree between any two clients in the edge device cluster using the quantum states of multiple clients to obtain a parameter difference matrix of the edge device cluster;

[0009] Based on the double pruning strategy, the spatio-temporal correlation matrix and the parameter difference matrix are used to screen out a target client set from the plurality of clients, and a dynamic aggregation operation is performed on the local model parameters of the target client set to generate a global model;

[0010] If it is detected that the global model meets the model convergence condition, the global model is taken as a target global model of the edge device cluster;

[0011] Real-time spatio-temporal data of the edge device cluster is obtained, and the real-time spatio-temporal data is input into the target global model for prediction to obtain an edge scene prediction result of the edge device cluster.

[0012] By means of the technical solutions described above, the technical solutions provided by the embodiments of the present application have at least the following advantages:

[0013] The prediction method for the edge scene based on the spatio-temporal quantum federated learning provided by the present application obtains the spatio-temporal data of the edge device cluster and the local model parameters of each client in the edge device cluster, constructs a double graph structure based on geographical distance and time similarity by using the spatio-temporal data, obtains the spatio-temporal graph of the edge device cluster, and calculates the spatio-temporal correlation matrix of the edge device cluster. The spatio-temporal data includes geographical coordinates and time series data, which can accurately capture the spatio-temporal dynamic pattern of the edge data, break through the limitations of traditional methods that ignore spatio-temporal heterogeneity, and improve the generalization performance of the global model. Then, the local model parameters of each client are encoded into quantum states, the model parameter difference degree between any two clients in the edge device cluster is calculated by using the quantum states of the plurality of clients, and the parameter difference matrix of the edge device cluster is obtained. The edge devices generally have narrow bandwidth and weak computing power. By quantum state encoding, the high-dimensional model parameters are compressed into low-dimensional quantum state vectors, reducing the parameter transmission amount. At the same time, the parameter difference matrix only transmits the difference degree instead of the complete parameters, further reducing the communication overhead, and greatly reducing the network resource consumption of the edge devices. Then, based on the double pruning strategy, the spatio-temporal correlation matrix and the parameter difference matrix are used to screen out a target client set from the plurality of clients, and a dynamic aggregation operation is performed on the local model parameters of the target client set to generate a global model for scene prediction. The double pruning strategy includes pruning from the parameter abnormal dimension and pruning from the space + time dimension, effectively solving the spatio-temporal non-independent and identically distributed problem of the edge device data distribution. If it is detected that the global model meets the model convergence condition, the global model is taken as a target global model of the edge device cluster, which has significant advantages in the edge computing architecture, can effectively handle the spatio-temporal heterogeneity between devices, reduce the communication overhead, and improve the model convergence speed and generalization performance.

[0014] The above description is only a summary of the technical solutions of the present application. In order to enable the technical means of the present application to be more clearly understood and implemented according to the content of the description, and in order to enable the above and other purposes, characteristics and advantages of the present application to be more apparent and easy to understand, the following specific embodiments of the present application are described. BRIEF DESCRIPTION OF DRAWINGS

[0015] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments. The accompanying drawings are included to provide a description of the preferred embodiments and are not meant to limit the present application. Moreover, the same reference numerals in different drawings represent the same or similar elements. In the drawings:

[0016] Figure 1 A flowchart of a prediction method based on spatiotemporal quantum federated learning for edge scenarios is shown;

[0017] Figure 2 Another flowchart of a prediction method based on spatiotemporal quantum federated learning for edge scenarios is shown;

[0018] Figure 3 A logical architecture diagram of a spatiotemporal quantum federated learning algorithm for edge scenarios is shown. DETAILED DESCRIPTION

[0019] In the description of the present application, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as limiting the present application.

[0020] In addition, the terms "first", "second" are only for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features. In the description of the present application, the meaning of "multiple" is two or more, unless otherwise specifically limited.

[0021] In this application, unless specifically defined and limited otherwise, the terms "mount", "connect", "connection", "fixed", and like terms should be construed as broadly as possible, for example, can be fixed connection, can also be detachable connection, or integrally connected; can be mechanical connection, can also be electrical connection; can be directly connected, can also be indirectly connected through an intermediate medium, can be the communication inside two elements. For those skilled in the art, the specific meaning of the above terms in this application can be understood according to the specific circumstances.

[0022] Exemplary embodiments of the present application will be described in more detail with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that the present application can be more thoroughly understood and the scope of the present application can be accurately conveyed to those skilled in the art.

[0023] The prior art mainly includes three types of classical federated average algorithm (such as FedAvg), adaptive federated optimization scheme and quantum heuristic federated learning, all of which have significant adaptability defects: the classical federated average algorithm only aggregates model updates by weighted average, completely ignores the geographical proximity and time dynamic correlation of edge data, and cannot cope with spatio-temporal heterogeneous non-independent and identically distributed (Non-IID) data, resulting in insufficient model generalization ability; adaptive federated optimization introduces client adaptive learning rate, but cannot solve the problem of unstable global model convergence caused by non-uniform data distribution; quantum heuristic federated learning attempts to use quantum bits to represent parameters, but only stays at the theoretical simulation level, relies on special quantum hardware, and has poor compatibility with conventional edge computing architecture, high deployment threshold and insufficient quantum resource utilization, making it difficult to land and use.

[0024] Traditional federated learning also faces core challenges and technical and engineering dual bottlenecks: at the core level, the resource constraints of spatio-temporal heterogeneous data, narrowband communication pressure, and low computing power and limited storage of edge devices caused by the explosive growth of edge devices make it difficult to adapt to actual scene requirements; at the technical level, the cross-theory of quantum and federated learning is not perfect, the algorithm complexity grows super-linearly due to spatio-temporal modeling, and the convergence under nonlinear quantum operation lacks rigorous demonstration; at the engineering level, the deployment standardization is difficult due to the heterogeneity of edge devices, the monitoring and debugging tools for quantum-classical hybrid systems are missing, and quantum state transmission may introduce privacy leakage risks, and the compliance and operational feasibility in sensitive scenarios such as medical treatment and intelligent driving are insufficient.

[0025] To solve the above problems, the present application aims to break through the basic limitations of traditional federated learning in edge computing scenarios, and build a new generation of edge intelligent basic framework through the deep integration of quantum computing ideas and space-time modeling. Not only does it solve the practical engineering problems faced by current federated learning, but it also explores a feasible path for the application of quantum-classical hybrid computing paradigm in distributed machine learning in the future. Therefore, the present application proposes a spatio-temporal quantum federated learning algorithm (Spatio-Temporal Quantum Amplitude Federated Aggregation, ST-QAFA) for edge computing scenarios. This algorithm is innovative, combining quantum amplitude estimation with the federated learning framework, and effectively solving the problem of spatio-temporal non-independent and identically distributed data distribution of edge devices by introducing a spatio-temporal dynamic pruning mechanism. The breakthrough of its core technology lies in parallel processing of multiple client model updates using the superposition property of quantum states, combined with a spatio-temporal attention weight matrix, which significantly improves the convergence speed and generalization performance of the model. Through rigorous theoretical verification, this algorithm can achieve a 3.8-fold improvement in training efficiency and a 67% reduction in communication overhead in typical edge scenarios, and is fully compatible with existing classical computing architectures. The execution subject of the present application can be a medical service platform, which relies on the computing power of a server to provide services to users. The server can be a standalone server, or it can provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms, among other basic cloud computing servers.

[0026] The present application embodiment provides a prediction method based on spatio-temporal quantum federated learning for edge scenarios, as shown in Figure 1 The method comprises:

[0027] 101, obtaining the spatio-temporal data of the edge device cluster and the local model parameters of each client in the edge device cluster.

[0028] In the embodiments of the present application, the spatio-temporal data and local model parameters of each client in the edge device cluster are collected, wherein the spatio-temporal data includes geographic coordinate data and business time series data. Specifically, the geographic coordinate data refers to the longitude and latitude information of the location of each edge device, and the business time series data records the business activity data of the device within a certain time period, which can reflect the working state and business process of the device. In addition, the local model parameters refer to the model parameters generated by each edge client based on the local data collected by itself after training. In order to ensure the consistency and processability of the data, the format of these model parameters needs to be unified as tensor type, so that efficient data processing and model optimization can be performed in the subsequent data analysis and model fusion process. In this way, the comprehensive collection and effective utilization of the data of each client in the edge device cluster can be realized, providing a solid data foundation for subsequent intelligent analysis and decision-making.

[0029] 102. Construct a dual graph structure based on geographic distance and time similarity using spatio-temporal data to obtain a spatio-temporal graph of the edge device cluster and calculate a spatio-temporal correlation matrix of the edge device cluster.

[0030] In the embodiments of the present application, in the traditional implementation of federated learning, the geographic proximity and time dynamics of edge data are often ignored, which has an important influence on the performance of the model in the edge scenario. In order to solve this problem, a dual graph structure is used to accurately capture the spatio-temporal correlation between data. This structure not only better adapts to the Non-IID data characteristics in the edge scenario, but also quantifies the contribution of each client to the data through multiple spatio-temporal correlation weights in the spatio-temporal correlation matrix. This method avoids the precision loss caused by traditional equal-weight aggregation, thereby significantly improving the generalization performance of the global model. In this way, edge data can be more effectively utilized, and the performance and accuracy of the model in the edge scenario can be improved.

[0031] 103. Encode the local model parameters of each client into quantum states, and use the quantum states of multiple clients to calculate the model parameter difference between any two clients in the edge device cluster to obtain a parameter difference matrix of the edge device cluster.

[0032] In the embodiments of the present application, quantum state encoding is an efficient data compression technique that can compress high-dimensional parameter information into low-dimensional quantum state vectors. This compression method greatly reduces the amount of parameters required during data transmission, especially suitable for edge computing scenarios, because in these scenarios, bandwidth resources are often strictly limited. By compressing high-dimensional data into low-dimensional quantum states, quantum state encoding technology can effectively reduce the bandwidth demand of data transmission, making it possible to process large data on edge devices. In addition, the quantum state encoding process can be completely simulated by a classical GPU cluster, which means that in practical applications, there is no need to rely on dedicated quantum hardware devices. This simulation method not only reduces the deployment cost of quantum state encoding technology, but also improves its accessibility, because GPU clusters are quite popular in the current computing environment. Therefore, quantum state encoding technology can be directly interfaced with mainstream edge devices without additional hardware investment, thereby accelerating the application and promotion of quantum state encoding technology in the field of edge computing. Therefore, quantum state encoding technology effectively reduces the amount of parameter transmission by compressing high-dimensional parameters into low-dimensional quantum state vectors, and is particularly suitable for the narrow bandwidth constraints of edge scenarios. At the same time, using a classical GPU cluster to simulate quantum encoding does not require dedicated quantum hardware, reducing deployment costs, so that quantum state encoding technology can be directly interfaced with mainstream edge devices, providing an efficient data processing solution for edge computing.

[0033] 104、Based on the double pruning strategy, the spatiotemporal correlation matrix and the parameter difference matrix are used to screen out a target client set from multiple clients, and a dynamic aggregation operation is performed on the local model parameters of the target client set to generate a global model.

[0034] In the embodiments of the present application, the first pruning in the double pruning strategy, also known as parameter anomaly pruning, mainly functions to carefully screen the parameter difference matrix, specifically, to select from it those clients whose average difference degree reaches or exceeds a preset threshold. The purpose of this step is to preliminarily filter out those clients whose parameter changes are not significant, thereby reducing the complexity of subsequent processing. Then, the second pruning in the double pruning strategy, i.e., spatiotemporal importance pruning, further screens the remaining clients, and the screening criterion this time is the spatiotemporal correlation weight. Only those clients whose spatiotemporal correlation weight is greater than a preset threshold can be retained. The intersection of the results of the two screenings is taken, and finally obtained is a target client set. The determination of this set can effectively reduce invalid aggregation operations and improve overall efficiency.

[0035] After the target client set is determined, the next step is to renormalize the initial aggregation weights of these target clients to ensure that the weights of each client are reasonable and balanced in the new aggregation process. Subsequently, an element-level weighted average operation is performed on the local model parameters of these clients, and the global model parameters generated in this way can ensure that clients with higher contributions to the model are allocated higher weights. The introduction of this weight allocation mechanism not only more fairly reflects the contributions of each client, but also significantly improves the convergence speed of the global model, thereby accelerating the training process of the entire model and improving the final performance of the model.

[0036] 105. If it is determined through detection that the global model meets the model convergence condition, the global model is taken as the target global model of the edge device cluster.

[0037] In the embodiments of the present application, the L2 norm difference between the global model parameters of the current round and the global model parameters of the last round needs to be calculated in each round of training process. The L2 norm difference is a method for measuring the difference between two vectors, which can reflect the degree of change of model parameters in space. Specifically, the L2 norm difference is obtained by calculating the square sum of the difference of corresponding elements of two vectors, and then taking the square root. If the L2 norm difference is less than the preset convergence threshold, it can be determined that the model has converged. In this case, the target global model is output. However, if the L2 norm difference is greater than the preset convergence threshold, the model has not yet reached the expected performance, and the global model needs to be fed back to the client and the next round of training cycle is started. In this way, the model parameters can be continuously optimized to better adapt to the spatio-temporal dynamic changes of edge data.

[0038] 106. Real-time spatio-temporal data of the edge device cluster is acquired, and the real-time spatio-temporal data is input into the target global model for prediction to obtain an edge scene prediction result of the edge device cluster.

[0039] In the embodiments of the present application, the real-time spatio-temporal data of the edge device cluster includes real-time geographic coordinates of the device (such as real-time latitude and longitude of the vehicle-mounted device, machine position coordinates of the wind turbine) and business time series data (such as real-time road condition perception sequence of intelligent driving, real-time vital sign sequence of ICU patients, real-time operation parameter sequence of wind turbine in wind farm). The real-time spatio-temporal data is input into the target global model, and the model calls the spatio-temporal correlation features learned in the training stage + quantum encoding optimization parameters to analyze and calculate the input data, and generates a prediction result corresponding to the edge scene, for example, the intelligent driving scene outputs the obstacle category (pedestrian / vehicle), distance, and motion trajectory prediction of the real-time road condition; the cross-hospital ICU scene outputs the patient sepsis risk level and risk occurrence time window; the distributed wind farm scene outputs the wind turbine fault type, fault probability, and predicted fault time.

[0040] The embodiment of the application provides a prediction method for an edge scene based on spatio-temporal quantum federated learning, compared with the prior art, the embodiment of the application obtains spatio-temporal data of an edge device cluster and local model parameters of each client in the edge device cluster, constructs a double graph structure based on geographical distance and time similarity by using the spatio-temporal data, obtains a spatio-temporal graph of the edge device cluster, and calculates a spatio-temporal correlation matrix of the edge device cluster, the spatio-temporal data includes geographical coordinates and time series data, can accurately capture the spatio-temporal dynamic mode of edge data, breaks through the limitation of ignoring spatio-temporal heterogeneity in the traditional method, and improves the generalization performance of a global model. Then, the local model parameters of each client are encoded into quantum states, the model parameter difference degree between any two clients in the edge device cluster is calculated by using the quantum states of the plurality of clients, a parameter difference matrix of the edge device cluster is obtained, the edge device generally has the problems of narrow bandwidth and weak computing power, the application compresses high-dimensional model parameters into low-dimensional quantum state vectors by quantum state encoding, reduces the parameter transmission amount, and simultaneously, the parameter difference matrix only transmits the difference degree instead of the complete parameters, further reduces the communication overhead, and the network resource consumption of the edge device is greatly reduced. Then, based on a double pruning strategy, the spatio-temporal correlation matrix and the parameter difference matrix are used to screen a target client set in the plurality of clients, and a dynamic aggregation operation is performed on the local model parameters of the target client set, and a global model is generated, wherein the double pruning strategy includes pruning from a parameter abnormal dimension and pruning from a space+time dimension, and the problem of spatio-temporal non-independent and identically distributed of the data distribution of the edge device is effectively solved. If it is detected that the global model satisfies a model convergence condition, the global model is used as a target global model of the edge device cluster, and the edge computing architecture has significant advantages, can effectively process the spatio-temporal heterogeneity between devices, reduce the communication overhead, and improve the model convergence speed and generalization performance.

[0041] Further, as a refinement and expansion of the above embodiment, in order to completely describe the specific implementation process of the embodiment, the embodiment of the application provides another prediction method for an edge scene based on spatio-temporal quantum federated learning, as shown in Figure 2 The method comprises the following steps.

[0042] 201, obtaining spatio-temporal data of an edge device cluster and local model parameters of each client in the edge device cluster.

[0043] In the embodiments of the present application, before starting to acquire data, a detailed and comprehensive check of the client list is first required to verify its validity and integrity, ensuring that each client in the list strictly has the following four indispensable attributes: device_id (device identifier), spatial_coords (spatial coordinate information), temporal_data (time data), and model_params (model parameters). The completeness and accuracy of these four attributes are the basis for subsequent data processing and analysis. In this stage, an empty undirected graph object is constructed, which will serve as a basic framework for subsequent reception and storage of node and edge information. It should be noted that the choice of undirected graph is based on its flexibility and applicability, which can effectively represent and process the complex relationships between clients. In order to further improve system performance, especially in the face of repeated computing scenarios, a client metadata caching mechanism can be added. This mechanism avoids the overhead of repeatedly acquiring and processing the same data in multiple calculations by caching the metadata information of the client, thereby significantly improving the overall computing efficiency and response speed, not only optimizing resource utilization, but also greatly improving user experience.

[0044] Then the system will traverse all connected client devices one by one, carefully extracting the spatial coordinate information provided by each device, and converting these raw data into efficient numpy array format for subsequent efficient data processing and analysis. During data extraction, strict assurance is made that all acquired coordinate data follows the standard format of (longitude, latitude). Specifically, the longitude value must be strictly limited to the range of [-180, 180], while the latitude value needs to be controlled within the interval of [-90, 90] to ensure the accuracy and consistency of the data. In order to further improve data quality, the system introduces an advanced coordinate correction algorithm. This algorithm can intelligently identify and automatically repair common geographic coordinate format errors, such as longitude and latitude reversal, exceeding the reasonable range, etc., thereby ensuring that the preprocessed coordinate data is more accurate and reliable, and providing a basis for subsequent spatial analysis and application.

[0045] Subsequently, time series data is extracted from each client system one by one, and during the extraction process, the length of each piece of data is carefully determined to meet the preset requirements, and the uniformity of the data format is strictly checked to ensure that all data meet the established standard specifications, thereby laying a solid foundation for subsequent data processing and analysis. After the data extraction is completed, the data is comprehensively scanned and reviewed, and once any abnormal data is found, such as data missing, format errors, or numerical abnormalities, the system will immediately start the log recording mechanism to record the detailed information of the abnormal data, including the time, location, and specific performance of the abnormality. Then the necessary standardization processing is performed on these abnormal data, through a series of data cleaning, conversion, and correction operations, the abnormal data is transformed into valid data that meets the standard, to ensure the quality and integrity of the overall data set. In order to further improve the usability and diversity of the data, the system supports multiple types of time series formats. Specifically, it includes both common equidistant sampling sequences, which collect data at fixed time intervals, and event-driven sequences, which record data based on the occurrence of specific events. By supporting these two main time series formats, the system can provide a solid and rich data foundation for various application scenarios, thereby greatly expanding the possibilities of data analysis and application.

[0046] These detailed spatial coordinate information, together with accurate time series data, i.e., a series of data points arranged in chronological order, are organically integrated to form a complete and multi-dimensional spatio-temporal data system. This spatio-temporal data not only accurately describes the spatial position at a specific time, but also dynamically presents the spatial evolution process over time, providing a solid foundation for data analysis and decision support in various complex scenarios.

[0047] 202、In the spatio-temporal data, spatial coordinate data and time series data are obtained.

[0048] In the embodiments of the present application, spatial coordinate data and time series data are obtained in the spatio-temporal data, wherein the spatial coordinate data includes the geographic position coordinates of each client, and the time series data includes the time series of each client, as shown in the following formula 1:

[0049] Formula 1:

[0050]

[0051]

[0052] wherein, represents the spatial coordinate data, represents the time series data, This represents the geographic coordinates of the i-th client. This represents the longitude of the i-th client. This represents the dimension of the i-th client. , This represents the time series of the i-th client. Represents the time series of the i-th client. Data points, This represents the client's sequence number index. This represents the length of the time series for the i-th client. This represents the time series of the j-th client. This represents the time series of the j-th client. Data points, This represents the length of the time series for the j-th client. Indicates the number of clients.

[0053] 203. Construct a spatial similarity matrix for the edge device cluster using spatial coordinate data, and construct a spatial graph of the edge device cluster using the spatial similarity matrix.

[0054] In this embodiment, the `cdist` function from the scientific computing library `scipy` can be used to accurately calculate the Euclidean distance between any two clients within an edge device cluster by inputting spatial coordinate data. This process generates a matrix containing multiple distance values, each representing the spatial distance between a pair of clients, ensuring computational stability and efficiency, thereby providing basic data support for subsequent data analysis and processing. The calculation formula is shown in Formula 2 below:

[0055]

[0056] in, This represents the distance between the i-th client and the j-th client. This represents the longitude of the i-th client. This represents the dimension of the i-th client. This represents the longitude of the j-th client. This represents the dimension of the j-th client. In the specific technical implementation, the distance matrix used is designed as a... The distance matrix is a square structure with dimensions, where each element precisely represents the distance metric between two data points at the corresponding position. The diagonal elements of the matrix are uniformly set to 0, indicating that the distance of any data point to itself is naturally zero. This representation not only conforms to the actual situation but also effectively simplifies the calculation process, avoiding unnecessary redundant operations. To further enhance the system's running efficiency under high load conditions, especially in scenarios that require processing large-scale client requests, a block calculation strategy can be adopted for performance optimization. Specifically, this strategy divides the entire distance matrix into multiple smaller sub-matrix blocks, and then independently calculates and processes these sub-blocks. This not only significantly reduces the memory resources required for single calculation, but also fully utilizes the multi-core parallel processing capabilities of modern computers, thereby greatly improving memory utilization efficiency while ensuring calculation accuracy, ensuring that the system remains efficient and stable in the face of large-scale client concurrent access.

[0057] Next, the multiple distance values are converted into corresponding similarity values, ensuring that all output similarity values are strictly within the range of (0, 1], and a spatial similarity matrix is constructed. The calculation formula is as follows:

[0058] Formula 3:

[0059] Where, represents the spatial similarity value of the i-th client and the j-th client in the spatial similarity matrix, represents the distance value of the i-th client and the j-th client, represents the distance scaling factor, with a default value of 1. To further optimize the algorithm, a diversified similarity conversion function option mechanism can be systematically introduced and integrated, aiming to fully meet and flexibly adapt to the specific needs of various differentiated application scenarios, thereby effectively improving the universality and practicality of the algorithm. Users can choose the most appropriate similarity conversion function based on actual application background and specific needs, ensuring that the algorithm can perform optimally in different situations, thereby significantly improving work efficiency and result accuracy.

[0060] Subsequently, based on the spatial edge generation decision, a spatial graph is constructed using the spatial similarity matrix. The calculation formula is as follows:

[0061] Formula 4:

[0062] Where, represents the edge set in the spatial graph, represents the edge between the i-th client and the j-th client, represents the spatial similarity value of the i-th client and the j-th client in the similarity matrix, A space similarity threshold value in space edge generation decision, with a default value of 1, used to control the generation of space edges. In specific technical implementation, the system will traverse all client pairs one by one, calculate their spatial similarity, and when the similarity value exceeds the pre-set threshold, the system will add an edge to the graph and record the specific value of the similarity. This edge generation mechanism based on pre-set threshold can effectively control the sparsity of the graph, ensuring that the structure of the graph is neither too dense nor too sparse, thereby optimizing the storage and computing efficiency of the graph. To further enhance the intelligent level of the system, a local density perception mechanism is introduced, which can automatically adjust the threshold strategy when the clients are unevenly distributed in space, so that the threshold can dynamically change according to the changes in local density, thereby more accurately reflecting the actual similarity relationship between clients. In addition, the system also adds support for geographic hash coding, which greatly improves the query efficiency of neighboring clients, making it possible to quickly find neighboring nodes in a large number of clients, further optimizing the overall performance and response speed of the system.

[0063] 204、Based on the improved DTW algorithm, a time similarity matrix of the edge device cluster is constructed using time series data, and a time graph of the edge device cluster is constructed using the time similarity matrix.

[0064] In the embodiments of the present application, the improved DTW algorithm is used to calculate the morphological similarity distance between the time series of any two clients in the time series data, obtaining a plurality of dynamic time warping distances, and the calculation formula is as follows Formula 5:

[0065] Formula 5:

[0066] Wherein, represents the dynamic time warping distance between the time series of the i-th client and the time series of the j-th client, represents the warping path of the time series of the i-th client and the time series of the j-th client, represents a point on the warping path, represents the p-th data point in the time series of the i-th client, represents the qth data point in the time series of the jth client. It is worth noting that the core idea of the improved DTW algorithm mainly revolves around three aspects: path constraint, local weight adjustment, and optimization of the algorithm itself. The introduction of path constraint is to limit the slope range of the regular path, which can avoid the regular path being too steep or flat, thus more accurately reflecting the similarity between sequences. Secondly, the purpose of local weight adjustment is to dynamically adjust the weight according to the importance of different time points, so that the algorithm can pay more attention to the key feature points in the sequence, thereby improving the accuracy of matching. The improved DTW algorithm greatly alleviates the sensitivity of the traditional method to sequence length, making the algorithm applicable to sequence matching problems of different lengths. In terms of technical innovation, the fast DTW approximation algorithm is introduced, which can significantly improve the calculation efficiency while ensuring accuracy, making the algorithm faster in completing sequence matching tasks.

[0067] Then the multiple dynamic time warping distances are normalized to obtain multiple time series difference degrees, and the calculation formula is as follows formula 6:

[0068] Formula 6:

[0069] Wherein, represents the time series difference degree between the time series of the ith client and the time series of the jth client, represents the dynamic time warping distance between the time series of the ith client and the time series of the jth client, represents the length of the time series of the ith client, represents the length of the time series of the jth client. Normalization is an effective technical means that can eliminate the bias problem that may occur when comparing sequences of different lengths. For various sequences of different lengths encountered in practical applications, the system provides a variety of flexible normalization strategies for users to choose from, including but not limited to maximum value-based normalization, minimum-maximum value-based normalization, mean value-based normalization, etc. Users can choose the most suitable normalization method according to specific needs and data characteristics, which can better adapt to data processing needs in different scenarios and further improve the accuracy and efficiency of sequence comparison.

[0070] Then the multiple time series difference degrees are converted into similarity values to construct a time similarity matrix, and the calculation formula is as follows formula 7:

[0071] Formula 7:

[0072] Wherein, represents the time similarity value between the ith client and the jth client in the time similarity matrix, denotes the time series difference degree between the time series of the i-th client and the time series of the j-th client, denotes the decay coefficient, and the default value is 2, which more accurately reflects the timeliness characteristics of the time series.

[0073] Then, a decision of time edge generation is made, a time graph is constructed using the time similarity matrix, and the calculation formula is as follows Formula 8:

[0074] Formula 8:

[0075] wherein, denotes the edge set in the time graph, denotes the edge between the i-th client and the j-th client, denotes the time similarity value of the i-th client and the j-th client in the time similarity matrix, denotes the time similarity threshold in the time edge generation decision, and the default value is 0.15, which is used to control the generation of the time edge.

[0076] 205, spatiotemporal feature fusion is performed using the spatial similarity matrix and the time similarity matrix to obtain a spatiotemporal correlation matrix of the edge device cluster, and the spatial graph and the time graph are fused based on the spatiotemporal correlation matrix of the edge device cluster to obtain a spatiotemporal graph of the edge device cluster.

[0077] In the embodiments of the present application, spatiotemporal feature fusion is performed using the spatial similarity matrix and the time similarity matrix to obtain a spatiotemporal correlation matrix of the edge device cluster, and the calculation formula is as follows Formula 9:

[0078] Formula 9:

[0079] wherein, denotes the spatiotemporal correlation weight of the i-th client and the j-th client in the spatiotemporal correlation matrix of the edge device cluster, denotes the spatial similarity value of the i-th client and the j-th client in the spatial similarity matrix, denotes the spatial weight coefficient, and the default value is 0.5, denotes the time similarity value of the i-th client and the j-th client in the time similarity matrix, denotes the time weight coefficient, and the default value is 0.5. In order to meet different application scenarios and user needs, the system provides a variety of fusion algorithm options, which cover a wide range of choices from traditional statistical methods to cutting-edge machine learning techniques.

[0080] Then, based on the spatio-temporal correlation matrix of the edge device cluster, the spatial graph and the temporal graph are deeply fused to construct a spatio-temporal graph of the edge device cluster. Each node contains complete client metadata information, ensuring the comprehensiveness and accuracy of the data, and the weight of the edge can truly reflect the actual association strength between clients, providing reliable data support for subsequent analysis and decision-making. In addition, the spatio-temporal graph also supports dynamic graph updating function, which can flexibly adapt to the complex scene of dynamic joining and exiting of clients in the federated learning environment, ensuring the real-time and robustness of the system. Through this dynamic updating mechanism, the system can continuously maintain the latest state of the data, so as to better cope with the changing actual application requirements.

[0081] Optionally, in order to further improve the efficiency of the system in calculating the spatio-temporal correlation matrix and generating the spatio-temporal graph, an advanced lazy loading mechanism can be used, which can intelligently delay the loading of unnecessary resources during system operation, effectively avoiding the waste of resources and performance loss caused by premature loading; At the same time, an efficient incremental update algorithm can be implemented, which can accurately identify the changed part of the data and only recompute these changed parts, greatly reducing the invalid consumption of computing resources; In addition, a cache management interface can be designed, which not only supports temporary caching of calculation results, but also supports persistent storage of calculation results, so as to be quickly called in subsequent operations, further improving the system performance.

[0082] Therefore, in order to improve the system performance, the present application establishes a comprehensive performance monitoring system to ensure that the system can maintain stable and efficient performance under various operating environments. For data format exception problems, such as when the system detects that the coordinate format has errors, the default coordinate value will be automatically adopted, and detailed alarm information will be recorded synchronously for subsequent troubleshooting and repair. In terms of fault-tolerant design, common data abnormalities are analyzed in depth, and corresponding classification processing strategies are established. Through the effective implementation of these strategies, the robustness and anti-interference ability of the system are significantly improved. For threshold adjustment strategies, a flexible dynamic threshold adjustment interface can be provided, which can support adaptive threshold calculation based on the current data distribution, ensuring that the threshold setting always matches the actual data characteristics. At the same time, a threshold sensitivity analysis function can be introduced, which can help users understand the impact of threshold changes on system performance, thereby setting optimal parameters more accurately and improving the overall running efficiency of the system. In terms of runtime exception handling, multiple levels of protection measures can be taken. Specifically, for memory overflow risks, a block processing mechanism is implemented, which can decompose large-scale data processing tasks into multiple small blocks and process them one by one, effectively supporting stable operation in large-scale client scenarios. For calculation timeout problems, a strict upper limit for calculation time is set. Once it is detected that the calculation time of a single client exceeds the preset threshold, the system will take immediate measures to prevent it from affecting the overall calculation progress.

[0083] It should be particularly noted that in addition to basic spatiotemporal features, the system also supports the fusion of device types, network states and other types of features. Through multi-dimensional feature fusion, a more comprehensive and detailed client association network is constructed. The system also designs an innovative incremental learning mechanism that can dynamically update the spatiotemporal graph structure based on changes in client behavior patterns during federated learning training, ensuring the timeliness and accuracy of the model. To facilitate user use, the system also provides rich visualization tools that can help users intuitively understand the structural characteristics of the spatiotemporal graph, quickly diagnose potential problems, and optimize parameter configurations, thereby further improving the performance and user experience of the system.

[0084] 206、Encode the local model parameters of each client into a quantum state.

[0085] In the embodiments of the present application, for each client, the local model parameters of the client are subjected to L2 normalization processing to obtain a normalized model parameter vector, and the calculation formula is as follows Formula 10:

[0086] Formula 10:

[0087]

[0088] wherein, local model parameters of the client, the i-th component of the local model parameters, N represents the parameter dimension of the local model parameters, the L2 norm of the local model parameters, the normalized model parameter vector, a very small constant value to prevent division by zero, the default value .

[0089] The normalized model parameter vector is subjected to dimension adaptation processing to obtain a target parameter vector, and the calculation formula is as follows formula 11:

[0090] Formula 11:

[0091] wherein, the target parameter vector, the normalized model parameter vector, N represents the parameter dimension of the local model parameters, k represents the number of quantum bits, which is an integer type, and the default value is 8, corresponding to a quantum state dimension of 256. This parameter determines the encoding accuracy and computational complexity, and needs to be selected according to the parameter scale in actual application. In the field of quantum computing, the number of quantum bits is a crucial technical detail that directly relates to the processing power and accuracy of quantum computers. According to the scale of the model parameters, the number of quantum bits is controlled between 6 and 10. When the model parameter scale is small, the number of quantum bits can be appropriately reduced, which can effectively reduce the consumption of computing resources and avoid unnecessary computational overhead; when facing large-scale parameter scenarios, 8 to 10 quantum bits are used, which can ensure the accuracy of the encoding and thus guarantee the accuracy of the calculation results. In general, the selection of the number of quantum bits needs to be determined according to the specific application scenario and the scale of the model parameters, in order to achieve the purpose of meeting the calculation requirements and saving resources as much as possible.

[0092] Based on the above process, the input parameters are subjected to L2 normalization processing to ensure that the modulus of each parameter vector is 1, thereby meeting the basic requirements of quantum state representation and ensuring the accuracy and stability of the subsequent quantum computing process. If the number of input parameters is less than the dimension of the quantum state, zero padding processing will be performed to ensure that the parameter vector can completely cover all dimensions of the quantum state; if the number of input parameters exceeds the dimension of the quantum state, a truncation operation will be performed to remove the redundant parameters, so as to ensure that the parameter vector and the dimension of the quantum state are completely matched, thereby ensuring the smooth progress of quantum computing and the accuracy of the results.

[0093] Then the target parameter vector is converted into an array format to obtain the quantum state of the client, and the calculation formula is as follows formula 12:

[0094] Formula 12:

[0095]

[0096]

[0097] in, Represents the quantum state of the client. Represents the amplitude of the quantum state. This represents the i-th component in the target parameter vector. The base state represents the quantum state, and k represents the number of qubits. The preprocessed and adjusted parameter data is converted into an array format supported by the NumPy library to facilitate efficient construction and representation of the quantum state. To ensure the accuracy and reliability of the constructed quantum state, the quantum state vector needs to be carefully verified again, re-verifying its normalization properties to ensure its magnitude is 1. Only quantum state vectors that satisfy the normalization condition can conform to the fundamental constraints of quantum mechanics, thus guaranteeing their effectiveness and legitimacy in quantum computing and quantum information processing.

[0098] 207. Calculate the model parameter difference between any two clients in the edge device cluster using the quantum states of multiple clients, and obtain the parameter difference matrix of the edge device cluster.

[0099] In this embodiment of the application, the inner product of any two quantum states in the quantum states of multiple clients is calculated to obtain multiple inner products, and the calculation formula is as follows: Formula 13:

[0100] Formula 13:

[0101] in, Represents the quantum state of the i-th client. Quantum state with the j-th client The inner product of the two parameters is such that a larger inner product value indicates that the parameter distributions are more similar. Represents the quantum state of the i-th client. The complex conjugate of the amplitude in the r-th ground state Represents the quantum state of the j-th client. The amplitude in the r-th ground state.

[0102] The similarity between any two quantum states in the quantum states of multiple clients is calculated using multiple inner products, resulting in multiple parameter similarities. The calculation formula is shown in Formula 14 below:

[0103]

[0104] in, Represents the quantum state of the i-th client. The parameter similarity of the quantum state of the jth client The inner product of the quantum state of the ith client The inner product of the quantum state of the jth client The inner product of the quantum state of the jth client The inner product of the quantum state of the jth client

[0105] The amplitude difference degree of any two quantum states of multiple clients is calculated using multiple parameter similarities to obtain a parameter difference matrix of the edge device cluster, and the calculation formula is as follows formula 15:

[0106]

[0107] Wherein, The amplitude difference degree of the ith client and the jth client in the parameter difference matrix, the range is [0, 1], wherein 0 represents complete consistency, and 1 represents complete orthogonality, The inner product of the quantum state of the ith client The inner product of the quantum state of the jth client The inner product of the quantum state of the jth client During the calculation process, the system will generate a symmetric Difference matrix, and the elements on the diagonal of the matrix are all 0, indicating that the amplitude difference of each client with itself is 0.

[0108] Optionally, the system will calculate the average amplitude difference degree of each client with all other clients, effectively identifying those isolated nodes with large differences from other clients. After completing the calculation of the average difference degree, the system will determine according to the preset threshold. If the average difference degree of a certain client exceeds this threshold, it will be marked as an abnormal state. It is worth noting that this threshold is not fixed, but can be flexibly adjusted according to the actual security level requirements to ensure the accuracy and adaptability of detection.

[0109] The advantage of quantum amplitude encoding lies in its unique dimension expansion capability, which can map parameters from classical space to quantum space. This quantum space has exponentially growing dimensions, which can fully utilize the high-dimensional representation capability of quantum states. This high-dimensional representation capability makes quantum amplitude encoding have a significant advantage in handling complex problems, which can capture subtle features that classical methods cannot capture. In addition, the enhanced sensitivity of quantum amplitude encoding is also one of its important advantages. The inner product calculation of quantum states is more sensitive to the subtle changes of parameters, which can detect small abnormalities that traditional methods cannot find, making quantum amplitude encoding have wide application prospects in anomaly detection, fault diagnosis and other fields. In terms of security, quantum amplitude encoding also has significant advantages. Through anomaly detection, quantum amplitude encoding can effectively identify potential malicious model updates, thereby preventing security threats such as model poisoning. Compared with traditional similarity calculation methods, quantum amplitude encoding has a significant parallel computing advantage when handling high-dimensional parameters, making quantum amplitude encoding more efficient when handling large-scale data, thereby providing a new way of thinking for solving complex problems.

[0110] Parameter encoding time is one of the key indicators of computing efficiency. The optimized system achieves a breakthrough in average processing time of less than 50 milliseconds (based on CPU computing), which not only significantly improves data processing speed, but also enables the system to seamlessly support real-time federated learning scenarios, meeting high real-time requirements. Secondly, considering the scenario of large-scale client concurrent processing, the system optimizes the memory allocation strategy, effectively reducing resource occupation, ensuring efficient operation when multiple tasks are parallel. In addition, algorithm convergence directly affects the training effect of the model. The experimental results on the standard test dataset show that compared with traditional aggregation methods, the convergence speed of the algorithm of the present application is improved by 15-30%, which not only shortens the training time, but also improves the stability and reliability of the model. In addition, the accuracy of anomaly detection is an important indicator to ensure the security of the system. In the simulation attack test, the detection accuracy of the system for various model poisoning attacks reached more than 92%, effectively improving the defense capability of the system and ensuring data security.

[0111] To ensure that the system can run smoothly in different environments, the hardware aspect of the system supports mainstream CPU architectures without relying on special quantum hardware. Implementing quantum algorithms through classical simulation not only reduces the hardware threshold, but also improves the universality and ease of use of the system. In terms of software dependency, the system is based on the widely used PyTorch and NumPy libraries, which maintains good compatibility with mainstream federated learning frameworks, allowing users to easily integrate and use existing resources, reducing development costs and learning curve.

[0112] 208、Calculate the normalized weight of each client using the space-time correlation matrix and the parameter difference matrix.

[0113] In the embodiments of the present application, all clients are traversed, and the preliminary aggregation weight of each client is calculated by using the space-time correlation matrix and the parameter difference matrix, and the calculation formula is as follows formula 16:

[0114] Formula 16:

[0115]

[0116] wherein, denotes the preliminary aggregation weight of the i-th client, denotes the space-time correlation weight of the i-th client in the space-time correlation matrix, denotes the amplitude difference degree of the i-th client and the j-th client in the parameter difference matrix, denotes the attenuation coefficient, and the exponential function ensures that the greater the amplitude difference, the more significant the weight attenuation.

[0117] Then, the preliminary aggregation weight of each client is normalized to obtain the normalized weight of each client, and the sum of all weights is ensured to be 1, and the calculation formula is as follows formula 17:

[0118] Formula 17:

[0119] wherein, denotes the normalized weight of the i-th client, denotes the preliminary aggregation weight of the i-th client, denotes the preliminary aggregation weight of the j-th client, and N denotes the number of clients.

[0120] 209、Through the double pruning strategy, the target client set is screened out from multiple clients by using the normalized weight of each client and the parameter difference matrix.

[0121] In the embodiments of the present application, the first pruning condition and the second pruning condition are obtained in the double pruning strategy. The first pruning, i.e., amplitude threshold pruning, is determined by the condition that the maximum amplitude difference between clients must be less than a pre-set threshold, and the purpose is to retain those clients with relatively consistent parameter distribution and small fluctuation, so as to ensure the consistency and stability of data, effectively exclude those clients obviously deviating from the group and with abnormal amplitude fluctuation, and ensure the data quality of subsequent analysis. The second pruning, i.e., space-time importance pruning, is determined by the condition that the aggregation weight of the client must be greater than a set importance threshold, and the purpose is to filter out those marginal clients with low contribution and small influence on the overall analysis, and further improve the effectiveness of data and the accuracy of analysis. This pruning method can effectively avoid dilution of a small number of important clients by a large number of ordinary clients, ensure that the influence of key data is not weakened, and thus improve the accuracy and reliability of the overall analysis.

[0122] Specifically, the parameter difference matrix is used to screen a client set meeting the first heavy pruning condition from the plurality of clients, and the calculation formula is as follows Formula 18:

[0123] Formula 18:

[0124] Wherein, represents the client set meeting the first heavy pruning condition, represents the amplitude difference degree of the i-th client and the j-th client in the parameter difference matrix, represents the amplitude threshold, and the default value is 0.25.

[0125] The normalized weight of each client is used to screen a client set meeting the second heavy pruning condition from the plurality of clients, and the calculation formula is as follows Formula 19:

[0126] Formula 19:

[0127] Wherein, represents the client set meeting the second heavy pruning condition, represents the normalized weight of the j-th client, represents the importance threshold, and the default value is 0.1.

[0128] Then, the intersection of the client set meeting the first heavy pruning condition and the client set meeting the second heavy pruning condition is taken to ensure that no client meeting the condition is missed, and a target client set is obtained.

[0129] 210, performing a dynamic aggregation operation on the local model parameters of the target client set to generate a global model.

[0130] In the embodiment of the application, the normalized weight of each client in the target client set is normalized to obtain the final aggregation weight of each client in the target client set, and the calculation formula is as follows Formula 20:

[0131] Formula 20:

[0132] Wherein, represents the final aggregation weight of the i-th client in the target client set, represents the normalized weight of the i-th client in the target client set, represents the normalized weight of the j-th client in the target client set, and S represents the target client set.

[0133] Then, a dynamic aggregation operation is performed by using the final aggregation weight of each client in the target client set and the local model parameter of each client in the target client set to obtain the global model, and a calculation formula is as follows Formula 21:

[0134] Formula 21:

[0135] wherein, denotes the model parameter of the global model, S denotes the target client set, denotes the final aggregation weight of the i-th client in the target client set, denotes the local model parameter of the i-th client in the target client set.

[0136] 211. If it is detected that the global model meets the model convergence condition, the global model is taken as the target global model of the edge device cluster.

[0137] In the embodiment of the present application, if it is detected that the global model has completely met the preset model convergence condition, that is, the performance indicators and stability of the model have all reached the expected standard, then in this case, the system will formally designate this global model as the target global model that is commonly followed and used by the edge device cluster. In this way, it is ensured that the edge device cluster can work efficiently and stably in coordination when performing various tasks, further improving the performance and reliability of the overall system.

[0138] 212. If it is detected that the global model does not meet the model convergence condition, the global model is sent to the plurality of clients of the edge device cluster, so that the plurality of clients train the global model as the initial model of the next round of training to obtain the local model parameters of the next round of training of the plurality of clients.

[0139] In the embodiment of the present application, if it is detected that the global model does not meet the model convergence condition, the global model is sent to the plurality of clients of the edge device cluster, so that the plurality of clients train the global model as the initial model of the next round of training to obtain the local model parameters of the next round of training of the plurality of clients, and a calculation formula is as follows Formula 22:

[0140] Formula 22:

[0141] wherein, denotes the model parameter of the global model obtained in the current training round, denotes the model parameter of the global model obtained in the last training round, denotes the convergence threshold in the convergence condition, and the default value is .

[0142] 213. Obtain real-time spatio-temporal data of the edge device cluster, input the real-time spatio-temporal data into the target global model for prediction, and obtain edge scene prediction results of the edge device cluster.

[0143] In the embodiments of the present application, real-time spatio-temporal data of the edge device cluster is obtained, wherein the real-time spatio-temporal data includes spatial data and time series data. The spatial data refers to the real-time geographic coordinates (such as GPS latitude and longitude) and deployment position information of the edge device. The time series data refers to the data stream collected by the device sensor in real time and arranged in time sequence. For example, in the intelligent driving scenario, the time series data can include real-time speed, acceleration, camera video stream, and laser radar point cloud of the vehicle. In the industrial Internet of Things scenario, the time series data can include vibration frequency, temperature, and pressure readings of the device during operation. In the smart city scenario, the time series data can include real-time pictures of monitoring cameras and traffic flow sensor data. Then the real-time spatio-temporal data is input into the target global model for prediction to obtain edge scene prediction results of the edge device cluster. The prediction results are not isolated, but consider the overall state of the device cluster. For example, the overall traffic congestion situation in a certain area is predicted, rather than only the trajectory of a single vehicle. The target global model is based on the efficient design of spatio-temporal graph and quantum coding, and the inference time is extremely short, which can meet the needs of low-delay sensitive scenarios such as intelligent driving obstacle avoidance and medical real-time monitoring, and ensure the timeliness of decision-making. In the model training stage, the spatio-temporal heterogeneity of edge data is fully learned, and the input real-time data and the training data are homologous spatio-temporal data, so the model can adapt to prediction tasks in multiple edge scenarios such as intelligent driving, medical treatment, and industrial operation and maintenance, greatly reducing the development cost of multi-scene adaptation. Because the training stage captures spatio-temporal correlation through double graphs, guarantees parameter consistency through quantum coding, and selects high-quality updates through double pruning, the prediction accuracy of the model on real-time data is significantly higher than that of traditional methods.

[0144] As Figure 3The logic architecture of the edge scene-oriented spatiotemporal quantum federated learning algorithm shown comprises data layer collection, three core modules (a spatiotemporal graph construction module processes edge device clusters, constructs a spatiotemporal graph, and clusters clients; a quantum amplitude coding module performs local model training, gradient quantum coding, and amplitude difference calculation; and a dynamic weight adjustment module aggregates through a central server, dynamically adjusts weights, and completes global model updating), and finally realizes application layer output; the overall workflow is edge device cluster, spatiotemporal graph construction, client clustering, local model training, gradient quantum coding, amplitude difference calculation, central server aggregation, dynamic weight adjustment, and global model updating. The spatiotemporal graph construction module defines edge devices as graph nodes and constructs a dual graph structure based on geographical distance and temporal similarity. An improved DTW algorithm is used to calculate the time series similarity, and then a spatiotemporal correlation matrix is formed, and the client importance weight is output through a graph attention network. This module can effectively capture the spatiotemporal correlation between edge devices and provide an important basis for subsequent federated aggregation. The quantum amplitude coding module maps traditional model parameters to n-dimensional quantum states, estimates the parameter difference with the help of quantum amplitude, establishes a parameter difference matrix, and thus effectively identifies abnormal updates. With the superposition property of quantum states, this module can handle multiple client model updates in parallel, significantly improving computational efficiency. The dynamic weight adjustment module combines spatiotemporal weights and quantum amplitude differences to generate aggregation weights and implement a dual pruning strategy, namely amplitude threshold pruning and spatiotemporal importance pruning, and the global model is updated in a weighted average manner. This system exhibits significant advantages in the edge computing architecture, can effectively cope with the spatiotemporal heterogeneity between devices, reduce communication overhead, and improve model convergence speed and generalization performance.

[0145] Scenario one: urban area air quality collaborative prediction model.

[0146] Multiple air quality monitoring stations are deployed in the city as federated learning clients, and each station continuously collects time series data of local PM2.5, SO2, and other pollutants. The goal is to collaboratively train a global air quality prediction model under the premise of ensuring the privacy of each station's data, effectively addressing the problem of uneven data spatiotemporal distribution caused by meteorological conditions and geographical differences.

[0147] The core steps of data processing are as follows:

[0148] The spatio-temporal graph construction module inputs: the past 24 hours of pollutant concentration time series data of each monitoring site, and the GPS coordinate information of each site. Spatial edge: based on the geographical distance between sites, the Gaussian kernel function is used to convert the distance into a spatial connection weight. Time edge: the similarity of the time series patterns between sites is evaluated using an improved DTW algorithm to form the connection weight in the time dimension. Weight calculation: input the graph structure containing spatio-temporal double correlation into the graph attention network (GAT), and GAT assigns importance weights to each client by adaptive learning, such as the upwind site or the site with high synchronization with the multi-region mode may obtain higher weight.

[0149] The quantum amplitude encoding module inputs: each monitoring site uploads the model parameter updates (such as gradient information) based on local data training. Quantum encoding and difference estimation uses quantum amplitude encoding to map the parameter vectors of each client to quantum states, enabling parallel representation in a log2(N) scale quantum system. Through the quantum amplitude estimation algorithm, the distance between any two quantum states is calculated in parallel to construct an N x N parameter difference matrix to efficiently identify abnormal model updates caused by sensor failure or local pollution events.

[0150] The dynamic weight adjustment module inputs: the spatio-temporal weight output by the spatio-temporal graph construction module and the parameter difference matrix generated by the quantum amplitude encoding module. Aggregated weight synthesis and pruning generation: the final aggregated weight of each client is the product of its spatio-temporal weight and the "reliability factor" based on the parameter difference matrix. Amplitude threshold pruning: set the aggregated weight of the client with high parameter difference (abnormal update) to zero. Spatio-temporal importance pruning: weight pruning is performed on the client with marginal long-term contribution. Global update: the server performs weighted averaging based on the pruned weighted clients to generate a new generation of global prediction model.

[0151] Scenario two: early warning system for sepsis in ICU patients across multiple hospitals.

[0152] Multiple hospital ICUs as federated learning clients continuously collect patient vital sign time series data (such as heart rate, blood pressure, and oxygen saturation). The goal is to jointly train a high-precision sepsis early warning model without sharing patient sensitive information, and to overcome the statistical heterogeneity caused by differences in patient population and treatment between hospitals.

[0153] The core steps of data processing are as follows:

[0154] Temporal-spatial graph construction module input: de-identified physiological parameter time series of each hospital ICU, and inter-hospital geographical location or referral relationship data. Spatial edges: connections are constructed based on geographical distance or referral frequency between hospitals. Temporal edges: the overall shape similarity of patients' vital sign sequences in different hospitals is analyzed using an improved DTW algorithm to identify typical pathological waveforms. Weight calculation: the graph attention network assigns different weights to each hospital client based on data quality and the representativeness of pathological patterns.

[0155] Quantum amplitude encoding module input: each hospital's sepsis early warning model update trained based on local data. Quantum encoding and difference estimation: encode the model update as a quantum state. Quickly construct a difference panoramic map between model updates through quantum amplitude estimation to identify abnormal updates caused by inconsistent labeling standards or training bias.

[0156] Dynamic weight adjustment module input: hospital importance weight and model update difference matrix. Aggregate weight synthesis and pruning: generate the final aggregate weight of each hospital by combining temporal-spatial weight and quantum difference. Implement double pruning strategy: perform amplitude threshold pruning on updates with significant quantum amplitude difference; perform temporal-spatial importance pruning on long-term low-contribution clients. Global update: integrate each hospital's model update through weighted averaging to improve the generalization ability and robustness of the global model.

[0157] Scenario three: predictive maintenance of wind turbine failures in distributed wind farms.

[0158] In wind farms distributed in different regions, each wind turbine continuously collects vibration, temperature, speed, and other running state time series data as a client. The goal is to jointly train a fault prediction model for key components of wind turbines (such as gearboxes), reduce communication costs, and alleviate temporal and spatial heterogeneity caused by differences in wind speed, load, and other working conditions.

[0159] The core steps of data processing are as follows:

[0160] Temporal-spatial graph construction module input: vibration signal time series of each wind turbine and its location coordinates. Spatial edges: connections are constructed based on the physical distance between wind turbines, and adjacent wind turbines may exhibit correlated vibration characteristics due to similar wind conditions. Temporal edges: the similarity of vibration waveforms of different wind turbines is analyzed using an improved DTW algorithm to detect pre-fault patterns. Weight calculation: the graph attention network assigns weights based on the influence of wind turbines in the temporal-spatial graph, such as key position wind turbines or wind turbines that can predict rare failures.

[0161] Quantum amplitude encoding module input: each wind turbine's locally trained fault prediction model update. Quantum encoding and difference estimation: map model parameters to quantum states. Calculate the difference matrix of all wind turbine model updates in parallel through quantum amplitude estimation to identify abnormal updates caused by sensor drift or sudden changes in wind turbine health status.

[0162] The dynamic weight adjustment module inputs: the spatio-temporal weight of the fan and the quantum amplitude difference matrix. The aggregate weight synthesis and pruning: the aggregate weight of each fan is generated by integrating the two aspects of information. The double pruning strategy is executed: the amplitude threshold pruning is implemented for abnormal updates; the spatio-temporal importance pruning is performed for the fans with low contribution degree. Global update: the global fault prediction model is updated by weighted average, realizing the reduction of communication overhead and acceleration of model convergence.

[0163] The beneficial effects of the present application include the following aspects:

[0164] Technical benefit aspects. First, the convergence speed is significantly improved, specifically in the test process of the standard data set, the training rounds required for the model to reach the predetermined target accuracy are greatly reduced, which means a significant reduction in training time, improving the overall research and development efficiency; second, the communication efficiency is significantly optimized, through efficient data compression technology, the amount of data that needs to be transmitted in each training process is effectively reduced, not only reducing the burden on network bandwidth, but also improving the data transmission rate; finally, the generalization ability of the model is significantly enhanced, and when the test set outside the spatio-temporal distribution is verified, the performance degradation of the model is significantly improved, which indicates that the model has stronger adaptability and higher stability under different environments and conditions.

[0165] Economic benefit aspects. First, direct costs are significantly saved, in typical edge computing deployment scenarios, due to technical optimization and improved resource utilization, the total cost of ownership is significantly reduced, saving a large amount of capital investment for the enterprise; second, opportunity costs are effectively transformed, the shortening of the model development cycle enables the enterprise to respond more quickly to market changes and customer needs, seize market opportunities, and improve the competitiveness and market share of the enterprise; finally, incremental revenue is created, in multiple different application scenarios, due to the improvement of model accuracy, more accurate services and higher user satisfaction are brought, thereby creating additional service value and economic benefits.

[0166] Social benefit aspects. First, the privacy protection ability is significantly enhanced, the application of quantum amplitude coding technology naturally protects the security and privacy of the original data, fully complies with relevant legal regulations, and improves user trust; second, energy consumption is significantly reduced, the average energy consumption of edge devices is significantly reduced after technical optimization, which not only reduces operating costs, but also actively responds to the concept of sustainable development, contributing to environmental protection; finally, digital inclusiveness is effectively promoted, through technical innovation, resource-constrained areas can also participate in the collaborative training of high-quality AI models, narrowing the digital divide, promoting the popularization and application of artificial intelligence technology, and promoting the balanced development of society.

[0167] The embodiment of the application provides a prediction method for an edge scene based on spatiotemporal quantum federated learning, compared with the prior art, the embodiment of the application obtains spatiotemporal data of an edge device cluster and local model parameters of each client in the edge device cluster, constructs a double graph structure based on geographical distance and time similarity by using the spatiotemporal data, obtains a spatiotemporal graph of the edge device cluster, and calculates a spatiotemporal correlation matrix of the edge device cluster, the spatiotemporal data includes geographical coordinates and time series data, and can accurately capture the spatiotemporal dynamic mode of edge data, break through the limitation of ignoring spatiotemporal heterogeneity in traditional methods, and improve the generalization performance of a global model. Then, the local model parameters of each client are encoded into quantum states, the model parameter difference degree between any two clients in the edge device cluster is calculated by using the quantum states of the plurality of clients, a parameter difference matrix of the edge device cluster is obtained, and the edge device generally has the problems of narrow bandwidth and weak computing power. The application compresses high-dimensional model parameters into low-dimensional quantum state vectors through quantum state encoding, reduces the parameter transmission amount, at the same time, the parameter difference matrix only transmits the difference degree instead of the complete parameters, further reduces the communication overhead, and greatly reduces the network resource consumption of the edge device. Then, based on a double pruning strategy, the spatiotemporal correlation matrix and the parameter difference matrix are used to screen a target client set in a plurality of clients, and a dynamic aggregation operation is performed on the local model parameters of the target client set, to generate a global model, wherein the double pruning strategy includes pruning from a parameter abnormal dimension and pruning from a space + time dimension, and effectively solves the problem of spatiotemporal non-independent and identically distributed distribution of edge device data. If it is detected that the global model meets the model convergence condition, the global model is taken as a target global model of the edge device cluster, and the edge computing architecture has significant advantages, can effectively handle the spatiotemporal heterogeneity between devices, reduce the communication overhead, and improve the model convergence speed and generalization performance.

[0168] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the application are information and data authorized by the user or authorized by all parties.

[0169] The technical features of the above embodiments can be combined in any way, in order to make the description concise, not all possible combinations of the technical features in the above embodiments are described, however, as long as the combinations of the technical features do not exist contradictions, they should be considered as the scope of the description.

[0170] The above-described embodiments are merely illustrative of several embodiments of the present application, which are described in more detail and in a specific manner, but should not be construed as limiting the scope of the patent of the present application. It should be noted that, for those skilled in the art, several modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

[0171] In an example embodiment, a computer device is also provided, which comprises a bus, a processor, a memory, and a communication interface, and can further comprise an input / output interface and a display device, wherein the communication between various functional units can be completed through the bus. The memory stores a computer program, and the processor is configured to execute the program stored in the memory to execute the prediction method for edge scene based on spatiotemporal quantum federated learning in the above-described embodiments.

[0172] A computer readable storage medium, which stores a computer program, the computer program being executed by a processor to implement the prediction method for edge scene based on spatiotemporal quantum federated learning.

[0173] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by hardware, or can be implemented by means of software and a necessary general hardware platform. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a U disk, a mobile hard disk, etc.), and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application.

[0174] Those skilled in the art can understand that the accompanying drawings are only a schematic diagram of a preferred embodiment, and the modules or processes in the drawings are not necessarily required for implementing the present application.

[0175] Those skilled in the art can understand that the modules in the device in the embodiments can be distributed in the device in the embodiments as described, or can be changed and located in one or more devices different from the embodiments. The modules in the above-described embodiments can be combined into one module, or can be further split into a plurality of sub-modules.

[0176] The above-mentioned serial numbers of the present application are only for description, and do not represent the advantages or disadvantages of the embodiments.

[0177] The above disclosure is only for several specific embodiments of the present application, but the present application is not limited thereto, and any changes that can be thought of by those skilled in the art should fall within the protection scope of the present application.

Claims

1. An edge scene-oriented spatiotemporal quantum federated learning-based prediction method, characterized in that, The method comprises the following steps: acquire spatio-temporal data of an edge device cluster and local model parameters of each client in the edge device cluster; construct a dual graph structure based on geographical distance and time similarity using the spatio-temporal data, obtain a spatio-temporal graph of the edge device cluster, and calculate a spatio-temporal correlation matrix of the edge device cluster; encode the local model parameters of each client into quantum states, and calculate the model parameter difference between any two clients in the edge device cluster using the quantum states of multiple clients to obtain a parameter difference matrix of the edge device cluster; based on a dual pruning strategy, filter out a target client set from the multiple clients using the spatio-temporal correlation matrix and the parameter difference matrix, and perform a dynamic aggregation operation on the local model parameters of the target client set to generate a global model, including calculating the preliminary aggregation weight of each client using the spatio-temporal correlation matrix and the parameter difference matrix, wherein, represents a preliminary aggregation weight of the i-th client, represents a spatiotemporal correlation weight of the i-th client in the spatiotemporal correlation matrix, represents an amplitude difference degree of the i-th client and the j-th client in the parameter difference matrix, represents a decay coefficient; normalizing the preliminary aggregation weight of each client to obtain a normalized weight of each client, wherein, denotes the normalized weight of the i-th client, denotes the preliminary aggregated weight of the i-th client, denotes the preliminary aggregated weight of the j-th client, and N denotes the number of clients; by means of the double pruning strategy, the target client set is screened out from the plurality of clients by using the normalized weight of each client and the parameter difference matrix. if it is detected that the global model meets the model convergence condition, the global model is taken as a target global model of the edge device cluster; acquire real-time spatio-temporal data of the edge device cluster, input the real-time spatio-temporal data into the target global model for prediction, and obtain an edge scene prediction result of the edge device cluster.

2. The method of claim 1, wherein, The method comprises the following steps: acquire spatial coordinate data and time series data in the spatio-temporal data, the spatial coordinate data comprising geographical position coordinates of each client, and the time series data comprising time series of each client, wherein, in, This represents the spatial coordinate data. This refers to the time series data. This represents the geographic coordinates of the i-th client. This represents the longitude of the i-th client. This represents the dimension of the i-th client. , This represents the time series of the i-th client. Represents the time series of the i-th client. Data points, This represents the client's sequence number index. This represents the length of the time series for the i-th client. This represents the time series of the j-th client. This represents the time series of the j-th client. Data points, This represents the length of the time series for the j-th client. Indicates the number of clients; construct a spatial similarity matrix of the edge device cluster using the spatial coordinate data, and construct a spatial graph of the edge device cluster using the spatial similarity matrix; construct a time similarity matrix of the edge device cluster using the time series data based on an improved DTW algorithm, and construct a time graph of the edge device cluster using the time similarity matrix; perform spatio-temporal feature fusion using the spatial similarity matrix and the time similarity matrix to obtain a spatio-temporal correlation matrix of the edge device cluster, wherein, represents a spatio-temporal correlation weight of an i-th client and a j-th client in a spatio-temporal correlation matrix of the edge device cluster, represents a spatial similarity value of an i-th client and a j-th client in the spatial similarity matrix, represents a spatial weight coefficient, represents a temporal similarity value of an i-th client and a j-th client in the temporal similarity matrix, represents a temporal weight coefficient; fuse the spatial graph and the time graph based on the spatio-temporal correlation matrix of the edge device cluster to obtain a spatio-temporal graph of the edge device cluster.

3. The method of claim 2, wherein, The method comprises the following steps: calculate the Euclidean distance between any two clients in the edge device cluster using the spatial coordinate data to obtain multiple distance values, wherein, represents a distance value of the i-th client to the j-th client, represents a longitude of the i-th client, represents a latitude of the i-th client, represents a longitude of the j-th client, represents a latitude of the j-th client; convert the multiple distance values into similarity values to construct a spatial similarity matrix, wherein, denotes a spatial similarity value of the i-th client with the j-th client in the spatial similarity matrix, denotes a distance value of the i-th client with the j-th client, denotes a distance scaling factor; construct the spatial graph using the spatial similarity matrix based on a spatial edge generation decision, wherein, represents a set of edges in the spatial graph, represents an edge between the i-th client and the j-th client, represents a spatial similarity value of the i-th client and the j-th client in the similarity matrix, represents a spatial similarity threshold in the spatial edge generation decision.

4. The method of claim 3, wherein, The improved DTW algorithm is used to construct a time similarity matrix of the edge device cluster based on the time series data, and a time graph of the edge device cluster is constructed based on the time similarity matrix, including: The improved DTW algorithm is used to calculate the shape similarity distance between the time series of any two clients in the time series data, obtaining a plurality of dynamic time warping distances, wherein, denotes the dynamic time warping distance between the time series of the i-th client and the time series of the j-th client, denotes the warping path of the time series of the i-th client and the time series of the j-th client, denotes a point on the warping path, denotes the p-th data point in the time series of the i-th client, denotes the q-th data point in the time series of the j-th client; The plurality of dynamic time warping distances are normalized to obtain a plurality of time series difference degrees, wherein, denotes the time series difference between the time series of the i-th client and the time series of the j-th client, denotes the dynamic time warping distance between the time series of the i-th client and the time series of the j-th client, denotes the length of the time series of the i-th client, denotes the length of the time series of the j-th client; The plurality of time series difference degrees are converted into similarity values to construct a time similarity matrix, wherein, denotes a time similarity value of the ith client and the jth client in the time similarity matrix, denotes a time series difference between the time series of the ith client and the time series of the jth client, denotes a decay coefficient; Based on the time edge generation decision, the time graph is constructed based on the time similarity matrix, wherein, denotes a set of edges in the time graph, denotes an edge between the i-th client and the j-th client, denotes a time similarity value of the i-th client and the j-th client in the time similarity matrix, denotes a time similarity threshold in the time edge generation decision.

5. The method of claim 1, wherein, The local model parameters of each client are encoded into a quantum state, including: For each client, the local model parameters of the client are subjected to L2 normalization processing to obtain a normalized model parameter vector, wherein, denotes the local model parameters of the client, denotes the i-th component of the local model parameters, N denotes the parameter dimension of the local model parameters, denotes the L2 norm of the local model parameters, denotes the normalized model parameter vector, denotes a very small constant value; The normalized model parameter vector is subjected to dimension adaptation processing to obtain a target parameter vector, wherein, denotes the target parameter vector, denotes the normalized model parameter vector, k denotes the number of qubits, and N denotes the parameter dimension of the local model parameters; The target parameter vector is converted into an array format to obtain the quantum state of the client, wherein, denotes a quantum state of the client, denotes a quantum state amplitude, , denotes the i-th component of the target parameter vector, denotes a basis state of the quantum state, k denotes the number of quantum bits.

6. The method of claim 1, wherein, The quantum state of each client is used to calculate the model parameter difference degree between any two clients in the edge device cluster to obtain a parameter difference matrix of the edge device cluster, including: The inner product of any two quantum states in the plurality of quantum states of the plurality of clients is calculated to obtain a plurality of inner products, wherein denotes the quantum state of the i-th client denotes the quantum state of the j-th client the inner product of the i-th client's quantum state denotes the quantum state of the i-th client the complex conjugate of the amplitude of the i-th client's quantum state in the r-th basis state denotes the quantum state of the j-th client the amplitude of the j-th client's quantum state in the r-th basis state The similarity between any two quantum states in the plurality of quantum states of the plurality of clients is calculated using the plurality of inner products to obtain a plurality of parameter similarities, wherein, denotes the quantum state of the i-th client denotes the quantum state of the j-th client the parameter similarity of denotes the quantum state of the i-th client denotes the quantum state of the j-th client the inner product of The amplitude difference degree of any two quantum states in the plurality of quantum states of the plurality of clients is calculated using the plurality of parameter similarities to obtain the parameter difference matrix of the edge device cluster, wherein, denotes the amplitude difference degree of the i-th client with the j-th client in the parameter difference matrix, denotes the parameter similarity of the quantum state of the i-th client with the quantum state of the j-th client, denotes the inner product of the quantum state of the i-th client with the quantum state of the j-th client.

7. The method of claim 1, wherein, The double pruning strategy is used to screen the target client set from the plurality of clients using the normalized weight of each client and the parameter difference matrix, including: In the double pruning strategy, a first pruning condition and a second pruning condition are obtained; The parameter difference matrix is used to screen a client set that meets the first pruning condition from the plurality of clients, wherein, denotes a set of clients satisfying the first heavy pruning condition, denotes the amplitude difference degree of the i-th client and the j-th client in the parameter difference matrix, denotes an amplitude threshold value; The normalized weight of each client is used to screen a client set that meets the second pruning condition from the plurality of clients, wherein, denotes the set of clients that meet the second heavy pruning condition, denotes the normalized weight of the jth client, denotes the importance threshold value; The intersection of the client set that meets the first pruning condition and the client set that meets the second pruning condition is taken to obtain the target client set.

8. The method of claim 1, wherein, The dynamic aggregation operation is performed on the local model parameters of the target client set to generate a global model, including: The normalized weight of each client in the target client set is normalized to obtain the final aggregation weight of each client in the target client set, wherein, denotes the final aggregated weight of the i-th client in the target client set, denotes the normalized weight of the i-th client in the target client set, denotes the normalized weight of the j-th client in the target client set, and S denotes the target client set. The final aggregation weight of each client in the target client set and the local model parameters of each client in the target client set are used to perform a dynamic aggregation operation to obtain the global model, wherein, denote model parameters of the global model, S denotes the set of target clients, denote the final aggregated weight of the i-th client in the set of target clients, denote local model parameters of the i-th client in the set of target clients.

9. The method of claim 1, wherein, The method further includes: if it is determined that the global model does not satisfy the model convergence condition, sending the global model to a plurality of clients of the edge device cluster to enable the plurality of clients to train the global model as an initial model for a next round of training to obtain local model parameters of the plurality of clients for the next round of training, wherein, denotes model parameters of the global model obtained in the current training round, denotes model parameters of the global model obtained in the previous training round, denotes a convergence threshold in the convergence condition.

Citation Information

Patent Citations

  • Data anomaly detection method and system based on quantum graph federated learning

    CN118964626A

  • Edge node heterogeneous network heterogeneous platform access management method and system

    CN120090889A