Adaptive federated family load prediction method and system driven by air computing

By employing an over-the-air computing-driven adaptive federated household load forecasting method, which utilizes logarithmic transformation and peak resampling, combined with Q-learning and the superposition characteristics of wireless multiple access channels, the problem of insufficient peak forecasting and low communication efficiency in household load forecasting is solved, achieving efficient and robust load forecasting results.

CN121749129APending Publication Date: 2026-03-27SUZHOU UNIV OF SCI & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing federal load forecasting methods suffer from insufficient peak forecasting sensitivity, low communication efficiency, and insufficient robustness in unstable cooperative network environments when dealing with household load scenarios.

Method used

An adaptive federal household load forecasting method driven by over-the-air computing is adopted. Standardized training samples are formed through logarithmic transformation and peak resampling. The action value function is updated online by combining Q-learning algorithm to generate weighted temperatures. Parallel weighted aggregation is performed using the linear superposition characteristics of wireless multiple access channels to ensure the accuracy and robustness of the model in the peak region.

Benefits of technology

It significantly improves peak tracking performance, enhances model convergence speed and communication resource utilization, provides a feasible path for large-scale terminal deployment, and ensures prediction accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121749129A_ABST
    Figure CN121749129A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of family load prediction, and relates to a self-adaptive federated family load prediction method driven by air computation, which comprises the following steps of: performing logarithmic transformation and normalization on family load data, calculating a quantile threshold by using peak-valley distribution, dividing peak section and valley section samples, and forming a training sample through peak section resampling; deploying a prediction model by the client, taking the training sample as input, performing local training on the prediction model by adopting a combined loss function, and outputting a prediction value and a prediction residual error; a server collects an upper quantile predicted value, a link quality index and a sample size to construct a state vector, updates an action value function through Q-learning to generate a weighted temperature, determines an aggregation weight, pre-scales a local model update amount, realizes parallel weighted aggregation through a linear superposition characteristic, receives a superposed signal to recover digital domain averaging, and finally performs data processing. And completing global model updating in combination with the global step length and the feasible region projection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of household load forecasting technology, specifically relating to an over-the-air computing-driven adaptive federal household load forecasting method and system. Background Technology

[0002] my country's power system is rapidly evolving from a traditional model dominated by centralized power sources and characterized by unidirectional transmission to a complex system primarily based on renewable energy sources, with multi-level coordination among power generation, grid, load, and storage. Under the constraints of "dual carbon" (carbon dioxide, carbon emissions, and carbon sequestration), the power distribution system faces challenges such as a high proportion of renewable energy, increased random load fluctuations, and widespread access to diverse and flexible resources. The mission and positioning of the power distribution system's operation and regulation are undergoing profound changes. With the large-scale deployment of smart meters and various smart home devices, residential electricity consumption behavior is being perceived with fine granularity, and household load exhibits typical characteristics such as high sampling frequency, diverse behavioral models, complex environmental scenarios, and a "high-valley, low-peak" pattern.

[0003] In the centralized learning paradigm, numerous studies have incorporated deep learning structures into short-term residential load forecasting. By combining deep models, generative models, and attention mechanisms, these methods characterize complex electricity consumption behaviors and temporal dependencies, achieving good predictive results with centralized data aggregation. However, such methods typically rely on massive amounts of centralized training data, which not only makes them susceptible to data silos and privacy concerns in practical deployments, but also limits their scalability and generalization capabilities in the context of a large number of terminals and highly heterogeneous data distribution in home scenarios.

[0004] To mitigate the risks of "data silos" and privacy breaches, federated learning has been increasingly introduced into short-term residential load forecasting. Related research deploys local models on the household side and aggregates sampled parameters to achieve multi-household collaborative modeling. Further, mechanisms such as differential privacy, secure aggregation, or homomorphic encryption are layered on top, and the trade-offs between privacy budget, noise intensity, and prediction performance are systematically analyzed. Other works, from a distributed learning perspective, combine layered security protocols with federated optimization to construct a privacy-preserving distributed learning framework for short-term residential load forecasting. Review studies systematically summarize the research progress and open issues of federated learning from multiple aspects, including optimization algorithms, system architecture, and privacy mechanisms, pointing out that achieving a reasonable balance between communication costs, personalized modeling, and privacy protection remains a significant challenge for federated learning. In the domestic context, significant progress has been made in federated forecasting models for integrated energy and residential load: existing works combine federated learning with LSTM for integrated energy multi-variable load forecasting under limited data conditions, and have made beneficial explorations in model structure and privacy protection protocols; literature (Yuan Yu, Yang Chao, Zheng Weiming, Lin Junpeng, Chen Xin,. Federated learning regional power short-term load forecasting model based on Bi-SRNN [J]. Power Grid and Clean Energy, 2023, (10): 45-55. and Zhu Songyang, Zhang Ge, Jia Yujing, Bai Xiaoqing,. Federated learning residential load forecasting model based on long short-term memory network model) Load Prediction [J]. Modern Electric Power, 2025, (01): 129-136.) proposes a federated prediction model for residential load based on Bi-SRNN or LSTM, which effectively characterizes the temporal and spatial correlation of terminal loads in the distribution network; A federated learning framework for power metering systems [Zheng Kaihong, Xiao Yong, Wang Xin, Chen Wei, A federated learning framework for power metering systems [J]. Proceedings of the CSEE, 2020, (S1): 122-133.] demonstrates the feasibility and advantages of the federated learning framework for power metering systems in the context of the power Internet of Things for metering data analysis. Building upon this foundation, the survey on federated learning for edge intelligence (Zhang Xueqing, Liu Yanwei, Liu Jinxia, ​​Han Yanni. A survey on federated learning for edge intelligence [J]. Computer Research and Development, 2023, 60(6):1276–1295) further systematically reviewed the key technologies of federated learning from the perspectives of client selection, communication compression, asynchronous aggregation, and personalized modeling. This demonstrates that there is a relatively complete technical foundation for deploying federated load prediction in edge scenarios, and also provides an overall background and theoretical support for the federated learning framework for edge-side household load in this invention.

[0005] With the development of edge intelligence and 5G / 6G wireless communication technologies, how to efficiently complete federated aggregation under conditions of bandwidth constraints, multi-terminal access, and random channel fading has become one of the key factors restricting the engineering implementation of load forecasting models. Over-the-Air Computation (Air Comp) utilizes the waveform superposition characteristics of multiple access channels to directly realize the parallel aggregation of multi-terminal signals in the analog domain, and is widely regarded as an effective way to alleviate the uplink communication bottleneck in federated learning. In the literature (K. Yang, T. Jiang, Y. Shi and Z. Ding, "Federated Learning via Over-the-Air Computation," in IEEE Transactions on Wireless Communications, vol. 19, no. 3, pp. 2022-2035, March 2020, doi: 10.1109 / TWC.2019.2961673.), Yang et al. constructed a typical Air Comp-FL framework around gradient aggregation. By jointly designing terminal selection and beamforming, they significantly reduced uplink latency and bandwidth ratio while ensuring convergence performance. Building upon this, the literature [D. Zhang, M. Xiao, Z. Pang, L. Wang and HV Poor, "IRS Assisted Federated Learning: ABroadband Over-the-Air Aggregation Approach," in IEEE Transactions on Wireless Communications, vol. 23, no. 5, pp. 4069-4082, May 2024,] further developed this framework. [doi:10.1109 / TWC.2023.3313968.] Federated learning is deployed in IRS-assisted UAV communication systems. By jointly optimizing parameters such as IRS phase, noise suppression factor, terminal transmit power, and UAV trajectory, the worst-case mean square error of the aggregation process is minimized, thereby improving the reliability and coverage of over-the-air aggregation.For general wireless network environments, the paper [HYOksuz, F. Molinari, H. Sprekeler and J. Raisch, "Federated Learning in Wireless Networks via Over-the-Air Computations," 2023 62nd IEEE Conference on Decision and Control (CDC), Singapore, Singapore, 2023, pp. 4379-4386, doi: 10.1109 / CDC49753.2023.10384001.] proposes an Air Comp federation scheme that does not require explicit channel reconstruction, performing parameter aggregation directly at the physical layer to reduce channel estimation and coding overhead. To adapt to large-scale cellular edge systems, the paper [Qiao L, Gao Z, Mashhadi MB, Gündüz D. Massive digital over-the-aircomputation for communication-efficient federated edge learning[J]. IEEE Journal on Selected Areas in Communications, 2024] designs Massive Digital... The AirComp mechanism enables a large number of terminals to achieve low-distortion parameter aggregation in the digital domain, thereby improving the communication efficiency and convergence speed of federated edge learning.References [Xiao B, Yu X, Ni W, Wang X, Poor H V. Over-the-air-federated learning: Status quo, open challenges, and future directions[J]. Fundamental Research, 2025, 5(4):1710–1724., Chen Z, Cai L, Zhao C. Convergence time of average consensus with heterogeneous random link failures[J]. Automatica, 2021, 127:109496] summarize the key technologies and open issues of Over-the-Air Federated Learning from a systems perspective, pointing out that noise amplification, channel uncertainty, and terminal heterogeneity are challenges that AirComp-FL must focus on when moving towards complex real-world scenarios. Overall, existing Air Comp-FL research mainly focuses on general tasks such as image recognition as the main verification object, with relatively insufficient attention paid to household load scenarios with "many valleys and few peaks" characteristics and high sensitivity to peak prediction. This provides research space for this invention to introduce Air Comp and adaptive aggregation mechanisms into federal household load prediction.

[0006] At the security and privacy level, federated learning also faces a variety of potential threats such as model poisoning, backdoor attacks, gradient leakage, and member inference. Related research systematically sorts out the above-mentioned attack methods and their impact mechanisms, and summarizes typical defense technologies such as differential privacy, secure multi-party computation, robust aggregation, and trusted execution environment, providing a methodological reference for the secure implementation of federated learning in highly sensitive businesses such as energy and electricity [Qiu Xiaohui, Yang Bo, Zhao Mengchen, Hu Shiyang, Sun Pu. Research on security defense and privacy protection technology of federated learning [J]. Computer Applications Research, 2022, 39(11):3220–3231, Gu Yuhao, Bai Yuebin. Research progress on security and privacy of federated learning models [J]. Journal of Software, 2023, 34(6):2833–2864].

[0007] The above research reveals two main issues: First, existing federated load forecasting efforts primarily focus on minimizing overall error, neglecting the forecasting sensitivity of the peak segments in the "valley-heavy, sparse-peak" household load curve, which can lead to situations where the overall error is small but the peak deviation is large. Second, current Air Comp-FL research focuses more on communication efficiency and average convergence performance, failing to adequately consider the peak-driven time-series characteristics of household loads, leaving gaps in areas such as peak resampling, task importance modeling, and adaptive aggregation weight design. Furthermore, for unstable cooperative network environments such as random terminal disconnections and link interruptions, there is a lack of systematic research on federated forecasting mechanisms within the Air Comp aggregation framework that balance robustness and efficiency for specific business scenarios. Summary of the Invention

[0008] The purpose of this invention is to overcome the shortcomings of the prior art and provide an adaptive federal household load forecasting method and system driven by over-the-air computing.

[0009] To achieve the objectives of this invention, the following technical solutions are adopted.

[0010] An over-the-air computing-driven adaptive federal household load forecasting method includes the following steps:

[0011] Logarithmic transformation and normalization preprocessing are performed on the original household load data. Based on the peak and valley distribution predicted by the sliding window, quantile thresholds are calculated to divide the peak and valley samples. Standardized training samples are formed by resampling the peak segments.

[0012] Deploy the prediction model on each client, use standardized training samples as input, and train the prediction model locally using a combined loss function. Do not upload the original household load data. Output the predicted value and prediction residual.

[0013] The server collects the upper quantile prediction values, link quality indicators and sample size of online clients to construct a state vector. It updates the action value function online through the Q-learning algorithm, generates a weighted temperature, and determines the set of clients participating in the aggregation and the aggregation weight of each client based on the state vector and the weighted temperature.

[0014] The client prescales the local model update based on the aggregation weights and uses the linear superposition characteristics of the wireless multiple access channel to achieve parallel weighted aggregation through over-the-air computation. After receiving the superimposed signal, the server restores the digital domain average by normalizing the ratio and applies a uniform global step size and feasible region projection to complete the global model update. When the channel conditions do not meet the requirements of over-the-air computation, it falls back to digital domain weighted average aggregation.

[0015] Furthermore, the specific preprocessing steps are as follows: first perform a logarithmic mapping, then normalize to... Assume the original household load is ,but:

[0016] ;

[0017] In the formula: The logarithmically mapped value of the original household load. This is the final output value obtained after normalization following the logarithmic transformation. and These are the minimum and maximum values ​​in the training set, respectively.

[0018] Furthermore, the specific implementation process of the peak segment resampling is as follows:

[0019] The threshold is given by the target distribution of the training set:

[0020] , ;

[0021] Define the sampling weights as ,in ;

[0022] Weighted sampling with replacement is used, with a sampling probability of... .

[0023] Furthermore, the prediction model includes a bidirectional BiLSTM, an attention mechanism, a normalization and multilayer perceptron module, and a residual correction module; wherein:

[0024] Bidirectional BiLSTM is used to encode the input sequence of a sliding window and output a bidirectional hidden state sequence.

[0025] The attention mechanism calculates the attention weights for each time dimension using dot product scoring and SoftMax, and then performs a weighted summation of the hidden state sequence to obtain the context vector.

[0026] The normalization and multilayer perceptron module normalizes and performs nonlinear transformations on the context vector, and outputs the model's preliminary prediction value after Sigmoid activation.

[0027] The residual correction module adaptively compromises between the model's initial predictions and the most recent observation baseline using learnable gating, and outputs the model's predicted values.

[0028] Furthermore, the expression for the combined loss function is:

[0029] ;

[0030] In the formula:

[0031] The mean square error is a two-sided weighted average. , ,in: The mixing coefficient, Control the degree of emphasis on peak values. Here, is the numerical stability constant, N is the number of samples used for calculation in this round, and t is the time index, corresponding to the end time of the current sliding window. Let be the total weight of the t-th sample in the baseline mean squared error. For valley segment weighting function, For peak segment weighting function, Let be the normalized predicted load corresponding to the t-th sample;

[0032] The mean square error of peak re-penalty. , ;

[0033] To underestimate the punishment, ;

[0034] For quantile loss, .

[0035] Furthermore, the process for determining the aggregation weights is as follows:

[0036] The server first collects information from online clients, organized as follows:

[0037] , ;

[0038] in: Let be the set of clients that are online in round t; The upper quantile is used for prediction to characterize the strength of short-term price increases; This is a link quality metric, given by the normalized instantaneous signal-to-noise ratio; This represents the effective sample size used for training in this round; based on this state vector, a comprehensive index is constructed that simultaneously characterizes risk, trend, link availability, and sample representativeness. The participation priority of client i in round t:

[0039] ;

[0040] To meet the size and availability constraints of each round, the server is based on... With link quality Generate a candidate set:

[0041] ;

[0042] in: , These are thresholds for urgency and link quality, respectively.

[0043] After obtaining the candidate set Then, Q-learning without function approximation is introduced to learn weighted temperatures online. To adaptively adjust the concentration of weight allocation; actions are taken from the discrete set A according to... -Greedy rule selection; immediate reward:

[0044] ;

[0045] in: Let be the combined loss on the validation set after the t-th round of aggregation. To normalize communication overhead, The balancing coefficient; a larger reward indicates a more significant reduction in validation loss with less communication cost; the Q-learning standard is used to update the rules:

[0046] ;

[0047] in: For learning rate, The discount factor is used; through multiple rounds of interaction, Q-learning gradually learns a temperature selection strategy that matches the task objective.

[0048] Given the current action, the urgency will be adjusted by temperature. Mapped to normalized importance weights:

[0049] ;

[0050] The larger the value, the more concentrated the weight becomes on key clients; The smaller the value, the more uniform the distribution; thus, the training layer's preference for "peak segments being more important and underestimation being more expensive" is systematically transmitted to the two stages of decision-making and weight allocation; subsequently, the server... broadcast Each client optimizes its local data using the combined loss function to obtain... To facilitate parallel aggregation, the update amount in this round is recorded as follows:

[0051] ;

[0052] A global update can directly target... Weighted averages can also be used for... For weighted averages, the two are numerically equivalent, differing only in that they involve a shift. .

[0053] Furthermore, the specific implementation process of the global model update is as follows:

[0054] Let the equivalent positive real gain of the uplink channel after phase compensation be... Define the equivalent aggregate weight on the end side. ;

[0055] because The time-varying channel gain coefficient is unknown and difficult to measure accurately; at the client end, two signals are transmitted simultaneously on the same time-frequency resource: one is the vector to be aggregated. The other path is to set a weighted scalar. Suppose the server receives the following response:

[0056] , A1;

[0057] When synchronization is good and SNR is high, ratio normalization is performed to eliminate unknown common scales, resulting in:

[0058] , A2;

[0059] Then apply a uniform server-wide step size. and by function Project to feasible ,get:

[0060] A3;

[0061] Equation (A3) is numerically equivalent to weighting the update amounts of each client in the numeric domain. Weighted averaging simply involves moving the weighted summation operation to the WMAC channel in the analog domain. If the receiver observes that the WMAC condition is not met, the server abandons the current round of over-the-air aggregation and reverts to the traditional digital domain weighted averaging form.

[0062] ;

[0063] Thus, a parallel weighted aggregation is achieved using over-the-air computing: the server provides the importance weights of each terminal, the terminal multiplies them by the link gain and transmits them on the same time-frequency resource, realizing weighted multiplication followed by summation; the server normalizes the ratio of the received vector sum to the weighted sum, applies a uniform step size and projects according to constraints to complete the global update.

[0064] An over-the-air computing-driven adaptive federated household load forecasting system includes a data processing layer, a client modeling layer, a communication and over-the-air aggregation layer, and a server-side aggregation and scheduling layer; wherein:

[0065] The data processing layer, deployed on the client, performs logarithmic transformation and normalization preprocessing on the raw household load data. It calculates quantile thresholds based on the peak-valley distribution predicted by the sliding window to divide peak and valley samples. It then forms standardized training samples through peak resampling.

[0066] The client-side modeling layer is deployed on the client side. It models the prediction model on each client, takes the training samples as input, uses the combined loss function to train the prediction model locally, forms the local model, and outputs the prediction value and prediction residual.

[0067] In the communication and air aggregation layer, the server collects the upper quantile prediction values, link quality indicators and sample size of online clients to construct a state vector. It updates the action value function online through the Q-learning algorithm, generates a weighted temperature, and determines the set of clients participating in the aggregation and the aggregation weight of each client based on the state vector and the weighted temperature.

[0068] The server-side aggregation and scheduling layer, deployed on the server, prescales the local model update based on the aggregation weight. It utilizes the linear superposition characteristics of the wireless multiple access channel to achieve parallel weighted aggregation through over-the-air computation. After receiving the superposition signal, the server restores the digital domain average by normalizing the ratio and applies a uniform global step size and feasible region projection to complete the global model update.

[0069] Furthermore, in the prediction model, the number of hidden units in the bidirectional BiLSTM layer is 128, and the number of layers is 3; the attention mechanism adopts dot product attention; the normalization and multilayer perceptron module includes a batch normalized BN layer and a fully connected layer with ReLU activation function; the learnable gating of the residual correction module is generated through the Sigmoid activation function.

[0070] Furthermore, the channel condition judgment criteria supported by the communication and air aggregation layer are as follows: when the signal-to-noise ratio of the received signal is higher than a preset threshold and the synchronization error meets the requirements, air aggregation is enabled; otherwise, it automatically switches to digital domain weighted average aggregation.

[0071] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:

[0072] 1) By leveraging logarithmic transformation, peak resampling, and a combined loss that balances peak penalty and risk underestimation, the constructed BiLSTM edge model significantly improves peak tracking performance at critical times such as evening rush hour while maintaining global errors such as MAE and RMSE that are no worse than those of mainstream CNN-RNN baselines.

[0073] 2) The introduction of Q-learning's adaptive client selection and weighting strategy enables the federated training loss curve to exhibit faster convergence speed and lower steady-state level under the same number of communication rounds, thereby improving the utilization rate of sample and bandwidth resources.

[0074] 3) Weighted aggregation achieved by combining over-the-air computing effectively reduces uplink communication overhead while ensuring model accuracy, providing a feasible path for practical deployment on a large scale of terminals. Attached Figure Description

[0075] Figure 1 A Q-learning adaptive federal household load forecasting system driven by in-flight computing;

[0076] Figure 2 For Res-BiLSTM attention residual prediction model;

[0077] Figure 3 For the Air Comp aggregation process;

[0078] Figure 4 Figure 1 shows the user load curves before and after peak resampling; where: (a) is the user load curve before alignment, and (b) is the user load curve after peak resampling.

[0079] Figure 5 This is a chart showing the percentage of peak segments with consistent performance.

[0080] Figure 6 This is a curve showing the underestimation rate of the peak segment.

[0081] Figure 7 Figure (a) in the figure is a comparison of the loss function curves, and Figure (b) is a comparison of the mean absolute error.

[0082] Figure 8 A comparison chart of loss function curves;

[0083] Figure 9 A comparison chart of prediction curves from different models. Detailed Implementation

[0084] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and should not be used to limit the scope of protection of the present invention.

[0085] 1. Structure and Theoretical Basis of Residential Load Forecasting System

[0086] 1.1 Overall Architecture and Data Processing of the Prediction System

[0087] Unlike traditional single-family prediction models, this invention's system achieves cross-family collaborative prediction within a federated learning framework, completing global modeling while keeping the original data at the edge. At the communication layer, it introduces Air Comp, utilizing the linear superposition of wireless multiple access channels to achieve parallel aggregation of parameter updates, thereby simultaneously reducing uplink bandwidth consumption and round latency, and enhancing privacy protection. The overall system architecture is as follows: Figure 1 As shown, the system consists of four sub-layers: data processing, client-side modeling, communication and over-the-air aggregation, and server-side aggregation and scheduling. First, the data processing layer performs logarithmic transformation and normalization on the original electricity consumption sequence at the client side, and resamples it based on peak-valley distribution, outputting standardized samples for local training. Second, client-side modeling uses a time-series model composed of bidirectional LSTM, attention mechanism, and residual correction module to represent multi-scale time-series patterns and abrupt changes. The client side only retains the model update quantity and necessary state quantity related to aggregation, without uploading the original data. Third, the communication and over-the-air aggregation layer uses Air Comp for parallel aggregation. When link conditions are poor or synchronization is limited, the system switches to digital domain weighted averaging to ensure robustness. Fourth, the server-side aggregation and scheduling layer completes global model updates after receiving the superposition results and introduces Q-learning. Based on the state constituted by link quality, sample representativeness, and historical returns, it adaptively selects participating clients and sets weights, thereby achieving better results in terms of communication efficiency, training fairness, and peak prediction stability. Overall, the above mechanism achieves collaborative optimization of the terminal, network, and cloud without leaking the original data, effectively improving the accuracy and robustness of short-term load forecasting.

[0088] 1.2 Data Preprocessing and Sample Construction

[0089] In the data processing stage, a logarithmic mapping is first performed, followed by normalization to [0,1]: Let the original load be... ,

[0090] ;

[0091] The input, objective, and loss function of the subsequent model are all calculated in [0,1], and the evaluation index is the statistics on the restored physical dimensions.

[0092] This invention constructs a sliding window input sequence of length T. Predicting targets Distributed statistics on the training target reveal a high proportion of low-load samples and a scarcity of high-load samples. Therefore, if uniform sampling is performed directly, the gradient will be dominated by troughs, leading the model to tend towards "conservative estimation," resulting in low peak prediction accuracy. To alleviate the "many zero-load, few peaks" characteristic, this invention utilizes peak resampling to increase the probability of selecting samples with higher target values. First, a threshold is given based on the target distribution of the training set:

[0093] , ;

[0094] Define the sampling weights as ,in Weighted sampling with replacement is used during training, with a sampling probability of... This approach, without altering the loss function, stably increases the frequency of peak samples while maintaining valley information. The resulting batches of training samples are then input into the local model described in Section 2.1 for optimization.

[0095] This chapter clarifies the system architecture of federal load forecasting and, based on the characteristics of load data—"many valleys and few peaks, large numerical differences"—completes preprocessing such as logarithmic normalization and sliding windowing, laying a practical foundation for stable training and peak sensitivity. Building on the above, Chapter 2 will delve into the methodology.

[0096] 2. Method Design

[0097] Considering the significant differences in data scale, network quality, and electricity consumption behavior among different households in real-world scenarios, simply adopting a "full-client, equal-weight" federated training strategy would not only waste bandwidth but also slow down the overall convergence speed and amplify peak prediction errors. This chapter focuses on three aspects: edge-side modeling, peak-sensitive loss construction, and server-side scheduling and air aggregation. It presents the Res-BiLSTM edge prediction model, a combined loss function that balances peak security and upper quantile calibration, and a Q-learning-driven client selection and Air Comp weighted aggregation mechanism, thus forming a complete AirQ-FedHF methodology.

[0098] 2.1 Local Model and Residual Correction

[0099] To address the sparse and intermittent characteristics of household load sequences (predominantly zero load with significant peak-valley differences), a bidirectional LSTM-attention-residual correction prediction structure is constructed at the terminal side, such as... Figure 2 As shown, BiLSTM captures the temporal dependencies and periodicity within the window, and the attention adaptively highlights key segments at the time level. The residual correction balances the "most recent observation baseline" and "model output" through learnable gating, thereby reducing the interference of the underestimated valley segment on the overall training and improving the fitting ability to the peak segment.

[0100] Will The hidden state sequence is obtained by feeding it into BiLSTM. ,in It is composed of forward and backward LSTM concatenation. In the time dimension, dot product scoring and SoftMax attention mechanisms are used to adaptively focus on key segments.

[0101] , , ;

[0102] Subsequently The model output is obtained by applying batch normalization (BN) and multilayer perceptron (MLP).

[0103]

[0104] Use Sigmoid to ensure consistency with the normalization domain. Attention weights. The contribution of key segments such as "climbing / peaking" can be interpreted in time dimension.

[0105] To reduce the dominance of valley segments in model training and improve peak segment response, learnable gating is introduced. Adaptive trade-off between recent observation baseline and model output:

[0106] , , ;

[0107] When the scene is in a valley. Smaller More reliant on recent observations; when a transition or near-peak occurs, Increase It relies more on model representation. To connect with the objective function in Chapter 2, the prediction residual is defined.

[0108] ;

[0109] Used to characterize the deviation between the model output and the final corrected prediction, and serves as an interface between peak re-penalty and quantiles. Final prediction Located in the [0,1] normalization domain, after inference, it is restored to physical dimensions according to the anti-normalization rule in Section 1.2.

[0110] In summary, BiLSTM handles temporal encoding, the attention mechanism highlights key segments in the time dimension, and the gated residuals adaptively balance between the baseline and the model. This structure provides stability in the troughs and enhances sensitivity in the peaks, complementing the "two-sided weighting + peak penalty + quantile loss" approach in Section 2.2; its end-side output ( The update quantity will be used for client scheduling and over-the-air aggregation in Section 2.3.

[0111] 2.2 Construction of the loss function

[0112] To ensure the prediction model covers the peak region and is stable in the valley region, the following four addable objective functions are constructed.

[0113] 2.2.1 Mean Square Error of Two-Sided Weighted Scaling

[0114] To reduce the relative error for low-load samples and avoid excessive penalty for high-load samples, this invention constructs two weights at the sample layer: one for the peak side and one for the valley side, and obtains the overall weight by linear combination.

[0115] ;

[0116] in The mixing coefficient, Control the degree of emphasis on peak values. Let be the numerical stability constant. The corresponding fundamental loss function is:

[0117] ;

[0118] 2.2.2 Mean Square Error of Peak Re-punishment

[0119] To make the model pay more attention to peak samples during the optimization process, Additional secondary penalties will be triggered at that time:

[0120] , ;

[0121] This feature is only activated during peak periods, and... The superposition effectively assigns a larger weight to the peak error, thus increasing its relative contribution to the overall objective. For In this respect, the term is differentiable when If that happens, the item will not participate.

[0122] 2.2.3 Underestimating the punishment

[0123] In real-world load forecasting scenarios, underestimating peak loads often incurs costs far exceeding those of overestimating them by the same magnitude. Therefore, a one-sided linear penalty is applied to negative errors using a positive part operator. remember:

[0124] ;

[0125] The equivalent segmented form is:

[0126] ;

[0127] This term only takes effect when the peak value is underestimated; otherwise, it is zero and does not affect the remaining training samples. Since there is an inflection point at zero, we use subgradient updates to stably optimize while concentrating the penalty in the critical region.

[0128] 2.2.4 Quantile Loss

[0129] The above This approach only considers cases where the peak value is underestimated, representing a localized optimization that can only suppress underestimation in key regions but is not comprehensive enough. To enable up-calibration across the entire sample and enhance robustness, an asymmetric L1 quantile loss is further introduced. The two complement each other: the former ensures that the peak value is not underestimated, while the latter makes the overall prediction closer to the upper quantile, reducing systematic underestimation. Asymmetric L1 loss:

[0130] ;

[0131] when This factor makes the prediction closer to the upper quantile target, less sensitive to extreme errors, and helps to improve the stability of training convergence.

[0132] From 2.2.1 to 2.2.4 above, the overall objective L is obtained by linear addition:

[0133] ;

[0134] The combined loss function constructed in this section ensures the stability of valley prediction and suppresses the risk of peak underestimation, while providing overall stability through the quantile term. This objective complements the peak weighted sampling in Section 1.2 and is consistent with the gated residual structure in Section 2.1. Therefore, the Q-learning×Air Comp in Section 2.3 can achieve adaptive adjustment and weighted aggregation based on more discriminative local updates and importance weights.

[0135] 2.3 Server-side scheduling and over-the-air aggregation

[0136] 2.3.1 Server-side selection and adaptive weighting

[0137] After clarifying the local training objectives and risk preferences in Section 2.2, the server needs to determine the participating clients and aggregate weights in each round t, and perform a global update accordingly. The specific steps are as follows: The server first collects a small amount of information from online clients and organizes it into...

[0138] , ;

[0139] in Let the geometry of the clients that are online in round t be denoted by ; The upper quantile forecast obtained in Section 2.2.4 characterizes the strength of the short-term upward trend; This is a link quality metric, given by the normalized instantaneous signal-to-noise ratio; This represents the effective sample size used for training in this round. Based on this state vector, a comprehensive index is constructed that simultaneously characterizes risk, trend, link availability, and sample representativeness. The participation priority of client i in round t:

[0140] ;

[0141] To meet the size and availability constraints of each round, the server is based on... With link quality Generate candidate set

[0142] ;

[0143] in , Thresholds for urgency and link quality, respectively.

[0144] After obtaining the candidate set Then, Q-learning without function approximation is introduced to learn weighted temperatures online. This adaptively adjusts the concentration of weight allocation. Actions are selected from the discrete set A according to the e-greedy rule; the immediate reward is:

[0145] ;

[0146] in Let be the combined loss on the validation set after the t-th round of aggregation (calculated using the objective function constructed in Section 2.2). To normalize communication overhead, This is a balancing factor. A larger reward indicates a more significant reduction in validation loss with lower communication costs. The update rule is based on the Q-learning standard.

[0147] ;

[0148] in For learning rate, This serves as a discount factor. Through multiple rounds of interaction, Q-learning gradually learns a temperature selection strategy that matches the task objective.

[0149] Given the current action, the urgency will be adjusted by temperature. Mapping to normalized importance weights

[0150] ;

[0151] The larger the value, the more the weight is concentrated in the hands of a few key clients; The smaller the value, the more uniform the distribution. This training layer's preference for "peak segments being more important, and underestimation being more expensive" is systematically transmitted to the decision-making and weight allocation stages. Subsequently, the server... broadcast Each client optimizes its local data using the combined loss from Section 2.3, resulting in... To facilitate parallel aggregation, the update amount in this round is recorded.

[0152] ;

[0153] A global update can directly target... Weighted averages can also be used for... For weighted averages, the two are numerically equivalent, differing only in that they involve a shift. .

[0154] 2.3.2 Aerial Convergence

[0155] like Figure 3 As shown, to achieve parallel weighted aggregation within a single uplink transmission, this invention employs over-the-air computation at the communication layer. This method leverages the linear superposition characteristic of wireless multiple access channels, enabling multiple terminals to transmit simultaneously on the same time-frequency resources, with their signals linearly superimposed at the receiving end via the wireless multiple access channel.

[0156] To maintain consistency with the selection and authorization process in Section 2.3.1, the server first obtains the authorization for this round. and distribute global parameters. Each terminal completes a limited number of optimization steps locally, recording the update volume. To ensure that aggregation reflects the importance of each endpoint and link availability, the terminal side pre-scales according to the weights set by the server side and then superimposes the data in the air: Let the equivalent positive real gain of the uplink channel after phase compensation be... Define the equivalent aggregate weight on the end side. .

[0157] because The time-varying channel gain coefficient is unknown and difficult to measure accurately; the client simultaneously transmits two signals on the same time-frequency resource: one is the vector to be aggregated. The other path is to set a weighted scalar. Suppose the server receives the following response:

[0158] , A1;

[0159] When synchronization is good and SNR is high, ratio normalization is performed to eliminate unknown common scales, resulting in...

[0160] , A2;

[0161] Then apply a uniform server-wide step size. and by function Project to feasible ,get

[0162] A3;

[0163] Equation (A3) is numerically equivalent to weighting the update amounts of each client in the numeric domain. Weighted averaging simply involves moving the weighted summation operation to the WMAC channel in the analog domain. If the receiver observes that the WMAC condition is not met, the server discards the current round of over-the-air aggregation results and reverts to the traditional digital domain weighted averaging form.

[0164] ;

[0165] Thus, a parallel weighted aggregation is achieved using over-the-air computation: the server assigns importance weights to each terminal, which multiplies these weights by the link gain and then transmits them on the same time-frequency resource, achieving weight multiplication followed by summation. The server normalizes the ratio of the received vector sum to the weighted sum, applies a uniform step size, and projects according to constraints to complete the global update.

[0166] Numerical simulation results and analysis

[0167] Numerical validation of the proposed AirQ-FedHF (Air Computation-Driven Q-learning Adaptive Federated Household Load Forecasting) system was conducted on a Python simulation platform. First, the dataset and preprocessing methods, model structure and key parameter settings, and the evaluation metrics used were introduced. Then, the system's prediction accuracy and robustness were comprehensively evaluated through comparison with baseline methods such as centralized bidirectional LSTM, traditional federated learning (FedAvg), and Air Comp-FL (Air Computation-FL), combined with ablation experiments involving peak resampling and Q-learning adaptive weighting strategies.

[0168] 3.1 Parameter Settings and Dataset

[0169] The prediction model and its comparison method proposed in this invention were both completed in an experimental environment with NVIDIA GeForce RTX4060 hardware and Python 3.8, Pytorch 1.8 and Anaconda3 deep learning frameworks.

[0170] This invention uses residential smart meter data from a certain region as the research sample. Data is sampled at 30-minute intervals, recording active power consumption over half an hour according to the "resident-timestamp" format, covering various energy consumption scenarios including weekdays, weekends, and public holidays. To focus on the effectiveness of the method itself, this invention uses only the load as a single variable for modeling and evaluation, without incorporating meteorological or electricity price features, and all data has been de-identified, containing no personal or geographical location information.

[0171] Simultaneously, the system's hyperparameters are primarily characterized and used in all comparative experiments under unified settings. After multiple model pre-training and hyperparameter optimization adjustments, the selection and settings of the relevant hyperparameters for the prediction model proposed in this invention are shown in Table 1.

[0172] Table 1 Hyperparameter Settings

[0173] submodule parameter Optimal settings BiLSTM Hidden unit number 128 LSTM layers 3 Fully connected activation function ReLU Optimizer Adam Batch size 32 Local training rounds 5 Local learning rate 0.0001 Air Comp Mixing coefficient 0.2 Channel weight 2 Q-learning Learning rate 0.1 Exploration rate 0.2 Discount factor 0.9

[0174] 3.2 Evaluation Indicators

[0175] To comprehensively characterize the overall fit, peak segment safety, and upper quantile calibration, this invention statistically analyzes the following indicators on the inversely normalized physical dimensions. Let the number of test samples be N, and the actual and predicted values ​​be respectively... and .

[0176] Mean Absolute Error (MAE)

[0177] ;

[0178] Root Mean Square Error (RMSE)

[0179] ;

[0180] Symmetric Mean Absolute Percentage Error (sMAPE)

[0181] ;

[0182] Corrected Mean Absolute Percentage Error (MMAPE)

[0183] ;

[0184] in To avoid a smoothing constant where the denominator is close to zero, MMAPE does not produce excessive percentage errors during extremely low load periods at night, making it more suitable for sequences with prolonged low load periods, such as residential load.

[0185] Peak-mean absolute error (P-MAE)

[0186] ;

[0187] Peak segment underestimation rate (UER)

[0188] ;

[0189] 3.3 Performance Evaluation

[0190] 3.3.1 Sensitivity Analysis and Optimal Setting of Quantile Threshold q

[0191] To improve the comparability of cross-day comparisons and reduce the time migration effect caused by home-based work and rest schedules, this invention performs uniform preprocessing on the load sequences of the same household: missing and outlier corrections are completed, logarithmic mapping is applied followed by Min-Max normalization, and the data is sliced ​​by day (48 points / day). Figure 4 As shown, "Time / h" is used as the horizontal axis and "Normalized load" as the vertical axis. Then, peak alignment resampling is introduced as a key step: the main peak of the daily curve and its surrounding area are aligned with the plotted points, and local resampling and interpolation are performed on the time axis. Only the timing is corrected and the amplitude structure is not changed, so as to align the main peak and its upper and lower slopes without destroying the overall structure. Figure 4 Figure (a) shows the situation before alignment. The two-day curves show significant peak misalignment and slope misalignment in the morning and evening, which can easily lead to deviations when directly compared. Figure 4 In (b) of the figure, after map alignment, the main peak in the 8-11h region overlaps well with the adjacent slope segment, and sporadic jitter in non-critical periods is mildly suppressed. This processing significantly enhances the interpretability of cross-day comparisons and subsequent modeling, and provides a unified reference for the robust evaluation of peak sensitivity indicators and federated training phases.

[0192] To quantify the positive effect of peak alignment resampling on cross-day stability, this invention introduces Peak Segment Consistency (PACR) as a stability measure of daily load patterns for a single household. Specifically, the daily load curves of the same household are first aligned with peaks. A representative day is selected under a peak segment weighted distance metric. The proportion of days sufficiently similar to the representative day to the number of valid days is recorded as the PACR of that household, with a value range of [0,1]. Figure 5 The distribution of the aligned PACR is presented: the PACR of most households is concentrated in the upper-middle range, with a significant portion exceeding two-thirds, and only a small number falling into the low-consistency range. This indicates that peak alignment significantly mitigates the cross-day shift caused by the shift in work and rest schedules without altering the amplitude structure, enhancing the comparability and stability of key peak segments; while the few low-consistency samples are mostly related to abnormal days, holiday disturbances, or equipment replacements. These results validate the effectiveness of the data preprocessing used and provide a reliable basis for subsequent peak-sensitive loss design and robust weighting in federated aggregation.

[0193] For candidate quantile threshold set Discretize the values. For each candidate q, calculate the quantile threshold based on the actual load sequence. Based on this, a peak segment sample set is constructed. Subsequently, P-MAE and UER were calculated for each client. The minimum, maximum, and average values ​​at the user level are summarized in Table 2, and a curve showing UER as a function of q is plotted. Figure 6 As shown. To establish a unified decision-making criterion among "error, risk, and coverage," a comprehensive score is introduced:

[0194] ;

[0195] in: Coverage In this invention .

[0196] Table 2 P-MAE at different quantile thresholds

[0197] 0.50 0.60 0.70 0.80 0.90 Minimum value (Min P) 0.2415 0.2211 0.2127 0.1812 0.2113 Maximum value (Max P) 0.3912 0.3714 0.3395 0.2897 0.3092 Mean P 0.3430 0.2950 0.2529 0.2228 0.2622 Overall score (J) 0.9080 0.6840 0.4239 0.1605 0.3524

[0198] From Table 2 and Figure 6 As can be seen, with the increase of the quantile threshold q, the peak underestimation rate (UER) shows a consistent downward trend across all five clients, reaching a trough at q=0.8; when q is further increased to 0.90, the UER generally rebounds slightly. Correspondingly, the user mean P-AME decreases monotonically with increasing q to q=0.80, then rebounds to 0.2622 at q=0.90; the maximum value also decreases by 25.9%. This indicates that increasing q leads to a decrease in the peak set... Non-peak sample interference is reduced, thus simultaneously mitigating systematic underestimation and amplitude error; however, when coverage... When the value is too small, the statistical variance increases and the index becomes more sensitive to extreme points, thus causing a rebound. Based on a comprehensive evaluation considering error, risk, and coverage, the minimum criterion is used to determine q=0.8 as the optimal quantile threshold.

[0199] To evaluate the impact of peak resampled (PR) on model training stability and error performance based on this optimal quantile threshold, the performance of the baseline model BiLSTM and BiLSTM+PR was compared over 50 rounds of federated training. Figure 7 As shown in Figure (a), the descent trajectories of the loss values ​​for both are basically consistent with the steady-state fluctuation amplitude, indicating that the introduction of PR did not change the numerical properties of the optimization process, nor did it introduce any additional oscillations or unstable behaviors. Correspondingly, Figure 7Figure (b) shows that BiLSTM+PR experiences a faster decrease in MAE and converges earlier in the early rounds; during the steady-state period, its mean MAE is generally lower than that of the baseline model. The combination of faster convergence and a smaller mean indicates that by reconstructing the training sample distribution and increasing the frequency of peak samples, PR effectively alleviates the gradient bias caused by the "many valleys and few peaks" phenomenon, further reducing the absolute error without worsening the overall loss. These results are consistent with the design goal of the "quantile threshold-driven resampling and loss weighting framework" described in Chapter 2, which aims to improve the representation and fitting ability of peak segments while maintaining global convergence stability, ultimately leading to better error performance.

[0200] 3.3.2 Effectiveness Analysis of Q-learning Strategy

[0201] After completing the sensitivity analysis of the quantile threshold q and determining q=0.8 as the unified standard for peak segmentation and loss weighting, we further examine the actual effect of the client selection and adaptive weighting strategy based on Q-learning proposed in Section 2.3.1 in the federated training process. Consistent with the previous section, all experiments were conducted under the same data preprocessing and local model structure, only changing the client selection and aggregation weighting methods on the server side. We compared and analyzed the convergence characteristics of the federated training loss after introducing Q-learning to verify the effectiveness of the proposed adaptive weighting strategy.

[0202] To highlight the role of the Q-learning strategy itself, two schemes were compared: the first was a traditional scheme combining fixed weights and random selection, where online terminals were randomly selected to participate in aggregation each round while meeting the minimum number of participating clients, and the model updates were normalized and weighted based on local samples; the second was a Q-learning adaptive weighting scheme, which used state vectors to encode information such as the client's quantile status, link quality, and effective samples, and updated the action value function online, outputting the corresponding weighting temperature in each training round, thereby achieving adaptive adjustment of participating clients and their aggregation weights. Apart from this difference, the two schemes were completely identical in terms of loss function construction, optimizer parameters, and the number of training rounds.

[0203] From the perspective of the global loss convergence process, such as Figure 8As shown, in the first few rounds of training, the Q-learning weighted approach exhibits a more pronounced loss reduction trend: with the same number of federated rounds, its global loss curve decreases faster, leaving the initial steep drop phase earlier and entering the stable convergence region. Using a target loss threshold as a reference, it can be observed that the number of federated rounds required for the Q-learning method to reach this threshold is significantly less than that of the fixed-weighted approach. This indicates that, given a certain communication overhead, Q-learning can more efficiently utilize the gradient information uploaded by each terminal, accelerating the convergence of the global model.

[0204] Further comparison of the steady-state performance of the loss curves from the two days in the later stages reveals that the convergence value of the Q-learning scheme is slightly lower than that of the fixed-weight scheme, meaning that its final stable loss level is smaller under the same number of training epochs. Simultaneously, the oscillation amplitude of the Q-learning curve in the stable phase is also relatively smaller, exhibiting a narrower fluctuation range and significantly suppressed violent fluctuations. This phenomenon indicates that Q-learning does not simply select a certain group of clients, but rather gradually learns a stable weighting strategy that matches the task objective during the cumulative interaction process. This makes the global model update direction obtained in each round of aggregation more consistent, thereby reducing the ineffective oscillations caused by random selection and fixed weighting.

[0205] In summary, without altering the local model structure and loss function design, the introduced Q-learning adaptive weighting mechanism has demonstrated stable benefits across multiple dimensions of the global loss curve: First, it significantly shortens the training time from the initial state to the low-loss interval for a given number of communication rounds, manifested as a faster descent slope; second, it reduces the steady-state loss level in the later stages of training, resulting in an overall downward shift in the convergence value; and third, it effectively suppresses ineffective oscillations in the mid-to-late stages, making the loss trajectory smoother. These results indicate that the proposed Q-Learning weighting strategy can systematically improve the sample utilization efficiency and convergence quality of federated training, laying a solid foundation for subsequent comparisons of the overall model performance advantages.

[0206] 3.4 Comparative Experiment

[0207] After separately validating the peak resampling mechanism and the Q-learning adaptive weighting strategy described above, we further compared the comprehensive performance of different network structures in the short-term load forecasting task under unified experimental conditions. All models adopted the same data preprocessing procedure and training configuration described in 3.1, with a time-series load of length 24 as input. The optimizer and learning rate settings were exactly the same, with differences only in network structure and introduction mechanism, thus ensuring the comparability of the comparison results and the reliability of the conclusions.

[0208] The selected comparison models include: CNN-BiLSTM, CNN-LSTM, LSTM-Attention, BiLSTM-LSTM, and the five structures proposed in this invention, AirQ-FedHF. Figure 9 A comparison of the predicted curves of each model with the actual load curves within the same time period shows that, during the relatively flat load troughs, most models can fit the actual curves well, with highly overlapping prediction results, indicating that the modeling capabilities of different structures are not significantly different under low load conditions. The main differences are concentrated in the peak intervals of sudden load increases: CNN-LSTM and CNN-BiLSTM exhibit significant amplitude compression and phase lag in multiple peak segments, with the peak position deviating somewhat from the actual curve; LSTM-Attention and BiLSTM-LSTM are more sensitive to the peak occurrence time, partially recovering the peak shape, but still underestimating the peak amplitude to some extent. In contrast, the AirQ-FedHF, corresponding to the orange curve, not only accurately tracks the peak occurrence time in most late peaks and abrupt change intervals, but also closely approximates the actual curve in terms of peak height and transition shape, with the peak waveform almost overlapping with the measured curve, intuitively demonstrating its advantage in peak capture.

[0209] To further analyze the performance differences of each model from a quantitative perspective, Table 3 lists various error metrics for the five models on the test set. It can be seen that, in terms of overall error, BiLSTM-LSTM slightly outperforms the other models in MAE, MSE, and RMSE, with values ​​of 0.1441, 0.0786, and 0.2804, respectively. The AirQ-FedHF model proposed in this invention is almost on par with it in these three metrics, with an MAE of 0.1442 and an RMSE of 0.2816, with the difference between the two being within 10... -4 The magnitude of the difference suggests that the method described in this invention does not sacrifice overall performance in terms of global fitting accuracy. In contrast, the MAE and RMSE of CNN-LSTM and CNN-BiLSTM are significantly larger, indicating that relying solely on convolutional feature extraction is insufficient to fully characterize the long-term dependency structure of resident workload.

[0210] Table 3 Comparison of evaluation indicators for different models

[0211] Model MAE MSE RMSE sMAPE MMAPE P-MAE CNN-BiLSTM 0.1524 0.0858 0.2929 35.6586 27.6898 0.1661 BiLSTM-LSTM 0.1441 0.0786 0.2804 32.3150 25.2853 0.1139 AirQ-FedHF 0.1442 0.0789 0.2816 31.0618 24.4255 0.1040 LSTM-Attention 0.1461 0.0789 0.2809 31.9229 25.0987 0.1071 CNN-LSTM 0.1485 0.0810 0.2846 36.7698 28.2028 0.1585

[0212] More discriminative is the relative error metric, which is more sensitive to peak segments. Table 3 shows that AirQ-FedHF achieves the best performance in all three metrics: sMAPE, MMAPE, and P-MAE. sMAPE is reduced to 31.06, and MMAPE to 24.43, both the lowest among all models. The P-MAE for high-load samples is 0.104, which is approximately 37% and 34% lower than CNN-BiLSTM and CNN-LSTM, respectively. These results demonstrate that, while maintaining overall MAE and RMSE scores no worse than the optimal baseline, the method of this invention can significantly reduce the relative error and average bias in high-load peak segments, effectively alleviating the problems of peak underestimation or response lag commonly found in traditional models during peak intervals.

[0213] Combination Figure 9 The comparison of the model prediction curves and the evaluation index results in Table 3 lead to the following conclusions: On the one hand, the introduction of bidirectional LSTM and attention structures can improve the peak segment response to some extent, but without a corresponding sample reweighting mechanism designed for the "more valleys and fewer peaks" distribution characteristics, the peak segment error is still difficult to reduce further. On the other hand, the peak resampling and Q-learning adaptive weighting superimposed on the BiLSTM structure in this invention enable the model to pay more attention to peak samples during the training phase and prioritize the absorption of updates on peak segment representativeness during the federated aggregation phase. Ultimately, this results in a more prominent peak segment prediction capability in both the prediction curve and evaluation index dimensions. Therefore, AirQ-FedHF not only maintains a level comparable to the baseline in global fitting accuracy but also significantly outperforms various mainstream CNN-RNN models in peak prediction accuracy and peak segment shape recovery, fully validating the effectiveness and superiority of this method in the task of short-term load forecasting for residents.

[0214] Summarize

[0215] Addressing the characteristics of residential load under the "dual-carbon" goals and the new power system context, such as "frequent valleys and sparse peaks, non-stationary loads, and non-independent distribution across users," this invention constructs an over-the-air computing-driven adaptive federated household load forecasting framework. While balancing privacy protection and communication constraints, it improves the peak load characterization capability and overall stability of short-term load forecasting. Experimental numerical results lead to the following conclusions:

[0216] 1) By leveraging logarithmic transformation, peak resampling, and a combined loss that balances peak penalty and risk underestimation, the constructed BiLSTM edge model significantly improves peak tracking performance during critical moments such as evening rush hour while maintaining global errors such as MAE and RMSE that are no worse than mainstream CNN-RNN baselines.

[0217] 2) The introduction of Q-learning's adaptive client selection and weighting strategy enables the federated training loss curve to exhibit faster convergence speed and lower steady-state level under the same number of communication rounds, thereby improving the utilization rate of sample and bandwidth resources.

[0218] 3) Weighted aggregation achieved by combining over-the-air computing effectively reduces uplink communication overhead while ensuring model accuracy, providing a feasible path for practical deployment on a large scale of terminals.

[0219] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. An over-the-air computing-driven adaptive federal household load forecasting method, characterized in that: Includes the following steps: Logarithmic transformation and normalization preprocessing are performed on the original household load data. Based on the peak and valley distribution predicted by the sliding window, quantile thresholds are calculated to divide the peak and valley samples. Peak resampling is then used to form standardized training samples. Deploy the prediction model on each client, use standardized training samples as input, and train the prediction model locally using a combined loss function. Do not upload the original household load data. Output the predicted value and prediction residual. The server collects the upper quantile prediction values, link quality indicators and sample size of online clients to construct a state vector. It updates the action value function online through the Q-learning algorithm, generates a weighted temperature, and determines the set of clients participating in the aggregation and the aggregation weight of each client based on the state vector and the weighted temperature. The client prescales the local model update based on the aggregation weights and uses the linear superposition characteristics of the wireless multiple access channel to achieve parallel weighted aggregation through over-the-air computation. After receiving the superimposed signal, the server restores the digital domain average by normalizing the ratio and applies a uniform global step size and feasible region projection to complete the global model update. When the channel conditions do not meet the requirements of over-the-air computation, it falls back to digital domain weighted average aggregation.

2. The over-the-air computing-driven adaptive federal household load forecasting method according to claim 1, characterized in that: The specific preprocessing steps are as follows: first perform a logarithmic mapping, then normalize to... : Let the original household load be... ,but: ; In the formula: The logarithmically mapped value of the original household load. This is the final output value obtained after normalization following the logarithmic transformation. and These are the minimum and maximum values ​​in the training set, respectively.

3. The over-the-air computing-driven adaptive federal household load forecasting method according to claim 2, characterized in that: The specific process of peak segment resampling is as follows: The threshold is given by the target distribution of the training set: , ; The sampling weights are defined as follows: ,in ; Weighted sampling with replacement is used, with a sampling probability of... .

4. The over-the-air computing-driven adaptive federal household load forecasting method according to claim 3, characterized in that: The prediction model includes a bidirectional BiLSTM, an attention mechanism, a normalization and multilayer perceptron module, and a residual correction module; wherein: Bidirectional BiLSTM is used to encode the input sequence of a sliding window and output a bidirectional hidden state sequence. The attention mechanism calculates the attention weights for each time dimension using dot product scoring and SoftMax, and then performs a weighted summation on the hidden state sequence to obtain the context vector. The normalization and multilayer perceptron module normalizes and performs nonlinear transformations on the context vector, and outputs the model's preliminary prediction value after Sigmoid activation. The residual correction module adaptively compromises between the model's initial predictions and the most recent observation baseline using learnable gating, and outputs the model's predicted values.

5. The over-the-air computing-driven adaptive federal household load forecasting method according to claim 4, characterized in that: The combined loss function is: ; In the formula: The mean square error is a two-sided weighted average. , ,in: The mixing coefficient, Control the degree of emphasis on peak values. Here, is the numerical stability constant, N is the number of samples used for calculation in this round, and t is the time index, corresponding to the end time of the current sliding window. Let be the total weight of the t-th sample in the baseline mean squared error. For valley segment weighting function, For peak segment weighting function, Let be the normalized predicted load corresponding to the t-th sample; The mean square error of peak re-penalty. , ; To underestimate the punishment, ; For quantile loss, .

6. The over-the-air computing-driven adaptive federal household load forecasting method according to claim 5, characterized in that: The process of determining the aggregation weight: The server first collects information from online clients, organized as follows: , ; in: Let t be the set of clients that are online in round t. The upper quantile is used for prediction to characterize the strength of short-term price increases; This is a link quality metric, given by the normalized instantaneous signal-to-noise ratio; This represents the effective sample size used for training in this round; based on this state vector, a comprehensive index is constructed that simultaneously characterizes risk, trend, link availability, and sample representativeness. The participation priority of client i in round t: ; To meet the size and availability constraints of each round, the server is based on... With link quality Generate a candidate set: ; in: , These are thresholds for urgency and link quality, respectively. After obtaining the candidate set Then, Q-learning online learning with weighted temperature is introduced without function approximation. To adaptively adjust the concentration of weight allocation; actions are taken from the discrete set A according to... -Greedy rule selection; immediate reward: ; in: Let be the combined loss on the validation set after the t-th round of aggregation. To normalize communication overhead, The balancing coefficient; a larger reward indicates a more significant reduction in validation loss with less communication cost; the Q-learning standard is used to update the rules: ; in: For learning rate, The discount factor is used; through multiple rounds of interaction, Q-learning gradually learns a temperature selection strategy that matches the task objective. Given the current action, the urgency will be adjusted by temperature. Mapped to normalized importance weights: ; The larger the value, the more concentrated the weight becomes on key clients; The smaller the value, the more uniform the distribution; thus, the peaks in the training layer are important, and the expensive training bias is underestimated, systematically transmitted to the two stages of decision-making and weight allocation; subsequently, the server... broadcast Each client further optimizes the data using the combined loss function on its local data, resulting in... To facilitate parallel aggregation, the update amount in this round is recorded as follows: ; Global updates can directly affect A weighted average can also be used to... For weighted averages, the two are numerically equivalent, differing only in that they involve a shift. .

7. The over-the-air computing-driven adaptive federal household load forecasting method according to claim 6, characterized in that: The specific process of global model update is as follows: Let the equivalent positive real gain of the uplink channel after phase compensation be... Define the equivalent aggregate weight on the end side. ; because The time-varying channel gain coefficient is unknown and difficult to measure accurately; at the client end, two signals are transmitted simultaneously on the same time-frequency resource: one is the vector to be aggregated. The other path is to set a weighted scalar. Suppose the server receives the following response: , A1; When synchronization is good and SNR is high, ratio normalization is performed to eliminate unknown common scales, resulting in: , A2; Then apply a uniform server-wide step size. and by function Project to feasible ,get: A3; Equation (A3) is numerically equivalent to weighting the update amounts of each client in the numeric domain. Weighted averaging simply involves moving the weighted summation operation to the WMAC channel in the analog domain. If the receiver observes that the WMAC condition is not met, the server abandons the current round of over-the-air aggregation and reverts to the traditional digital domain weighted averaging form. ; Thus, a parallel weighted aggregation is achieved using over-the-air computing: the server provides the importance weights of each terminal, the terminal multiplies them by the link gain and transmits them on the same time-frequency resource, realizing weighted multiplication followed by summation; the server normalizes the ratio of the received vector sum to the weighted sum, applies a uniform step size and projects according to constraints to complete the global update.

8. An over-the-air computing-driven adaptive federal household load forecasting system, characterized in that, It includes a data processing layer, a client-side modeling layer, a communication and over-the-air aggregation layer, and a server-side aggregation and scheduling layer; among which: The data processing layer, deployed on the client, performs logarithmic transformation and normalization preprocessing on the raw household load data. It calculates quantile thresholds based on the peak-valley distribution predicted by the sliding window to divide peak and valley samples. It then forms standardized training samples through peak resampling. The client-side modeling layer is deployed on the client side. It models the prediction model on each client, takes the training samples as input, and uses the combined loss function to train the prediction model locally. It does not upload the original household load data and outputs the predicted value and prediction residual. In the communication and air aggregation layer, the server collects the upper quantile prediction values, link quality indicators and sample size of online clients to construct a state vector. It updates the action value function online through the Q-learning algorithm, generates a weighted temperature, and determines the set of clients participating in the aggregation and the aggregation weight of each client based on the state vector and the weighted temperature. The server-side aggregation and scheduling layer, deployed on the server, prescales the local model update based on the aggregation weight. It utilizes the linear superposition characteristics of the wireless multiple access channel to achieve parallel weighted aggregation through over-the-air computation. After receiving the superposition signal, the server restores the digital domain average by normalizing the ratio and applies a uniform global step size and feasible region projection to complete the global model update.

9. An over-the-air computing-driven adaptive federal household load forecasting system according to claim 8, characterized in that: In the prediction model, the number of hidden units in the bidirectional BiLSTM layer is 128, and the number of layers is 3; the attention mechanism adopts dot product attention; the normalization and multilayer perceptron module includes a batch normalized BN layer and a fully connected layer with ReLU activation function; the learnable gating of the residual correction module is generated through the Sigmoid activation function.

10. An over-the-air computing-driven adaptive federal household load forecasting system according to claim 8, characterized in that: The channel condition judgment criteria supported by the communication and air aggregation layer are as follows: when the signal-to-noise ratio of the received signal is higher than a preset threshold and the synchronization error meets the requirements, air aggregation is enabled; otherwise, it automatically switches to digital domain weighted average aggregation.

Citation Information

Patent Citations

  • Building space unit cooling load prediction method and system based on federated learning

    CN113240184A

  • Method and apparatus for supporting federated machine learning operations in communication network

    CN118715762A

  • Household load peak regulation and control method based on deep reinforcement learning

    CN120090248A