Artificial intelligence-based meter reading scheduling method

By constructing an electrical heterogeneous graph and an edge-time-varying topology-physical consistency model, combined with a communication conflict hypergraph and reinforcement learning scheduling, an efficient meter reading scheduling plan is generated, solving the problems of inaccurate communication prediction and low scheduling efficiency in complex electromagnetic environments, and realizing efficient and reliable data acquisition.

CN121390798BActive Publication Date: 2026-03-24NANJING XINLIAN ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-23
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing PLC meter reading and scheduling methods suffer from insufficient accuracy in communication prediction, low efficiency in scheduling strategy decision-making, and inability to cope with asynchronous physical events when facing complex and ever-changing electromagnetic environments, resulting in low acquisition efficiency and poor data integrity.

Method used

We construct an electrical heterogeneous graph and edge time-varying drive, adopt a topology-physical consistent continuous-time distribution prediction model to generate a communication performance prediction distribution, and generate a scheduling plan based on a graph structure reinforcement learning scheduling model with hypergraph constraints, using a communication conflict hypergraph and system configuration. The plan is then adjusted through online closed-loop and dual-time-scale self-optimization.

Benefits of technology

It improves the accuracy of communication forecasting and the robustness of scheduling decisions, ensuring efficient data acquisition and data integrity in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121390798B_ABST
    Figure CN121390798B_ABST
Patent Text Reader

Abstract

The application discloses a meter reading scheduling method based on artificial intelligence, comprising: acquiring power grid operation data, constructing an electrical heterogeneous graph and an edge time-varying drive; based on the electrical heterogeneous graph and the edge time-varying drive, using a topological-physical consistent continuous-time distribution prediction model, generating a communication performance prediction distribution containing a confidence band; based on the distribution and the electrical graph, constructing a dynamic communication conflict hypergraph; based on the hypergraph and the prediction distribution, using a hypergraph-constrained graph structure reinforcement learning scheduling model to generate a scheduling plan under the satisfaction of constraints (such as time, concurrency and mutual exclusion). The application improves the accuracy and safety of the prediction model by using physical consistency constraints and conformal calibration, and ensures the robustness and efficiency of the scheduling decision by using the hypergraph-constrained reinforcement learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to meter reading scheduling methods, and more particularly to a meter reading scheduling method based on artificial intelligence. Background Technology

[0002] The stable operation and refined management of smart grids heavily rely on massive amounts of timely electricity consumption data. Power line carrier PLC communication is a crucial technology for data acquisition in the last mile of automatic meter reading. Its reliability, timeliness, and integrity form the foundation for supporting upper-level services such as real-time billing, load analysis, topology identification, and fault location. Researching and optimizing PLC meter reading scheduling strategies in complex power grid environments to maximize data acquisition efficiency and quality has significant theoretical and engineering application value.

[0003] Currently, research on PLC meter reading scheduling mainly focuses on several aspects. Some solutions employ scheduling mechanisms based on fixed polling or static priority tables, executing meter readings according to a preset order or time window. Other solutions introduce dynamic adjustment mechanisms, such as adjusting the meter reading sequence or retry count based on historical communication success rate statistics. Regarding communication quality prediction, some research has begun to attempt to use conventional time series models, such as Long Short-Term Memory (LSTM) networks or recurrent neural networks, to predict the single-point communication success rate at a future point in time by learning from historical communication logs, time, load, and other data. In terms of conflict handling, static conflict diagrams based on phase or transformer area topology are mainly used to avoid known, fixed concurrent conflicts (such as three parallel lines A / B / C).

[0004] However, the existing solutions mentioned above have several insurmountable technical problems when facing the complex and ever-changing electromagnetic environment of real distribution areas. These problems are mainly reflected in the insufficient accuracy and robustness of communication prediction, as well as the low decision-making efficiency of scheduling strategies under constraints.

[0005] Existing prediction models (such as LSTM) are data-driven black-box models. Their prediction logic is disconnected from physical reality, and they cannot understand the causal impact of electrical topology (such as line branches and impedance) or transmission line physical laws (such as frequency-dependent attenuation and impedance mismatch reflection) on the state of the communication channel. When asynchronous and sudden physical events occur in the channel, such as switching or high-power load start-up and shutdown, black-box models have difficulty generalizing accurately, leading to prediction failure.

[0006] These models typically only provide average point estimates, such as average success rate or average time taken, lacking quantification of the uncertainty or confidence interval of the predicted results. For scheduling decisions, a prediction of an average time taken of 5 seconds is highly ambiguous; the scheduler cannot distinguish between a (low-risk) task that is stable between 4.9 and 5.1 seconds and a (high-risk) task that fluctuates wildly between 1 and 50 seconds. The lack of a safety margin in the prediction prevents the scheduler from performing robust planning based on worst-case scenarios, leading to frequent time budget overruns.

[0007] The deficiencies at the prediction level have spilled over to the scheduling level. On the one hand, existing conflict modeling is mostly static and paired, failing to describe the dynamic, multi-faceted (three or more meters) concurrent crosstalk generated by the combined effects of time-varying noise and multipath attenuation. On the other hand, existing scheduling algorithms (especially reinforcement learning) typically treat resource constraints such as a total time budget not exceeding 900 seconds as a soft penalty in the reward function. This not only leads to instability in the training process but also fails to guarantee that the constraint will not be violated during actual execution, easily causing systemic scheduling failures under resource pressure. Summary of the Invention

[0008] The purpose of this invention is to provide an artificial intelligence-based meter reading and scheduling method to solve the aforementioned problems existing in the prior art.

[0009] According to one aspect of this application, a meter reading scheduling method based on artificial intelligence includes:

[0010] Acquire power grid operation data to construct an electrical heterogeneity diagram and edge time-varying drive;

[0011] Based on electrical heterogeneous graphs and edge time-varying drive, a topology-physical consistent continuous-time distribution prediction model is used to generate a communication performance prediction distribution.

[0012] Based on the electrical heterogeneity graph and the predicted distribution of communication performance, a communication conflict hypergraph is constructed;

[0013] A graph-structured reinforcement learning scheduling model with hypergraph constraints is proposed to generate scheduling plans based on communication conflict hypergraphs, communication performance prediction distributions, and system configuration.

[0014] According to one aspect of this application, based on an electrical heterogeneous graph and edge time-varying drive, a topology-physical consistent continuous-time distribution prediction model is used to generate a communication performance prediction distribution, including:

[0015] By applying neural control differential equations, edge time-varying driving is evolved to generate temporal representations;

[0016] Based on the physical attributes and time sequence representation of electrical heterogeneity graphs, physical consistency constraints of transmission lines are constructed, and physical consistency loss is derived.

[0017] By using a distributed regression head, time-series representations are mapped, and the distribution of communication performance predictions is defined in a parameterized manner.

[0018] Furthermore, during model training, physical consistency loss is used to regularize the model's representation space.

[0019] According to one aspect of this application, based on the physical properties and time-series characterization of electrical heterogeneity diagrams, physical consistency constraints for transmission lines are constructed, and physical consistency losses are derived, including:

[0020] Calculate the frequency-dependent attenuation coefficient based on the per-unit-length parameter of the electrical heterogeneity diagram;

[0021] Based on the device parameters of the electrical heterogeneity diagram, determine the device transfer function;

[0022] Based on the load impedance and characteristic impedance of the electrical heterogeneity diagram, determine the termination reflection coefficient;

[0023] By combining the frequency-dependent attenuation coefficient, device transfer function, and termination reflection coefficient, the physical desired signal-to-noise ratio is derived.

[0024] A physical consistency loss is constructed based on the model prediction signal-to-noise ratio obtained by mapping the physical expected signal-to-noise ratio to the temporal representation.

[0025] According to one aspect of this application, neural control differential equations are applied to evolve edge time-varying drives and generate temporal representations, including:

[0026] An adaptive step-size solver is used to numerically solve the time-varying edge drive and obtain intermediate solution results.

[0027] At the switch switching event included in the edge time-varying drive, forced point marking and state update are performed on the intermediate solution results;

[0028] Along the electrical topology path of the electrical heterogeneous graph, the updated intermediate solution results are aggregated to form a time-series representation.

[0029] According to one aspect of this application, a scheduling plan is generated using a graph structure reinforcement learning scheduling model with hypergraph constraints, based on a communication conflict hypergraph, a communication performance prediction distribution, and system configuration. The plan includes:

[0030] Based on the communication conflict hypergraph, the communication performance prediction distribution and system configuration, an action feasible domain mask and candidate set are generated.

[0031] Using the structure of the communication conflict hypergraph as input, construct a graph structure policy network;

[0032] Graph-structured policy networks sample candidate sets and generate scheduling plans under the constraint of action-feasible domain masks.

[0033] According to one aspect of this application, a graph-structured policy network samples a candidate set under the constraint of an action-feasible region mask, including:

[0034] Based on the communication conflict hypergraph, mutually exclusive clusters are derived and the target cluster is sampled, which is constrained by the action feasible domain mask.

[0035] Within the target cluster, target nodes are sampled, constrained by the action-feasible domain mask.

[0036] After sampling the target node, based on the upper bound of communication time in the communication performance prediction distribution, an immediate time knapsack check is performed to ensure that the cumulative communication time does not exceed the time budget in the system configuration, and a scheduling plan is generated.

[0037] According to one aspect of this application, training a graph-structured policy network includes:

[0038] Identify homogeneous mutually exclusive clusters in the communication conflict hypergraph and enable these clusters to share policy sub-network parameters.

[0039] Compute the cluster-level policy entropy of the graph structured policy network on isomorphic mutually exclusive clusters, and obtain the structure entropy regularization term;

[0040] When training a graph structure policy network, a structural entropy regularization term is added to the training loss function.

[0041] According to one aspect of this application:

[0042] The communication performance prediction distribution includes a confidence band generated by a conformal calibration method that preserves coverage;

[0043] A graph structure reinforcement learning scheduling model with hypergraph constraints is used to perform safety-side filtering on electricity meters using confidence bands to generate a candidate set.

[0044] Furthermore, the scheduling model also utilizes the upper bound of communication time in the confidence band to calculate the constraint violation degree for Lagrange constraint optimization.

[0045] According to one aspect of this application, Lagrange-constrained optimization includes:

[0046] Obtain the real-time feedback logs after executing the scheduling plan, and calculate the actual batch rewards accordingly.

[0047] Construct PPO loss based on actual batch returns;

[0048] The PPO loss is combined with the constraint violation weighted by the Lagrange multipliers to form the total loss;

[0049] A graph structure reinforcement learning scheduling model that updates hypergraph constraints using total loss;

[0050] The Lagrange multipliers are adaptively updated based on the degree of constraint violation.

[0051] According to one aspect of this application, training of a topology-physically consistent continuous-time distribution prediction model further includes topology intervention data augmentation:

[0052] Simulate topology changes in electrical heterogeneity diagrams and generate a list of topology interventions;

[0053] Based on the topological intervention list, the process of deriving the physical expected signal-to-noise ratio is reused to re-estimate the counterfactual communication metric.

[0054] Generate counterfactual samples based on counterfactual communication metrics;

[0055] We use counterfactual samples and real samples to train the topologically and physically consistent continuous-time distribution prediction model alternately.

[0056] Beneficial effects: Through the above technical solutions, this invention solves the problem of low data collection efficiency caused by inaccurate communication prediction and rigid conflict modeling in existing scheduling methods, improves the accuracy and security of the prediction model, and ensures the robustness and efficiency of scheduling decisions. Attached Figure Description

[0057] Figure 1 This is a schematic diagram of the overall process of the meter reading and scheduling method based on artificial intelligence.

[0058] Figure 2 This is a flowchart illustrating the process of constructing physical consistency constraints for transmission lines and deriving physical consistency losses.

[0059] Figure 3 This is a flowchart illustrating the process of generating time-series representations.

[0060] Figure 4 This is a schematic diagram of the training process of a graph-structured policy network. Detailed Implementation

[0061] To enable those skilled in the art to better understand the present invention, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0062] Example 1: Explaining the technical background and problems to be solved by the present invention.

[0063] Existing electricity consumption information collection systems typically consist of a master station, concentrators, and electricity meters. The master station is responsible for issuing collection tasks and plans to the concentrators. The concentrators, based on the task plan, network with the electricity meters in the distribution area and perform periodic meter readings via power line carrier PLC communication.

[0064] However, in actual operation, especially under conditions of poor field carrier networking and complex and variable channel environments, such as high noise, strong attenuation, and impedance mismatch, the quality of power line carrier PLC communication links is relatively unstable. Existing data acquisition methods mostly employ fixed scheduling strategies or simple retry mechanisms, leading to problems such as untimely acquisition and low acquisition integrity. For example, a concentrator might attempt to read a critical meter during the period of worst channel quality, resulting in multiple failures; or it might attempt to read two meters simultaneously that would cause severe crosstalk on the physical line, resulting in both failures. This makes it difficult to meet the increasingly stringent requirements of smart grids for data integrity and timeliness.

[0065] To address the aforementioned technical issues, this invention proposes an artificial intelligence-based self-optimization method for meter reading scheduling. This method aims to improve data collection efficiency and data integrity in unstable communication environments by dynamically predicting communication quality and utilizing reinforcement learning models to optimize scheduling decisions within limited resources and time.

[0066] Example 2: A detailed description of the overall framework of the artificial intelligence-based meter reading and scheduling method proposed in this invention, such as... Figure 1 As shown, the specific steps include the following:

[0067] Step 1: Acquire power grid operation data and construct an electrical heterogeneity graph and edge time-varying drive. This step requires accessing and fusing multi-source heterogeneous data from multiple data sources within the power information system. Specifically, the acquired data includes, but is not limited to: historical communication sequences, such as the success or failure of each data reading, communication time, error rate, retransmission count, and timestamp; power grid topology data, such as the connectivity between transformers, branch boxes, meters, and lines; line physical attributes, such as line length, wire diameter, material, and equivalent impedance spectrum; and operating environment sequences, such as temperature, humidity, load curves, noise peaks, and switchover events. Based on the above data, an electrical heterogeneity graph G is constructed through topology analysis and parameter fusion. _elec The graph structure uses electrical equipment (such as transformers and meters) as nodes and lines or coupling relationships as edges, assigning physical feature vectors to the edges. Simultaneously, it extracts continuous-time driving factors from the operating environment sequence and historical communication sequences to construct an edge-time-varying driving mechanism X. _edge (t) reflects the dynamic changes in the environment in which each communication link is located.

[0068] Step two involves generating a communication performance prediction distribution based on the electrical heterogeneity graph and edge time-varying drive, using a topology-physical consistent continuous-time distribution prediction model. In this step, a deep learning model is constructed to predict the communication performance of each meter over a future period (e.g., the next 15 minutes). This model is topology-physical consistent, learning not only the temporal correlations in the data but also adhering to physical constraints such as transmission line theory. Specifically, the electrical heterogeneity graph G generated in step one... _elec and edge time-varying drive X _edge Using (t) as input, graph time series modeling, for example combining graph neural networks and neural control differential equations, outputs the communication performance prediction distribution F for each meter. _pred The distribution is not a single predicted value, but a probability distribution that includes uncertainty. For example, for each meter, a Beta distribution of success rate and a log-normal distribution of communication time can be provided, from which a lower bound p of the success rate for security decision-making can be obtained. _lower and the upper bound of communication time τ _upper .

[0069] Specifically, based on the electrical heterogeneity graph and edge time-varying drive, a topology-physical consistent continuous-time distribution prediction model is used to generate a communication performance prediction distribution, including:

[0070] By applying neural control differential equations and evolving edge time-varying drivers, time-series representations are generated. Based on the physical properties and time-series representations of electrical heterogeneity graphs, physical consistency constraints of transmission lines are constructed, and physical consistency loss is derived. Through distributed regression heads, time-series representations are mapped, and the distribution of communication performance predictions is defined in a parameterized manner. During model training, the physical consistency loss is used to regularize the model's representation space.

[0071] Step 3: Construct a communication conflict hypergraph based on the electrical heterogeneity graph and the predicted distribution of communication performance. In this step, the electrical heterogeneity graph G from Step 1 is used. _elec The communication performance prediction distribution F in step two _pred This is used to identify which meters are experiencing concurrent communication conflicts. Traditional conflict graphs can only represent pairwise mutual exclusion, while communication conflict hypergraphs (H) can... _conf This can represent multi-factor conflicts, such as the failure of simultaneous communication between any two tables in {Table A, Table B, Table C}. Specifically, this is achieved by analyzing the spectral overlap and physical path overlap between meters (based on the electrical heterogeneity graph G). _elec (Judgment), and combine with the prediction model in step two, for example, use the model to predict the lower bound p of the success rate of each when {A,B} are concurrent. _lower The degree of descent is used to determine the set of meters that may simultaneously experience significant crosstalk, and these sets are defined as hyperedges in the hypergraph.

[0072] Step four: Based on the communication conflict hypergraph, the predicted distribution of communication performance, and the system configuration, a graph structure reinforcement learning scheduling model with hypergraph constraints is used to generate a scheduling plan; wherein, the system configuration includes at least the total time budget of the scheduling window and the upper limit of concurrent meter reading in a single batch.

[0073] System Configuration C _sys It is the set of constraint parameters necessary for the operation of the scheduling model, including at least the total time budget T of the scheduling window. _window This refers to the maximum cumulative duration allowed for meter reading communication within a single scheduling cycle, such as 900 seconds; and the maximum concurrent meter reading limit K for a single batch. _max This refers to the maximum number of meters allowed to initiate communication simultaneously within the same time slot, for example, eight. The system configuration may further include a fairness budget. _budget Minimum success rate threshold p _min Optional parameters, etc.

[0074] In this step, a reinforcement learning decision model (agent) is constructed to replace traditional scheduling logic. This model is used to address system configuration constraints (such as total time budget T). _window Concurrency limit K _max Under the premise of maximizing data collection benefits, such as overall success rate and timeliness, this model is hypergraph-constrained. When making decisions, it must adhere to the communication conflict hypergraph H constructed in step three. _conf The defined mutual exclusion constraint, i.e., a scheduling batch A _batch A graph cannot contain multiple meters originating from the same hyperedge. Specifically, the model uses a communication conflict hypergraph H. _conf Communication performance prediction distribution F _pred (Especially the lower bound of success rate p) _lower and the upper bound of communication time τ _upper Using the state input as the basis, a graph-structured policy network (such as a graph attention network) is used to select a group of meters from the available actions for batch meter reading, ultimately generating a scheduling plan. _t The plan specifies in detail which meters are read in which time slot within the current scheduling window.

[0075] Furthermore, this method also includes an online closed-loop and dual-time-scale self-optimization step. Real-time feedback logs L are collected and recorded after the scheduling plan is executed. _exec This includes actual success / failure, actual time consumption, and error rate. At shorter timescales, such as daily, newly collected real-time feedback logs are used. _exec The prediction model from step two is fine-tuned and calibrated to quickly correct prediction biases. At longer timescales, such as weeks, accumulated real-time feedback logs are utilized. _execThe data is used to retrain and update the policy of the reinforcement learning scheduling model in step four offline, enabling it to adapt to long-term environmental changes. Through self-optimization on two time scales, the system achieves adaptive capability that improves with use.

[0076] Example 3 details the construction of input data, input processing, and time-series evolution of a topologically and physically consistent continuous-time distribution prediction model, addressing the definition and implementation of key concepts such as electrical heterogeneous graphs, edge time-varying drive, and Neural-CDE solution.

[0077] The prediction model in this embodiment takes as input a detailed implementation of the steps for acquiring power grid operation data, constructing an electrical heterogeneity diagram, and implementing edge time-varying drive.

[0078] Specifically, electrical heterogeneity diagram G _elec The construction process includes:

[0079] Node V definition: Define physical entities in a power system as nodes of different types.

[0080] Transformer node _transformer Attributes may include rated capacity (Sn) _kVA (unit: kVA), transformer ratio a, leakage reactance Xσ _pu Parallel magnetization admittance Y _m_s And parting ways.

[0081] Branch box node _branchbox The attribute can include the number of ports n. _ports Dynamic switch state vector _s and grounding status.

[0082] Electricity meter node _meter Attributes may include meter model and coupling circuit type (coupler). _type Installation phase and historical health status.

[0083] Concentrator node _concentrator Attributes may include hardware concurrency capabilities and clock precision.

[0084] Edge E and edge property _attr Definition: Electrical connections between nodes are defined as edges of different types.

[0085] line edge _line edge attribute _attr This can include line length, conductor material (such as copper, aluminum), cross-sectional area, phase, number of branches, and the crucial RLGC parameter per unit length, i.e., the resistance function R. _prime_f (A function of frequency f, Hz -> Ω / km), inductance L_prime (H / km), Capacitance C _prime (F / km) and conductivity G _prime (S / km).

[0086] Transformer coupling edge _transformer Attributes may include turns ratio α, primary / secondary leakage inductance (Lσp, Lσs), and parallel magnetization admittance Y. _m and parasitic capacitance C _par .

[0087] Coupler / filter edge _coupler The attribute can include the passband range B. _pass and insertion loss IL as a function of frequency _f Based on the above definitions, the final generated electrical heterogeneity diagram G _elec =(V,E,node _attr ,edge _attr phase _map ), where node _attr For node attributes, phase _map This phase mapping provides the foundation for subsequent physical constraints and topological aggregation.

[0088] Edge time-varying drive X _edge The construction process of (t) includes:

[0089] Collect dynamic environmental factors that affect PLC channel quality. For example, collect temperature (T). _degC (t), unit °C, humidity H _pct (t), active power of load P _load_W (t) or current I _load_A (t), voltage offset ΔU _V (t), Narrowband noise peak count (N) _peak (t), derived from modem spectrum statistics.

[0090] Data Acquisition Switch Switching Event E _switch (t) represents the 0 / 1 event stream recorded with precision down to the second.

[0091] The above multi-source time series data are time-aligned (based on the concentrator timescale), missing value imputed (e.g., by local regression), and each channel is normalized, for example, to zero mean and unit variance, and the normalization parameter Norm is saved. _edge So that it can be reused during online reasoning.

[0092] In obtaining the electrical heterogeneity diagram G _elec and edge time-varying drive X _edge Following (t), this embodiment describes the preparation stage for applying neural control differential equations:

[0093] Before evolution, piecewise cubic spline fitting is performed on the edge time-varying drive to generate continuous-time control curves.

[0094] Specifically, the discrete-time series edge time-varying driving X obtained in the previous step is used... _edge (t), including channels such as temperature, load, and noise, generates a differentiable, continuous-time control curve X through piecewise cubic spline interpolation. _edge_spline (t). This allows the model to evaluate and differentiate the environmental state at any time point t, and is a necessary input for Neural-CDE.

[0095] Based on the static physical features contained in the electrical heterogeneous graph, the initial hidden state of the neural control differential equation is generated by mapping.

[0096] Specifically, extract the electrical heterogeneity graph G. _elec Static physical characteristics of each edge in the middle. _attr Examples include the length, material, and RLGC parameters of the edge. Static features are input into, for example, a multilayer perceptron, which maps these features to generate the initial hidden state h of the neural control differential equation for that edge. _edge (t0).

[0097] Based on this, this embodiment illustrates the temporal evolution process of applying neural control differential equations, such as... Figure 3 As shown, specifically, by applying neural control differential equations to evolve edge time-varying drivers and generate temporal representations, the following steps are included:

[0098] An adaptive step-size solver is used to numerically solve the edge time-varying drive and obtain intermediate solution results.

[0099] This embodiment preferably uses Neural-CDE (Neural Control Differential Equation) to encode the temporal sequence. The evolution equation can be expressed as: dh _edge (t)=f _θ (h _edge (t))dX _edge_spline (t).

[0100] Among them, h _edge (t) is the hidden state of edge e at time t (the initial hidden state is at t=t0); dX _edge_spline (t) is the derivative of the continuous-time control curve; f _θ It is a vector field defined by the parameter θ.

[0101] Preferably, the vector field f _θ Implemented as a two-layer, multi-layer perceptron, its network structure can be exemplarily set as: d _h ->2d _h ->d_h d _h For the hidden state dimension, the activation function is SiLU (Swish function), and layer normalization is used between layers.

[0102] The adaptive step-size solver preferably uses the Dormand–Prince (DP5) solver.

[0103] To ensure the accuracy and efficiency of the solution, i.e., through an error control mechanism, the solver is configured with strict tolerances. For example, a relative tolerance rel is set. _tol =1e-4, Absolute tolerance abs _tol =1e-6. Set the minimum step size h. _min =0.1 seconds and maximum step size h _max =60 seconds.

[0104] At each step of the numerical solution, the solver predicts the new state h. _pred And estimate the local truncation error err. If err is greater than max(rel) _tol *|h _pred |,abs _tol If the current step size does not meet the accuracy requirements, the solver will reject this step, backtrack, and halve the step size (but not less than h). _min Retry if h is not found; otherwise, accept this step and dynamically adjust the step size for the next step (but not greater than h). _max ).

[0105] At the time of the switch switching event included in the edge time-varying drive, forced point marking and state update are performed on the intermediate solution results.

[0106] Edge time-varying drive X _edge (t) contains the switch switching event E. _switch (t).

[0107] In the evolution of the adaptive solver h _edge During process (t), when the solver's timestamp is about to cross a known event.time (switchover time), the solver is forced to stop at event.time, i.e., forced to mark the point, and a state update is performed on the hidden state h, for example, h=project. _to_event (h,event) is used to inject the impulse effect of the discrete event into the continuously evolving hidden state in real time, and then continue to solve.

[0108] Along the electrical topology path of the electrical heterogeneous graph, the updated intermediate solution results are aggregated to form a time-series representation.

[0109] Through the above evolution, at the target time, such as t at the end of the meter reading window... _end The model is an electrical heterogeneous diagram G._elec Each edge e in the process generates an edge-time embedding Z that incorporates time-varying driving. _edge (t _end ).

[0110] Next, the model is based on the electrical heterogeneity diagram G. _elec Provided electrical topology path P _paths That is, the set of edges traversed from the concentrator to each meter, and the time-series embedding of the edges on the path is Z. _edge (t _end Aggregate.

[0111] Preferably, the aggregation method is topological attention aggregation. For example, the attention weights of the aggregation can be based on the electrical distance d. _elec Alternatively, path impedance summaries can be used for gating and attenuation, with attention scores applied. _edge =q T •W•z _edge -γ•d _elec Where q is the query vector, q T W is the transpose of the query vector, W is the weight matrix, and γ is the decay coefficient.

[0112] Where the electrical distance d _elec Used to measure the degree of electrical coupling between two nodes in a power grid. In this invention, the electrical distance d _elec (i,j) is defined as the cumulative impedance value along the electrical topology path from node i to node j, which is a formula and will not be elaborated here.

[0113] The aggregated result is the final time-series representation E that represents the complete state of the meter's communication link. _path This is used for subsequent distribution prediction and physical consistency constraints.

[0114] Example 4 elaborates on how the constructed prediction model achieves the physical consistency constraint in topology-physical consistency, and describes the implementation details of the physical model H(f,d,Z) and its various formulas.

[0115] In this embodiment, the process of constructing transmission line physical consistency constraints and deriving physical consistency loss based on the physical attributes and time series representation of the electrical heterogeneity diagram is used to ensure the model (which generates the time series representation E) _path The predicted results are consistent with or maintain the order of the expected results derived from physics (especially transmission line theory of power line carrier).

[0116] like Figure 2 As shown, specifically, the construction process includes:

[0117] Step 4.1: Calculate the frequency-related attenuation coefficient based on the per-unit-length parameter of the electrical heterogeneity diagram.

[0118] This step utilizes the electrical heterogeneity diagram G. _elec line edge _line edge property _attr That is, the resistance R per unit length (e.g., per kilometer). _prime_f Inductor L _prime Capacitor C _prime and conductivity G _prime .

[0119] At a given operating frequency f (or frequency grid freq covering the operating bandwidth) _grid Let the angular frequency ω = 2πf.

[0120] First, calculate the complex propagation constant γ(f), γ(f) = sqrt((R _prime_f (f)+jωL _prime )*(G _prime +jωC _prime )); where R _prime_f (f) is a function of frequency to account for the skin effect, for example, R'(f)≈R _dc •(1+k _skin •sqrt(f)), where R'(f) is the AC resistance per unit length of the transmission line at frequency f, R _dc It is the DC resistance per unit length of the transmission line, k. _skin It is the skin effect coefficient.

[0121] Extracting the real part, we obtain the frequency-dependent attenuation coefficient α(f) (unit: Neper / meter): α(f) = Re(γ(f)).

[0122] For a line of length d, its amplitude squared attenuation (power attenuation) is H. _line (f)=exp(-2*α(f)•d).

[0123] Step 4.2: Determine the device transfer function based on the device parameters of the electrical heterogeneity diagram.

[0124] This step utilizes the electrical heterogeneity diagram G. _elec edge in _transformer (Transformer) and edge _coupler (Coupler) edge _attr (Edge attribute).

[0125] Calculate the transformer transmission T _i (f) Based on its small-signal equivalent circuit parameters, such as turns ratio a, leakage inductance Lσp / Lσs, and parallel admittance Y _m Parasitic capacitance C _parCalculate its transfer function. For example, consider the downstream load impedance Z. _down Mapping to the primary side yields the load impedance Z mapped to the primary side. _Lp =a 2 *Z _down And consider leakage inductance Z _leak and parallel impedance Z _par Calculate the voltage transfer ratio H _tr_amp =|Z _Lp / (Z _source +Z _leak +(Z _par ||Z _Lp ))|, where Z _source The source impedance is denoted as .

[0126] Calculate the transfer T of the coupler / filter _i (f), based on its edge attribute edge _attr Insertion loss IL _f (Unit: dB), converted to linear amplitude IL _linear =10 (-IL_f(f) / 20) .

[0127] Along electrical topology path P _paths Multiplying the transfer functions (amplitude squared) of all devices on the device gives the total device transfer attenuation H. _dev (f)=Π(T _i (f) 2 ).

[0128] Step 4.3: Determine the termination reflection coefficient based on the load impedance and characteristic impedance of the electrical heterogeneity diagram.

[0129] Calculate the characteristic impedance Z of the line _c (f), Z _c (f)=sqrt((R _prime_f (f)+jωL _prime ) / (G _prime +jωC _prime ));

[0130] Obtain the load impedance Z of the communication link terminal _load (f) This impedance can be derived from the electrical heterogeneity diagram G. _elec electricity meter node _meter The attribute, or a typical value for the station configuration.

[0131] Calculate the voltage reflection coefficient Γ(f), Γ(f) = (Z _load (f)-Z _c (f)) / (Z _load (f)+Z _c (f));

[0132] Calculate the termination reflection coefficient (here referring to the power attenuation caused by reflection) A _ref (f), A _ref (f)=1-|Γ(f)| 2 .

[0133] Step 4.4: Combine the frequency-dependent attenuation coefficient, device transfer function, and termination reflection coefficient to derive the physical expected signal-to-noise ratio.

[0134] Combining all the attenuation terms above, we obtain the total path power gain (attenuation) H. _total_mag2 (f):

[0135] H _total_mag2 (f)=H _line (f)×H _dev (f)×A _ref (f).

[0136] Introducing the transmit power spectrum S _tx (f) (e.g., the concentrator's nominal transmit power) and the ambient noise power spectrum N(f) (which can be obtained from the edge-time-varying drive X) _edge Noise channel estimation in (t).

[0137] Calculate the received signal power S _rx and received noise power N _rx That is, in the working bandwidth B _pass Points accumulation:

[0138] S _rx =∑ _f (H _total_mag2 [f]*S _tx_f (f)*Δf); N _rx =∑ _f (N _f (f)*Δf); S _tx_f (f) is the transmit power spectrum at frequency f, N _f (f) represents the ambient noise power spectrum at frequency f.

[0139] The physical expected signal-to-noise ratio (SNR) is derived. _phys (Unit: dB): SNR _phys =10*log10(S _rx / N _rx ).

[0140] Step 4.5: Based on the model predicted SNR obtained by mapping the physical expected SNR to the temporal representation, construct the physical consistency loss. The order-preserving constraint only penalizes cases where the model predicted SNR is higher than the physical expected SNR.

[0141] The model has generated a time series representation E _path.

[0142] Construct simple mapping layers, such as linear layers g(), to extract time-series representations E _path Predictive model predicts signal-to-noise ratio (SNR) _pred (Logarithmic scale).

[0143] Constructing physical consistency loss L _phy Preferably, the loss is an order-preserving constraint loss:

[0144] L _phy =mean([log(SNR _pred )-log(SNR _phys )] + );

[0145] in[...] + It's the ReLU function, i.e., max(0,...). The meaning of this loss is: the SNR predicted by the model. _pred It can be lower than the physical expected SNR _phys This is because the model can learn instantaneous disturbances not included in the physical model; however, if the SNR... _pred Higher than SNR _phys This violates the laws of physics, meaning the model imagines a non-existent channel gain, in which case L... _phy This will result in a positive penalty.

[0146] Building upon this, during model training, the physical consistency loss is used to regularize the model's representation space. Specifically, the physical consistency loss L... _phy (may include a weighting coefficient λ) _phy Added to the total loss function L of the model _total middle.

[0147] To illustrate the calculations of the physical model more concretely, a simplified example is provided. Assume a certain line segment has the following characteristics: AC resistance R' = 0.4 Ω / km, inductance L' = 0.7 mH / km, capacitance C' = 0.05 μF / km, conductance G' = 1e⁻⁶ S / km, path length d = 0.25 km, and target frequency f = 100 kHz. Calculations in step 4.1 show that α(f) is approximately 0.015 dB / km. Therefore, the line attenuation H for this segment... _line Approximately 0.0038 dB (amplitude ratio approximately 0.9996). Assume the transmission amplitude T of a transformer along the path at 100 kHz. _i The value is 0.92. Assume the termination reflection attenuation A calculated in step 4.3 is... _ref The value is 0.90. Therefore, the total path magnitude transfer |H _total|≈0.9996×0.92×0.90≈0.828. |H _total | Will be used in step 4.4 to calculate SNR _phys And as the physical consistency loss L in step 4.5 _phy The basis for the constraint.

[0148] If a symmetrical mean square error loss L is used _phy_mse =mean((SNR _pred -SNR _phys ) 2 If the actual channel has transient interference not considered in the physical model, such as impulse noise generated by a neighbor's appliance starting up, the actual SNR will be lower than the physical expected SNR. _phys If the model predicts a lower SNR at this time _pred Instead, it will be penalized, preventing the model from learning these additional interfering factors.

[0149] Using order-preserving constraints, only SNR is penalized. _pred >SNR _phys In certain situations, the model is allowed to learn negative factors not covered by the physical model, such as unknown noise sources; the model is prohibited from predicting optimistic results beyond physical limits, such as transmission without attenuation; and scheduling decisions based on model predictions are guaranteed not to fail due to over-optimism.

[0150] Assume the physical expected SNR of a certain electricity meter link _phys =15dB.

[0151] Scenario 1: Model predicts SNR _pred =12dB (below physical expectations).

[0152] Order-preserving loss: [log(12)-log(15)] + =[-0.22] + =0 (no penalty).

[0153] In this situation, it is reasonable for the model to have identified additional sources of interference.

[0154] Scenario 2: Model Predicts SNR _pred =18dB (higher than physical expectation).

[0155] Order-preserving loss: [log(18)-log(15)] + =[0.18] + =0.18 (with penalty).

[0156] In this case, the model predicts a physically impossible channel gain, which must be corrected.

[0157] Example 5: Detailed Explanation of the Topologically and Physically Consistent Continuous-Time Distribution Prediction Model Θ _A How to generate the final prediction output, and how to train it using a preferred data augmentation method.

[0158] In this embodiment, the model maps time-series representations through a distributed regression head, defining the distribution of communication performance predictions in a parameterized manner.

[0159] Specifically, time series representations (e.g., time series representation E) _path The input is one or more distribution regression heads, such as those implemented by a multilayer perceptron. Its goal is not to predict a single numerical value, such as a success rate of 0.8, but rather to predict the parameters of a complete probability distribution.

[0160] In a preferred implementation:

[0161] Regarding communication success rate, the distributed regression head represents the time series E _path The mapping is to two positive parameters, α _meter and β _meter α can be output through a multilayer perceptron. _meter_raw and β _meter_raw Then, the softplus activation function is applied to ensure that it is greater than zero. These two parameters define a Beta(α) value. _meter ,β _meter The Beta distribution is naturally applicable to the [0,1] interval and is suitable for analyzing success rates p. _meter Modeling is then performed. At this point, the estimated success rate p predicted by the model is... _hat Can be derived from α _meter / (α _meter +β _meter (This is given.)

[0162] Regarding communication latency, the distributed regression head will represent the time series E _path (or its temporal embedding with edge Z) _edge (t _end The combination of ) is mapped to two parameters, μ _meter and σ _meter (σ) _meter (Must be greater than 0). These two parameters define a LogNormal(μ) _meter ,σ _meter The log-normal distribution is suitable for modeling communication latency τ, which is always positive and may have a long tail. _meter At this point, the estimated time τ predicted by the model... _hat It can be derived from exp(μ) _meter +0.5*σ _meter 2 (This is given.)

[0163] Using the parameterization method described above, the model not only provides the predicted mean but also a measure of the uncertainty surrounding that prediction, such as the variance or width of the distribution. This uncertainty information is used for subsequent robust scheduling.

[0164] The prediction loss L used during model training _pred Preferably, it is a hybrid loss function. For example:

[0165] L _pred =NLL _β +NLL _lognormal +λ _mse *(MSE _p +MSE _τ );

[0166] Among them, NLL _β Is the observed true success or failure (0 or 1) relative to Beta(α)? _meter ,β _meter Negative log-likelihood loss of NLL distribution; _lognormal It is the observed actual communication time relative to LogNormal(μ _meter ,σ _meter Negative log-likelihood loss of the MSE distribution; _p and MSE _τ They are point estimates p _hat τ _hat Mean squared error loss between the true value and the actual value; λ _mse It is a weighting coefficient used to balance the distribution likelihood and the accuracy of point estimation.

[0167] Furthermore, the topologically and physically consistent continuous-time distribution prediction model Θ _A The training also includes topology intervention data augmentation. In power systems, topology changes (such as switch opening and closing) are low-frequency events, resulting in a severe lack of such samples in the training data. This makes it difficult for the model to learn the real impact of topology changes on communication, and it may only learn some spurious correlations.

[0168] To address this issue, this embodiment employs the following data augmentation steps:

[0169] Simulate topology changes in electrical heterogeneity diagrams and generate a list of topology interventions.

[0170] Specifically, based on the electrical heterogeneity diagram G _elec It programmatically simulates various possible topology changes that conform to electrical operating rules.

[0171] Case 1: Simulating a branch bin node _branchbox A vector of port switch states _sChanging from 1 (closed) to 0 (open) will disconnect all meters under that branch.

[0172] Case 2: Simulating an edge of a certain line _line The impedance parameter experiences a small disturbance, for example, the resistance per unit length R... _prime_f Increase by 5% to simulate line aging or temporary splicing.

[0173] Filter out all simulated operations that could lead to illegal states such as ring networks or short circuits, and generate a topology intervention list I. _list .

[0174] Based on the topological intervention list and by reusing the process of deriving the physical expected signal-to-noise ratio, the counterfactual communication metric is re-estimated.

[0175] For topological intervention list I _list For each topological intervention, the model does not rely on (and cannot rely on) real observation data, but reuses the physical model H(f,d,Z).

[0176] Specifically, the model re-runs the process of deriving the physical expected signal-to-noise ratio on the intervened topology, and recalculates the total path power gain H of the affected path. _total_mag2 (f) and derive the new physical expected signal-to-noise ratio (SNR). _phys .

[0177] Newly derived SNR _phys The success rate and time consumption of the approximate correction are known as counterfactual communication metrics.

[0178] Counterfactual samples are generated based on counterfactual communication metrics.

[0179] The counterfactual communication metric is combined with the original environment-driven edge time-varying driver X. _edge By combining (t), a new training sample, namely the counterfactual sample E, is synthesized. _counterfactual .

[0180] Preferably, a credibility weight w can be assigned to the counterfactual sample. _cf For example, 0.8, to indicate that it is based on a physical model simulation rather than actual observation.

[0181] We use counterfactual samples and real samples to train the topologically and physically consistent continuous-time distribution prediction model alternately.

[0182] In the training loop, the real sample E _comm Counterfactual sample E _counterfactual Mixed phases.

[0183] A preferred training strategy is alternating training; for example, the model is trained N times on real samples, e.g., N=10, and then on counterfactual samples (using confidence weights w). _cf (Weighted loss) Training step 1.

[0184] This training forces the model to learn the patterns in real data, and also to ensure that its predictions remain consistent with the derivation of the physical model H(f,d,Z) when topological changes occur.

[0185] Data augmentation through topological intervention improves the model's generalization ability and prediction accuracy when faced with low-frequency topological changes in the real world.

[0186] Example 6: A detailed explanation of the specific implementation method for constructing a communication conflict hypergraph. The input for this step is based on the electrical heterogeneity graph G. _elec And communication performance prediction distribution F _pred .

[0187] In traditional scheduling, conflicts are typically modeled as graphs, where a conflict exists between meter A and meter B (one edge). However, in PLC communication, conflicts are often multi-faceted. For example, A and B can coexist alone, A and C can coexist, and B and C can coexist, but simultaneous coexistence of A, B, and C will lead to channel collapse. Traditional graphs cannot represent such ternary conflicts.

[0188] This embodiment uses a hypergraph to solve this problem. (Communication conflict hypergraph H) _conf =(V _meters E _hyper ), its vertex V _meters It is all the meter nodes, and its superedge E _hyper It is a set of sets, E _hyper ={ε _1 ,ε _2 ,...}, where each hyperedge ε _i (e.g., ε) _1 ={tableA,tableB,tableC}) represents a conflict cluster, meaning that the meters in this set are not allowed to be read concurrently in the same scheduling batch, or the number of concurrent reads is limited, for example, at most 1.

[0189] Constructing a communication conflict hypergraph includes:

[0190] Step 6.1: Assess the physical path overlap between the meters.

[0191] Specifically, based on the electrical heterogeneity diagram G _elec and electrical topology path P _paths Calculate the communication path P of any two meters i and j originating from the concentrator. _i and P _j Shared physical line (edge)_line )length.

[0192] Define physical path overlap _ij =len(P _i ∩P _j ) / min(len(P _i ),len(P _j )).

[0193] Set threshold θ _overlap For example, θ _overlap =0.5. If overlap _ij ≥θ _overlap If i and j are considered to be strongly path-coupled, then i and j are considered to be strongly path-coupled.

[0194] Step 6.2: Assess the spectral overlap between the meters.

[0195] Specifically, obtain the operating frequency band B of meters i and j. _pass Or its power spectral density PSD function PSD _i (f) and PSD _j (f) (Available from device metadata or historical estimates).

[0196] Calculate the spectral overlap S _ov_ij =∫ _B PSD _i (f)•PSD _j (f)df.

[0197] Set the normalization threshold θ _spec For example, θ _spec =0.3. If S _ov_ij ≥θ _spec If i and j are considered spectrum competitors, then i and j are considered spectrum competitors.

[0198] Step 6.3: Using a topology-physical consistent continuous-time distribution prediction model, estimate the predicted signal-to-noise ratio drop or the predicted collision probability during concurrent communication.

[0199] Conflict determination is no longer static but dynamic, based on a topologically and physically consistent continuous-time distribution prediction model Θ. _A Make a judgment.

[0200] Specifically, the model can be configured to predict counterfactual concurrent scenarios: when meters i and j communicate simultaneously, the model predicts the signal-to-noise ratio (SNR) of i. _pred (or the lower bound of success rate p) _lower How much will it decrease? Obtain the predicted signal-to-noise ratio decrease value ΔSNR. _pred_ij .

[0201] Set threshold θ_snr For example, θ _snr =-3dB. If ΔSNR _pred_ij ≤θ _snr If so, it is believed that the concurrency of i and j will lead to significant crosstalk.

[0202] In another alternative approach, the model can also directly estimate the predicted collision probability P. _col_ij (Based on historical slot conflict replay data trained), and a threshold θ is set. _col For example, θ _col =0.2.

[0203] Step 6.4: Based on physical path overlap, spectral overlap, and predicted signal-to-noise ratio decrease or predicted collision probability, determine the conflict relationship to construct a communication conflict hypergraph.

[0204] A preferred rule for constructing superedges is: if two meters i and j satisfy overlap _ij ≥θ _overlap And (S) _ov_ij ≥θ _spec (or phases are consistent) and (ΔSNR) _pred_ij ≤θ _snr or P _col_ij ≥θ _col If i and j are in a strong conflict, then i and j are determined to be in a strong conflict.

[0205] Place all meter pairs {i,j} with strong conflicts into an initial conflict set.

[0206] Furthermore, using the maximal clique algorithm in graph theory, we find cliques on the conflict graph, that is, nodes in the subgraph that are in pairwise conflict. For example, if {i,j}, {j,k}, and {i,k} all satisfy the above conflict condition, then {i,j,k} constitutes a ternary maximal clique.

[0207] Each found maximal clique, or a multi-set sharing the same bottleneck (like a branch box or the same phase), is defined as a hyperedge ε. _k And added to the communication conflict supergraph H _conf The set of superedges E _hyper middle.

[0208] Preferably, the communication conflict supergraph H _conf It is not static, but dynamically updated. For example, a topologically and physically consistent continuous-time distribution prediction model Θ is updated daily or whenever needed. _A After the update, a supermap reconstruction will be triggered.

[0209] Completed communication conflict hypergraph H _conf It will serve as a key structured input to the reinforcement learning scheduling model, used to constrain its action space.

[0210] Example 7 details the implementation of a graph-structured reinforcement learning scheduling model employing hypergraph constraints. The model predicts the distribution F of received communication performance. _pred Hypergraph H with communication conflicts _conf As input.

[0211] The method in this embodiment includes the following steps:

[0212] Based on the communication conflict hypergraph, communication performance prediction distribution, and system configuration, an action-feasible domain mask and candidate set are generated. This process includes:

[0213] A mutual exclusion matrix is ​​derived from the communication conflict hypergraph and used as the first layer of constraint for the action feasible region mask.

[0214] Based on the upper bound of communication time in the communication performance prediction distribution and the time budget in the system configuration, the time feasibility constraint is calculated as the second layer of constraint in the action feasibility domain mask.

[0215] Based on the concurrency limit in the system configuration, a capacity constraint is set as the third layer of constraint in the action feasible domain mask;

[0216] Generate an action-feasible domain mask by combining three layers of constraints.

[0217] Specifically, the system configuration C _sys Including total time budget T _window (e.g., 900 seconds), concurrency limit K _max (e.g., 8 concurrent connections) etc. Candidate set S _cand The generation preferably embodies security-side filtering: the system traverses all meters to be read and only reads meters whose confidence bands meet the security threshold (e.g., p). _lower ≥p _min The minimum success rate threshold p _min It can be set to 0.6; and the uncertainty U _uncert ≤u _max ) included in candidate set S _cand This step can be seen as an early failure elimination mechanism, preventing the reinforcement learning model from wasting exploration budget on obviously infeasible meters. Action feasible region mask M _mask It is a dynamically updated Boolean matrix or vector used to shield illegal actions in real time during the sampling process, combining (1) the communication conflict hypergraph H _conf The derived mutual exclusion matrix C _mutex (1) Do not select conflicting nodes in the selected batch; (2) The total time budget T _window and the upper bound of communication time τ _upper (From the communication performance prediction distribution F) _pred The time feasibility of the calculation C _time; and (3) the concurrency limit K _max .

[0218] Using the structure of the communication conflict hypergraph as input, a graph structure policy network is constructed. Specifically, this policy network is preferably implemented using a graph neural network (GNN) or a graph attention network (GAT). Its adjacency matrix (or graph structure) is the communication conflict hypergraph H. _conf The induced graph. The input feature vector for each node (meter) can be constructed as follows:

[0219] [p _lower ,τ _upper ,uncert,staleness,node _degree ]; where the lower bound of the success rate p _lower Upper bound of communication time τ _upper Uncertainty stems from the communication performance prediction distribution F _pred Data freshness (staleness) comes from historical communication sequences; node degree _degree From the communication conflict hypergraph H _conf After being aggregated by graph convolution or graph attention layers, the network outputs node policy embeddings H. _actor and node value embedding H _critic .

[0220] Graph-structured policy networks sample candidate sets under the constraint of action feasible region masks to generate scheduling plans. The sampling process preferably employs a two-stage sampling strategy, first clusters then nodes, to achieve action space compression.

[0221] First, based on the mutual exclusion clusters derived from the communication conflict hypergraph, and constrained by the action feasible region mask, the target cluster is sampled. Here, the mutual exclusion cluster is... _id It is a communication conflict hypergraph H _conf One approach is to decompose (e.g., a set of meters sharing the same branch box or the same superedge). The policy network first calculates the probability π of each cluster being selected. _cluster For example, based on the lower bound p of the maximum success rate of nodes within the cluster. _lower Alternatively, the maximum policy logit synthesis (logit refers to the raw score or unnormalized log probability calculated by the policy network for each basic unit (such as a single meter node)) is used to sample a target cluster.

[0222] Within the selected target cluster, and constrained by the action-feasible region mask, target nodes are sampled. Within the target cluster selected in the previous step, the policy network samples nodes based on the node selection probability π. _node Sampling target node i. Here, π _node Through H _actor The logit obtained from the mapping is generated using the mask Softmax, i.e., the logit_masked =logit(assign value), and set the action feasible region mask M. _mask All illegal nodes in the set S (e.g., already selected, conflicting, or not in the candidate set S) _cand The logit value of (in the middle) is set to negative infinity (-∞).

[0223] After sampling the target nodes, based on the upper bound of the communication time in the confidence band of the communication performance prediction distribution, an immediate time knapsack check is performed to ensure that the cumulative communication time does not exceed the time budget in the system configuration, and a scheduling plan is generated. This is an iterative construction of scheduling batch A. _batch The process. Let A _batch ={} ({} is an empty set), T _remain =T _window T _remain The remaining time. After sampling the target node i, perform the following check: if τ _upper_i ≤T _remain , τ _upper_i To account for time costs, i will be added to scheduling batch A. _batch and update T _remain =T _remain -τ _upper_i Simultaneously update the action feasible region mask M. _mask If the condition is false (meaning insufficient time remains), then node i is rejected, and (optionally) resampled in the current cluster or all clusters. This process is iterated until the remaining time T is reached. _remain Exhausted or |A _batch | Reached the concurrency limit K _max .

[0224] Preferably, to ensure the robustness of the algorithm, an upper limit is set for the number of resampling attempts, for example, 5 times. If consecutive resampling attempts fail, or if it is found that the remaining time T is... _remain It is already smaller than the candidate set S _cand τ of all remaining nodes _upper_min (Upper bound of communication time τ) _upper If the minimum value is not reached, it is considered a deadlock or the knapsack is full, and the construction of this batch is immediately terminated. The final generated scheduling batch A _batch The execution sequence of these components constitutes the scheduling plan. _t .

[0225] In a preferred embodiment, to further improve exploration efficiency, the logit computation of the policy network can introduce graph prior-driven optimistic exploration. Specifically, logit _i =logit _i_base +β•b _i Among them, logit _i_base It is the basic output of the policy network; β is the exploration coefficient; b_i It's about exploring the dividends. _i =c1•sqrt(U _uncert_i )+c2•centrality _i Among them, U _uncert_i That is, predicting uncertainty; centrality _i Is node i in the communication conflict hypergraph H? _conf The centrality of nodes, such as node degree; c1 and c2 are weight coefficients. Through the above process, the model prioritizes exploring (1) nodes with high prediction uncertainty (high risk and high return) and (2) nodes that are critical in the topology (high centrality), thus accelerating the learning of effective strategies.

[0226] like Figure 4 As shown, the training of the graph-structured policy network includes:

[0227] Identify homogeneous mutually exclusive clusters in the communication conflict hypergraph and enable these clusters to share policy sub-network parameters. _conf Clusters with the same topology and similar physical properties are considered in a graph neural network, such as multiple star-shaped structures with one branch box and three meters. By calculating the structural signatures of these clusters, clusters with the same signature can share the same sub-network weights (policy network Φ) within the graph neural network. _B This (part of the process) reduces the number of model parameters and improves the model's generalization ability to newly added meters with similar topologies. The structural signature is a part of the communication conflict hypergraph H. _conf The topology of each mutually exclusive cluster is encoded into a feature vector, which is used to identify clusters with the same or similar structures, enabling them to share policy subnetwork parameters. The structural signature typically includes a combination of encoded features such as node degree sequence statistics, subgraph size, and connection patterns. These will not be elaborated upon here.

[0228] The cluster-level policy entropy of the graph-structured policy network on isomorphic mutually exclusive clusters is calculated to obtain the structure entropy regularization term. During training of the graph-structured policy network, the structure entropy regularization term is added to the training loss function. An additional structure entropy regularization term L is added to the training objective function. _struct =-τ•H(π _over_clusters H(π) _over_clusters ) represents the selection entropy of the policy network at the cluster level; τ is the weight coefficient. This process encourages the policy network to explore diverse clusters rather than repeatedly optimizing nodes within a single cluster, thus improving the exploration coverage at the structural level.

[0229] According to one aspect of this application, the action-feasibility field mask M _mask It can also be a dynamically maintained Boolean vector with a length equal to the total number of meters N. _mask[i]=True indicates that meter i is currently selectable, while False indicates that it is blocked.

[0230] M _mask The initialization and update follow the logic and operation process of three-level constraints as follows:

[0231] First layer: Hypergraph mutual exclusion constraint M _mutex .

[0232] From the communication conflict hypergraph H _conf Extracting the mutual exclusion matrix C _mutex [i,j];

[0233] When meter j has been selected into the current batch A _batch When, all satisfying C _mutex The meter i with [i,j]=1 is blocked;

[0234] M _mutex [i]=∏ j∈A_batch (1-C _mutex [i,j]).

[0235] Second layer: Time feasibility constraint M _time .

[0236] For each candidate meter i, check its upper bound τ for communication time. _upper [i] Whether the remaining time T has been exceeded _remain ;

[0237] M _time [i]=1 if τ _upper [i]≤T _remain else 0.

[0238] Third layer: Capacity constraint M _cap .

[0239] When the selected batch size | A _batch | Reached the concurrency limit K _max At that time, all nodes are blocked.

[0240] M _cap [i]=1 if |A _batch | <K _max else 0.

[0241] Composite mask: M _mask =M _mutex and M _time and M _cap and M _cand ;and is the logical AND operator.

[0242] Where M _candIt is the candidate set S _cand The indicator vector.

[0243] In a certain scenario, there are 5 electricity meters {A, B, C, D, E}, with a maximum concurrent connection K. _max =3, Time Budget T _window =100 seconds.

[0244] Hypergraph hyperedges: ε1={A,B,C} (the three are mutually exclusive), ε2={D,E} (the two are mutually exclusive).

[0245] Upper bound of predicted communication time: τ _upper ={A:30,B:25,C:40,D:35,E:20}.

[0246] Initial state:

[0247] A _batch ={},T _remain =100; M _mask =[1,1,1,1,1] (all optional).

[0248] Step 1: Select meter A for sampling.

[0249] A _batch ={A},T _remain =100-30=70;

[0250] Update M _mutex A is mutually exclusive with B and C -> B and C are blocked;

[0251] Update M _time Check the τ of each node. _upper Is it less than T? _remain =70;

[0252] τ _upper [C]=40<70? Yes, the timeframe is feasible;

[0253] τ _upper [D]=35<70? Yes, the timeframe is feasible;

[0254] τ _upper [E]=20<70? Yes, time is feasible;

[0255] M _mask =[0,0,0,1,1] (D and E are optional only).

[0256] Step 2: Select meter E for sampling.

[0257] A _batch ={A,E},T _remain =70-20=50; Update M _mutexE and D are mutually exclusive -> D is masked;

[0258] M _mask =[0,0,0,0,0] (No optional nodes);

[0259] Step 3: M is detected _mask All zeros, terminate this batch build.

[0260] Output scheduling batch A _batch ={A,E}, with an upper limit of cumulative time of 50 seconds.

[0261] According to one aspect of this application, the time knapsack check in this scheme uses an upper bound τ for communication time consumption. _upper (The upper boundary of the confidence band), rather than the predicted mean τ _hat .

[0262] The predicted distribution of meter X is LogNormal (μ=3.0, σ=0.5); the predicted mean is τ. _hat =exp(3.0+0.5×0.25)=22.3 seconds; 90% confidence upper bound τ _upper =exp(3.0+1.28×0.5)=38.1 seconds;

[0263] If the remaining time T _remain =35 seconds:

[0264] Using τ _hat Judgment: 22.3 < 35, selection allowed -> there is approximately a 10% probability of timeout;

[0265] Using τ _upper Judgment: 38.1>35, reject selection -> guarantee 90% probability of not timeout.

[0266] This scheme adopts τ _upper The conservative strategy ensures that the scheduling plan meets the time constraints under high confidence.

[0267] Example 8 details the process of secure coupling and closed-loop training between the prediction model and the scheduling model through confidence bands and Lagrange constraints.

[0268] The communication performance prediction distribution includes a confidence band generated through a conformal calibration method that guarantees coverage. This is a post-processing step for the parameterized distribution Beta(α,β) and the log-normal distribution LogNormal(μ,σ). Specifically, the conformal calibration does not rely on the accuracy of the distribution assumptions but provides a statistical guarantee of coverage using historical data.

[0269] In other words, the conformal calibration method for ensuring coverage includes: calculating the residual between the model prediction and the actual observation on the calibration dataset, determining the quantile of the residual at a preset coverage rate as a confidence coefficient, and generating a confidence band with statistical coverage guarantee based on the confidence coefficient;

[0270] The conformal calibration method includes the following steps: (1) dividing the historical training data into a training subset Train _set and calibration subset Cal _set (2) In the training subset Train _set Training a topologically and physically consistent continuous-time distribution prediction model Θ _A (3) In the calibration subset Cal _set Above, calculate the point estimate p of the model. _hat μ _meter Compared with the actual observed value p _obs τ _obs The residual measure between. For example, the success rate residual r. _p =|p _obs -p _hat |;Time-consuming residual r _τ =|log(τ _obs )-μ _meter |。 (4) Set the desired coverage rate 1-δ (e.g., 90%, i.e., δ=0.1). Calculate r _p and r _τ In Cal _set The confidence coefficient q is obtained from the (1-δ) quantile (i.e., the 90th quantile) on the right. _p and q _τ (5) During online prediction, the model outputs p _hat and μ _meter and using the calibrated q _p q _τ Generate a confidence band with a 1-δ coverage guarantee.

[0271] For example, the confidence band includes:

[0272] Lower bound of success rate: p _lower =max(0,p _hat -q _p );

[0273] Upper bound of success rate: p _upper =min(1,p _hat +q _p );

[0274] Lower bound of communication time: τ _lower =exp(μ _meter -q _τ );

[0275] Upper bound of communication time: τ _upper =exp(μ _meter +q _τ );

[0276] Uncertainty Measure U _uncert It can also be defined in this way, for example:

[0277] U _uncert =(p _upper -p _lower )+k k •(log(τ _upper )-log(τ _lower )), k k It is the adjustment coefficient.

[0278] The aforementioned confidence band is used in the scheduling model.

[0279] On one hand, a graph-structured reinforcement learning scheduling model with hypergraph constraints uses a confidence band to perform safety-side filtering on electricity meters to generate a candidate set. The candidate set S _cand The generation depends on the lower bound of the success rate p. _lower and uncertainty measure U _uncert This ensures that the scheduler only selects from candidates with statistical certainty.

[0280] On the other hand, the scheduling model also utilizes the upper bound of communication time in the confidence band to calculate constraint violation degree for Lagrangian constraint optimization. In the time knapsack check and subsequent training, the model does not use the average prediction time τ. _hat Instead, it uses an upper bound τ for communication time with a 90% coverage guarantee. _upper .

[0281] Lagrange constraint optimization is the training process of a graph structure reinforcement learning scheduling model that uses hypergraph constraints. Before training begins, the reward function r(X) and the constraint violation degree g(X) are defined.

[0282] Preferably, the batch revenue r(X) is defined as:

[0283] r(X)=Σ _i∈X [w _p •p _lower_i -w _τ •(τ _upper_i / T _window )-w _u •uncert _i -w _s •staleness _i ];

[0284] Where, Σ _i∈XThis represents the summation of all meters i in batch X; w _p ,w _τ ,w _u ,w _s These are the weighting coefficients for each item; p _lower_i It is a successful gain; τ _upper_i It's the time cost; uncert _i It is a risk penalty; staleness _i It is the cost of timeliness.

[0285] Preferably, the constraint violation degree g _k (X) is defined as:

[0286] Time budget constraint violation degree: g _time (X)=max(0,Σ _i∈X τ _upper_i -T _window );

[0287] Communication conflict constraint violation degree: g _conf (X) represents the hypergraph H that violates communication conflict in batch X. _conf Counting the superedges of a mutual exclusion constraint;

[0288] Concurrency upper limit constraint violation degree: g _cap (X)=max(0,|X|-K _max ), where |X| is the number of nodes contained in batch X (the cardinality of the set);

[0289] Degree of violation of fairness constraints: g _fair (X)=max(0,Σ _i∈X fairness _deficit_i -Fair _budget ), fairness _deficit_i It is the fairness deficit of node i, Fair _budget It is a fairness budget.

[0290] The training process specifically includes:

[0291] Real-time feedback log L after obtaining the execution scheduling plan _exec L _exec Recorded scheduling batch A _batch The actual success or failure of each meter, and the actual communication time τ _actual wait.

[0292] The actual batch return r is calculated based on the real-time feedback logs. _real r _real The calculation formula is similar to that of batch revenue r(X), but it uses real-time feedback logs L. _exec The actual observed value in the data._success (The true Boolean result of whether a node succeeded in this scheduling) Replacement success rate lower bound p _lower , τ _actual Replace the upper bound τ of communication time _upper .

[0293] PPO loss L is constructed based on actual batch returns. _PPO Utilizing actual batch returns r _real Calculate the advantage function Adv, for example using generalized advantage estimation (GAE). Construct the pruning target loss L for PPO. _clip That is, L _PPO :

[0294] L _PPO =E[min(r _t •Adv,clip(r _t ,1-ε,1+ε)•Adv)];where, E[...] represents the expectation; r _t It is the probability ratio of the new and old strategies; ε is the pruning coefficient, for example, 0.2; Adv is the advantage function.

[0295] The PPO loss is combined with the constraint violation weighted by Lagrange multipliers to form the total loss. The Lagrange multiplier λ is a vector, λ = {λ...} _time ,λ _conf ,λ _cap ,λ _fair}, λ _time , λ _conf , λ _cap , λ _fair These are the Lagrange multipliers corresponding to the time budget constraint, communication conflict constraint, concurrency cap constraint, and fairness constraint, respectively. Total loss L _total Defined as:

[0296] L _total =L _PPO +Σ _k λ _k •E[g _k (X)]. In practical gradient descent, the goal is to maximize the gain and minimize constraint violations, therefore the total loss function is L. _total =-L _PPO +Σ _k λ _k •E[g _k (X)].

[0297] The graph structure reinforcement learning scheduling model for hypergraph constraints is updated using the total loss. That is, the total loss L is backpropagated. _total The gradient is used to update the policy network Φ. _B Value Network V _ The parameter of φ.

[0298] The Lagrange multipliers are adaptively updated based on the constraint violation degree. Multiplier λ _k These are not fixed hyperparameters, but are learned (updated) synchronously. The update rule is (after performing one model update):

[0299] λ _k <-max(0,λ _k +η _λ •E[g _k (X)]), where η _λ It is the learning rate of the multiplier. The rule of the adaptive update mechanism is that if constraint k (e.g., time budget constraint violated g) is updated, the update will continue until the constraint k is updated. _time This batch was violated (E[g) _time If λ > 0, then λ _time It will automatically increase. In the next round of training, λ _time •E[g _time In L _total The penalty weight in the network will automatically increase, forcing the policy network Φ to... _B Pay extra attention to avoid violating time constraints.

[0300] Example 9 illustrates the process by which the method of the present invention forms a complete, online-running self-optimizing closed loop.

[0301] This embodiment utilizes real-time feedback log L _exec The two models in the system, namely the prediction model and the scheduling model, are continuously updated on a rolling basis with two time scales.

[0302] Specifically, this self-optimizing closed loop includes:

[0303] First, perform continuous archiving of data logs. Real-time feedback log L _exec This includes timestamps, meter ID, batch ID, actual success / failure, actual time elapsed, bit error rate, number of retransmissions, etc., which are written back in real time and merged into the long-term historical communication sequence D. _comm In the database. D _comm This provides the data foundation for the two subsequent updates.

[0304] Perform daily (short-cycle) fine-tuning to correct prediction biases. This step is to quickly respond to short-term environmental drift, such as minor changes in network-wide noise levels due to weather variations. The updated data is a topology-physically consistent continuous-time distribution prediction model. This update cycle can be set to run once daily, for example, during off-peak hours in the early morning, using only the most recent (e.g., the past 24-72 hours) real-time feedback logs. _exec data.

[0305] The specific actions of the daily-level fine-tuning include: (a) Fine-tuning: freezing the vector field f _θ (a) Fine-tuning: Using the same batch of recent data as the new calibration subset Cal. The underlying parameters, such as the graph encoder, are fine-tuned only for the distribution regression head (i.e., the mapping layer between the Beta and LogNormal distributions). (b) Recalibration: Using the same batch of recent data as the new calibration subset Cal. _set Repeat the quantile calibration steps to calculate the new confidence coefficient q. _p and q _τ The above actions can quickly correct recent prediction biases of the model without destroying the complex physical-temporal representations already learned by the model, ensuring the lower bound p of the success rate. _lower and the upper bound of communication time τ _upper The confidence band always maintains its statistical validity.

[0306] Perform weekly (long-term) retraining to optimize the scheduling strategy. This step adapts to long-term environmental changes, such as the addition of new meters in the distribution area, permanent line alterations, or the strategy getting stuck in a local optimum. The update target is the hypergraph-constrained graph-structured reinforcement learning scheduling model. The update cycle can be set to once a week or month, for example, on weekends, using real-time feedback logs L accumulated over the past week or month. _exec Serves as a buffer for experience replay.

[0307] The specific actions of weekly retraining include: (a) offline retraining, fully implementing the Lagrangian constraint optimization PPO-Lagrangian training process, and retraining the policy network Φ. _B Value Network V _ φ and Lagrange multipliers are fully updated over multiple rounds. (b) (Optional) Hypergraph reconstruction, which can be triggered once before retraining (communication conflict hypergraph H). _conf This reflects newly discovered conflict relationships within the cycle, enabling the scheduling strategy itself to evolve, learn to adapt to structural and permanent changes in the environment, and find new and better scheduling patterns.

[0308] Preferably, to ensure the stability of online self-optimization and prevent performance degradation of newly trained models (whether daily or weekly), this embodiment introduces a versioned baseline rollback mechanism. The system retains the verified, stable model version B. _baseline The newly generated model (Θ) _A_new or Φ _B_new Before deploying to the production environment, first evaluate the model on a reserved validation dataset. Calculate key performance indicators (KPIs), such as success rate or completeness of data collection. If the new model's KPI is below B for K consecutive windows (e.g., 3),... _baseline The KPI minus a tolerance threshold δ, for example, KPI _new <KPI_baseline If -δ is selected, the update is rejected, the system automatically rolls back and continues using B. _baseline This will trigger an alarm so that manual intervention and analysis can be performed.

[0309] By employing a dual-timescale closed loop of daily fine-tuning prediction, weekly retraining strategy, and stability rollback, the method of this invention achieves the beneficial effect of continuous self-adaptation and self-optimization while maintaining high robustness and stability.

[0310] According to one aspect of this application, in the hypergraph-constrained graph structure reinforcement learning scheduling, the graph structure policy and value network, i.e., ensemble sampling with action masks, can also be:

[0311] Obtain the candidate set S _cand Action Feasible Domain Mask M _mask For illegal nodes, set the logit value to -∞, and apply softmax to obtain the node selection probability π. _node .

[0312] Get Mutual Exclusive Cluster _id Node selection probability π _node First, select each cluster based on its probability of being chosen, π. _cluster (From the maximum π within the cluster) _node (Sampling clusters with cluster weights), then sampling node i within the cluster; forming candidate actions, and temporarily reducing the weight of nodes in the same cluster to obtain the cluster-level sampling trajectory. _cluster Cluster weights are used in a two-stage sampling strategy that prioritizes clusters over nodes, comprehensively considering factors such as cluster size and node quality within a cluster to determine the basic selection tendency for each mutually exclusive cluster. The technical process of cluster weights can be implemented using existing technologies.

[0313] Acquiring candidate actions and time cost τ _upper_i Remaining time T _remain Mutual exclusion matrix C _mutex If τ _upper_i ≤T _remain If there is no conflict, then add i to the scheduling batch A. _batch And update the remaining time T _remain With action feasible region mask M _mask Otherwise, refuse and re-extract.

[0314] Obtain cluster-level sampling trace _cluster Calculate the cluster-level entropy H(π) _over_clusters ), and add the structural entropy regularization term L. _struct =-τ•H(π _over_clusters ); H(π _over_clusters ) represents the selection entropy of the policy network at the cluster level; τ is the weight coefficient. For homogeneous clusters, subnetwork parameters are shared, improving generalization and reducing redundant exploration. L_struct Used for subsequent total loss summation.

[0315] To address the disconnect between the predictive model and physical reality, this scheme constructs an electrical heterogeneity diagram containing detailed RLGC physical parameters and introduces a physical consistency loss derived from transmission line theory (such as the attenuation coefficient α(f) and reflection coefficient Γ(f)). This physical constraint, combined with topology intervention data augmentation—that is, using the physical model to generate counterfactual samples—forces the predictive model to obey physical laws, enabling it to understand and generalize low-frequency topology changes such as switching.

[0316] To address the issues of insufficient uncertainty quantification in forecasting and its inability to support robust planning, this scheme outputs a parameterized probability distribution, such as a Beta distribution, rather than a single point estimate. This scheme employs a coverage-preserving conformal calibration method, generating a confidence band with strict statistical coverage guarantees through residual quantile calibration, particularly regarding the upper bound of communication time τ. _upper And the lower bound of success rate p _lower This provides a solid data foundation for the scheduler to make worst-case safety decisions.

[0317] To address the issues of rigid conflict modeling and the inability of scheduling to strictly enforce hard constraints, this scheme constructs a communication conflict hypergraph. Its hyperedges can describe multi-dimensional (three or more meters) concurrent crosstalk caused by dynamic channel changes, overcoming the limitations of static pairwise modeling. Based on this, the scheduling model employs Lagrangian constraint optimization of the PPO-Lagrangian model, combined with an action-feasible region mask and an upper bound τ on communication time. _upper The real-time knapsack check transforms resource constraints such as total time budget from traditional soft penalties into hard constraints that must be strictly adhered to, ensuring the feasibility and robustness of the scheduling plan during execution.

[0318] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the protection scope of the present invention.

Claims

1. A meter reading scheduling method based on artificial intelligence, characterized in that, include: Acquire power grid operation data, construct an electrical heterogeneity diagram, and extract continuous-time driving factors from the operation environment sequence and historical communication sequence in the power grid operation data to construct a side-time variable drive. Based on electrical heterogeneous graphs and edge time-varying drive, a topology-physical consistent continuous-time distribution prediction model is used to generate a communication performance prediction distribution. Based on the electrical heterogeneity graph and the predicted distribution of communication performance, a communication conflict hypergraph is constructed; Based on the communication conflict hypergraph, the predicted distribution of communication performance, and system configuration, a graph structure reinforcement learning scheduling model with hypergraph constraints is used to generate a scheduling plan. Among them, based on the electrical heterogeneity graph and edge time-varying drive, a topology-physical consistent continuous-time distribution prediction model is adopted to generate a communication performance prediction distribution, including: By applying neural control differential equations, edge time-varying driving is evolved to generate temporal representations; Based on the physical attributes and time sequence representation of electrical heterogeneity graphs, physical consistency constraints of transmission lines are constructed, and physical consistency loss is derived. By using a distributed regression head, time-series representations are mapped, and the distribution of communication performance predictions is defined in a parameterized manner. Furthermore, during model training, physical consistency loss is used to regularize the model's representation space. Based on communication conflict hypergraphs, communication performance prediction distributions, and system configuration, a graph-structured reinforcement learning scheduling model with hypergraph constraints is used to generate scheduling plans, including: Based on the communication conflict hypergraph, the communication performance prediction distribution and system configuration, an action feasible domain mask and candidate set are generated. Using the structure of the communication conflict hypergraph as input, construct a graph structure policy network; Graph-structured policy networks sample candidate sets and generate scheduling plans under the constraint of action-feasible domain masks.

2. The method according to claim 1, characterized in that, Based on the physical attributes and time-series representations of electrical heterogeneity graphs, physical consistency constraints for transmission lines are constructed, and physical consistency losses are derived, including: Calculate the frequency-dependent attenuation coefficient based on the per-unit-length parameter of the electrical heterogeneity diagram; Based on the device parameters of the electrical heterogeneity diagram, determine the device transfer function; Based on the load impedance and characteristic impedance of the electrical heterogeneity diagram, determine the termination reflection coefficient; By combining the frequency-dependent attenuation coefficient, device transfer function, and termination reflection coefficient, the physical desired signal-to-noise ratio is derived. A physical consistency loss is constructed based on the model prediction signal-to-noise ratio obtained by mapping the physical expected signal-to-noise ratio to the temporal representation.

3. The method according to claim 1, characterized in that, Applying neural control differential equations, we evolve edge time-varying drivers to generate temporal representations, including: An adaptive step-size solver is used to numerically solve the time-varying edge drive and obtain intermediate solution results. At the switch switching event included in the edge time-varying drive, forced point marking and state update are performed on the intermediate solution results; Along the electrical topology path of the electrical heterogeneous graph, the updated intermediate solution results are aggregated to form a time-series representation.

4. The method according to claim 1, characterized in that, Graph-structured policy networks sample the candidate set under the constraint of action-feasibility region masks, including: Based on the communication conflict hypergraph, mutually exclusive clusters are derived and the target cluster is sampled, which is constrained by the action feasible domain mask. Within the target cluster, target nodes are sampled, constrained by the action-feasible domain mask. After sampling the target node, based on the upper bound of communication time in the communication performance prediction distribution, an immediate time knapsack check is performed to ensure that the cumulative communication time does not exceed the time budget in the system configuration, and a scheduling plan is generated.

5. The method according to claim 1, characterized in that, Training a graph-structured policy network includes: Identify homogeneous mutually exclusive clusters in the communication conflict hypergraph and enable these clusters to share policy sub-network parameters. Compute the cluster-level policy entropy of the graph structured policy network on isomorphic mutually exclusive clusters, and obtain the structure entropy regularization term; When training a graph structure policy network, a structural entropy regularization term is added to the training loss function.

6. The method according to claim 1, characterized in that: The communication performance prediction distribution includes a confidence band generated by a conformal calibration method that preserves coverage; A graph structure reinforcement learning scheduling model with hypergraph constraints is used to perform safety-side filtering on electricity meters using confidence bands to generate a candidate set. Furthermore, the scheduling model also utilizes the upper bound of communication time in the confidence band to calculate the constraint violation degree for Lagrange constraint optimization.

7. The method according to claim 6, characterized in that, Lagrangian-constrained optimization includes: Obtain the real-time feedback logs after executing the scheduling plan, and calculate the actual batch rewards accordingly. Construct PPO loss based on actual batch returns; The PPO loss is combined with the constraint violation weighted by the Lagrange multipliers to form the total loss; A graph structure reinforcement learning scheduling model that updates hypergraph constraints using total loss; The Lagrange multipliers are adaptively updated based on the degree of constraint violation.

8. The method according to claim 2, characterized in that, Training of the topology-physically consistent continuous-time distribution prediction model also includes topology intervention data augmentation: Simulate topology changes in electrical heterogeneity diagrams and generate a list of topology interventions; Based on the topological intervention list, the process of deriving the physical expected signal-to-noise ratio is reused to re-estimate the counterfactual communication metric. Generate counterfactual samples based on counterfactual communication metrics; We use counterfactual samples and real samples to train the topologically and physically consistent continuous-time distribution prediction model alternately.

Citation Information

Patent Citations

  • Power distribution network intelligent scheduling method and system based on data enhancement and topology awareness dynamic partitioning

    CN120582130A

  • Distributed source-load collaborative optimization method based on high-order topology and multi-scale attention

    CN121032068A