A reinforcement learning-based distributed cooperative scheduling method for energy nodes

CN122533152APending Publication Date: 2026-08-07YANGZHOU POLYTECHNIC INST +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
YANGZHOU POLYTECHNIC INST
Filing Date
2026-05-14
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

一方面,异构采集节点的物理通信链路客观存在衰减与阻滞,传统架构在强行执行全量特征对齐时,极易因输入端严重的时间戳错位而引发业务处理通道的逻辑停滞;另一方面,现有技术(如公开号为CN121076815A的发明申请,其构建了层次化深度强化学习控制框架以优化虚拟电厂的储能与机组出力策略)虽在多时间尺度动态模型与经济调度维度取得了技术突破,但该类架构将所有未经过滤的业务请求统一推入高复杂度的强化学习计算组件

Benefits of technology

[0015]Based on steps S1 and S2, the main-end feature latency and branch flow load rate are extracted, and the two are weighted and biased to overcome the limitations of traditional architectures that forcibly perform full feature clock alignment. This method jointly represents the timing decay of physical links and the resource occupation of core computing power channels, dynamically updates the flow degradation trigger threshold, thereby shrinking the global scheduling flow window and avoiding logical stagnation and state synchronization failure caused by data out-of-order issues from the source.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122533152A_ABST
    Figure CN122533152A_ABST
Patent Text Reader

Abstract

The application belongs to the field of artificial intelligence and distribution network control technology, and discloses a kind of energy node distributed collaborative scheduling method based on reinforcement learning.The method contains: checking node identity and obtaining voltage frequency and power sampling characteristics, extracting its main end feature delay quantity to arrive service interface;Extract reinforcement learning unit and send buffer depth to calculate evaluation branch load rate;Based on the weighted bias dynamic update of delay quantity and load rate, the transfer degradation trigger critical value is updated;According to this, the binary out-of-limit logic condition is judged to determine whether the sampling characteristics are normally routed to the reinforcement learning unit output optimal switching strategy, or forced to route to the secondary safety isolation component output static degradation quota, and then generate inverter adjustment instructions.The application effectively resolves the communication blockage and computing power overflow conflict under multiple concurrent requests through space-time decoupling and double-track routing mechanism, and improves the robustness of smart grid collaborative scheduling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and power distribution network control technology, and in particular to a distributed collaborative scheduling method for energy nodes based on reinforcement learning. Background Technology

[0002] With the continuous expansion of distributed energy access and the in-depth evolution of underlying computing models such as deep learning and reinforcement learning on the business side, the control architecture is shifting from the traditional centralized model to a distributed collaborative scheduling paradigm. Under this trend, high-dimensional data feature extraction and optimization mechanisms, represented by neural networks, are widely deployed in the intelligent node control of AC distribution network applications to address the challenges of state-space mapping and policy optimization in dynamic environments. Under the constraints of conventional server architectures, existing collaborative scheduling systems face severe spatiotemporal matching challenges and architectural limitations at the processing logic layer when dealing with massive, high-frequency requests. On the one hand, the physical communication links of heterogeneous acquisition nodes are subject to attenuation and obstruction. When traditional architectures forcibly perform full feature alignment, they are prone to causing logical stagnation in the business processing channel due to severe timestamp misalignment at the input end. On the other hand, while existing technologies (such as the invention application with publication number CN121076815A, which constructs a hierarchical deep reinforcement learning control framework to optimize the energy storage and unit output strategies of virtual power plants) have achieved technological breakthroughs in multi-timescale dynamic models and economic dispatch dimensions, this type of architecture uniformly pushes all unfiltered business requests into highly complex reinforcement learning computing components. This mechanism, which relies solely on static physical computing resources to fully bear the load, will lead to severe concurrent data accumulation and over-limit states in the core business flow buffer channel, ultimately inducing over-limit instruction issuance delays and the failure of the global business state synchronization mechanism.

[0003] This invention introduces an active sensing and degradation routing control strategy to break the signaling queuing bottleneck under high-frequency concurrency conditions without increasing hardware overhead, thereby achieving a high degree of coordination and spatiotemporal decoupling between the control cycle and the physical execution layer. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a distributed collaborative scheduling method for energy nodes based on reinforcement learning. A dynamic cross-mapping hub is constructed between the input end of the business application data stream and the evaluation and transfer layer. By extracting the temporal discrete deviation of the front-end input and the load loss state of the back-end reinforcement learning evaluation branch, an adaptive traffic splitting gateway controlled by multi-dimensional spatiotemporal characteristics is built. This solves the problems mentioned in the background technology.

[0005] The objective of this invention is achieved as follows: a distributed cooperative scheduling method for energy nodes based on reinforcement learning, executed by a cooperative scheduling center, includes the following steps:

[0006] Step S1: Obtain the voltage frequency offset and real-time active and reactive power sample values ​​of the parallel nodes that characterize the operating load of the parallel nodes in the AC distribution network, and extract the timestamp discrete difference of the voltage frequency offset and real-time active and reactive power sample values ​​reaching the service interface of the collaborative dispatch center, and label the timestamp discrete difference as the main end characteristic delay.

[0007] Step S2: Extract the business concurrency buffer depth of the pre-configured reinforcement learning evaluation unit and multi-dimensional rule evaluation component, and calculate the evaluation branch flow load rate based on the business concurrency buffer depth;

[0008] Step S3: Calculate the weighted bias value of the main end characteristic delay and the evaluation branch flow load rate, and update the weighted bias value to the flow degradation trigger threshold value;

[0009] Step S4: Determine whether the currently acquired evaluation branch flow load rate meets the binary over-limit logic condition set based on the flow degradation trigger threshold.

[0010] In response to the fulfillment of the binary over-limit logic condition, route redirection is triggered, and the voltage frequency offset and real-time active and reactive power sampling values ​​of the parallel node are forcibly routed to the secondary basic safety isolation component to perform discrete over-limit verification and matching based on the preset physical boundary threshold, and output the logic matching result carrying the static degradation quota parameter.

[0011] In response to the failure to meet the binary over-limit logic condition, the voltage frequency offset of the parallel node and the real-time active and reactive power sampled values ​​are routinely routed to the pre-trained reinforcement learning evaluation unit.

[0012] The reinforcement learning evaluation unit is configured to take the voltage frequency offset of the parallel node and the real-time active and reactive power sampling values ​​as inputs, and output an evaluation decision matrix that represents the optimal power switching strategy of each parallel node.

[0013] Step S5: Based on the evaluation decision matrix or the logical matching result of the secondary basic security isolation component, generate the corresponding inverter active and reactive power output adjustment command.

[0014] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0015] Based on steps S1 and S2, the main-end feature latency and branch flow load rate are extracted, and the two are weighted and biased to overcome the limitations of traditional architectures that forcibly perform full feature clock alignment. This method jointly represents the timing decay of physical links and the resource occupation of core computing power channels, dynamically updates the flow degradation trigger threshold, thereby shrinking the global scheduling flow window and avoiding logical stagnation and state synchronization failure caused by data out-of-order issues from the source.

[0016] To address the challenge of high-frequency concurrency easily causing overflow in core business buffer channels, this invention constructs a binary overload logic condition judgment model in step S4. When the system detects an overload, the mechanism proactively triggers route redirection, forcibly diverting high-dimensional sampling features to a lightweight secondary basic security isolation component for extreme value matching, effectively eliminating unnecessary computational resource pressure on the reinforcement learning evaluation unit. This mechanism avoids congestion and blockage in the instruction queue, improving the core computing hub's anti-disturbance capability and operational robustness under extreme conditions.

[0017] In step S5 of this invention, a differentiated inverter output adjustment command is generated by combining the reinforcement learning evaluation decision matrix and the static degradation quota parameters. Under normal operating conditions, the neural network provides a high-precision optimal switching strategy to maintain the grid frequency safety tolerance; under stalled operating conditions, it directly issues the bottom-line safety constant and triggers the degradation reset of the source-end data acquisition frequency. This collaborative mechanism combines the high-dimensional analytical sensitivity of deep learning with the absolute security of physical boundaries, achieving optimal comprehensive performance of the business flow closed loop. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0019] Figure 1 This is a technical architecture diagram of a distributed collaborative scheduling method for energy nodes based on reinforcement learning.

[0020] Figure 2 This is a schematic diagram of the technical route for steps S1 to S5 of the present invention. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] like Figures 1 to 2 As shown, a distributed cooperative scheduling method for energy nodes based on reinforcement learning is implemented by the cooperative scheduling center and includes the following steps:

[0023] Step S1: Parse the node permission credentials of the parallel node requesting access to extract the structured identity identifier, and perform discrete feature comparison and verification with the pre-set underlying communication whitelist. The structured identity identifier is essentially a composite service data carrier of the logical address allocation code and encrypted timestamp of the underlying interactive node at the network access layer, which is used to characterize the communication legitimacy status of the node.

[0024] After the verification is passed, the voltage frequency offset of the parallel node and the real-time active and reactive power sampling values ​​that characterize the operating load of the parallel node in the AC distribution network are intercepted and obtained. The timestamp discrete difference of the voltage frequency offset of the parallel node and the real-time active and reactive power sampling values ​​reaching the business interface of the collaborative dispatch center is extracted, and the timestamp discrete difference is marked as the main end characteristic delay.

[0025] Step S2: Extract the business concurrency buffer depth of the pre-configured reinforcement learning evaluation unit and multi-dimensional rule evaluation component, and calculate the evaluation branch flow load rate based on the business concurrency buffer depth;

[0026] Step S3: Calculate the weighted bias value of the main end characteristic delay and the evaluation branch flow load rate, and update the weighted bias value to the flow degradation trigger threshold value;

[0027] Step S4: Determine whether the currently acquired evaluation branch flow load rate meets the binary over-limit logic condition set based on the flow degradation trigger threshold.

[0028] In response to the fulfillment of the binary over-limit logic condition, route redirection is triggered, and the voltage frequency offset and real-time active and reactive power sampling values ​​of the parallel node are forcibly routed to the secondary basic safety isolation component to perform discrete over-limit verification and matching based on the preset physical boundary threshold, and output the logic matching result carrying the static degradation quota parameter.

[0029] In response to the failure to meet the binary over-limit logic condition, the voltage frequency offset of the parallel node and the real-time active and reactive power sampled values ​​are routinely routed to the pre-trained reinforcement learning evaluation unit.

[0030] The reinforcement learning evaluation unit is configured to take the voltage frequency offset of the parallel node and the real-time active and reactive power sampling values ​​as inputs, and output an evaluation decision matrix that represents the optimal power switching strategy of each parallel node.

[0031] Step S5: Based on the evaluation decision matrix or the logical matching result of the secondary basic security isolation component, generate the corresponding inverter active and reactive power output adjustment command.

[0032] Following step S5, the following is also included:

[0033] Based on the underlying communication network, the active and reactive power output adjustment command of the inverter is sent to the corresponding parallel node of the AC power distribution network. The active and reactive power output adjustment command of the inverter is encapsulated with operating status adjustment parameters.

[0034] The operating status adjustment parameters are configured for parsing by the parallel nodes of the AC distribution network at the receiving end, so as to guide the solid-state circuit breakers and inverter control units within the parallel nodes of the AC distribution network to asynchronously perform state switching and power parameter adjustment actions, so that the real-time power load of the parallel nodes of the AC distribution network converges to the preset grid frequency safety tolerance baseline.

[0035] When parsing the operating status adjustment parameters of the parallel node, it also includes extracting the physical warning information of the over-boundary physical state that has reached the safe operating limit state of the local system, and performing interception rule verification based on the safe operating boundary conditions pre-built into the local logical interaction layer of the parallel node.

[0036] In response to the interception rule verification prompt that there is an over-limit conflict in the current control state, the processing pipeline of the received inverter active and reactive power output adjustment command is forcibly blocked, and a local fallback state control parameter independent of the remote collaborative scheduling is generated to replace the operation state adjustment parameter to execute the underlying logic safety closed loop.

[0037] Extract the discrete time stamp difference between the voltage frequency offset of the parallel nodes and the real-time active and reactive power sample values ​​reaching the service interface of the collaborative scheduling center, and label the discrete time stamp difference as the master-end characteristic delay, including:

[0038] Open the data aggregation time window; extract the first arrival timestamp of the voltage frequency offset and real-time active and reactive power sampling values ​​of each parallel node of the AC distribution network that arrive at the business interface of the collaborative dispatch center.

[0039] Extract the largest and smallest first arrival timestamps representing extreme value arrival records within the same data aggregation time window, and perform discrete absolute difference calculation based on the largest and smallest first arrival timestamps. Extract the output discrete absolute difference as the master-end feature delay.

[0040] Extract the business concurrency buffer depth of the pre-configured reinforcement learning evaluation unit and multi-dimensional rule evaluation component, and calculate the evaluation branch flow load rate based on the business concurrency buffer depth, including:

[0041] The system periodically extracts the backlog of pending business and the concurrent execution channel size of the reinforcement learning evaluation unit;

[0042] The backlog of pending business is weighted and summed with the concurrent execution channel size to generate a real-time backlog load value;

[0043] The evaluation branch flow load rate is generated by dividing the real-time backlog load value by the preset node baseline throughput threshold.

[0044] Calculate the weighted bias value between the main-end characteristic latency and the evaluation branch flow load rate, and update the weighted bias value to the flow degradation trigger threshold, including:

[0045] When the extracted master-end feature delay exceeds the system's preset asynchronous timing tolerance threshold, the absolute difference between the master-end feature delay and the asynchronous timing tolerance threshold is calculated, and the absolute difference is divided by the preset limit timing tolerance span to eliminate the dimension. The pure numerical ratio of the output is then extracted as the excess gain coefficient.

[0046] A pre-configured smoothing penalty constant is introduced, and the excess gain coefficient, the evaluation branch flow load rate, and the smoothing penalty constant are multiplied together to generate a negative bias attenuation term;

[0047] Obtain the preset basic branch switching baseline, and subtract the negative bias attenuation term from the basic branch switching baseline to obtain the updated flow degradation trigger threshold.

[0048] Based on the evaluation decision matrix or the logical matching result of the secondary basic security isolation component, a corresponding inverter active and reactive power output adjustment command is generated, including:

[0049] When data is routed to the reinforcement learning evaluation unit, a corresponding dynamic optimal inverter active and reactive power output adjustment command is generated based on the evaluation decision matrix.

[0050] When data is routed to the secondary basic security isolation component, the logical matching result is parsed to extract discrete over-limit identifiers that characterize physical boundary over-limits, and static degradation quota parameters that are pre-hard bound to the discrete over-limit identifiers are extracted. Based on the static degradation quota parameters, corresponding static degradation inverter active and reactive power output adjustment commands are generated.

[0051] In response to the inverter's active and reactive power output adjustment command including the static degradation quota parameter, the adjustment command is configured as follows: after determining that the solid-state circuit breaker of the target AC distribution network parallel node has completed the state transition and outputs a physical latch response, the degradation reset of the reporting cycle of the source-end data acquisition interface of the parallel node is asynchronously triggered based on the timing decoupling logic.

[0052] The steps for obtaining the evaluation decision matrix specifically include:

[0053] The voltage frequency offset and real-time active and reactive power sampling values ​​of each parallel node are obtained. The pre-set limit boundary extreme values ​​of each parameter are extracted. Based on the limit boundary extreme values, the voltage frequency offset and real-time active and reactive power sampling values ​​of the parallel node are normalized and mapped in a dimensionless manner. The normalized data structure is then subjected to cross-dimensional feature splicing to construct a node state observation space feature matrix that characterizes the objective operating state of the global power grid.

[0054] To enhance the configuration of the deep neural network feature extraction pipeline pre-built in the learning evaluation unit, optimization convergence constraint mapping rules based on the global business reward function are applied.

[0055] The node state observation space feature matrix is ​​input into the deep neural network feature extraction pipeline. Feature extraction and dimensionality reduction mapping are performed based on the optimization convergence constraint mapping rule. The output is the evaluation decision matrix composed of discrete optimal switching state identifiers and continuous output power quota coefficients.

[0056] Feature dimensionality reduction transformation is performed based on the pre-constructed node state observation space feature matrix and optimization convergence constraint mapping rules, including:

[0057] Extract the master-end characteristic delay of each parallel node of the AC distribution network, and perform time-series inverse feature mapping to generate the time-degradation penalty index of each parallel node of the AC distribution network.

[0058] The state-action value calculation logic inside the optimization convergence constraint mapping rule is triggered to perform backpropagation of value error, so as to extract the initial expected cumulative reward gradient representing the optimization convergence direction. The time decay penalty index is used to perform cross-dimensional multiplication suppression operation on the initial expected cumulative reward gradient to obtain the updated state-action optimization gradient.

[0059] Based on the state action optimization gradient, the action space is traversed, and discrete power adjustment actions with negative gradient amplification are eliminated to complete the feature dimensionality reduction transformation.

[0060] Obtain the initial expected cumulative reward gradient output by the optimization convergence constraint mapping rule, and perform a cross-dimensional multiplication suppression operation on the initial expected cumulative reward gradient using the time decay penalty exponent to obtain the updated state-action optimization gradient, including:

[0061] Determine whether the extracted time-deterioration penalty index falls below the preset feature confidence limit value;

[0062] In response to the fact that the time-decrease penalty index does not fall below the feature confidence limit, the time-decrease penalty index and the initial expected cumulative reward gradient are multiplied by a matrix in the corresponding dimension.

[0063] In response to the time decay penalty index falling below the feature confidence limit, the absolute blocking control logic is activated to generate a state over-limit isolation flag. Based on the state over-limit isolation flag, the state action optimization gradient corresponding to the node is forcibly reset to a zero vector, thereby blocking the erroneous weight allocation of high-latency features in the evaluation decision matrix.

[0064] Obtain the voltage frequency offset and real-time active and reactive power sample values ​​of the parallel nodes, which characterize the operating load of the parallel nodes in the AC distribution network. Specifically, this includes:

[0065] Obtain discrete voltage sampling sequences and discrete current sampling sequences from the underlying interaction interface of the parallel nodes of the AC power distribution network;

[0066] Extract the continuous zero-crossing time-scale features from the discrete voltage sampling sequence and perform reciprocal mapping to derive the real-time operating frequency parameter. Then, perform a differential operation between the real-time operating frequency parameter and the system rated frequency reference, and extract the absolute value of the difference as the voltage frequency offset of the parallel node.

[0067] A fixed-time sliding window is established, and time-domain interpolation alignment mapping is performed on the introduced discrete voltage sampling sequence and the discrete current sampling sequence based on the local reference synchronization time scale. Then, within the time sliding window, point-by-point discrete multiplication and integral mean extraction are performed on the aligned discrete voltage sampling sequence and the discrete current sampling sequence to generate the real-time active power sampling value.

[0068] The discrete voltage sampling sequence is reconstructed by an orthogonal phase shift to generate an orthogonal voltage sequence. The orthogonal voltage sequence is then multiplied point-by-point by the discrete current sampling sequence and the mean of the integral is extracted to generate the real-time reactive power sampling value.

[0069] The following are specific implementation instructions for the above content:

[0070] This embodiment has an asynchronous timing tolerance threshold. With limit timing tolerance span During the system initialization phase, the gateway baseline configuration library at the collaborative scheduling center is pre-mounted and persistently stored. This includes the asynchronous timing tolerance threshold. This indicates that the business data is still within the safe physical tolerance edge of the effective calculation window; the extreme time-series tolerance span. This defines the absolute collapse and stagnation dead zone boundary that the system can tolerate after crossing the safety red line.

[0071] In response to the access of real-time dynamic business data streams, the system reads the configuration library to obtain these two threshold scalars as reference baselines for pre-logic judgment and out-of-bounds normalization calculations, ensuring that the underlying architecture gateway always executes a set of global physical timing calibration benchmarks with constant scale, regardless of the type of inverter node accessed.

[0072] Limiting time tolerance span In this embodiment, the preferred range is [50ms, 200ms], and the preferred value is 100ms.

[0073] Master-end characteristic delay (denoted as) By obtaining the extreme timestamp values ​​of the first and last data packets within the convergence time window, discrete difference calculation is performed for extraction. In this embodiment, the preferred range of the master-end characteristic delay is [0ms, preset maximum network tolerance delay window]. In this embodiment, the relative deviation rate is normalized and mapped to [0, 1.0], and the preferred value is a reference synchronization state value that approximates 0.

[0074] Assess branch flow load rate (denoted as ), a dimensionless state quantification coefficient characterizing the congestion level of pending business data. It is the saturation ratio relative to the baseline throughput capacity, calculated by weighting the concurrent backlog size and concurrent channel size based on reinforcement learning units. In this embodiment, the preferred range is [0%, 100%], or the normalized floating-point interval [0, 1], with a preferred value maintained within the interval [40%, 75%].

[0075] The flow degradation trigger threshold (denoted as) This is a dynamic gateway threshold used to control the degradation switching of dual-track routing branches in the business flow panorama. It is dynamically adjusted based on the cross-domain multiplicative penalty of physical latency exceedance deviation and system concurrent load.

[0076] Time-degradation penalty index (denoted as This represents a dynamic constraint scalar that characterizes the timeliness and confidence of high-frequency concurrent business data. It is obtained by performing reciprocal feature mapping based on the master-end feature delay of each parallel node, and is forced as a prior weight for evaluating the state confidence of the RL network.

[0077] Status exceeding the limit isolation flag (denoted as) This represents a binary structured Boolean logic state bit that triggers the circuit breaker of the underlying logic of the high-dimensional action matrix. It is generated when the aforementioned attenuation penalty exponent falls below a preset limit and is used to block error weight allocation. Its state transition value is a structured Boolean logic control word of [0 (normal release security code), 1 (over-limit circuit breaker / physical blocking code)].

[0078] State-action optimization gradient (denoted as) ), representing the updated policy gradient that guides the convergence direction of optimization within the reinforcement learning evaluation unit. It is extracted by performing a matrix multiplication suppression operation on the initial expected cumulative reward gradient and the time decay penalty exponent.

[0079] This embodiment opens and maintains a data aggregation time window, extracts extreme arrival time records of business data arriving at the collaborative scheduling center's business interface, performs absolute difference calculation, and generates the master-end characteristic latency. Its calculation logic is as follows: ;in and These are the maximum and minimum first arrival timestamps of records within the same data aggregation time window.

[0080] Master-end characteristic delay The preferred value is a baseline synchronization state value close to 0, whose actual value characterizes the asynchronous jitter span caused by physical distance or network congestion in the communication link. The backlog of pending services and the scale of concurrent execution channels are periodically acquired. A weighted mapping factor to eliminate differences in data units is introduced to integrate heterogeneous parameters, generating a real-time backlog load value. This value is then divided by a preset node baseline throughput threshold to map to an evaluation branch throughput rate. : ;

[0081] in The backlog of pending business is a sequence of positive integers [0, maximum stack depth constant of the buffer queue]. The concurrent execution channel size is a discrete positive integer in the range [1, N1] (where N1 is the maximum number of parallel processing branches allocated). and For the corresponding dimensionless weighting factor, Set a baseline throughput threshold for the node to ensure the output evaluation branch flow load rate. It is within [0,1]. and In this embodiment, it is responsible for weight allocation and unit conversion, and its value range is from 0 to 1.

[0082] Extract the difference exceeding the system timing tolerance, divide it by the limiting timing tolerance span to eliminate physical dimensions, and generate a dimensionless excess gain coefficient; introduce a smoothing penalty constant, perform a multiplication operation, and generate an updated flow degradation trigger threshold. :

[0083] in, This is the limit of the time series tolerance span used to eliminate the time dimension. This is a preset smoothing penalty constant. Representing the excess gain coefficient, in this embodiment, a dimensionless pure numerical ratio with a preferred range of (0, 1.0) is used to ensure that its absolute value does not exceed 1.0, thus eliminating the physical dimension of time.

[0084] This represents the negative bias attenuation term, and in this embodiment, the preferred range is [0, 0.4]. The preferred value is achieved by introducing a smoothing penalty constant. The leveled product, where the smoothing penalty constant is preferably selected within a closed interval ranging from 0.3 to 0.5: Ensure that the dimensions and scale of the attenuation term are absolutely aligned with the baseline.

[0085] Generate the time decay penalty index based on the mapping rule with safety bias. To prevent zero overflow: ;

[0086] in To prevent the denominator from being zero and to adjust the minimum positive real number compensation coefficient of the attenuation slope; the preferred range in this embodiment is

[10] . -4 10 -2 Between ], the preferred value is 10. -3 .

[0087] This embodiment has a delay due to the main-end feature. In the ideal zero-latency state, the value is 0, and a minimum positive real number compensation coefficient is forcibly introduced. Not only does it eliminate the risk of computational crashes and downtime due to division by zero, but it also makes the time decay penalty index represented by the confidence scalar of (0,1] more effective. It can smoothly characterize the degree of marginal deterioration of business latency over time.

[0088] against Obtain the initial expected cumulative reward gradient of the backpropagation output of the deep network. In deep neural networks, the initial expected cumulative reward gradient is represented as a high-dimensional feature tensor and participates in subsequent cross-dimensional matrix suppression operations. An element-wise dot product operator is then applied to it using a time-decay penalty exponent. The multiplicative suppression operation can directly map the latency degradation state of the physical communication network as an attention mask onto the weight update chain of the deep neural network. This allows the weights of the erroneous states of high-risk edge nodes with high latency to be passively weakened mathematically in the decision evaluation space by the underlying calculus suppression action, without being removed from the control loop. This achieves self-healing correction of physical latency in the algorithm optimization space.

[0089] As the baseline for basic branch switching, in this embodiment, the preferred system concurrency saturation red line is in the range of [0.65, 0.95], and the preferred constant warning level is 0.85. For asynchronous timing tolerance threshold, the preferred range in this embodiment is [10ms, 100ms], and the preferred value is 10ms, which corresponds to the extreme value defense line of a specific control cycle of the AC power grid.

[0090] This is the feature confidence limit value. In this embodiment, the preferred range is [0.1, 0.5], and the preferred value is 0.2;

[0091] In high-frequency request scenarios of smart grids, traditional collaborative scheduling mechanisms rely solely on statically set load allocation warning lines. This one-dimensional blocking flow and open-loop mechanism has serious local limitations: when heterogeneous nodes experience severe timing misalignment at the high-dimensional input due to differences in physical communication link distance and attenuation, the static architecture forcibly performs full feature alignment, causing the underlying buffer queue to overflow and the digital model to become distorted.

[0092] This embodiment forcibly removes physical communication attenuation in the time domain, mapping it to a purely digital parameter, "master-end characteristic latency." In the spatial / resource domain, it extracts the actual backlog data structure state within the logical execution channel and reorganizes it into "evaluation branch flow load rate." The degree of deterioration of timing misalignment is used as the excess gain coefficient represented by the penalty amplification lever, dynamically reducing the flow degradation gateway boundary. In the optimization and convergence business processing logic, the latency feature is inversely mapped to the "time-degradation penalty index." When this mechanism encounters business boundary conditions (e.g., sudden network outages or high-frequency pulse congestion causing data latency to penetrate the bottom line limit), it executes a logic circuit breaker: if the time-degradation penalty index falls below the feature confidence limit, the system refuses to execute regular calculations to prevent dirty data from polluting the action value matrix, and instead directly triggers the over-limit isolation mechanism, forcibly resetting the state action optimization gradient to a zero vector, thereby depriving the high-latency node of its global scheduling decision weight.

[0093] Based on the above logic and parameter definitions, the specific implementation steps of this invention are as follows:

[0094] Step S1: Perform underlying communication credential verification and multi-source service feature cleaning for access: The system gateway intercepts access request frames from the underlying layer, extracts the logical address sequence and encrypted authentication timestamp from the data frame header, and constructs a structured identity identifier. In the secure and isolated operating environment at the collaborative scheduling center, it traverses the pre-persistently stored underlying communication whitelist. In response to the determination that the structured identity identifier matches any legitimate entry in the underlying communication whitelist with an exact literal match, a status release signal is triggered; otherwise, a blocking status word is generated to intercept illegal or unauthorized underlying network collaborative control requests. This achieves hard filtering of the legitimacy of out-of-order high-frequency requests before parsing, thereby improving the system communication architecture's concurrency fault tolerance boundary in the face of abnormal peak traffic.

[0095] After the verification is passed, an interception-type data acquisition operation is performed to collect the voltage frequency offset of the parallel nodes and the real-time active and reactive power sampling values ​​that characterize the operating load of the parallel nodes in the AC distribution network.

[0096] Furthermore, discrete voltage sampling sequences and discrete current sampling sequences are obtained from the underlying sensing interfaces of parallel nodes in the AC power distribution network.

[0097] To extract the voltage frequency offset of parallel nodes, the following steps are performed:

[0098] Extract the time-stamp features of consecutive zero-crossing points in the discrete voltage sampling sequence; perform reciprocal mapping based on the time difference between adjacent zero-crossing time-stamps to generate a real-time operating frequency parameter characterizing the current dynamic response of the power grid at that node; obtain the preset system rated frequency reference, and perform differential operation between the real-time operating frequency parameter and the system rated frequency reference; extract the absolute value of the difference between the two and output it directly as the voltage frequency offset of the parallel node.

[0099] The above execution logic is as follows: ;in This represents the voltage frequency offset at parallel nodes. The real-time operating frequency parameters are derived from discrete sampling sequences. The static system's rated frequency serves as the reference. For the extraction of real-time active and reactive power samples, the following feature recombination actions are performed based on the same discrete time-series reference:

[0100] A time sliding window is established for synchronous feature observation. Within the time sliding window, point-by-point discrete multiplication and integral averaging operations are performed on the discrete voltage sampling sequence and the discrete current sampling sequence, and the result is extracted as the real-time active power sampling value. To address the asynchronous timing drift problem caused by the inconsistency of the underlying physical crystal oscillator frequencies of multi-source sensing nodes, underlying timing alignment preprocessing is forcibly enabled before power dimensionality reduction calculation. Specifically, the first data timestamp of the discrete voltage sampling sequence reaching the center is extracted as the local reference synchronization timescale within the sliding window. Using a first-order linear or polynomial time-domain interpolation algorithm, the timescale of the discrete current sampling sequence is forcibly resampled and mapped to the discrete time grid anchored by the local reference synchronization timescale, completing the cross-sequence time-domain interpolation alignment mapping. Within the time sliding window, point-by-point discrete multiplication and integral averaging operations are performed on the aligned discrete voltage and current sequences, and the result is extracted as the real-time active power sampling value.

[0101] Using a preset orthogonal transform logic-represented discrete Hilbert transform, a 90-degree phase shift reconstruction is performed on the discrete voltage sampling sequence to generate an orthogonal voltage sequence. The orthogonal voltage sequence is then multiplied point-by-point with the original discrete current sampling sequence, and the result is extracted as the real-time reactive power sampling value. The calculation logic is represented as follows:

[0102]

[0103]

[0104] in Represents the real-time active power sample value. M represents the real-time reactive power sample value; M represents the total number of discrete sampling points within the time sliding window. and Representing the first The instantaneous discrete voltage and instantaneous discrete current values ​​at each time step. This represents the orthogonal discrete voltage value after orthogonal phase shift reconstruction.

[0105] In this embodiment, the sensing interface is configured as a functional data acquisition center with standardized clock synchronization capabilities. This is for the voltage-frequency offset of parallel nodes. The generation of real-time active and reactive power sample values ​​(P, Q) originates from feature dimensionality reduction mapping. Specifically, the input data is discretized to form a one-dimensional numerical sequence, which is then truncated within a preset time sliding window. This process maps physical electrical parameters (instantaneous alternating states of voltage and current) into continuous numerical state scalars that can be directly analyzed by the system application layer, filtering out interference from high-frequency service noise.

[0106] In response to the continuous access of multi-source business data streams, a fixed data aggregation time window is opened; the initial arrival time records of the business data streams uploaded by each parallel node of the AC power distribution network to the business interface at the collaborative dispatch center are extracted. Specifically, this time record is a timestamp log automatically generated by the operating system kernel for concurrent data packets in the underlying buffer. This timestamp log is obtained and marked as the first arrival timestamp; the absolute difference between the largest and smallest first arrival timestamps within the same data aggregation time window is calculated.

[0107] Based on the absolute difference extraction results, it is calibrated as the master-end characteristic delay quantity used to characterize the asynchronous jitter span of the underlying communication link, thereby accurately converting the physical communication attenuation state of multi-source nodes into a dimensionless unified service timing benchmark parameter.

[0108] For step S2, perform concurrent state monitoring and load rate calculation of the business processing architecture: periodically scan and extract the business concurrency buffer depth at the bottom layer of the pre-configured reinforcement learning evaluation unit and multi-dimensional rule evaluation component; the business concurrency buffer depth is characterized by monitoring the backlog queuing index of the business data to be consumed.

[0109] Extract the current business flow interface status of the reinforcement learning evaluation unit, specifically obtain the backlog of pending business signals that have not entered the matching stage (representing the scale of pending business backlog), and the number of concurrent flow branches that are currently active and assigned to execute computation tasks (representing the scale of concurrent execution channels).

[0110] The backlog of pending business and the concurrent execution channel size are obtained. A standardized dimensionless mapping weight is introduced to perform heterogeneous weighted summation, generating a real-time backlog load value representing the combined tension of static computing power and dynamic stacking. A preset node baseline throughput threshold representing the maximum processing capacity of a single node is obtained. The calculated real-time backlog load value is divided by the node baseline throughput threshold. Based on the division result, the evaluation branch flow load rate is generated and output.

[0111] For step S3, perform cross-domain load penalty mapping and flow degradation trigger threshold update: obtain the determined master-end characteristic latency and evaluate the branch flow load rate. Extract the preset asynchronous timing tolerance threshold to characterize the maximum allowable delay time for normal physical network attenuation. Determine whether the master-end characteristic latency exceeds the system's preset asynchronous timing tolerance threshold; in response to the determination that the master-end characteristic latency exceeds the safe time domain boundary, calculate the absolute difference between the master-end characteristic latency and the asynchronous timing tolerance threshold.

[0112] Obtain a preset limiting time-series tolerance span for eliminating physical time dimensions; divide the calculated absolute difference by the limiting time-series tolerance span to obtain a normalized pure numerical ratio, and extract this pure numerical ratio as the excess gain coefficient. Obtain a pre-configured smoothing penalty constant, and perform a multiplication operation on the extracted excess gain coefficient, the evaluated branch flow load rate, and the smoothing penalty constant; based on the scale alignment and positive cross-fusion of heterogeneous parameters, generate a negative bias attenuation term that is limited by the absolute safety margin and dynamically amplifies with the degree of disorder.

[0113] Obtain the pre-fixed basic branch switching baseline that represents the warning line of business load distribution; subtract the negative bias attenuation term from the basic branch switching baseline to calculate the updated flow degradation trigger threshold.

[0114] In the system's global static runtime architecture, the basic branch switching baseline And subsequent confidence limits of features used to determine the degradation state of high-frequency data. All are defined as static baseline parameters. During the system startup initialization phase, they are statically mounted and locked in the global state configuration registry at the collaborative scheduling center. The basic branch switching baseline is one of them. It is a dimensionless warning constant characterizing the upper limit of the system's logical processing concurrency capacity. The cross-domain negative bias decay rule is only triggered when the dynamic flow load rate calculated by the main business approaches this level; while the feature confidence limit value It is the absolute confidence lower limit that characterizes the bottom line of data timeliness and failure, and serves as the hard trigger limit for deep decision nodes in the network.

[0115] This represents the updated flow degradation trigger threshold for the final output. In this embodiment, the preferred range is [0.25, 0.95], and the preferred value is the value at which the baseline is switched in the base branch. Based on this, subtract the normalized negative bias attenuation term. The obtained dynamic floating value ensures that the system gateway state transition logic is 100% self-consistent.

[0116] Excess gain coefficient in this embodiment and the threshold for triggering degradation during circulation These are the core control parameters constituting adaptive route redirection. The excess gain coefficient is defined as a unitless relative penalty multiplier, with a constant value within the closed interval (0,1). By obtaining the out-of-bounds difference in the master-end characteristic delay and forcibly dividing it by the system's preset maximum timing tolerance span, a clean dimensionless mechanism is achieved in the computational process.

[0117] Flow degradation trigger threshold The configuration logic is constrained by the gateway boundary, which shrinks spontaneously with the degree of concurrency and out-of-order processing. Under severe high-concurrency data access conditions, the acceptance level of the reinforcement learning unit is actively lowered, and excess load is drained to the bypass isolation component without loss, thereby preventing the risk of cascading crashes of global application services.

[0118] For step S4, execute architecture-level dual-track traffic splitting and model optimization / dimensionality reduction / static discrete matching mapping; continuously acquire the timing and concurrency characteristics of the current business data flow. The system gateway continuously executes objective binary condition judgments. In response to the judgment that the currently acquired evaluation branch flow load rate meets the binary over-limit logic condition set based on the flow degradation trigger threshold (if and only if the real-time calculated output...), the system gateway performs objective binary condition judgments. When established, the system status flag flips, triggering gateway-level bypass redirection.

[0119] Perform a cross-domain business logic branch switching action, and force the parallel node voltage frequency offset and real-time active and reactive power sampling values ​​to the backup secondary basic safety isolation component.

[0120] Within the execution pipeline of the secondary basic safety isolation component, a set of discrete preset physical boundary thresholds are directly extracted. These preset physical boundary thresholds include, but are not limited to, overvoltage protection thresholds and frequency limit offsets that characterize the objective physical limit red line of the device.

[0121] In this embodiment, the overvoltage protection threshold is the maximum voltage limit that the power grid or equipment can withstand. The frequency limit offset is the maximum physically permissible range of the power grid frequency deviating from the rated frequency (50Hz). See Table 0 below for examples;

[0122] Table 0: Example of mapping between preset physical boundary thresholds and static degradation quotas for secondary basic security isolation components.

[0123]

[0124] It should be further explained that in the above discrete limit check matching logic, the parameter definitions on which the preset physical boundary threshold and the static degradation quota mapping table depend are as follows: where The reference value of the system's rated operating voltage, which is pre-configured in the target AC power distribution network, represents the effective value of the instantaneous voltage, which is acquired and mapped in real time through the underlying sensing interface. S represents the rated limit apparent power output capacity of the underlying power modules such as inverters within the parallel nodes of the AC distribution network; S represents the operating apparent power calculated in real time by the system.

[0125] This method extracts the real-time active power sample value P and the real-time reactive power sample value Q generated by the pre-dimensionality reduction and conventional routing, and then uses the vector orthogonal summation rule ( We perform algebraic derivation to generate the apparent power S.

[0126] When the sampled data of the parallel nodes falls within any of the aforementioned preset physical boundary threshold ranges, the secondary basic safety isolation component does not need to perform high-dimensional feature dimensionality reduction or reward gradient calculation. Instead, it directly matches the corresponding discrete over-limit identifier and extracts the static degradation quota parameters (discrete machine codes such as 0x00, 0x01, and 0x02) that are rigidly bound to it. This quota parameter is directly encapsulated into the inverter's active and reactive power output adjustment command and issued. By configuring the direct mapping between the aforementioned absolute physical extreme values ​​and static constant quotas, this embodiment not only ensures that when concurrent congestion causes the RL evaluation unit to be bypassed, the solid-state circuit breaker and inverter can still achieve millisecond-level rigid self-healing and disaster prevention network disconnection based on safety rules.

[0127] Based on a preset physical boundary threshold, a one-dimensional extreme value comparison operation is performed on the voltage / power sample values ​​routed here to characterize discrete out-of-bounds verification matching.

[0128] Based on the extreme value comparison results, a structured signaling is output. This structured signaling represents the logical matching result, which includes an over-limit status flag and discrete polymorphic control constants (static degradation quota parameters such as power output forced halving coefficient or forced cut-off status code) that are pre-binded to it. The state transition value of the static degradation quota parameter is a preset set of safe operating baseline constants (e.g., [0x01: forced 20% power output quota code, 0x02: forced 50% power degradation code, 0x00: solid-state circuit breaker direct disconnect blocking code]).

[0129] Status Exceeds Limit Isolation Flag This is a binary structured Boolean logic state bit, mounted in the internal scheduling pipeline of the reinforcement learning evaluation unit. It is initialized by default to the normal open state (logic 0); if the internal checker responds to the condition that the exponent of the time decay penalty falls below the extreme value, this flag immediately undergoes a state flip (flows to logic 1), and the gradient update operation of the contaminated node is absolutely blocked through the interruption mechanism of the state machine.

[0130] In the secondary basic security isolation component, the data flow format of the static degradation quota parameters is represented by a set of discrete polymorphic control constants. Based on the pre-computed one-dimensional extreme value comparison results, matching items are extracted from the set of constants pre-persisted in the local security mapping library. These constants not only include source-end degradation state prefixes used to suppress front-end throughput concurrency, but also directly map to the structured instruction set of the end control interface (e.g., discrete state codes representing forced power output suppression, or latching control bits that directly execute physical cutoff).

[0131] Conversely, in response to the absence of a degradation-based traffic diversion condition, the business data flow is routinely routed to the pre-trained reinforcement learning evaluation unit. Specifically, the extreme values ​​of the boundary conditions of various underlying business parameters pre-embedded in the control node (the maximum allowed frequency of exceeding the system's limits) are extracted. and the inverter's rated maximum output power Based on the extreme boundary values, a dimensionless normalization mapping is performed on the voltage frequency offset of parallel nodes and the real-time active and reactive power sample values. The implementation method is as follows: ;in Generally refers to the currently acquired physical sample values. and These are the corresponding limit boundary extrema and bottom limit extrema. This transformation ensures that all data structures input to the deep neural network are compressed into the dimensionless pure numerical state interval of [0,1] or [-1,1], in order to eliminate the model convergence bias caused by the difference in absolute numerical magnitude.

[0132] Within the reinforcement learning evaluation unit, the obtained parallel node voltage frequency offset and real-time active and reactive power sampling values ​​are numerically normalized and then concatenated in a high-dimensional vector space to construct a continuous observation space feature matrix that characterizes the current objective operating state of the power grid.

[0133] In this embodiment, the reinforcement learning evaluation unit preferably adopts a Deep Deterministic Policy Gradient (DDPG) or Soft Action Evaluation (SAC) network architecture adapted to a continuous action space as the carrier of the optimization convergence constraint mapping rules. The reinforcement learning evaluation unit is specifically instantiated as a deep neural network feature extraction pipeline.

[0134] This feature extraction pipeline consists of a multilayer fully connected perceptron (MLP) or deep convolutional module. It serves as the Actor-Critic network foundation for reinforcement learning (DDPG or SAC algorithms) and is embedded and solidified within the computational hub of the reinforcement learning evaluation unit. All feature matrices that are routinely routed to the reinforcement learning evaluation unit are carried by this deep neural network feature extraction pipeline, which performs optimization convergence and feature dimensionality reduction operations.

[0135] The node state observation space feature matrix serves as the input tensor. This vector undergoes forward propagation and nonlinear activation (ReLU or Tanh activation function) through a multi-layer fully connected perceptron network for deep business feature extraction. In the flow mapping of the policy network, the core driving force for optimizing the convergence constraint mapping rule originates from the preset global business reward function. Specifically, the global business reward function integrates the convergence objective and the negative penalty constraint term into a global continuous cost evaluation equation. At the end of each control cycle t, the issued switching action and power state are evaluated and calculated to obtain a clear scalar reward value. The mapping logic of this state-action reward value is as follows:

[0136]

[0137] in It is a single-step state-action reward value that represents the output of the current control step t in the deep network service flow; in this embodiment, the preferred range is a continuous space of negative real numbers, and the closer the value is to 0, the better the global convergence state is. Characterizes the absolute error of the frequency offset of the i-th parallel node; In the evaluation of system control effect, the transient fluctuation spatial variance matrix characteristics of the active and reactive power sampling of global parallel nodes in this cycle are respectively characterized. The discrete characteristic indicator function; its state transition value is 1 (when the switching action of adjacent time series is identified). When a flip jump occurs (or 0 when no switch jump is performed), the mechanical and communication overhead penalty for quantifying asynchronous actions is used.

[0138] In the network convergence constraint, the multidimensional positive adjustment bias weight coefficient represents the corresponding state deviation; in this embodiment, the preferred range is an empirical constant in the interval (0, 1.0); In the action penalty mechanism, the jitter penalty gain scalar represents the suppression of frequent switching; in this embodiment, the preferred range is [0.1, 0.5], and the preferred value is 0.2.

[0139] The evaluation decision matrix in this embodiment includes discrete Boolean variables (optimal switching state identifier, controlling the opening / closing of solid-state circuit breakers) that guide the response of the underlying actuators, and continuous floating-point variables (output power rating coefficient, characterizing the inverter's load regulation ratio from 0% to 100%).

[0140] For the i-th parallel node, its independent one-dimensional normalized state scalar is stacked as row vectors according to a preset business topology dimension to generate a node-level multi-dimensional state vector. Then, the vectors of the N effective parallel nodes in the entire network are orthogonally connected to construct a node state observation space feature matrix representing the objective operating state of the global power grid. The execution logic is represented as follows: ;

[0141] in, In the business data flow, the node status observation space feature matrix represents the current control step t and the entire network. In this embodiment, it is preferably a high-dimensional two-dimensional continuous feature tensor with a dimension of N×3. N represents the total number of valid parallel nodes that have successfully accessed and have not been isolated and released in the business mechanism. Its value is a set of discrete positive integers greater than or equal to 1. and In the data structure mapping, the row traversal stacking operator across node dimensions and the cross-domain scalar horizontal concatenation operator within a single node feature vector are respectively represented.

[0142] These represent the frequency offset, active power, and reactive power normalized characteristic scalars, respectively, which are input to the i-th parallel node after the preceding extremum processing and whose values ​​are forced to converge to the dimensionless interval [0,1].

[0143] Furthermore, the main-end characteristic delay of each parallel node is extracted, a minimal positive real number compensation coefficient for preventing overflow is introduced, and a time-series reciprocal feature mapping algorithm with safety bias is executed to generate a time-decay penalty index that characterizes the time-sensitivity confidence of each node's data and whose value is constant in the (0,1) interval.

[0144] Determine whether the time-deterioration penalty index falls below the preset feature confidence limit value.

[0145] In response to the fact that the time decay penalty exponent does not fall below the aforementioned limit value, the state-action value calculation logic within the optimization convergence constraint mapping rule is triggered. This logic calculates the time difference loss based on the evaluation error of the current state and action combination, and performs a backpropagation action of the value error along the network layers, thereby extracting and outputting the initial expected cumulative reward gradient representing the optimization convergence direction. The time decay penalty exponent is used to perform a matrix dot product suppression operation on the initial expected cumulative reward gradient in the corresponding dimension to obtain the updated state-action optimization gradient.

[0146] In response to the time decay penalty index falling below the feature confidence limit, the absolute blocking control logic is activated to generate a state over-limit isolation flag. Based on the state over-limit isolation flag, the state action optimization gradient corresponding to the node is forcibly reset to a zero vector to block the update of erroneous weights. Simultaneously, in the current output evaluation decision matrix, a keep-alive control word output indicating the maintenance of its preceding steady-state operating load is assigned to the node.

[0147] It should be noted that the zero vector of the gradient for state-action optimization is... This only deprives the abnormal node of its ability to interfere with the iterative learning of deep network parameters. Simultaneously, in the evaluation decision matrix output by the current forward inference, a keep-alive control word instructing the node to maintain its preceding steady-state operating load is assigned. Feature dimensionality reduction is performed based on the gradient optimization of the state-action after constraint updates. Negative gain actions that cause an increase in frequency offset are forcibly filtered out, and the state-action gradient direction that minimizes the expected power fluctuation is selected for output mapping, generating an evaluation decision matrix composed of the optimal switching state and output power coefficient of each node. The mathematical logic and constraint boundaries of the above state determination and dimensionality reduction suppression are as follows:

[0148]

[0149] in In network flow, the learnable continuous weight distribution matrix in each hidden layer of the current policy network is represented; In the optimization mechanism, the characterization is... The parameterized optimal switching strategy function maps the current grid state to an action probability distribution. This function is instantiated as a deep neural network containing a multilayer perceptron (MLP). Its input receives the feature matrix s of the node state observation space. After nonlinear feature dimensionality reduction through the hidden layer, its output is based on the joint action probability distribution and outputs two decision loads a in parallel and decoupled (the aforementioned discrete optimal switching state identifier and continuous output power quota coefficient).

[0150] s represents the node state observation space feature matrix extracted within the current control cycle in the business data source; a represents the inverter active and reactive power output adjustment action vector output by the model in response to the current state at the instruction issuing layer. Representation in strategy Under constraints, the expected state-action value function evaluation quantity is the result of the current node state observation space feature matrix s executing the inverter active and reactive power output adjustment action vector a. The global objective optimization function representing the reinforcement learning evaluation unit represents the state-action reward value mapping logic. Characterizing the continuous weight distribution matrix The gradient calculation of the differential operator is used to indicate the optimal approximation direction for parameter iterative updates; The optimal switching policy function is represented by the function that follows the current network policy. Under this premise, calculate the mathematical expectation value of all valid state-action interaction trajectories generated by the system exploration; : Characterizes the conditional probability distribution function of the deep neural network that calculates and outputs the corresponding action a (representing the optimal switching state identifier and output power quota coefficient) under the given input of the node state observation space feature matrix s;

[0151] Furthermore, in step S5, the output evaluation decision matrix or the logical matching result of the secondary basic security isolation component is obtained. The routing source branch of the current business data flow is determined, and the differentiated instruction assembly logic is executed:

[0152] In response to the data being routinely routed to the reinforcement learning evaluation unit, the optimal power switching strategy for each parallel node in the evaluation decision matrix is ​​extracted; based on the high-dimensional feature mapping result of this strategy, the corresponding dynamic optimal inverter active and reactive power output adjustment command is generated, and the optimized dynamic control load is then issued.

[0153] In response to data being degraded to a secondary basic security isolation component under abnormal operating conditions, the structured signaling of the logical matching result is parsed.

[0154] The signaling returned from the aforementioned routing branches is structurally decomposed, and the logical matching results are parsed to extract the out-of-limit status flag bit represented by the discrete out-of-limit identifier that indicates a physical boundary violation. The static degradation quota parameter (pre-set discrete polymorphic control constant) that is pre-hard-bound to this out-of-limit identifier is also extracted. In the actual business boundary, the discrete polymorphic control constant is specifically represented as a setpoint constant used to force the inverter to operate at a minimum safety baseline of 20% or to directly trigger the solid-state circuit breaker to lock out.

[0155] Based on deterministic discrete multi-state control constants, the most conservative active and reactive power output regulation commands of the inverter are directly assembled and generated. This completes the shortest operational closed loop from state matching to physical regulation command generation.

[0156] After generating the active and reactive power output adjustment commands for the inverter, the underlying mapping and response for cross-domain collaboration are executed.

[0157] Based on the underlying control transmission link, the active and reactive power output adjustment commands of the inverter are sent to the corresponding parallel nodes of the AC distribution network. In response to the inclusion of the aforementioned static degradation quota parameters in the inverter active and reactive power output adjustment commands, the degradation status code carried in the prefix of the same business logic indicator business control word is extracted.

[0158] The degradation status code is passed through to the front-end intelligent edge acquisition module of the parallel node for extraction, which directly triggers the extension configuration of the clock frequency of the data acquisition interface of the node (synchronously triggering the degradation reset of the reporting cycle of the source data acquisition interface), thereby achieving absolute containment of the input digital concurrent traffic from the root.

[0159] Constrained by the objective physical execution boundaries of solid-state circuit breakers and inverter control units within parallel nodes of the AC power distribution network, the active and reactive power output adjustment commands of the inverter are parsed through the underlying network communication interface.

[0160] Extract the running status adjustment parameters contained in the instruction (the business logic indicator and the running status adjustment parameters are specifically represented by the underlying communication message control word and the corresponding protocol parsing field that carry the switching control signal at the underlying mechanism).

[0161] When the inverter's active and reactive power output adjustment commands are packaged and sent out, they contain encapsulated operating status adjustment parameters. These parameters serve as the system-level control signaling carrier. After being transparently transmitted to the target node, they are parsed by the parallel nodes of the AC distribution network at the receiving end. This guides the internal end-effectors (solid-state circuit breakers and inverter control units) to asynchronously execute the corresponding state switching and power parameter adjustment actions based on the parameter values. This achieves physical load offloading at the execution end, forcing the real-time power load of the parallel nodes of the AC distribution network to converge towards the preset grid frequency safety tolerance baseline.

[0162] In response to the target parallel node's underlying communication protocol stack receiving and parsing the operating status adjustment parameters, the node's built-in local microprocessor distributes the structured parameters as Boolean-state switching signals and continuous-state power setpoints. Asynchronous timing latching logic is forcibly introduced: the local microprocessor sends the switching signal to the solid-state circuit breaker driver, triggering it to perform a zero-position closing or opening action; the local state machine maintains an active polling waiting cycle until it captures a physical latching response frame returned by the underlying hardware, indicating that the circuit breaker state has stabilized and flipped; after acquiring this response frame, the data channel is unlocked, and the power setpoint is transparently transmitted to the PWM duty cycle register module of the inverter control unit to execute the power parameter adjustment action; ultimately, the real-time power load of the parallel nodes in the AC distribution network is forced to smoothly converge towards the preset grid frequency safety tolerance baseline. By introducing the above-mentioned asynchronous timing decoupling and state confirmation mechanism of switching before power is applied, logical conflicts caused by concurrent control signaling leading to idling overload of the underlying power module before grid connection are avoided, thus ensuring the absolute timing safety of high-frequency scheduling commands at the physical end.

[0163] The specific configuration of the operating state adjustment parameters and the preset grid frequency safety tolerance baseline in this embodiment is described as follows: The operating state adjustment parameters are structured data frames serialized and spliced ​​from the composite decision results output by the reinforcement learning evaluation unit (or secondary basic safety isolation component) in the preceding steps. The frame header of this parameter encapsulates a discrete optimal switching state identifier, which is mapped to a Boolean control word, specifically for the solid-state circuit breaker at the receiving end to parse in order to perform zero-position closing or opening action; the frame body of this parameter encapsulates a continuous output power rating coefficient (or static degradation rating parameter), which is mapped to a power setpoint ratio, specifically for the inverter control unit to parse in order to perform flexible adjustment of the PWM duty cycle. The preset grid frequency safety tolerance baseline refers to the reference frequency of the static system of the AC distribution network (… The absolute physical safety tolerance band is set at 50Hz or 60Hz. In this embodiment, the baseline is quantized and configured as an interval. ,in The preset safety offset limit threshold is used (preferably 0.2Hz~0.5Hz in this embodiment). When the parallel nodes of the AC distribution network analyze the operating status adjustment parameters, the solid-state circuit breaker is driven to complete the physical grid connection according to the frame header control word, and then the inverter output is adjusted according to the ratio given in the frame body. The direct physical constraint target of this asynchronous closed-loop action is the voltage frequency offset of the parallel node obtained locally by the forced voltage drop. This forces the dynamic frequency of the entire network to regress and stabilize at the aforementioned levels. Within the tolerance baseline range, global convergence of distributed collaborative scheduling is achieved.

[0164] Furthermore, it acquires out-of-bounds physical warning information that triggers the absolute frequency safety tolerance red line (safe operating limit state) of the AC power distribution network; before the parallel nodes parse the remotely issued instructions, the system background daemon process reads the out-of-bounds physical warning information in parallel; when the aforementioned verification pipeline is triggered, the following multi-branch judgment is executed at the underlying interaction node:

[0165] In response to the detection that the local real-time frequency exceeds the physical limit hard interrupt threshold set by the inherent grid connection protocol of the node, or if it is determined that the timestamp of the received remote instruction has exceeded the absolute failure window compared to the local clock, the system triggers the highest priority logical circuit breaker flag. If the flag is effective, it forcibly intercepts and blocks the abnormal asynchronous collaborative data stream in the current receive stack. The system blocks write permissions to the external communication interface, retrieves the baseline disaster prevention quota table embedded in the local security ROM, and forcibly replaces the original remote output control parameters with independent self-healing control instructions generated based on the disaster prevention quota (local fallback state control parameters independent of remote collaborative scheduling). Then, through this independent instruction, it forcibly disconnects the connection or maintains an extremely low keep-alive load, thereby achieving absolute self-protection and secure closed-loop control of the underlying physical node state when abnormal conditions such as remote scheduling blockage or malicious attack and congestion of the communication channel occur in the core application logic.

[0166] In the core business execution chain, this method outputs an evaluation decision matrix consisting of continuous output power quota coefficients and discrete optimal switching state identifiers for the access of multi-source parallel nodes. Under over-limit conditions, it also outputs discrete static degradation quota parameters. When these quota parameters switch in different routing branches, and when the state action optimization gradient is forcibly reset to the direction of the zero vector evolution, it signifies a smooth transition of the underlying business state from conventional refined digital twin feature optimization to a degraded physical bottom-line defense state.

[0167] There is a non-linear negative bias relationship between the master-end characteristic latency and the threshold for triggering process degradation. When the extracted latency exceeds the asynchronous timing tolerance threshold, a cross-domain multiplication operation is triggered internally. The resulting negative bias decay term monotonically increases as the deviation widens, leading to a continuous reduction in the threshold for triggering process degradation from the basic branch switching baseline. Designing this cross-domain mapping relationship as a non-linear downward pressure logic aligns with the throughput characteristics of distributed services. When congestion and jitter occur in the communication link, the timeliness of the data to be processed deteriorates marginally. This negative correlation design of the parameter supports the core beneficial effect of this method in effectively preventing the risk of cascading collapse of global application services.

[0168] The time-decrease penalty index exerts absolute multiplicative inhibition and hard truncation on the state-action optimization gradient. This method extracts the inverse feature mapping of high-frequency concurrent business data to generate a time-decrease penalty index representing confidence, and uses it as an attention mask to perform a matrix dot product with the initial expected cumulative reward gradient along the corresponding dimension. Under this mapping constraint, the higher the data latency, the closer the time-decrease penalty index is to the feature confidence limit, and the optimization gradient vector space is forcibly compressed proportionally; if it falls below this feature confidence limit, the gradient is immediately reset to a zero vector. The design of hard-binding network backpropagation to the physical temporal decay rate cuts off the path of dirty data contaminating the action value matrix, ensuring strong convergence of reinforcement learning evaluation decisions under complex constraints.

[0169] The method in this embodiment is deployed in a distributed AC power distribution network high-frequency data collaborative interaction service execution environment. The central service interface needs to process multiple service data streams carrying discrete voltage frequency offsets and power sampling characteristics in real time and in parallel. Moreover, the physical links from each parallel node to the underlying gateway have objective asymmetric communication attenuation and sudden concurrent surges. Based on the aforementioned service environment, when the core data acquisition interface introduces external delay disturbances and internal computing power fluctuations, the data flow logic within this method will undergo dynamic transitions. Based on the pre-configured extreme timing tolerance span and basic branch switching baseline, as the delay characteristics and backlog scale increase, the normalized penalty multiplier maintained within the method will be proportionally amplified, thereby triggering cross-dimensional suppression of the initial expected cumulative reward gradient, and forcing the service flow pipeline to forcibly switch from the deep neural network forward propagation branch to the pre-set one-dimensional physical safety threshold extreme value comparison branch at the gateway decision layer.

[0170] By setting approximation parameters that cover the boundary from normal security to extreme deterioration, the internal deduction state of the execution logic judgment and routing of this method is as follows.

[0171] Table 1: Example of predictive numerical simulation based on this embodiment under the superimposed degradation of multi-source communication latency and internal concurrent load.

[0172]

[0173] The spatiotemporal mapping transition trends shown in the table are controlled by the asynchronous timing tolerance threshold baseline and the extreme timing tolerance span constant fixed in the system's global static runtime configuration parameter pool.

[0174] In response to the aforementioned deteriorating working conditions, traditional distributed network architectures rely on statically fixed resource scheduling full-load warning levels when handling massive, high-frequency burst requests from heterogeneous edge nodes. When encountering sudden high-frequency congestion and timing drift as set in input group two (latency surges to 80ms) in the table above, the static threshold model of traditional solutions inevitably continues to import dirty data into the core optimization link, causing the global model parameter update chain to suffer severe pollution from high-latency incomplete features, ultimately falling into computing queue deadlock and convergence divergence. In contrast, the method in this embodiment, when the load reaches 0.8, utilizes the dynamically leveraged, lowered gateway threshold of 0.626 to preemptively puncture and cut off the decision network, cutting off the redundant complex multiplication and addition calculations of the high-dimensional optimization matrix before the reinforcement learning business flow buffer channel experiences substantial memory stack overflow.

[0175] In this embodiment, the optimization convergence constraint mapping rule is a composite training and weight update mechanism configured within the feature extraction pipeline of the deep neural network. This rule includes the following two levels of logical constraints:

[0176] Basic Optimization Layer. In this layer, the Critic network branch inside the pipeline evaluates the temporal difference loss of the current state-action combination based on a preset global business reward function (a comprehensive cost model that integrates frequency offset, power fluctuation variance, and high-frequency switching penalty), and performs backpropagation of the error along the network layer to calculate the initial expected cumulative reward gradient representing the ideal optimization direction.

[0177] Dynamic constraint intervention layer. After obtaining the initial expected cumulative reward gradient, this mapping rule forcibly interrupts the regular gradient update process, introducing a time-decay penalty index that characterizes the degree of obstruction in underlying physical communication. This rule stipulates that the time-decay penalty index is used to perform a cross-dimensional multiplicative suppression operation on the pre-calculated initial expected cumulative reward gradient; and when the penalty index falls below the feature confidence limit, the optimization gradient of the corresponding node is forcibly reset to a zero vector. This ensures that the optimization convergence process of the deep neural network conforms to the freshness confidence limit of the business data (physical constraint mapping), thereby blocking the path of high-latency dirty data contaminating the global evaluation decision matrix at the underlying architecture level.

[0178] Figure 1 In the diagram, Path1 and Path2 represent the two paths generated by the binary over-limit logic gateway.

[0179] PEN stands for parallel node in AC distribution network. It represents the underlying execution object and heterogeneous data acquisition source of smart grid, and is responsible for continuously reporting the discrete voltage and current sampling sequence and operating load status of the node.

[0180] SSCB stands for Solid State Circuit Breaker, which represents front-end physical protection hardware built into the parallel node. It is used to receive asynchronous commands from the upper layer and perform real-time physical actions such as closing at zero position or forcibly disconnecting.

[0181] ICU stands for Inverter Control Unit, which is located within the same node as the solid-state circuit breaker. After capturing the physical latch response, it is responsible for unlocking the data channel and performing flexible adjustment of active and reactive power output parameters.

[0182] CDC stands for Collaborative Dispatch Center Business Interface. It represents the system control access layer, which is responsible for intercepting and obtaining underlying communication requests, extracting structured identity identifiers, and performing discrete comparison and verification of the underlying communication whitelist.

[0183] MFLC represents the master-side latency and load balancing calculation logic, which is the core preprocessing flow used to extract the master-side characteristic latency. And assess branch flow load rate Based on the calculation results, the flow degradation trigger threshold is dynamically updated and pushed down. .

[0184] BLGW stands for Binary Cross-Limit Logic Gateway, representing the core control center's branch point, based on the currently acquired assessment branch flow load rate. With dynamically updated flow degradation trigger threshold Based on the algebraic comparison results, perform cross-domain dual-track bypass route redirection actions to adjust system load.

[0185] RLEU stands for Reinforcement Learning Evaluation Unit, representing a deep neural network pipeline that carries out dimensionality reduction of high-dimensional features and policy optimization. When the degradation and splitting conditions are not triggered, it extracts the expected cumulative reward gradient and outputs the optimal switching policy based on the node state observation space feature matrix.

[0186] SBSIC stands for Secondary Basic Security Isolation Component, representing a disaster recovery defense logic block. In response to the triggering of concurrency over-limit conditions, it bypasses high-dimensional calculations and directly performs a one-dimensional comparison based on internally preset physical boundary thresholds, outputting a logical matching result carrying static degradation quota parameters.

[0187] RIG stands for Regulation Instruction Generation and Distribution Pipeline. It represents a closed-loop output node responsible for fusing reinforcement learning evaluation decision matrices or safety isolation logic results to generate inverter active and reactive power output regulation instructions containing timing decoupling prefixes, and asynchronously distributing them to parallel nodes.

[0188] The distributed collaborative scheduling method of this invention achieves spatiotemporal decoupling in the underlying business flow. It intercepts and verifies access requests from a massive number of parallel nodes through the business interface of the collaborative scheduling center, synchronously acquiring timing and power characteristics. Based on this, the business pipeline is pushed into the main-end latency and flow load calculation logic to measure the current communication timing jitter and system computing stack load, thereby deduce the flow degradation trigger threshold in real time. When the data flow arrives at the binary over-limit logic gateway, this gateway performs a core dual-track diversion action based on the aforementioned dynamic critical conditions: under normal safe operating conditions, data is smoothly diverted to the reinforcement learning evaluation unit for global optimization and complex dimensionality reduction mapping; however, once it is determined that the concurrent load exceeds the limit or the timing is distorted, the gateway instantly cuts off the optimization path, forcibly redirecting the signaling to the secondary basic security isolation component, performing discrete over-limit verification with simplified rules to avoid computing resource collapse. The two branches merge into the adjustment command generation pipeline based on their own judgment results, assemble and generate asynchronous control signals, and guide the solid-state circuit breaker and inverter control unit at the physical layer to complete the closed-loop switching and power regulation actions in sequence.

[0189] The above description of the embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. It should be noted that those skilled in the art can make several improvements and modifications to the present invention without departing from the principles of the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

Claims

1. A distributed cooperative scheduling method for energy nodes based on reinforcement learning, wherein the method is executed by a cooperative scheduling center, characterized in that, The specific steps include: Step S1: Obtain the voltage frequency offset and real-time active and reactive power sample values ​​of the parallel nodes that characterize the operating load of the parallel nodes in the AC distribution network, and extract the timestamp discrete difference of the voltage frequency offset and real-time active and reactive power sample values ​​reaching the service interface of the collaborative dispatch center, and label the timestamp discrete difference as the main end characteristic delay. Step S2: Extract the business concurrency buffer depth of the pre-configured reinforcement learning evaluation unit and multi-dimensional rule evaluation component, and calculate the evaluation branch flow load rate based on the business concurrency buffer depth; Step S3: Calculate the weighted bias value of the main end characteristic delay and the evaluation branch flow load rate, and update the weighted bias value to the flow degradation trigger threshold value; Step S4: Determine whether the currently acquired evaluation branch flow load rate meets the binary over-limit logic condition set based on the flow degradation trigger threshold. In response to the fulfillment of the binary over-limit logic condition, route redirection is triggered, and the voltage frequency offset and real-time active and reactive power sampling values ​​of the parallel node are forcibly routed to the secondary basic safety isolation component to perform discrete over-limit verification and matching based on the preset physical boundary threshold, and output the logic matching result carrying the static degradation quota parameter. In response to the failure to meet the binary over-limit logic condition, the voltage frequency offset of the parallel node and the real-time active and reactive power sampled values ​​are routinely routed to the pre-trained reinforcement learning evaluation unit. The reinforcement learning evaluation unit is configured to take the voltage frequency offset of the parallel node and the real-time active and reactive power sampling values ​​as inputs, and output an evaluation decision matrix that represents the optimal power switching strategy of each parallel node. Step S5: Based on the evaluation decision matrix or the logical matching result of the secondary basic security isolation component, generate the corresponding inverter active and reactive power output adjustment command.

2. The energy node distributed cooperative scheduling method based on reinforcement learning according to claim 1, characterized in that: Following step S5, the following is also included: The active and reactive power output adjustment command of the inverter is sent to the corresponding parallel node of the AC power distribution network. The active and reactive power output adjustment command of the inverter is encapsulated with operating status adjustment parameters. The operating status adjustment parameters are configured for parsing by the parallel nodes of the AC distribution network at the receiving end, so as to guide the solid-state circuit breakers and inverter control units within the parallel nodes of the AC distribution network to asynchronously perform state switching and power parameter adjustment actions, so that the real-time power load of the parallel nodes of the AC distribution network converges to the preset grid frequency safety tolerance baseline.

3. The energy node distributed cooperative scheduling method based on reinforcement learning according to claim 2, characterized in that: Extract the discrete time stamp difference between the voltage frequency offset of the parallel nodes and the real-time active and reactive power sample values ​​reaching the service interface of the collaborative scheduling center, and label the discrete time stamp difference as the master-end characteristic delay, including: Open the data aggregation time window; extract the first arrival timestamp of the voltage frequency offset and real-time active and reactive power sampling values ​​of each parallel node of the AC distribution network that arrive at the business interface of the collaborative dispatch center. Extract the largest and smallest first arrival timestamps representing extreme value arrival records within the same data aggregation time window, and perform discrete absolute difference calculation based on the largest and smallest first arrival timestamps. Extract the output discrete absolute difference as the master-end feature delay.

4. The energy node distributed cooperative scheduling method based on reinforcement learning according to claim 1 or 3, characterized in that: Extract the business concurrency buffer depth of the pre-configured reinforcement learning evaluation unit and multi-dimensional rule evaluation component, and calculate the evaluation branch flow load rate based on the business concurrency buffer depth, including: The system periodically extracts the backlog of pending business and the concurrent execution channel size of the reinforcement learning evaluation unit; The backlog of pending business is weighted and summed with the concurrent execution channel size to generate a real-time backlog load value; The evaluation branch flow load rate is generated by dividing the real-time backlog load value by the preset node baseline throughput threshold.

5. The energy node distributed cooperative scheduling method based on reinforcement learning according to claim 4, characterized in that: Calculate the weighted bias value between the main-end characteristic latency and the evaluation branch flow load rate, and update the weighted bias value to the flow degradation trigger threshold, including: When the extracted master-end feature delay exceeds the system's preset asynchronous timing tolerance threshold, the absolute difference between the master-end feature delay and the asynchronous timing tolerance threshold is calculated, and the absolute difference is divided by the preset limit timing tolerance span to eliminate the dimension. The pure numerical ratio of the output is then extracted as the excess gain coefficient. A pre-configured smoothing penalty constant is introduced, and the excess gain coefficient, the evaluation branch flow load rate, and the smoothing penalty constant are multiplied together to generate a negative bias attenuation term; Obtain the preset basic branch switching baseline, and subtract the negative bias attenuation term from the basic branch switching baseline to obtain the updated flow degradation trigger threshold.

6. The energy node distributed cooperative scheduling method based on reinforcement learning according to claim 5, characterized in that: Based on the evaluation decision matrix or the logical matching result of the secondary basic security isolation component, a corresponding inverter active and reactive power output adjustment command is generated, including: When data is routed to the reinforcement learning evaluation unit, a corresponding dynamic optimal inverter active and reactive power output adjustment command is generated based on the evaluation decision matrix. When data is routed to the secondary basic security isolation component, the logical matching result is parsed to extract discrete over-limit identifiers that characterize physical boundary over-limits, and static degradation quota parameters that are pre-hard bound to the discrete over-limit identifiers are extracted. Based on the static degradation quota parameters, corresponding static degradation inverter active and reactive power output adjustment commands are generated. In response to the inverter's active and reactive power output adjustment command including the static degradation quota parameter, the adjustment command is configured as follows: after determining that the solid-state circuit breaker of the target AC distribution network parallel node has completed the state transition and outputs a physical latch response, the degradation reset of the reporting cycle of the source-end data acquisition interface of the parallel node is asynchronously triggered based on the timing decoupling logic.

7. The energy node distributed cooperative scheduling method based on reinforcement learning according to claim 6, characterized in that: The steps for obtaining the evaluation decision matrix specifically include: Obtain the voltage frequency offset and real-time active and reactive power sampling values ​​of each parallel node, extract the pre-set limit boundary extreme values ​​of each parameter, perform dimensionless normalization mapping on the voltage frequency offset and real-time active and reactive power sampling values ​​based on the limit boundary extreme values, and construct a node state observation space feature matrix that characterizes the objective operating state of the global power grid. To enhance the configuration of the deep neural network feature extraction pipeline pre-built in the learning evaluation unit, optimization convergence constraint mapping rules based on the global business reward function are applied. The node state observation space feature matrix is ​​input into the deep neural network feature extraction pipeline. Feature extraction and dimensionality reduction mapping are performed based on the optimization convergence constraint mapping rule. The output is the evaluation decision matrix composed of discrete optimal switching state identifiers and continuous output power quota coefficients.

8. The energy node distributed cooperative scheduling method based on reinforcement learning according to claim 7, characterized in that: Feature dimensionality reduction transformation is performed based on the pre-constructed node state observation space feature matrix and optimization convergence constraint mapping rules, including: Extract the master-end characteristic delay of each parallel node of the AC distribution network, and perform time-series inverse feature mapping to generate the time-degradation penalty index of each parallel node of the AC distribution network. The state-action value calculation logic inside the optimization convergence constraint mapping rule is triggered to perform backpropagation of value error, so as to extract the initial expected cumulative reward gradient representing the optimization convergence direction. The time decay penalty index is used to perform cross-dimensional multiplication suppression operation on the initial expected cumulative reward gradient to obtain the updated state-action optimization gradient. Based on the state action optimization gradient, the action space is traversed, and discrete power adjustment actions with negative gradient amplification are eliminated to complete the feature dimensionality reduction transformation.

9. The energy node distributed cooperative scheduling method based on reinforcement learning according to claim 8, characterized in that: Obtain the initial expected cumulative reward gradient output by the optimization convergence constraint mapping rule, and perform a cross-dimensional multiplication suppression operation on the initial expected cumulative reward gradient using the time decay penalty exponent to obtain the updated state-action optimization gradient, including: Determine whether the extracted time-deterioration penalty index falls below the preset feature confidence limit value; In response to the fact that the time-decrease penalty index does not fall below the feature confidence limit, the time-decrease penalty index and the initial expected cumulative reward gradient are multiplied by a matrix in the corresponding dimension. In response to the time decay penalty index falling below the feature confidence limit, the absolute blocking control logic is activated to generate a state over-limit isolation flag. Based on the state over-limit isolation flag, the state action optimization gradient corresponding to the node is forcibly reset to a zero vector, thereby blocking the erroneous weight allocation of high-latency features in the evaluation decision matrix.

10. The energy node distributed cooperative scheduling method based on reinforcement learning according to claim 9, characterized in that: Obtain the voltage frequency offset and real-time active and reactive power sample values ​​of the parallel nodes, which characterize the operating load of the parallel nodes in the AC distribution network. Specifically, this includes: Obtain discrete voltage sampling sequences and discrete current sampling sequences from the underlying interaction interface of the parallel nodes of the AC power distribution network; Extract the continuous zero-crossing time-scale features from the discrete voltage sampling sequence and perform reciprocal mapping to derive the real-time operating frequency parameter. Then, perform a differential operation between the real-time operating frequency parameter and the system rated frequency reference, and extract the absolute value of the difference as the voltage frequency offset of the parallel node. A fixed-time sliding window is established, and time-domain interpolation alignment mapping is performed on the introduced discrete voltage sampling sequence and the discrete current sampling sequence based on the local reference synchronization time scale. Then, within the time sliding window, point-by-point discrete multiplication and integral mean extraction are performed on the aligned discrete voltage sampling sequence and the discrete current sampling sequence to generate the real-time active power sampling value. The discrete voltage sampling sequence is reconstructed by an orthogonal phase shift to generate an orthogonal voltage sequence. The orthogonal voltage sequence is then multiplied point-by-point by the discrete current sampling sequence and the mean of the integral is extracted to generate the real-time reactive power sampling value.

Citation Information

Patent Citations

  • Virtual power plant collaborative optimization scheduling method and system based on deep reinforcement learning

    CN121076815A