A cloud-edge collaborative based HPLC dual-mode communication method and system

By employing a cloud-edge collaborative HPLC dual-mode communication method, which utilizes real-time acquisition of physical layer indicators on the terminal side, DDQN model decision-making on the edge side, and cloud-based federated distillation optimization, the reliability and latency issues of HPLC and HRF dual-mode communication in complex electromagnetic environments are resolved, achieving highly reliable and low-latency power IoT communication.

CN122160023APending Publication Date: 2026-06-05CHINA SOUTHERN POWER GRID COMPANY
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA SOUTHERN POWER GRID COMPANY
Filing Date
2026-02-11
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing HPLC and HRF dual-mode communication solutions face problems such as cross-network co-frequency interference, strong electromagnetic interference, and insufficient adaptive capabilities in complex electromagnetic environments, resulting in fluctuating communication reliability and excessively high response latency, making it difficult to guarantee the continuity of terminal services.

Method used

A cloud-edge collaborative HPLC dual-mode communication method is adopted. The terminal side collects physical layer indicators in real time to generate a state vector, and the edge side uses the DDQN model to make real-time decisions. The model is optimized by combining cloud federated distillation technology to form a three-level collaborative architecture, which realizes highly reliable and low-latency communication.

Benefits of technology

Maintaining high communication success rate and low bit error rate in strong electromagnetic interference scenarios, reducing response latency to millisecond level, meeting the real-time requirements of the power Internet of Things, and improving the generalization capability across regional scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122160023A_ABST
    Figure CN122160023A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of electric power communication, in particular to an HPLC dual-mode communication method and system based on cloud-edge cooperation, which generates a state vector through real-time collection of physical layer indexes on a terminal side, generates an optimal action according to the state vector and a DDQN model on an edge side, feeds back a communication result and a state vector of a next time slot to the edge side after the terminal side executes the optimal action, and calculates an instant reward and a target Q value of the DDQN model on the edge side, so that a closed-loop decision flow of "state-action-feedback" is formed; the edge side updates parameters of the DDQN model through global model parameters issued by the cloud, realizes a closed loop of "local training-cloud aggregation-global optimization", and thus can still maintain a high communication success rate and a low bit error rate in a strong electromagnetic interference scene such as lightning stroke and large-current switching; and a dual-mode communication solution with high reliability, low time delay and low power consumption is provided for a power internet of things.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of power communication technology, and in particular to an HPLC dual-mode communication method and system based on cloud-edge collaboration. Background Technology

[0002] To mitigate the uncertainties of a single HPLC (High-speed Power Line Communication) communication link, a dual-mode communication scheme combining HPLC and HRF (High-Frequency Radio Frequency) achieves complementary redundancy between the power line and wireless channels through a "one-chip dual-mode" baseband design. Compared to single-mode HPLC, this scheme improves data acquisition efficiency by over 40%. However, in practical applications, this scheme still faces three challenges: First, parallel operation of both modes is prone to cross-network co-frequency interference, with signals from different or adjacent areas interfering with each other, significantly reducing the signal-to-noise ratio. Second, strong electromagnetic interference (such as lightning strikes, arc discharges, and high-current switching) can simultaneously affect both the power line and wireless links, causing simultaneous degradation of the dual-mode system and compromising the continuity of terminal services. Third, existing link switching and anti-interference strategies mainly rely on fixed rules or threshold control, lacking adaptability to complex dynamic environments, making it difficult to guarantee system performance. Although the patent application with publication number CN120263228A proposed a hybrid data transmission control method that combines four-dimensional QoS dynamic weight allocation with fractional-order HJB equation optimization, achieving multi-dimensional service quality collaborative control and dynamic optimization, it still has prominent problems such as communication reliability fluctuations and excessively high response delays under strong electromagnetic interference environments. Summary of the Invention

[0003] Therefore, it is necessary to provide a cloud-edge collaborative HPLC dual-mode communication method and system to address the issue of ensuring high reliability and real-time performance of dual-mode communication in complex electromagnetic environments.

[0004] Firstly, this application provides a cloud-edge collaborative HPLC dual-mode communication method. The method includes:

[0005] Step S1: The terminal generates the state vector of the current time slot based on the real-time collected physical layer indicators;

[0006] Step S2: The edge side generates the optimal action for the current time slot based on the state vector and DDQN model of the current time slot, and sends the optimal action to the terminal side;

[0007] Step S3: After the terminal side completes the optimal action, it feeds back the communication result of the current time slot and the state vector of the next time slot to the edge side.

[0008] Step S4: The edge side calculates the instant reward of the current time slot based on the communication result of the current time slot, and calculates the target Q value of the DDQN model based on the instant reward of the current time slot and the state vector of the next time slot.

[0009] Step S5: When the preset update time is reached, the edge side selects samples from the experience replay pool to calculate the output probability distribution of the current network, and uploads the output probability distribution of the current network to the cloud.

[0010] Step S6: The cloud generates the output probability distribution of the global model based on the output probability distribution of the current network, updates the global model parameters based on the output probability distribution of the current network and the output probability distribution of the global model, and sends the global model parameters to the edge side.

[0011] Step S7: The edge side updates the network parameters of the DDQN model based on the global model parameters.

[0012] Furthermore, the state vector includes channel gain, interference power, buffer queue length, remaining energy state, and delay spread characteristics.

[0013] Furthermore, the optimal dynamic is a triplet comprising a frequency modulation channel index, a transmit power level, and a dual-mode communication standard selection.

[0014] Furthermore, the communication results include information age, energy consumption, and signal-to-noise ratio.

[0015] Furthermore, the formula for calculating the instant reward is as follows:

[0016]

[0017] In the formula, The signal-to-noise ratio for the current time slot, The information age in the current time slot. The energy consumption for the current time slot, , and All weights are adjustable.

[0018] Furthermore, the formula for calculating the signal-to-noise ratio is:

[0019]

[0020] In the formula, Indicates the actual delivered power. For channel gain, Let I be the noise power and I be the interference power; wherein, the expression for the actual delivered power is:

[0021]

[0022] In the formula, For transmission power, This is the reflection coefficient.

[0023] Furthermore, the expression for the distillation loss function in the cloud is:

[0024]

[0025] In the formula, This represents the distillation loss value. This represents the weight of each edge node in the edge side, and KL(·) is the Kullback-Leibler divergence. This represents the output probability distribution of the current network at the i-th edge node. This represents the output probability distribution of the global model.

[0026] Furthermore, the output probability distribution of the current network is the output probability distribution after introducing noise.

[0027] Secondly, this application also provides a cloud-edge collaborative HPLC dual-mode communication system. The system includes a terminal side, an edge side, and a cloud. Each terminal node on the terminal side includes a data acquisition module and an action feedback module. The data acquisition module generates a state vector for the current time slot based on real-time acquired physical layer indicators. The action feedback module, after executing the optimal action, feeds back the communication result of the current time slot and the state vector for the next time slot to the edge side.

[0028] Each edge node in the edge side includes an action selection module, a Q-value calculation module, and a network parameter update module. The action selection module generates the optimal action for the current time slot based on the state vector of the current time slot and the DDQN model, and sends the optimal action to the terminal side. The Q-value calculation module calculates the instant reward for the current time slot based on the communication result of the current time slot, and calculates the target Q-value of the DDQN model based on the instant reward of the current time slot and the state vector of the next time slot. The network parameter update module selects samples from the experience replay pool to calculate the output probability distribution of the current network when a preset update time is reached, uploads the output probability distribution of the current network to the cloud, and updates the network parameters of the DDQN model based on the global model parameters issued by the cloud.

[0029] The cloud is used to generate the output probability distribution of the global model based on the output probability distribution of the current network, update the global model parameters based on the output probability distribution of the current network and the output probability distribution of the global model, and send the global model parameters to the edge side.

[0030] Furthermore, each of the terminal nodes is equipped with a programmable impedance matching module and a Turbo code rate linkage module.

[0031] The aforementioned HPLC dual-mode communication method and system based on cloud-edge collaboration achieves highly reliable and low-latency communication through a three-level collaborative architecture of terminal-edge-cloud. Specifically, the terminal side collects physical layer indicators in real time to generate a state vector, which serves as the state input for the DDQN model on the edge side. The edge side generates the optimal action based on the state vector and the DDQN model and sends it to the terminal side for execution. After execution, the terminal side feeds back the communication result and the state vector of the next time slot to the edge side, forming a closed-loop decision flow of "state-action-feedback" to ensure that the decision adapts quickly to environmental changes. Meanwhile, the edge stores historical samples through an experience replay pool, selects samples at preset update times to calculate the current network output probability distribution, and uploads it to the cloud. The cloud uses federated distillation technology to fuse the distribution of multiple edge nodes, generate a global model output probability distribution, and update global parameters. Finally, it sends the data to the edge to update the DDQN model, realizing a closed loop of "local training - cloud aggregation - global optimization". This ensures high communication success rate and low bit error rate even in scenarios with strong electromagnetic interference such as lightning strikes and high-current switches. By delegating decision-making to edge nodes, data backhaul delay is avoided, and response latency is reduced to the millisecond level, meeting the real-time requirements of the power IoT. Furthermore, cloud federated distillation can achieve continuous evolution of the global model without uploading raw data, improving the generalization capability across regional scenarios and providing a highly reliable, low-latency, and low-power dual-mode communication solution for the power IoT. Attached Figure Description

[0032] Figure 1 This is a schematic diagram of the structure of an HPLC dual-mode communication system based on cloud-edge collaboration in one embodiment;

[0033] Figure 2 This is a flowchart illustrating a cloud-edge collaborative HPLC dual-mode communication method in one embodiment. Detailed Implementation

[0034] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0035] Example 1

[0036] The cloud-edge collaborative HPLC dual-mode communication method provided in this embodiment can be applied to, for example... Figure 1 In the application environment shown, the terminal nodes on the terminal side can be, but are not limited to, smart meters, distributed photovoltaic monitoring devices, energy storage management modules, electric vehicle charging facilities, etc., and are connected to the edge side via high-speed power lines or high-speed wireless radio frequency. The edge side includes multiple edge nodes deployed in distribution substations or regional aggregation points, and each edge node connects to multiple terminal devices. The cloud can be a central server for the power Internet of Things or a regional data center, connected to each edge node on the edge side via a broadband network.

[0037] like Figure 2 As shown, the HPLC dual-mode communication method based on cloud-edge collaboration provided in this embodiment specifically includes the following steps:

[0038] Step S1: The terminal generates the state vector of the current time slot based on the physical layer indicators collected in real time.

[0039] On the terminal side, high-precision measurement modules integrated into terminal devices such as smart meters and sensors are used to collect physical layer indicators that characterize the communication environment and device status in real time, including key parameters such as channel gain, interference power, buffer queue length, remaining energy status, delay spread characteristics, Age of Information (AoI), and energy consumption. Among them, channel gain is a parameter describing the change in signal strength during transmission, which is quantified by a power detection circuit after the signal attenuates through the transmission medium; interference power refers to the power of other signals or noise in the communication system besides the useful signal, which is obtained by capturing the noise energy of non-target frequency bands by a spectrum analyzer; buffer queue length refers to the size of the buffer space used to store data to be transmitted in the communication system or the amount of data currently stored, which is obtained by calculating the real-time data backlog by the difference between the memory read and write pointers; remaining energy status is usually used to describe the remaining battery power or energy level of wireless communication devices (such as sensor nodes), which is calculated based on the voltage monitoring chip; delay spread characteristics refer to the distribution characteristics of signal delay caused by factors such as multipath effects during signal transmission, which is measured by the channel sounding reference signal to measure the multipath delay distribution; information age (AoI) is the time interval from the generation of information to its reception and processing, reflecting the "freshness" of the information, which can be calculated by the difference between the data generation timestamp and the current system clock; energy consumption is the energy consumed by the device during communication, which directly affects the device's battery life and operating costs, and is statistically analyzed by a power integration circuit to determine the real-time energy consumption of the communication module.

[0040] Meanwhile, as the core unit for millisecond-level decision-making, the edge side faces the unique multipath fading effect of power line transmission, random interference caused by multi-user contention, and the dual constraints of energy consumption and service latency. To accurately characterize this dynamic and complex electromagnetic environment and resource constraints, this embodiment innovatively constructs a five-dimensional state vector from the core parameters directly affecting decision-making among the indicators collected by the terminal side—channel gain, interference power, buffer queue length, remaining energy state, and delay spread characteristics—according to Equation (1). Each terminal node initiates a synchronous measurement process in each time slot, encapsulates the state vector into a data packet, and transmits it to the edge side via a dual-mode communication module (HPLC / HRF), providing a real-time environmental awareness basis for edge nodes to achieve low-latency and high-reliability decision-making in complex electromagnetic environments.

[0041] (1)

[0042] In the formula, This represents the state vector of the current time slot. This represents the channel gain of the current time slot. This indicates the interference power in the current time slot. This indicates the length of the buffer queue for the current time slot. This indicates the remaining energy state of the current time slot. This indicates the delay spread characteristic of the current time slot.

[0043] Step S2: The edge side generates the optimal action for the current time slot based on the state vector of the current time slot and the DDQN model, and sends the optimal action to the terminal side.

[0044] To address the policy bias problem caused by Q-value overestimation in traditional DQN, this embodiment employs a DDQN (Double Deep Q Network) model at the edge to achieve real-time decision-making in complex electromagnetic environments. The DDQN model comprises a current network and a target network. The current network is responsible for selecting the optimal action based on the current state and calculating the instantaneous Q-value. It typically consists of multiple fully connected layers and is optimized using gradient descent. The optimal action is a triplet comprising the frequency modulation channel index, transmit power level, and dual-mode communication selection, expressed as equation (2):

[0045] (2)

[0046] In the formula, The optimal action for the current time slot. This is a frequency modulation channel index used to specify the communication frequency band to avoid multipath fading interference. This is the transmit power setting, used to dynamically adjust the signal strength to balance communication distance and power consumption. It offers dual-mode communication options, including HPLC mode and HRF mode.

[0047] Each terminal node on the terminal side is equipped with a programmable impedance matching module and a Turbo code rate linkage module. The optimal action is sent from the edge side and directly applied to the hardware registers on the terminal side. The adjustment process can be completed in microseconds. The terminal node synchronously adjusts the hardware circuit parameters and encoding configuration by parsing the triple instruction shown in Equation (2), thereby achieving a precise mapping from the upper layer decision to the physical layer.

[0048] For example, the signal-to-noise ratio calculation formula on the terminal side is:

[0049] (3)

[0050] In the formula, The actual delivered power is expressed as shown in equation (4), where h is the channel gain. I represents noise power, and I represents interference power.

[0051] (4)

[0052] In the formula, For transmission power, The reflection coefficient is the amplitude of the reflection coefficient. Directly determines signal coupling efficiency, when The smaller the value, the lower the reflection loss. The expression for the reflection coefficient is shown in equation (5):

[0053] (5)

[0054] In the formula, For equivalent input impedance, , Equivalent resistance For equivalent reactance, This is the characteristic impedance of the channel.

[0055] The programmable impedance matching module minimizes the impedance by adjusting the equivalent input impedance to make it as close as possible to the channel characteristic impedance. This maximizes the signal-to-noise ratio at a given transmit power, thereby improving the communication's anti-interference capability.

[0056] System effective throughput Defined as:

[0057] (6)

[0058] In the formula, Here, B represents the Turbo coding rate, and B represents the channel bandwidth. When the signal-to-noise ratio (SNR) is low, the terminal reduces the Turbo coding rate to increase redundant bits, thereby improving error correction capabilities and ensuring transmission success rate. Conversely, when the SNR is high, the terminal increases the Turbo coding rate to reduce redundancy overhead and improve the overall effective transmission rate. The adjustment of the Turbo coding rate is achieved through software parameter configuration, working synchronously with the hardware module. This dynamic adjustment, combined with the optimization of the programmable impedance matching module, forms a closed loop: improved channel quality increases the SNR, triggering a coding rate upgrade, ultimately maximizing throughput.

[0059] Step S3: After the terminal side completes the optimal action, it feeds back the communication result of the current time slot and the state vector of the next time slot to the edge side.

[0060] The terminal side completes data transmission on the selected communication mode and frequency point based on the optimal action issued, and obtains the actual link performance of the current time slot, such as signal-to-noise ratio, effective throughput, bit error rate, service success rate, AoI, and energy consumption. After the current time slot ends, the terminal side sends these statistical results back to the edge side to form the state vector of the next time slot and the instant reward of the current time slot, thereby explicitly incorporating the execution results of the terminal side into the optimization target of the edge side and completing the local closed loop of "decision-execution-feedback". For example, the communication results in this embodiment include information age, energy consumption, and signal-to-noise ratio. The terminal side feeds back the signal-to-noise ratio, AoI, and energy consumption calculated after executing the optimal action to the edge side, and constructs the state vector s of the next time slot according to equation (1) after executing the optimal action. t+1 And send it to the edge side.

[0061] Step S4: The edge side calculates the instant reward of the current time slot based on the communication results of the current time slot, and calculates the target Q value of the DDQN model based on the instant reward of the current time slot and the state vector of the next time slot.

[0062] Specifically, the formula for calculating instant rewards is as follows:

[0063] (7)

[0064] In the formula, The signal-to-noise ratio of the current time slot. The information age of the current time slot, The energy consumption for the current time slot. , and All weights are adjustable. By dynamically adjusting the weights in different scenarios, throughput can be prioritized in data-intensive scenarios, information latency can be reduced in scenarios with high real-time requirements, and power consumption can be suppressed in energy-constrained scenarios.

[0065] The target network of the DDQN model is used to calculate the target Q-value, avoiding training instability caused by self-feedback in the current network. The formula for calculating the target Q-value is:

[0066] (8)

[0067] In the formula, The target Q value for the current time slot, The immediate reward for the current time slot. As a discount factor, For the current network parameters, For the network parameters of the target network, For the action selection section, it indicates that the network parameters of the current network are... At that time, the state vector for the next time slot Choose the action that maximizes the Q value. , This is the part for calculating the Q-value of the target network, representing the network parameters of the target network. At that time, based on the state vector of the next time slot and actions The calculated Q-value of the target network. The DDQN model effectively mitigates estimation bias and makes the training process more stable by separating action selection and action evaluation through this "dual network" approach.

[0068] Next, the predicted Q-value of the current network will be... Substitute the target Q-value of the target network into the loss function shown in Equation (9), and calculate the gradient of the loss function with respect to the network parameters of the current network through backpropagation, and update the network parameters. At the same time, periodically copy the network parameters of the current network to the target network to maintain the stability of the target network.

[0069] (9)

[0070] In the formula, D represents the experience replay pool. N is the total number of samples, and E[·] represents the mean squared error function. The introduction of the empirical replay pool makes the training samples more uniform in time, reducing the instability caused by data correlation.

[0071] Furthermore, in practical deployments, edge nodes are limited by the limited computing power of embedded devices, requiring the adoption of lightweight neural network structures to reduce computational complexity. For example, methods such as network pruning to remove redundant connections, quantization to compress model parameter accuracy, and distillation to transfer knowledge to small-scale models are used to ensure that the inference process is completed within milliseconds, avoiding decision delivery from becoming a system latency bottleneck. At the same time, the current network of the DDQN model uses an ε-greedy strategy to dynamically balance and utilize the optimal action, where the ε value decays with training time to gradually converge to the optimal strategy. The target network maintains its difference from the current online network through soft updates (such as weighted averaging of the parameters of the current network and the target network) or periodic hard updates (such as completely copying the parameters of the online network every C steps). This decoupled design makes action selection and value evaluation independent of each other, improving the convergence stability and decision robustness of the DDQN model in scenarios such as dynamic channel selection and energy consumption optimization.

[0072] Step S5: When the preset update time is reached, the edge side selects samples from the experience replay pool to calculate the output probability distribution of the current network and uploads the output probability distribution of the current network to the cloud.

[0073] In large-scale power IoT scenarios, directly uploading raw collected data not only consumes significant communication bandwidth but also poses data privacy risks. Therefore, this embodiment proposes a cloud optimization scheme based on federated distillation. The cloud input comes from the output probability distribution of the current network in the DDQN model of each edge node, rather than uploading raw data. Specifically, the output probability distribution of the current network of the i-th edge node... The expression is:

[0074] (10)

[0075] In the formula, is the original score vector logits output by the last layer of the DDQN model, which is the Q value of all actions, T is the distillation temperature parameter, and softmax(·) is the normalization function.

[0076] To further enhance privacy protection, a noise injection method based on differential privacy can be introduced when uploading the current network's output probability distribution at edge nodes. For example, random noise following a Gaussian distribution can be added to the current network's output probability distribution to generate a perturbed output probability distribution. ,in This indicates that the mean is zero and the variance is... The Gaussian noise is used. This ensures that even if an attacker obtains the uploaded data, they cannot reconstruct the original training samples, thus meeting the stringent privacy compliance requirements of the power IoT scenario while guaranteeing the model's collaborative optimization performance.

[0077] Step S6: The cloud generates the output probability distribution of the global model based on the current network's output probability distribution, updates the global model parameters based on the current network's output probability distribution and the global model's output probability distribution, and sends the global model parameters to the edge side.

[0078] The expression for the output probability distribution of the global model is as follows:

[0079] (11)

[0080] In the formula, The original score vector logits is the output of the last layer of the global model, which represents the global model parameters.

[0081] The expression for the distillation loss function in the cloud is:

[0082] (12)

[0083] In the formula, This represents the distillation loss value. This represents the weight of each edge node on the edge side. KL(·) is the Kullback-Leibler divergence, used to measure the difference between the edge node distribution and the global distribution. This represents the current network output probability distribution for the i-th edge node. This represents the output probability distribution of the global model. By minimizing this loss function, the cloud can integrate knowledge from multiple edge nodes, such as differences in the operating modes of power equipment in different regions or differences in user electricity consumption habits, thereby generating a global model that takes into account the characteristics of each edge node, resulting in stronger generalization ability. The update process of the global model parameters can be represented as:

[0084] (13)

[0085] In the formula, These are global model parameters. For parameters The cloud-based global model output function takes x as the reference input for distillation (such as a typical state vector / feature sample) and outputs a logits vector (unnormalized score). The prediction probability distribution of the global model is obtained through softmax(· / T) mapping, which is also the output probability distribution of the global model. T is the distillation temperature coefficient; is the weight coefficient of the i-th edge node; KL(·‖·) represents the Kullback-Leibler divergence.

[0086] Step S7: The edge side updates the network parameters of the DDQN model based on the global model parameters.

[0087] When communication resources are limited, the model synchronization cycle between the cloud and the edge must be planned reasonably. Too frequent synchronization will waste bandwidth, while too long synchronization intervals will cause the edge model to deviate from the global optimum. This embodiment adopts a dynamic scheduling mechanism to adjust the distribution frequency and bandwidth budget according to the network conditions and task load of each edge node, ensuring that the system achieves optimal global evolution under limited resources.

[0088] Specifically, after obtaining new global model parameters through gradient updates, the cloud periodically distributes them to each edge node. This allows for the direct replacement or fine-tuning of the network parameters θ in the current DDQN model, thereby altering the network's decision output for the same state vector and correcting edge decision-making behavior over the long term. Simultaneously, edge nodes can quickly adapt to their local environment and share global knowledge from other edge nodes. Thus, through bidirectional interaction between the edge and the cloud, adaptive and continuous evolution capabilities are achieved.

[0089] This embodiment of the cloud-edge collaborative HPLC dual-mode communication method employs a DDQN model at the edge, with a reward function that integrates signal-to-noise ratio, information timeliness, and energy consumption. This allows the system to dynamically select frequency hopping, transmit power, and communication mode even under strong electromagnetic disturbances such as lightning strikes and high-current switching, improving anti-interference capabilities and ensuring high communication success rates and low bit error rates. The decision-making process is decentralized to the edge nodes, avoiding data backhaul delays and reducing response latency to milliseconds, meeting the real-time requirements of the power IoT. The cloud integrates edge node model knowledge through federated distillation, enabling continuous global model evolution without uploading raw data, ensuring overall optimality and improving cross-regional scenario generalization capabilities. The terminal side employs a programmable impedance matching module and a Turbo code rate linkage module. The system forms a closed loop of "perception-decision-execution," enhancing link stability and resource utilization. The reward function incorporates AoI (Aspect-of-Information) and energy consumption constraints, prioritizing the real-time nature of critical information and reducing energy consumption for the same workload. The system can dynamically and adaptively adjust based on power line channel impedance, interference intensity, and service load, avoiding communication islands and data acquisition failures during sudden interference or spectrum congestion. Ultimately, through a long-term closed loop of "edge strategy output uplink—cloud fusion optimization—global model downlink—edge decision update," it maintains optimal overall performance even under cross-regional deployment and changing interference environments. This successfully overcomes the shortcomings of existing technologies, such as unstable communication quality, high response latency, and insufficient energy efficiency, providing a solid technical guarantee for high-reliability, low-latency, and low-power communication in the power Internet of Things and smart grids.

[0090] Example 2

[0091] like Figure 1 As shown, this embodiment provides a cloud-edge collaborative HPLC dual-mode communication system, including a terminal side, an edge side, and a cloud.

[0092] Each terminal node on the terminal side includes a data acquisition module and an action feedback module. The data acquisition module is used to generate the state vector of the current time slot based on the physical layer indicators collected in real time. The expression of the state vector is shown in Equation (1). Among them, the physical layer indicators include key parameters such as channel gain, interference power, buffer queue length, remaining energy state, delay spread characteristics, information age (AoI), and energy consumption.

[0093] The action feedback module is used to feed back the communication result of the current time slot and the state vector of the next time slot to the edge side after executing the optimal action. The expression of the optimal action is shown in Equation (2). Each terminal node is equipped with a programmable impedance matching module and a Turbo code rate linkage module. The action feedback module parses the triple instruction shown in Equation (2) and synchronously adjusts the hardware module parameters and encoding configuration to achieve a precise mapping from the upper layer decision to the physical layer. The specific adjustment process can be referred to in Embodiment 1, which will not be elaborated here. In this embodiment, the action feedback module feeds back the signal-to-noise ratio, AoI and energy consumption calculated after executing the optimal action to the edge side, and constructs the state vector s of the next time slot according to Equation (1) after executing the optimal action. t+1 And send it to the edge side.

[0094] Each edge node on the edge side includes an action selection module, a Q-value calculation module, and a network parameter update module. The action selection module generates the optimal action for the current time slot based on the state vector of the current time slot and the DDQN model, and sends the optimal action to the action feedback module on the terminal side. The Q-value calculation module calculates the instant reward for the current time slot based on the communication results of the current time slot, and calculates the target Q-value of the DDQN model based on the instant reward of the current time slot and the state vector of the next time slot. The instant reward is calculated using Equation (7), and the target Q-value is calculated using Equation (8). The network parameter update module selects samples from the experience replay pool to calculate the output probability distribution of the current network when the preset update time is reached, uploads the output probability distribution of the current network to the cloud, and updates the network parameters of the DDQN model based on the global model parameters issued by the cloud. The network parameter update module is also used to update the network parameters of the current network based on the loss function shown in Equation (9).

[0095] The cloud is used to generate the output probability distribution of the global model based on the current network's output probability distribution, and to update the global model parameters based on the current network's output probability distribution and the global model's output probability distribution, and then distribute the global model parameters to the edge side. The expression for the current network's output probability distribution is shown in Equation (10), the expression for the global model's output probability distribution is shown in Equation (11), and the update process for the global model parameters is shown in Equation (13).

[0096] Furthermore, the system performance metrics may include average information timeliness as shown in equation (14). The average throughput (Throughput) shown in Equation (15) and the packet loss rate (PLR) shown in Equation (16) can not only be used to evaluate the system's performance, but also be incorporated into the reward function and cloud optimization objectives to keep the system's performance objectives consistent with actual performance metrics.

[0097] (14)

[0098] In the formula, This indicates the information age of the current time slot, where T is the preset time length.

[0099] (15)

[0100] In the formula, N represents the effective data volume, where N is the number of effective data points within the time period T.

[0101] (16)

[0102] In the formula, The number of packets successfully sent. This represents the total number of packets sent.

[0103] This embodiment of the cloud-edge collaborative HPLC dual-mode communication system employs a hierarchical design to achieve collaboration between global optimization and local decision-making: the cloud is responsible for global optimization and continuous model evolution, edge nodes complete local optimal decisions on a millisecond timescale, and the terminal is responsible for closed-loop execution at the physical layer. Unlike traditional fragmented hierarchical designs, this embodiment constructs a closed-loop system through the tight coupling of data flow, decision flow, and feedback flow: the physical layer indicators measured in real-time by the terminal directly constitute the state input of the edge-side DDQN model; the action commands output by the edge side (such as frequency hopping selection, transmit power adjustment, and communication mode switching) serve as control signals for the terminal's programmable impedance matching module and Turbo code rate linkage module; the actual effect feedback after terminal execution flows back to the edge nodes, forming the reward signal for reinforcement learning and the input for the next state. Simultaneously, the output probability distribution generated by the edge nodes during local training is uploaded to the cloud at preset intervals. The cloud, through federated distillation technology, integrates the knowledge of multiple edge nodes, generates global model parameters, and sends them back to the edge nodes, thereby dynamically adjusting subsequent edge decision-making behavior. Through this closed-loop interaction of top-down global optimization and bottom-up local feedback, the system can still achieve highly reliable, low-latency, and low-power dual-mode communication under complex conditions such as strong electromagnetic interference, sudden events, and cross-regional deployment.

[0104] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0105] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A dual-mode communication method for HPLC based on cloud-edge collaboration, characterized in that, The method includes: Step S1: The terminal generates the state vector of the current time slot based on the real-time collected physical layer indicators; Step S2: The edge side generates the optimal action for the current time slot based on the state vector and DDQN model of the current time slot, and sends the optimal action to the terminal side; Step S3: After the terminal side completes the optimal action, it feeds back the communication result of the current time slot and the state vector of the next time slot to the edge side. Step S4: The edge side calculates the instant reward of the current time slot based on the communication result of the current time slot, and calculates the target Q value of the DDQN model based on the instant reward of the current time slot and the state vector of the next time slot. Step S5: When the preset update time is reached, the edge side selects samples from the experience replay pool to calculate the output probability distribution of the current network, and uploads the output probability distribution of the current network to the cloud. Step S6: The cloud generates the output probability distribution of the global model based on the output probability distribution of the current network, updates the global model parameters based on the output probability distribution of the current network and the output probability distribution of the global model, and sends the global model parameters to the edge side. Step S7: The edge side updates the network parameters of the DDQN model based on the global model parameters.

2. The HPLC dual-mode communication method based on cloud-edge collaboration according to claim 1, characterized in that, The state vector includes channel gain, interference power, buffer queue length, remaining energy state, and delay spread characteristics.

3. The HPLC dual-mode communication method based on cloud-edge collaboration according to claim 1, characterized in that, The optimal dynamic is a triplet that includes a frequency modulation channel index, a transmit power level, and a dual-mode communication standard selection.

4. The HPLC dual-mode communication method based on cloud-edge collaboration according to any one of claims 1-3, characterized in that, The communication results include information age, energy consumption, and signal-to-noise ratio.

5. The HPLC dual-mode communication method based on cloud-edge collaboration according to claim 4, characterized in that, The formula for calculating the instant reward is as follows: In the formula, The signal-to-noise ratio for the current time slot, The information age in the current time slot. The energy consumption for the current time slot, , and All weights are adjustable.

6. The HPLC dual-mode communication method based on cloud-edge collaboration according to claim 5, characterized in that, The formula for calculating the signal-to-noise ratio is: In the formula, Indicates the actual delivered power. For channel gain, Let I be the noise power and I be the interference power; wherein, the expression for the actual delivered power is: In the formula, For transmission power, This is the reflection coefficient.

7. The HPLC dual-mode communication method based on cloud-edge collaboration according to claim 1, characterized in that, The expression for the distillation loss function in the cloud is: In the formula, This represents the distillation loss value. This represents the weight of each edge node in the edge side, and KL(·) is the Kullback-Leibler divergence. This represents the output probability distribution of the current network at the i-th edge node. This represents the output probability distribution of the global model.

8. The HPLC dual-mode communication method based on cloud-edge collaboration according to claim 1, characterized in that, The output probability distribution of the current network is the output probability distribution after introducing noise.

9. A cloud-edge collaborative HPLC dual-mode communication system, characterized in that, The system includes a terminal side, an edge side, and a cloud. Each terminal node on the terminal side includes a data acquisition module and an action feedback module. The data acquisition module is used to generate a state vector for the current time slot based on the physical layer indicators collected in real time. The action feedback module is used to feed back the communication result of the current time slot and the state vector of the next time slot to the edge side after executing the optimal action. Each edge node in the edge side includes an action selection module, a Q-value calculation module, and a network parameter update module; the action selection module is used to generate the optimal action for the current time slot based on the state vector of the current time slot and the DDQN model, and send the optimal action to the terminal side; The Q-value calculation module is used to calculate the instant reward of the current time slot based on the communication result of the current time slot, and to calculate the target Q-value of the DDQN model based on the instant reward of the current time slot and the state vector of the next time slot. The network parameter update module is used to select samples from the experience replay pool to calculate the output probability distribution of the current network when the preset update time is reached, and upload the output probability distribution of the current network to the cloud, and update the network parameters of the DDQN model based on the global model parameters issued by the cloud. The cloud is used to generate the output probability distribution of the global model based on the output probability distribution of the current network, update the global model parameters based on the output probability distribution of the current network and the output probability distribution of the global model, and send the global model parameters to the edge side.

10. The HPLC dual-mode communication system based on cloud-edge collaboration according to claim 9, characterized in that, Each of the terminal nodes is equipped with a programmable impedance matching module and a Turbo code rate linkage module.

Citation Information

Patent Citations

  • HPLC and HRF dual-mode communication hybrid data transmission control method

    CN120263228A