Heterogeneous communication network relay selection optimization method, system, device and medium

By introducing a learning-based inverse selection mechanism and a state-action value function into a heterogeneous communication network, the relay node selection is optimized, solving the communication response delay problem caused by unreasonable relay node selection in existing technologies, and realizing efficient and reliable data transmission of distributed resources.

CN116456420BActive Publication Date: 2026-07-21CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD
Filing Date
2023-04-28
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing relay selection methods in heterogeneous communication networks lack an optimal relay node selection strategy during communication between multiple source nodes and destination nodes. This makes it difficult for some distributed resource source nodes to obtain communication access responses in real time, especially when a large number of distributed resource source nodes make simultaneous requests, making it impossible to achieve the optimal communication strategy.

Method used

A relay selection optimization method for heterogeneous communication networks is adopted. By judging the HPLC channel state, relay nodes are selected for data transmission. When a conflict occurs, a learning inverse selection mechanism is introduced to optimize the selection of relay nodes. The goal is to minimize the delay and bit error rate. The optimization is carried out by action selection and state-action value function.

Benefits of technology

It improves the reliability of data transmission and reduces network latency, effectively avoids conflicts in relay node selection, and ensures the real-time performance and reliability of distributed resource collaboration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116456420B_ABST
    Figure CN116456420B_ABST
Patent Text Reader

Abstract

The application discloses a relay selection optimization method, system, device and medium in a heterogeneous communication network, judges the HPLC channel state between a source node and a destination node, selects HPLC for data transmission if the HPLC channel state estimation value meets the requirement, and enters the next time slot; otherwise, each source node selects a relay node for data transmission according to the current time slot service data volume and the state action value function of the source node; whether the conflict situation that different source nodes select the same relay node occurs is judged, if yes, the counter-selection of the relay node to the source node is carried out based on a learning counter-selection mechanism, other source nodes reselect a relay node for service data transmission according to the state action value function of the source node; and whether the conflict situation that different source nodes select the same relay node occurs is continuously judged; if no conflict occurs, the state action value function is updated. The application considers the optimal selection of the relay node, increases the reliability of data transmission, and reduces the network delay.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of power grid operation and dispatching technology, specifically relating to a method, system, equipment, and medium for optimizing relay selection in heterogeneous communication networks. Background Technology

[0002] On-demand access and flexible regulation of demand-side distributed resources such as distributed generation, adjustable loads, and user-side energy storage are effective means to improve the balance and regulation capabilities of new power systems and alleviate the power deficit caused by high renewable energy ratios on large power grids. Specifically, the aggregation and control gateway, based on information such as electrical quantities, micro-meteorological data, and operating conditions of distributed resource equipment uploaded by distributed resource communication terminals, regulates distributed resources to participate in grid peak shaving, frequency regulation, and demand response, thereby improving the safe and stable operation of the power system. Therefore, the stability of the network and the real-time communication between distributed resources and the aggregation and control gateway are crucial for the smooth implementation of various interactive functions such as peak shaving, frequency regulation, and demand response.

[0003] Chinese patent publication number CN111866985A, entitled "Hybrid Relay Selection Method for Dense Communication Networks," includes: using a group of switching nodes as source and destination nodes; excluding relay nodes in unfavorable positions by setting and adjusting timer parameters; and distributing the channel state information monitoring burden from source nodes to various relay nodes to reduce the burden on these monitoring nodes; and achieving a dynamic balance between system stability, latency, and signaling burden by adjusting the timer parameters. However, this invention lacks a strategy for relay nodes to counter source nodes when source node selection conflicts with relay nodes, leading to situations where some distributed resource source nodes struggle to obtain real-time communication access responses.

[0004] Existing technical solutions, designed for resource allocation scenarios involving relay nodes during communication between multiple source and destination nodes, lack consideration for the impact of relay node selection on overall communication quality, thus failing to achieve optimal communication strategies. Furthermore, when a large number of distributed resource source nodes simultaneously request relay nodes, the relay nodes cannot reverse-select source nodes based on actual factors such as service priorities, resulting in some distributed resource source nodes failing to receive real-time communication access responses. Summary of the Invention

[0005] To achieve fast and reliable transmission of distributed resource collaborative interaction business data, this invention proposes a method, system, device, and medium for optimizing relay selection in heterogeneous communication networks. This invention considers the optimal selection of relay nodes, increases the reliability of data transmission, and reduces network latency.

[0006] To achieve the above objectives, the present invention employs the following technical solution:

[0007] A relay selection optimization method for heterogeneous communication networks includes:

[0008] The HPLC channel state between the source node and the destination node is determined. If the HPLC channel state estimate meets the requirements, HPLC is selected for data transmission and the process moves to the next time slot. If the channel state estimate does not meet the requirements, each source node selects a relay node for data transmission based on the current time slot's service data volume and its own state action value function.

[0009] The system determines whether there is a conflict where different source nodes select the same relay node. If a conflict occurs, the relay node performs a reverse selection of the source node based on the learning reverse selection mechanism. Other source nodes reselect relay nodes for business data transmission according to their own state action value function. The system then continues to determine whether there is a conflict where different source nodes select the same relay node. If no conflict occurs, the system proceeds to the next time slot and updates the state action value function according to the business data volume of the next time slot.

[0010] Determine if the current time slot has reached the maximum time slot. If it has, end the process; otherwise, return to the step of determining the HPLC channel status between the source node and the destination node.

[0011] As a further improvement of the present invention, each source node selects a relay node for data transmission based on the current time slot service data volume and its own state action value function, including:

[0012] Each source node queries the action based on the current time slot service data volume and its own status action value function table. Select the action with the highest state-action value function, and calculate the action reward for the current time slot based on the selected action.

[0013] As a further improvement of the present invention, the method for selecting the action with the highest state-action value function is as follows:

[0014]

[0015] In the formula, S i (t) represents the source node b i The state space, source node b i The action space is X i =[χ1,K,χ j ,K,χ J ], χ i (t)∈X represents the source node b in the t-th time slot. i The selected action is to select the j-th relay node for business data transmission; For source node b i The state-action value function has the following properties:

[0016]

[0017] Where β is the learning efficiency and γ is the discount factor.

[0018] As a further improvement of the present invention, the step of performing relay node deselection of source nodes based on the learning deselection mechanism, and other source nodes reselecting relay nodes for business data transmission according to their own state-action value functions, includes:

[0019] When different source nodes choose the same action, a learning-based inversion mechanism is introduced; in this mechanism, as of the current time slot, source node b... i Select relay node g j If the average reward is empirical performance, then the source node b i The current experience learning performance is represented by source node b. i The priority is multiplied by the empirical performance, and the specific formula is:

[0020]

[0021] When different source nodes select the same relay node, the relay node is deselected. The source node with the larger value transmits the business data. Other source nodes select relay nodes for business data transmission based on their suboptimal state action value function. If selection conflicts still exist, the selection of relay nodes continues according to the learning reverse selection mechanism until all source nodes have selected relay nodes for data transmission.

[0022] As a further improvement of the present invention, each source node selects a relay node for data transmission based on the current time slot service data volume and its own state action value function, including:

[0023] Multi-service data transmission employs a latency model and a bit error rate model. A model function for selecting relay nodes for data transmission is constructed with the objective of minimizing the weighted average of latency and bit error rate. The model function is as follows:

[0024] P1:

[0025]

[0026]

[0027] Among them, L i Indicates source node b i priority, D i (t) represents the time delay, P i (t) represents the bit error rate, α represents the weighting parameter, and C1 represents the limit that the total latency of business data transmission for each distributed resource node cannot exceed its limit. The total bit error rate of business data transmission on each distributed resource node must not exceed its limit value P. i th (t); C2 indicates that each source node selects only one relay node in each round;

[0028] The selection of relay nodes for source node service data transmission is optimized using an action selection approach; the state space represents the amount of service data that each source node needs to transmit, and source node b... i The state space is denoted as S i (t); Source node b i The action space is X i =[χ1,K,χ j ,K,χ J ], χ i (t)∈X represents the source node b in the t-th time slot. i The selected action is to select the j-th relay node for business data transmission;

[0029] and source node b i The reward for selecting a relay node to transmit service data is converted into a weighted sum of latency and bit error rate, and expressed as:

[0030] r i [S i (t),χ i [(t)]=L i D i (t)+αP i (t)

[0031] Among them, source node b i The state-action value function is And there are:

[0032]

[0033] Where β is the learning efficiency; γ is the discount factor, representing the degree of influence of state rewards;

[0034] Each source node selects a relay node based on its own optimal state action value function.

[0035] As a further improvement to the present invention, the bit error rate model is specifically as follows:

[0036] Source node b i The total bit error rate P of data transmitted to destination node a i (t) is represented as:

[0037]

[0038] In the formula, x i,j(t) is a binary indicator variable, and J is the number of relay nodes;

[0039] P j,a (t) represents the relay node g j The bit error rate of data transmitted to destination node a

[0040]

[0041] Among them, C j,a (t) represents the relay node g j Data transfer rate to destination node a; ω j,a These represent the channel bandwidth radiated during the data transmission of services;

[0042] In the formula, P i,j (t) represents the source node b i To relay node g j The data transmission error rate;

[0043]

[0044] Where SF is the spreading factor, H(g) is the harmonic number, and Q(g) is the tail integral function of the standard Gaussian distribution, r i,j For source node b i To relay node g j Signal-to-noise ratio of transmitted data.

[0045] As a further improvement to the present invention, the time delay model is specifically as follows:

[0046] The latency model is used to represent the end-to-end transmission latency of distributed resource data, including the communication latency from the source node to the relay node and the communication latency from the relay node to the destination node. The total latency D for service data transmission is... i Represented as:

[0047]

[0048] In the formula, D i,j (t) represents the source node b i To relay node g j Data transmission latency, λ i It is a constant; source node b i To relay node g j Data transmission rate C i,j (t) is represented as:

[0049]

[0050] Where N0 represents the channel white noise power; V i,j and ω i,jThese represent the electromagnetic interference radiated during service data transmission using HRF and the channel bandwidth, respectively; p i H represents the source node's signal transmission power; i,j (t) represents source node b within time slot t. i To relay node g j Channel gain when transmitting data;

[0051] D j,a (t) represents the relay node g j Data transmission delay to destination node a,

[0052] relay node g j Data transfer rate C to destination node a j,a (t) is:

[0053]

[0054] Among them, V j,a and ω j,a These represent the electromagnetic interference radiated during service data transmission using HRF and the channel bandwidth, respectively; p j H represents the source node's signal transmission power; j,a (t) represents source node b within time slot t. j To relay node g a Channel gain when transmitting data.

[0055] As a further improvement of the present invention, the channel state adopts a quasi-static time slot model, with T equal-length time slots. The channel state information remains unchanged within one time slot and changes in different time slots. Within each time slot, each relay node can only be selected by one source node.

[0056] A relay selection optimization system for heterogeneous communication networks includes:

[0057] The channel state determination module is used to determine the HPLC channel state between the source node and the destination node. If the HPLC channel state estimation meets the requirements, HPLC is selected for data transmission and the process moves to the next time slot. If the channel state estimation does not meet the requirements, each source node selects a relay node for data transmission based on the current time slot's service data volume and its own state action value function.

[0058] The relay node judgment module is used to determine whether there is a conflict in which different source nodes select the same relay node. If a conflict occurs, the relay node will reverse the selection of the source node based on the learning reverse selection mechanism. Other source nodes will reselect the relay node for business data transmission according to their own state action value function. The module will continue to determine whether there is a conflict in which different source nodes select the same relay node. If no conflict occurs, the module will proceed to the next time slot and update the state action value function according to the business data volume of the next time slot.

[0059] The maximum time slot determination module is used to determine whether the current time slot has reached the maximum time slot. If it has, the process ends; if it has not reached the maximum time slot, the process returns to the step of determining the HPLC channel status between the source node and the destination node.

[0060] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the relay selection optimization method for the heterogeneous communication network.

[0061] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the relay selection optimization method for heterogeneous communication networks.

[0062] Compared with the prior art, the present invention has the following beneficial effects:

[0063] This invention proposes an optimization method for relay selection in heterogeneous communication networks that supports collaborative interaction of distributed resources. It employs a local multimodal, multi-relay node heterogeneous network for transmission, introducing relay nodes to forward various types of service data. Based on this, with the optimization objective of minimizing the weighted average of latency and bit error rate, a learning-based inverse selection mechanism is used to consider the optimal selection of relay nodes, increasing data transmission reliability and reducing network latency. This invention introduces a learning-based inverse selection mechanism for relay nodes to select source nodes. This addresses the problem in traditional channel selection methods where multiple source nodes select the same relay node for data transmission, leading to difficulties for some distributed resource source nodes in obtaining real-time communication access responses. When different source nodes select the same relay node, the learning-based inverse selection mechanism is triggered, and the relay node selection is repeated until all source nodes have selected their service data, effectively avoiding relay node selection conflicts. Attached Figure Description

[0064] Figure 1 This is a flowchart of a relay selection optimization method for heterogeneous communication networks according to the present invention;

[0065] Figure 2 This is a system architecture diagram provided in an embodiment of the present invention;

[0066] Figure 3 This is a flowchart of a method for optimizing relay selection in heterogeneous communication networks that supports distributed resource collaborative interaction, as provided in an embodiment of the present invention.

[0067] Figure 4 This is a comparison of the total latency of business data transmission given in the embodiments of the present invention;

[0068] Figure 5 This is a comparison of the total bit error rate of business data transmission given in the embodiments of the present invention;

[0069] Figure 6 A schematic diagram of the relay selection optimization system for heterogeneous communication networks provided by the present invention;

[0070] Figure 7 This is a schematic diagram of an electronic device provided by the present invention. Detailed Implementation

[0071] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0072] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0073] Currently, distributed resource communication terminals (source nodes) and aggregation and control gateways (destination nodes) mostly employ multimodal heterogeneous networking for communication. When the HPLC link quality between the source and destination nodes is poor, HRF or LoRA communication can be switched to send source node data to potential relay nodes. Relay node optimization is then performed based on the HRF / LoRA channel quality between the source and relay nodes, and the HPLC channel quality between the relay and destination nodes, thereby improving the reliability and real-time performance of data transmission. However, local multimodal heterogeneous networking optimization supporting distributed resource collaborative interaction still faces the following challenges:

[0074] On the one hand, due to electromagnetic interference radiated during the operation of distributed resource devices, the channel state information between source nodes and relay nodes, as well as between relay nodes and destination nodes, is time-varying and difficult for source nodes and relay nodes to accurately obtain. How to optimize relay nodes when channel state information is incompletely available is a pressing issue. On the other hand, due to the limited transmission capacity of relay nodes and channels, they cannot simultaneously receive and forward data from multiple source nodes, leading to conflicts in relay node selection schemes among multiple source nodes. Therefore, it is necessary to design a relay node selection conflict mitigation mechanism based on the quality of service requirements of each source node's interactive functions, thereby improving the reliability and real-time performance of system data transmission.

[0075] like Figure 1 As shown, the first objective of this invention is to provide a relay selection optimization method for heterogeneous communication networks, comprising:

[0076] S1, determine the HPLC channel status between the source node and the destination node. If the HPLC channel status estimate meets the requirements, HPLC is selected for data transmission and the process moves to the next time slot. If the channel status estimate does not meet the requirements, each source node selects a relay node for data transmission based on the current time slot's service data volume and its own status action value function.

[0077] S2, determine whether there is a conflict where different source nodes select the same relay node. If a conflict occurs, the relay node will reverse the selection of the source node based on the learning reverse selection mechanism. Other source nodes will reselect the relay node for business data transmission according to their own state action value function. Continue to determine whether there is a conflict where different source nodes select the same relay node. If no conflict occurs, proceed to the next time slot and update the state action value function according to the business data volume of the next time slot.

[0078] S3, determine whether the current time slot has reached the maximum time slot. If it has, end; if it has not reached the maximum time slot, return to the step of determining the HPLC channel status between the source node and the destination node.

[0079] Furthermore, in this embodiment of the invention, the weighted value of minimizing latency and bit error rate is used to select all available relay node transmission schemes for each source node. By analyzing the impact of relay node selection on the communication access of distributed resource source nodes, the optimal relay selection scheme is obtained, ensuring real-time collaborative interaction between distributed resources and the power grid.

[0080] In step S1, each source node selects a relay node for data transmission based on the current time slot service data volume and its own state action value function, including:

[0081] Each source node queries the action based on the current time slot service data volume and its own status action value function table. Select the action with the highest state-action value function, and calculate the action reward for the current time slot based on the selected action.

[0082] As a further improvement of the present invention, the method for selecting the action with the highest state-action value function is as follows:

[0083]

[0084] In the formula, S i(t) For source node b i The state space, source node b i The action space is X i =[χ1,K,χ j ,K,χ J ], χ i (t)∈X represents the source node b in the t-th time slot. i The selected action is to select the j-th relay node for business data transmission; For source node b i The state-action value function has the following properties:

[0085]

[0086] Where β is the learning efficiency and γ is the discount factor.

[0087] In step S1, each source node selects a relay node for data transmission based on the current time slot service data volume and its own state action value function, including:

[0088] Multi-service data transmission employs a latency model and a bit error rate model. A model function for selecting relay nodes for data transmission is constructed with the objective of minimizing the weighted average of latency and bit error rate. The model function is as follows:

[0089] P1:

[0090]

[0091]

[0092] Among them, L i Indicates source node b i priority, D i (t) represents the time delay, P i (t) represents the bit error rate, α represents the weighting parameter, and C1 represents the limit that the total latency of business data transmission for each distributed resource node cannot exceed its limit. The total bit error rate of business data transmission on each distributed resource node must not exceed its limit value P. i th (t); C2 indicates that each source node selects only one relay node in each round;

[0093] The selection of relay nodes for source node service data transmission is optimized using an action selection approach; the state space represents the amount of service data that each source node needs to transmit, and source node b... i The state space is denoted as S i (t); Source node b i The action space is X i =[χ1,K,χ j ,K,χ J ], χ i (t)∈X represents the source node b in the t-th time slot. i The selected action is to select the j-th relay node for business data transmission;

[0094] and source node b i The reward for selecting a relay node to transmit service data is converted into a weighted sum of latency and bit error rate, and expressed as:

[0095] r i [S i (t),χ i [(t)]=L i D i (t)+αP i (t)

[0096] Among them, source node b i The state-action value function is And there are:

[0097]

[0098] Where β is the learning efficiency; γ is the discount factor, representing the degree of influence of state rewards;

[0099] Each source node selects a relay node based on its own optimal state action value function.

[0100] The bit error rate model is as follows:

[0101] Source node b i The total bit error rate P of data transmitted to destination node a i (t) is represented as:

[0102]

[0103] In the formula, x i,j (t) is a binary indicator variable, and J is the number of relay nodes.

[0104] The delay model is as follows:

[0105] The latency model is used to represent the end-to-end transmission latency of distributed resource data, including the communication latency from the source node to the relay node and the communication latency from the relay node to the destination node. The total latency D for service data transmission is... i Represented as:

[0106]

[0107] In the formula, D i,j (t) represents the source node b i To relay node g j Data transmission latency.

[0108] The channel state proposed in this invention adopts a quasi-static time slot model, with T equal-length time slots. The channel state information remains unchanged within one time slot but changes in different time slots. Within each time slot, each relay node can only be selected by one source node.

[0109] As a specific embodiment, the above-mentioned relay node deselection of source nodes based on the learning deselection mechanism, and other source nodes reselecting relay nodes for business data transmission according to their own state-action value functions, includes:

[0110] When different source nodes choose the same action, a learning-based inversion mechanism is introduced; in this mechanism, as of the current time slot, source node b... i Select relay node g j If the average reward is empirical performance, then the source node b i The current experience learning performance is represented by source node b. i The priority is multiplied by the empirical performance;

[0111] When different source nodes select the same relay node, the relay node is deselected. The source node with the larger value transmits the business data. Other source nodes select relay nodes for business data transmission based on their suboptimal state action value function. If selection conflicts still exist, the selection of relay nodes continues according to the learning reverse selection mechanism until all source nodes have selected relay nodes for data transmission.

[0112] The method of the present invention will be described in detail below with reference to specific implementations and the following contents.

[0113] The specific methods for constructing the system model are as follows:

[0114] This invention considers scenarios where distributed resource aggregation and regulation participate in grid interaction. Numerous communication devices (source nodes) are deployed on distributed resources such as distributed power sources, controllable loads, and distributed energy storage to monitor and control the operational status of these distributed resources. Data is then uploaded in real-time to an edge aggregation and regulation gateway (destination node) for the formulation of distributed resource operation optimization decisions. Communication between source and destination nodes primarily employs heterogeneous networking methods such as HPLC, HRF, and LoRa.

[0115] Some relay nodes are equipped with HPLC+HRF dual-mode communication chips, supporting both HPLC and HRF communication methods simultaneously. Other relay nodes integrate a LoRa communication interface on top of an HPLC single-mode communication module, supporting both HPLC and LoRa communication methods simultaneously. When the HPLC link quality between the source and destination nodes is poor, the system can switch to HRF or LoRa communication to send data from the source node to potential relay nodes. Relay node selection is optimized based on the HRF or LoRa channel quality between the source and relay nodes, and the HPLC channel quality between the relay and destination nodes, thereby improving the reliability and real-time performance of data transmission.

[0116] Assume there are I source nodes and J relay nodes in this scenario, and the sets are defined as B = {b1, K, b}. i ,K b I} and G={g1,K,g j ,K g J The system architecture is as follows: Figure 2 As shown.

[0117] This invention employs a quasi-static time-slot model with T equal-length time slots, denoted as T = {1, K, t, K, T}. Channel state information remains constant within a single time slot but varies across different time slots. Within each time slot, each relay node can only be selected by one source node. Let source node b be defined. i The relay node is selected as a binary indicator variable x i,j (t), where x i,j (t) = 1 indicates that bi Choose g in the t-th time slot j Otherwise x i,j (t) = 0. The delay model and bit error rate model for multi-service data transmission are as follows.

[0118] (1) Delay Model

[0119] The end-to-end transmission latency of distributed resource data includes two parts: the communication latency from the source node to the relay node and the communication latency from the relay node to the destination node. Source node b i To relay node g j The data transmission rate can be expressed as:

[0120]

[0121] Where N0 represents the channel white noise power; V i,j and ω i,j These represent the electromagnetic interference radiated during service data transmission using HRF and the channel bandwidth, respectively; p i H represents the source node's signal transmission power; i,j (t) represents source node b within time slot t. i To relay node g j Channel gain when transmitting data.

[0122] If relay node g j If HRF communication is supported, then the source node b i To relay node g j The channel gain model can be modeled as follows:

[0123]

[0124] Where d i,j (t) represents the source node b i With relay node g j The communication distance between them is given by l, where l is the height of the relay node's location, and d0 is the given distance.

[0125] If relay node g j If LoRa communication is supported, then the source node b i To relay node g j The channel gain model can be modeled as follows:

[0126] H i,j (t) = 69.55 + 26.16lgf - 13.82lgh b +(44.9-6.55lgh b )·lgd i,j (t)-α(h m (3)

[0127] Where f is the communication frequency, h b h is the effective height of the transmitting antenna. m For the effective height of the receiving antenna, α(h) m ) is the effective height correction factor for the receiving antenna.

[0128] Therefore, source node b i To relay node g j The communication delay can be expressed as:

[0129]

[0130] Similarly, relay node g j The data transmission rate to destination node a is:

[0131]

[0132] Among them, the channel characteristics of HPLC are related to network topology, cable parameters, electrical load, and other parameters. Therefore, relay node g j The communication delay D to the destination node a j,a (t) can be expressed as:

[0133]

[0134] Therefore, the total latency D of business data transmission i It can be represented as:

[0135]

[0136] (2) Bit Error Rate Model

[0137] The bit error rate (BER) of information during transmission is related to the demodulation method and signal-to-noise ratio (SNR). Currently, HRF typically uses 16-QAM digital modulation, so this invention uses 16-QAM digital modulation as an example to construct the HRF BER model. However, the proposed scheme is also applicable to other digital modulation methods. (The source node b...) i To relay node g j The bit error rate of transmitted data is defined as:

[0138]

[0139] Where Q(g) is the tail integral function of the standard Gaussian distribution, r i,j For source node b i To relay node g j The signal-to-noise ratio of the transmitted data can be obtained by combining formulas (1) and (8) for the source node b under QAM digital modulation. i To relay node g j The expression for the data transmission error rate:

[0140]

[0141] If LoRa communication is used, source node b i To relay node g j The expression for the data transmission error rate is:

[0142]

[0143] Where SF is the spreading factor and H(g) is the harmonic number, by combining formulas (1) and (10), we can obtain the source node b under LoRa communication mode. i To relay node g j The expression for the data transmission error rate:

[0144]

[0145] Currently, HPLC typically employs QPSK digital modulation. Therefore, this invention uses QPSK digital modulation as an example to construct an HPLC bit error rate model. However, the proposed scheme is also applicable to other digital modulation methods. The relay node g... j The bit error rate of data transmitted to destination node a is defined as:

[0146]

[0147] Where erfc(g) is the complementary error function, r j,a For relay node g j The signal-to-noise ratio of the data transmitted to the destination node a. Combining equations (5) and (12), we can obtain the signal-to-noise ratio of the relay node g under QPSK digital modulation. j The expression for the data transmission error rate to destination node a:

[0148]

[0149] Then source node b i The total bit error rate P of data transmitted to destination node a i (t) can be expressed as:

[0150]

[0151] To achieve superior service transmission performance, this invention selects relay nodes for data transmission. By optimizing the relay node selection variable, the objective function is to minimize the weighted value of latency and bit error rate. The optimization problem can be modeled as follows:

[0152]

[0153] Among them, L i Indicates source node bi The priority is defined by α, which represents the weight parameter. C1 indicates that the total latency of business data transmission for each distributed resource node cannot exceed its limit. The total bit error rate of business data transmission on each distributed resource node must not exceed its limit value P. i th (t). C2 indicates that each source node selects only one relay node in each round.

[0154] This invention employs an action selection method to optimize the selection of relay nodes for source node service data transmission. The state space of this optimization algorithm is defined as the amount of service data that each source node needs to transmit, and source node b... i The state space is denoted as S i (t). Define source node b i The action space is X i =[χ1,K,χ j ,K,χ J ], χ i (t)∈X represents the source node b in the t-th time slot. i The selected action is to choose the j-th relay node for business data transmission.

[0155] Define source node b i The reward for selecting a relay node to transmit service data is the weighted sum of latency and bit error rate, which can be expressed as:

[0156] r i [S i (t),χ i [(t)]=L i D i (t)+αP i (t) (16)

[0157] Assume source node b i The state-action value function is And there are:

[0158]

[0159] Where β is the learning efficiency, the higher the learning efficiency, the faster the convergence; γ is the discount factor, representing the degree of influence of state rewards, the higher the discount factor, the greater the influence of state rewards. Each source node selects a relay node based on its own optimal state-action value function.

[0160] When different source nodes select the same action, a learning-based inversion mechanism is introduced. In this mechanism, source node b is defined up to the current time slot. i Select relay node g j If the average reward is empirical performance, then the source node bi The current performance of experience learning can be represented by the source node b. i The priority is multiplied by the empirical performance, that is:

[0161]

[0162] When different source nodes select the same relay node, the relay node is deselected. The source node with the larger value transmits the service data. At this time, other source nodes select relay nodes for service data transmission based on their suboptimal state action value function. If selection conflicts still exist, the selection of relay nodes continues according to the learning inverse selection mechanism until all source nodes have selected relay nodes for service data transmission.

[0163] The specific process of the relay selection optimization method for heterogeneous communication networks supporting distributed resource collaborative interaction is as follows: Figure 3 As shown, the steps are described below:

[0164] Step 1: Determine the HPLC channel status between the source node and the destination node. If the channel status is good, use HPLC for data transmission and proceed to the next time slot; otherwise, proceed to Step 2 and select an appropriate relay node for data transmission.

[0165] Step 2: Initialize the state action value function table Time slot t=0;

[0166] Step 3: Each source node queries the value function table based on the current time slot service data volume and its own status action. The action with the highest state-action value function is selected, and the action reward for the current time slot is calculated based on the selected action. The action selection strategy can be expressed as:

[0167]

[0168] Step 4: Determine if there is a conflict where different source nodes select the same relay node. If a conflict occurs, the relay node will reverse the selection of the source node based on the learning reverse selection mechanism, and then proceed to Step 5; if no conflict occurs, proceed directly to Step 6.

[0169] Step 5: At this point, other source nodes select a relay node for business data transmission based on their suboptimal state action value function, and then return to step 4 to continue to determine whether there is a conflict in which different source nodes select the same relay node.

[0170] Step 6: Enter the next time slot, and update the status action value function according to the business data volume of the next time slot and formula (14).

[0171] Step 7: Determine if the current time slot has reached the maximum time slot T. If it has, the algorithm ends; otherwise, the source node uses the current action value function table. Select a new relay node.

[0172] This invention presents a simulation experiment on the proposed relay selection optimization method for heterogeneous communication networks supporting distributed resource collaborative interaction. The simulation results are as follows:

[0173] Figure 4 This section compares the total latency of business data transmission. Simulation results show that, compared with other algorithms, the algorithm proposed in this invention can effectively reduce the latency of distributed resource nodes selecting relay nodes for business data transmission. This is because this invention comprehensively considers the source node's selection of relay nodes, aiming to minimize the weighted value of latency and bit error rate. Through continuous learning and optimization, each source node will select the optimal transmission scheme for the relay node, thereby reducing the latency of distributed resource nodes selecting relay nodes for business data transmission.

[0174] Figure 5 This section compares the overall bit error rate (BER) of business data transmission. Simulation results show that, compared with other algorithms, the algorithm proposed in this invention can effectively reduce the BER when distributed resource nodes select relay nodes for business data transmission. This is because the algorithm uses minimizing the weighted value of latency and BER as its optimization objective and introduces a learning-based inverse selection mechanism to solve the problem of multiple source nodes selecting the same relay node for data transmission. This ensures that all source nodes have selected a relay node for all business data transmission, thereby reducing the BER of distributed resource nodes transmitting business data.

[0175] like Figure 6 As shown, the present invention provides a relay selection optimization system for heterogeneous communication networks, comprising:

[0176] The channel state determination module is used to determine the HPLC channel state between the source node and the destination node. If the HPLC channel state estimation meets the requirements, HPLC is selected for data transmission and the process moves to the next time slot. If the channel state estimation does not meet the requirements, each source node selects a relay node for data transmission based on the current time slot's service data volume and its own state action value function.

[0177] The relay node judgment module is used to determine whether there is a conflict in which different source nodes select the same relay node. If a conflict occurs, the relay node will reverse the selection of the source node based on the learning reverse selection mechanism. Other source nodes will reselect the relay node for business data transmission according to their own state action value function. The module will then continue to determine whether there is a conflict in which different source nodes select the same relay node. If no conflict occurs, the module will proceed to the next time slot and update the state action value function according to the business data volume of the next time slot.

[0178] The maximum time slot determination module is used to determine whether the current time slot has reached the maximum time slot. If it has, the process ends; if it has not reached the maximum time slot, the process returns to the step of determining the HPLC channel status between the source node and the destination node.

[0179] like Figure 7 As shown, the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the heterogeneous communication network relay selection optimization method that takes into account the latency characteristics of multiple interactive functions.

[0180] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the method for optimizing relay selection in heterogeneous communication networks that takes into account the latency characteristics of multiple interactive functions.

[0181] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0182] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0183] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0184] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0185] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A relay selection optimization method for heterogeneous communication networks, characterized in that, include: Determine the HPLC channel status between the source node and the destination node. If the HPLC channel status estimate meets the requirements, select HPLC for data transmission and proceed to the next time slot. If the channel state estimate does not meet the requirements, each source node selects a relay node for data transmission based on the current time slot service data volume and its own state action value function. Determine if there is a conflict where different source nodes select the same relay node. If a conflict occurs, the relay node will reverse the selection of the source node based on the learning reverse selection mechanism. Other source nodes will reselect the relay node for business data transmission based on their own state action value function. And continue to determine whether there is a conflict where different source nodes select the same relay node; If no conflict occurs, proceed to the next time slot and update the status action value function according to the amount of business data in the next time slot; Determine if the current time slot has reached the maximum time slot. If it has, end the process; otherwise, return to the step of determining the HPLC channel status between the source node and the destination node. Each source node selects a relay node for data transmission based on the current time slot service data volume and its own state action value function, including: Each source node queries the action based on the current time slot service data volume and its own status action value function table. Select the action with the highest state-action value function, and calculate the action reward for the current time slot based on the selected action; The learning-based inverse selection mechanism performs inverse selection of source nodes by relay nodes, and other source nodes reselect relay nodes for business data transmission based on their own state-action value functions, including: When different source nodes choose the same action, a learning-based inversion mechanism is introduced; in this mechanism, up to the current time slot, the source nodes... Selecting relay nodes The average reward is the empirical performance, then the source node The current experience learning performance is represented by the source node. The priority is multiplied by the empirical performance, and the specific formula is: When different source nodes select the same relay node, the relay node is deselected. The source node with the larger value transmits the business data. Other source nodes select relay nodes for business data transmission based on their suboptimal state action value function. If selection conflicts still exist, the selection of relay nodes continues according to the learning reverse selection mechanism until all source nodes have selected relay nodes for business data transmission.

2. The relay selection optimization method for heterogeneous communication networks according to claim 1, characterized in that, The method for selecting the action with the highest state-action value function is as follows: In the formula, For the source node State space, source node The action space is , Indicates the first Source nodes within each time slot Select the action, select the first Each relay node transmits business data; For the source node The state-action value function has the following properties: in, It's about learning efficiency. This is the discount factor.

3. The relay selection optimization method for heterogeneous communication networks according to claim 1, characterized in that, Each source node selects a relay node for data transmission based on the current time slot service data volume and its own state action value function, specifically including: Multi-service data transmission employs a latency model and a bit error rate model. A model function for selecting relay nodes for data transmission is constructed with the objective of minimizing the weighted average of latency and bit error rate. The model function is as follows: in, Indicates the source node priority, For time delay, For bit error rate, This represents the weighting parameter; C1 indicates that the total latency of business data transmission on each distributed resource node cannot exceed its limit. The total bit error rate of business data transmission on each distributed resource node must not exceed its limit. C2 indicates that each source node selects only one relay node in each round; The selection of relay nodes for source node service data transmission is optimized using an action selection approach; the state space represents the amount of service data that each source node needs to transmit. The state space is denoted as ;Source node The action space is , Indicates the first Source nodes within each time slot Select the action, select the first Each relay node transmits business data; and source node The reward for selecting a relay node to transmit service data is converted into a weighted sum of latency and bit error rate, and expressed as: Among them, the source node The state-action value function is And there are: in, It's about learning efficiency; This is a discount factor, representing the degree of influence of state returns; Each source node selects a relay node based on its own optimal state action value function.

4. The method for optimizing relay selection in heterogeneous communication networks according to claim 3, characterized in that, The bit error rate model is as follows: Source node To the destination node Total bit error rate of transmitted data Represented as: In the formula, For binary indicator variables, This represents the number of relay nodes; For relay nodes To the destination node Bit error rate of transmitted data in, For relay nodes to the destination node The data transmission rate; These represent the channel bandwidth radiated during the data transmission of services; In the formula, For the source node To relay node The data transmission error rate; in For spreading factor, Here is the harmonic number, where The tail integral function of the standard Gaussian distribution. For the source node To relay node Signal-to-noise ratio of transmitted data.

5. The relay selection optimization method for heterogeneous communication networks according to claim 3, characterized in that, The delay model is as follows: The latency model is used to represent the end-to-end transmission latency of distributed resource data, including the communication latency from the source node to the relay node and the communication latency from the relay node to the destination node, and the total latency of business data transmission. Represented as: In the formula, For the source node To relay node Data transmission latency, , Constant; source node To relay node Data transmission rate Represented as: in, Indicates the channel white noise power; and These represent the electromagnetic interference radiated during service data transmission using HRF and the channel bandwidth, respectively. Indicates the signal transmission power of the source node; Indicates time slot Intrinsic nodes To relay node Channel gain when transmitting data; For relay nodes to the destination node Data transmission latency, , relay node to the destination node Data transmission rate for: in, and These represent the electromagnetic interference radiated during service data transmission using HRF and the channel bandwidth, respectively. Indicates the signal transmission power of the source node; Indicates time slot Intrinsic nodes To relay node Channel gain when transmitting data.

6. The relay selection optimization method for heterogeneous communication networks according to claim 1, characterized in that, The channel state adopts a quasi-static time-slot model, and there exists... Each relay node has an equal-length time slot. The channel state information remains unchanged within one time slot but changes in different time slots. Within each time slot, each relay node can only be selected by one source node.

7. A relay selection optimization system for heterogeneous communication networks, based on the relay selection optimization method for heterogeneous communication networks according to any one of claims 1 to 6, characterized in that, include: The channel state determination module is used to determine the HPLC channel state between the source node and the destination node. If the HPLC channel state estimation meets the requirements, HPLC is selected for data transmission, and the process proceeds to the next time slot. If the channel state estimation does not meet the requirements, each source node selects a relay node for data transmission based on the current time slot's service data volume and its own state action value function. The selection of a relay node based on the current time slot's service data volume and its own state action value function includes: Each source node queries the action based on the current time slot service data volume and its own status action value function table. Select the action with the highest state-action value function, and calculate the action reward for the current time slot based on the selected action; The relay node judgment module is used to determine whether there is a conflict in which different source nodes select the same relay node. If a conflict occurs, the relay node will reverse the selection of the source node based on the learning reverse selection mechanism. Other source nodes will reselect the relay node for business data transmission according to their own state action value function. The module will continue to determine whether there is a conflict in which different source nodes select the same relay node. If no conflict occurs, the module will proceed to the next time slot and update the state action value function according to the business data volume of the next time slot. The maximum time slot determination module is used to determine whether the current time slot has reached the maximum time slot. If it has, the process ends; if it has not reached the maximum time slot, the process returns to the step of determining the HPLC channel status between the source node and the destination node.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the steps of the relay selection optimization method for heterogeneous communication networks according to any one of claims 1-6.

9. A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the relay selection optimization method for heterogeneous communication networks according to any one of claims 1-6.