Routing Selection Method, System, Device and Equipment Applicable to Electric Power Communication Services
By obtaining and processing the status information of routing nodes in power communication services, combining Q learning algorithm and graph search algorithm to calculate channel similarity and random process weights, the problem of low efficiency of Q learning algorithm in routing is solved, and the rationality and communication quality of routing are improved.
Patent Information
- Application Number
- CN202510329270.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-20
AI Technical Summary
When the Q learning algorithm is executed in routing, the difference between the simulated data and the actual solution is too large, resulting in low reinforcement learning efficiency, and the algorithm responds slowly to changes in the routing network state, resulting in unreasonable routing results.
By obtaining the node occupancy, bandwidth and communication delay of the routing nodes in the power communication service, the graph search algorithm is used to obtain the shortest communication delay, and dimensionality reduction is used to obtain the routing node coordinate set. Combining the Q learning algorithm, the channel similarity and random process weight are calculated, the R table is constructed and the Q table is updated to improve the rationality of routing.
The Q-learning algorithm response speed to changes in the routing network state is improved, the rationality of routing selection is enhanced, and the communication quality of power communication services is improved.
Smart Images

Figure CN119854198B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of power communication technology, and in particular to a routing selection method, system, device and equipment applicable to power communication services. Background Art
[0002] The power communication service is a key service used to transmit control information in the power grid system. It relies on the collaborative work with the power grid to build the industrial Internet and realize information transmission. This type of service has very high requirements for reliability, security and latency consistency to ensure that the power system can operate stably and efficiently. At present, the power communication network is mainly based on optical fiber communication. In this process, different routers need to be controlled as communication transfer stations to ensure the high reliability of the communication process.
[0003] When two routing nodes in the power system are communicating by routing, a secure and reliable router link needs to be selected for communication. Traditional methods usually aim to optimize the balance and efficiency of the entire communication network, and use the Q learning algorithm to determine the router link to complete data communication. This method uses the node utilization as the evaluation criterion and obtains the final path selection Q table by running the Q learning algorithm. During the generation of the Q table, the system randomly controls the router to simulate the formation of the link, and updates and generates the Q table by simulating the communication effect of the link.
[0004] In this process, the simulated link does not fully consider the communication direction between routers, resulting in a large difference between the simulated link generated by the Q learning algorithm and the actual link when it is executed. This gap reduces the convergence speed of Q learning, thereby slowing down the algorithm's response to changes in router network status. This situation may lead to unreasonable routing selection, which in turn affects the communication quality of power communication services. Summary of the invention
[0005] In view of the above, it is necessary to provide a routing selection method, system, device and equipment suitable for power communication business to solve the problem that when the Q learning algorithm is executed in routing selection, the reinforcement learning efficiency is low due to the large difference between the simulation data and the actual scheme habits, which makes the algorithm respond slowly to the changes in the routing network state, resulting in unreasonable routing selection results.
[0006] In a first aspect, the present application provides a routing selection method applicable to a power communication service, the method comprising:
[0007] Obtain the node occupancy rate and node bandwidth of each routing node used for data transmission in the power communication service at each time, as well as the communication delay between routing nodes;
[0008] Adopt a graph search algorithm for the communication delay between routing nodes, obtain the shortest communication delay between any two routing nodes, and obtain a distance matrix; reduce the dimension of the distance matrix to obtain a routing node coordinate set;
[0009] In the routing node coordinate set, use the vector formed by the starting points of each simulated link as the communication direction vector, and combine the Q-learning algorithm with the similarity degree between the vector formed by each routing node and its directly transmitted routing node and the communication direction vector to obtain the channel similarity of each channel corresponding to each routing node based on each simulated link;
[0010] Denote all routing nodes other than each simulated link as the bandwidth distribution nodes of each simulated link, analyze the similarity degree between the vector from the starting point of each simulated link to each bandwidth distribution node and the communication direction vector to obtain each communication distribution weight, and combine the node bandwidths of all bandwidth distribution nodes at each moment to obtain the random process weight of each simulated link at each moment;
[0011] Based on the Q-learning algorithm, according to the channel similarity of each channel corresponding to each routing node selected by each simulated link, and in combination with the random process weight, obtain the channel learning rate of each channel of each simulated link at each moment; use whether the routing nodes can directly transmit as the judgment criterion, and in combination with the distribution of the node occupancy rate between the routing nodes, construct an R table, and replace the original learning rate in the Q-learning algorithm with the channel learning rate to update the Q table;
[0012] Select the routing for each moment according to the updated Q table at each moment and the channel similarity.
[0013] Among them, the specific method for obtaining the channel similarity of each channel corresponding to each routing node based on each simulated link is as follows:
[0014] Obtain all routing nodes that directly transmit with each routing node, and denote them as the communication nodes of each routing node;
[0015] Denote the vector from the routing node to each of its communication nodes as each communication vector; use the normalized value of the cosine similarity between each communication vector and the communication direction vector corresponding to each simulated link as the communication direction rationality of each communication vector based on each simulated link;
[0016] Replace the node selection probability in the Q-learning algorithm with the communication direction rationality; for each simulated link, use the communication direction rationality corresponding to the channel selected by each routing node based on the Q-learning algorithm as the channel similarity of each routing node based on each simulated link.
[0017] Among them, the specific method for obtaining each communication distribution weight is as follows:
[0018] Denote the vector formed by the starting point of each analog link to each bandwidth distribution node as each communication distribution vector;
[0019] Take the normalized value of the cosine similarity between each communication distribution vector and the communication direction vector of each analog link as each communication distribution weight.
[0020] Among them, the specific method for obtaining the random process weight of each analog link at each moment is as follows:
[0021] For each analog link, obtain the node bandwidth occupancy ratio of each bandwidth distribution node at each moment; based on each communication distribution weight, perform weighted summation on the corresponding node bandwidth occupancy ratio to obtain the random process weight of each analog link at each moment.
[0022] Among them, the specific formula for obtaining the channel learning rate of each channel of each analog link at each moment is: ; where is the channel learning rate of the g-th channel of the analog link at time t, is the random process weight of the analog link at time t, is the channel similarity of the g-th channel of the analog link ;
[0023] Among them, the specific method for constructing the R table is as follows:
[0024] If it can be directly transmitted from the routing node s to the routing node a, calculate the average value of the node occupancy ratios of the two routing nodes, and take the difference between the first preset value and the average value as the element in the s-th row and a-th column of the R table; if it cannot be directly transmitted, take the second preset value as the element in the s-th row and a-th column of the R table;
[0025] Among them, the first preset value is greater than or equal to 1; the second preset value is less than 0.
[0026] Among them, the method for selecting the route at each moment according to the updated Q table at each moment and the channel similarity includes:
[0027] For any routing node at each moment, obtain the Q value on the Q table with other routing nodes, fuse it with the negative correlation mapping result of the channel similarity with other routing nodes, and calculate the routing control judgment value;
[0028] Select the routing node with the smallest routing judgment value among the routing nodes of the any routing node as the next routing node for information transmission.
[0029] In a second aspect, the present application provides a routing selection system applicable to power communication services, and the system includes:
[0030] A router status acquisition module, which is used to acquire the node occupancy rate, node bandwidth of each routing node for data transmission in power communication services at each moment, and the communication delay between routing nodes;
[0031] Adopt a graph search algorithm for the communication delay between routing nodes to obtain the shortest communication delay between any two routing nodes, and obtain a distance matrix; reduce the dimension of the distance matrix to obtain a routing node coordinate set;
[0032] A Q-table generation module, which is used to use the vector formed by the starting points of each simulated link in the routing node coordinate set as the communication direction vector, and combine the Q-learning algorithm with the similarity degree between the vector formed by each routing node and its directly transmitted routing node and the communication direction vector to obtain the channel similarity of each channel corresponding to each routing node based on each simulated link;
[0033] Denote all routing nodes other than each simulated link as the bandwidth distribution nodes of each simulated link, analyze the similarity degree between the vector from the starting point of each simulated link to each bandwidth distribution node and the communication direction vector to obtain each communication distribution weight, and combine the node bandwidth of all bandwidth distribution nodes at each moment to obtain the random process weight of each simulated link at each moment;
[0034] Based on the Q-learning algorithm, according to the channel similarity of each channel corresponding to each routing node selected by each simulated link, combine the random process weight to obtain the channel learning rate of each channel of each simulated link at each moment; use whether the routing nodes can directly transmit as a judgment criterion, combine the distribution of the node occupancy rate between the routing nodes, construct an R-table, and replace the original learning rate in the Q-learning algorithm with the channel learning rate to update the Q-table;
[0035] A routing selection module, which is used to select the routing at each moment according to the updated Q-table at each moment and the channel similarity.
[0036] In a third aspect, the present application provides a routing selection device applicable to power communication services, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the method described in any one of the above are implemented.
[0037] In a fourth aspect, the present application provides a routing selection device applicable to power communication services. A computer program is stored in the device, and when the computer program is executed by a processor, the method described in any one of the above is implemented.
[0038] In the above solution, first, obtain the node occupancy rate, node bandwidth, and communication delay between routing nodes at each moment for data transmission in power communication services; since direct communication between two routing nodes is not possible, the graph search algorithm is used for the communication delay between routing nodes to obtain the shortest communication delay between any two routing nodes, resulting in a distance matrix. The beneficial effect is that the distance between routing nodes is characterized by the shortest communication delay; the distance matrix is dimensionally reduced to obtain a set of routing node coordinates, which is convenient for subsequently judging the directional differences between different channels and the final transmission target based on the coordinates of the routing nodes, guiding the selection of the simulation operation scheme of the Q-learning algorithm, and obtaining the channel similarity of each routing node corresponding to the channel based on each simulated link. The beneficial effect is to improve the quality of the generated simulation data, thereby improving the response speed of the Q-learning algorithm to changes in the routing network state.
[0039] Furthermore, through the node bandwidth and the set of routing node coordinates, calculate the random process weight, which is used to judge the similarity between the simulation data and the actual communication situation, and filter out the simulation data close to the actual transmission process; combine the channel similarity and the random process weight to calculate the channel learning rate. The higher the channel learning rate, the more it indicates that it conforms to the channel control result of the actual communication link scheme. When updating the Q-table, increase the learning rate for the corresponding channel, so that the Q-learning algorithm can learn more quickly to conform to the actual communication link scheme; finally, perform routing selection through the Q-table, and assign different learning rates to the simulation data according to the actual router information transmission habit, accelerating the learning rate of the Q-learning algorithm for changes in the routing network state and the response speed to changes in the routing network state, making the selected router communication link more reasonable and improving the communication quality. Brief Description of the Drawings
[0040] Figure 1 It is a flowchart of the steps of the routing selection method applicable to power communication services provided by this application;
[0041] Figure 2 It is a network schematic diagram provided by this application;
[0042] Figure 3 It is a block diagram of the routing selection system applicable to power communication services provided by this application. Detailed Embodiments
[0043] In the description of the embodiments of this application, words such as "exemplary", "or", and "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary", "or", and "for example" is intended to present relevant concepts in a specific manner.
[0044] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs. The terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application.
[0045] In addition, it should be noted that the terms "first" and "second" in this application and the accompanying drawings are used to distinguish similar objects and are not used to describe a specific order or sequence. The methods disclosed in the embodiments of this application or shown in the flowcharts of the methods include one or more steps for implementing the methods. Without departing from the scope of this application, the execution order of multiple steps can be interchanged with each other, and some steps can also be deleted.
[0046] Please refer to Figure 1 , which shows the step flowchart of a routing selection method applicable to power communication services provided by an embodiment of this application. The method includes the following steps:
[0047] Step 1: Obtain the node occupancy rate, node bandwidth, and communication delay between routing nodes at each moment for each routing node used for data transmission in power communication services.
[0048] This application uses the Q-learning algorithm as the algorithm for routing control in power communication services. In this process, it is necessary to obtain the status of the router network in real time, construct a corresponding R table according to the router network status, use the R table as the input, generate a Q table using the Q-learning algorithm, and select the corresponding communication link through the Q table. The routing network status information obtained in this embodiment includes: the occupancy rate of the node, the bandwidth of the node, and the communication delay between nodes.
[0049] In this embodiment, there are a total of N routing nodes used for data transmission in power communication services. Record the communication delay obtained between routing nodes through the power communication system, and denote the communication delay between routing node A and routing node B as . In particular, two routers may not be able to communicate directly due to excessive distance, and in this case, the communication delay between the two routing nodes cannot be directly obtained.
[0050] For the nth routing node, through the power communication system, obtain the occupancy rate of this node at time t, denoted as the node occupancy rate , and obtain the bandwidth of the nth routing node, denoted as the node bandwidth .
[0051] Therefore, for N routing nodes in a routing network, they form a network consisting of N routing nodes. The nth routing node in the network represents the nth router, and each routing node has its node occupancy rate. At the same time, there are connections between nodes, indicating that two routing nodes can communicate with each other. The attribute value of the connection is the communication delay, representing the time required for two routing nodes to communicate. Among them, the network schematic diagram is as shown in Figure 2 shown Figure 2 a network composed of 6 routing nodes. In the figure, a circle represents a routing node with an attribute value. If two nodes can communicate directly, there is a connection, and the attribute value of the connection is the communication delay.
[0052] Finally, the occupancy rates and node bandwidths of N routing nodes, as well as several communication delays, are obtained. The number of communication delays is determined by the number of routers that can communicate directly with each other. It should be specifically noted that the node occupancy rate and bandwidth will change with the state of the routing network, while the communication delay is the communication delay obtained when the routing network is operating normally and will not change with time.
[0053] Step 2: Apply a graph search algorithm to the communication delays between routing nodes to obtain the shortest communication delay between any two routing nodes, and obtain a distance matrix. Then, reduce the dimension of the distance matrix to obtain a set of routing node coordinates.
[0054] When the traditional Q-learning algorithm is applied to the routing selection of power communication services, two routing nodes are randomly selected from N routing nodes as simulated communication tasks. In this embodiment, taking node BG and node ED as examples, the communication process from BG to ED is simulated.
[0055] In an actual communication link scheme, a communication link scheme with a shorter total delay time is usually selected. Therefore, in this embodiment, the shortest communication delay between the communication link scheme and two routing nodes is obtained, and the shortest communication delay is compared with the communication delay of the simulated link to determine whether a communication link scheme meets the actual communication link scheme.
[0056] This application takes routing node A and routing node B as examples for analysis: First, if two routing nodes A and B can communicate directly, their shortest communication delay is If two routing nodes A and B cannot communicate directly, taking the communication structure data set as the input, the Dijkstra algorithm is used for calculation. The parameters are that the two connected routing nodes are A and B, and the output is the shortest communication delay between the two routing nodes. Finally, there is a shortest communication delay between any two of the N routing nodes, and a total of H shortest communication delays are obtained. Among them, the Dijkstra algorithm is a graph search algorithm, which is a well-known existing technology, and this application will not elaborate on it here.
[0057] It should be understood that for any simulated link, the closer its communication delay is to the shortest communication delay, the more it conforms to the actual communication link scheme, and the simulated data generated by it can make the Q-table converge faster, improving the response speed of the Q-learning algorithm to the changes in the routing network state.
[0058] First, the shortest communication delays are processed by the manifold learning method, and the shortest communication delays and the corresponding routing nodes are mapped to the coordinate system, that is, the H shortest communication delays form a distance matrix, as shown below:
[0059]
[0060] Among them, , both represent the shortest communication delay between the first routing node and the second routing node; , both represent the shortest communication delay between the Nth routing node and the first routing node; , both represent the shortest communication delay between the Nth routing node and the second routing node.
[0061] It should be understood that the distance matrix is a square matrix with side length N. The element in the nth row and the ith column of the distance matrix represents the shortest communication delay between the nth routing node and the ith routing node. The elements on the main axis of the matrix are the shortest communication delays of the node communicating with itself, so the values are 0. It should be noted that the elements symmetric along the main diagonal in the distance matrix are equal, for example = .
[0062] Furthermore, taking the distance matrix as the input, the Multiple Dimensional Scaling (MDS) algorithm is used to obtain a two-dimensional coordinate point set, denoted as the routing node coordinate set. The routing node coordinate set contains N points, where the nth point corresponds to the nth routing node. Among them, the MDS algorithm is a commonly used technology in the field of data processing and is a dimensionality reduction algorithm for manifold learning, which can map the nodes to a two-dimensional coordinate system through the node distance information.
[0063] Step 3: In the routing node coordinate set, use the vector formed by the starting points of each simulated link as the communication direction vector. Combine the similarity between the vector formed by each routing node and its directly transmitting routing node and the communication direction vector with the Q-learning algorithm to obtain the channel similarity of each routing node corresponding to the channel based on each simulated link.
[0064] This application aims to reduce the generated unrealistic communication links by intervening in the generation process of the simulated link scheme, thereby improving the quality of the simulated data of Q-learning and accelerating its convergence speed. Therefore, this application needs to evaluate the impact of different generation paths on the rationality of the final communication link during the simulated link generation process.
[0065] When using the shortest communication delay to determine whether a simulated link conforms to the actual communication link, the evaluation can only be carried out for the entire simulated link, and it is impossible to evaluate the performance of a specific channel alone. In this regard, this application maps the shortest communication delay and nodes into the coordinate system, and judges the rationality of a single channel through the direction of the single channel and the direction of the simulated link. The specific process is as follows:
[0066] This application takes the simulated link from routing node BG to routing node ED as an example for analysis. In the routing node coordinate set, with routing node BG as the starting point and routing node ED as the end point, construct a vector denoted as the communication direction vector.
[0067] When the Q-learning algorithm randomly generates the simulated link between routing node BG and routing node ED, it will use routing node BG as the first node with existing information, and then find other nodes that can directly transmit with routing node BG. Randomly select a routing node with the same probability among these directly transmitting nodes as the second node with existing information; further obtain the third node with existing information in the same way through the second node with existing information, and so on, until the next routing node found is node ED.
[0068] In this application, taking a node C with existing information as an example for analysis, denote the routing nodes that can directly transmit with the node C with existing information as communication nodes e. There are a total of E communication nodes that can directly transmit with node C, and this is used to describe the process of generating the simulated link.
[0069] When the Q-learning algorithm randomly selects a communication node from the E communication nodes of node C to form a simulated channel with node C, the probability of selecting a node from the E communication nodes is equal. However, not all of the E communication nodes conform to the communication habits. Therefore, the method of selecting nodes with the same selection probability will result in the simulated link not conforming to the actual communication link scheme, leading to a decline in the quality of the generated simulated data. In response to this, starting from node C, this application constructs E communication vectors with the E communication nodes as the endpoints respectively. The cosine similarity between the e-th communication vector and the communication direction vector is calculated. The greater the cosine similarity, the more the single channel formed by node C and its e-th communication node conforms to the actual communication link scheme.
[0070] Furthermore, the E cosine similarities of node C are normalized. In this embodiment, the normalization method is as follows: First, add 1 to each of the E cosine similarities to complete the non-negativity processing. Then, sum the non-negatively processed cosine similarities, and then divide each of the non-negatively processed cosine similarities by the sum value to obtain the normalized cosine similarity. The normalized cosine similarity is denoted as the communication direction rationality. In other embodiments, the maximum-minimum method can be used for normalization. It should be understood that the greater the communication direction rationality, the more the single channel formed by node C and its e-th communication node conforms to the actual communication link scheme.
[0071] When the Q-learning algorithm randomly selects a routing node from the E communication nodes at node C to form a channel, the communication direction rationality is calculated by the method described herein. When selecting a node, the e-th communication direction rationality is used as the selection probability of the e-th communication node, replacing the same selection probability of the E communication nodes in the original Q-learning algorithm, and then a channel is selected to obtain a simulated link. Compared with the traditional simulated communication method, the simulated link construction scheme described in this embodiment can make the generated simulated link conform to the actual communication link scheme, improve the quality of the simulated data, and improve the convergence speed of the Q-learning algorithm. At the same time, for the channel selected by the Q-learning algorithm based on the simulated link, the corresponding communication direction rationality is obtained and denoted as the channel similarity.
[0072] Step Four: Denote all routing nodes other than each simulated link as the bandwidth distribution nodes of each simulated link, analyze the similarity between the vector from the starting point of each simulated link to each bandwidth distribution node and the communication direction vector to obtain each communication distribution weight, and combine the node bandwidths of all bandwidth distribution nodes at each moment to obtain the random process weight of each simulated link at each moment.
[0073] In the Q-learning algorithm, the link selection process from routing node BG to routing node ED is random, but in the actual routing network, the links for communication are not random, but there are obvious tendencies. For example, most of the nodes that issue communication requests and the final communication nodes are fixed, and some routers only serve as transit nodes.
[0074] Therefore, for the selection of routing nodes BG and routing nodes ED, this application characterizes whether the selection of routing nodes conforms to the actual communication situation by calculating the random process weight. The specific method is as follows:
[0075] There are N routing nodes in the routing node coordinate set. The nodes other than routing node BG and routing node ED are recorded as simulated links. bandwidth distribution node; taking the routing node BG as the starting point and the mth bandwidth distribution node as the end point, the mth communication distribution vector is obtained; the mth cosine similarity of the mth communication distribution vector and the communication direction vector is calculated and normalized to obtain the mth communication distribution weight. In this embodiment, the method for normalizing the cosine similarity of the communication distribution vector and the communication direction vector is: adding 1 to the cosine similarity and then dividing it by 2 to obtain the mth communication distribution weight; the implementer may also use other normalization methods, which are not limited in this application.
[0076] It should be understood that the smaller the cosine similarity between the mth communication distribution vector and the communication direction vector, the closer the mth routing node is to the connecting line segment between routing node BG and routing node ED. When routing node BG and routing node ED are selected for simulated communication, the simulated link is performed through the mth bandwidth distribution node. The more communication can achieve analog link The lower the latency of the communication, the more it conforms to the communication habits of the nodes, and the more likely the mth bandwidth distribution node is to participate in the communication between the two. Therefore, the greater the mth communication distribution weight, the more likely the mth bandwidth distribution node is to participate in the communication between the two routing nodes, which means that the bandwidth distribution node is more likely to become one of the routing nodes in the communication link connecting the two routing nodes.
[0077] Finally, taking the communication distribution weight as the weight, the node bandwidth of all bandwidth distribution nodes is weighted and summed, and then divided by the sum of the bandwidth of all bandwidth distribution nodes to obtain the random process weight. The larger the value, the more the bandwidth of the entire routing network is distributed in the simulated link. The more frequently the routing nodes BG and ED communicate in the current routing network, the more important the simulated data they generate. In particular, the random process weight is related to the time t at which the node bandwidth is obtained. The random process weight is recorded as .
[0078] Step 5: Based on the Q - learning algorithm, according to the channel similarity of the channels corresponding to each routing node selected by each simulated link, combined with the random process weight, obtain the channel learning rate of each channel of each simulated link at each moment; use whether the routing nodes can directly transmit as the judgment criterion, combined with the distribution of the node occupancy rate between the routing nodes, construct the R - table, and replace the original learning rate in the Q - learning algorithm with the channel learning rate to update the Q - table.
[0079] After the Q - learning algorithm generates the simulated link, each channel can be regarded as a state - action pair in the Q - table. In each simulation operation, the calculation result updates the value of the corresponding element in the Q - table, enabling the system to learn how to optimize the selection of communication links according to the reward feedback. In this way, the R - table can be used to record the reward information and help the system optimize during the continuous trial - and - error process. When the Q - value in the Q - table converges, it indicates that the Q - learning has successfully learned the optimal strategy, and at this time, the routing control can be performed through the Q - table.
[0080] To accelerate the convergence speed of the R - table, this application calculates the channel learning rate. For channels that are more in line with the actual routing control, the higher their channel learning rate and the higher their weight when constructing the R - table. The calculation method is as follows:
[0081] Still use the simulated link for analysis. Assume that a total of G channels are obtained in this simulation process. For the g - th channel among them, obtain the corresponding channel similarity ; finally, calculate the channel learning rate: ; where is the channel learning rate of the g - th channel of the simulated link at time t, is the random process weight of the simulated link at time t, is the channel similarity of the g - th channel of the simulated link
[0082] It should be understood that the greater the channel similarity, the more the data obtained by the g - th channel in this simulated communication process conforms to the actual communication situation, and the better the simulated data generated. Therefore, the data generated in this communication process should have a higher learning rate.
[0083] Furthermore, this application improves the transfer rule formula of the Q - learning algorithm through the obtained channel learning rate to accelerate the convergence speed of the Q - table.
[0084] This application obtains the node occupancy rate and node bandwidth of the routing network in real time and constructs an R table: At time t, the element in the s-th row and a-th column of the R table represents the channel from the s-th routing node to the a-th routing node. Therefore, the node occupancy rates of the s-th routing node and the a-th routing node are obtained and averaged, and then the first preset value is subtracted from the average value to obtain the channel idle rate, which is used as the element in the s-th row and a-th column of the R table; the larger the value, the more communication information can be accommodated by the channel between the s-th routing node and the a-th routing node, and this channel should be used; if the s-th routing node and the a-th routing node cannot directly transmit, the second preset value is used as the channel idle rate between the s-th routing node and the a-th routing node, which is used as the element in the s-th row and a-th column of the R table; in this embodiment, the first preset value is greater than or equal to 1 and takes the value of 1; the second preset value is less than 0 and takes the value of -1, and the implementer can adjust it according to the actual situation.
[0085] It should be understood that there is a corresponding R table at each moment. The larger the value of a certain element in the R table, the higher the idle rate between the two routing nodes. Selecting the corresponding nodes for communication will balance the load of the entire routing network and improve network reliability.
[0086] At time t, the node occupancy rate and node bandwidth of N routing nodes can be obtained, and at the same time, the communication delay of the entire routing network is obtained; when using Q-learning for simulation operations, during a simulation process, the communication task of the simulated link BG→ED is analyzed, and a channel is selected for transmission. After that, the in the Q table needs to be updated. In this embodiment, the channel learning rate is used to replace the fixed learning rate in the traditional transition rule formula.
[0087] It should be noted that compared with the traditional fixed learning rate, the higher the channel learning rate of the channel , the more it conforms to the channel control result of the actual communication link scheme. The data generated by the simulation operation is more credible, its impact on the Q table is greater, and it can make the learning result of Q-learning converge faster to a result similar to the actual communication link scheme, improving the response speed of the Q-learning algorithm to the change of the routing network state when performing routing selection.
[0088] Step 6: Select the routing for each moment according to the updated Q table at each moment and the channel similarity.
[0089] Taking the R table as the input, the Q-learning algorithm is used for calculation to obtain the corresponding Q table; among them, the transition rule formula in the Q-learning algorithm is replaced according to the method described in this embodiment. The specific calculation process of the Q-learning algorithm is a commonly used technique in the field of reinforcement learning and will not be elaborated here;
[0090] Finally, obtain the corresponding Q-table in real time. When performing route selection, the current information is at the routing node s. At this time, obtain the Q-values of the s-th row of the Q-table, and obtain a total of H non-zero Q-values, corresponding to H channels. If the h-th channel is in the a-th column, its corresponding channel is ; calculate the channel similarity of the h-th channel, and then divide the Q-value of the h-th channel by the sum of the channel similarity and z to obtain the routing control judgment value. In this embodiment, z is a constant to prevent the denominator from being zero, and its value is 0.01. Select the channel composed of the routing nodes with the smallest routing judgment value of the routing node s as the channel for information transmission.
[0091] It should be understood that the smaller the Q-value of the h-th channel, the more the selection of this channel can balance the load of the routing network; the greater the channel similarity, the more the h-th channel is the channel for transmitting in the direction of the target routing node, and the more this channel should be selected.
[0092] Based on the same concept as the method embodiment of the present application, a route selection system applicable to power communication services is proposed, including:
[0093] A router state acquisition module, configured to acquire the node occupancy rate, node bandwidth, and communication delay between routing nodes of each routing node used for data transmission in power communication services at each moment;
[0094] Adopt a graph search algorithm for the communication delay between routing nodes to obtain the shortest communication delay between any two routing nodes, and obtain a distance matrix; perform dimensionality reduction on the distance matrix to obtain a routing node coordinate set;
[0095] A Q-table generation module, configured to use the vector formed by the starting points of each simulated link in the routing node coordinate set as the communication direction vector, and combine the Q-learning algorithm with the similarity between the vector formed by each routing node and its directly transmitted routing node and the communication direction vector to obtain the channel similarity of each routing node corresponding to the channel based on each simulated link;
[0096] Denote all routing nodes other than each simulated link as the bandwidth distribution nodes of each simulated link, analyze the similarity between the vector from the starting point of each simulated link to each bandwidth distribution node and the communication direction vector to obtain each communication distribution weight, and combine the node bandwidths of all bandwidth distribution nodes at each moment to obtain the random process weight of each simulated link at each moment;
[0097] Based on the channel similarity of the channels corresponding to each routing node selected by each simulated link according to the Q-learning algorithm, combined with the weight of the stochastic process, the channel learning rate of each channel of each simulated link at each moment is obtained; using whether the routing nodes can directly transmit as the judgment criterion, combined with the distribution of the node occupancy rates between the routing nodes, an R table is constructed, and the channel learning rate is used to replace the original learning rate in the Q-learning algorithm to update the Q table.
[0098] A routing selection module, configured to select the routing at each moment according to the Q table updated at each moment and the channel similarity.
[0099] Among them, the block diagram of the routing selection system applicable to power communication services is as Figure 3 shown.
[0100] Based on the same concept as the method embodiment of the present application, a routing selection device applicable to power communication services is proposed, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the method described in any one of the above are implemented.
[0101] Based on the same concept as the method embodiment of the present application, a routing selection device applicable to power communication services is proposed. A computer program is stored in the device, and when the computer program is executed by the processor, the method described in any one of the above is implemented.
[0102] In summary, the present application first obtains the node occupancy rate, node bandwidth of each routing node for data transmission in power communication services at each moment, and the communication delay between the routing nodes. Since two routing nodes cannot directly communicate, the graph search algorithm is used for the communication delay between the routing nodes to obtain the shortest communication delay between any two routing nodes, and a distance matrix is obtained. The beneficial effect is that the distance between the routing nodes is characterized by the shortest communication delay; the distance matrix is dimensionally reduced to obtain a routing node coordinate set, which is convenient for subsequently judging the direction difference between different channels and the final transmission target through the coordinates of the routing nodes, guiding the selection of the Q-learning algorithm simulation operation scheme, and obtaining the channel similarity of each routing node corresponding to the channel based on each simulated link. The beneficial effect is to improve the generation quality of the simulated data, and further improve the response speed of the Q-learning algorithm to the changes in the routing network state.
[0103] Further, based on the node bandwidth and the set of routing node coordinates, calculate the weight of the stochastic process, which is used to determine the similarity between the simulated data and the actual communication situation, and filter out the simulated data that is close to the actual transmission process. Combine the channel similarity and the weight of the stochastic process to calculate the channel learning rate. The higher the channel learning rate, the more it conforms to the channel control result of the actual communication link scheme. When updating the Q-table, increase the learning rate for the corresponding channel, so that the Q-learning algorithm can learn more quickly to conform to the actual communication link scheme. Finally, perform routing selection through the Q-table, assign different learning rates to the simulated data according to the actual router information transmission habit, accelerate the learning rate of the Q-learning algorithm for the change of the routing network state and the response speed to the change of the routing network state, make the selected router communication link more reasonable, and improve the communication quality.
[0104] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the systems, methods, and computer program products according to the embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the block may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description. Sometimes, there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. Each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0105] The above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A routing selection method applicable to power communication services, characterized in that: The method comprises the following steps: Obtain the node occupancy rate and node bandwidth of each routing node used for data transmission in the power communication service at each time, as well as the communication delay between routing nodes; A graph search algorithm is used to calculate the communication delay between routing nodes to obtain the shortest communication delay between any two routing nodes and obtain a distance matrix. The distance matrix is reduced in dimension to obtain a set of routing node coordinates. In the routing node coordinate set, the vector formed by the starting point of each simulated link is used as the communication direction vector, and the similarity between the vector formed by each routing node and the routing node directly transmitted by it and the communication direction vector is combined with the Q learning algorithm to obtain the channel similarity of each routing node corresponding to the channel based on each simulated link; All routing nodes except for each simulated link are recorded as bandwidth distribution nodes of each simulated link, and the similarity between the vector from the starting point of each simulated link to each bandwidth distribution node and the communication direction vector is analyzed to obtain each communication distribution weight, and the random process weight of each simulated link at each moment is obtained by combining the node bandwidth of all bandwidth distribution nodes at each moment; Based on the Q learning algorithm, according to the channel similarity of the corresponding channels of each routing node selected by each simulated link, combined with the random process weight, the channel learning rate of each channel of each simulated link at each moment is obtained; whether direct transmission between routing nodes can be used as a judgment standard, combined with the distribution of node occupancy between routing nodes, an R table is constructed, the channel learning rate is used to replace the original learning rate in the Q learning algorithm, and the Q table is updated; The route at each moment is selected according to the Q table updated at each moment and the channel similarity.
2. The routing selection method applicable to power communication services according to claim 1, characterized in that: The channel corresponding to each routing node is obtained based on the channel similarity of each simulated link, specifically: Get all routing nodes that directly transmit to each routing node, and record them as the communication nodes of each routing node; The vector from the routing node to each of its communication nodes is recorded as each communication vector; the normalized value of the cosine similarity between each communication vector and the communication direction vector corresponding to each simulated link is used as the rationality of the communication direction of each communication vector based on each simulated link; The communication direction rationality is used to replace the node selection probability in the Q learning algorithm; For each simulated link, the rationality of the communication direction corresponding to the channel selected by each routing node based on the Q learning algorithm is used as the channel similarity of each routing node based on each simulated link.
3. The routing selection method applicable to power communication services according to claim 1, characterized in that: The communication distribution weights are obtained as follows: The vectors from the starting point of each simulated link to each bandwidth distribution node are recorded as each communication distribution vector; The normalized value of the cosine similarity between each communication distribution vector and the communication direction vector of each simulated link is used as each communication distribution weight.
4. The routing selection method applicable to power communication services according to claim 1, characterized in that: The random process weight of each simulated link at each time is obtained as follows: For each simulated link, the node bandwidth proportion of each bandwidth distribution node at each time is obtained; based on each communication distribution weight, the corresponding node bandwidth proportion is weighted and summed to obtain the random process weight of each simulated link at each time.
5. The routing selection method applicable to power communication services according to claim 1, characterized in that: The specific formula for obtaining the channel learning rate of each channel of each simulated link at each moment is: ;in, is the simulated link at time t The channel learning rate of the g-th channel, is the simulated link at time t The random process weights, It is an analog link The channel similarity of the g-th channel.
6. The routing selection method applicable to power communication services according to claim 1, characterized in that: The construction of the R table is specifically as follows: If direct transmission is possible from routing node s to routing node a, the average of the node occupancy rates from routing node s to routing node a is calculated, and the difference between the first preset value and the average is used as the element of the sth row and the ath column of the R table; If it cannot be directly transmitted, the second preset value is used as the element of the sth row and the ath column of the R table; Among them, the first preset value is greater than or equal to 1; the second preset value is less than 0.
7. The routing selection method applicable to power communication services according to claim 1, characterized in that: The selecting of the route at each moment according to the Q table updated at each moment and the channel similarity comprises: For any routing node at each moment, obtain the Q value with other routing nodes in the Q table, merge the negative correlation mapping results with the channel similarity of other routing nodes, and calculate the routing control judgment value; The routing node with the smallest routing judgment value with any of the routing nodes is selected as the next routing node for information transmission.
8. A routing system for electric power communication services, characterized in that: The system comprises: A router status acquisition module is used to obtain the node occupancy rate and node bandwidth of each routing node used for data transmission in the power communication business at each moment, as well as the communication delay between routing nodes; A graph search algorithm is used to calculate the communication delay between routing nodes to obtain the shortest communication delay between any two routing nodes and obtain a distance matrix. The distance matrix is reduced in dimension to obtain a set of routing node coordinates. A Q table generation module is used to use the vector formed by the starting points of each simulated link as the communication direction vector in the routing node coordinate set, and to obtain the channel similarity of each routing node corresponding to the channel based on each simulated link by combining the Q learning algorithm with the similarity between the vector formed by each routing node and the routing node directly transmitted to it and the communication direction vector; All routing nodes except for each simulated link are recorded as bandwidth distribution nodes of each simulated link, and the similarity between the vector from the starting point of each simulated link to each bandwidth distribution node and the communication direction vector is analyzed to obtain each communication distribution weight, and the random process weight of each simulated link at each moment is obtained by combining the node bandwidth of all bandwidth distribution nodes at each moment; Based on the Q learning algorithm, according to the channel similarity of the corresponding channels of each routing node selected by each simulated link, combined with the random process weight, the channel learning rate of each channel of each simulated link at each moment is obtained; according to the transmission situation between the routing nodes, combined with the distribution of the node occupancy rate between the routing nodes, an R table is constructed, the channel learning rate is used to replace the original learning rate in the Q learning algorithm, and the Q table is updated; The route selection module is used to select the route at each moment according to the Q table updated at each moment and the channel similarity.
9. A routing device for electric power communication services, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A routing device suitable for electric power communication services, characterized in that: The device stores a computer program, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method for Constructing Energy-efficient Network Content Distribution Mechanism Based on Edge Intelligent Caches
AU2020103384A4
Power communication access network service route planning method based on multi-agent reinforcement learning
CN118524477A