Communication Path Selection Method and Device for Device Node to Transmit Data
By calculating virtual force values based on geographical location and traffic carrying capacity, using the LSTM model to predict path performance index, and dynamically adjust path selection, the problem of inaccurate path selection in the existing technology is solved, and efficient and reliable data transmission is achieved.
Patent Information
- Application Number
- CN202510030056.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-01-08
AI Technical Summary
The prior art lacks full utilization of historical data in data communication networks, resulting in inaccurate path selection and inability to meet the needs of efficient and reliable data transmission. Especially in complex network environments, problems such as inaccurate path transmission performance prediction, data loss or insufficiency in transmission efficiency are prone to problems.
By obtaining the geographical location and traffic carrying capacity of the device node, calculating virtual force values, predicting the path performance index with the LSTM model, dynamically adjusting the path selection, using historical transmission records and real-time channel quality to prioritize the path, and designing a redundant transmission mechanism to ensure the reliability and efficiency of data transmission.
The scientificity and dynamic nature of path selection are realized, the accuracy of path selection and the reliability of data transmission are improved, and data recovery is ensured through redundant paths when a single path fails, ensuring the efficiency and reliability of data transmission.
Smart Images

Figure CN119892705B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of communication paths for device nodes to transmit data, and specifically to a method and device for selecting a communication path for device nodes to transmit data. Background Art
[0002] In a data communication network, selecting an optimal path for data transmission is the core issue for improving network performance. However, traditional path selection methods usually rely only on static evaluations of the geographical distance between nodes, the carrying capacity of nodes, or the channel quality to determine the path. This path selection method based on a single metric lacks comprehensive consideration of the dynamic characteristics of the network, resulting in the selected path being unable to ensure the high efficiency and reliability of data transmission during actual transmission due to problems such as channel quality fluctuations and node overload. In addition, the utilization of historical transmission data in the prior art is limited, staying only at the level of simple recording and statistics, and failing to fully exploit the potential laws in historical data to guide future path selection. Therefore, the prior art has obvious deficiencies in the dynamics and scientific nature of path selection. Especially in a complex network environment, problems such as inaccurate prediction of path transmission performance, data loss, or low transmission efficiency are likely to occur.
[0003] In the prior art, communication path selection methods usually rely on static routing protocols or simple dynamic adjustment mechanisms. For example, traditional shortest path algorithms (such as Dijkstra's algorithm) select paths based on a preset network topology, but this method has limited effectiveness in a dynamic network. Dynamic routing protocols, such as OSPF and AODV, adapt to network changes by periodically updating routing information, but often suffer from problems such as update delays and high overheads. In addition, some path selection methods based on greedy algorithms, although improving a certain efficiency, still have difficulty dealing with frequent topological changes in the face of a complex network environment.
[0004] Through in-depth analysis of the prior art, it can be found that the core difficulty in data transmission path selection lies in how to comprehensively consider the performance of historical paths and the dynamic changes in real-time channel quality, and on this basis, conduct scientific path optimization. However, the existing methods lack a path selection mechanism that can take into account historical priority and real-time channel quality while dynamically evaluating path performance, resulting in inaccurate path selection and being unable to meet the requirements of modern communication networks for efficient and reliable data transmission.
[0005] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure, and thus it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0006] The object of the present invention is to provide a method and device for selecting a communication path for a device node to transmit data, so as to solve the problems raised in the above-mentioned background technology.
[0007] To achieve the above object, the present invention provides the following technical solutions:
[0008] A method for selecting a communication path for a device node to transmit data, the specific steps include:
[0009] Step 1: Based on the device source node and the target node, determine a set of possible paths for transmitting data, and obtain the geographical location information and traffic carrying capacity of each node in the set. The geographical location information is the longitude and latitude coordinate data of the node, and the Haversine formula is used to calculate the geographical distance data between the nodes;
[0010] Step 2: Based on the traffic carrying capacity of the nodes and the geographical distance data between the nodes, calculate the virtual force value of each node, calculate the priority value of each possible path according to the virtual force value of the node, and perform priority sorting on the possible paths based on the priority value;
[0011] Step 3: Obtain the channel quality of the nodes in the historical path at the historical sampling moment, and the transmission record after the historical sampling moment. Generate a data transmission index for the historical path based on the transmission record, construct an LSTM model, and use the channel quality as the input and the data transmission index as the label to train the LSTM model;
[0012] Step 4: Input the channel quality of the nodes in the possible path at the current moment into the LSTM model to obtain the data transmission index of each possible path, and combine the data transmission index and the priority value to calculate the path priority value of the possible path. Sort the possible paths based on the path priority value, and output a path priority list, and use the possible path with the highest path priority value as the optimal path for data transmission.
[0013] Further, the logic for obtaining the geographical location information and traffic carrying capacity of each node in the set is:
[0014] Determine all potential paths from the source node to the target node derived based on the network topology structure:
[0015] Obtain the geographical location information and traffic carrying capacity of each node in the path. The geographical location information is directly measured by a GPS device or extracted from an existing geographical database. Determine the maximum amount of data that each node can forward or process per unit time, and calibrate it as the traffic carrying capacity;
[0016] The logic for calculating the geographical distance data between the nodes using the Haversine formula is:
[0017] Calibrate the longitude and latitude of the first point as (φ1, λ1), and the longitude and latitude of the second point as (φ2, λ2); when calculating the geographical distance between two points, convert the latitude and longitude from degrees to radians, and the conversion formula is:
[0018]
[0019] φ'1 and φ'2 represent the radians after the longitude conversion of the first point and the second point;
[0020]
[0021] λ'1 and λ'2 represent the radians after the latitude conversion of the first point and the second point;
[0022] Then, use the average radius r of the earth to calculate the distance d between two points, where r = 6371 km, and the Haversine formula is:
[0023]
[0024] Among them, Δφ represents the latitude difference, that is, Δφ = φ'2 - φ'1, and Δλ represents the longitude difference, that is, Δλ = λ'2 - λ'1.
[0025] Furthermore, the logic for calculating the virtual force value of each node is:
[0026] Obtain the traffic carrying capacity of each node, which is the maximum data traffic that the node can process per unit time. For each node i, determine the node weight based on the traffic carrying capacity, and the formula is:
[0027]
[0028] Among them, w i is the weight of node i, c i is the traffic carrying capacity of node i, is the sum of the traffic carrying capacities of all nodes, and n is the total number of nodes in the entire network;
[0029] The virtual force value between nodes is defined as the product of the node weight and the reciprocal of the geographical distance, and the formula is:
[0030]
[0031] Among them, w i and w j are the weight data of node i and node j respectively, d ij represents the geographical distance data between node i and node j, F ij represents the virtual force value between node i and node j, ∈ is a constant, and ∈ > 0;
[0032] Create an array G to store the total virtual force value of each node; for each node i, traverse all other nodes j, and accumulate the virtual force value F between node i and other nodes j. Since the virtual force between a node and itself is usually not counted, the case of i = j is ignored. The formula for calculating the total virtual force value G[i] is as follows: ij Accumulate. Since the virtual force between a node and itself is usually not counted, the case of i = j is ignored. The formula for calculating the total virtual force value G[i] is as follows:
[0033]
[0034] Among them, G[i] represents the total virtual force value of node i, and n represents the total number of nodes in the entire network; a larger G[i] indicates that node i has a higher "attraction" or "importance" in the network;
[0035] The logic of calculating the priority value of each possible path based on the virtual force value of the node and performing priority sorting on the possible paths is as follows: Initialize the path priority value P k to zero, where k represents the path number, and k = 1, 2,..., m, and m represents the number of all possible paths;
[0036] Determine the number of nodes n on path k k , traverse each node i on the path, and accumulate the total virtual force value G[i] of each node to the path priority value P k , the formula for calculating the average priority value of the path is as follows:
[0037]
[0038] Among them, n k represents the total number of all nodes on path k, and Nodes(k) represents the set composed of all nodes on path k;
[0039] According to the calculated path priority value P k , for the m paths in all possible path sets, sort them in descending order according to the priority value P k , the formula is as follows:
[0040] P1≥P2≥…≥P m
[0041] Among them, P1 represents the path with the highest priority value, and P m represents the path with the lowest priority value.
[0042] Further, the logic for obtaining the channel quality of nodes in the historical path at the historical sampling moment and the transmission records after the historical sampling moment is as follows: At the historical sampling moment t, extract the channel quality data of each node from the network monitoring system; the channel quality comes from the time window from the historical moment t - T to t - 1, denoted as [t - T, t - 1];
[0043] The channel quality reflects the network performance state before the sampling moment. For each node i on the path P, collect its channel quality data within the historical time window [t - 1, t - T] to form a time series, which is expressed in the form of:
[0044] Q i,h =[q (i,t-T) ,q (i,t-T+1) ,…,q (i,t-1)
[0045] where T is the length of the historical time window of the channel quality;
[0046] The channel quality data of path L is the set of historical time series of the channel quality of all nodes on the path, expressed as:
[0047]
[0048] where n L is the number of nodes in path L;
[0049] For the same historical sampling moment t, the transmission records come from the time window of the historical moment t + Δt, denoted as [t, t + Δt], which is a time range with a total length of Δt, representing the network performance state after the sampling moment;
[0050] Within the historical time window [t, t + Δt], collect the transmission records of each node and record three main indicators: the transmission success rate, and the transmission success rate SR i,h represents the proportion of successful transmissions of node i within the historical time window [t, t + Δt], and the formula is:
[0051]
[0052] where ST i,h is the number of successful transmissions of node i within the historical time window, and TT i,h is the total number of transmission attempts of node i within the historical time window;
[0053] The average delay data, and the average delay data AD i,h reflects the average time required for each data transmission of node i within the historical time window [t, t + Δt], and the formula is:
[0054]
[0055] where N is the number of successful transmissions of node i within the time window, De i,h,o The delay of node i in the o-th transmission;
[0056] Transmission jitter data, transmission jitter data Ji i,h is an important indicator to measure the volatility of delay, representing the standard deviation of delay, and the formula is:
[0057]
[0058] Based on the above three indicators, define the data transmission performance index TI of the historical path of node i i,h , and the formula is:
[0059]
[0060] where w1, w2, and w3 represent the weights of historical transmission success rate, average delay, and jitter data respectively, and w1 + w2 + w3 = 1, w1 > w2, w3;
[0061] For the historical transmission performance index TI of path L L,h , it is the average value of the transmission performance indices of all nodes in path L, and the formula is:
[0062]
[0063] where n L represents the total number of nodes in path L, and TI i,h represents the historical transmission performance index of node i in the path, which is used to measure the comprehensive performance index of path L within the historical time window [t, t+Δt].
[0064] Furthermore, the logic of constructing the LSTM model with channel quality as the input and data transmission index as the label to train the LSTM model is as follows: build a model based on a deep neural network, use the LSTM network architecture to process time series features, set one or two layers of LSTM units, and set each layer to include 50 to 100 LSTM units; use ReLU as the activation function for the hidden layer and a linear activation function for the output layer; use the channel quality at historical sampling moments as input features, and generate the data transmission index of the historical path based on transmission records as the label to train the LSTM model;
[0065] The time series data is divided into an 80% training set and a 20% validation set. Using the batch training method, the entire training set is divided into multiple small batches. In each training iteration, the input data is fed into the LSTM model to calculate its output prediction value. The MSE loss function is used to calculate the difference between the predicted value and the true value of the model. The gradient of the loss function with respect to the model parameters is calculated through the backpropagation algorithm to update the model parameters;
[0066] During the training process, the hyperparameters of the LSTM model are set. The hyperparameters of the LSTM model include: the number of network layers, the number of iterations, the learning rate, the batch size, the number of training times, the batch processing quantity, and the number of neurons in the hidden layer. Among them, the number of network layers is set to a two-layer network structure, the number of iterations is set to 200, the learning rate is set to 0.001, the batch size is set to 64, the number of training times is set to 500, the batch processing quantity is set to 256, and the number of neurons in the hidden layer is 50 for each layer. The model parameters are gradually adjusted to improve the performance, and the RMSE is used to evaluate the model performance until the model training is completed.
[0067] Further, the logic of inputting the channel quality of the nodes in the possible paths at the current moment into the LSTM model to obtain the data transmission index of each possible path is: inputting the current channel quality data of each node in the possible paths into the pre-trained LSTM model, and the model outputs the predicted data transmission index of each possible path where L k represents the k-th possible path;
[0068] The logic of calculating the path priority value of the possible path by combining the data transmission index and the priority value is: based on the obtained data transmission index of each path and the previously calculated path priority value P k , calculate the path priority value of each possible path. The formula is:
[0069]
[0070] where α and β are weighting coefficients used to adjust the influence of real-time transmission performance and historical priority on path selection, and α + β = 1; when real-time is more important, α > β, and when stability is more important, β > α;
[0071] The logic of sorting the possible paths based on the path priority value, outputting the path priority list, and taking the possible path with the highest path priority value as the optimal path for data transmission is: sorting all possible paths in descending order according to the calculated path priority value PV k to generate a path priority list. Define the set of all possible paths as {L1, L2, …, L m}, where each path L kis a possible path from the source node to the target node. There are a total of m paths, and their corresponding path priority values are: {PV1, PV2, …, PV m};
[0072] Arrange all paths in descending order according to the priority values {PV1, PV2, …, PV m} to obtain a path priority list, such that:
[0073] PV (1) ≥ PV (2) ≥ … ≥ PV (m)
[0074] Among them, PV (1) is the path with the highest comprehensive priority value, and PV (m) is the path with the lowest comprehensive priority value;
[0075] After sorting, the path priority list is [L (1) , L (2) , …, L (m) , where L (1) is the path with the highest priority value corresponding to PV (1) , and L (m) is the path with the lowest priority value corresponding to PV (m) ; Select the path L (1) with the highest priority value from the sorted path priority list as the main data transmission path;
[0076] To improve the reliability and efficiency of data transmission, select the first few paths in the path priority list for parallel transmission. In addition to the main transmission path, the selected parallel paths will also contain redundant data. Even if a certain path fails, the data on the redundant path can compensate for the lost information.
[0077] The present invention also provides a communication path device for a device node to transmit data. The system is used to implement the above-mentioned communication path method for a device node to transmit data, and specifically includes:
[0078] A node calculation module, which is used to determine a set of possible paths for data transmission based on the device source node and the target node, and obtain the geographical location information and traffic carrying capacity of each node in the set. The geographical location information is the longitude and latitude coordinate data of the node, and the Haversine formula is used to calculate the geographical distance data between nodes;
[0079] A virtual force value module, which is used to calculate the virtual force value of each node based on the traffic carrying capacity of the node and the geographical distance data between nodes, calculate the priority value of each possible path according to the virtual force value of the node, and perform priority sorting on the possible paths based on the priority value;
[0080] A model building model is used to obtain the channel quality of nodes in the historical path at the historical sampling moment and the transmission records after the historical sampling moment. Based on the transmission records, a data transmission index of the historical path is generated. An LSTM model is constructed, with the channel quality as the input and the data transmission index as the label to train the LSTM model.
[0081] A path selection module is used to input the channel quality of nodes in the possible paths at the current moment into the LSTM model to obtain the data transmission indices of the respective possible paths, and combine the data transmission indices and the priority values to calculate the path priority values of the possible paths. Based on the path priority values, the possible paths are sorted, and a path priority list is output, and the possible path with the highest path priority value is used as the optimal path for data transmission.
[0082] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0083] The present invention uses the geographical location information of nodes, namely longitude and latitude coordinates and traffic carrying capacity, and combines the Haversin formula to calculate the geographical distance, providing accurate basic data for the construction of the path set, and solving the problem of unreasonable path construction caused by the mismatch between node distance and carrying capacity in the background art.
[0084] Based on the virtual force values of nodes, the priority values of possible paths are calculated, and combined with the geographical distribution and traffic carrying capacity of the paths, the preliminary sorting of the paths is completed, solving the problem of lack of scientific basis for path priority calculation in the background art and laying a foundation for subsequent dynamic decision-making; based on the transmission records and channel quality of historical paths, the LSTM model is used to model the path performance, providing a reliable tool for predicting real-time path performance, solving the problem of lack of full utilization of historical data in the background art, and making the path performance prediction more accurate.
[0085] The present invention dynamically calculates the path priority value by combining the real-time data transmission index predicted by the LSTM model and the historical priority value of the path to ensure the scientific nature of path selection. Through the sorting of path priority values, the best path can be quickly screened out, avoiding the problem of low path selection efficiency in the background art. The design of the parallel transmission mechanism enables the system to still rely on redundant paths for data recovery in the case of single-path failure, ensuring the high efficiency and reliability of data transmission, and solving the problems of insufficient path selection dynamics and data loss caused by single-path failure in the background art. BRIEF DESCRIPTION OF THE DRAWINGS
[0086] Figure 1 It is a schematic diagram of the overall method flow of the present invention;
[0087] Figure 2 It is a schematic diagram of the overall device module of the present invention. Detailed implementation manners
[0088] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with specific embodiments.
[0089] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the present invention should be of the ordinary meanings understood by those with ordinary skills in the field to which the present invention pertains. The "first", "second" and similar terms used in the present invention do not denote any order, quantity or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left" and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0090] Embodiment:
[0091] Please refer to Figure 1 , the present invention provides a technical solution:
[0092] A method for selecting a communication path for an equipment node to transmit data, the specific steps include:
[0093] Step 1: Based on the equipment source node and the target node, determine a set of possible paths for transmitting data, and obtain the geographical location information and traffic carrying capacity of each node in the set. The geographical location information is the longitude and latitude coordinate data of the node, and the Haversine formula is used to calculate the geographical distance data between the nodes;
[0094] Determine all potential paths from the source node to the target node derived based on the network topology structure:
[0095] Obtain the geographical location information and traffic carrying capacity of each node in the path. The geographical location information is directly measured by a GPS device or extracted from an existing geographical database, determine the maximum amount of data that each node can forward or process per unit time, and calibrate it as the traffic carrying capacity;
[0096] The set of possible paths contains all potential communication paths from the source node to the destination node, which is the alternative path set for network transmission. The size / number of the set of possible paths reflects the connectivity of the network. The more paths there are, the higher the communication redundancy between the source node and the destination node, and the stronger the fault tolerance and stability of the network. The more complex the topological structure, that is, the more nodes and edges there are, the larger the set of possible paths; conversely, the simpler the topological structure, the smaller the set of possible paths.
[0097] The geographical location information of a node is provided by a GPS device or a database, determined by external factors, and does not depend on the network itself. It is usually determined by the hardware performance of the node, such as bandwidth and computing power. The stronger the hardware performance, the higher the traffic carrying capacity.
[0098] The logic for calculating the geographical distance data between nodes using the Haversine formula is as follows:
[0099] Calibrate the latitude and longitude of the first point as (φ1, λ1), and the latitude and longitude of the second point as (φ2, λ2); when calculating the geographical distance between two points, convert the latitude and longitude from degrees to radians. The conversion formula is:
[0100]
[0101] φ'1 and φ'2 represent the radians after the longitude conversion of the first point and the second point;
[0102]
[0103] λ'1 and λ'2 represent the radians after the latitude conversion of the first point and the second point;
[0104] Converting the angle to radians is for calculating the geographical distance between nodes because the spherical geometry calculation formula, that is, the Haversine formula, requires the input parameters to be in radians. The larger the converted radian value, the closer the geographical location is to the extreme regions of the earth's sphere, such as near the North and South Poles or the International Date Line.
[0105] Then, use the average radius r of the earth to calculate the distance d between two points, where r = 6371 km. The Haversine formula is:
[0106]
[0107] Among them, Δφ represents the latitude difference, that is, Δφ = φ'2 - φ'1, and Δλ represents the longitude difference, that is, Δλ = λ'2 - λ'1;
[0108] The Haversine formula is used to calculate the shortest path between two points on the surface of a sphere, i.e., the great circle distance, which takes into account the spherical geometry of the Earth; d represents the geographical distance between two points (unit: kilometers), Δφ is the difference in latitudes between the two points, reflecting the geographical location difference in the north-south direction, and Δλ is the difference in longitudes between the two points, reflecting the geographical location difference in the east-west direction;
[0109] reflects the spherical distance between two points in the latitude direction. The larger the value, the greater the latitude difference between the two points; cos(φ1)·cos(φ2) reflects the cosine relationship of the latitudes of the two points and is used to correct the influence of the latitude difference on the spherical distance; reflects the spherical distance between two points in the longitude direction. The larger the value, the greater the longitude difference between the two points;
[0110] The larger d is, the farther the distance between the two geographical locations. The smaller d is, the closer the two points are, and they may even coincide.
[0111] Step 2: Based on the traffic-carrying capacity of the nodes and the geographical distance data between the nodes, calculate the virtual force values of each node, calculate the priority values of each possible path according to the virtual force values of the nodes, and perform priority sorting on the possible paths based on the priority values;
[0112] Obtain the traffic-carrying capacity of each node, which is the maximum data traffic that the node can process per unit time. For each node i, determine the node weight based on the traffic-carrying capacity, and the formula used is:
[0113]
[0114] where, w i is the weight of node i, c i is the traffic-carrying capacity of node i, is the sum of the traffic-carrying capacities of all nodes, and n is the total number of nodes in the entire network;
[0115] The traffic-carrying capacity c i of node i represents the maximum data traffic that the node can process per unit time, and w i represents the weight of node i, reflecting the importance of the node relative to other nodes in the entire network. The larger the weight, the higher the proportion of the node's traffic-carrying capacity in the entire network; is the sum of the traffic-carrying capacities of all nodes and is the denominator for normalization to ensure that the weight value w i is between 0 and 1;
[0116] When w i = 1, it means that node i is the only node in the network and undertakes the load of the entire network; when w i= 0 indicates that node i has no traffic - carrying capacity at all; the weight value ranges between 0 and 1, and the larger the value, the stronger the traffic - carrying capacity of the node; the weight value can be used to measure the importance of the node, and nodes with larger weights are more likely to be given priority in the path - selection process;
[0117] w i is a linear function of c i : when the carrying capacity c i of node i increases, the weight w i increases; when c i is 0, the weight w i = 0; w i is inversely proportional to if the total carrying capacity increases while c i remains unchanged, then the weight w i of node i will decrease;
[0118] The logic for calculating the virtual - force value is as follows: The virtual - force value between nodes is defined as the product of the node weight and the reciprocal of the geographical distance, representing that the closer the distance and the greater the traffic capacity, the greater the "attraction" of the node pair to the path. The formula is:
[0119]
[0120] where w i and w j are the weight data of node i and node j respectively, d ij represents the geographical - distance data between node i and node j, and the virtual - force value F ij between nodes is used to measure the "attraction" of node i to node j. ∈ is a constant to prevent the denominator from being zero, and ∈>0;
[0121] ∈ is usually set to a positive number much smaller than the minimum geographical distance d ij : when d ij is in kilometers, ∈ = 0.0001 or smaller can be selected; or ∈ can be set to 1 / 100 of the minimum value among all geographical distances to ensure that ∈ is very small and adapts to the actual distance range, that is, ∈ = min(d ij ) / 100;
[0122] The larger the virtual - force value F ij , the closer the connection between node i and node j. The meaning of this formula is that the "virtual attraction" between two nodes is determined by the product of the node weights and is inversely proportional to the square of the geographical distance; the larger the weight, that is, the stronger the node ability, the larger the virtual - force value F ij ; the smaller the geographical distance, that is, the closer the nodes are to each other, the larger the virtual - force value F ijThe larger it is; the virtual force value is used to measure the connection strength between nodes, and the larger the value, the more important the connection between two nodes is;
[0123] F ij is proportional to w i and w j The higher the weight of the node, the larger the virtual force value; F ij is inversely proportional to d ij When the geographical distance increases, the virtual force value decreases sharply; the role of ∈ is to avoid the denominator being zero. When the distance is very small, that is, d ij →0, ∈ ensures the stability of the calculation;
[0124] Create an array G to store the total virtual force value of each node; for each node i, traverse all other nodes j, and accumulate the virtual force value F ij between node i and other nodes j. Since the virtual force between a node and itself is usually not counted, the case of i = j is ignored. The formula for calculating the total virtual force value G[i] is:
[0125]
[0126] Among them, G[i] represents the total virtual force value of node i, and n represents the total number of nodes in the entire network; a larger G[i] indicates that node i has a higher "attraction" or "importance" in the network;
[0127] The formula represents the sum of all virtual force values of node i, reflecting the overall "attraction" of node i to other nodes; the larger the total virtual force value G[i], the higher the centrality or importance of node i in the network; G[i] is the ij accumulated value of F. When the virtual force value between node i and other nodes is larger, G[i] is larger; if the weight w i of node i increases, or the geographical distance between node i and other nodes decreases, then G[i] increases;
[0128] The logic of calculating the priority value of each possible path based on the virtual force value of the node and performing priority sorting on the possible paths is: initialize the path priority value P k to zero, where k represents the number of the path, and k = 1, 2,..., m, and m represents the number of all possible paths;
[0129] Determine the number of nodes n k on path k, traverse each node i on the path, and accumulate the total virtual force value G[i] of each node to the path priority value P k The formula for calculating the average priority value of the path is:
[0130]
[0131] Among them, n k represents the total number of all nodes on path k, Nodes(k) represents the set composed of all nodes on path k, and the calculated P k The larger the value, the more prioritized the path is;
[0132] The path priority is the average value of the virtual force values of all nodes on the path, which is used to measure the importance of the path. The higher the priority value, the more prioritized the path. P k is the average value of G[i]: The higher the total virtual force value G[i] of the nodes on the path, the higher the path priority P k The higher it is, the more nodes n k there are on the path, and the priority value P k will decrease relatively;
[0133] According to the calculated path priority value P k , for the m paths in all possible path sets, sort them in descending order according to the priority value P k , and the formula based on which is:
[0134] P1≥P2≥…≥P m
[0135] Among them, P1 represents the path with the highest priority value, and P m represents the path with the lowest priority value.
[0136] Step 3: Obtain the channel quality of the nodes in the historical path at the historical sampling moment, as well as the transmission records after the historical sampling moment, generate the data transmission index of the historical path based on the transmission records, construct an LSTM model, use the channel quality as the input, and the data transmission index as the label to train the LSTM model;
[0137] The logic of obtaining the channel quality of the nodes in the historical path at the historical sampling moment, as well as the transmission records after the historical sampling moment is: At the historical sampling moment t, extract the channel quality data of each node from the network monitoring system; The channel quality comes from the time window from historical moment t - T to t - 1, denoted as [t - T, t - 1], which is a historical time window with a total length of T, representing the network state before the historical sampling moment t;
[0138] The channel quality reflects the network performance state before the sampling moment. For each node i on path P, collect its channel quality data within the historical time window [t - 1, t - T] to form a time series, and its form is expressed as:
[0139] Q i,h =[q (i,t-T) ,q (i,t-T+1) ,…,q(i,t-1 )]
[0140] Among them, T is the length of the historical time window of the channel quality;
[0141] Q i,h reflects the change of the channel quality of node i over a period of time, indicating the stability and reliability of the node in transmitting data during this period. The channel quality time series Q i,h is a data set at multiple moments, showing the performance state of node i within the time window; by observing the channel quality time series, it can be judged whether the performance of the node is stable, whether there are fluctuations, and whether it can continuously maintain high-quality communication; Q i,h is a set of q (i,t) The higher the channel quality q (i,t) is, the higher the average value in the sequence may be;
[0142] The larger the historical time window T is, the longer the length of the sequence increases, indicating the data quality over a longer time, but it may reduce the sensitivity to short-term trends; when the channel quality changes violently, that is, the variance is large, shorten T, and when the channel quality changes smoothly, that is, the variance is small, extend T; during the model training and verification process, test multiple different T values, such as T = 5, 10, 20, 30, and select the T with the optimal model performance through the validation set;
[0143] The channel quality data of path L is the set of the channel quality historical time series of all nodes on the path, denoted as:
[0144]
[0145] Among them, n L is the number of nodes in path L;
[0146] Q L,h reflects the overall channel quality state of path L, which is the set of the channel quality of all nodes on the path. The path channel quality time series represents the overall communication state of the path, summarizing the channel quality of each node in the path. By aggregating the channel quality of all nodes in the path, the stability and reliability of the path can be evaluated; Q L,h is a set of Q i,h and is proportional to the number of nodes n L in the path; if the channel quality of the nodes in the path fluctuates greatly, the overall channel quality of the path will also fluctuate greatly;
[0147] For the same historical sampling moment t, the transmission record comes from the time window of historical moment t + Δt, denoted as [t, t + Δt], which is a time range with a total length of Δt, indicating the network performance state after the sampling moment;
[0148] Within the historical time window [t, t + Δt], collect the transmission records of each node and record three main metrics: the transmission success rate, and the transmission success rate SR i,h It represents the proportion of successful transmissions of node i within the historical time window [t, t + Δt], and the formula is as follows:
[0149]
[0150] where ST i,h is the number of successful transmissions of node i within the historical time window, and TT i,h is the total number of transmission attempts of node i within the historical time window; when ST i,h increases, SR i,h increases, and when TT i,h increases but ST i,h remains unchanged, SR i,h decreases
[0151] SR i,h represents the transmission success rate of node i, and its value ranges from [0, 1]. The transmission success rate is the proportion of the number of successful transmissions to the total number of transmission attempts. The higher the success rate, the better the communication quality of the node and the more reliable the transmission link; SR i,h = 1 indicates that all transmission attempts of the node are successful and the communication quality is ideal; SR i,h = 0 indicates that the node has no successful transmissions;
[0152] Average delay data, average delay data AD i,h reflects the average time required for each data transmission of node i within the historical time window [t, t + Δt], and the formula is as follows:
[0153]
[0154] where N is the number of successful transmissions of node i within the time window, and De i,h,o is the delay of node i in the o-th transmission;
[0155] AD i,h reflects the average transmission delay of node i. The average delay represents the average time-consuming during node transmission. The smaller the value, the higher the transmission efficiency of the node; AD i,h is the average value of De i,h,o If the single delay increases, the average delay increases. If the number of successful transmissions N increases but the delay value remains stable, the average delay remains unchanged;
[0156] Transmission jitter data, transmission jitter data Ji i,h is an important metric to measure the volatility of delay, representing the standard deviation of delay, and the formula is as follows:
[0157]
[0158] Ji i,h Reflects the transmission delay volatility of node i. Transmission jitter is the standard deviation of the delay, indicating the degree of delay fluctuation. The smaller the jitter, the more stable the node's delay; Ji i,h Changes with De i,h,o and AD i,h increases as the difference between them increases. If the delay fluctuation is large, then Ji i,h tends to increase, indicating unstable network quality;
[0159] Based on the above three metrics, define the data transmission performance index TI of the historical path of node i i,h , and the formula is:
[0160]
[0161] where w1, w2, and w3 represent the weights of the historical transmission success rate, average delay, and jitter data respectively, and w1 + w2 + w3 = 1, w1 > w2, w3;
[0162] TI i,h Comprehensively measures the data transmission performance of node i. By using the weighted method, it comprehensively considers the success rate, delay, and jitter to evaluate the transmission performance of the node. The higher the performance index, the better the transmission performance of the node; TI i,h Changes with SR i,h and increases as SR increases. TI i,h Changes with AD i,h and Ji i,h and increases as they decrease;
[0163] For the historical transmission performance index TI of path L L,h , it is the average of the transmission performance indices of all nodes in path L, and the formula is:
[0164]
[0165] where n L represents the total number of nodes in path L, and TI i,h represents the historical transmission performance index of node i in the path, which is used to measure the comprehensive performance index of path L within the historical time window [t, t + Δt]. The higher the value, the more excellent the comprehensive performance of the path, and the more reliable and efficient the transmission;
[0166] This formula calculates the overall transmission performance of path L by averaging the transmission performance indices of all nodes on the path. The higher the performance index TI L,h of all nodes, the better the overall performance of the path; If TI L,hIf it is relatively high, this path can be preferentially selected for data transmission; TI L,h The larger the value is, the higher the average transmission performance of all nodes on the path is, and the higher the overall communication ability of the path is. If TI L,h is close to the maximum value, it indicates that the path has an extremely high transmission success rate, a relatively low average delay, and a small jitter;
[0167] The logic of constructing the LSTM model with the channel quality as the input and the data transmission index as the label to train the LSTM model is as follows: establish a model based on a deep neural network, use the LSTM network architecture to process time series features, set one or two layers of LSTM units, and set each layer to include 50 to 100 LSTM units; use ReLU as the activation function for the hidden layer and a linear activation function for the output layer; use the channel quality at historical sampling moments as input features, and generate the data transmission index of the historical path based on the transmission record as the label to train the LSTM model;
[0168] For each path L, the input feature is the time series Q of the channel quality of the path nodes L,g , where the channel quality reflects the communication performance of the path in the historical time window [t, t+Δt], and the target output is the transmission performance index TI L,h of path L, which is obtained by averaging the transmission performance indices of each node in the path, and the label reflects the overall performance of the path within the historical time window [t, t+Δt];
[0169] Divide the time series data into an 80% training set and a 20% validation set, and use the Batch training method. Divide the entire training set into multiple small batches. In each training iteration, input the data into the LSTM model, calculate its output prediction value, use the MSE loss function to calculate the difference between the predicted value and the true value of the model, and calculate the gradient of the loss function with respect to the model parameters through the backpropagation algorithm to update the model parameters;
[0170] During the training process, set the hyperparameters of the LSTM model. The hyperparameters of the LSTM model include: the number of network layers, the number of iterations, the learning rate, the batch size, the number of training times, the batch processing quantity, and the number of neurons in the hidden layer; among them, the number of network layers is set to a two-layer network structure, the number of iterations is set to 200, the learning rate is set to 0.001, the batch size is set to 64, the number of training times is set to 500, the batch processing quantity is set to 256, and the number of neurons in the hidden layer is 50 for each layer; gradually adjust the model parameters to improve the performance, and use RMSE to evaluate the model performance until the model training is completed;
[0171] During the training process, the loss function value of the model, such as MSE, will gradually decrease as the training progresses and tend to stabilize. If, in the last several iterations of training, the change amplitude of the loss function is very small, less than the set threshold, such as 10 -4 , it can be considered that the model has converged; if the loss function of the model tends to stabilize on both the training set and the validation set, and the loss of the validation set does not increase significantly, that is, there is no overfitting, it can be considered that the model has been trained;
[0172] Suppose a two-layer LSTM model is constructed based on the given logic to predict the transmission performance index TI of path L L,h , and the hyperparameters of the model are set as follows: number of network layers, a two-layer LSTM network, 50 neurons in each layer, number of iterations 200, learning rate 0.001, batch size 64, number of batches 256; using the historical channel quality time series Q L,h as the input feature and the data transmission index TI L,h as the label to train the model and evaluate its performance;
[0173] According to the calculation results, it can be obtained that after the number of training rounds reaches 150, the loss function values (MSE) of the training set and the validation set decrease significantly slower and tend to stabilize after 200 iterations. The training set MSE decreases from 0.015 to 0.012, and the validation set MSE decreases from 0.017 to 0.014, with a change amplitude <0.01, indicating that the loss function of the model has basically converged; the RMSE of the validation set decreases from the initial 0.205 to 0.118, reaching a relatively low error level. If the expected RMSE target in the business scenario is less than 0.12, then the current model meets the task requirements.
[0174] Step 4: Input the channel quality of the nodes in the possible paths at the current moment into the LSTM model to obtain the data transmission indices of each possible path, and combine the data transmission index and the priority value to calculate the path priority value of the possible paths. Sort the possible paths based on the path priority value and output the path priority list, and use the possible path with the highest path priority value as the optimal path for data transmission;
[0175] The logic of inputting the channel quality of the nodes in the possible paths at the current moment into the LSTM model to obtain the data transmission indices of each possible path is: input the current channel quality data of each node in the possible path into the pre-trained LSTM model, and the model outputs the predicted data transmission index of each possible path where L k represents the k-th possible path, and the formula is:
[0176]
[0177] Among them, represents the channel quality data of all nodes on path L k , which may be characteristics such as channel bandwidth, delay, packet loss rate, etc. f(·) is the mapping function of the LSTM model, used to process the channel quality data and output the transmission performance index of the path;
[0178] represents the data transmission performance of path L k at the current moment, predicting the transmission reliability and efficiency of the path, The larger it is, the better the channel quality of path L k is, and it can support more efficient and reliable data transmission; the better the channel quality data, such as the higher the bandwidth, the lower the delay, and the lower the packet loss rate, the larger the predicted by the model; the LSTM model f(·) will learn the non-linear relationship between channel quality and transmission performance;
[0179] Combining the data transmission index and the priority value, the logic for calculating the path priority value of possible paths is: based on the data transmission index of each path obtained and the previously calculated path priority value P k , calculate the path priority value of each possible path, and the formula is:
[0180]
[0181] Among them, α and β are weighting coefficients, used to adjust the influence of real-time transmission performance and historical priority on path selection, and α + β = 1; PV k The larger it is, the better the path performs in the comprehensive evaluation, and this path can be preferentially selected for data transmission; is the real-time transmission performance index of path L k , reflecting the quality of its current channel, The larger it is, the larger the path PV k is, indicating that the real-time performance of the path makes an important contribution to the final priority; P k is the historical priority of path L k , which may be determined by the virtual force value of the nodes in the path or other historical factors. The larger P k is, the larger the path PV k is, indicating that the historical priority of the path is relatively high; the dynamic change of PV k can reflect the real-time fluctuation of channel quality and the long-term accumulation of path historical priority;
[0182] α determines the influence weight of real-time performance on PV k , and β determines the historical priority Pk The influence weight on PV k ; when α > β, the system pays more attention to the real-time transmission performance of the path, and when β > α, the system pays more attention to historical priority;
[0183] The logic of sorting the possible paths based on the path priority value, outputting the path priority list, and taking the possible path with the highest path priority value as the optimal path for data transmission is as follows: Sort all possible paths in descending order according to the calculated path priority value PV k to generate a path priority list. Define the set of all possible paths as {L1, L2, …, L m}, where each path L k is a possible path from the source node to the target node, with a total of m paths, and their corresponding path priority values are: {PV1, PV2, …, PV m};
[0184] Sort all paths in descending order according to the priority values {PV1, PV2, …, PV m} to obtain a path priority list, such that:
[0185] PV (1) ≥ PV (2) ≥ … ≥ PV (m)
[0186] where PV (1) is the path with the highest comprehensive priority value, and PV (m) is the path with the lowest comprehensive priority value;
[0187] After sorting, the path priority list is [L (1) , L (2) , …, L (m) , where L (1) is the path with the highest priority value corresponding to PV (1) , and L (m) is the path with the lowest priority value corresponding to PV (m) ; Select the path L (1) with the highest priority value from the sorted path priority list as the main data transmission path;
[0188] To improve the reliability and efficiency of data transmission, select the first few paths in the path priority list for parallel transmission. In addition to the main transmission path, the selected parallel paths will also include redundant data, so that even if a certain path fails, the data on the redundant path can compensate for the lost information;
[0189] Select the first few paths from the priority list, including L (1), as a parallel transmission path, partial data is transmitted in the redundant path to improve the robustness of the system. Even if the main path fails, the redundant path can compensate for data loss, and the number of redundant paths can be dynamically adjusted according to network resources and reliability requirements.
[0190] Please refer to Figure 2 , the present invention also provides a communication path device for a device node to transmit data. The system is used to implement the above-mentioned communication path method for a device node to transmit data, and specifically includes:
[0191] A node calculation module, configured to determine a set of possible paths for data transmission based on a device source node and a target node, and obtain the geographical location information and traffic carrying capacity of each node in the set. The geographical location information is the longitude and latitude coordinate data of the node, and the Haversine formula is used to calculate the geographical distance data between nodes;
[0192] A virtual force value module, configured to calculate the virtual force value of each node based on the traffic carrying capacity of the node and the geographical distance data between nodes, calculate the priority value of each possible path according to the virtual force value of the node, and perform priority sorting on the possible paths based on the priority value;
[0193] A model building model, configured to obtain the channel quality of the nodes in the historical path at the historical sampling moment and the transmission record after the historical sampling moment, generate a data transmission index of the historical path based on the transmission record, construct an LSTM model, and use the channel quality as the input and the data transmission index as the label to train the LSTM model;
[0194] A path selection module, configured to input the channel quality of the nodes in the possible paths at the current moment into the LSTM model to obtain the data transmission index of each possible path, combine the data transmission index and the priority value, calculate the path priority value of the possible paths, sort the possible paths based on the path priority value, and output a path priority list, and use the possible path with the highest path priority value as the optimal path for data transmission.
[0195] The above formulas are all dimensionless and take their numerical calculations. The formula is obtained by collecting a large amount of data for software simulation to obtain a formula closest to the actual situation. The preset parameters in the formula are set by those skilled in the art according to the actual situation.
[0196] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those skilled in the art will appreciate that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed by hardware or software methods depends on the specific application and design constraints of the technical solution.
[0197] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units. They may be located in one place or distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment. As described above, this is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application.
Claims
1. A method for selecting a communication path for a device node to transmit data, characterized in that, The specific steps include: Step 1: Based on the source node and the target node of the device, determine the set of possible paths for data transmission, and obtain the geographical location information and traffic carrying capacity of each node in the set. The geographical location information is the longitude and latitude coordinate data of the node, and use the Haversine formula to calculate the geographical distance data between nodes; Step 2: Based on the traffic carrying capacity of the nodes and the geographical distance data between nodes, calculate the virtual force value of each node, calculate the priority value of each possible path according to the virtual force value of the node, and perform priority sorting on the possible paths based on the priority value; Step 3: Obtain the channel quality of the nodes in the historical path at the historical sampling moment, and the transmission record after the historical sampling moment. Generate the data transmission index of the historical path based on the transmission record, construct an LSTM model, use the channel quality as the input, and the data transmission index as the label to train the LSTM model; Step 4: Input the channel quality of the nodes in the possible path at the current moment into the LSTM model to obtain the data transmission index of each possible path, combine the data transmission index and the priority value, calculate the path priority value of the possible path, sort the possible paths based on the path priority value, and output the path priority list, and use the possible path with the highest path priority value as the optimal path for data transmission; The logic for calculating the virtual force value of each node is: Obtain the traffic carrying capacity of each node, which is the maximum data traffic that the node can process per unit time. For each node i, determine the node weight based on the traffic carrying capacity. The formula is: where, w i is the weight of node i, c i is the traffic-carrying capacity of node i, is the sum of the traffic-carrying capacities of all nodes, and n is the total number of nodes in the entire network; The virtual force value between nodes is defined as the product of the node weight and the reciprocal of the geographical distance. The formula is: where, w i and w j are the weight data of node i and node j respectively, d ij represents the geographical distance data between node i and node j, F ij represents the virtual force value between node i and node j, ∈ is a constant and ∈ > 0; Create an array G to store the total virtual force value of each node; for each node i, traverse all other nodes j, and accumulate the virtual force value F between node i and other nodes j. Since the virtual force between a node and itself is usually not counted, the case of i = j is ignored. The formula for calculating the total virtual force value G[i] is as follows: ij Accumulate. Since the virtual force between a node and itself is usually not counted, the case of i = j is ignored. The formula for calculating the total virtual force value G[i] is as follows: Where G[i] represents the total virtual force value of node i, and n represents the total number of nodes in the entire network; a larger G[i] indicates that node i has a higher "attraction" or "importance" in the network.
2. The communication path selection method for a device node to transmit data according to claim 1, wherein: The logic for obtaining the geographical location information and traffic carrying capacity of each node in the set is: Determine all potential paths from the source node to the target node derived based on the network topology structure: Obtain the geographical location information and traffic carrying capacity of each node in the path. The geographical location information is directly measured by a GPS device or extracted from an existing geographical database. Determine the maximum data volume that each node can forward or process per unit time, and calibrate it as the traffic carrying capacity; The logic for calculating the geographical distance data between nodes using the Haversine formula is: Calibrate the longitude and latitude of the first point as (φ1, λ1), and the longitude and latitude of the second point as (φ2, λ2); when calculating the geographical distance between two points, convert the latitude and longitude from degrees to radians. The conversion formula is: φ'1 and φ'2 represent the radians after the longitude conversion of the first point and the second point; λ'1 and λ'2 represent the radians after the latitude conversion of the first point and the second point; Then use the average radius r of the earth to calculate the distance d between the two points, where r = 6371 km, and the Haversine formula is: Among them, Δφ represents the latitude difference, that is, Δφ = φ'2 - φ'1, and Δλ represents the longitude difference, that is, Δλ = λ'2 - λ'1.
3. A method for selecting a communication path for data transmission of a device node according to claim 2, characterized in that: The logic for calculating the priority value of each possible path based on the node-based virtual force value and prioritizing the possible paths based on the priority value is as follows: Initialize the path priority value P k to zero, where k represents the path number and k = 1, 2, …, m, and m represents the number of all possible paths; Determine the number of nodes \(n\) on path \(k\). k Traverse each node \(i\) on the path, and accumulate the total virtual force value \(G[i]\) of each node to the path priority value \(P\). k The formula for calculating the average priority value of the path is as follows: where n k represents the total number of all nodes on path k, and Nodes(k) represents the set composed of all nodes on path k; According to the calculated path priority value P k , for the m paths in all possible path sets, according to the priority value P k perform a descending sort, and the formula based on is: P1≥P2≥…≥P m Among them, P1 represents the path with the highest priority value, and P m represents the path with the lowest priority value.
4. The communication path selection method for the device node to transmit data according to claim 3, wherein: The logic of obtaining the channel quality of nodes in the historical path at the historical sampling moment and the transmission record after the historical sampling moment is as follows: at the historical sampling moment t, extract the channel quality data of each node from the network monitoring system; the channel quality comes from the time window from historical moment t - T to t - 1, denoted as [t - T, t - 1]; The channel quality reflects the network performance state before the sampling moment. For each node i on the path P, collect its channel quality data within the historical time window [t - 1, t - T] to form a time series, and its form is expressed as: Q i,h = [q (i,t-T) , q (i,t-T+1) , …, q (i,t-1) Among them, T is the length of the historical time window of the channel quality; The channel quality data of path L is the set of historical time series of the channel quality of all nodes on the path, expressed as: where n L is the number of nodes in path L; For the same historical sampling moment t, the transmission record comes from the time window of historical moment t + Δt, denoted as [t, t + Δt], which is a time range with a total length of Δt, representing the network performance state after the sampling moment; Within the historical time window [t, t + Δt], collect the transmission records of each node and record three main metrics: the transmission success rate, the transmission success rate SR i,h Indicates the proportion of successful transmissions of node i within the historical time window [t, t + Δt], and the formula is as follows: Among them, ST i,h is the number of successful transmissions of node i within the historical time window, and TT i,h is the total number of transmission attempts of node i within the historical time window; Average delay data, average delay data AD i,h It reflects the average time required for each data transmission of node i within the historical time window [t, t+Δt], and the formula is as follows: where N is the number of successful transmissions of node i within the time window, and De i,h,o is the delay of the o-th transmission of node i; Transmit jitter data, transmit jitter data ji i,h Is an important indicator to measure the volatility of latency, representing the standard deviation of latency, and the formula is as follows: Based on the above three metrics, define the data transmission performance index TI of the historical path of node i i,h , and the formula is as follows: Among them, w1, w2, and w3 respectively represent the weights of the historical transmission success rate, average delay, and jitter data, and w1 + w2 + w3 = 1, w1 > w2, w3; For the historical transmission performance index TI of path L L,h , which is the average of the transmission performance indices of all nodes in path L, and the formula is as follows: where n L represents the total number of nodes in path L, and TI i,h represents the historical transmission performance index of node i in the path, which is used to measure the comprehensive performance index of path L within the historical time window [t, t+Δt].
5. A method for selecting a communication path for data transmission of a device node according to claim 4, characterized in that: The logic of constructing the LSTM model, using the channel quality as the input and the data transmission index as the label to train the LSTM model is as follows: establish a model based on a deep neural network, use the LSTM network architecture to process time series features, set one to two layers of LSTM units, and set each layer to include 50 to 100 LSTM units; use ReLU as the activation function for the hidden layer and the linear activation function for the output layer; use the channel quality at the historical sampling moment as the input feature, and generate the data transmission index of the historical path based on the transmission record as the label to train the LSTM model; Divide the time series data into an 80% training set and a 20% validation set, use the Batch training method, divide the entire training set into multiple small batches, in each training iteration, input the data into the LSTM model, calculate its output prediction value, use the MSE loss function to calculate the difference between the prediction value and the true value of the model, and calculate the gradient of the loss function with respect to the model parameters through the backpropagation algorithm to update the model parameters; During the training process, set the hyperparameters of the LSTM model. The hyperparameters of the LSTM model include: the number of network layers, the number of iterations, the learning rate, the batch size, the number of training times, the batch processing quantity, and the number of neurons in the hidden layer. Among them, the number of network layers is set to a two-layer network structure, the number of iterations is set to 200, the learning rate is set to 0.001, the batch size is set to 64, the number of training times is set to 500, the batch processing quantity is set to 256, and the number of neurons in the hidden layer is 50 for each layer. Gradually adjust the model parameters to improve performance, and use RMSE to evaluate the model performance until the model training is completed.
6. A method for selecting a communication path for an equipment node to transmit data according to claim 5, characterized in that: The logic of inputting the channel quality of nodes in possible paths at the current moment into the LSTM model to obtain the data transmission indices of each possible path is as follows: Input the current channel quality data of each node in the possible paths into the pre-trained LSTM model, and the model outputs the predicted data transmission indices of each possible path. where L k represents the k-th possible path; The logic for calculating the path priority value of possible paths by combining the data transmission index and the priority value is as follows: Based on the data transmission index of each obtained path and the previously calculated path priority value P k , calculate the path priority value of each possible path, and the formula is as follows: Among them, α and β are weighting coefficients used to adjust the influence of real-time transmission performance and historical priority on path selection, and α + β = 1. When real-time is more important, α > β; when stability is more important, β > α. The logic of sorting the possible paths based on the path priority value, outputting a path priority list, and taking the possible path with the highest path priority value as the optimal path for data transmission is as follows: all possible paths are sorted in descending order according to the calculated path priority value PV k to generate a path priority list. Define the set of all possible paths as {L1, L2, …, L m}, where each path L k is a possible path from the source node to the destination node, with a total of m paths, and their corresponding path priority values are: {PV1, PV2, …, PV m}; Arrange all paths in descending order according to the priority values {PV1, PV2, …, PV m}, obtaining a path priority list such that: PV (1) ≥ PV (2) ≥ … ≥ PV (m) Among them, PV (1) is the path with the highest comprehensive priority value, and PV (m) is the path with the lowest comprehensive priority value; The sorted path priority list is [L (1) , L (2) , …, L (m) , where L (1) is the path with the highest priority value corresponding to PV (1) , and L (m) is the path with the lowest priority value corresponding to PV (m) ; Select the path L (1) with the highest priority value from the sorted path priority list as the main data transmission path; To improve the reliability and efficiency of data transmission, select the first few paths in the path priority list for parallel transmission. In addition to the main transmission path, the selected parallel paths will also include redundant data. Even if a certain path fails, the data on the redundant path can compensate for the lost information.
7. A communication path device for a device node to transmit data, characterized in that: The device is used to execute the communication path selection method for an equipment node to transmit data according to any one of claims 1-6, including: A node calculation module, configured to determine a set of possible paths for data transmission based on the device source node and the target node, and obtain the geographical location information and traffic carrying capacity of each node in the set. The geographical location information is the longitude and latitude coordinate data of the node, and the Haversine formula is used to calculate the geographical distance data between nodes. A virtual force value module, configured to calculate the virtual force value of each node based on the traffic carrying capacity of the node and the geographical distance data between nodes, calculate the priority value of each possible path according to the virtual force value of the node, and perform priority sorting on the possible paths based on the priority value. A model building module, configured to obtain the channel quality of the nodes in the historical path at the historical sampling moment and the transmission record after the historical sampling moment, generate a data transmission index of the historical path based on the transmission record, build an LSTM model, and use the channel quality as the input and the data transmission index as the label to train the LSTM model. A path selection module, configured to input the channel quality of the nodes in the possible paths at the current moment into the LSTM model to obtain the data transmission index of each possible path, combine the data transmission index and the priority value, calculate the path priority value of the possible paths, sort the possible paths based on the path priority value, and output a path priority list, and use the possible path with the highest path priority value as the optimal path for data transmission. The logic for calculating the virtual force value of each node is: Obtain the traffic carrying capacity of each node, which is the maximum data traffic that the node can process per unit time. For each node i, determine the node weight based on the traffic carrying capacity, and the formula is: where w i is the weight of node i, c i is the traffic carrying capacity of node i, is the sum of the traffic carrying capacities of all nodes, and n is the total number of nodes in the entire network; The virtual force value between nodes is defined as the product of the node weight and the reciprocal of the geographical distance, and the formula is: where, w i and w j are the weight data of node i and node j respectively, d ij represents the geographical distance data between node i and node j, F ij represents the virtual force value from node i to node j, ∈ is a constant and ∈ > 0; Create an array G to store the total virtual force value of each node; for each node i, traverse all other nodes j, and accumulate the virtual force value F between node i and other nodes j. Since the virtual force between a node and itself is usually not counted, the case of i = j is ignored. The formula for calculating the total virtual force value G[i] is as follows: ij Accumulate, because the virtual force between a node and itself is usually not counted, so the case of i = j is ignored. The formula for calculating the total virtual force value G[i] is: Among them, G[i] represents the total virtual force value of node i, and n represents the total number of nodes in the entire network; a larger G[i] indicates that node i has a higher "attraction" or "importance" in the network.
Citation Information
Patent Citations
Optical network routing optimization method based on LSTM deep learning and related device thereof
CN112560204A
Unmanned aerial vehicle positioning navigation and path planning optimization method based on machine learning
CN119105524A