Multi-path redundant communication control method and system, storage medium and electronic equipment
Through the multi-path redundant communication control method, the reinforcement learning algorithm is used to select the optimal backup communication path, which solves the problems of single communication and low reliability of traditional column tail devices, and improves communication stability and reliability.
Patent Information
- Application Number
- CN202510148328.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-23
AI Technical Summary
The communication method of traditional column tail devices is single and susceptible to environmental interference, resulting in communication fluctuations and low reliability, and it is impossible to report column tail information in real time.
The multi-path redundant communication control method is adopted to detect the original communication path abnormality and use reinforcement learning algorithm to select the optimal backup communication path for switching based on the network status and historical expected values of the backup communication path to ensure stable communication.
It realizes automatic switching to the optimal backup communication path when fluctuations or interference occur during communication, ensuring the stability and reliability of the column tail device and ensuring real-time information reporting.
Smart Images

Figure CN120034836A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of train communications, and in particular relates to a multi-path redundant communication control method, system, storage medium and electronic equipment. Background Art
[0002] Traditional tail equipment is a special transportation safety device installed at the rear of a train to improve the safety of railway transportation when the guard car and the operating conductor are removed from freight or passenger trains and the rear of the train is unmanned. The current commonly used railway freight train integrity detection method is to install a traditional tail device at the rear of the train, communicate with the locomotive station at the head of the train to report the tail pressure of the train air duct to the driver, and then the driver manually judges the integrity of the train.
[0003] Traditional tail-of-train equipment only uses 400M analog or digital radio as a communication method, with a single detection method and low reliability. When affected by the environment (such as signal interference), communication fluctuations are prone to occur, and the tail-of-train information cannot be reported in real time. It has low reliability and poor stability. Summary of the invention
[0004] The present invention provides a multi-path redundant communication control method, system, storage medium and electronic equipment.
[0005] A multi-path redundant communication control method of the present invention comprises:
[0006] Determine the original communication path is abnormal through detection;
[0007] Using a reinforcement learning algorithm, according to the network status of the backup communication path and the historical expected value of the backup communication path, an optimal backup communication path is selected, and the original communication path is switched to the optimal backup communication path;
[0008] The optimal backup communication path is switched back to the original communication path according to a predetermined rule.
[0009] Further,
[0010] Using the reinforcement learning algorithm, the optimal backup communication path is selected according to the network status of the backup communication path and the historical expected value of the backup communication path, including:
[0011] Define state space and action space;
[0012] A plurality of state-action pairs are defined according to the network state in the state space and the action in the action space, each state-action pair has a corresponding Q value, and all the state-action pairs and all the Q values form a Q table;
[0013] Based on the Q table, select the best action from all the actions corresponding to the current network state, and switch the original communication path to the backup communication path corresponding to the best action;
[0014] Record a network status;
[0015] calculating a total reward after switching to the backup communication path;
[0016] Based on the next network state and the total reward, the current network state and the Q value corresponding to the best action are updated using the Bellman equation.
[0017] Further,
[0018] Using a reinforcement learning algorithm, selecting an optimal backup communication path according to a network state of the backup communication path and a historical expected value of the backup communication path, further comprising:
[0019] A predetermined number of iterations or an iteration target is set, and the Q value is continuously updated until the predetermined number of iterations or the iteration target is reached, and a maximum Q value is selected from historical Q values, where the maximum Q value corresponds to the optimal backup communication path.
[0020] Further,
[0021] The network status includes bandwidth, delay, packet loss rate and load balancing.
[0022] Further,
[0023] Based on the Q table, selecting the best action from all the actions corresponding to the current network state, and switching the original communication path to the backup communication path corresponding to the best action, includes:
[0024] When the Q value of each of the actions corresponding to the current network state is the same, randomly selecting an action from all the actions corresponding to the current network state as the optimal action;
[0025] When the Q value of each of the actions corresponding to the current network state is not completely the same, the action with the largest Q value is selected from all the actions corresponding to the current network state as the optimal action.
[0026] Further,
[0027] The Bellman equation is:
[0028] Q(s,a)←Q(s,a)+α[r+γmax a′ Q(s′,a′)-Q(s,a)],
[0029] Where r is the total reward; α is the learning rate; γ is the discount factor; s′ is the next network state; a′ is all possible actions; max a′ Q(s′,a′) represents the maximum Q value corresponding to all actions in the next network state.
[0030] Further,
[0031] Switching the optimal backup communication path back to the original communication path according to a predetermined condition includes:
[0032] After the original communication path is restored to an available state, it is monitored that no abnormality occurs on the original communication path within a predetermined time period, and the optimal backup communication path is switched back to the original communication path.
[0033] Further,
[0034] Switching the optimal backup communication path back to the original communication path according to a predetermined condition also includes:
[0035] The optimal backup communication path is continuously monitored, and when an abnormality is detected in the optimal backup communication path, the optimal backup communication path is switched back to the original communication path.
[0036] A multi-path redundant communication control system of the present invention is used to implement the aforementioned multi-path redundant communication control method, and the system comprises:
[0037] A detection module, used to determine whether the original communication path is abnormal through detection;
[0038] A first switching module is used to select an optimal backup communication path by using a reinforcement learning algorithm according to a network status of the backup communication path and a historical value of the backup communication path, and switch the original communication path to the optimal backup communication path;
[0039] The second switching module is used to switch the optimal backup communication path back to the original communication path according to a predetermined rule.
[0040] A computer-readable storage medium of the present invention stores a program or instruction. When the program or instruction is executed on a computer, the computer executes the aforementioned multi-path redundant communication control method.
[0041] An electronic device of the present invention comprises a processor, wherein the processor is coupled to a memory; the processor is used to read and execute a computer program stored in the memory to implement the aforementioned multi-path redundant communication control method.
[0042] Compared with the prior art, the present invention has the following beneficial effects:
[0043] The integrated tail-of-train equipment used in the heavy-load train control system adds 4G and 5G communication paths on the basis of the traditional tail-of-train equipment to realize communication between the head and the tail of the train. The multi-path redundant communication control method of the present invention ensures that when the current line fluctuates or is disturbed during the communication process, a reinforcement learning algorithm is used to select a backup line based on the current network status and historical network records, and seamlessly switch to the line without data loss, thereby ensuring the stability and reliability of the tail-of-train equipment. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0045] Figure 1 This is a schematic diagram of the communication method between the first and last columns of the present invention;
[0046] Figure 2 It is a flow chart of the multi-path redundant communication control method of the present invention;
[0047] Figure 3 It is a structural schematic diagram of a multi-path redundant communication control system of the present invention;
[0048] Figure 4 It is a schematic structural diagram of the electronic device of the present invention.
[0049] Description of reference numerals:
[0050] 201 - detection module, 202 - first switching module, 203 - second switching module, 301 - processor, 302 - memory. DETAILED DESCRIPTION
[0051] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0052] like Figure 1As shown, three communication paths, 5G, 4G, and 400M, are used between the column head and the column tail to realize communication. When a communication path fails, a redundant path failure recovery algorithm (RAP) is used to automatically switch the original communication path (also called "failed communication path" or "abnormal communication path") to other communication paths to ensure stable communication. The embodiment of the present invention provides a multi-path redundant communication control method to solve this problem. Figure 2 As shown, in one embodiment, the steps include:
[0053] S101: Determine through detection whether the original communication path is abnormal.
[0054] Check the original communication path according to the following rules:
[0055] 1. For any one of the three communication paths, when the reception of the heartbeat packet sent by any one of the three communication paths according to the predetermined period times out, or the number of consecutive heartbeat packet losses is greater than or equal to a first predetermined value, it is considered abnormal. Preferably, the first predetermined value is 10 times.
[0056] 2. For 5G communication paths, when the RSSI value (Received Signal Strength Indicator, used to measure the received signal strength, the same below) is lower than the second predetermined value, or the heartbeat packet loss rate is higher than the third predetermined value, or the bandwidth is lower than the fourth predetermined value, it is considered abnormal. Preferably, the second predetermined value is -90dBm, the third predetermined value is 0.1%, and the fourth predetermined value is 50Mbps.
[0057] 3. For 4G communication paths, when the RSSI value is lower than the fifth predetermined value, or the heartbeat packet loss rate is higher than the sixth predetermined value, or the bandwidth is lower than the seventh predetermined value, it is considered abnormal. Preferably, the fifth predetermined value is -95dBm, the sixth predetermined value is 0.5%, and the seventh predetermined value is 10Mbps.
[0058] 4. For a 400M communication path, when the RSSI value is lower than the eighth predetermined value, or the heartbeat packet loss rate is higher than the ninth predetermined value, or the bandwidth is lower than the tenth predetermined value, it is considered abnormal. Preferably, the eighth predetermined value is -80dBm, the sixth predetermined value is 2%, and the seventh predetermined value is 20Mbps.
[0059] In summary, triggering the abnormal conditions listed in any of the above rules will identify the original communication path as abnormal.
[0060] S102: Using a reinforcement learning algorithm, according to the network status of the backup communication path and the historical value of the backup communication path, an optimal backup communication path is selected, and the original communication path is switched to the optimal backup communication path.
[0061] The reinforcement learning algorithm is also called the "Q-learning algorithm". The process of selecting the optimal backup communication path using the reinforcement learning algorithm is as follows:
[0062] S102-1: Define state space and action space. The state space refers to the set of all network states of the backup communication path, including bandwidth, delay, packet loss rate, and load balancing; the action space refers to the set of actions to switch the original communication path to any backup communication path other than the original communication path.
[0063] S102-2: Define multiple state-action pairs according to the network state and action, each state-action pair has a corresponding Q value, and all state-action pairs and their corresponding Q values together form a Q table. The Q value refers to the expected value of the long-term return that can be obtained after taking an action. The Q value of each state-action pair is represented by Q(s, a), and the Q value of each state-action pair in the Q table is initialized to an eighth predetermined value. Preferably, the eighth predetermined value is 0.
[0064] S102-3: Based on the Q table, select the best action from all actions corresponding to the current network state, specifically: determine whether the Q values of each action corresponding to the current network state are the same. If they are the same, randomly select an action from all actions corresponding to the current network state as the best action, and switch the original communication path to the backup communication path corresponding to the randomly selected action; if they are not exactly the same, select the action with the largest Q value from all actions corresponding to the current network state as the best action, and switch the original communication path to the backup communication path corresponding to the action with the largest Q value.
[0065] S102-4: Record the next network state and define the next network state as s′. For example, if the current network state is bandwidth, the next network state is delay, packet loss rate, or load balancing. The selection of the next network state is determined according to the setting order of the network state in the Q table.
[0066] S102-5: Calculate the total reward after switching to the backup communication path.
[0067] 1. The bandwidth before switching to the backup communication path is referred to as the original bandwidth, defined as B; the bandwidth after switching to the backup communication path is referred to as the new bandwidth, defined as B'. If the new bandwidth is improved compared to the original bandwidth, a positive reward will be given, otherwise, the reward is 0. The calculation formula is as follows:
[0068] r bandwidth=max(0,B′-B) (1),
[0069] In the formula, r bandwidth is the reward corresponding to the bandwidth; B′ is the new bandwidth; B is the original bandwidth.
[0070] 2. The delay before switching to the backup communication path is referred to as the original delay, defined as D; the delay after switching to the backup communication path is referred to as the new delay, defined as D'. If the new delay is improved compared to the original delay, a positive reward is given, otherwise, the reward is 0. The calculation formula is as follows:
[0071] r latency =max(0,D′-D) (2),
[0072] In the formula, r latency is the reward corresponding to the delay; D′ is the new delay; D is the original delay.
[0073] 3. The packet loss rate before switching to the backup communication path is referred to as the original packet loss rate, defined as P; the packet loss rate after switching to the backup communication path is referred to as the new packet loss rate, defined as P′. If the new packet loss rate is lower than the original packet loss rate, a positive reward is given, otherwise, the reward is 0. The calculation formula is as follows:
[0074] r packet_loss =max(0,PP′) (3),
[0075] In the formula, r packet_loss is the reward corresponding to the packet loss rate; P′ is the new packet loss rate; P is the original packet loss rate.
[0076] 4. The load balance degree before switching to the backup communication path is referred to as the original load balance degree, defined as L; the load balance degree after switching to the backup communication path is referred to as the new load balance degree, defined as L'. If the new load balance degree is improved compared to the original load balance degree, a positive reward is given, otherwise, the reward is 0. The calculation formula is as follows:
[0077] r load_balance =max(0,L′-L) (4),
[0078] In the formula, r load_balance is the reward corresponding to the load balance; L′ is the new load balance; L is the original load balance.
[0079] The total reward is calculated based on the reward corresponding to bandwidth, the reward corresponding to delay, the reward corresponding to packet loss rate, and the reward corresponding to load balancing. The calculation formula is as follows:
[0080] r=ω 1 ·r bandwidth +ω 2 ·rlatency +ω 3 ·r packet_loss +ω 4 ·r packet_loss (5)
[0081] In the formula, ω i is the weight value; r bandwidth is the reward corresponding to bandwidth; r latency is the reward corresponding to the delay; r packet_loss is the reward corresponding to the packet loss rate; r load_balance The reward corresponding to the load balancing degree.
[0082] ω i Make settings based on actual site conditions.
[0083] S102-6: Based on the next network state and the total reward, the Bellman equation is used to update the Q value corresponding to the current network state and the best action. The formula is as follows:
[0084] Q(s,a)←Q(s,a)+α[r+γmax a′ Q(s′,a′)-Q(s,a)] (6),
[0085] Where r is the total reward; α is the learning rate, which is used to control the speed at which new information covers old information; γ is the discount factor, which is used to control the weight of the immediate reward; s′ is the next network state; a′ is all possible actions; max a′ Q(s′,a′) represents the maximum Q value corresponding to all actions in the next network state.
[0086] Preferably, the learning rate ranges from 0.1 to 0.5, and the discount factor ranges from 0.8 to 1.
[0087] It is worth noting that for max a′ Q(s′,a′), in the early stage of the reinforcement learning algorithm, the maximum Q value cannot be obtained. a′ Q(s′,a′) defaults to 0.
[0088] By setting a predetermined number of iterations or an iteration target, S102-1 to S102-6 are iteratively executed to continuously update the Q value of the state-action pair until the predetermined number of iterations or the iteration target is reached. After the iteration is completed, the maximum Q value is selected from the historical Q values, and the maximum Q value corresponds to the optimal backup communication path.
[0089] Then the original communication path is switched to the optimal backup communication path.
[0090] S103: Switch the optimal backup communication path back to the original communication path according to a predetermined rule.
[0091] After the original communication path is switched to the optimal alternative communication path, the optimal alternative communication path is switched back to the original communication path according to the following predetermined rules:
[0092] 1. After the original communication path is restored to an available state, monitor whether the original communication path is abnormal within a predetermined time period according to the method described in S101. If no abnormality occurs within the predetermined time period, switch the optimal backup communication path back to the original communication path; if an abnormality occurs within the predetermined time period, do not switch the optimal backup communication path back to the original communication path. Afterwards, when the original communication path is restored to an available state, judge and switch according to this rule. Preferably, the predetermined time period is 30s-60s to avoid frequent switching.
[0093] Specifically, the available status refers to:
[0094] 1) For any one of the three communication paths, the heartbeat packets sent by any one of the three communication paths according to the predetermined period are received without timeout and the number of consecutive packet losses of the heartbeat packets is less than a first predetermined value. Preferably, the first predetermined value is 10 times.
[0095] 2) For the 5G communication path, the RSSI value is greater than or equal to the second predetermined value, the heartbeat packet loss rate is less than or equal to the third predetermined value, and the bandwidth is greater than or equal to the fourth predetermined value. Preferably, the second predetermined value is -90dBm, the third predetermined value is 0.1%, and the fourth predetermined value is 50Mbps.
[0096] 3) For the 4G communication path, the RSSI value is greater than or equal to the fifth predetermined value, the heartbeat packet loss rate is less than or equal to the sixth predetermined value, and the bandwidth is greater than or equal to the seventh predetermined value. Preferably, the fifth predetermined value is -95dBm, the sixth predetermined value is 0.5%, and the seventh predetermined value is 10Mbps.
[0097] 4) For a 400M communication path, the RSSI value is greater than or equal to the eighth predetermined value, the heartbeat packet loss rate is less than or equal to the ninth predetermined value, and the bandwidth is greater than or equal to the tenth predetermined value. Preferably, the eighth predetermined value is -80dBm, the sixth predetermined value is 2%, and the seventh predetermined value is 20Mbps.
[0098] 2. Continuously monitor the optimal backup communication path according to the method described in S101, and when an abnormality occurs in the optimal backup communication path, switch the optimal backup communication path back to the original communication path.
[0099] If any one of the above two rules is met, the optimal backup communication path can be switched back to the original communication path.
[0100] In addition, the load balance of the optimal backup communication path is continuously monitored. When the load balance is higher than a predetermined threshold, part of the communication traffic in the optimal backup communication path is allocated to other communication paths outside the original communication path. If the load balance of the optimal backup communication path is still higher than the predetermined threshold after allocation, part of the communication traffic in the optimal backup communication path is continuously allocated to other communication paths outside the original communication path to avoid overloading of a single path. Preferably, the predetermined threshold is 90%. It is worth noting that the rules described in this paragraph are applicable at the same time as the above two rules.
[0101] The embodiment of the present invention also provides a multi-path redundant communication control system, such as Figure 3 As shown, including:
[0102] The detection module 201 is used to determine whether the original communication path is abnormal through detection.
[0103] The first switching module 202 is used to select the optimal backup communication path by using a reinforcement learning algorithm according to the network status of the backup communication path and the historical value of the backup communication path, and switch the original communication path to the optimal backup communication path.
[0104] The second switching module 203 is used to switch the optimal backup communication path back to the original communication path according to a predetermined rule.
[0105] It should be noted here that the above-mentioned detection module 201, the first switching module 202 and the second switching module 203 correspond to steps S101 to S103 in the embodiment of the multi-path redundant communication control method, and the examples and application scenarios implemented by the above-mentioned modules and corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned embodiments.
[0106] An embodiment of the present invention further provides a computer-readable storage medium storing a program or instruction. When the program or instruction is executed on a computer, the computer executes the multi-path redundant communication control method as described in the above method embodiment.
[0107] like Figure 4 As shown, an embodiment of the present invention further provides an electronic device, including: a processor 301, the processor 301 is coupled to a memory 302, and the processor 301 is used to read and execute a computer program stored in the memory 302 to implement the multi-path redundant communication control method as described in the above method embodiment.
[0108] Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent substitutions for some of the technical features therein; and these modifications or substitutions do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-path redundant communication control method, characterized in that: include: Determine the original communication path is abnormal through detection; Using a reinforcement learning algorithm, according to the network status of the backup communication path and the historical expected value of the backup communication path, an optimal backup communication path is selected, and the original communication path is switched to the optimal backup communication path; The optimal backup communication path is switched back to the original communication path according to a predetermined rule.
2. The method according to claim 1, characterized in that Using the reinforcement learning algorithm, the optimal backup communication path is selected according to the network status of the backup communication path and the historical expected value of the backup communication path, including: Define state space and action space; A plurality of state-action pairs are defined according to the network state in the state space and the action in the action space, each state-action pair has a corresponding Q value, and all the state-action pairs and all the Q values form a Q table; Based on the Q table, select the best action from all the actions corresponding to the current network state, and switch the original communication path to the backup communication path corresponding to the best action; Record a network status; calculating a total reward after switching to the backup communication path; Based on the next network state and the total reward, the current network state and the Q value corresponding to the best action are updated using the Bellman equation.
3. The method according to claim 2, characterized in that Using a reinforcement learning algorithm, selecting an optimal backup communication path according to a network state of the backup communication path and a historical expected value of the backup communication path, further comprising: A predetermined number of iterations or an iteration target is set, and the Q value is continuously updated until the predetermined number of iterations or the iteration target is reached, and a maximum Q value is selected from historical Q values, where the maximum Q value corresponds to the optimal backup communication path.
4. The method according to claim 2, characterized in that: The network status includes bandwidth, delay, packet loss rate and load balancing.
5. The method according to claim 2, characterized in that: Based on the Q table, selecting the best action from all the actions corresponding to the current network state, and switching the original communication path to the backup communication path corresponding to the best action, includes: When the Q value of each of the actions corresponding to the current network state is the same, randomly selecting an action from all the actions corresponding to the current network state as the optimal action; When the Q value of each of the actions corresponding to the current network state is not completely the same, the action with the largest Q value is selected from all the actions corresponding to the current network state as the optimal action.
6. The method according to claim 2, characterized in that The Bellman equation is: Q(s,a)←Q(s,a)+α[r+γmax a′ Q(s′,a′)-Q(s,a)], Where r is the total reward; α is the learning rate; γ is the discount factor; s′ is the next network state; a′ is all possible actions; max a′ Q(s′,a′) represents the maximum Q value corresponding to all actions in the next network state.
7. The method according to claim 1, characterized in that Switching the optimal backup communication path back to the original communication path according to a predetermined condition includes: After the original communication path is restored to an available state, it is monitored that no abnormality occurs on the original communication path within a predetermined time period, and the optimal backup communication path is switched back to the original communication path.
8. The method according to claim 1, characterized in that Switching the optimal backup communication path back to the original communication path according to a predetermined condition also includes: The optimal backup communication path is continuously monitored, and when an abnormality is detected in the optimal backup communication path, the optimal backup communication path is switched back to the original communication path.
9. A multi-path redundant communication control system, characterized in that: include: A detection module, used to determine whether the original communication path is abnormal through detection; A first switching module is used to select an optimal backup communication path by using a reinforcement learning algorithm according to a network status of the backup communication path and a historical value of the backup communication path, and switch the original communication path to the optimal backup communication path; The second switching module is used to switch the optimal backup communication path back to the original communication path according to a predetermined rule.
10. A computer-readable storage medium, characterized in that: A program or instruction is stored, and when the program or instruction is executed on a computer, the computer is caused to execute the multi-path redundant communication control method according to any one of claims 1 to 8.
11. An electronic device, characterized in that: comprising a processor coupled to a memory; The processor is used to read and execute the computer program stored in the memory to implement the multi-path redundant communication control method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Adaptive switching method for redundant channels
CN115001896A
Method for dynamically optimizing communication path of multistage wireless ad hoc network
CN117915422A
System and method for providing fault detection and structure switching in redandancy systemic structure communication system
CN1409494A