Self-learning algorithm for collecting field communication network quality
Through the Actor-Critic algorithm and QoS evaluation self-learning algorithm, multi-dimensional QoS data is collected in real time and network selection is optimized. The problems of communication instability and improper energy consumption management caused by the RSSI-based switching method in the prior art are solved, and efficient and stable network switching and low-power intelligent switching are achieved.
Patent Information
- Application Number
- CN202510444823.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, the RSSI-based network switching method cannot fully reflect communication quality, which can easily lead to Ping-Pong effect and unnecessary disconnection, and fail to effectively manage energy consumption, affecting communication stability and equipment battery life.
Actor-Critic algorithm and QoS evaluation are adopted to collect multi-dimensional QoS data in real time, implement strategy learning and value evaluation through neural networks, optimize network selection, and introduce a switching cost penalty mechanism to reduce unnecessary switching, improve network stability and energy consumption management.
Accurate network switching is achieved, unnecessary switching is reduced, network stability and communication reliability are improved, equipment battery life is extended, and different network environments and business needs are adapted to different network environments and business needs.
Smart Images

Figure CN120343639A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data communication in the power industry, and particularly relates to a self-learning algorithm for collecting the quality of on-site communication networks. Background Art
[0002] With the development of power automation, the requirements for power control systems are getting higher and higher, and high requirements are put forward for the real-time performance, reliability, etc. of the control system network. In order to improve the reliability of the communication network, most systems adopt an architecture of multiple communication networks (multiple communication modules or multiple SIM cards), and the multiple communication networks increase the communication cost of the system. Problems such as poor network signals and signal drift caused by overlapping on-site base stations lead to a decline in network communication quality and an increase in communication delay. In severe cases, communication may not be possible.
[0003] Traditional network switching is usually based on RSSI (Received Signal Strength Indicator). When the RSSI of the current network is lower than a certain threshold, the switching logic is triggered to select the network with the strongest RSSI as the new connection target. However, this method has some disadvantages:
[0004] 1. It cannot comprehensively reflect the network quality: RSSI only measures the signal strength, but it does not equal the communication quality (QoS). In many cases, even if the RSSI is high, the data transmission quality may still be very poor. In an environment with dense base stations or APs (access points), although the RSSI is high, the wireless interference is large, resulting in an increase in the error rate and a low actual throughput. Some operators will also limit the bandwidth of high-load base stations, resulting in a low actual available bandwidth even though the RSSI is high;
[0005] 2. It is easy to cause the Ping-Pong effect (frequent switching): Switching only based on RSSI easily causes the device to switch back and forth between two base stations or networks, resulting in an unstable connection state. RSSI is affected by many short-term environmental changes, such as weather changes (heavy rain and thunderstorms will cause signal attenuation); slight changes in the position of the mobile device (when the device rotates slightly or approaches a metal object, the RSSI will fluctuate); temporary obstacles (the movement of vehicles and personnel may cause RSSI fluctuations). If there is no hysteresis processing in the switching strategy, the device may continuously switch between multiple networks due to RSSI fluctuations within a short period of time, affecting the continuity of data transmission;
[0006] 3. Pseudo-RSSI fluctuations caused by base station load balancing mechanisms: In some cellular networks (such as 4G / 5G), operators dynamically adjust the power and load of base stations. If a base station has a high load, it may reduce the signal strength (RSSI drops), attracting some devices to switch to other base stations to relieve the pressure. However, if the device finds that the new base station has a higher load and the actual QoS deteriorates after switching, it will switch back to the original base station, forming a Ping-Pong effect.
[0007] 4. Handoff decision lag leading to unnecessary disconnections: Handoff based on RSSI usually requires setting a handoff threshold (e.g., RSSI < -85 dBm triggers handoff), but there are the following problems with the setting of the threshold: Too high a threshold → premature handoff, where the device may switch to a network with slightly stronger signal but actually worse QoS in advance, resulting in performance degradation. Too low a threshold → late handoff, where the device will switch when the signal is very poor, leading to packet loss, connection interruption, and in some cases, it cannot switch in time due to signal loss, causing disconnection. Summary of the Invention
[0008] The object of the present invention is to solve the above problems. This application proposes a self-learning algorithm for collecting the quality of on-site communication networks. By collecting multi-dimensional QoS data in real time and using neural networks to achieve policy learning and value evaluation, it optimizes network selection in the long term, taking into account both communication quality and energy consumption management.
[0009] To achieve the above object, the present invention provides the following technical solutions: A self-learning algorithm for collecting the quality of on-site communication networks, based on the Actor-Critic algorithm and QoS evaluation, collects multi-dimensional QoS data in real time, uses neural networks to achieve policy learning and value evaluation, and optimizes network selection in the long term, including the following steps:
[0010] (1) Data collection and preprocessing
[0011] Collect multi-dimensional QoS metrics of the network in real time, normalize, filter, and fuse the data to construct a standardized state vector. The evaluation formula for QoS is
[0012]
[0013] where: D is the delay, T is the bandwidth / throughput, P is the packet loss rate, J represents the jitter, i represents the i-th network, and ω1, ω2, ω3, ω4 are the weight parameters of each index, satisfying ω1 + ω2 + ω3 + ω4 = 1;
[0014] (2) Model training based on Actor-Critic
[0015] Actor Network: Generates the probability distribution of each candidate network switching action based on the current state. The goal is to learn a parameterized policy πθ(a|s) such that, given the state s, selecting the action a can maximize the future cumulative reward;
[0016] The update of Actor is based on the policy gradient method, and the update formula is:
[0017]
[0018] In the formula: θ is the Actor network parameter, α is the learning rate, which determines the step size of parameter update, πθ(a|s) is the probability of selecting the action a in the state s, output by the Actor network, represents taking the gradient with respect to θ, which is used to indicate the parameter adjustment direction, and A(s,a) is the advantage function, which is used to measure the advantage of selecting the action a in the state s relative to the average, and A(s,a) is defined as:
[0019] A(s,a) = r + γV φ (s') - V φ (s)
[0020] r is the current reward, γ is the discount factor, s' is the next state after executing the action a, and V φ (s) is the value evaluation of the state s by the Critic network;
[0021] Critic Network: Estimates the value of the current state, outputs the state value function Vφ(s), calculates the state value error by comparing the actually obtained reward with the expected value of the current state, and the Critic network is trained using the mean square error loss function, and the update formula is:
[0022] L(φ) = (r + γV φ (s') - V φ (s)) 2
[0023] The update method of gradient descent is:
[0024]
[0025] In the formula: φ is the Critic network parameter, β is the learning rate of the Critic network, r + γV φ (s') represents the target value, and V φ (s) is the predicted value of the current state;
[0026] (3) Decision Execution and Feedback Optimization Mechanism
[0027] Dynamically select the optimal network switching action based on the Actor network output and business requirements. After the switch is executed, collect the new state and actual QoS data, calculate the reward feedback, and use the feedback data to update the Actor and Critic network parameters online, continuously optimizing the switching strategy.
[0028] The reward function is defined as:
[0029] r = Qos selected -(η + θ c ·|QoS target -QoS cuerent |)-λ·N switch -δ·E
[0030] In the formula, QoS selected is the comprehensive QoS index calculated under the new network, η is the basic switching cost, θ c is the adjustment factor, N switch represents the number of switches within a certain window, λ is the corresponding penalty coefficient, E represents the power consumption index under the new network, and δ is the energy consumption penalty weight.
[0031] (4) Energy consumption management and switching cost control
[0032] Introduce the switching cost and power consumption index into the reward function to suppress the frequent switching Ping-Pong effect caused by short-term fluctuations and maintain the balance between communication quality and power consumption.
[0033] Furthermore, the training process of the Actor-Critic network is as follows:
[0034] S1. State collection: The device periodically collects the current network state sss, including signal strength, throughput, delay, jitter, and packet loss rate indicators, and forms a state vector.
[0035] S2. Policy selection: The Actor network outputs the probability distribution of each candidate network switching action a according to the current state s, and then selects a specific action a according to the policy.
[0036] S3. Action execution and feedback: The device executes the selected action a to perform network switching, obtains the new state s′, and calculates the immediate reward r according to the actual QoS change and energy consumption index.
[0037] S4. Parameter update: Use the policy gradient and advantage function to perform online adaptive learning on the model, and update the corresponding Actor and Critic networks.
[0038] S5. Loop iteration: Repeat the above steps, continuously collect new states, execute actions, and feedback updates to ensure that the model continuously learns and adaptively optimizes in a changing network environment.
[0039] Furthermore: The multi-dimensional QoS metrics include signal strength, throughput, latency, jitter, and packet loss rate.
[0040] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0041] Based on the Actor-Critic learning algorithm + QoS evaluation as the reward function, and introducing a handover cost penalty mechanism, the present invention comprehensively evaluates the network quality, realizes precise handover, effectively reduces unnecessary handovers, and improves network stability. It solves the problem in the existing solutions that only relying on signal strength for handover is likely to lead to selecting a network with high RSSI but low QoS, resulting in a decline in communication quality.
[0042] By adopting the Actor-Critic algorithm, learning the historical handover decision benefits, avoiding unnecessary handovers caused by short-term signal fluctuations, setting a handover cost penalty mechanism, and avoiding overly frequent network handovers, the present invention improves the continuity of data transmission. It solves the problem that frequent Ping-Pong handovers affect the data transmission stability of service terminals.
[0043] Through the Actor-Critic algorithm, the device can adapt to environmental changes and optimize the handover strategy. By continuously learning historical QoS data, it finds the long-term optimal network handover scheme. According to the network characteristics at different time periods, it autonomously adjusts the handover strategy to improve the long-term benefits; the present invention supports multiple networks such as 4G, 5G, Wi-Fi, NB-IoT, LoRa, etc., and performs intelligent matching in combination with service types (SCADA, video surveillance, low-power sensing) to ensure that different services select the optimal network. At the same time, it uses historical data to optimize the long-term strategy and introduces a power consumption factor to achieve low-power intelligent handover, extend the device battery life, and improve the communication reliability and stability of network devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only for more clearly illustrating the technical solutions of the embodiments of the present invention or the prior art. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0045] Figure 1 It is a flowchart of the Actor-Critic network training of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] To enable those skilled in the art to better understand and implement the technical solution of the present invention, the present invention will be further described below in conjunction with specific embodiments. However, the embodiments cited are only for the illustration of the present invention and do not limit the present invention.
[0047] A self-learning algorithm for collecting the quality of on-site communication networks, based on the Actor-Critic algorithm and QoS evaluation, uses a neural network to achieve policy learning and value evaluation, thereby optimizing network selection in the long term, taking into account both communication quality and energy consumption management. Its core technical solutions include
[0048] (1) Data collection and preprocessing
[0049] Collect multi-dimensional QoS metrics of the network in real time, such as signal strength, throughput, delay, jitter, packet loss rate, etc.; normalize, filter, and fuse the data to construct a standardized state vector. The evaluation formula for QoS is
[0050]
[0051] where: D is the delay (Delay, D),
[0052] T is the bandwidth / throughput (Throughput, T),
[0053] P is the packet loss rate (Packet Loss, P),
[0054] J represents jitter (Jitter, J),
[0055] i represents the i-th network,
[0056] ω1, ω2, ω3, ω4 are the weight parameters of each index, satisfying ω1 + ω2 + ω3 + ω4 = 1.
[0057] (2) Model training based on Actor-Critic
[0058] Actor network: Generate the probability distribution of each candidate network switching action according to the current state. The goal is to learn a parameterized policy πθ(a|s) such that, given the state s, selecting the action a can maximize the future cumulative reward;
[0059] The update of Actor is based on the policy gradient method, and the update formula is:
[0060]
[0061] where: θ is the Actor network parameter,
[0062] α is the learning rate, which determines the step size of parameter update,
[0063] πθ(a|s) is the probability of selecting action a in state s, output by the Actor network,
[0064] denotes taking the gradient with respect to θ, which is used to indicate the direction of parameter adjustment,
[0065] A(s,a) is the advantage function, which is used to measure the advantage of selecting action a in state s relative to the baseline. A(s,a) is defined as:
[0066] A(s,a) = r + γV φ (s') - V φ (s) (3)
[0067] r is the current reward,
[0068] γ is the discount factor,
[0069] s' is the next state after executing action a,
[0070] V φ (s) is the value evaluation of state s by the Critic network;
[0071] Critic network: Estimates the value of the current state, outputs the state value function Vφ(s), calculates the state value error by comparing the actually obtained reward with the expected value of the current state. The Critic network is trained using the mean squared error loss function, and the update formula is:
[0072] L(φ) = (r + γV φ (s') - V φ (s)) 2 (4)
[0073] The update method of gradient descent is:
[0074]
[0075] where: φ are the parameters of the Critic network,
[0076] β is the learning rate of the Critic network,
[0077] r + γV φ (s') represents the target value, that is, the reward r obtained after taking action a from the current state plus the expected value of the discounted next state,
[0078] V φ (s) is the predicted value of the current state,
[0079] The gradient descent update process aims to minimize the error between the actual value and the predicted value, ensuring that the Critic can accurately evaluate the state value.
[0080] The training process of the Actor-Critic network is as follows Figure 1 shown below:
[0081] S1. State acquisition: The device periodically acquires the current network state sss, including signal strength, throughput, latency, jitter, and packet loss rate metrics, and forms a state vector;
[0082] S2. Policy selection: The Actor network outputs the probability distribution of each candidate network switching action a according to the current state s, and then selects a specific action a according to the policy;
[0083] S3. Action execution and feedback: The device executes the selected action a for network switching, obtains the new state s′, and calculates the immediate reward r according to the actual QoS change and energy consumption metrics;
[0084] S4. Parameter update: Use policy gradient and advantage function to perform online adaptive learning on the model, and update the corresponding Actor and Critic networks;
[0085] S5. Loop iteration: Repeat the above steps, continuously acquire new states, execute actions, and perform feedback updates to ensure that the model continuously learns and adaptively optimizes in a changing network environment.
[0086] (3) Decision execution and feedback optimization mechanism
[0087] Dynamically select the optimal network switching action according to the output of the Actor network and business requirements. After the switch is executed, acquire the new state and actual QoS data, calculate the reward feedback, and use the feedback data to online update the parameters of the Actor and Critic networks to continuously optimize the switching strategy.
[0088] The reward function is defined as:
[0089] r = QoS selected -(η + θ c ·|QoS target - QoS cuerent |)-λ·N switch -δ·E (6)
[0090] QoS selected is the comprehensive QoS metric calculated under the new network,
[0091] η is the basic switching cost, reflecting the fixed overhead of physical layer switching,
[0092] θ c is the adjustment factor, measuring the QoS metrics of the old and new networks (calculated according to formula (1)),
[0093] N switchIndicates the number of handovers within a certain window (used to penalize overly frequent handovers)
[0094] λ is the corresponding penalty coefficient
[0095] E represents the power consumption index under the new network
[0096] δ is the power consumption penalty weight
[0097] (4) Power consumption management and handover cost control
[0098] Introduce handover cost and power consumption index into the reward function to suppress frequent handovers (Ping-Pong effect) caused by short-term fluctuations, achieve a balance between communication quality and power consumption, and ensure energy consumption reduction while meeting QoS
[0099] Preferably: The present invention supports adaptive handover between multiple networks, is applicable to multiple networks such as 4G, 5G, Wi-Fi, NB-IoT, LoRa, industrial Wi-Fi, etc., and intelligently matches the optimal network according to service requirements (high throughput / low latency) to optimize power consumption and performance
[0100] Key points and protection points of the present invention
[0101] Adopt Actor-Critic algorithm + QoS evaluation to achieve more efficient, intelligent and stable network handover
[0102] 1) Intelligent network handover mechanism based on QoS evaluation
[0103] Adopt QoS as the core decision-making basis. Key QoS indicators include but are not limited to throughput, latency, jitter, packet loss rate, network load, etc. The device calculates the reward value according to QoS parameters, optimizes the handover strategy, and improves network stability and service adaptability
[0104] 2) Use Actor-Critic for self-learning to optimize handover decisions
[0105] The device adopts the Actor-Critic learning algorithm to dynamically learn the network environment and avoid inefficient handovers caused by fixed rules: train the network through Actor-Critic, and optimize the handover strategy according to historical data. Combine with the e-greedy exploration strategy to dynamically balance between long-term optimality and short-term exploration, set the handover cost penalty term, reduce unnecessary handovers, and avoid the Ping-Pong effect
[0106] 3) Support adaptive handover between multiple networks
[0107] Adapted to multiple networks such as 4G, 5G, Wi-Fi, NB-IoT, LoRa, and Industrial Wi-Fi, it intelligently matches the optimal network according to service requirements (high throughput / low latency), and optimizes power consumption and performance.
[0108] 4) Long-term optimal strategy learning
[0109] Through network learning, the device can optimize the handover strategy based on historical data: dynamically adjust the handover strategy according to time period (day / night) and location. Combine long-term QoS data to train the actor network to achieve long-term optimal decision-making.
[0110] Compared with the traditional RSSI-based handover scheme, the present invention has higher accuracy, more stable handover strategy, stronger service adaptability, supports multiple network modes, and optimizes power consumption management. The traditional RSSI scheme only makes handovers based on signal strength, which is prone to misjudging network quality, resulting in service interruptions and frequent Ping-Pong handovers. The present invention adopts the Actor-Critic learning algorithm, combines QoS evaluation (throughput, latency, packet loss rate, etc.) as the core decision-making basis, and introduces a handover cost penalty mechanism to effectively reduce unnecessary handovers and improve network stability. In addition, the present invention supports multiple networks such as 4G, 5G, Wi-Fi, NB-IoT, and LoRa, and intelligently matches them according to service types (SCADA, video surveillance, low-power sensing) to ensure that different services select the optimal network. At the same time, historical data is used to optimize long-term strategies, and a power consumption factor is introduced to achieve low-power intelligent handover, extend the device's battery life, and improve the communication reliability and stability of network devices.
[0111] The content not described in detail in the present invention is prior art.
[0112] The above are only the preferred embodiments of the present invention, and are not limited to those described in the specification and implementation manners. Therefore, any equivalent changes or modifications made according to the structure, features, and principles described in the scope of the patent application of the present invention shall be included in the scope of the patent application of the present invention.
Claims
1. A self-learning algorithm for collecting the quality of on-site communication networks, based on the Actor-Critic algorithm and QoS evaluation, realizes policy learning and value evaluation by using a neural network through real-time collection of multi-dimensional QoS data, and optimizes network selection in the long term, characterized in that: It includes the following processes (1) Data collection and preprocessing Collect multi-dimensional QoS metrics of the network in real time, normalize, filter and fuse the data, and construct a standardized state vector. The evaluation formula for QoS is where: D is the delay, T is the bandwidth / throughput, P is the packet loss rate, J represents the jitter, i represents the i-th network, and ω1, ω2, ω3, ω4 are the weight parameters of each metric, satisfying ω1 + ω2 + ω3 + ω4 = 1; (2) Model training based on Actor-Critic Actor network: Generate the probability distribution of each candidate network switching action according to the current state. The goal is to learn a parameterized policy πθ(a|s) such that, given the state s, selecting the action a can maximize the future cumulative reward; The update of Actor is based on the policy gradient method, and the update formula is: where: θ is the Actor network parameter, α is the learning rate, which determines the step size of parameter update, πθ(a|s) is the probability of selecting action a in state s, output by the Actor network, denotes taking the gradient with respect to θ, used to indicate the parameter adjustment direction, A(s,a) is the advantage function, used to measure the advantage of selecting action a in state s relative to the baseline, and A(s,a) is defined as: A(s,a) = r + γV φ (s') - V φ (s) r is the current reward, γ is the discount factor, s' is the next state after executing action a, and V φ (s) is the value evaluation of state s by the Critic network; Critic network: Estimate the value of the current state and output the state value function Vφ(s). Calculate the state value error by comparing the actually obtained reward with the expected value of the current state. The Critic network is trained using the mean squared error loss function, and the update formula is: L(φ) = (r + γV φ (s') - V φ (s)) 2 The update method of gradient descent is: Where: φ is the parameter of the Critic network, β is the learning rate of the Critic network, r + γV φ (s') represents the target value, V φ (s) is the predicted value of the current state; (3) Decision execution and feedback optimization mechanism Dynamically select the optimal network switching action according to the output of the Actor network and service requirements. After the switch is executed, collect the new state and actual QoS data, calculate the reward feedback, and use the feedback data to online update the parameters of the Actor and Critic networks to continuously optimize the switching strategy. The reward function is defined as: r = QoS selected -(η + θ c |·QoS target -QoS cuerent |)-λ·N switch -δ·E where QoS selected is the comprehensive QoS index calculated under the new network, η is the basic handover cost, θ c is the adjustment factor, N switch represents the number of handovers within a certain window, λ is the corresponding penalty coefficient, E represents the power consumption index under the new network, and δ is the energy consumption penalty weight; (4) Energy consumption management and switching cost control Introduce switching cost and power consumption metrics into the reward function to suppress the frequent switching Ping-Pong effect caused by short-term fluctuations and maintain the balance between communication quality and power consumption.
2. The self-learning algorithm for collecting the quality of the on-site communication network according to claim 1, characterized in that: The training process of the Actor-Critic network is as follows: S1. State collection: The device periodically collects the current network state sss, including signal strength, throughput, delay, jitter, and packet loss rate metrics, and forms a state vector; S2. Policy selection: The Actor network outputs the probability distribution of each candidate network switching action a according to the current state s, and then selects a specific action a according to the policy; S3. Action execution and feedback: The device executes the selected action a to perform network switching, obtains the new state s′, and calculates the immediate reward r according to the actual QoS change and energy consumption metrics; S4. Parameter update: Use the policy gradient and advantage function to perform online adaptive learning on the model, and update the corresponding Actor and Critic networks; S5. Loop iteration: Repeat the above steps, continuously collect new states, execute actions and feedback updates to ensure that the model continuously learns and adaptively optimizes in a changing network environment.
3. The self - learning algorithm for collecting the quality of on - site communication network according to claim 1, characterized in that: The multi-dimensional QoS metrics include signal strength, throughput, delay, jitter, and packet loss rate.
Citation Information
Cited By
Communication equipment multi-link automatic switching control method based on reinforcement learning
CN121396704A