Optimization method for information age and energy consumption in hybrid industrial wireless network

Through the optimized and improved dual-head Actor network of Lyapunov, combined with virtual queues and dynamic channel allocation, the optimization problems of information age and energy consumption in hybrid industrial wireless networks are solved, low overdue rate, high throughput and energy consumption balance are achieved, and network performance and equipment life are improved.

CN120456062APending Publication Date: 2025-08-08CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510574420.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

In hybrid industrial wireless networks, the prior art is difficult to achieve low overdue rate of periodic data, high throughput of non-periodic data under the conditions of restricted spectrum resources, time-varying channel state and energy constraints, and minimize the optimization of the system's long-term average information age and energy consumption.

Method used

Using the scheduling framework combined with the improved double-headed Actor network, the scheduling and channel allocation decisions are decoupled and the information age and energy consumption are optimized by building virtual queues and dynamic channel allocation strategies.

Benefits of technology

Significantly reduce the periodic data overdue rate by 40%-60%, improve the non-periodic data throughput by 25%, and at the same time reduce node energy consumption by 15%-20%, extend the life of battery-powered nodes by 30%-50%, improve channel utilization to more than 85%, and reduce algorithm robustness fluctuations by 60%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120456062A_ABST
    Figure CN120456062A_ABST
Patent Text Reader

Abstract

The invention relates to an optimization method for information age and energy consumption in a hybrid industrial wireless network, and belongs to the technical field of industrial Internet of Things. Aiming at the problems of low scheduling efficiency, high periodic data overdue rate, insufficient aperiodic data throughput and unbalanced energy consumption under a dynamic channel condition, a joint optimization scheme based on Lyapunov optimization and a double-end Actor network is provided. By constructing an overdue rate and throughput virtual queue, converting long-term constraint into a real-time optimization problem, establishing a Markov decision model, adopting a double-end Actor network to decouple and update scheduling and channel allocation decision, and combining a conflict detection mechanism and dynamic weight adjustment, the periodic data overdue rate is reduced by 40-60%, and the real-time performance of the system is improved. And the aperiodic data throughput is improved by more than 25%. The method solves the problem of collaborative optimization of information freshness and energy consumption in a mixed service scene, and is suitable for real-time control and monitoring scenes of the industrial Internet of Things.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of industrial Internet of Things and relates to a method for optimizing information age and energy consumption in a hybrid industrial wireless network. Background Art

[0002] With the rapid development of the Industrial Internet of Things (IIoT), the importance of industrial wireless networks in smart manufacturing, real-time monitoring, and automated control is becoming increasingly prominent. Nodes need to transmit device status data to a central controller or the cloud in real time to ensure the system's stringent requirements for information freshness. Especially in hybrid industrial scenarios, the network needs to support both periodic control data (such as device status sampling) and bursty monitoring data (such as fault alarms). The two types of data have significant differences in quality of service (QoS) requirements: periodic data must be strictly guaranteed to be delivered within the deadline to avoid control delays, while non-periodic data must maintain high throughput to cope with real-time responses to emergencies.

[0003] However, existing technologies face multiple challenges in coping with such heterogeneous data mixed transmission scenarios:

[0004] (1) Spectrum resources in industrial environments are limited and channel conditions vary (e.g., multipath fading, electromagnetic interference), leading to fluctuations in transmission success rates. Traditional static scheduling strategies are difficult to adapt to dynamic channel conditions and can easily lead to periodic data overdue or non-periodic data congestion.

[0005] (2) Age of Information (AoI), as a key indicator for measuring data freshness, needs to be optimized through frequent updates. However, frequent communications will increase the energy consumption of nodes. Industrial nodes are often deployed in power-constrained scenarios (such as battery power), and there is an urgent need to achieve an efficient balance between AoI and energy consumption.

[0006] (3) There is a natural conflict between the overdue rate constraints of periodic data and the throughput requirements of non-periodic data. Existing research has mostly focused on optimizing a single data type, such as optimizing throughput through queuing theory or reducing overdue rates through deadline-aware scheduling, but lacks a collaborative optimization mechanism for both types of constraints.

[0007] Some studies have attempted to address the above issues:

[0008] Long-term constraints are converted into instantaneous optimization objectives through virtual queues, but traditional methods face the problem of computational complexity explosion in high-dimensional action spaces (such as multi-channel allocation).

[0009] Deep reinforcement learning (DRL) is used to adaptively learn channel status and business needs. However, the standard actor-critic algorithm is prone to policy oscillation in the mixed action space (discrete scheduling + continuous power control), and no dedicated network structure is designed for the periodic / aperiodic data characteristics of industrial scenarios.

[0010] Existing AoI optimization solutions mostly assume ideal channel conditions or a single service type, and do not consider the dynamic coupling effects of multi-channel contention and data queues in actual industrial networks.

[0011] Therefore, designing an optimization method that simultaneously achieves low overdue rates for periodic data, high throughput for aperiodic data, and minimizes the system's long-term average information age and energy consumption, while addressing limited spectrum resources, time-varying channel states, and energy constraints, has become a core challenge in the field of industrial wireless networks. Addressing this technological gap, this paper proposes a hybrid scheduling framework that integrates Lyapunov optimization and improved reinforcement learning. This framework decouples scheduling and channel allocation decisions through a dual-headed strategy network, achieving efficient optimization under complex constraints. Summary of the Invention

[0012] In view of this, the purpose of the present invention is to provide a method for optimizing information age and energy consumption in hybrid industrial wireless networks. By optimizing the update scheduling and channel allocation strategies, an optimization problem is constructed with the goal of minimizing the long-term average information age and energy consumption of the system, while satisfying the overdue rate of periodic data and the throughput constraints of non-periodic data.

[0013] In order to achieve the above object, the present invention provides the following technical solutions:

[0014] A method for optimizing information age and energy consumption in a hybrid industrial wireless network, the method specifically comprising the following steps:

[0015] S1: Obtain system information of the hybrid industrial wireless network, establish a system model based on the network structure, and construct an optimization problem to minimize the system's long-term average information age and energy consumption under the constraints of periodic data overdue rate and aperiodic data throughput;

[0016] S2: The optimization problem is converted using the Lyapunov optimization method. The state space, action space, and cost function of the system are then established based on the relevant system parameters. This transforms the link scheduling optimization problem into a constrained Markov decision process.

[0017] S3: Use an improved two-headed actor network to obtain strategies and utilize a critic network to evaluate value. Specifically, the current state is input into the two-headed actor network to obtain updated scheduling and channel allocation strategies. Decisions are obtained through policy distribution sampling and actions are executed to obtain the next state. The link scheduling method is obtained by waiting for the network's loss function to stabilize.

[0018] Furthermore, the S1 specifically includes the following steps:

[0019] S11: The hybrid industrial wireless network system includes a base station and N sensor nodes; the sensor nodes are composed of K periodic sampling nodes and NK random sampling nodes; using i∈φ P ={1,2,…,K} and j∈φ R ={K+1,K+2,…,N} denote the indices of periodic sampling nodes and random sampling nodes respectively, and n∈φ={1,2,…,N} is used to denote the index of all nodes; assuming that the time in the network is divided into multiple time slots and the maximum running time of the entire network is T time slots, let t∈{1,2,…,T} denote the index of the time slot; there are M orthogonal subchannels in the network that do not interfere with each other, and these subchannels are denoted by m∈{1,2,…,M}; spectrum resources are limited, and the number of nodes K in the network is greater than the number of subchannels M; let U n (t)∈{0,1} indicates whether the base station selects node n to send data at the beginning of time slot t, U n (t) = 1 means node n is scheduled and U n (t) = 0 means that node n is not scheduled; in order to prevent channel interference and thus transmission conflicts, it is assumed that the node update scheduling strategy in each time slot is limited by the number of available sub-channels; the node channel scheduling strategy uses V n (t)∈{0,1,2,…M} represents, V n (t) = m means that node n uses subchannel m to send data in time slot t. If node n does not send data, then V n (t) = 0; each node can only send a data packet through one sub-channel in each time slot;

[0020] S12: If the node collects data, then g n (t)=1, otherwise g n (t) = 0, and at most one data packet is generated in each time slot; the sampling period of the periodic sampling node is H i ; The periodic sampling node uses a single packet buffer, and the newly generated data will replace the old data; let g n (t)∈{0,1} represents whether node n generates and collects data packets in time slot t; for random sampling nodes, it is assumed that their data sampling method obeys the parameter λj ∈(0,1] Bernoulli distribution; random sampling nodes store data packets through the data queue, let Q j (t) represents the amount of data in the transmission queue of the randomly sampled node j at time slot t; the data buffer follows the first-in-first-out principle and gives priority to the first-arrived data. The evolution process of the data queue is as follows:

[0021] Q j (t+1)=max(Q j (t)+g j (t)L j -d j (t)l j ,Q max )

[0022] where l n Indicates the data length of the node sampling, Q max Indicates the upper limit of the data queue; d n (t)∈{0,1} indicates whether the base station receives the data of node n in time slot t. If the base station receives the data of node n, then d n (t)=1, otherwise d n (t) = 0;

[0023] S13: Assuming that the channel state is time-varying, the M orthogonal channels in the network are modeled as the Gilbert-Elliott channel model; we let h n,m (t)∈{0,1} represents the state information of channel m in the time slot before time slot t; the channel state is divided into good state and bad state. If the channel state between node n and subchannel m is in good state, then h n,m (t)=1, otherwise h n,m (t) = 0; in the "bad" state, the channel is assumed to be in deep fading, so that the transmission fails with probability 1, while in the "good" state, the transmission attempt is always successful; the channel transition probability is given by P(h n,m (t+1)=1|h n,m (t) = 0) = p 01 and P(h n,m (t+1)=0|h n,m (t) = 1) = p 10 given;

[0024] S14: The AoI value of node n at the base station is represented by a n (t); If the base station receives the data packet from node n in time slot t, then a n The value of (t) will be updated to the time that has passed since the packet was generated, otherwise, a n (t+1)=a n (t)+1;an The update process of (t) is as follows:

[0025]

[0026] where z n (t) is the AoI value at node n at time slot t. The AoI update method is different at different types of nodes. The AoI update at the periodic sampling node is related to the sampling period. Each time the sampling is completed, the AoI is updated to 0. The AoI value at the periodic sampling node is z i The iterative process of (t) is expressed as follows:

[0027]

[0028] There is a data queue in the random sampling node, and the generation time of each data packet in the data queue needs to be considered when updating; if the data queue Q j (t) is greater than l j , then the information age of the subsequent data of the head data in the data queue is Where s is the subsequent data identifier; the value of AoI at the randomly sampled node is z j (t) The iterative process is expressed as follows:

[0029]

[0030] S15: Assume that the energy consumption of the node is mainly consumed by the process of sending data packets to the base station through the sub-channel, ignoring the power consumption of the node in idle mode; the transmission rate of the sub-channel is fixed to R n , let P n They represent the transmission power of node n respectively; then the energy consumption of node n in time slot t is:

[0031]

[0032] S16: The data sampled by the periodic sampling node needs to be delivered with a delay of k as much as possible. i (t) is less than the deadline X i (t), if k i (t)≤X i , which means that the data is delivered within the deadline, otherwise it is considered overdue; define c i (t)∈{0,1} is whether a data overdue event occurs at the periodic sampling node i in time slot t. i (t) = X i +1 when let c i (t)=1, it is overdue, otherwise c i (t) = 0; data delivery delay k i (t) The change process is shown as follows:

[0033] k i (t+1)=(1-d i (t))(k i (t)+1)

[0034] Order o i (t) represents the overdue rate of node i from the initial time slot to time slot t, which is expressed as:

[0035]

[0036] S17: For data generated by random sampling nodes, the long-term average throughput constraint needs to be considered to maintain the update quality; the long-term average throughput of random sampling nodes in the network is defined as:

[0037]

[0038] The average throughput constraint of randomly sampled nodes is expressed as:

[0039]

[0040] where δ min represents the minimum throughput requirement of randomly sampled node n;

[0041] S18: Under the constraints of the probability of periodic data delivery delay exceeding the deadline and the minimum average throughput of randomly arriving nodes, the optimization problem of optimizing the long-term average information age and energy consumption of the system is expressed as follows:

[0042]

[0043] where Γ n represents the overdue rate threshold of periodic sampling node n, β is the energy consumption weight coefficient, I{V n (t) = m} is V n (t) = indicator function of event occurrence, if V n (t)=m, then I{V n (t)=m}=1, otherwise I{V n (t)=m}=0.

[0044] Furthermore, the step S2 specifically includes the following steps:

[0045] S21: According to the Lyapunov optimization method, define a virtual queue Z with an overdue rate for each periodic sampling node. i (t), represents the accumulation of overdue events; its update process is as follows:

[0046] Z i(t+1)=max{Z i (t)+c i (t)-Γ i s i (t),0}

[0047] S22: Define a throughput virtual queue Y for each randomly sampled node j (t), represents the gap between the average throughput and the minimum throughput requirement; its update rule is:

[0048] Y j (t+1)=max{Y j (t)+δ min -d j (t)l j ,0}

[0049] S23: Based on the states of the two types of queues, a composite Lyapunov function L(t) is defined to measure the congestion level of the queue:

[0050]

[0051] where ω j >0 is the weight coefficient of the throughput virtual queue; the Lyapunov drift term is defined as the expected value of the change of the Lyapunov function in a time slot, which is expressed as follows:

[0052]

[0053] S24: Based on Lyapunov optimization theory, a control strategy that can minimize the original optimization problem with constraints is obtained by minimizing the penalty plus drift function. The penalty plus drift function is expressed as follows:

[0054]

[0055] Where V represents the weight parameter that measures the importance of the penalty function. By changing the value of V, the desired trade-off between the size of the queue backlog and the objective function value is obtained;

[0056] S25: According to Lyapunov optimization theory, the objective function of the optimization problem P1 is minimized under the constraints by minimizing the upper bound of the drift plus penalty. The optimization problem is transformed into the following form:

[0057]

[0058] S26: According to the above optimization problem, the state space s(t) of the established system is:

[0059] S(t)={z(t),a(t),e(t),Z(t),Y(t)}

[0060] Where z(t)=(z1(t),z2(t),...z n (t)), a(t)=(a1(t),a2(t),...a n (t)), e(t)=(e1(t),e2(t),...e n (t)), Z(t)=(Z1(t),Z2(t),...Z i (t)), Y(t)=(Y1(t),Y2(t),...Y j (t));

[0061] S27: The action space A(t) of the established system is:

[0062] A(t)={U(t),V(t)}

[0063] Among them, U(t)=(U1(t),U2(t),...U n (t)), V(t)=(V1(t),V2(t),...V n (t));

[0064] S28: The reward function r(t) of the established system is:

[0065]

[0066] Furthermore, the step S3 specifically includes the following steps:

[0067] S31: Based on a dual-head actor network containing an update scheduling head and a channel allocation head, the current system state s(t) is input into the network and features f(S(t)) are extracted through a shared feature extraction layer; the update scheduling head outputs the strategy distribution of the updated scheduling action based on f(S(t)). Sampling the strategy distribution to obtain the updated scheduling action U(t), and then input the feature f(S(t)) and the scheduling action to obtain the channel allocation strategy distribution

[0068] S32: Sampling strategy distribution obtains actions and executes to obtain the next state;

[0069] Calculate the Actor network loss function Loss through the current state, next state, and action a (θ) and the loss function of the Critic network Loss c (θ):

[0070]

[0071] Loss c (θ)=Ω 2(t)=(r(t)+γQ(S(t+1),A(t+1))-Q(S(t),A(t))) 2

[0072] in To update the scheduling head gradient, The gradient of the scheduling head is assigned to the channel, Ω(t) is the TD error, Q(S, A) is the Q value estimated by the critic network; B(t) is the advantage function, which is approximated by the TD error;

[0073] S33: Update the network parameters according to the gradient descent method, wait for the network to reach the termination condition, and obtain a scheduling method for hybrid industrial wireless networks.

[0074] Furthermore, the channel allocation strategy includes a conflict detection mechanism when it is executed. Specifically, when multiple nodes are assigned to the same sub-channel, the node that makes Z n (t)(c n (t)-Γ n )+ω n Y n (t)(δ min -d n (t)l n ) value obtains the channel usage right, and the remaining nodes are set to the unscheduled state in this time slot.

[0075] Furthermore, the weighted coefficient ω of the throughput virtual queue is j The determination method is: according to the random sampling node data queue length Q j (t) Dynamic adjustment, when Q n (t)>0.8Q max When ω j =2ω0; when 0.5Q max <Q j (t)≤0.8Q max When, ω j =1.5ω0; other casesω j =ω0, where ω0 is the preset basic weight.

[0076] Furthermore, during the training process of the two-headed Actor network, the policy update priority of updating the scheduling head is higher than that of the channel allocation head. Specifically, during gradient backpropagation, the learning rate of updating the scheduling head is set to 1.2 to 1.5 times the learning rate of the channel allocation head.

[0077] Furthermore, the energy consumption optimization weight β is determined as follows: when the base station detects that the remaining power of a node is less than 20%, the β value corresponding to the node is automatically increased to β max , and according to the formula β=β max×(1-SOC / 100) dynamic adjustment, where SOC is the current power percentage of the node.

[0078] Furthermore, the sampling process of the strategy distribution includes an exploration-exploitation balance mechanism, specifically: ε=0.3 in the ε-greedy strategy at the beginning of training, and decays according to ε←0.99ε after every 1000 training iterations until ε drops to 0.05 and remains constant.

[0079] Furthermore, the overdue rate virtual queue Z i (t) and throughput virtual queue Y j The initialization method of (t) is: Z i (t) = ceil(Γ i T total / H i ), Y j (0) = floor(δ min T total / l j ), where T total The total number of running slots preset by the system, ceil() is the rounding up function, and floor() is the rounding down function

[0080] The beneficial effects of the present invention are:

[0081] (1) By constructing a virtual queue for periodic data overdue rates and a virtual queue for aperiodic data throughput, long-term statistical constraints are transformed into an immediate queue stability problem. Combined with the design of a Lyapunov drift plus penalty term, the system automatically balances the conflicting demands of the two types of data under dynamic channel conditions. Actual measurements show that, under the same channel conditions, the periodic data overdue rate is reduced by 40%-60%, and the aperiodic data throughput is increased by more than 25%. At the same time, the compliance rate for the C1 and C2 constraints exceeds 98%.

[0082] (2) Innovatively integrate AoI and energy consumption into a unified optimization framework and introduce a dynamic energy consumption weight coefficient β. By decoupling scheduling decisions and channel allocation through a dual-headed actor network, the strategy oscillation problem of traditional methods in high-dimensional action space is avoided. Experiments have shown that compared with solutions that only optimize AoI or energy consumption, this method reduces node energy consumption by 15%-20% while ensuring a 30% reduction in the average AoI, significantly extending the service life of battery-powered nodes.

[0083] (3) A spatiotemporal feature prediction mechanism is constructed based on the Gilbert-Elliott channel model. Through conflict detection and dynamic priority scheduling, channel utilization is increased to over 85%. In a typical scenario with M = 8 subchannels and K = 15 nodes, the transmission conflict rate is reduced from 22% in traditional random allocation to below 5%, and the channel allocation decision delay is controlled at the order of 0.1ms, meeting industrial real-time requirements.

[0084] (4) A dual-head actor network structure with hierarchical feature sharing reduces the number of parameters by 40% by sharing the underlying feature extraction layer, and increases the training speed by 2 times. Combining a dynamic learning rate mechanism with an exploration-exploitation balance strategy, the number of iterations required to achieve stable convergence in complex industrial scenarios is reduced by 50%, and the strategy fluctuation amplitude is reduced by 60%, significantly improving the robustness of the algorithm.

[0085] (5) Differentiated energy consumption control is achieved through a dynamic energy consumption weight coefficient β and a node power sensing mechanism. When a node's remaining power falls below 20%, its β value is automatically raised to a preset upper limit, prioritizing the reduction of high-energy consumption operations. Actual measurement data shows that this mechanism can extend the life cycle of low-power nodes by 30%-50% and improve the energy balance of the entire network by 35%.

[0086] (6) The hardware implementation of the CPU+NPU heterogeneous computing architecture compresses the inference latency of the dual-head actor network to less than 1 / 20 of the time slot length (typical value <50μs), supporting millisecond-level scheduling response. Through shared memory pools and pipeline optimization, resource overhead is reduced to 1 / 3 of traditional DRL solutions, making it suitable for resource-constrained industrial edge devices.

[0087] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0088] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:

[0089] Figure 1 This is a schematic diagram of the hybrid industrial wireless network structure provided by the present invention;

[0090] Figure 2 This is a schematic diagram of the channel model provided by the present invention;

[0091] Figure 3 This is a flow chart of the method for optimizing scheduling of information age and energy consumption in a hybrid industrial wireless network of the present invention. DETAILED DESCRIPTION

[0092] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0093] Among them, the accompanying drawings are only for illustrative purposes and represent only schematic diagrams rather than actual pictures, and should not be understood as limiting the present invention. In order to better illustrate the embodiments of the present invention, some parts of the accompanying drawings may be omitted, enlarged or reduced, and do not represent the dimensions of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions may be omitted in the accompanying drawings.

[0094] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "back", etc. indicating directions or positional relationships, they are based on the directions or positional relationships shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, the terms describing the positional relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.

[0095] A method for optimizing information age and energy consumption in hybrid industrial wireless networks is proposed. This approach aims to minimize the system's long-term average information age and energy consumption by optimizing update scheduling and channel allocation strategies for industrial wireless networks with a coexistence of periodic and random sampling nodes. This optimization problem is formulated to minimize the system's long-term average information age and energy consumption while satisfying the overdue rate of periodic data arrivals and the throughput constraints of aperiodic data. The long-term constraints are transformed into immediate subproblems using the Lyapunov optimization method, and the optimization problem is modeled as a Markov decision process. The method innovatively employs a two-headed actor network based on the actor-critic algorithm to decouple update scheduling and channel allocation tasks, significantly reducing the dimensionality of the action space, improving learning efficiency and policy quality, and ultimately achieving a link scheduling method.

[0096] The method specifically comprises the following steps:

[0097] S1: Obtain system information of the industrial wireless network, establish a system model based on the network structure, and construct an optimization problem to minimize the system's long-term average information age and energy consumption under the constraints of periodic data overdue rate and aperiodic data throughput;

[0098] S2: The optimization problem is converted using the Lyapunov optimization method. The state space, action space, and cost function of the system are then established based on the relevant system parameters. This transforms the link scheduling optimization problem into a constrained Markov decision process.

[0099] S3: Use an improved two-headed actor network to obtain strategies and utilize a critic network to evaluate value. Specifically, the current state is input into the two-headed actor network to obtain updated scheduling and channel allocation strategies. Decisions are obtained through policy distribution sampling and actions are executed to obtain the next state. The link scheduling method is obtained by waiting for the network's loss function to stabilize.

[0100] Furthermore, in S1, establishing the system model and constructing the optimization problem specifically includes the following steps:

[0101] S11: The hybrid industrial wireless network system includes a base station and N sensor nodes; the sensor nodes are composed of K periodic sampling nodes and NK random sampling nodes; using i∈φ P ={1,2,…,K} and j∈φ R ={K+1,K+2,…,N} denote the indices of periodic sampling nodes and random sampling nodes respectively, and n∈φ={1,2,…,N} is used to denote the index of all nodes; assuming that the time in the network is divided into multiple time slots and the maximum running time of the entire network is T time slots, let t∈{1,2,…,T} denote the index of the time slot; there are M orthogonal subchannels in the network that do not interfere with each other, and these subchannels are denoted by m∈{1,2,…,M}; spectrum resources are limited, and the number of nodes K in the network is greater than the number of subchannels M; let U n (t)∈{0,1} indicates whether the base station selects node n to send data at the beginning of time slot t, U n (t) = 1 means node n is scheduled and U n (t) = 0 means that node n is not scheduled; in order to prevent channel interference and thus transmission conflicts, it is assumed that the node update scheduling strategy in each time slot is limited by the number of available sub-channels; the node channel scheduling strategy uses V n (t)∈{0,1,2,…M} represents, V n (t) = m means that node n uses subchannel m to send data in time slot t. If node n does not send data, then V n (t) = 0; each node can only send a data packet through one sub-channel in each time slot;

[0102] S12: If the node collects data, then g n (t)=1, otherwise g n (t) = 0, and at most one data packet is generated in each time slot; the sampling period of the periodic sampling node is H i ; The periodic sampling node uses a single packet buffer, and the newly generated data will replace the old data; let g n (t)∈{0,1} represents whether node n generates and collects data packets in time slot t; for random sampling nodes, it is assumed that their data sampling method obeys the parameter λ j ∈(0,1] Bernoulli distribution; random sampling nodes store data packets through the data queue, let Q j (t) represents the amount of data in the transmission queue of the randomly sampled node j at time slot t; the data buffer follows the first-in-first-out principle and gives priority to the first-arrived data. The evolution process of the data queue is as follows:

[0103] Q j (t+1)=max(Q j (t)+g j (t)L j -d j (t)l j ,Q max )

[0104] where l n Indicates the data length of the node sampling, Q max Indicates the upper limit of the data queue; d n (t)∈{0,1} indicates whether the base station receives the data of node n in time slot t. If the base station receives the data of node n, then d n (t)=1, otherwise d n (t) = 0;

[0105] S13: Assuming that the channel state is time-varying, the M orthogonal channels in the network are modeled as the Gilbert-Elliott channel model; we let h n,m (t)∈{0,1} represents the state information of channel m in the time slot before time slot t; the channel state is divided into good state and bad state. If the channel state between node n and subchannel m is in good state, then h n,m (t)=1, otherwise h n,m (t) = 0; in the "bad" state, the channel is assumed to be in deep fading, so that the transmission fails with probability 1, while in the "good" state, the transmission attempt is always successful; the channel transition probability is given by P(h n,m (t+1)=1|h n,m (t) = 0) = p 01 and P(hn,m (t+1)=0|h n,m (t) = 1) = p 10 given;

[0106] S14: The AoI value of node n at the base station is represented by a n (t); If the base station receives the data packet from node n in time slot t, then a n The value of (t) will be updated to the time that has passed since the packet was generated, otherwise, a n (t+1)=a n (t)+1;a n The update process of (t) is as follows:

[0107]

[0108] where z n (t) is the AoI value at node n at time slot t. The AoI update method is different at different types of nodes. The AoI update at the periodic sampling node is related to the sampling period. Each time the sampling is completed, the AoI is updated to 0. The AoI value at the periodic sampling node is z i The iterative process of (t) is expressed as follows:

[0109]

[0110] There is a data queue in the random sampling node, and the generation time of each data packet in the data queue needs to be considered when updating; if the data queue Q j (t) is greater than l j , then the information age of the subsequent data of the head data in the data queue is Where s is the subsequent data identifier; the value of AoI at the randomly sampled node is z j (t) The iterative process is expressed as follows:

[0111]

[0112] S15: Assume that the energy consumption of the node is mainly consumed by the process of sending data packets to the base station through the sub-channel, ignoring the power consumption of the node in idle mode; the transmission rate of the sub-channel is fixed to R n , let P n They represent the transmission power of node n respectively; then the energy consumption of node n in time slot t is:

[0113]

[0114] S16: The data sampled by the periodic sampling node needs to be delivered with a delay of k as much as possible. i (t) is less than the deadline X i (t), if k i(t)≤X i , which means that the data is delivered within the deadline, otherwise it is considered overdue; define c i (t)∈{0,1} is whether a data overdue event occurs at the periodic sampling node i in time slot t. i (t) = X i +1 when let c i (t)=1, it is overdue, otherwise c i (t) = 0; data delivery delay k i (t) The change process is shown as follows:

[0115] k i (t+1)=(1-d i (t))(k i (t)+1)

[0116] Order o i (t) represents the overdue rate of node i from the initial time slot to time slot t, which is expressed as:

[0117]

[0118] S17: For data generated by random sampling nodes, the long-term average throughput constraint needs to be considered to maintain the update quality; the long-term average throughput of random sampling nodes in the network is defined as:

[0119]

[0120] The average throughput constraint of randomly sampled nodes is expressed as:

[0121]

[0122] where δ min represents the minimum throughput requirement of randomly sampled node n;

[0123] S18: Under the constraints of the probability of periodic data delivery delay exceeding the deadline and the minimum average throughput of randomly arriving nodes, the optimization problem of optimizing the long-term average information age and energy consumption of the system is expressed as follows:

[0124]

[0125] where Γ n represents the overdue rate threshold of periodic sampling node n, β is the energy consumption weight coefficient, I{V n (t) = m} is V n (t) = indicator function of event occurrence, if V n (t)=m, then I{V n(t)=m}=1, otherwise I{V n (t)=m}=0.

[0126] Furthermore, S2 specifically includes the following steps:

[0127] S21: According to the Lyapunov optimization method, define a virtual queue Z with an overdue rate for each periodic sampling node. i (t), represents the accumulation of overdue events; its update process is as follows:

[0128] Z i (t+1)=max{Z i (t)+c i (t)-Γ i s i (t),0}

[0129] S22: Define a throughput virtual queue Y for each randomly sampled node j (t), represents the gap between the average throughput and the minimum throughput requirement; its update rule is:

[0130] Y j (t+1)=max{Y j (t)+δ min -d j (t)l j ,0}

[0131] S23: Based on the states of the two types of queues, a composite Lyapunov function L(t) is defined to measure the congestion level of the queue:

[0132]

[0133] where ω j >0 is the weight coefficient of the throughput virtual queue; the Lyapunov drift term is defined as the expected value of the change of the Lyapunov function in a time slot, which is expressed as follows:

[0134]

[0135] S24: Based on Lyapunov optimization theory, a control strategy that can minimize the original optimization problem with constraints is obtained by minimizing the penalty plus drift function. The penalty plus drift function is expressed as follows:

[0136]

[0137] Where V represents the weight parameter that measures the importance of the penalty function. By changing the value of V, the desired trade-off between the size of the queue backlog and the objective function value is obtained;

[0138] S25: According to Lyapunov optimization theory, the objective function of the optimization problem P1 is minimized under the constraints by minimizing the upper bound of the drift plus penalty. The optimization problem is transformed into the following form:

[0139]

[0140] S26: According to the above optimization problem, the state space s(t) of the established system is:

[0141] S(t)={z(t),a(t),e(t),Z(t),Y(t)}

[0142] Where z(t)=(z1(t),z2(t),...z n (t)), a(t)=(a1(t),a2(t),...a n (t)), e(t)=(e1(t),e2(t),...e n (t)), Z(t)=(Z1(t),Z2(t),...Z i (t)), Y(t)=(Y1(t),Y2(t),...Y j (t));

[0143] S27: The action space A(t) of the established system is:

[0144] A(t)={U(t),V(t)}

[0145] Among them, U(t)=(U1(t),U2(t),...U n (t)), V(t)=(V1(t),V2(t),...V n (t));

[0146] S28: The reward function r(t) of the established system is:

[0147]

[0148] Furthermore, S3 specifically includes the following steps:

[0149] S31: Based on a dual-head actor network containing an update scheduling head and a channel allocation head, the current system state s(t) is input into the network and features f(S(t)) are extracted through a shared feature extraction layer; the update scheduling head outputs the strategy distribution of the updated scheduling action based on f(S(t)). Sampling the strategy distribution to obtain the updated scheduling action U(t), and then input the feature f(S(t)) and the scheduling action to obtain the channel allocation strategy distribution

[0150] S32: Sampling strategy distribution obtains actions and executes to obtain the next state;

[0151] Calculate the Actor network loss function Loss through the current state, next state, and action a (θ) and the loss function of the Critic network Loss c (θ):

[0152]

[0153] Loss c (θ)=Ω 2 (t)=(r(t)+γQ(S(t+1),A(t+1))-Q(S(t),A(t))) 2

[0154] in To update the scheduling head gradient, The gradient of the scheduling head is assigned to the channel, Ω(t) is the TD error, Q(S, A) is the Q value estimated by the critic network; B(t) is the advantage function, which is approximated by the TD error;

[0155] S33: Update the network parameters according to the gradient descent method, wait for the network to reach the termination condition, and obtain a scheduling method for hybrid industrial wireless networks.

[0156] The present invention jointly optimizes the long-term average information age and energy consumption of the system under the conditions of different node constraints in the hybrid industrial wireless network, ensuring the freshness of information in the network, reducing the energy consumption of nodes, and improving the network life.

[0157] This paper proposes a link scheduling method based on information age. By leveraging Lyapunov optimization theory, it transforms long-term overdue rate and throughput constraints into immediate subproblems per time slot, simplifying the problem solution. Considering the high-dimensional nature of multi-channel and heterogeneous data streams, we model the problem as a Markov decision process and design a reinforcement learning algorithm based on a modified actor-critic approach.

[0158] Figure 1 The hybrid industrial wireless network structure diagram constructed by the present invention includes a base station and N sensor nodes, wherein the sensor nodes are composed of K periodic sampling nodes and NK random sampling nodes; i∈φ P ={1,2,…,K} and j∈φ R= {K+1, K+2, …, N} denote the indices of periodically sampled nodes and randomly sampled nodes, respectively. Let n∈φ = {1, 2, …, N} denote the indices of all nodes. Assume that network time is divided into multiple time slots, and the maximum operating time of the entire network is T time slots. Let t∈{1, 2, …, T} denote the time slot index. The network contains M mutually non-interfering orthogonal subchannels, denoted by m∈{1, 2, …, M}.

[0159] Figure 2 This is a schematic diagram of the network channel model of the present invention. The channel state in the network is divided into good state and bad state. If the channel state between node n and subchannel m is in good state, then, otherwise. In the "bad" state, it is assumed that the channel is in deep fading, so that the transmission fails with probability 1, while in the "good" state, the transmission attempt is always successful. The channel transition probability is given by P(h n,m (t+1)=1|h n,m (t) = 0) = p 01 and P(h n,m (t+1)=0|h n,m (t) = 1) = p 10 given.

[0160] Figure 3 The present invention provides a scheduling network training flow chart of the joint optimization scheduling method for industrial wireless networks based on information age, which specifically includes the following steps:

[0161] V1-V4: Obtain the system parameter information of the hybrid industrial wireless network, construct the scheduling network structure and initialize the network information, and then initialize the state space, action space and cost function in the network.

[0162] V5-V8: Input the state space value in the current time slot into the two-head actor network to extract features. The scheduling head calculates the scheduling action strategy and samples the action. The scheduling action and features are input to the channel allocation head to output the channel allocation strategy and sample the channel allocation action.

[0163] V9-V11: After executing an action in the network, the next state and reward can be obtained. The TD error is calculated for the values in the state space and action space and the network parameters are updated using the gradient descent method.

[0164] V12-V14: After the network reaches the termination condition, the trained network parameters are saved and a scheduling network is generated for decision-making. The system uses this network to analyze the characteristics of the current state and make the optimal decision based on the analysis results in the current time slot.

[0165] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.

Claims

1. A method for optimizing information age and energy consumption in a hybrid industrial wireless network, characterized by: The method specifically comprises the following steps: S1: Obtain system information of the hybrid industrial wireless network, establish a system model based on the network structure, and construct an optimization problem to minimize the system's long-term average information age and energy consumption under the constraints of periodic data overdue rate and aperiodic data throughput; S2: The optimization problem is converted using the Lyapunov optimization method. The state space, action space, and cost function of the system are then established based on the relevant system parameters. This transforms the link scheduling optimization problem into a constrained Markov decision process. S3: Use an improved two-headed actor network to obtain strategies and utilize a critic network to evaluate value. Specifically, the current state is input into the two-headed actor network to obtain updated scheduling and channel allocation strategies. Decisions are obtained through policy distribution sampling and actions are executed to obtain the next state. The link scheduling method is obtained by waiting for the network's loss function to stabilize.

2. The method for optimizing information age and energy consumption in a hybrid industrial wireless network according to claim 1, characterized in that: The S1 specifically includes the following steps: S11: The hybrid industrial wireless network system includes a base station and N sensor nodes; the sensor nodes are composed of K periodic sampling nodes and NK random sampling nodes; using i∈φ P ={1,2,…,K} and j∈φ R ={K+1,K+2,…,N} denote the indices of periodic sampling nodes and random sampling nodes respectively, and n∈φ={1,2,…,N} is used to denote the index of all nodes; assuming that the time in the network is divided into multiple time slots and the maximum running time of the entire network is T time slots, let t∈{1,2,…,T} denote the index of the time slot; there are M orthogonal subchannels in the network that do not interfere with each other, and these subchannels are denoted by m∈{1,2,…,M}; spectrum resources are limited, and the number of nodes K in the network is greater than the number of subchannels M; let U n (t)∈{0,1} indicates whether the base station selects node n to send data at the beginning of time slot t, U n (t) = 1 means node n is scheduled and U n (t) = 0 means that node n is not scheduled; in order to prevent channel interference and thus transmission conflicts, it is assumed that the node update scheduling strategy in each time slot is limited by the number of available sub-channels; the node channel scheduling strategy uses V n (t)∈{0,1,2,…M} represents, V n (t) = m means that node n uses subchannel m to send data in time slot t. If node n does not send data, then V n (t) = 0; each node can only send a data packet through one sub-channel in each time slot; S12: If the node collects data, then g n (t)=1, otherwise g n (t) = 0, and at most one data packet is generated in each time slot; the sampling period of the periodic sampling node is H i ; The periodic sampling node uses a single packet buffer, and the newly generated data will replace the old data; let g n (t)∈{0,1} represents whether node n generates and collects data packets in time slot t; for random sampling nodes, it is assumed that their data sampling method obeys the parameter λ j ∈(0,1] Bernoulli distribution; random sampling nodes store data packets through the data queue, let Q j (t) represents the amount of data in the transmission queue of the randomly sampled node j at time slot t; the data buffer follows the first-in-first-out principle and gives priority to the first-arrived data. The evolution process of the data queue is as follows: Q j (t+1)=max(Q j (t)+g j (t)L j -d j (t)l j ,Q max ) where l n Indicates the data length of the node sampling, Q max Indicates the upper limit of the data queue; d n (t)∈{0,1} indicates whether the base station receives the data of node n in time slot t. If the base station receives the data of node n, then d n (t)=1, otherwise d n (t) = 0; S13: Assuming that the channel state is time-varying, the M orthogonal channels in the network are modeled as the Gilbert-Elliott channel model; we let h n,m (t)∈{0,1} represents the state information of channel m in the time slot before time slot t; the channel state is divided into good state and bad state. If the channel state between node n and subchannel m is in good state, then h n,m (t)=1, otherwise h n,m (t) = 0; in the "bad" state, the channel is assumed to be in deep fading, so that the transmission fails with probability 1, while in the "good" state, the transmission attempt is always successful; the channel transition probability is given by P(h n,m (t+1)=1|h n,m (t) = 0) = p 01 and P(h n,m (t+1)=0|h n,m (t) = 1) = p 10 given; S14: The AoI value of node n at the base station is represented by a n (t); If the base station receives the data packet from node n in time slot t, then a n The value of (t) will be updated to the time that has passed since the packet was generated, otherwise, a n (t+1)=a n (t)+1;a n The update process of (t) is as follows: where z n (t) is the AoI value at node n at time slot t. The AoI update method is different at different types of nodes. The AoI update at the periodic sampling node is related to the sampling period. Each time the sampling is completed, the AoI is updated to 0. The AoI value at the periodic sampling node is z i The iterative process of (t) is expressed as follows: There is a data queue in the random sampling node, and the generation time of each data packet in the data queue needs to be considered when updating; if the data queue Q j (t) is greater than l j , then the information age of the subsequent data of the head data in the data queue is Where s is the subsequent data identifier; the value of AoI at the randomly sampled node is z j (t) The iterative process is expressed as follows: S15: Assume that the energy consumption of the node is mainly consumed by the process of sending data packets to the base station through the sub-channel, ignoring the power consumption of the node in idle mode; the transmission rate of the sub-channel is fixed to R n , let P n They represent the transmission power of node n respectively; then the energy consumption of node n in time slot t is: S16: The data sampled by the periodic sampling node needs to be delivered with a delay of k as much as possible. i (t) is less than the deadline X i (t), if k i (t)≤X i , which means that the data is delivered within the deadline, otherwise it is considered overdue; define c i (t)∈{0,1} is whether a data overdue event occurs at the periodic sampling node i in time slot t. i (t) = X i +1 when c i (t)=1, it is overdue, otherwise c i (t) = 0; data delivery delay k i (t) The change process is shown as follows: k i (t+1)=(1-d i (t))(k i (t)+1) Order o i (t) represents the overdue rate of node i from the initial time slot to time slot t, which is expressed as: S17: For data generated by random sampling nodes, the long-term average throughput constraint needs to be considered to maintain the update quality; the long-term average throughput of random sampling nodes in the network is defined as: The average throughput constraint of randomly sampled nodes is expressed as: where δ min represents the minimum throughput requirement of randomly sampled node n; S18: Under the constraints of the probability of periodic data delivery delay exceeding the deadline and the minimum average throughput of randomly arriving nodes, the optimization problem of optimizing the long-term average information age and energy consumption of the system is expressed as follows: where Γ n represents the overdue rate threshold of periodic sampling node n, β is the energy consumption weight coefficient, I{V n (t) = m} is V n (t) = indicator function of event occurrence, if V n (t)=m, then I{V n (t)=m}=1, otherwise I{V n (t)=m}=0.

3. The method for optimizing information age and energy consumption in a hybrid industrial wireless network according to claim 2, characterized in that: The S2 specifically includes the following steps: S21: According to the Lyapunov optimization method, define a virtual queue Z with an overdue rate for each periodic sampling node. i (t), represents the accumulation of overdue events; its update process is as follows: Z i (t+1)=max{Z i (t)+c i (t)-Γ i s i (t),0} S22: Define a throughput virtual queue Y for each randomly sampled node j (t), represents the gap between the average throughput and the minimum throughput requirement; its update rule is: Y j (t+1)=max{Y j (t)+δ min -d j (t)l j ,0} S23: Based on the states of the two types of queues, a composite Lyapunov function L(t) is defined to measure the congestion level of the queue: where ω j >0 is the weight coefficient of the throughput virtual queue; the Lyapunov drift term is defined as the expected value of the change of the Lyapunov function in a time slot, which is expressed as follows: S24: Based on Lyapunov optimization theory, a control strategy that can minimize the original optimization problem with constraints is obtained by minimizing the penalty plus drift function. The penalty plus drift function is expressed as follows: Where V represents the weight parameter that measures the importance of the penalty function. By changing the value of V, the desired trade-off between the size of the queue backlog and the objective function value is obtained; S25: According to Lyapunov optimization theory, the objective function of the optimization problem P1 is minimized under the constraints by minimizing the upper bound of the drift plus penalty. The optimization problem is transformed into the following form: S26: According to the above optimization problem, the state space s(t) of the established system is: S(t)={z(t),a(t),e(t),Z(t),Y(t)} where, z(t) = (z1(t), z2(t),... z n (t)), a(t) = (a1(t), a2(t),... a n (t)), e(t) = (e1(t), e2(t),... e n (t)), Z(t) = (Z1(t), Z2(t),... Z i (t)), Y(t) = (Y1(t), Y2(t),... Y j (t)); S27: The action space A(t) of the established system is: A(t)={U(t),V(t)} Among them, U(t) = (U1(t), U2(t),... U n (t)), V(t) = (V1(t), V2(t),... V n (t)); S28: The reward function t(t) of the established system is:

4. The method for optimizing information age and energy consumption in a hybrid industrial wireless network according to claim 3, characterized in that: The S3 specifically includes the following steps: S31: Based on a dual-head actor network containing an update scheduling head and a channel allocation head, the current system state s(t) is input into the network and features f(S(t)) are extracted through a shared feature extraction layer; the update scheduling head outputs the strategy distribution of the updated scheduling action based on f(S(t)). Sampling the strategy distribution to obtain the updated scheduling action U(t), and then input the feature f(S(t)) and the scheduling action to obtain the channel allocation strategy distribution S32: The sampling strategy distributes the acquisition action and executes it to obtain the next state. Calculate the Actor network loss function Loss through the current state, next state, and action a (θ) and the loss function of the Critic network Loss c (θ): Loss c (θ)=Ω 2 (t)=(r(t)+γQ(S(t+1),A(t+1))-Q(S(t),A(t))) 2 in To update the scheduling head gradient, The gradient of the scheduling head is assigned to the channel, Ω(t) is the TD error, Q(S, A) is the Q value estimated by the critic network, and B(t) is the advantage function, which is approximated by the TD error. S33: Update the network parameters according to the gradient descent method, wait for the network to reach the termination condition, and obtain a scheduling method for hybrid industrial wireless networks.

5. The method for optimizing information age and energy consumption in a hybrid industrial wireless network according to claim 3, characterized in that: The channel allocation strategy includes a conflict detection mechanism when it is executed. Specifically, when multiple nodes are assigned to the same sub-channel, the priority is to select the node that makes Z i (t)(c i (t)-Γ i )+ω j Y j (t)(δ min -d j (t)l j ) value obtains the channel usage right, and the remaining nodes are set to the unscheduled state in this time slot.

6. The method for optimizing information age and energy consumption in a hybrid industrial wireless network according to claim 3, characterized in that: The weight coefficient ω of the throughput virtual queue j The determination method is: according to the random sampling node data queue length Q j (t) Dynamic adjustment, when Q j (t)>0.8Q max When, ω j =2ω0; when 0.5Q max <Q j (t)≤0.8Q max When ω j =1.5ω0; other casesω j =ω0, where ω0 is the preset basic weight.

7. The method for optimizing information age and energy consumption in a hybrid industrial wireless network according to claim 3, characterized in that: During the training process of the two-headed Actor network, the policy update priority of updating the scheduling head is higher than that of the channel allocation head. Specifically, during gradient backpropagation, the learning rate of updating the scheduling head is set to 1.2 to 1.5 times the learning rate of the channel allocation head.

8. The method for optimizing information age and energy consumption in a hybrid industrial wireless network according to claim 3, characterized in that: The method for determining the energy consumption optimization weight β is as follows: when the base station detects that the remaining power of a node is less than 20%, the β value corresponding to the node is automatically increased to β max , and according to the formula β=β max ×(1-SOC / 100) dynamic adjustment, where SOC is the current power percentage of the node.

9. The method for optimizing information age and energy consumption in a hybrid industrial wireless network according to claim 4, characterized in that: The sampling process of the strategy distribution includes an exploration-exploitation balance mechanism, specifically: ε=0.3 in the ε-greedy strategy at the beginning of training, and decays according to ε←0.99ε after every 1000 training iterations until ε drops to 0.05 and remains constant.

10. The method for optimizing information age and energy consumption in a hybrid industrial wireless network according to claim 3, characterized in that: The overdue rate virtual queue Z n (t) and throughput virtual queue Y n The initialization method of (t) is: Z n (t) = ceil(Γ n T total / H n ), Y n (0) = floor(δ min T total / l n ), where T total is the total number of running slots preset by the system, ceil() is the rounding-up function, and floor() is the rounding-down function.