An edge device intelligent sensing method and system based on information freshness

By optimizing the perception strategy of edge devices through Markov decision process and deep Q learning algorithm, the problems of information transmission timeliness and energy constraints in passive Internet of Things are solved, and a balance between real-time information and energy efficiency is achieved in complex environments.

CN119907021BActive Publication Date: 2025-10-14HUAZHONG NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411752487.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-02
Publication Date
2025-10-14
Estimated Expiration
2044-12-02

AI Technical Summary

Technical Problem

In the passive Internet of Things, how to design intelligent perception scheduling strategies for IoT edge devices to maximize the timeliness and freshness of information transmission under time-varying energy constraints and noise interference environments, especially how to accurately calculate and optimize information age under energy-constrained conditions.

Method used

Markov decision process modeling is adopted, combined with Lagrange multiplier method and deep Q-learning algorithm to optimize the sampling frequency, sample size, transmission power and bandwidth allocation of edge devices. The balance between energy consumption and information age is dynamically adjusted through Lagrange multipliers, and the optimal perception strategy is learned through deep Q-learning algorithm.

Benefits of technology

Under energy-constrained conditions, it effectively reduces the age of information, ensures the real-time and freshness of data, and is suitable for IoT applications in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119907021B_ABST
    Figure CN119907021B_ABST
Patent Text Reader

Abstract

The application discloses an edge device intelligent sensing method and system based on information freshness, constructs a Markov decision process model by analyzing key factors in the edge device intelligent sensing system. Secondly, the system state is updated in real time through dynamic adjustment of the sensing state, the parameters of the device are adjusted according to the new environmental changes, so that the system can quickly respond in a complex environment. In the aspect of joint optimization, the Lagrange multiplier method is used to jointly optimize the information age and energy consumption, so as to maximize the freshness of information while meeting the energy constraint. Finally, through the deep Q network, the system can continuously optimize the decision strategy through the experience replay mechanism, so that the edge device can minimize the information age under the condition of energy limitation in the changing environment. The application improves the timeliness of the sensing information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to fields such as passive Internet of Things and artificial intelligence, and in particular to an edge device intelligent perception method and system based on information freshness. Background Art

[0002] For many IoT applications, obtaining the latest information is crucial. With the prevalence of ubiquitous sensing devices and wireless data connections, real-time condition monitoring has become a reality in large-scale cyber-physical systems such as power grids, manufacturing facilities, and intelligent transportation systems. However, the unprecedented high dimensionality and speed at which sensor data is generated also pose severe challenges to its timely delivery. At the edge of the network, multiple source nodes of different applications collect information samples from the physical environment and forward this sample information to the edge server. The collected information can be processed and stored locally (edge ​​computing) and / or forwarded to the cloud. Since many applications in the upper layer rely on the timeliness of the sampled information, it is critical to forward the freshest information to the destination as quickly as possible, which is the problem of minimizing the age of information.

[0003] In the application of passive sensor monitoring, consider such a scenario (such as Figure 1 As shown in Figure 1: One or more edge devices (energy harvesting sensors) continuously monitor a system and send timestamped status updates (perception information) to the destination. The destination tracks the system status through the received perception information. We use the age of information (AoI), which is the time that has passed since the last update was received, to measure the freshness of the state information available to the destination. However, due to time-varying energy constraints and battery limitations, the calculation of the freshness of perception information is challenging, and the channel between the source and the destination is noisy, and each transmitted update may fail independently with a constant probability. Therefore, how to accurately design an intelligent perception scheduling strategy for IoT edge devices based on the time-varying energy constraints, AoI, the influence of various interference environments and their own characteristics in passive IoT is a crucial issue. Summary of the Invention

[0004] The present invention relates to an intelligent perception method for edge devices based on information freshness, aiming to optimize the information update strategy of edge devices under energy-constrained conditions to maximize the timeliness of data transmission.

[0005] To achieve the above objectives, the present invention adopts a technical solution: an edge device intelligent perception method based on information freshness, comprising the following steps:

[0006] Step 1: Analyze the key factors in the edge device intelligent perception system and then model the Markov decision process based on these factors.

[0007] Step 2, optimize and calculate the state in the Markov decision process, update the state space through PDS, model and update the perception state in advance using known information, and obtain the post-decision state;

[0008] Step 3, add the energy consumption factor to the reward in the Markov decision process through the Lagrange multiplier, jointly optimize the energy consumption and the information age AoI target using the Lagrange multiplier method, integrate the energy constraint and the AoI target into the same Lagrange optimization problem in the optimization process, and dynamically adjust the value of the Lagrange multiplier to effectively control the AoI; the Lagrange multiplier is updated in each iteration process to achieve dynamic control of energy consumption;

[0009] Step 4, solve and calculate the established Markov decision process model, learn the optimal perception strategy through the deep Q learning algorithm, in each time step, the deep Q learning algorithm estimates the action value through the neural network, and selects the action to be executed according to the action value, thereby updating the state and generating the optimal strategy to minimize the information age and energy consumption.

[0010] Further, the key factors in step 1 include task queue, sampling model, energy consumption model, energy arrival model, and channel transmission model, wherein the task queue includes the remaining data amount of the head task and the queue length, the sampling model includes the sampling frequency and the sample size, and the channel transmission model includes the bandwidth, the number of discrete levels / modulation order; the energy consumption model includes the perception energy consumption and the transmission energy consumption.

[0011] Further, the four elements of the Markov decision process include state, action, reward, and state transition, the action space includes sampling frequency, sample size, number of discrete levels, and bandwidth, and the state vector includes the remaining data amount of the head task, the task queue length, and the information age, wherein the state vector is s=[d r ,A,q], wherein d r represents the remaining data amount of the head task, A represents the information age, and q represents the task queue length; the energy consumption of the device is considered in the reward process, including the perception energy consumption and the transmission energy consumption;

[0012] The perception energy consumption is represented as:

[0013] E 采样 =P 采样 ·T 采样

[0014] Wherein P 采样 is the power consumption of the sensor, and T 采样 is the total sampling time;

[0015]

[0016] where f 采样 is the sampling frequency, N 样本 is the total number of samples;

[0017] The energy consumption of data transmission is expressed as:

[0018] E 传输 =P 发送 ·T 传输

[0019] Among them, P 发送 is the power consumption of the communication module, T 传输 is the data transmission time, D 样本 is the sample size, P 传输速率 is the transmission rate.

[0020] Furthermore, the specific implementation process of PDS is divided into the following steps:

[0021] First, extract the remaining data d of the first task from the input state vector and action vector r , task queue length q, information age A and sampling frequency f, sample size N, transmission power P and bandwidth allocation W;

[0022] Then, the transmission rate is calculated to determine whether the task is completed. Without considering noise, the maximum transmission rate is determined by the number of discrete levels and bandwidth. According to the Nyquist criterion, it can be expressed as: c = 2W*log2M;

[0023] Where W is the bandwidth allocated to the device and M is the number of discrete levels of the signal;

[0024] Then, the state variables are updated according to whether the task is completed, the post-decision state is generated, and the task demand is updated: if the task is completed, the remaining data volume d r Return to zero, otherwise update to d r -d, that is, the amount of remaining data of the new first task in the queue, d is the amount of task completion, d = Δt*2W*log2M, Δt is the interval time; update information age: if the task is completed, the information age A is reset to 0, otherwise it is accumulated by 1; update task queue information: if the task is completed, the task queue information q is reduced by 1, indicating that a task is completed; the channel gain remains unchanged: the channel gain h is affected by the external environment, but remains unchanged in a short time; finally, the decision state after combination generation: the updated variable pds A , pds q Combined into a new state vector s′, the specific formula is as follows:

[0025]

[0026] pdsA =A+1

[0027] pds q =q-1

[0028]

[0029] Furthermore, the specific implementation of step 3 is as follows:

[0030] First, a cost function C is defined to quantify the total cost of information age and energy consumption. The cost function is defined as follows:

[0031] C = ∑(A + λ·(E-Emax))

[0032] Where A represents the information age, E represents the current energy consumption, and Emax is the maximum energy threshold; λ is the Lagrange multiplier used to dynamically adjust the penalty intensity for energy overspending;

[0033] Then update the Lagrange multiplier λ. At each time step, adjust the value of λ according to the deviation of energy consumption relative to the maximum energy threshold. The update formula of the Lagrange multiplier is:

[0034] λ′=λ+θ(t)·(E-Emax)

[0035] where θ(t) is a step-size control parameter that is dynamically adjusted with time step t; if E > Emax, λ increases to suppress future energy overspending; otherwise, λ decreases;

[0036] In order to ensure the rationality of the Lagrange multiplier value, it is constrained at each time step to keep it within a certain range. This process is expressed as:

[0037] λ″=max(0,min(λ′,λmax))

[0038] where λmax is the upper limit of the Lagrange multiplier.

[0039] Furthermore, in step 4, the deep Q-learning algorithm DQN uses a deep neural network to estimate the Q-value function Q(s, a; θ), where s is the current state, a is the possible action combination, and θ is the network parameter. The specific implementation process is as follows:

[0040] First, define a deep neural network denoted as Q-network, whose input layer receives the current state s, including task demand quantity dr, information age A, and task queue information q. After the input layer, there are multiple fully connected layers, each with a number of neurons, using ReLU activation function to introduce nonlinearity and prevent gradient vanishing. The fully connected layers abstract and combine input features continuously, enabling the network to accurately estimate Q values for different state-action pairs. The final output layer contains nodes corresponding to the action space, and the value of each node is the Q value of a specific action under the current state.

[0041] Then define the target Q value: after performing the current action, the system will reach the next state s' and obtain the immediate reward r. Update the Q value using the Bellman equation, and define the target Q value as:

[0042]

[0043] where γ is the discount factor, used to weigh future rewards and immediate rewards; Q(s',a';θ - ) is the maximum Q value of the next state s' calculated by the target Q network, θ - represents the parameters of the target Q network, and a' represents the set of all possible actions in the next state s';

[0044] Then calculate TD 误差 : use mean square error MSE to measure the difference between the Q network output and the target Q value, and define the time difference TD 误差 as:

[0045] TD 误差 = (y - Q(s, a; θ)) 2

[0046] Here, y is the target Q value, and Q(s, a; θ) is the predicted value of the Q network.

[0047] Finally, update the Q network by backpropagation, and use gradient descent to minimize TD 误差 , adjust the network weights by calculating the gradient of the loss function with respect to the network parameters θ:

[0048]

[0049] where β is the learning rate. Through backpropagation, the Q network gradually learns and updates its parameters to improve the accuracy of Q value estimation, is the gradient, representing the derivative of the loss function TD 误差 2 with respect to the network parameters θ.

[0050] Furthermore, in order to improve the training efficiency and stability, an experience replay mechanism is adopted to store the state, action, reward, and next state quadruple (s, a, r, s′) of each step into the experience replay buffer, from which small batch samples are randomly extracted for training.

[0051] Furthermore, every fixed number of steps, the parameters of the current Q network are synchronized to the target Q network to ensure the stability of the target Q value, thereby preventing instability during training.

[0052] Furthermore, a ∈-greedy strategy is adopted during the training process. Specifically, an action is randomly selected with probability ∈, and the action with the largest Q value is selected with probability 1-∈. As the training progresses, ∈ is gradually reduced to make more use of the learned strategy.

[0053] The present invention also provides an edge device intelligent perception system based on information freshness, including a processor and a memory, the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute an edge device intelligent perception method based on information freshness as described in the above technical solution.

[0054] This invention analyzes key factors in edge device intelligent perception systems and constructs a Markov decision process model to intelligently optimize decision variables such as device sampling frequency, sample size, transmission power, and bandwidth allocation. Secondly, the method dynamically adjusts the perception state to update the system state in real time, adjusting device parameters based on new environmental changes, enabling the system to respond quickly in complex environments. Regarding joint optimization, the invention utilizes the Lagrange multiplier method to jointly optimize information age and energy consumption, striving to maximize information freshness while meeting energy constraints. Based on a set cost function, the system dynamically adjusts the Lagrange multiplier to balance information updates with energy consumption, avoiding performance degradation caused by excessive energy consumption. Finally, combined with a deep Q-learning algorithm, the invention can adaptively learn optimal perception strategies in practical applications. Through the deep Q-network, the system can continuously optimize decision strategies through an experience replay mechanism, enabling edge devices to minimize information age within energy-constrained conditions in a constantly changing environment. This invention improves the timeliness of perception information. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 This is an application scenario diagram in the background technology;

[0056] Figure 2 This is a flow chart of an embodiment of the present invention. DETAILED DESCRIPTION

[0057] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings.

[0058] like Figure 2 As shown, the present invention provides an intelligent edge device perception method based on information freshness. This method optimizes the information update strategy of edge devices under energy-constrained conditions to maximize the timeliness of transmitted data, ensure real-time data, and meet the demand for up-to-date information in IoT applications. This method achieves intelligent control of information freshness through four steps: analytical modeling of key factors affecting information age, perception state adjustment, joint optimization of energy consumption and information age, and deep learning strategies.

[0059] First, key factors in the intelligent perception system of edge devices (such as those affecting the age of information (AoI)) are analyzed and modeled. Then, based on these factors, a Markov decision process is modeled. Key factors include the task queue (the amount of remaining data for the first task in the queue, queue length), sampling model (sampling frequency and sample size), energy consumption model, energy arrival process, and channel transmission model (bandwidth, number of discrete levels / modulation order). The core elements of the Markov decision process include state, action, state transition, and reward. The state is calculated using the energy arrival model, channel transmission model, and battery model. The state vector includes the amount of remaining data for the first task in the queue, task queue length, and AoI. The action is derived from the intelligent perception strategy, and the action vector includes sampling frequency, sample size, number of discrete levels, and bandwidth. By modeling these factors, a model describing the changes in different device states is constructed, providing a foundation for subsequent perception strategy optimization.

[0060] Next, the method constructs a post-decision state (PDS) by sensing state adjustments to achieve dynamic updates of the system state. After each decision is executed, such as adjusting parameters such as frequency, power, or bandwidth, the system will be updated according to the new state. With the introduction of PDS, the system can flexibly adjust the state without relying on complete information, so that it can adapt to changes in the external environment in a timely manner. The state space includes variables such as task requirements, battery power, and channel capacity. The use of PDS allows the system to respond to different combinations of states and actions, achieving efficient modeling of the state space. In this way, the system can ensure the stability of state changes in complex environments and lay the foundation for subsequent optimization steps.

[0061] In the joint optimization step, the method introduces the Lagrange multiplier method to reduce the information age under energy constraints. Specifically, the system combines information age and energy consumption into a cost function, which includes the cost of information age and a penalty term for exceeding the energy limit. When energy consumption exceeds the set threshold, the system increases the penalty, driving it to reduce energy consumption to meet the energy constraint. In addition, by dynamically adjusting the value of the Lagrange multiplier, the system can appropriately increase the penalty when energy consumption is too high, and reduce it when it is too low, to balance the relationship between energy and information age. This method ensures that under energy constraints, the real-time nature of the data is maintained at a lower energy consumption, so that it can find the optimal balance between information updating and energy consumption.

[0062] Finally, this method uses a deep Q-learning algorithm to learn the optimal perception policy. A deep Q-network (DQN) estimates the value of each possible action through a neural network and selects the optimal action that maximizes the reward. During training, the system records data for each state, action, reward, and next state. Using an experience replay mechanism, it randomly extracts data for training, thereby breaking down correlations between data and improving training stability. The DQN network continuously updates network weights and gradually optimizes the policy by calculating temporal differential errors and performing backpropagation. An epsilon-greedy strategy is employed during training to strike a balance between exploring new actions and leveraging existing knowledge, gradually reducing the emphasis on exploration as training progresses. Through deep Q-learning, the system gradually learns the optimal decision under different states, achieving dynamic policy optimization in complex environments, effectively reducing information age, and meeting energy constraints.

[0063] This approach forms a complete intelligent perception method through analytical modeling of key influencing factors, dynamic adjustment of post-decision states, joint optimization using the Lagrange multiplier method, and policy optimization using deep Q-learning. It optimizes data update efficiency and maximizes information freshness for edge devices under energy constraints and fluctuating channel conditions, making it particularly suitable for IoT applications requiring real-time data support, such as smart transportation, power grid monitoring, and industrial automation.

[0064] An embodiment of the present invention provides an edge device intelligent perception method based on information freshness, such as Figure 1 As shown, the implementation of the method includes the following steps:

[0065] Step 1: Analyze and model the key factors in the edge device intelligent perception system.

[0066] Step 1: In-depth analysis and modeling of key factors in the edge device intelligent sensing system. This process is the foundation of the entire sensing strategy optimization. By accurately identifying and modeling the main factors in the edge device intelligent sensing system, data support can be provided for subsequent decision-making and optimization steps. The information age, which is the time delay from data generation to use, reflects the timeliness of the data. The system considers AoI as the target variable for optimization and considers multiple key factors that affect it, including task queue (remaining data amount of the first task in the queue, queue length), sampling model (sampling frequency and sample size), energy consumption model, energy arrival process, and channel transmission model (bandwidth, discrete level number). By analyzing the changes and interactions of these factors in different states, the system constructs a comprehensive action space and state space model to provide accurate basis for intelligent sensing strategy decision-making.

[0067] In detail, first is the action space (action vector). First is the sampling frequency of the task. The sampling frequency determines the frequency of data updates. The higher the frequency, the fresher the information obtained by the system, and thus the lower the information age. However, too high a sampling frequency will cause rapid energy consumption of the edge device, making it difficult for the system to continuously update data under energy constraints. Therefore, there is a trade-off between sampling frequency and energy consumption. Let the sampling frequency be f, the information age a is inversely proportional to the sampling frequency, which can be expressed as:

[0068]

[0069] That is, increasing the sampling frequency can reduce AoI, thereby improving the timeliness of information. Second is the sample size, third is the transmission power, and fourth is the bandwidth allocated to the device.

[0070] Second is the state space (state vector). First is the remaining data amount, second is the task queue information, and third is the information age. The channel capacity (maximum transmission rate) is calculated as:

[0071] c = 2Wlog2M

[0072] where W is the bandwidth and M is the discrete level number of the signal, also known as the symbol number or modulation order. The larger the channel capacity c, the faster the data transmission, which can reduce the information age. When constructing the state space model, the system combines the above factors to form a mathematical model that describes the key factors. Let the state vector be s = [d r ,a,q], where d r represents the remaining data amount, a represents the information age, and q represents the task queue length. The remaining data amount d rand the task queue length q are used to represent the urgency of the task and the number of current pending tasks. By constantly updating these state variables, the system can dynamically reflect the changes in the environment and select the optimal perception strategy in different states.

[0073] Finally, the energy consumption of the device. The energy consumption of the device is not considered in the action space and state space, but in the reward (state, action, reward, state transition are the four elements of Markov decision process, after establishing the Markov decision process, through deep reinforcement learning algorithm to learn Markov decision process, the edge device perception strategy can be obtained). The energy consumption of the edge device directly affects its endurance, and further affects the frequency of information update and the stability of data transmission. In the case of energy limitation, the device cannot continuously collect and transmit data at high frequency, thereby increasing the information age. The energy consumption of the device mainly includes perception energy consumption and transmission energy consumption. The perception energy consumption can be represented as:

[0074] E 采样 =P 采样 ·T 采样

[0075] where P 采样 is the power consumption of the sensor (usually in the order of milliwatt), T 采样 is the total sampling time.

[0076]

[0077] where f 采样 is the sampling frequency, N 样本 is the total sampling number. High sampling frequency will increase the number of samples per unit time, thereby increasing the sampling power consumption.

[0078] The data transmission energy consumption can be represented as:

[0079] E 传输 =P 发送 ·T 传输

[0080] where P 发送 is the power consumption of the communication module, which is affected by transmission power, channel quality, etc. T 传输 is the data transmission time, where N 样本 is the sample number, D 样本 is the sample size, P 传输速率 is the transmission rate, and the formula of the maximum transmission rate has been given in the previous text. The higher the sampling frequency, the more data generated per unit time, increasing the transmission power consumption. The larger the sample size, the more data transmitted each time, increasing the transmission time and power consumption.

[0081] To extend the effective working time of the device, the system needs to balance the energy consumption and the frequency of information update to ensure the timeliness of information.

[0082] Step 2: Perform perception state adjustment, model the state space a priori with known information, and obtain the post-decision state.

[0083] In step 2, the system state space is updated using the post-decision state (PDS). In the system state update iteration, although some factors such as the energy arrival process are unknown and cannot be calculated, some factors can be directly calculated and updated. The calculable part does not need to be learned by a deep reinforcement learning algorithm, avoiding the waste of computing resources and time resources. In particular, in scenarios that require fast transmission, low power consumption, and do not require high data reliability, the sender will send the message without waiting for the receiver's confirmation or retransmitting the message, such as the message delivery method of the MQTT-SN protocol, which is consistent with the goal of pursuing information freshness. Therefore, the channel model for quantitative analysis only considers the sender's state, such as bandwidth and the number of discrete signal levels, without considering the influence of noise, channel attenuation, etc. Since the bandwidth and the number of discrete signal levels are known, the channel model can also calculate the PDS. All key variables are extracted from the current state vector and action vector, including the remaining data volume, battery power, channel state, etc. Then, the updated values of these variables are calculated immediately after each decision, for example, the new remaining data volume, task queue length, information age, battery energy consumption are calculated by the current remaining data volume, task queue length, information age, battery energy consumption, power and bandwidth parameters. Then, based on the calculation results, part of the factors in the state vector are updated. Finally, the updated state and the unupdated state are combined into a new state vector as the post-decision state (PDS), which is passed to the subsequent step for further decision-making.

[0084] The specific implementation process of PDS can be divided into several clear steps, which gradually calculate and update the state variables to form the post-decision state, and the specific implementation is as follows:

[0085] First, extract the current state and action variables: First, we extract the remaining data volume d r , information age A, task queue information q, and sampling frequency f, sample size N, transmission power P, and bandwidth allocation W from the input state vector s = [d r , A, q] and action vector α = [f, N, P, W].

[0086] Then, calculate the transmission rate and determine whether the task is completed. Without considering noise, the maximum transmission rate is determined by the number of discrete levels and bandwidth, which can be expressed according to the Nyquist criterion: c = 2W*log2M;

[0087] Where W is the bandwidth allocated to the device, and M is the number of discrete signal levels. The Nyquist criterion states that in a noise-free channel of bandwidth W, the maximum symbol rate is 2W symbols per second. Each symbol can carry log2M bits of information (based on the number of signal levels M).

[0088] Then, the state variables are updated according to whether the task is completed, and the post-decision state is generated. Update the remaining data volume: If the task is completed (i.e. b = 1), the remaining data volume d r Return to zero, otherwise update to d r -d, that is, the amount of remaining data of the new first task in the queue. Update information age: If the task is completed, the information age A is reset to 0, otherwise it is accumulated by 1. Update task queue information: If the task is completed, the task queue information q is reduced by 1, indicating that a task is completed. Channel gain remains unchanged: The channel gain h is affected by the external environment, but it usually remains unchanged in a short time. Finally, the decision state after combination generation: the updated variable pds A , pds q Combined into a new state vector s′, the specific formula is as follows:

[0089]

[0090] pds A =A+1

[0091] pds q =q-1

[0092] s′=[pds dr , pds A , pds q ]

[0093] The generated post-decision state s′ will replace state s and provide a real-time updated state basis for the perception strategy selection in the next time step. This mechanism of updating the state space through PDS enables the system to quickly respond to changes in the external environment, thereby maintaining the efficiency of the perception strategy in complex dynamic environments.

[0094] Step 3: Use the Lagrange multiplier method to jointly optimize the system energy consumption and AoI target.

[0095] Step 3 uses the Lagrange multiplier method to jointly optimize the system's energy consumption and age of information (AoI) objectives. The main goal of this step is to minimize the information age as much as possible under energy constraints, ensuring the timeliness of the data while satisfying the energy constraint. By introducing Lagrange multipliers, we integrate the goal of minimizing the information age and the energy constraint into the same optimization problem, allowing the system to find the optimal perception strategy under the constraints. The core of this method is to construct a cost function that includes information age and energy consumption. In this function, when energy consumption exceeds a set threshold, the system will increase the penalty, thereby forcing it to optimize information freshness while limiting energy consumption.

[0096] First, we define a cost function C (one of the four elements of the Markov decision process, also known as the reward) to quantify the total cost of information age and energy consumption. Age of Information (AoI) is used to measure the freshness of the data, and we usually want to minimize AoI to ensure the real-time nature of the data. Energy consumption represents the total energy consumed by the edge device during the perception and transmission process. In practical systems, energy consumption is often strictly limited, so we construct a cost function that includes information age and a penalty term for energy exceeding the limit to balance these two objectives during the optimization process. The cost function is defined as follows:

[0097] C = ∑(A + λ·(E-Emax))

[0098] Where A represents the current AoI value, reflecting the timeliness of information in the edge device's perception system. E represents the current energy consumption, and Emax is the maximum energy threshold allowed by the system. When energy consumption E exceeds this threshold, the system applies a penalty term λ·(E-Emax), where λ is a Lagrange multiplier that dynamically adjusts the penalty intensity for energy overruns. If energy consumption exceeds the limit significantly, the Lagrange multiplier λ increases, thereby increasing the total cost and prompting the system to reduce energy consumption.

[0099] During the optimization process, updating the Lagrange multiplier is a key step. At each time step, we adjust the value of λ based on the deviation of energy consumption from the maximum energy threshold. Specifically, if the current energy consumption E exceeds the set threshold Emax, λ is increased, causing the system to focus more on energy conservation in the next step. Conversely, when energy consumption falls below the threshold, λ is appropriately decreased, allowing the system to focus more on optimizing the AoI. The update formula for the Lagrange multiplier is:

[0100] λ=λ+σ(t)·(E-Emax)

[0101] where σ(t) is a step size control parameter that dynamically adjusts at time step t, usually gradually decreasing during the training process to ensure the optimization process tends to be stable. This formula means that when E > Emax, the Lagrange multiplier will gradually increase, thus imposing a greater energy penalty; while when E≤Emax, the Lagrange multiplier will gradually decrease, reducing the constraint of energy, so that the system will use more resources for the optimization of information age.

[0102] To ensure the reasonableness of the value of the Lagrange multiplier, we also constrain it to remain within a certain range at each time step. The following restrictions are used: λ = max(0, λ) and λ = min(λ, λmax). Here, λmax is the upper limit of the Lagrange multiplier, preventing the penalty term of the cost function from being too large and losing balance. In this way, λ is limited to a range that is non-negative and not too large, thus controlling the energy consumption of the system.

[0103] Throughout the optimization process, the system will continuously adjust the Lagrange multiplier at each time step according to the current state and consumption, dynamically balancing the relationship between information age minimization and energy constraints through the Lagrange multiplier method. This joint optimization method ensures that the system can still obtain the freshest data information with lower energy consumption under energy constraints.

[0104] Step 4: Learn the optimal perception strategy through deep Q-learning algorithm.

[0105] The goal of MDP (Markov Decision Process) is to maximize the cumulative reward R (cost function) by selecting an action α from the current state s through a policy π(α∣s). Q-learning is a commonly used reinforcement learning method that guides action selection by learning the Q function Q(s,α). The formula is:

[0106] Q(s,α) = E[r + γa'maxQ(s',a')]

[0107] where r is the immediate reward of the current action; γ is the discount factor, which weighs the current and future rewards; s' is the next state (state transition). DQN uses neural network approximation Q(s,α) to solve the storage and calculation problems of Q-learning in high-dimensional state space.

[0108] Step 4 employs a deep Q-learning (DQN) algorithm to learn the optimal sensing policy for the edge device, aiming to minimize the age of information (AoI) and satisfy the energy consumption constraint. Through the DQN network, the system can select the optimal action combination in different states, effectively balancing the timeliness of information update and the restriction of energy consumption. In this process, the previously defined post-decision state (PDS) and Lagrange multiplier are introduced, enabling the DQN network to be more targeted in terms of state update and energy optimization. DQN is a deep learning-based reinforcement learning algorithm specifically designed to solve the optimal policy learning problem in Markov decision processes (MDPs). Overall, the DQN network estimates the Q-value function π(α∣s) through a deep neural network, where s is the current state, α is the possible action combination (action vector), and θ is the network parameter. At each time step, the system selects an action combination based on the current state, observes the reward and new state after execution, and stores this information in the experience replay buffer. By introducing PDS, we can update the state space immediately after each action execution without relying on complete external feedback, thereby improving the accuracy of Q-value estimation. Additionally, through the Lagrange multiplier method, we add a penalty term for energy consumption in the cost function to ensure that the AoI is minimized while satisfying the energy constraint. The Q-value network gradually approximates the optimal Q-value during training, enabling the learning of the best sensing policy in complex environments.

[0109] The following are the core steps of DQN learning MDP:

[0110] (1) Initialization. First, construct the neural network structure of DQN. The input layer of this network receives the current state s, including the remaining data volume dr, the information age A, and the task queue information q. These information describe the current state of the edge device, providing comprehensive state information for the network. After the input layer, multiple fully connected layers are connected, each with a number of neurons, usually using ReLU activation function to introduce nonlinearity and prevent gradient vanishing. Fully connected layers abstract and combine input features continuously, enabling the network to accurately estimate the Q-value of different state-action pairs. The final output layer contains nodes corresponding to the action space, and the value of each node is the Q-value of a specific action under the current state.

[0111] (2) Data collection. At each time step t, an action is selected from the current state, the action is executed, and the environment is observed, storing the experience. At each time step, the system selects an action α based on the current state s t Select an action α through the ε-greedy strategy t, that is, randomly select actions with probability ∈ to increase the diversity of exploration; select the action with the highest Q value in the current Q network with probability 1-∈ to utilize existing knowledge. Execute action α t After that, the system observes the immediate reward r t and the next state s t+1 . Instant Rewards t Typically defined as a penalty term that includes both the negative AoI and energy consumption, it encourages the system to choose actions that optimize information freshness while conserving energy. In the reward function, the Lagrange multiplier λ adjusts the energy consumption penalty term, allowing the system to dynamically balance AoI and energy consumption under varying energy conditions.

[0112] In order to achieve dynamic update of PDS, the system updates the state variables immediately after each action is executed. r For example, the system calculates the amount of tasks d that the device can complete in the current time step based on the current bandwidth W and the number of discrete levels M:

[0113] d=Δt*2W*log2M

[0114] Δt is the interval time. By calculating the task completion amount d, the system can determine the current remaining data amount d r The updated state is calculated and the updated PDS is calculated to ensure that the latest state is used in the next step.

[0115] (3) Network update. Repeat the following training: first, sample from the buffer, then calculate the target value, then update the Q network, and then update the target network. During the Q network update process, the system randomly extracts small batches of samples from the experience buffer (in order to improve training efficiency and stability, the experience replay mechanism is adopted to store the state, action, reward, and next state quadruple (s, α, r, s′) of each step into the experience replay buffer, and randomly extracts small batches of samples for training. This approach can break the correlation between samples and increase the convergence speed.) for training. Calculate the target Q value y for each sample t :

[0116]

[0117] Where γ is a discount factor that balances the impact of immediate rewards and future rewards. 误差 Defined as the difference between the current Q network output and the target Q value:

[0118] TD 误差 =(yQ(s,α;θ)) 2

[0119] By minimizing TD 误差, update the parameters of the Q network so that the network gradually approaches the optimal Q value. The parameter update formula is:

[0120]

[0121] Where β is the learning rate. After each update, the system synchronizes the Q network parameters to the target network within a fixed number of steps to ensure the stability of the target Q value and prevent fluctuations during network training.

[0122] Throughout the training process, Lagrange multiplier updates ensure dynamic regulation of energy consumption. Specifically, if current energy consumption exceeds a set threshold, λ is increased to increase the penalty for energy consumption; otherwise, λ is decreased, allowing the system to focus more on AoI optimization. This dynamic adjustment process ensures that the system can intelligently select the optimal perception strategy in complex environments while maintaining a balance between energy consumption and information freshness.

[0123] The simulation experiment flow is as follows: First, the edge device is in state 1, then performs action 1, receives reward 1, and enters state 2, and this process repeats. The post-decision state adds "state 1.5" between states 1 and 2. The Lagrange multiplier rule redefines the reward. The DQN algorithm controls the edge device to perform actions and observes the rewards after performing these actions. It selects the action that maximizes the reward as the optimal action. As the experimental loop repeats, the optimal action gradually converges to the optimal action, which is the optimal perception method / strategy.

[0124] An embodiment of the present invention also provides a system including a processor and a memory, wherein the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute an edge device intelligent perception method based on information freshness as described in the above technical solution.

[0125] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Persons skilled in the art may make various modifications, additions, or substitutions to the described specific embodiments without departing from the spirit of the present invention or exceeding the scope of the appended claims.

Claims

1. An edge device intelligent perception method based on information freshness, characterized in that: The steps include: Step 1: Analyze the key factors in the edge device intelligent perception system and then model the Markov decision process based on these factors. Step 2: Optimize and calculate the state in the Markov decision process, update the state space through PDS, use the known information to model and update the perception state a priori, and obtain the post-decision state; Step 3: Energy consumption is factored into the reward in the Markov decision process using Lagrange multipliers. Energy consumption and the Age of Information (AoI) objective are jointly optimized using the Lagrange multiplier method. During the optimization process, the energy constraint and the AoI objective are integrated into a single Lagrange optimization problem. The Lagrange multiplier value is dynamically adjusted to achieve effective control of the AoI. The Lagrange multiplier is updated during each iteration to achieve dynamic control of energy consumption. The specific implementation of step 3 is as follows: The cost function C is defined to quantify the total cost of information age and energy consumption. The cost function is defined as follows: C = ∑(A + λ·(E-Emax)) Where A represents the information age, E represents the current energy consumption, and Emax is the maximum energy threshold; λ is the Lagrange multiplier used to dynamically adjust the penalty intensity for energy overspending; Update the Lagrange multiplier λ. At each time step, adjust the value of λ according to the deviation of energy consumption from the maximum energy threshold. The update formula of the Lagrange multiplier is: λ′=λ+θ(t)·(E-Emax) Where θ(t) is a step size control parameter that is dynamically adjusted with time step t; At each time step, λ′ is constrained to keep it within a certain range, and λ″ is obtained. This process is expressed as: λ″=max(0,min(λ′,λmax)) Where λmax is the upper limit of the Lagrange multiplier; Step 4: Solve the established Markov decision process model and learn the optimal perception strategy through the deep Q learning algorithm. In each time step, the deep Q learning algorithm estimates the action value through the neural network and selects the action to be performed based on these action values, thereby updating the state and generating the optimal strategy to minimize information age and energy consumption.

2. The method for intelligent edge device perception based on information freshness according to claim 1, characterized in that: The key factors in step 1 include the task queue, sampling model, energy consumption model, energy arrival model, and channel transmission model. The task queue includes the remaining data amount of the first task in the queue and the queue length. The sampling model includes the sampling frequency and sample size. The channel transmission model includes the bandwidth, number of discrete levels, and modulation order. The energy consumption model includes perception energy consumption and transmission energy consumption.

3. The method for intelligent edge device perception based on information freshness according to claim 1, characterized in that: The four elements of the Markov decision process include state, action, reward, and state transition. The action space includes sampling frequency, sample size, number of discrete levels, and bandwidth. The state vector includes the amount of remaining data of the first task, the length of the task queue, and the age of information. Let the state vector be s = [d r ,A,q], where d r represents the amount of remaining data for the first task in the queue, A represents the information age, and q represents the length of the task queue. The energy consumption of the device is considered in the reward process, including perception energy consumption and transmission energy consumption. The perceived energy consumption is expressed as: E 采样 =P 采样 ·T 采样 Among them, P 采样 is the power consumption of the sensor, T 采样 is the total sampling time; where f 采样 is the sampling frequency, N 样本 is the total number of samples; The energy consumption of data transmission is expressed as: E 传输 =P 发送 ·T 传输 Among them, P 发送 is the power consumption of the communication module, T 传输 is the data transmission time, D 样本 is the sample size, P 传输速率 is the transmission rate.

4. The method for intelligent edge device perception based on information freshness according to claim 1, characterized in that: The specific implementation process of PDS is divided into the following steps: Extract the remaining data d of the first task from the input state vector and action vector r , task queue length q, information age A and sampling frequency f, sample size N, transmission power P and bandwidth allocation W; Calculate the transmission rate and determine whether the task is completed. Without considering noise, the maximum transmission rate is determined by the number of discrete levels and bandwidth. According to the Nyquist criterion, it can be expressed as: c = 2W*log2M; Where W is the bandwidth allocated to the device and M is the number of discrete levels of the signal; Update the state variables according to whether the task is completed, generate the decision state, and update the task demand: if the task is completed, the remaining data volume d r Return to zero, otherwise update to d r -d, that is, the amount of remaining data of the new first-in-line task, d is the amount of task completion, d = Δt*2W*log2M, Δt is the interval time; update information age: if the task is completed, the information age A is reset to 0, otherwise it is accumulated by 1; update task queue information: if the task is completed, the task queue information q is reduced by 1, indicating that one task is completed; the channel gain remains unchanged; After combination, the decision state is generated: the updated variables pds A , pds q Combined into a new state vector s′, the specific formula is as follows: pds A =A+1 pds q =q-1 5. The method for intelligent edge device perception based on information freshness according to claim 1, characterized in that: In step 4, the deep Q-learning algorithm DQN uses a deep neural network to estimate the Q-value function Q(s, a; θ), where s is the current state, a is the possible action combination, and θ is the network parameter. The specific implementation process is as follows: The deep neural network is defined as a Q-network. Its input layer receives the current state s, including the task demand dr, the information age A, and the task queue information q. The input layer is followed by multiple fully connected layers, each with several neurons, using the ReLU activation function to introduce nonlinearity and prevent gradient vanishing. The fully connected layers continuously abstract and combine input features to enable the network to estimate the Q value of different state-action pairs. The final output layer contains nodes corresponding to the action space, and the value of each node is the Q value of a specific action in the current state. Define the target Q value: After executing the current action, the system will reach the next state s′ and obtain an immediate reward r. Use the Bellman equation to update the Q value and define the target Q value as: Where γ is a discount factor used to weigh future rewards and immediate rewards; Q(s′,a′; θ - ) is the maximum Q value of the next state s′ calculated by the target Q network, θ - represents the parameters of the target Q network, and α′ represents the set of all possible actions in the next state s′; Calculate TD 误差 :Use mean square error MSE to measure the difference between the Q network output and the target Q value, and define time difference TD 误差 for: TD 误差 =(y-Q(s,a;θ)) 2 y is the target Q value, Q(s,a;θ) is the predicted value of the Q network; Back propagation updates the Q network and minimizes TD using gradient descent 误差 , by calculating the gradient of the loss function with respect to the network parameters θ, the network weights are adjusted: Among them, β is the learning rate. Through back propagation, the Q network gradually learns and updates its parameters to improve the accuracy of Q value estimation. is the gradient.

6. The method for intelligent edge device perception based on information freshness according to claim 5, characterized in that: The experience replay mechanism is used to store the state, action, reward, and next state quadruple (s, a, r, s′) of each step into the experience replay buffer, from which small batch samples are randomly extracted for training.

7. The method for intelligent edge device perception based on information freshness according to claim 5, characterized in that: Every fixed number of steps, the parameters of the current Q network are synchronized to the target Q network to ensure the stability of the target Q value.

8. The method for intelligent edge device perception based on information freshness according to claim 5, characterized in that: The ∈-greedy strategy is adopted during the training process. Specifically, an action is randomly selected with probability ∈, and the action with the largest Q value is selected with probability 1-∈.

9. An edge device intelligent perception system based on information freshness, characterized by: It includes a processor and a memory, the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute an edge device intelligent perception method based on information freshness as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • State updating method based on related information age in energy harvesting Internet of Things

    CN116056033A

  • Constraint reinforcement learning-based communication perception joint optimization method and system

    CN116367337A