Data updating optimization method and system based on reinforcement learning and potential diffusion model

By using a data update optimization method based on reinforcement learning and a potential diffusion model, the problems of information content relevance and dynamic environmental changes in IoT systems are solved, thereby improving resource utilization and system stability. In particular, it significantly improves the efficiency and success rate of data updates in strongly correlated environments.

CN120897203APending Publication Date: 2025-11-04NORTHWEST A & F UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511047667.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

In existing IoT control systems, traditional data update methods fail to effectively consider the relevance of information content and dynamic environmental changes, and cannot adapt to hard constraints such as bandwidth and energy consumption. This makes it difficult to achieve synergistic optimization of resource utilization and information value. In particular, in multi-sensor scenarios, scheduling strategies often fall into the contradiction of update redundancy or insufficient timeliness.

Method used

We adopt a data update optimization method based on reinforcement learning and latent diffusion model. By constructing a resource-constrained Internet of Things system model, we define the context-aware information timeliness index VoCAI and combine Lagrange relaxation and Lyapunov stability theory to model the optimization problem as CMDP. We use the SR-PPO algorithm for policy optimization and introduce the latent diffusion model LDM to jointly optimize the dynamic information correlation matrix and scheduling strategy.

Benefits of technology

In medium-to-large scale scenarios, the SR-PPO algorithm improves the average reward by 1.2-2 times, while the LDM-PPO algorithm improves it by 4.87 times in strongly correlated environments. The information equivalent scheduling mechanism reduces redundant transmission, improves resource utilization by 30%, and maintains system stability and efficiency within the transmission success probability range of 0.85-0.99.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120897203A_ABST
    Figure CN120897203A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of Internet of Things data processing, discloses a data updating optimization method and system based on reinforcement learning and a potential diffusion model, and provides a VoCAI-based SR-PPO algorithm and LDM-PPO combined learning framework for the problems of information timeliness optimization and dynamic correlation modeling in the resource-constrained Internet of Things. Bandwidth and power constraints are processed through Lagrange relaxation and a Lyapunov stability theory, time-varying correlation between sensors is dynamically modeled by using LDM, an equivalent scheduling mechanism is designed to improve the resource utilization efficiency, and experiments show that in a strong correlation environment, the average reward of the LDM-PPO algorithm is improved by 4.87 times compared with that of a traditional algorithm, and the reward of the LDM-PPO algorithm is greatly improved. The problem of efficient scheduling of the multi-source sensor under resource limitation is effectively solved, and the method is suitable for industrial monitoring, intelligent transportation and other scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of Internet of Things (IoT) data processing technology, specifically to a data update optimization method and system based on reinforcement learning and a potential diffusion model. Background Technology

[0002] In Internet of Things (IoT) control systems, wireless sensor networks, as a core component of the perception layer, rely heavily on the timeliness of their dynamic data updates to ensure efficient system decision-making. However, existing technologies suffer from the following key shortcomings:

[0003] 1. Traditional metrics (such as AoI) generally do not take into account the relevance of information content and dynamic environmental changes, and rely on ideal resource conditions, which cannot adapt to the hard constraints such as bandwidth and energy consumption in actual IoT systems.

[0004] 2. In multi-sensor related source scenarios, existing methods use static correlation matrices to characterize sensor dependencies, which are difficult to adapt to dynamic environmental evolution (such as state dependency changes in scenarios like target movement and fire spread).

[0005] 3. Faced with practical constraints such as unreliable channels and limited power, traditional scheduling strategies often fall into the contradiction of "update redundancy" or "insufficient timeliness", and cannot achieve the coordinated optimization of resource utilization and information value. Summary of the Invention

[0006] The purpose of this invention is to provide a data update optimization method and system based on reinforcement learning and potential diffusion models to solve the problems mentioned in the background art.

[0007] To achieve the above objectives, the present invention provides the following technical solution: a data update optimization method and system based on reinforcement learning and a latent diffusion model, comprising: Construct a resource-constrained IoT system model, including N sensors and a single receiving node, and set unreliable channels and bandwidth constraints. and power constraints .

[0008] A context-aware information timeliness index, VoCAI, is defined. The optimization problem is modeled as CMDP through Lagrange relaxation and Lyapunov stability theory, and solved using the SR-PPO algorithm, which introduces a power Lyapunov drift penalty term.

[0009] Introducing the Latent Diffusion Model (LDM), a dynamic information correlation matrix is ​​constructed based on the system context state. Design an information equivalent scheduling mechanism The correlation source scheduling problem is modeled as CMDP, and the LDM-PPO algorithm is used to achieve joint optimization of correlation modeling and scheduling strategy.

[0010] Preferably, the VoCAI is defined as wherein is the context weight, is the content utility function, .

[0011] Preferably, the state space of the SR-PPO algorithm is wherein is the power difference sequence, updated as The reward function is wherein is the power Lyapunov drift.

[0012] Preferably, the input of the LDM is the system context state , and the output is the dynamic information correlation matrix , and the generation process includes encoder mapping, forward diffusion, conditional denoising and decoder reconstruction.

[0013] Preferably, the joint optimization objective of the LDM-PPO algorithm is wherein is the LDM denoising loss, is the long / short-term feature decoupling loss.

[0014] The system model construction module is configured to configure a multi-sensor network topology, set unreliable channel parameters and resource constraint conditions; the VoCAI optimization module is configured to construct a VoCAI index based on a context state and an information age, and perform CMDP modeling and policy optimization through an SR-PPO algorithm; the dynamic correlation modeling module is configured to receive a system context state t, and generate a dynamic information correlation matrix t through a latent diffusion model LDM; and the joint optimization execution module is configured to jointly optimize a scheduling policy based on LDM modeling using an LDM-PPO algorithm, and output a scheduling action sequence.

[0015] The context-aware optimization process adopts an SR-PPO algorithm, which includes: initializing policy network and value network parameters; constructing state input based on sensor state and power drift, and sampling scheduling action; obtaining state transition and reward data through environment interaction, and constructing an experience pool; calculating advantage value based on a generalized advantage estimation (GAE) method; and optimizing the policy network using a policy gradient method of clipping probability ratio.

[0016] Compared with the prior art, the present application has the following beneficial effects: 0、The SR-PPO algorithm of the present application can improve the average reward by 1.2-2 times compared with traditional DRL algorithms in medium and large-scale scenarios, verifying the mining ability of the VoCAI index for context correlation.​

[0017] 1、The LDM-PPO in the strong correlation environment is 4.87 times higher than the average reward of the SR-PPO, which proves the capture ability of the time-varying dependence between sensors.

[0018] 2、The information equivalent scheduling mechanism makes the sensors not directly scheduled to realize data update through structural compensation, reduces redundant transmission, and improves resource utilization by more than 30%.

[0019] 3、Through the dynamic modeling of LDM and the strategy optimization of SR-PPO, the system maintains stable and efficient performance in the transmission success probability range of 0.85-0.99. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor. Figure 1 The figure is a target-oriented data dynamic updating method in the Internet of Things control system of the present application; Figure 2 The figure is a LDM-PPO algorithm structure diagram of the present application; Figure 3 The figure is a SR-PPO algorithm structure diagram of the present application; Figure 4 The figure is a multi-source wireless transmission Internet of Things system schematic diagram of the present application; Figure 5 The figure is a DQN basic idea schematic diagram of the present application. DETAILED DESCRIPTION

[0021] The technical solutions in the embodiments of the present application will be described clearly and completely in the following with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0022] Please refer to Figures 1-5 The present application provides a technical solution: a data update optimization method and system based on reinforcement learning and latent diffusion model.

[0023] Embodiment 1: Resource-constrained data dynamic update optimization method based on VoCAI 1. System model construction and VoCAI index definition In this embodiment, a wireless IoT system containing 6 sensor nodes (N=6) is constructed, with a maximum of 3 sensors scheduled per time slot (M=3), satisfying the bandwidth constraint , the channel transmission success probability between the sensor and the receiving end obeys a uniform distribution (0.85, 0.99), the power consumed when each sensor is scheduled , the long-term average power constraint is .

[0024] The context awareness information timeliness index VoCAI is defined as follows: wherein the context state , represents a high-value state (such as temperature anomaly) 1 represents a normal state. The context awareness weight is set as: when , 0; when , 5.

[0025] The content utility function adopts a segmented form: The information age is defined as the difference between the current time slot and the generation time of the latest successfully received data.

[0026] 2. CDP modeling and Lagrangian relaxation The data update scheduling problem is modeled as a constrained Markov decision process (CMDP), with the state space defined as wherein is the power drift sequence, and is constructed through Lyapunov stability theory: wherein is the power threshold, .

[0027] The bandwidth constraint is converted into a penalty term through Lagrangian relaxation method, and the Lagrangian function is: wherein is the Lagrangian multiplier, =3 is the upper limit of the bandwidth.

[0028] 3. Specific implementation of SR-PPO algorithm (1) State and action space State input: information age of each sensor (Value range 1-10), Context state (1 or 2) Power drift (Normalized to [0,1]), forming an 18-dimensional state vector.

[0029] Action space: Scheduling decisions for each sensor , A combination of actionable actions.

[0030] (2) Reward function design in, This is the bandwidth constraint penalty coefficient. The power drift penalty coefficient is the power Lyapunov drift. The calculation is as follows: (3) Strategy optimization process Initialization: Both the policy network (Actor) and the value network (Critic) use a 2-layer fully connected network (64 neurons per layer, ReLU activation), parameters... Random initialization, old strategy parameters Experience pool Clear.

[0031] Data acquisition: 200 time slots are executed per training round, based on the old strategy. Sampling action Receive a reward after execution and new status Stored in the experience pool.

[0032] Advantage estimation: Calculating time difference error discount factor The generalized advantage estimate is 0.995. in 0.95 is the GAE coefficient.

[0033] Strategy update: Replace the objective function with a truncation function. in clip parameters 0.2, gradient descent is performed 4 times per update round, learning rate .

[0034] (4) Experimental verification After training for 1000 rounds on the PyTorch platform, the SR-PPO algorithm converges to an average reward of 480±20, which is 1.1 times higher than the traditional DQN algorithm (average reward 230±15), and the VoCAI index is 35% higher than the AoI index while meeting the power constraint.

[0035] Embodiment 2: LDM-PPO data update optimization method for correlation-awareness 1. Dynamic correlation modeling based on LDM (1) System parameter setting 5 sensor nodes (N=5) are set, the bandwidth constraint M=2, a strong correlation environment is constructed, and the real correlation matrix The off-diagonal elements of the matrix are subject to .

[0036] The history observation window The state contains: The current VoCAI state of each sensor ; The scheduling trajectory of the past 10 time slots ; The reception record of the past 10 time slots .

[0037] (2) Pseudo-label matrix generation Fusion of three types of features to construct the pseudo-label matrix : VoCAI proximity: , where ; Scheduling synergy: ; Reception similarity: , where is an indicator function and is an indicator function.

[0038] The weight vector is generated by the subnetwork : Where the state embedding is extracted by a 2-layer CNN, outputting a 3-dimensional weight vector that satisfies .

[0039] (3) LDM model architecture Encoder: map to a 64-dimensional latent space: ; Forward diffusion: add noise according to , where , ; Conditional denoising predictor: with Given the condition, noise is predicted using the UNet network. ; Decoder: Reconstructing the correlation matrix , where D is a 2-layer fully connected network.

[0040] (4) Training loss function in, , , This is a projection of short-term features into a long-term feature subspace, achieved by decoupling long-term features. Compared with short-term characteristics Achieve dynamic adaptability.

[0041] 2. Information Equivalent Scheduling and LDM-PPO Joint Optimization (1) Calculation of equivalent scheduling state when When adjacent sensors are scheduled, the scheduling of adjacent sensors can trigger an equivalent update of the current sensor, reducing redundant scheduling.

[0042] (2) Joint optimization objective in, As a long-term cumulative reward for the scheduling strategy, For LDM denoising loss, For feature decoupling loss, weights .

[0043] (3) LDM-PPO algorithm flow Initialization: LDM parameters PPO strategy network parameters ψ, value network parameters Random initialization, experience pool Clear; Collaborative training: The LDM model is updated every 50 time slots, reconstructed based on the latest 200 state-action trajectories. ; PPO policy network input includes Output scheduling action ; Reward function: in Based on equivalent scheduling state calculate; Parameter update: Jointly optimize LDM and PPO networks using the Adam optimizer, learning rate... Each training round performs 10 joint gradient descent iterations.

[0044] (4) Experimental verification In the strong correlation environment, the average reward of LDM-PPO algorithm reaches 475±15 after convergence, which is 3.8 times higher than that of SR-PPO algorithm (average reward 98±8), and in the N=5, M=2 scenario, the number of sensor updates is reduced by 28%, which proves the effectiveness of dynamic correlation modeling and equivalent scheduling mechanism.

[0045] Embodiment 3: Specific implementation of system modules 1. Model construction module VoCAI calculation unit: real-time calculation of each sensor , support custom context weight and utility function; CMDP modeling unit: generate state transition equation and constraint condition through Lagrange relaxation and Lyapunov operator.

[0046] 2. SR-PPO algorithm module Actor network: input 18-dimensional state, output 20-dimensional action probability distribution, use PPO-Clip policy update; Critic network: estimate state value , support GAE advantage calculation.

[0047] 3. LDM modeling module Feature fusion unit: combine VoCAI proximity, scheduling synergy, and reception similarity features according to weights; Diffusion-de-noising unit: perform 100-step diffusion process to generate dynamic correlation matrix .

[0048] 4. Joint optimization module Equivalent scheduling calculator: based on generate , update VoCAI and reward; Joint trainer: synchronize LDM parameter and PPO policy optimization, support decoupling regularization.

[0049] The application has been implemented in a discrete-time simulation environment constructed on a Python platform, a multi-source perception system with different information dependence structures is generated through ER random graph and Watts-Strogatz small-world network, and various sensor scales (N=3, 5, 7) and bandwidth limits (M=1, 2, 3) are set. Experimental results show that in the environment of a medium correlation matrix, LDM-PPO achieves an average reward of 475±15 under the configuration of N=5, M=2, which is significantly better than SR-PPO (about 98±8) and traditional algorithms such as DQN, reflecting the synergistic advantages of the method in information structure modeling and scheduling strategy optimization. Especially in the strong correlation environment, the average reward of the method is nearly 4 times that of SR-PPO, further verifying the effectiveness of the structure perception mechanism in the resource-constrained multi-source Internet of Things system.

[0050] Although embodiments of the application have been shown and described, it is to be understood that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the application, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A data update optimization method based on reinforcement learning and a latent diffusion model, characterized in that, Includes the following steps: a) Construct a resource-constrained IoT system model, including N sensors and a single receiving node, setting an unreliable channel and bandwidth constraints. and power constraints ; b. Define the context-aware information timeliness index VoCAI, model the optimization problem as CMDP through Lagrange relaxation and Lyapunov stability theory, and solve it using the SR-PPO algorithm, which introduces a power Lyapunov drift penalty term. c. Introduce the Potential Diffusion Model (LDM) and construct a dynamic information correlation matrix based on the system context state. Design an information equivalent scheduling mechanism The correlation source scheduling problem is modeled as CMDP, and the LDM-PPO algorithm is used to achieve joint optimization of correlation modeling and scheduling strategy.

2. The data update optimization method based on reinforcement learning and latent diffusion model according to claim 1, characterized in that, The VoCAI is defined as in For context weights, For content utility function, .

3. The data update optimization method based on reinforcement learning and latent diffusion model according to claim 1, characterized in that, The state space of the SR-PPO algorithm is: ,in For power difference sequences, according to renew; The reward function is ,in For power Lyapunov drift.

4. The data update optimization method based on reinforcement learning and latent diffusion model according to claim 1, characterized in that, The input to the LDM is the system context state. Output dynamic information correlation matrix Its generation process includes encoder mapping, forward diffusion, conditional denoising, and decoder reconstruction.

5. The data update optimization method based on reinforcement learning and latent diffusion model according to claim 1, characterized in that, The joint optimization objective of the LDM-PPO algorithm is: ,in For LDM denoising loss, The loss is used to decouple long-term / short-term features.

6. The data update optimization method based on reinforcement learning and latent diffusion model according to claim 1, characterized in that, This method includes the following modular execution process: a. System model building module, used to configure multi-sensor network topology, set unreliable channel parameters and resource constraints; b. VoCAI optimization module, used to construct VoCAI index based on context state and information age, and to perform CMDP modeling and strategy optimization through SR-PPO algorithm; c. Dynamic correlation modeling module, used to receive the system context state t and generate a dynamic information correlation matrix Ê through the Latent Diffusion Model (LDM). t ; d. Joint optimization execution module, used to jointly optimize the scheduling strategy based on LDM modeling using the LDM-PPO algorithm, and output the scheduling action sequence.

7. The data update optimization method based on reinforcement learning and latent diffusion model according to claim 1, wherein the context-aware optimization process adopts the SR-PPO algorithm, which includes: a. Initialize the parameters of the policy network and value network; b. Construct state inputs based on sensor status and power drift, and sample and schedule actions; c. Obtain state transition and reward data through environmental interaction to build an experience pool; d. Calculate the dominance value based on the generalized dominance estimation (GAE) method; e. Optimize the policy network using the policy gradient method with shearing probability ratios.