A multi-access point edge computing service state synchronization method and system based on a deep Q network

CN121462596BActive Publication Date: 2026-09-25OCEAN UNIV OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202512049492.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-31
Publication Date
2026-09-25
Estimated Expiration
2045-12-31

AI Technical Summary

Technical Problem

[0009]为了解决现有状态度量方法未能充分体现接入点的实际决策逻辑与用户行为模式、从而导致信息新鲜度刻画不足的问题,本发明提供一种基于深度Q网络的多接入点边缘计算服务状态同步方法及系统

Benefits of technology

为在边缘节点更新能力受限的条件下提高多接入点卸载决策的准确性与系统处理效率,本发明设计了MH-D3QN架构,并在此基础上提出了一种接入点状态更新优化方法。现有技术中存在:仅依赖AoI无法刻画状态错误对卸载决策的实际影响、组合动作空间在多接入点场景下呈指数级增长导致策略难以求解、以及更新资源受限情况下难以高效分配的问题。本发明围绕上述问题进行了系统性的设计优化,主要有益效果如下:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121462596B_ABST
    Figure CN121462596B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of edge state synchronization, and especially relates to a multi-access point edge computing service state synchronization method and system based on a deep Q network. The method comprises the following steps: acquiring access point data; constructing an EDPT index based on the acquired access point data; constructing an MH-D3QN network model, wherein MDP modeling is performed based on the constructed EDPT index; calculating the Q value of the access point based on the MH-D3QN algorithm; performing income sorting according to the Q value of the access point; and performing model training based on the constructed MH-D3QN network model. The method effectively avoids the difficulty in solving caused by the combination explosion, and makes the update strategy still computable and scalable in the multi-access point scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of edge state synchronization technology, and in particular to a method and system for synchronizing the state of multi-access point edge computing services based on deep Q networks. Background Technology

[0002] Edge computing alleviates the high latency and centralized processing bottlenecks faced by traditional networks in low-latency applications such as drone control and autonomous driving by bringing computing power closer to terminal devices or users. However, in traditional edge computing architectures, network routing and computational decisions remain decoupled, making global optimization difficult. To address this limitation, Computing First Networking (CFN) integrates computation and network control at the network layer. By enabling compute-aware routers to consider network conditions and computing resource status during packet forwarding, CFN supports more flexible task offloading, more efficient resource scheduling, and improved service reliability. However, CFN still faces a key challenge: how to maintain the service status information of compute-aware routers or access points (APs) to ensure the correctness of decisions. Inaccurate status information can lead to incorrect offloading decisions, misallocation of resources, or even service failures, thus affecting the overall performance and reliability of the system. For example, in a smart factory, there is one edge node (or central service node) and multiple access points. Edge nodes update their status information to the access point. Users (or smart devices) connect to the access point and send computing requests. The access point decides whether to have the user process the task locally or offload it to an edge node based on the recorded status information. However, when the status information recorded by the access point is inaccurate, it may offload the task to a service node that is already heavily loaded, leading to resource overload, increased processing latency, and potentially even system failures or service interruptions, thus affecting overall productivity and service reliability.

[0003] Existing research typically uses Age of Information (AoI) to measure the freshness of service state information recorded by access points and optimizes AoI to ensure decision accuracy. However, these methods do not fully consider the actual decision-making logic of access points and user access behavior patterns. When a user visits again, even if the access point has a high AoI, it can still make the correct offloading decision, so updating its state at this time is actually unnecessary. Thus, limited update resources can be prioritized for access points that truly need updates. Given that the number of state updates that edge nodes can perform simultaneously under resource constraints is limited, how to rationally allocate update opportunities is particularly crucial for improving overall system performance and resource utilization efficiency.

[0004] Traditional optimization methods, such as queuing models, have limited ability to model system behavior and struggle to cope with random changes; mathematical programming methods rely on numerous parameters and incur high computational costs, making them unsuitable for real-time decision-making; while heuristic algorithms are efficient, they cannot guarantee global optimality. Overall, traditional methods in edge computing networks struggle to simultaneously ensure both service state freshness and scheduling efficiency.

[0005] In contrast, Deep Reinforcement Learning (DRL) methods, by learning decision-making policies, can automatically cope with complex dynamic environments and random changes, overcoming the limitations of traditional methods. However, in this scenario, DRL faces the problem of combinatorial actions, requiring the simultaneous selection of the optimal combination of multiple decision variables, which leads to an expansion of the decision space and increases computational complexity. Although some research has attempted to combine Deep Q-Networks (DQNs) with Mixed-Integer Linear Programming (MILP) to address this problem, MILP's high computational cost makes it unsuitable for real-time decision-making, thus limiting its application in this scenario.

[0006] In summary, existing optimization methods primarily focus on achieving task scheduling and resource allocation in edge computing networks through traditional queuing models, mathematical programming methods, heuristic algorithms, and deep reinforcement learning. However, the following prominent issues remain unresolved: 1) Insufficient measurement of information freshness: Existing methods typically measure the freshness of state information through AoI (Aspect-Oriented Intelligence), but do not fully consider the actual decision-making logic of the access point and user behavior patterns. Even with a large AoI, the access point can still make correct decisions, and excessive updates may waste resources.

[0007] 2) Unreasonable allocation of update resources: Due to the limited resources of edge nodes, the existing method has failed to effectively prioritize the allocation of limited update resources, resulting in some important access points not being updated in a timely manner, which affects system efficiency.

[0008] 3) Shortcomings of existing deep reinforcement learning methods: The application of existing deep reinforcement learning methods in edge computing faces significant challenges. Although techniques such as deep Q-networks have been introduced, they still suffer from problems such as excessively large state spaces and complex action combinations when dealing with complex decision spaces, making it difficult to achieve real-time, efficient decision-making and resource allocation. Summary of the Invention

[0009] To address the problem that existing state measurement methods fail to fully reflect the actual decision-making logic and user behavior patterns of access points, resulting in insufficient characterization of information freshness, this invention provides a method and system for synchronizing the state of multi-access point edge computing services based on deep Q networks.

[0010] Firstly, the present invention provides a method for synchronizing the state of multi-access point edge computing services based on deep Q networks, which adopts the following technical solution: A method for synchronizing the state of multi-access point edge computing services based on deep Q networks includes: Obtain access point data; The EDPT metric is constructed based on the acquired access point data; Construct an MH-D3QN network model, in which MDP modeling is performed based on the constructed EDPT index; the Q-value of the access point is calculated based on the MH-D3QN algorithm; and the revenue is ranked according to the Q-value of the access point. Model training was performed based on the constructed MH-D3QN network model; The trained MH-D3QN network model is used to synchronize the status of edge computing services with multiple access points.

[0011] Secondly, a multi-access point edge computing service state synchronization system based on a deep Q-network includes: The data acquisition module is configured to acquire access point data; The metric construction module is configured to construct EDPT metrics based on the acquired access point data; The model building module is configured to build an MH-D3QN network model, which includes MDP modeling based on the constructed EDPT metric; calculating the Q-value of the access point based on the MH-D3QN algorithm; and ranking the access points by their Q-values. The model training module is configured to train the model based on the constructed MH-D3QN network model; The optimization module is configured to use the trained MH-D3QN network model to synchronize the state of multi-access point edge computing services.

[0012] Thirdly, the present invention provides a computer-readable storage medium storing a plurality of instructions adapted for loading and execution by a processor of a terminal device of the aforementioned method for synchronizing the state of a multi-access point edge computing service based on a deep Q network.

[0013] Fourthly, the present invention provides a terminal device, including a processor and a computer-readable storage medium, wherein the processor is used to implement various instructions; the computer-readable storage medium is used to store multiple instructions, the instructions being adapted to be loaded and executed by the processor to provide a method for synchronizing the state of a multi-access point edge computing service based on a deep Q network.

[0014] In summary, the present invention has the following beneficial technical effects: To improve the accuracy of multi-access-point offloading decisions and system processing efficiency under conditions of limited edge node update capabilities, this invention designs the MH-D3QN architecture and proposes an access point state update optimization method based on it. Existing technologies suffer from several problems: relying solely on AoI cannot characterize the actual impact of state errors on offloading decisions; the combinatorial action space grows exponentially in multi-access-point scenarios, making strategy solving difficult; and efficient allocation is challenging under limited update resources. This invention systematically optimizes the design to address these issues, with the following main benefits: (1) It can accurately depict the real cost of decision-making errors and improve the effectiveness of state updates: This invention introduces EDPT as a new cost index, which quantifies the actual impact of state errors on unloading decisions into a directly optimizable processing time cost. Compared with traditional methods that rely solely on AoI, EDPT can explicitly reflect the impact of state expiration on task latency, enabling the system to identify the state deviations that have the most significant impact on decisions, avoid invalid updates, and thus improve the efficiency of update resource utilization.

[0015] (2) Achieve accurate allocation of update resources under resource-constrained conditions and improve decision-making efficiency: Each time slot of the edge node can only be updated To address the limitation of individual access points, this invention, based on MH-D3QN, accurately captures the dynamic environment of each access point and employs an action value evaluation and benefit ranking mechanism to prioritize updates for the access points with the greatest impact on system performance, while satisfying update constraints. This approach ensures the most efficient use of limited update resources, thereby improving overall task processing efficiency.

[0016] (3) Solving the problem of exponential expansion of the combined action space and significantly reducing the complexity of strategy solving: This invention reduces the complexity of strategy solving that was originally required to be solved by the exponential expansion of the combined action space. The update decision-making process for searching within a combinatorial action is decomposed into three computable steps: local optimum judgment, quantity constraint screening, and profit ranking. This reduces the action decision complexity from exponential to linear and ranking levels. This mechanism effectively avoids the solution difficulties caused by combinatorial explosion, ensuring that the update strategy remains computable and scalable in multi-access-point scenarios. Attached Figure Description

[0017] Figure 1 This is an overall flowchart of Embodiment 1 of the present invention.

[0018] Figure 2 This is a flowchart of step 3 of embodiment 1 of the present invention, which uses the MH-D3QN algorithm to calculate the Q value.

[0019] Figure 3 This is a flowchart of step 4 of embodiment 1 of the present invention, which is based on the Q-value selection behavior.

[0020] Figure 4This is a flowchart of step 5, model training, in Embodiment 1 of the present invention.

[0021] Figure 5 This is a comparative experimental diagram of Embodiment 1 of the present invention. Detailed Implementation

[0022] The present invention will be further described in detail below with reference to the accompanying drawings.

[0023] Example 1 Reference Figure 1 This embodiment presents a multi-access point edge computing service state synchronization method based on a deep Q-network, comprising: proposing an Error Decision Processing Time (EDPT) metric to comprehensively quantify the impact of state deviation from two dimensions: decision accuracy and request execution time, and avoiding unnecessary update operations triggered while the state is still valid. Addressing the problems of unreasonable update resource allocation and the difficulty of solving problems in the combined action space using deep reinforcement learning, this invention further proposes a Multi-Head Dueling Double DeepQ Network (MH-D3QN) algorithm. Through global feature extraction, local feature extraction, and feature fusion, it achieves accurate estimation of the potential update value of each access point. Simultaneously, to enable the above value estimation results to be transformed into effective update decisions under resource-constrained conditions, this invention designs a state update strategy dominated by reward ranking. The update reward of the access point is obtained based on the Q-value calculated by MH-D3QN, and candidate update objects are ranked and selected accordingly. This achieves computable update decisions and reasonable resource allocation in the combined action space, improving overall system performance.

[0024] Figure 1 The flowchart of this invention is illustrated. The service state update process shown in the diagram is the core processing content of this invention, while the time required for the user request process serves as a key indicator for evaluating system performance. The service state update process can be specifically divided into the following parts: 1) Constructing the EDPT metric; 2) Modeling the problem based on a Markov Decision Process (MDP); 3) Calculating the Q-value using the MH-D3QN algorithm; 4) Selecting the optimal update strategy based on the calculated Q-value using a state update strategy dominated by payoff ranking; 5) Model training module. These five parts will be described in detail below.

[0025] S1. Construct the EDPT metric This step addresses the problem that existing state measurement methods fail to adequately reflect the actual decision-making logic and user behavior patterns at access points, leading to irrationalities in state update triggering and resource allocation. To address this, an EDPT comprehensive index is proposed. This index quantifies the runtime corresponding to decision-making errors caused by state deviations and serves as the evaluation basis for subsequent optimization modeling.

[0026] This invention models the system as a discrete time slot, where the time slot is composed of... Index. The system consists of... It consists of an access point and an edge node. This invention uses... Indicates the edge node in time To the One access point sends an update. Due to resource limitations, a maximum of [number] updates can be sent within each time slot. This update satisfies the following constraints: This invention assumes that request transmission and response reception each consume one time slot, and that the user has a certain local computing capability, denoted as . The computing power of edge nodes is In time The state information of an edge node is its load, denoted as . And the first The state of the edge node sensed by an access point is denoted as... User visits the The data size carried by one access point is .

[0027] This invention designs an offloading strategy for access points, which selects the optimal offloading decision based on the principle of minimizing the estimated processing time. Specifically, it uses... Indicates the first Each access point in time Decision: in, Indicates the index of the access point. Indicates time, This indicates the time required for the user to process the task locally. This indicates the estimated time for edge nodes to process tasks.

[0028] Similarly, in time The load of edge nodes can be utilized. To calculate the actual processing time and the right decision Based on this, the present invention proposes the EDPT index: in, Indicates the index of the access point. Indicates time.

[0029] In this invention, indicator variables are used. To indicate the first Each access point in time Whether it has been accessed, among which This indicates that the site has been visited. Ultimately, the optimization objective proposed in this invention is: Optimizing the objective function described above allows for a balance between the error rate and average processing time of unloading decisions at the system level. Since the EDPT (Execution Time Per Minute) for a correct unloading decision is 0, optimizing the EDPT will naturally tend to improve the accuracy of decisions. Furthermore, for unavoidable incorrect decisions, reducing their corresponding processing time can further optimize the overall average processing time.

[0030] S2. Modeling the problem based on MDP This step constructs the corresponding MDP model based on the aforementioned EDPT index. By formally defining the reward function related to the system state, executable actions, and the runtime of erroneous decisions, a structured modeling foundation is provided for subsequent policy solving and reinforcement learning algorithm design.

[0031] state It is composed of the states of edge nodes and the states of each access point. The state of an edge node contains a single variable, namely its current load. The load on an edge node depends on the remaining load from the previous time step, the amount of tasks arriving, and the system's service capacity. Its variation can be represented as: Access point status consists of three parts: the load sensed by the access point. Freshness and user access interval Composition, in which It is time. This is the index of the access point. The access point is at time... Receive update actions from edge nodes When the load is updated, the perceived load will be updated to the latest load; otherwise, it will remain unchanged. The evolution formula is: Freshness It increments over time, and resets to 1 upon receiving an update. The formula for its change is: At that moment A user visited the When there are multiple access points (i.e.) The access interval is reset to 1; otherwise, it is incremented. By comparing the actual load of edge nodes Load recorded locally at the access point This can determine whether there are inconsistencies in the perception status of access points, thereby identifying potential decision-making errors. Freshness index The timeliness of load information, while the user access interval. This is used to characterize user access characteristics. System state space The state of an edge node and The state composition of each access point: Therefore, the state space yes .

[0032] action The dimension is Each component represents whether to send a status update to the corresponding access point. Within each time slot, the system determines whether the edge node should send a status update to the corresponding access point by selecting a binary decision from the action vector. The latest load information is sent to each access point. Specifically, the action vector... Each component Used to indicate time Update options: When At that time, the system sent to the first Each access point sends an update to obtain the latest load status of the edge nodes; when When the system does not send updates to that access point, its locally stored sensing state remains unchanged. Therefore, at time [time value missing], the system [action missing]. The complete action can be represented as a vector: The above action structure enables the system to selectively update based on the state freshness of each access point, user access frequency, and load consistency, thereby achieving a more efficient and accurate state maintenance strategy under the condition of limited update resources.

[0033] Reward (or cost) Defined as a dimension The vector, where the first... Each component corresponds to the user at time [time]. Visit the The reward generated when a user accesses the first access point. Specifically, the reward generated when a user accesses the first access point. Access points (i.e.) When this occurs, the reward for this dimension is determined by the EDPT generated by the access, which quantifies the runtime corresponding to decision errors caused by state bias. Furthermore, to suppress network resource consumption caused by frequent updates, the system introduces a penalty term each time a state update is sent to the access point. This encourages reducing unnecessary update costs while ensuring the correctness of decisions. In summary, the first... Each access point at time The reward can be represented as: This reward structure can strike a balance between processing latency and update costs, enabling the system to maintain a reasonable state update frequency while optimizing unloading decisions.

[0034] After completing the MDP modeling, the overall optimization objective of this invention is ( This can be expressed as: finding the optimal strategy. This minimizes the expected state value of the system under the constraints. in, It is the initial state, and the state value function. Defined in strategy Below, from the state The expected cumulative discount reward for departure is expressed as follows: Among the discount factors This is used to ensure the convergence of the value function while controlling the effective time range of the reward.

[0035] S3. Calculate the Q value using the MH-D3QN algorithm. This step addresses the problems of unreasonable resource allocation and the difficulty of solving problems in the combinatorial action space using deep reinforcement learning by proposing the MH-D3QN method. This method achieves accurate estimation of the Q-value of each access point through global feature extraction, local feature extraction, and feature fusion mechanisms, combined with an independent Dueling network structure, providing an efficient and reliable basis for subsequent state update decisions.

[0036] In the time slot The MH-D3QN algorithm proposed in this invention is used to analyze the current system state. Binary actions with each access point Cost estimation is performed. To do this, edge nodes are defined in time slots. For the first Each access point performs an action. The action value function at time is: The Q-value matrix output by the network has a dimension of , of which The rows correspond to and .

[0037] In the time slot State vector After inputting into the Q network, the following calculation process is performed sequentially to obtain the action cost of each access point.

[0038] (1) Global feature extraction First, let's define the global state. The input is processed by a global feature extraction network and mapped to a global feature vector through a two-layer feedforward fully connected structure. It can be formally represented as: in It is a global feature extraction network, specifically a multilayer perceptron with ReLU activation. Global features Then, replication and expansion are performed at the access point level, enabling the system to operate within time slots. All access points share the same set of global context information to reflect system-level factors such as overall load, resource constraints, and network status. The purpose of this step is to ensure that subsequent Q-value estimations for each access point can fully utilize global system information, thereby improving the overall consistency and stability of decision-making.

[0039] (2) Local feature extraction For each access point In the time slot Take its local state triples and denote them as: The local vector is input into a local feature encoding network, and after being encoded by a two-layer feedforward fully connected structure, the local feature vector is obtained. , represented as: in It is a local feature extraction network, specifically a multilayer perceptron with ReLU activation. By analyzing all... The above process is executed in parallel by multiple access points, and the system operates within a time slot. This step involves obtaining local feature representations that characterize the differences in load perception, freshness, and user access behavior among various access points. The purpose of this step is to capture the heterogeneity between different access points, enabling the Q network to finely differentiate the update actions of each access point.

[0040] (3) Feature fusion For each access point its local feature vector With global feature vectors The features are concatenated along the feature dimension to obtain a fused input, which is then fed into a feature fusion network to obtain a fused feature vector. : in It is a feature fusion network, a linear module with a ReLU activation function. It fuses features. It simultaneously incorporates system-level global information and access point-level local information, serving as a unified input for subsequent Q-value calculations in the Dueling structure. The purpose of this step is to jointly model both global and local information within the same feature space, thereby improving the ability to express action costs.

[0041] (4) Calculate the Q value of the Dueling network In the time slot fusion of feature vectors Input to the and the Each access point corresponds to a Dueling structure network. For each access point, the system sets up independent value branches and advantage branches, and their network parameters are independent of each other. Different access points do not share any parameters, thereby ensuring that each access point can achieve specialized action value estimation based on its own state characteristics.

[0042] Specifically, the The Dueling structure of each access point consists of two parts: the Value Branch and the Advantage Branch.

[0043] The state value branch is used to calculate the access point in the time slot. The state value is defined as: in Indicates the first Each access point corresponds to an independent value network.

[0044] Action advantage branch is used to calculate the access point in the action. and time slots The advantage value is defined as: in Indicates the first Each access point corresponds to an independent advantage network.

[0045] Based on the combination form of the Dueling structure, the edge nodes in the time slot are obtained. For the first Each access point performs an action. Action value function: After performing the above operations in parallel on all access points, in the time slot A complete action value matrix can be obtained: By configuring a Dueling structure network with independent parameters for each access point, this invention can estimate the action value of multiple access points in parallel within a unified network framework. This ensures that the model can fully express the differences in characteristics of different access points, and significantly improves the efficiency and accuracy of action value calculation.

[0046] S4. State update strategy dominated by profit ranking This step addresses the problems of unreasonable resource allocation and the difficulty of solving problems in the combinatorial action space using deep reinforcement learning by proposing a state update strategy dominated by reward ranking. Based on the Q-values ​​of each access point calculated by MH-D3QN, this strategy further obtains their corresponding update rewards and ranks candidate access points according to these rewards. Under update quota constraints, access points with higher rewards are prioritized for update, thereby achieving efficient decision-making and reasonable allocation of update resources in the combinatorial action space.

[0047] In both the training and operation phases of this invention, edge nodes are based on time slots. The generated action value matrix This invention uses Indicates the first... Line 1 The elements of a column, their meanings are: Based on the above description, the action selection process of the present invention consists of the following four steps: (1) Random exploration To enhance the strategy's ability to explore the action space, this invention employs the following during the training phase: - Greedy strategy, i.e., using probability Perform exploration, with probability Perform Q-value-based exploitation. In exploration mode, the system randomly generates values ​​that satisfy: Action vector This avoids the strategy getting stuck in local optima and directly terminates the action selection process for that time slot. In exploitation mode, the subsequent action selection step based on the Q-value is then initiated. During the operational phase after system deployment, this step is no longer executed, and action selection is entirely based on the exploitation mode.

[0048] (2) Local optimal action selection For each access point The system is in a binary action set Minimizing the execution cost yields the locally optimal action: Based on this, candidate action vectors are formed: And define a candidate update set: (3) Quantity judgment If the number of candidate updates is satisfied If all locally optimal updates are within the maximum number of updates that the system can execute, then the system directly adopts the following approach: The action selection process ends at this point.

[0049] (4) Ranking of earnings when At this point, the system needs to further filter the candidate set. This invention defines the updated revenue amount: in This indicates that the cost of updating is reduced compared to not updating. The system will collect... Access points in the middle are based on revenue Sort by largest to smallest, then select the first... The access point with the highest revenue constitutes the update set: The final action vector is defined as: Through the aforementioned action selection mechanism, this invention will eliminate the need for... The problem of searching within a combination of actions is decomposed into computable sub-steps such as local optimum determination, quantity constraint checking, and payoff ranking, thus avoiding the exponential complexity introduced by the combination of actions space. This method satisfies the condition of updating at most one... Given the constraints of multiple access points, this invention can quickly select the access point that contributes most to cost reduction, achieving an approximate globally optimal decision. Therefore, this invention effectively solves the problem of difficult combination action selection, significantly improving the efficiency and scalability of action selection.

[0050] S5. Model Training This step involves model training based on the aforementioned MH-D3QN network structure. MH-D3QN consists of a global feature extraction module, a local feature extraction module, a feature fusion module, and a multi-head Dueling structure configured independently for each access point, and is implemented through an online Q-network. With the target Q network The invention employs a dual-network mechanism for training and optimization. Through a reinforcement learning framework, the parameters of MH-D3QN are continuously updated during environmental interactions, enabling it to learn the optimal state update strategy that satisfies update constraints in a dynamic environment.

[0051] The overall training process includes the following steps: environment interaction and hierarchical experience collection, hierarchical experience replay sampling, calculation of target Q value and expected update quantity, loss function and gradient update, target network soft update, and exploration rate decay.

[0052] (1) Environmental interaction and hierarchical experience collection At the start of each training round, the environment generates an initial state vector. In the time slot The online Q network of MH-D3QN is based on the current exploration rate. Select Action : based on probability Explore and randomly select within the constraints. ; with probability Perform an action selection based on the Q-value. After executing the action, the environment returns to the next state. Reward Vector With termination mark The generated quintuple: Where the action vector Indicates whether a status update is performed on each access point. Defines the action resource consumption (or update scale): but This invention constructs A hierarchical experience queue: And based on Write the experience into the corresponding queue: Through this hierarchical writing mechanism, the replay buffer can organize and manage experience samples according to the network update resource occupancy, enabling the training process to learn the state transition characteristics under different resource occupancy conditions.

[0053] (2) Layered experience playback sampling Once the experience buffer has accumulated to a certain size, a training mini-batch is constructed from the hierarchical queue set. Let the mini-batch size be... This invention employs a balanced hierarchical sampling strategy: for each queue sampling 10 samples, and merge them to obtain the real mini-batch: Where B represents the batch size of each sample taken from the experience playback buffer. It is the sampled set: These samples serve as the training input for MH-D3QN. This mechanism stratifies experience based on network update resource occupancy and performs balanced sampling across resource occupancy layers during sampling, suppressing the bias of training samples in the resource consumption dimension, thereby improving the policy convergence and execution robustness under update budget constraints.

[0054] (3) Calculation of target Q value and expected update quantity For each batch of samples, this invention first utilizes an online Q-network. Calculate the next state The action value matrix is ​​used to determine the optimal action for the next state based on the action selection mechanism. Subsequently, a target Q-network was used. Calculate the target Q-value for the action corresponding to the next state, and thus construct the Bellman objective: in The index of the Q value, As a discount factor, and These are trainable parameters. The online network is used for action selection, and the target network is used for target value calculation, making the training of MH-D3QN more stable and accelerating convergence.

[0055] Then, for the first One access point, utilizing the online Q network. Calculate the benefit of performing an update compared to not updating: in Indicates the first Each access point performs an update. This indicates no update. An estimate of the expected number of updates is calculated based on the update benefits: in This is the Sigmoid function.

[0056] (4) Loss function and gradient update This invention employs a smoothed L1 loss (Huber Loss) to construct a multi-head time-series difference loss function to reduce the error between the predicted and target values: in, Indicates the first in the batch One sample.

[0057] Furthermore, this invention also constructs a budget-consistent regularization term: This feature ensures that the network output maintains a consistent resource consumption tendency under budget constraints during training, improving deployment controllability and convergence stability.

[0058] Ultimately, the loss was: in These are the weighting coefficients. Backpropagation is then performed, and the Adam optimization method is used to adjust the online Q-network parameters. Update the parameters to adaptively adjust the learning step size and accelerate convergence.

[0059] (5) Target network soft update To ensure the smoothness of the target Q-network and prevent large fluctuations in the target value from causing training instability, this invention employs a dual-timescale soft update mechanism for the shared backbone and multi-head branch structure of MH-D3QN. Specifically, the online network parameters are decomposed into shared backbone parameters. Multi-head branch parameters configured independently per access point The corresponding target network parameters are as follows: and Its soft update rules are as follows: in, and These are the soft update coefficients for the shared backbone and multi-headed branches, respectively, and are set accordingly. This allows the shared backbone representation to evolve smoothly at a slower pace to suppress global target value oscillations, while each access point branch tracks local state changes at a faster pace, thereby improving convergence stability and enhancing the policy's adaptability under dynamic network conditions during training.

[0060] (6) Exploration rate decay and training termination After each training round, the exploration rate is updated exponentially: in, This is the lower limit of the exploration rate. This represents the decay coefficient of the exploration rate. Through the above update method, the exploration rate will gradually decrease as training progresses, allowing the MH-D3QN policy to smoothly transition from a strong exploration phase in the early stages of training to a decision-making mode that is primarily based on exploitation in the later stages. This is beneficial for achieving the convergence and stable execution of the final policy.

[0061] During training, this invention incorporates an early termination mechanism based on performance evaluation metrics. When performance no longer improves after multiple consecutive evaluation rounds, training is immediately terminated, and the optimal MH-D3QN parameters are saved as the final deployment model.

[0062] Experimental verification This invention selects Max AoI, Random, and MH-D3QN-QAoI as comparison methods. Specifically, the Max AoI method selects the method with the highest AoI in each time slot. Each access point updates its state. Existing research largely focuses on minimizing the average AoI as the optimization objective, which in this scenario can be considered an update strategy equivalent to Max AoI. The Random method, on the other hand, randomly selects an access point within each time slot. Each access point is updated. The MH-D3QN-QAoI method, under the MH-D3QN framework proposed in this invention, replaces the EDPT metric in the reward function with the QAoI metric, and performs training and decision-making based on the AoI at the time of user access.

[0063] This invention evaluates performance in a small-scale edge computing scenario, considering the number of access points. The value ranges from 15 to 20. The user access process exhibits periodic characteristics, used to simulate the periodic computational needs of intelligent devices in a smart factory; simultaneously, to reflect random disturbances in the actual production environment, this invention superimposes a normally distributed value onto the periodicity. The deviation is calculated to simulate more realistic fluctuations in access intervals. The processing capacity of the edge nodes is set to... The local terminal's processing power is set to The system allows a maximum number of access points to be updated within each time slot. Set to 2. Task size Follow the interval The uniform distribution.

[0064] During model training, the learning rate is recorded as 0.0002. The exploration rate... The initial value is set to And at the end of each training round, according to the decay factor The update will be performed, and its lower limit will be set as follows: The soft update coefficient of the target network is denoted as... The batch size for experience replay is denoted as The batch size of each layer is The discount factor is denoted as Update the penalty item as follows Weighting coefficient .

[0065] The experimental results are shown in the figure. Under different access point numbers, the MH-D3QN method of this invention has a lower overall task processing time than the comparative methods. As the number of access points increases, the processing time of the comparative methods shows a more significant upward trend, while the increase in the method of this invention is relatively smaller. This result indicates that the state update decision based on EDPT of this invention can allocate update opportunities more effectively under resource-constrained conditions, thereby reducing the impact of erroneous decisions caused by state inconsistency on overall processing efficiency.

[0066] In terms of decision accuracy metrics, MH-D3QN also demonstrates a significant advantage. Whether considering the decision error rate or the processing cost of erroneous decisions, the method of this invention maintains optimal performance. Since EDPT directly reflects the true cost of erroneous decisions, the model can more effectively distinguish the importance of different access points during training, reducing the occurrence of high-cost errors. Therefore, the Max AoI and MH-D3QN-QAoI methods based on AoI perform worse than MH-D3QN in these metrics, while the error of the stochastic strategy is more pronounced.

[0067] Furthermore, the method of this invention is more efficient in updating resource usage. Compared with the comparative methods, MH-D3QN significantly reduces unnecessary state updates, allowing update operations to be more focused on key access points, effectively reducing resource consumption. Under conditions of limited update budget, this invention achieves precise scheduling of update resources through action value assessment and screening mechanisms, thereby further improving the overall system performance while ensuring decision quality. Overall, the experimental results verify the comprehensive advantages of the method of this invention in terms of processing efficiency, decision accuracy, and resource utilization.

[0068] Example 2 This embodiment provides a multi-access point edge computing service state synchronization system based on a deep Q network, including: The data acquisition module is configured to acquire access point data; The metric construction module is configured to construct EDPT metrics based on the acquired access point data; The model building module is configured to build an MH-D3QN network model, which includes MDP modeling based on the constructed EDPT metric; calculating the Q-value of the access point based on the MH-D3QN algorithm; and ranking the access points by their Q-values. The model training module is configured to train the model based on the constructed MH-D3QN network model; The optimization module is configured to use the trained MH-D3QN network model to synchronize the state of multi-access point edge computing services.

[0069] A computer-readable storage medium storing a plurality of instructions adapted for loading and execution by a processor of a terminal device of the aforementioned method for synchronizing the state of a multi-access point edge computing service based on a deep Q network.

[0070] A terminal device includes a processor and a computer-readable storage medium, the processor being configured to implement various instructions; the computer-readable storage medium being configured to store multiple instructions adapted for loading and execution by the processor of the aforementioned method for synchronizing the state of a multi-access point edge computing service based on a deep Q network.

[0071] The above are all preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Therefore, all equivalent changes made in accordance with the structure, shape and principle of the present invention should be covered within the scope of protection of the present invention.

Claims

1. A method for synchronizing the state of multi-access point edge computing services based on deep Q-networks, characterized in that, include: Obtain access point data; The EDPT metric is constructed based on the acquired access point data; Construct an MH-D3QN network model, in which MDP modeling is performed based on the constructed EDPT metric; Calculate the Q value of the access point based on the MH-D3QN algorithm; Rank the revenue based on the Q value of the access point; Model training was performed based on the constructed MH-D3QN network model; The trained MH-D3QN network model is used to synchronize the state of edge computing services at multiple access points. The construction of the EDPT metric based on the acquired access point data includes first performing discrete time slot modeling, where the time slot is composed of... Index, including One access point and one edge node, and adopt Indicates the edge node in time To the An access point sends an update, assuming that the request transmission and response reception each consume one time slot, and the user's local computing power is denoted as... The computing power of edge nodes is In time The state information of an edge node is its load, denoted as . And the first The state of the edge node sensed by an access point is denoted as... User visits the The data size carried by one access point is ; Then, based on the principle of minimizing the estimated processing time, the optimal offloading decision is selected using the offloading strategy of the access point. Indicates the first Each access point in time The decision is expressed as: in, Indicates the index of the access point. Indicates time, This indicates the time required for the user to process the task locally. This represents the estimated time for edge nodes to process tasks; in time... Utilizing the load of edge nodes Calculate the actual processing time and the right decision And construct the EDPT metric: in Indicates the index of the access point. To indicate time, finally use an indicator variable. Indicates the first Each access point in time Whether it has been accessed, among which This indicates that the site has been visited. The final optimization goal is: ; The MDP modeling based on the constructed EDPT index includes constructing a corresponding MDP model based on the EDPT index, wherein the state... It consists of the state of edge nodes and the state of each access point. The state of an edge node contains a single variable, namely the current load. The load on an edge node depends on the remaining load from the previous time step, the amount of tasks arriving currently, and the system's service capacity, and is expressed as: Access point status is determined by the load sensed by the access point. Freshness and user access interval Composition, in which It is time. It is an index of the access point, a freshness metric. The timeliness of load information, while the user access interval. This is used to characterize user access characteristics, system state space The state of an edge node and The state composition of each access point: Therefore, the state space yes ;action The dimension is Each component represents whether to send a state update to the corresponding access point. Within each time slot, a binary decision in the action vector is used to determine whether the edge node should send a state update to the corresponding access point. Each access point sends the latest load information; rewards Defined as a dimension The vector, where the first... Each component corresponds to the user at time [time]. Visit the The reward generated when the first access point is reached, the first Each access point at time The reward is represented as: After completing the MDP modeling, the overall optimization objective is ( This can be expressed as: finding the optimal strategy. This minimizes the expected state value of the system under the constraints. in, It is the initial state, and the state value function. Defined in strategy Below, from the state The expected cumulative discount reward for departure is expressed as: Among the discount factors This is used to ensure the convergence of the value function while controlling the effective time range of the reward.

2. The method for synchronizing the state of multi-access point edge computing services based on deep Q-networks according to claim 1, characterized in that, The calculation of the access point Q value based on the MH-D3QN algorithm includes time slots. The MH-D3QN algorithm is used to determine the current system state. Binary actions with each access point Cost estimation is performed, and edge nodes are defined in time slots. For the Each access point performs an action. The action value function at time is: The Q-value matrix output by the network has a dimension of , of which The rows correspond to and In the time slot State vector After inputting into the Q network, the action costs of each access point are obtained sequentially, including first setting the global state. The input is processed by a global feature extraction network and mapped to a global feature vector through a two-layer feedforward fully connected structure. Formal representation: ,in It is a global feature extraction network, global features Replication and expansion at the access point level enable the system to operate within time slots. All access points share the same set of global context information, and then local feature extraction is performed for each access point. In the time slot Take its local state triples and denote them as: The local vector is input into a local feature encoding network, and after being encoded by a two-layer feedforward fully connected structure, the local feature vector is obtained. , represented as: ,in It is a local feature extraction network, which extracts features from all... Feature extraction is performed in parallel at each access point within a time slot. Obtain local feature representations that characterize the differences in load perception, freshness, and user access behavior at each access point.

3. The method for synchronizing the state of multi-access point edge computing services based on deep Q-networks according to claim 2, characterized in that, The calculation of the access point Q value based on the MH-D3QN algorithm also includes processing each access point... Local feature vectors With global feature vectors The features are concatenated along the feature dimension to obtain a fused input, which is then fed into a feature fusion network to obtain a fused feature vector. : ,in It is a feature fusion network that fuses features. It includes both system-level global information and access point-level local information, and finally uses a Dueling network to calculate the Q-value, where the time slot... fusion of feature vectors Input to the and the The Dueling structure network corresponding to the access point, the th The Dueling structure for each access point consists of two parts: a state value branch and an action advantage branch. The state value branch is used to calculate the access point's performance in the time slot. The state value is defined as: in Indicates the first Each access point has its own independent value network, and the action advantage branch is used to calculate the access point's action advantage. and time slots The advantage value is defined as: in Indicates the first The independent advantage network corresponding to each access point, based on the combination form of the Dueling structure, yields the edge node's position in the time slot. For the first Each access point performs an action. Action value function: After performing calculations in parallel on all access points, in the time slot Obtain the complete action value matrix: By configuring a Dueling structure network with independent parameters for each access point, the action value of multiple access points can be estimated in parallel within a unified network framework.

4. The method for synchronizing the state of multi-access point edge computing services based on deep Q networks according to claim 3, characterized in that, The step of ranking access points based on their Q-values ​​includes obtaining updated revenue for each access point based on the Q-values ​​calculated using MH-D3QN, and then ranking and selecting candidate access points according to these updated revenues. Represents the matrix of the first Line 1 The elements of a column are represented as: Then adopt - Greedy strategy, based on probability Perform exploration, with probability Perform Q-value-based exploitation, randomly generating variables that satisfy the following in exploration mode: Action vector However, when utilizing the Q-value, a locally optimal action selection is performed for each access point. In the binary action set Minimizing the execution cost yields the locally optimal action: Based on this, candidate action vectors are formed: And define a candidate update set: If the number of candidate updates meets the requirement If all locally optimal updates are within the maximum number of updates that the system can execute, then we can directly adopt the following approach: The selection process is complete; when Then, further refine the candidate set by defining and updating the payout: in This indicates that the cost of updating the action is reduced compared to not updating the action, and the set... Access points in the middle are based on revenue Sort by largest to smallest, then select the first... The access point with the highest profit constitutes the update set: The final action vector is defined as: .

5. The method for synchronizing the state of multi-access point edge computing services based on deep Q-networks according to claim 4, characterized in that, The model training based on the constructed MH-D3QN network model includes updating the parameters of the MH-D3QN during environmental interactions, enabling it to learn the optimal state update strategy that satisfies update constraints in a dynamic environment. The training process includes: environmental interaction and hierarchical experience collection, hierarchical experience replay sampling, calculation of the target Q-value and expected update quantity, loss function and gradient update, target network soft update, and exploration rate decay. Specifically, environmental interaction and hierarchical experience collection includes generating an initial state vector at the beginning of each training round. In the time slot The online Q network of MH-D3QN is based on the current exploration rate. Select Action : based on probability Explore and randomly select within the constraints. ; with probability Perform an action selection based on the Q-value; after executing the action, the environment returns to the next state. Reward Vector With termination mark The generated quintuple: According to the update scale Write to the corresponding experience replay buffer Layered experience replay sampling specifically includes randomly sampling from each layer of the queue after the experience buffer has accumulated to a certain amount. These samples constitute a batch of samples: Where B represents the batch size of each sample taken from the experience playback buffer. It is the sampled set. The experience is stratified according to the network update resource consumption, and each resource consumption layer is sampled evenly during sampling to suppress the bias of training samples in the resource consumption dimension, thereby improving the policy convergence and execution robustness under update budget constraints.

6. The method for synchronizing the state of multi-access point edge computing services based on deep Q-networks according to claim 5, characterized in that, The model training based on the constructed MH-D3QN network model also includes calculating the target Q-value and the expected update quantity. For each batch of samples, an online Q-network is first used. Calculate the next state The action value matrix is ​​used to determine the optimal action for the next state based on the action selection mechanism. Subsequently, the target Q network was used. Calculate the target Q-value for the action corresponding to the next state, and thus construct the Bellman objective: in The index of the Q value, As a discount factor, and These are trainable parameters; then the access point is calculated. Update revenue And the estimated number of updates ,in The Sigmoid function is used; then the loss function is calculated, and a multi-head time-series difference loss function is constructed using Huber loss to reduce the error between the predicted and target values: in Indicates the first in the batch We then construct a consistent regularized loss for each sample to ensure a consistent resource consumption tendency under budget constraints. The final loss is: in These are the weighting coefficients; then backpropagation is performed, and the Adam optimization method is used to optimize the online Q-network parameters. Update the parameters to adaptively adjust the learning step size and accelerate convergence.

7. The method for synchronizing the state of multi-access point edge computing services based on deep Q-networks according to claim 6, characterized in that, The model training based on the constructed MH-D3QN network model also includes soft updates of the target network. To ensure the smoothness of the target Q-network and prevent large fluctuations in the target value from causing training instability, a dual-time-scale soft update mechanism is adopted for the shared backbone and multi-branch structure of MH-D3QN. This includes updating the shared backbone parameters of the online network. Multi-head branch parameters configured independently per access point and the corresponding target network parameters and Its update rules are as follows: in and Set the soft update coefficients for shared backbone and multi-headed branches respectively. This allows the shared backbone representation to evolve smoothly at a slower pace to suppress global target value oscillations, while each access point branch tracks local state changes at a faster pace; then, exploration rate decay and training termination are performed, with the exploration rate updated exponentially after each training round. in, This is the lower limit of the exploration rate. The decay coefficient of the exploration rate is updated to gradually reduce the exploration rate as training progresses, so that the MH-D3QN policy can smoothly transition from a strong exploration phase in the early stage of training to a decision mode that is mainly based on exploitation in the later stage, which is conducive to the convergence and stable execution of the final policy.

8. A multi-access point edge computing service state synchronization system based on a deep Q-network, executing the multi-access point edge computing service state synchronization method based on a deep Q-network as described in claim 1, characterized in that, include: The data acquisition module is configured to acquire access point data; The metric construction module is configured to construct EDPT metrics based on the acquired access point data; The model building module is configured to build an MH-D3QN network model, in which MDP modeling is performed based on the constructed EDPT metric; Calculate the Q value of the access point based on the MH-D3QN algorithm; Rank the revenue based on the Q value of the access point; The model training module is configured to train the model based on the constructed MH-D3QN network model; The optimization module is configured to use the trained MH-D3QN network model to synchronize the state of multi-access point edge computing services.

Citation Information

Patent Citations

  • D3QN-based edge enabling IIOT online computing migration method

    CN117196008A

  • Vehicle task unloading system and method based on deep Q network

    CN119938277A