Edge user allocation method and system based on sequential perception reinforcement learning
Through the method of sequentially perceptual reinforcement learning, predicting and optimizing the allocation order of edge users, the problems of low resource utilization and poor service experience in mobile edge computing environments are solved, and more efficient resource utilization and better service experience are achieved.
Patent Information
- Application Number
- CN202510165561.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-27
AI Technical Summary
In mobile edge computing environments, it is difficult for the prior art to effectively allocate edge users to edge servers, especially when the number of users and server resources change rapidly, resulting in low resource utilization, high latency and poor service experience.
The edge user allocation method based on sequential perception reinforcement learning is adopted. By constructing feature data and graph neural networks, the adjacency relationship between users and servers is processed, combined with multi-head attention mechanism and greedy strategy, the user allocation order is predicted and the allocation plan is optimized, and resource utilization and service experience are improved.
It improves the efficiency of network bandwidth, computing resources and storage resources utilization, and improves the overall network performance and service experience, especially when resources are limited and the number of users is large.
Smart Images

Figure CN120050719A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of edge user allocation, and particularly to an edge user allocation method and system based on sequential perception reinforcement learning. Background Art
[0002] With the rapid development of mobile computing technology, computing-intensive applications such as augmented reality, virtual reality, and artificial intelligence have become increasingly popular. However, due to the limitations of resources such as memory, computing power, and battery capacity of mobile devices, the performance and efficiency of mobile devices have been severely reduced. To solve this problem, many experts and scholars have outsourced computing tasks to the cloud to reduce the burden on mobile resources. For example, Google's Stadia project has built a cloud-based game platform that enables mobile users to enjoy high-definition video game experiences without downloading a complete game client. However, the method of outsourcing computing tasks to a remote cloud has problems such as large communication overhead, high energy consumption, and low efficiency. Fortunately, with the development of the fifth-generation communication technology, the emergence of Mobile Edge Computing (MEC) technology provides a solution to the above problems. In the MEC mode, the base station is not only responsible for wireless communication, but also equipped with an appropriate amount of computing infrastructure, that is, an MEC server, and allows mobile application providers such as Google, Facebook, or Netflix to deploy services on these servers, so that users can directly access application services without relying on a remote cloud. Compared with the traditional mode, MEC technology processes most tasks on the MEC server, greatly reducing communication overhead and energy consumption.
[0003] Although MEC technology has many advantages, it also faces some challenges in practical applications. In the MEC environment, each MEC server has a specific service coverage area and can only provide computing services for users within the service radius. Therefore, when allocating computing tasks, distance limitations must be considered to ensure the effectiveness of the service. In addition, the resource status (such as computing power, memory, etc.) of the MEC server determines its task processing ability. When making a decision on task offloading, it is necessary to consider whether the resources of the MEC server can meet the user's needs. Lai et al. defined such an allocation task as an Edge User Allocation (EUA) problem. Currently, the quality of EUA solutions is highly correlated with the investment and revenue of service providers.
[0004] In recent years, the EUA problem has attracted extensive attention from many researchers and practitioners at home and abroad. A series of research results and practical application progress have been achieved in the research on the EUA problem, which can be summarized into four categories: methods based on integer linear programming, heuristic methods, game theory methods, and reinforcement learning methods.
[0005] (1) Integer Linear Programming (ILP) is a method in mathematical optimization or operations research. It aims to find a set of integer solutions that can satisfy a series of linear equality or inequality constraints while maximizing or minimizing a given linear objective. The EUA problem involves a series of discrete decisions, such as whether to assign a specific user to a certain edge server. This is a typical binary decision (yes / no), which exactly matches the integer decision variables in integer linear programming. Moreover, the constraints in the EUA problem are usually linear, such as the capacity limit of the server or the service scope limit. Therefore, these constraints can be expressed as linear equalities or inequalities, falling within the scope of integer linear programming. Since integer linear programming can find the optimal solution by optimizing the linear objective function, for EUA problems such as maximizing service coverage, minimizing latency, or maximizing the quality of service experience, the method of integer linear programming can be used to solve them.
[0006] Lai et al. (LAIP, HE Q, ABDELRAZEK M, et al. Optimal edge user allocation in edge computing with variable sized vector bin packing[C] / / Service-Oriented Computing: 16th International Conference, ICSOC 2018, Hangzhou, China, November 12 - 15, 2018, Proceedings 16. Springer, 2018: 230 - 245.) first formulated the EUA problem as a variable sized vector bin packing problem and established an ILP-EUA (Integer Linear Programming for Edge User Allocation) model. This model, with the help of the IBM ILOG CPLEX optimizer, aims to maximize the number of mobile user accesses while minimizing the number of required servers, and solves the EUA problem.
[0007] Subsequently, Lai et al. (LAIP, HE Q, CUI G, et al. Edge user allocation with dynamic quality of service[C] / / ServiceOriented Computing: 17th International Conference, ICSOC 2019, Toulouse, France, October 28–31, 2019, Proceedings 17. Springer, 2019: 86-101.) further refined the problem, improved the ILP-EUA method, introduced Quality of Service (QoS) as a requirement, and changed the optimization goal from reducing the number of edge servers to improving the overall Quality of Experience (QoE).
[0008] However, the above methods did not consider users' preferences for QoS. Panda et al. (PANDA S P, RAY K, BANERJEE A. Dynamic edge user allocation with user specified qos preferences[C] / / Service-Oriented Computing: 18th International Conference, ICSOC 2020, Dubai, United Arab Emirates, December 14–17, 2020, Proceedings 18. Springer, 2020: 187-197.) proposed an allocation method centered on user needs, focused on users' preferences for service quality, and gave a better QoE optimization scheme.
[0009] Li et al. (LI B, HE Q, CUI G, et al. READ: robustness-oriented edge application deployment in edge computing environment[J]. IEEE Trans. Serv. Comput., 2022, 15(3): 1746-1759.) considered the stability of application services, defined it as a constrained optimization problem, and solved it precisely by integer programming methods. Liu et al. (LIU F, LV B, HUANG J, et al. Edge user allocation in overlap areas for mobile edge computing[J]. Mob. Networks Appl., 2021, 26(6): 2423-2433.) studied the user allocation problem in overlapping areas, modeled it as a multi-objective optimization model, and used the Pareto boundary search and principal component analysis algorithms to determine the allocation strategy, aiming to balance the server workload and minimize the access latency.
[0010] (2) Heuristic methods usually rely on experience, intuition, or specific rules to quickly generate feasible solutions without having to perform a complete problem-solving process. In the face of complex optimization problems, traditional exact algorithms may not be applicable due to excessive computation time. Heuristic algorithms search the solution space iteratively, using heuristic rules or evaluation metrics to filter and optimize candidate solutions, aiming to find a suboptimal solution within an acceptable time. Since the EUA problem is an NP (Non-deterministic Polynomial-time) hard problem, it is also effective to use heuristic algorithms to solve the EUA problem.
[0011] In addition to proposing an integer linear programming method, Lai et al. (LAIP, HE Q, CUI G, et al. Edge user allocation with dynamic quality of service[C] / / Service Oriented Computing: 17th International Conference, ICSOC 2019, Toulouse, France, October 28–31, 2019, Proceedings 17. Springer, 2019: 86-101.) also proposed the GA-EUA (Greedy Algorithm for Edge User Allocation) method and the heuristic QoEUA method to solve the large-scale EUA problem. When the GA-EUA algorithm aims to maximize the service coverage rate, it improves the number of allocated users by sequentially allocating users to the server with the most abundant nearby resources; when it aims to maximize the quality of service experience, it improves the quality of service experience by sequentially allocating users to the server with the most abundant nearby resources at the desired quality of service level. The QoEUA method is mainly used to optimize the quality of service experience. First, the users in the user set $U$ are sorted in ascending order according to the number of neighboring edge servers, and then the users are sequentially allocated to the server with the most abundant nearby resources at the lowest quality of service level, and the quality of service level is continuously increased through iteration until the resources on the server are exhausted or all users are allocated at the desired level.
[0012] In addition to the integer linear programming method, Panda et al. (PANDA S P, RAY K, BANERJEE A. Dynamic edge user allocation with user specified qos preferences[C] / / Service-Oriented Computing: 18th International Conference, ICSOC 2020, Dubai, United Arab Emirates, December 14–17, 2020, Proceedings 18. Springer, 2020: 187-197.) also proposed an IHA-EUA (I-factor Heuristic Algorithm for Edge User Allocation) algorithm to optimize resource allocation according to the time-varying QoS requirements of users. The IHA-EUA algorithm is equivalent to an upgrade based on the GA-EUA algorithm. It classifies users into single-server users (S-class) and multi-server users (M-class) according to the number of edge servers near the users, and preferentially allocates resources to S-class users. Then, according to the difference between the currently allocated quality of service level of the user and the desired quality of service level, its i-factor value is calculated, and the users are ranked according to this value. Users with lower i-factor values are given priority to improve their QoS levels, and this process continues until all users reach the desired QoS level or the server resources cannot support further improvement.
[0013] (3) Using game theory methods to solve the EUA problem is one of the main research methods currently. Game theory models can help analyze and design the interaction and decision-making processes among players (such as users, application providers, or edge server providers) to achieve the optimization of resource allocation or the maximization of user service experience.
[0014] He et al. (HE Q, CUI G, ZHANG X, et al. A game-theoretical approach for user allocation in edge computing environment[J]. IEEE Trans. Parallel Distributed Syst., 2020, 31(3): 515-529.) proposed an EUA Game method based on game theory. This method simulates the user allocation of application providers in the edge computing environment as a potential game, proves that there is at least one Nash equilibrium point in this game model, and then optimizes the service coverage rate through a distributed algorithm. This research breaks through the limitations of traditional centralized allocation strategies and improves the allocation efficiency and user satisfaction.
[0015] Since the above scheme does not consider the wireless interference generated by parallel communication between multiple devices and edge servers, Cui and Lai et al. (CUI G, HE Q, XIA X, et al. Interference-aware saas user allocation game for edge computing[J]. IEEE Trans. Cloud Comput., 2022, 10(3): 1888-1899.) focused on how to effectively allocate mobile devices to edge servers in the mobile edge computing environment. They considered the possible communication interference between devices and between devices and base stations, redefined the problem, and achieved Nash equilibrium through a decentralized method, improving the performance and efficiency of the system.
[0016] Lai et al. (LAI P, HE Q, CUI G, et al. Quality of experience-aware user allocation in edge computing systems: A potential game[C] / / 40th IEEE International Conference on Distributed Computing Systems, ICDCS 2020, Singapore, November 29 - December 1, 2020. IEEE, 2020: 223-233.) also changed the optimization goal from service coverage rate to quality of service experience and used a distributed algorithm under the same game theory framework to achieve the optimization of QoE.
[0017] Xia et al. (XIA X, CHEN F, HE Q, et al. Data, user and power allocations for caching in multi-access edge computing[J]. IEEE Trans. Parallel Distributed Syst., 2022, 33(5): 1144-1155.) made the first attempt to study the joint data, user and power allocation problem and proposed a two-stage decentralized algorithm called DUPAGame, aiming to maximize the user coverage and the overall data rate of users. This algorithm uses the concept of game theory, simulates all users as game participants, and makes decisions based on a specific payoff function.
[0018] To accelerate convergence, Kumar et al. (KUMAR S, GOSWAMI A, GUPTA R, et al. A cost-effective and qos-aware user allocation approach for edge computing enabled iot[J]. IEEE Internet Things J., 2023, 10(2): 1696-1710.) clustered the edge servers and ran the distributed edge server allocation algorithm in parallel in each cluster, so as to achieve rapid convergence to a pure Nash equilibrium in polynomial time.
[0019] (4) Reinforcement learning learns the optimal decision-making strategy by interacting with the environment. In the EUA problem, the environment is dynamically changing. For example, factors such as the number of users, demands, and edge server loads are constantly changing. Reinforcement learning can interact with this dynamic environment, continuously learn and adapt according to the feedback, and find the optimal user allocation strategy.
[0020] In recent years, some scholars have begun to use reinforcement learning to solve the EUA problem. Panda et al. (PANDAS P, BANERJEE A, BHATTACHARYA A. User allocation in mobile edge computing: A deep reinforcement learning approach[C] / / 2021 IEEE International Conference on Web Services(ICWS). IEEE, 2021: 447-458.) proposed the CLDQN-EUA (Considering Latency Deep Q Network for Edge User Allocation, CLDQN-EUA) algorithm. The CLDQN-EUA algorithm uses a Deep Q-Network (DQN) to predict the number of users that an edge server can serve under a given latency threshold, and then uses a greedy algorithm to allocate users. Their method reasonably evaluates and utilizes the service capabilities of the server, but in the subsequent allocation stage, a heuristic algorithm is still used to allocate users, and this method still has certain limitations in resource-scarce environments.
[0021] Therefore, Chang et al. (CHANG J, WANG J, LI B, et al. Attention-based deep reinforcement learning for edge user allocation[J]. IEEE Trans. Netw. Serv. Manage., 2024, 21(1): 590-604.) handed over the allocation process entirely to a deep reinforcement learning framework, proposed the ADRL-EUA (Attention-based Deep Reinforcement Learning for Edge User Allocation, ADRL-EUA) algorithm, and introduced an attention mechanism to dynamically select and focus on important input information. This can make the model pay more attention to features related to user allocation, thereby improving the learning ability and effect of the model. However, they only considered the EUA problem under the scenario of maximizing service coverage. In addition, their solution can only target the edge user allocation scenario with a fixed scale, that is, the scenario with a fixed number of users. Once the number of users changes, their network needs to be retrained, which takes a long time.
[0022] Li et al. (BAO L, GAO S, LI Z. Qos preferences edge user allocation using reinforcement learning[C] / / 2023 IEEE Cloud Summit. IEEE, 2023: 21-26.) aimed to maximize the quality of service experience and designed and implemented a hybrid edge user allocation algorithm based on reinforcement learning and greed, called GreedRL-EUA (Greed and Reinforcement Learning for Edge User Allocation, GreedRL-EUA). They first modeled the edge user allocation problem as a Markov decision model, and then decomposed the action into two parts. Action one was to select a suitable server for the user, and action two was to select a suitable quality of service level for the user. They selected a server with sufficient resources from many servers in a greedy manner, and the agent used a neural network to select a suitable quality level for the user and gave an allocation plan. Since their policy network directly predicted the allocation plan without considering the capacity constraint of the server, the allocation plan output by the policy network was very likely not an effective solution that met the constraints. The policy network needed a long time to learn, and their solution could only be applied to the edge user allocation scenario with a fixed scale. Moreover, since the target edge server of the user was pre-determined in the initial stage, this led to multiple users gathering on the same server, resulting in an imbalance in resource allocation. Even if reinforcement learning was subsequently used to optimize the quality of service level, its improvement space was limited by the unsatisfactory initial allocation and could only achieve relative optimization within a limited range. Therefore, the quality of the solution needed to be improved urgently.
[0023] In summary, although integer linear programming can effectively obtain the optimal solution in scenarios with a small number of edge servers and user scale. However, as the scale increases, the EUA problem becomes more complex, resulting in low efficiency in solving integer linear programming and difficulty in meeting the user's requirements for low latency. In addition, the method based on integer linear programming is more suitable for static or offline EUA scenarios and is not applicable to scenarios where users dynamically join or leave.
[0024] Although the heuristic-based method can ensure a feasible allocation plan quickly, the time required to find a high-quality solution is still too long, and the quality of the solution obtained in resource-scarce scenarios is poor.
[0025] The game theory-based method can allocate resources to users to meet their needs, but it is difficult for game theory methods to cope with the dynamics and uncertainties in the edge environment, including changes in user needs, fluctuations in available resources, etc.
[0026] Using reinforcement learning technology to solve the EUA problem is not yet mature, and a more appropriate research plan is urgently needed. Summary of the Invention
[0027] To solve the above problems, the present invention proposes an Order-Aware Reinforcement Learning for Edge User Allocation (OARL-EUA) method and system, which focuses on predicting the user allocation order based on a pointer mechanism, and then allocates users sequentially based on a greedy strategy to improve the user allocation ratio and the universality of the model. By reasonably allocating the requests and tasks of edge users, the present invention can effectively improve the utilization efficiency of network bandwidth, computing resources, and storage resources, thereby improving the overall network performance and service experience.
[0028] The technical solution adopted by the present invention is as follows:
[0029] An edge user allocation method based on order-aware reinforcement learning, comprising:
[0030] Constructing feature data for reinforcement learning based on edge user and server information, and performing user allocation based on the feature data, and optimizing through a policy network and a value network;
[0031] Converting the user allocation order output by the policy network into a detailed allocation plan through a greedy strategy, and calculating the reward; determining the core components of reinforcement learning and performing corresponding evaluations;
[0032] Optimizing the policy network and the value network, initializing the network parameters, and improving the network performance through training updates, thereby realizing edge user allocation.
[0033] Further, the constructing feature data for reinforcement learning based on edge user and server information includes:
[0034] Processing the information of edge users and servers, and converting it into feature data that can be utilized by reinforcement learning; using a graph neural network to process the adjacency relationship between users and servers, so as to capture more complex dependencies;
[0035] Adopting a method of vector normalization and concatenation to fuse information from different data sources, including user demand vectors, server resource vectors, and the adjacency relationship between edge users and edge servers; further enriching the user feature representation through a multi-head attention mechanism;
[0036] Extra features are introduced, including historical load features and user behavior pattern features. The historical load features represent the resource usage of the server over a past period of time, and the user behavior pattern features represent the usage behaviors and demand changes of users over a past period of time.
[0037] Furthermore, the historical load features include CPU usage rate, memory usage rate, and network bandwidth usage rate; the calculation method of the historical load features includes: regularly recording the resource usage of the server and representing the historical load features in the form of a time series; the user behavior pattern features include request frequency, data transfer volume, and average session duration; the calculation method of the user behavior pattern features includes: monitoring the usage of users, recording the request timestamps, data transfer volume, and session duration of users to generate user behavior logs; parsing the user behavior logs to extract the behavior patterns within a specific time window; calculating the statistical features including request frequency, average data transfer volume, and average session duration, and representing the user behavior patterns in the form of a time series.
[0038] Furthermore, the user allocation based on the feature data is optimized through a policy network and a value network, including:
[0039] The constructed feature data is passed as a state input to the reinforcement learning agent, and more comprehensive state inputs are provided by combining the environmental state and external factors. The environmental state and external factors include network bandwidth fluctuations, server health status, and user location.
[0040] The user allocation order is output through the policy network, and a confidence score is output. The allocation order is optimized by combining the confidence score. The confidence score represents the certainty of the policy network for each user allocation decision, and the maximum probability value is extracted from the probability distribution of each user as the confidence score for the user allocation decision.
[0041] The value network is used to evaluate the quality of the policy. Multi-objective optimization is introduced, and various factors are considered when evaluating the policy. The weighted sum method is used to assign different weights to different evaluation indicators and calculate the comprehensive evaluation value.
[0042] Furthermore, the user allocation order output by the policy network is converted into a detailed allocation plan through a greedy policy, and a reward is calculated, including:
[0043] The greedy algorithm is used to convert the user order output by the policy network into a detailed allocation plan, and the A* search algorithm is combined to further optimize the allocation result on the basis of the greedy policy.
[0044] After generating the allocation plan, calculate the rewards based on the number of users allocated in the plan, design a piecewise reward function, and provide fine-grained feedback; the piecewise reward function includes: dividing the possible number of user allocations into several intervals, defining corresponding reward values for each interval, and determining the interval where the actual number of user allocations is located and giving the corresponding reward value.
[0045] Further, determining the core components of the reinforcement learning and conducting corresponding evaluations includes:
[0046] Determining the state of the reinforcement learning, where the state is represented by the constructed feature data and extended to include time series information to capture dynamic changes;
[0047] Determining the actions of the reinforcement learning, where the actions include indirect actions and direct actions, the indirect actions include the user allocation order generated by the policy network, and a multi-level action space is introduced; the actions include the behaviors taken by the agent in a certain state;
[0048] Determining the rewards of the reinforcement learning, where the rewards are used to evaluate the quality of the actions, and thus guide the training of the agent. Design a multi-level reward function, consider short-term and long-term benefits, and introduce a penalty mechanism to avoid resource waste.
[0049] Further, the optimization of the policy network and the value network includes:
[0050] The optimized policy network includes a first encoder and a second decoder based on the long short-term memory network. The first encoder inputs user features one by one and converts them into a sequence of latent memory states; the second decoder maintains the sequence of latent memory states, captures relevant information within the sequence of latent memory states through the attention mechanism, and generates a probability distribution of the next user number to be allocated, and combines the copy mechanism to improve the adaptability of unknown user allocation.
[0051] The optimized value network includes a second encoder and a second decoder, and the second encoder is based on the long short-term memory network; the value network parametrically learns the expected number of users to be served by the current policy given the input, and is trained by the stochastic gradient descent method, and is trained on the mean square error between the predicted value and the number of user allocations obtained by the recent policy.
[0052] Further, the initialization of the network parameters includes initializing the parameters of the policy network and the value network through pre-training. The pre-training includes pre-training the policy network and the value network on the publicly available datasets in the relevant fields to make the initial parameters closer to the final optimal parameters, thereby accelerating convergence.
[0053] Further, the improvement of the network performance through training updates includes:
[0054] Sample a batch of users from the dataset as input and use an intelligent sampling strategy to improve sampling efficiency; the intelligent sampling strategy includes selecting representative data samples from the dataset through an optimized sampling method;
[0055] Perform forward propagation of the policy network and the value network to obtain the allocation priority and the expected number of served users of the users, and introduce a dropout layer to prevent overfitting;
[0056] Calculate the user allocation scheme based on the output of the policy network and the greedy policy, and use parallel computing to accelerate the execution of the greedy policy;
[0057] Update the parameters of the policy network and the value network through the gradient descent method, and use the self-supervised learning method to enhance the generalization ability of the value network.
[0058] An edge user allocation system based on sequential-aware reinforcement learning, comprising:
[0059] A first processing module, configured to construct feature data for reinforcement learning based on edge user and server information, perform user allocation based on the feature data, and optimize through a policy network and a value network;
[0060] A second processing module, configured to convert the user allocation order output by the policy network into a detailed allocation scheme through a greedy policy, and calculate the reward; determine the core components of the reinforcement learning, and perform corresponding evaluations;
[0061] A third processing module, configured to optimize the policy network and the value network, initialize the network parameters, and improve the network performance through training updates, so as to realize edge user allocation.
[0062] The beneficial effects of the present invention are as follows:
[0063] The present invention changes the strategy of directly predicting the allocation result, instead predicts the user allocation order, and by adopting a greedy algorithm, allocates users to the edge server with the most abundant resources one by one according to the predicted user allocation order, so as to obtain the actual allocation scheme; adding a pointer mechanism to the order prediction enables the model to effectively point to a specific position in the input sequence, rather than predicting index values from a fixed-size vocabulary. The present invention can reasonably allocate the requests and tasks of edge users, can effectively improve the utilization efficiency of network bandwidth, computing resources and storage resources, thereby improving the overall network performance and service experience.
[0064] The embodiments of the present invention conduct experiments using the real-world dataset EUA Datasets, which is widely used for researching EUA problems. The dataset covers a large amount of real mobile user and base station information in the metropolitan area of Melbourne, Australia. By varying three parameters: the number of edge users (n), the number of candidate edge servers (m), and the average resource capacity of edge servers (c), the experiments simulate different scenarios and evaluate the performance of the OARL-EUA algorithm, including the number of user allocations, the user allocation ratio, and the user allocation time. These experimental settings ensure a comprehensive evaluation of the applicability and robustness of the algorithm in various practical environments.
[0065] The present invention is compared with four existing EUA algorithms, including the Fruit Fly Optimization Algorithm-based Edge User Allocation (FOA-EUA) algorithm, the Greedy Algorithm-based Edge User Allocation (GA-EUA) algorithm, the Integer Linear Programming-based Edge User Allocation (ILP-EUA) algorithm, and the Attention Mechanism-incorporated Deep Reinforcement Learning-based Edge User Allocation (ADRL-EUA) algorithm. Through comparison with these algorithms, the performance of the present invention in different metrics is comprehensively demonstrated.
[0066] In the EUA problem of maximizing service coverage, the effectiveness of the algorithm is mainly reflected by the number of user allocations or the user allocation ratio. Figure 6 and Figure 7 The experimental results of
[0067] show that regardless of the scale of the dataset used in the experiment, the ILP-EUA algorithm has the best effect, and the OARL-EUA algorithm proposed by the present invention ranks second. However, the ILP-EUA algorithm takes an extremely long time and is usually only regarded as the theoretically optimal solution and is not actually adopted. The specific analysis is as follows: Figure 6 (a), Figure 6 (d), Figure 7 (a) and Figure 7 (d) show that with the number of servers and the amount of resources on the servers fixed, as the number of users increases, the overall number of user allocations is continuously increasing, but the user allocation ratio is continuously decreasing. This is because with limited resources, the number of users who can be allocated cannot increase indefinitely. The OARL-EUA algorithm has more obvious advantages when the number of users is larger and is superior to other non-ILP-EUA algorithms.
[0068] Change in the number of servers: Figure 6 (b), Figure 6 (e), Figure 7 (b) and Figure 7(e) shows that the number of fixed-edge users and the number of resources on the server increase as the number of servers increases. The number and proportion of users allocated increase continuously, but the increasing amplitude gradually decreases. The OARL-EUA algorithm has more obvious advantages when the number of servers is small.
[0069] Server resource changes: Figure 6 (c), Figure 6 (f), Figure 7 (c) and Figure 7 (f) show that as the number of resources on the server increases, the number and proportion of users allocated both increase, but the increasing amplitude decreases. The OARL-EUA algorithm has more obvious advantages when the number of resources is small.
[0070] Figure 8 Shows the user allocation time of the EUA algorithm. The experimental results show that when the EUA scenario becomes larger, the time consumption of the ILP-EUA algorithm increases rapidly, especially in Figure 8 (b) and Figure 8 (d) shows an exponential growth. This verifies that although the ILP-EUA algorithm can find the optimal solution, it is not efficient in solving practical large-scale EUA problems and is difficult to apply. In contrast, the running times of other methods such as FOA-EUA, GA-EUA, and ADRL-EUA algorithms are all within 1 second, with no significant difference, and are much faster than the ILP-EUA algorithm. The proposed OARL-EUA algorithm of the present invention shows higher efficiency and stability in practical applications, especially when the number of users is large and the server resources are limited, and the allocation effect is significantly better than other non-ILP-EUA algorithms.
[0071] In summary, the OARL-EUA algorithm of the present invention shows excellent performance in terms of the number of users allocated, the allocation proportion, and the allocation time. Especially when the resources are limited and the number of users is large, it is significantly better than the existing non-ILP-EUA algorithms, reflecting the innovation and practicality of the present invention in improving the allocation efficiency and effect. Brief Description of the Drawings
[0072] Figure 1 Is a flowchart of an edge user allocation method based on sequential perception reinforcement learning of the present invention.
[0073] Figure 2 Is a schematic diagram of reinforcement learning of the present invention.
[0074] Figure 3 Is a schematic diagram of an encoder-decoder of the present invention.
[0075] Figure 4 Is a structural diagram of a policy network of the present invention.
[0076] Figure 5 It is the value network structure diagram of the present invention.
[0077] Figure 6 It is the comparison result diagram of the change in the number of users allocated.
[0078] Figure 7 It is the comparison result diagram of the change in the proportion of users allocated.
[0079] Figure 8 It is the comparison diagram of the algorithm allocation time.
[0080] Reference numerals: Point - pointer module, Attention - attention module, LSTM - long short - term memory unit, FC1, FC2 - fully connected layers. Detailed implementation manners
[0081] In order to have a clearer understanding of the technical features, objectives, and effects of the present invention, the specific implementation manners of the present invention are now described. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention, that is, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0082] Embodiment 1
[0083] Numerous studies have shown that the problem of allocating edge users to maximize service coverage is a combinatorial optimization problem that is NP (Nondeterministic Polynomial time) - hard. The traditional methods for solving such NP - hard problems are mainly divided into three categories: exact algorithms, approximation algorithms, and heuristic algorithms. However, these three methods rarely utilize a common feature of practical optimization problems: problem instances of the same type are often repeatedly solved, and these problems maintain the same combinatorial structure but differ in data. That is to say, in many applications, the coefficient values in the objective function or constraints can be regarded as sampled from the same underlying distribution. Although there are inherent similarities between problem instances that appear in the same field, traditional algorithms do not systematically utilize this fact. In an industrial environment, if such a process can speed up real - time decision - making and improve quality, companies are willing to invest in upfront offline computing and learning.
[0084] Therefore, this embodiment considers using machine learning techniques to solve the edge user allocation problem in combinatorial optimization problems. Most successful machine learning techniques belong to supervised learning, which learns the mapping from input to output. However, since the optimal labels cannot be obtained for the edge user allocation problem, supervised learning is not applicable to the edge user allocation problem. However, a validator can be used to compare the quality of a set of solutions and provide some reward feedback to the learning algorithm. This embodiment considers using the reinforcement learning paradigm to solve the edge user allocation problem of maximizing service coverage. Based on the advantages of the Actor-Critic (policy network - value function network) method, the Actor-Critic framework in reinforcement learning is specifically adopted.
[0085] However, the existing reinforcement learning policy networks for solving the edge user allocation problem directly output the user allocation results without considering the capacity constraints of edge servers. The solutions provided by the policy network always exceed the capacity limit of the edge server, resulting in an infeasible allocation scheme, thus leading to a lower reward and hindering the effective learning of the policy network parameters. In addition, these methods are only applicable to the user allocation scenario with a fixed number of users. In the face of changes in the number of users, the network needs to be retrained, which is both time-consuming and inefficient.
[0086] To overcome the above problems, this embodiment designs and improves the policy network based on the Actor-Critic framework: (1) changes the policy of directly predicting the allocation result, instead predicting the user allocation order, and by adopting a greedy algorithm, users are allocated to the edge server with the most abundant resources one by one according to the predicted user allocation order, thus obtaining the actual allocation scheme; (2) adds a pointer mechanism to the order prediction so that the model can effectively point to a specific position in the input sequence, rather than predicting the index value from a fixed-size vocabulary.
[0087] Based on this, this embodiment provides an edge user allocation method based on sequential-aware reinforcement learning, which can effectively and quickly obtain the allocation result, that is, the user-server binding scheme, even when the environment changes. Specifically, the method of this embodiment includes:
[0088] Construct feature data for reinforcement learning based on edge user and server information, and perform user allocation based on the feature data, and optimize through the policy network and value network;
[0089] Convert the user allocation order output by the policy network into a detailed allocation scheme through a greedy strategy, and calculate the reward; determine the core components of reinforcement learning and conduct corresponding evaluations;
[0090] Optimize the policy network and value network, initialize the network parameters, and improve the network performance through training updates, thereby realizing edge user allocation.
[0091] As Figure 1 shown, the method of this embodiment predicts the allocation order of application users based on the Actor-Critic framework. After determining the allocation priority of users, users are sequentially allocated to the servers with the most abundant nearby resources according to the greedy strategy, so as to ensure that the allocation scheme is always feasible. In the face of the limitations of traditional sequence prediction models, that is, only being able to select index values from a fixed-size vocabulary, the method of this embodiment innovatively solves the challenge of the dynamically changing number of users by adopting a pointer mechanism. This mechanism endows the model with the ability to point to any position in the input sequence, rather than being limited to the index prediction of a fixed vocabulary. In order to further improve the performance of the model when processing sequence data, the method of this embodiment integrates an attention mechanism into the construction of its policy network and value network. This mechanism effectively captures the dependencies and key information within the sequence by dynamically weighting different parts of the input sequence, thus significantly improving the decision-making quality of the algorithm.
[0092] Preferably, the method of this embodiment can be implemented by the following steps:
[0093] Step 1: By processing the information of edge users and servers, construct the feature data that can be utilized by reinforcement learning, as Figure 1 shown in the upper part;
[0094] Step 2: Utilize the constructed feature data for user allocation and optimize it through the policy network and value network;
[0095] Step 3: Convert the user allocation order output by the policy network into a specific allocation scheme through the greedy strategy and calculate the reward;
[0096] Step 4: Define the core components of reinforcement learning: state, action, and reward, and conduct corresponding evaluations, as Figure 2 shown;
[0097] Step 5: Design and optimize the policy network and value network for user allocation, the policy network;
[0098] Step 6: Initialize the algorithm parameters and update and improve the model performance through training.
[0099] More preferably, Step 1 can be specifically implemented by the following sub-steps:
[0100] Step 1-1: Process the information of edge users and servers and convert it into the feature data that can be utilized by reinforcement learning. Use a graph neural network to process the adjacency relationship between users and servers, so as to capture more complex dependencies. The graph neural network can effectively represent the nodes and edges in the graph, and generate richer and more expressive node representations by propagating and aggregating the feature information of nodes and edges.
[0101] Step 1-2: Use the methods of vector normalization and concatenation to fuse the information from different data sources, including the user demand vector, the server resource vector, and the adjacency relationship between the edge user and the edge server. Further enrich the user feature representation through the multi-head attention mechanism. The multi-head attention mechanism can capture the dependencies between different parts of the input sequence, thereby generating a more expressive feature representation.
[0102] Step 1-3: Introduce additional features, such as historical load, user behavior patterns, etc. The historical load represents the resource usage of the server over a period of time in the past, including CPU usage rate, memory usage rate, network bandwidth usage rate, etc. By regularly recording the server resource usage, time series data is generated. The user behavior pattern feature describes the usage behavior and demand changes of users over a period of time in the past, including request frequency, data transfer volume, average session duration, etc. By monitoring the usage of users, information such as the request timestamp, data transfer volume, and session duration of users is recorded to generate user behavior logs. Parse the user behavior logs to extract the behavior patterns within a specific time window. Calculate statistical features such as request frequency, average data transfer volume, average session duration, etc., and represent the user behavior patterns in the form of time series.
[0103] More preferably, Step 2 can be specifically implemented by the following sub-steps:
[0104] Step 2-1: Use the user feature representation generated by feature construction as the state input and pass it to the reinforcement learning agent. Combine the environmental state and external factors (such as network bandwidth fluctuations, server health status, user location, etc.) to provide a more comprehensive state input. The network bandwidth fluctuation refers to the real-time usage of the network bandwidth, including the availability of the current bandwidth and the historical bandwidth usage record. The server health status includes key indicators such as CPU utilization rate, memory utilization rate, disk I / O, temperature, etc. The user location is the geographical location or relative location of the user, which is particularly important for edge computing because the geographical location will affect the latency and stability of data transmission.
[0105] Step 2-2: Output the user allocation order through the policy network and output the confidence score. Optimize the allocation order in combination with the confidence score. The confidence score reflects the certainty of the policy network for each user allocation decision. Extract the maximum probability value from the probability distribution of each user as the confidence score for the user allocation decision. When determining the allocation order, give priority to users with higher confidence scores, that is, users with higher certainty for the allocation decision are allocated first.
[0106] Step 2-3: Evaluate the quality of the strategy using the value network, introduce multi-objective optimization, and consider various factors such as resource utilization rate and user satisfaction when evaluating the strategy. The weighted sum method is adopted to assign different weights to different evaluation indicators and calculate the comprehensive evaluation value.
[0107] More preferably, Step 3 can be specifically implemented by the following sub-steps:
[0108] Step 3-1: Use the greedy algorithm to convert the user order output by the policy network into a specific allocation plan, and combine it with A* search to further optimize the allocation result on the basis of the greedy strategy. The A* search first defines a heuristic function for each user-server allocation, usually the difference between the remaining resource amount and the user demand. Use the initial allocation plan generated by the greedy algorithm as the starting point of the A* search, and then start from the initial node, expand all possible allocation plans of the current node, calculate the heuristic function value of each plan, and select the node with the smallest heuristic function value for expansion. Continue this process until the globally optimal allocation plan is found or the search depth limit is reached.
[0109] Step 3-2: After generating the allocation plan, calculate the reward based on the number of users allocated in the plan, design a piecewise reward function, and provide fine-grained feedback. The piecewise reward first divides the possible number of user allocations into several intervals, then defines the corresponding reward value for each interval, and determines the interval where the actual number of user allocations is located and gives the corresponding reward value.
[0110] More preferably, Step 4 can be specifically implemented by the following sub-steps:
[0111] The state is represented by the user information obtained through feature construction and is extended to include time series information to capture dynamic changes. The state is a comprehensive description of the environment by the agent at a certain moment. In the present invention, the state is represented by the user information obtained through feature construction, including the user demand vector, the server resource vector, and the adjacency relationship between the user and the server. After being processed by the graph neural network, these information form high-dimensional feature vectors. In addition, the state is also extended to include time series information to capture dynamic changes, such as the user behavior pattern and the historical load situation of the server resources.
[0112] Step 4-2: The action consists of two parts: an indirect action and a direct action. The policy network is responsible for generating the indirect action, i.e., the allocation order of edge users, and introduces a multi-level action space, such as primary allocation and fine-tuning allocation, to enhance the flexibility of the action. The action is the behavior taken by the agent in a certain state. In the present invention, the action consists of two parts: an indirect action and a direct action. The indirect action is the user allocation order generated by the policy network. The agent processes the current state through the policy network to generate a list of user allocation orders. This order list indicates the order of users to be preferentially allocated at the current moment. The direct action is the specific user-server allocation decision generated based on the indirect action. In the present invention, a multi-level action space is introduced to enhance the flexibility of the action. The primary allocation in the multi-level action space is the preliminary allocation of the user allocation order, which allocates users to the servers with the most abundant resources. The fine-tuning allocation is to further optimize and adjust the allocation scheme based on the primary allocation.
[0113] Step 4-3: The reward is used to evaluate the quality of the action, and thus guide the training of the agent. A multi-level reward function is designed, considering short-term and long-term benefits, and a penalty mechanism is introduced to avoid resource waste. The reward is a feedback signal used to evaluate the quality of the agent's action. In the present invention, the reward is designed as a multi-level reward function, considering short-term and long-term benefits, and a penalty mechanism is introduced to avoid resource waste. The reward function includes multiple levels to evaluate the benefits on different time scales. In the present invention, the reward function mainly includes short-term reward and long-term reward. The short-term reward is based on the resource utilization rate and user satisfaction of the current allocation scheme, while the long-term reward considers the long-term behavior patterns of users and the sustainable use of resources.
[0114] More preferably, Step 5 can be specifically implemented by the following sub-steps:
[0115] Step 5-1: The policy network adopts an Encoder-Decoder architecture as shown in Figure 3 and no longer directly predicts the allocation scheme, but predicts an allocation order of application users.
[0116] Step 5-2: The policy network consists of two recurrent neural network modules, an encoder and a decoder, both of which adopt long short-term memory units (LSTM), as shown in Figure 4 . The encoder network inputs user features one by one and converts them into a sequence of latent memory states. A bidirectional LSTM is used to better capture global information. The decoder also maintains its sequence of latent states, captures relevant information within the sequence through an attention mechanism, and generates a probability distribution of the next user number to be allocated. Combining with the Copy mechanism, the adaptability of unknown user allocation is improved.
[0117] Step 5-4: The value network consists of two neural network modules: an LSTM encoder and a decoder, as Figure 5 shown. It is parameterized by the parameter θ v to learn the number of users expected to be served by the current policy p θ given the input φ. The value network is trained using the stochastic gradient descent method, and its objective is to be trained on the mean squared error between its predicted value and the number of user allocations obtained by the most recent policy, i.e.:
[0118]
[0119] More preferably, Step 6 can be specifically implemented by the following sub-steps:
[0120] Step 6-1: Initialize the policy network parameters and the value network parameters. Use a pre-trained model to initialize the policy network and the value network parameters to accelerate the convergence speed. The pre-training refers to pre-training the policy network and the value network on a publicly available dataset in the relevant field so that their initial parameters are closer to the final optimal parameters, thereby accelerating the convergence.
[0121] Step 6-2: Sample a batch of users from the dataset as the input, and use an intelligent sampling strategy to improve the sampling efficiency. The intelligent sampling strategy refers to selecting representative data samples from the dataset through an optimized sampling method to improve the training efficiency and the generalization ability of the model. In the present invention, the importance sampling method is used for selection according to the representativeness of the samples.
[0122] Step 6-3: Perform the forward propagation of the policy network and the value network to obtain the allocation priorities of the users and the expected number of served users, and introduce a Dropout layer to prevent overfitting. The Dropout layer randomly discards a part of the neurons during the training process, so that the network does not overly rely on certain specific paths, thereby improving the generalization ability of the model.
[0123] Step 6-4: Calculate the user allocation scheme based on the output of the policy network and the greedy policy, and use parallel computing to accelerate the execution of the greedy policy.
[0124] Step 6-5: Update the parameters of the policy network and the value network by the gradient descent method, and use self-supervised learning techniques to enhance the generalization ability of the value network.
[0125] Specifically, in this embodiment, a real-world dataset, EUA Datasets, is used for experiments. This dataset has been widely used in the study of EUA problems. The dataset contains a large number of real mobile users and base stations in the metropolitan area of Melbourne, Australia. The area of this region is 6.2 square kilometers. There are 125 base stations in this region, corresponding to 125 edge servers. Each edge server has a specific coverage range, and its value is randomly set according to the edge server density. Generally speaking, the coverage radius of the edge server is between [200, 400] (in meters).
[0126] In each experiment, m base stations are randomly selected from the dataset as the locations of m candidate edge servers in the experiment, and then the resources of each edge server are randomly generated according to the normal distribution. For each type of resource, the average quantity on one edge server is c, and the variance is 1, that is, the resource distribution follows Meanwhile, n mobile users are randomly selected from the dataset, and their locations are used as the locations of edge users in the experiment. In this embodiment, the values of three parameters are varied to simulate different scenarios, and the performance of OARL-EUA is widely evaluated, including: the number of edge users (n); the number of candidate edge servers (m); the average resource capacity of edge servers (c). Two sets of data are taken for each parameter, and there are a total of six datasets, from Dataset 1.1 to Dataset 2.3. One parameter changes for each experimental dataset, as shown in Table 1. On these datasets, this experiment mainly compares the number of user allocations, the user allocation ratio, and the user allocation time.
[0127] Table 1 - Experimental Datasets
[0128] Dataset n m c Dataset 1.1 70,80,90,100,110,120 20 4 Dataset 1.2 100 5,10,15,20,25,30 4 Dataset 1.3 100 20 1,2,3,4,5,6 Dataset 2.1 300,400,500,600,700,800 80 5 Dataset 2.2 500 60,70,80,90,100,110 5 Dataset 2.3 500 80 2,3,4,5,6,7
[0129] (1) Experimental Setup
[0130] The hyperparameter settings of this embodiment are shown in Table 2.
[0131] Table 2 - Hyperparameter Settings
[0132] Hyperparameter Value Batch size batch_size 16 Initial learning rate learning_rate 1e-3 Hidden layer dimension hidden_size 128 Embedding layer dimension embed_size 128 Learning rate decay rate learning_rate_decay 0.9 Learning rate decay steps learning_rate_decay_step 3000 Total training steps total_timesteps 50000
[0133] (2) Comparative Algorithms
[0134] Edge User Allocation based on Fruit Fly Optimization Algorithm (FOA-EUA) (LI T, NIU W, JI C. Edge user allocation by FOA in edge computing environment[J]. J. Comput. Sci., 2021, 53: 101390.).
[0135] Greedy Algorithm-based Edge User Allocation (GA-EUA) algorithm (LAIP, HE Q, ABDELRAZEK M, et al. Optimal edge user allocation in edge computing with variable sized vector bin packing[C] / / Service-Oriented Computing: 16th International Conference, ICSOC 2018, Hangzhou, China, November 12 - 15, 2018, Proceedings 16. Springer, 2018: 230 - 245.).
[0136] Integer Linear Programming-based Edge User Allocation (ILP-EUA) algorithm (LAIP, HE Q, ABDELRAZEK M, et al. Optimal edge user allocation in edge computing with variable sized vector bin packing[C] / / Service-Oriented Computing: 16th International Conference, ICSOC 2018, Hangzhou, China, November 12 - 15, 2018, Proceedings 16. Springer, 2018: 230 - 245.).
[0137] Attention-based Deep Reinforcement Learning-based Edge User Allocation (ADRL-EUA) algorithm (CHANG J, WANG J, LI B, et al. Attention-based deep reinforcement learning for edge user allocation[J]. IEEE Trans. Netw. Serv. Manage., 2024, 21(1): 590 - 604.).
[0138] (3) Results
[0139] ① The effectiveness results are as follows (the higher the better):
[0140] In the EUA problem of maximizing service coverage, the effectiveness of the algorithm is mainly reflected by the number of users allocated or the user allocation ratio, as shown in Figure 6 、 Figure 7 . From Figure 6 and Figure 7It can be seen that, regardless of the scale of the dataset on which the experiment is conducted, the ILP-EUA algorithm has the best performance, followed by the OARL-EUA algorithm proposed in this embodiment. However, the ILP-EUA algorithm takes an extremely long time and is generally regarded as the theoretically optimal solution, but it is not actually adopted in practice.
[0141] Figure 6 (a), Figure 6 (d), Figure 7 (a) and Figure 7 (d) show that, with the number of servers and the amount of resources on the servers fixed, as the number of users increases, the overall number of users allocated increases continuously, while the user allocation ratio decreases continuously. This is because, with limited resources, the number of users who can be allocated cannot increase indefinitely. In addition, the OARL-EUA algorithm shows more obvious advantages when the number of users is larger and is superior to other non-ILP-EUA algorithms.
[0142] Figure 6 (b), Figure 6 (e), Figure 7 (b) and Figure 7 (e) show that, with the number of edge users and the amount of resources on the servers fixed, as the number of servers increases, the number of users who can be allocated increases continuously, and the user allocation ratio also increases continuously, but the increasing amplitude becomes smaller and smaller. The OARL-EUA algorithm shows more obvious advantages when the number of servers is small. This is because when the number of servers is too large, basically most users can receive services, and the user allocation order learned by the OARL-EUA algorithm is not so important, so the improvement effect is not obvious.
[0143] Figure 6 (c), Figure 6 (f), Figure 7 (c) and Figure 7 (f) show that, as the amount of resources on the servers increases, the number of users who can be allocated increases continuously, and the ratio of users who can be allocated also increases continuously, but the increasing amplitude becomes smaller and smaller. Moreover, the OARL-EUA algorithm shows more obvious advantages when the amount of resources is small. All in all, the stronger the user competition relationship and the scarcer the resources, the more obvious the advantages of the OARL-EUA algorithm.
[0144] ② The results of efficiency are as follows:
[0145] In this embodiment, the efficiency of the algorithms is mainly compared by the allocation time. Figure 8 Shows the user allocation time of the EUA algorithm. From Figure 8 (a), Figure 8 (b), Figure 8 (d) and Figure 8(e) It can be seen that when the EUA scenario becomes larger, the time consumption of the ILP-EUA algorithm increases rapidly, Figure 8 (b) and Figure 8 (d) show an exponential growth rate. This confirms the previous conclusion that the ILP-EUA algorithm cannot find the optimal solution for large-scale EUA problems. In other words, although ILP-EUA can always find the optimal solution, it is not efficient enough for solving practical EUA problems. This confirms that EUA is an NP-hard problem, and it is difficult to apply ILP-EUA to find the optimal solution for large-scale EUA problems. On the contrary, as Figure 8 shown, all other methods are very effective. Although the running times of the FOA-EUA, GA-EUA, and ADRL-EUA algorithms are shorter than that of the OARL-EUA algorithm in most cases, the running times of these four algorithms do not exceed 1 second. For application service providers, these differences are not significant, and they are all much faster than the optimal solution algorithm ILP-EUA.
[0146] In summary, the OARL-EUA algorithm of this embodiment shows superior performance in terms of the number of user allocations, allocation ratios, and allocation times. Especially in the case of limited resources and a large number of users, it is significantly better than the existing non-ILP-EUA algorithms, reflecting the innovation and practicality of the OARL-EUA algorithm of this embodiment in improving allocation efficiency and effectiveness.
[0147] Embodiment 2
[0148] Based on Embodiment 1, this embodiment:
[0149] This embodiment provides an edge user allocation system based on sequential perception reinforcement learning, including:
[0150] A first processing module, configured to construct feature data for reinforcement learning based on edge user and server information, and perform user allocation based on the feature data, and optimize through a policy network and a value network;
[0151] A second processing module, configured to convert the user allocation order output by the policy network into a detailed allocation plan through a greedy policy, and calculate the reward; determine the core components of reinforcement learning, and perform corresponding evaluations;
[0152] A third processing module, configured to optimize the policy network and the value network, initialize the network parameters, and improve the network performance through training updates, thereby realizing edge user allocation.
[0153] Embodiment 3
[0154] Based on Embodiment 1, this embodiment:
[0155] This embodiment provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the edge user allocation method based on sequential perception reinforcement learning in Embodiment 1. Among them, the computer program can be in the form of source code, object code, executable file, or some intermediate form, etc.
[0156] Embodiment 4
[0157] On the basis of Embodiment 1, this embodiment:
[0158] This embodiment provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements the edge user allocation method based on sequential perception reinforcement learning in Embodiment 1. Among them, the computer program can be in the form of source code, object code, executable file, or some intermediate form, etc. The storage medium includes: any entity or device capable of carrying computer program code, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the storage medium does not include electrical carrier signals and telecommunication signals.
[0159] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application. In addition, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
Claims
1. A method for allocating edge users based on sequential perceptual reinforcement learning, characterized in that: include: Construct feature data for reinforcement learning based on edge user and server information, assign users based on feature data, and optimize through policy network and value network; The user allocation sequence output by the policy network is converted into a detailed allocation plan through a greedy strategy, and the rewards are calculated; Identify the core components of reinforcement learning and evaluate them accordingly; Optimize the policy network and value network, initialize network parameters, and improve network performance through training updates to achieve edge user allocation.
2. According to the method for allocating edge users based on sequential perception reinforcement learning in claim 1, it is characterized in that: The feature data for reinforcement learning is constructed based on edge user and server information, including: Process edge user and server information and convert it into feature data that can be used by reinforcement learning; use graph neural networks to process the adjacency relationship between users and servers to capture more complex dependencies; The method of vector normalization and concatenation is used to fuse information from different data sources, including user demand vectors, server resource vectors, and the adjacency relationship between edge users and edge servers. The multi-head attention mechanism is used to further enrich the user feature representation. Additional features are introduced, including historical load features and user behavior pattern features. The historical load features represent the resource usage of the server in the past period of time, and the user behavior pattern features represent the usage behavior and demand changes of the users in the past period of time.
3. The edge user allocation method based on sequential perception reinforcement learning according to claim 2 is characterized in that: The historical load characteristics include CPU usage, memory usage and network bandwidth usage; the calculation method of the historical load characteristics includes: by regularly recording the resource usage of the server, expressing the historical load characteristics in the form of a time series; The user behavior pattern characteristics include request frequency, data transmission volume and average session duration; the calculation method of the user behavior pattern characteristics includes: monitoring the user's usage, recording the user's request timestamp, data transmission volume and session duration, and generating a user behavior log; parsing the user behavior log to extract the behavior pattern within a specific time window; calculating statistical characteristics including request frequency, average data transmission volume and average session duration, and representing the user behavior pattern in the form of a time series.
4. The edge user allocation method based on sequential perception reinforcement learning according to claim 1 is characterized in that: The user allocation based on feature data is optimized through a strategy network and a value network, including: Passing the constructed feature data as state input to the reinforcement learning agent, and providing a more comprehensive state input in combination with the environment state and external factors, including network bandwidth fluctuations, server health status, and user location; The user allocation order is output through the policy network, and the confidence score is output, and the allocation order is optimized in combination with the confidence score; the confidence score represents the certainty of the policy network on each user allocation decision, and the maximum probability value is extracted from the probability distribution of each user as the confidence score of the user allocation decision; The quality of the strategy is evaluated through the value network, multi-objective optimization is introduced, and multiple factors are considered when evaluating the strategy; the weighted sum method is used to assign different weights to different evaluation indicators and calculate the comprehensive evaluation value.
5. The edge user allocation method based on sequential perception reinforcement learning according to claim 1 is characterized in that: The user allocation sequence output by the policy network is converted into a detailed allocation plan through a greedy strategy, and the rewards are calculated, including: The user sequence output by the policy network is converted into a detailed allocation plan using a greedy algorithm, and the allocation result is further optimized based on the greedy strategy by combining the A* search algorithm. After the allocation plan is generated, the reward is calculated based on the number of users allocated in the allocation plan, and a segmented reward function is designed to provide fine-grained feedback; the segmented reward function includes: dividing the possible number of user allocations into several intervals, defining a corresponding reward value for each interval, determining the interval in which the user is located according to the actual number of user allocations and giving a corresponding reward value.
6. The edge user allocation method based on sequential perception reinforcement learning according to claim 1 is characterized in that: The core components of reinforcement learning are identified and evaluated accordingly, including: Determining a state of reinforcement learning, the state being represented by the constructed feature data and extended to include time series information to capture dynamic changes; Determine the action of reinforcement learning, the action includes indirect action and direct action, the indirect action includes the user allocation sequence generated by the policy network, and introduces a multi-level action space; the action includes the behavior taken by the intelligent agent in a certain state; Determine the rewards for reinforcement learning, which are used to evaluate the quality of actions and guide the training of the agent. Design a multi-level reward function, consider short-term and long-term benefits, and introduce a penalty mechanism to avoid wasting resources.
7. The edge user allocation method based on sequential perception reinforcement learning according to claim 1 is characterized in that: The optimization strategy network and value network include: The optimized policy network includes a first encoder and a second decoder based on a long short-term memory network. The first encoder inputs user features one by one and converts them into a potential memory state sequence. The second decoder maintains the potential memory state sequence, captures relevant information in the potential memory state sequence through an attention mechanism, and generates a probability distribution of the next user number to be assigned, and combines the replication mechanism to improve the adaptability of unknown user assignment. The optimized value network includes a second encoder and a second decoder, wherein the second encoder is based on a long short-term memory network; the value network learns the number of users expected to be served by the current strategy when a given input is parameterized, and is trained by a stochastic gradient descent method, and is trained on the mean square error between the predicted value and the number of user allocations obtained by the most recent strategy.
8. The edge user allocation method based on sequential perception reinforcement learning according to claim 1 is characterized in that: The initialization of network parameters includes initializing policy network and value network parameters through pre-training, and the pre-training includes pre-training the policy network and the value network on public data sets in related fields to make the initial parameters closer to the final optimal parameters, thereby accelerating convergence.
9. The edge user allocation method based on sequential perception reinforcement learning according to claim 1 is characterized in that: Improving network performance through training updates includes: Sampling a batch of users from the data set as input, and using an intelligent sampling strategy to improve sampling efficiency; the intelligent sampling strategy includes selecting representative data samples from the data set through an optimized sampling method; Perform forward propagation of the policy network and the value network to obtain the user's allocation priority and the expected number of service users, and introduce a dropout layer to prevent overfitting; Calculate the user allocation plan based on the policy network output and the greedy strategy, and use parallel computing to accelerate the execution of the greedy strategy; The parameters of the policy network and the value network are updated by the gradient descent method, and the generalization ability of the value network is enhanced using a self-supervised learning method.
10. An edge user allocation system based on sequential perception reinforcement learning, characterized in that: include: A first processing module is configured to construct feature data for reinforcement learning based on edge user and server information, and to perform user allocation based on the feature data, and to perform optimization through a policy network and a value network; The second processing module is configured to convert the user allocation sequence output by the policy network into a detailed allocation plan through a greedy strategy and calculate the reward; Identify the core components of reinforcement learning and evaluate them accordingly; The third processing module is configured to optimize the policy network and the value network, initialize network parameters, and improve network performance through training updates, thereby achieving edge user allocation.