A 5G network resource allocation method, system and device
By adjusting the resource allocation strategy through unsupervised feature extraction and reinforcement learning, the problem of user experience evaluation bias in 5G networks is solved, and the user experience is improved and resources are efficiently and dynamically allocated.
Patent Information
- Application Number
- CN202511061486.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-07-31
AI Technical Summary
Existing 5G network resource allocation technology cannot accurately capture millisecond-level fluctuations when users are engaged in services with high real-time requirements, such as short videos or games. This causes a deviation between the detection data on the network side and the user's actual experience, affecting the user experience.
By obtaining users' traffic data, unsupervised physical layer-application layer joint feature extraction is performed to predict network usage patterns. The resource allocation strategy is adjusted by combining the reinforcement learning strategy network to achieve dynamic resource allocation.
It improves the user's network experience, meets the user's personalized needs, and avoids resource waste while ensuring the overall efficiency of the network.
Smart Images

Figure CN120568501B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network resource allocation, and in particular to a 5G network resource allocation method, system, and device. Background Art
[0002] Currently, operators generally adopt a slice-based multi-service bearer approach for 5G networks. This approach uses a centralized controller to collect network-level data such as wireless channel quality, base station buffer occupancy, and spectrum utilization. This is then combined with a linear programming algorithm to allocate fixed resource quotas to each slice type. In recent years, with the widespread adoption of artificial intelligence (AI), deep reinforcement learning models are often used to directly output resource quotas to improve overall network efficiency and carrying capacity.
[0003] Related technologies typically quantify user-side perception data, such as business-layer MOS (Management Operation System) and lag rate, into a single score, which is then mapped to a weighting factor through a table lookup, and then fine-tuned to resource quotas. However, due to the slow sampling and limited dimensionality of user perception data, it is unable to detect millisecond-level fluctuations in latency and bitrates, such as those experienced when watching short videos or playing games. This leads to a discrepancy between the network-side detection data and the user's actual experience. Consequently, in pursuit of overall network efficiency, some users' network resources are sacrificed, resulting in a poor user experience. Summary of the Invention
[0004] The problem solved by the present invention is how to improve the network usage experience of users.
[0005] To solve the above problems, the present invention provides a 5G network resource allocation method, system and device.
[0006] In a first aspect, the present invention provides a 5G network resource allocation method, comprising:
[0007] Obtain the traffic data of each user in the 5G network during the current time period and the user's current resource allocation strategy;
[0008] Performing unsupervised physical layer-application layer joint feature extraction on the traffic data to obtain traffic characteristics of each user in the current time period;
[0009] Predicting the network behavior of the user in the current time period based on the traffic characteristics in the current time period to obtain a network usage pattern of the user;
[0010] Determining the user's experience score based on the network usage pattern and the current resource allocation strategy;
[0011] Determining whether it is necessary to adjust the current resource allocation strategy of the user according to the perception score;
[0012] When the current resource allocation policy of the user needs to be adjusted, the resource adjustment amount for the user in the next time period is obtained through a reinforcement learning strategy network based on the user's network usage pattern, the user's remaining bandwidth, and the current resource allocation policy;
[0013] A future resource allocation strategy for the user in the next time period is determined according to the resource adjustment amount, and resources are allocated to the user according to the future resource allocation strategy.
[0014] Optionally, obtaining traffic data of each user in the 5G network in a current time period and the user's current resource allocation strategy includes:
[0015] Continuously collecting the I / Q sampling stream, control layer scheduling log, protocol data unit timestamp, message arrival interval, and instantaneous throughput value of each user according to a preset granularity, and forming raw traffic data based on the I / Q sampling stream, the control layer scheduling log, the protocol data unit timestamp, the message arrival interval, and the instantaneous throughput value;
[0016] Performing timestamp alignment and slice-level indexing on the original traffic data to obtain the traffic data of the user in the current time period;
[0017] Obtaining, through the slice scheduler, the number of physical resource blocks, the remaining bandwidth, and the maximum schedulable modulation and coding strategy level of the network slice where the user is located in the current time period;
[0018] The current resource allocation strategy is determined according to the number of physical resource blocks, the remaining bandwidth, the maximum schedulable modulation and the coding strategy level.
[0019] Optionally, performing unsupervised physical layer-application layer joint feature extraction on the traffic data to obtain traffic features of each user in the current time period includes:
[0020] Dividing the traffic data into multiple continuous subsequences according to time windows;
[0021] Performing short-time Fourier transform on the I / Q sampling stream in each subsequence to obtain an in-phase orthogonal time-frequency diagram;
[0022] Time-aligning and concatenating the control layer scheduling log, the protocol data unit timestamp, the message arrival interval, and the instantaneous throughput value in the subsequence to obtain a cross-layer feature vector;
[0023] Inputting the in-phase orthogonal time-frequency map and the cross-layer feature vector into an unsupervised convolutional hybrid encoder to obtain an embedding vector of fixed dimension;
[0024] Time series pooling and normalization are performed on the embedding vectors corresponding to all the subsequences in the current time period to obtain a parameter-free traffic fingerprint of the user, and the parameter-free traffic fingerprint is used as the traffic feature.
[0025] Optionally, predicting the network behavior of the user in the current time period based on the traffic characteristics in the current time period to obtain the network usage pattern of the user includes:
[0026] Inputting the parameter-free traffic fingerprint into a network behavior prediction model;
[0027] Outputting a probability distribution of the network usage pattern of the user in the current time period through the network behavior prediction model;
[0028] Screening is performed according to the probability distribution of the network usage pattern, and the pattern with the highest probability value in the probability distribution is used as the network usage pattern of the user.
[0029] Optionally, determining the user's perception score based on the network usage pattern in combination with the current resource allocation strategy includes:
[0030] According to the network usage pattern, obtaining a network experience mapping table corresponding to the network usage pattern;
[0031] Determining a standard resource allocation strategy corresponding to the network usage pattern according to the network experience mapping table;
[0032] Mapping the number of physical resource blocks, the remaining bandwidth, the maximum schedulable modulation, and the coding strategy level in the standard resource allocation strategy as benchmark parameters;
[0033] Comparing the number of physical resource blocks, the remaining bandwidth, the maximum schedulable modulation, and the coding strategy level in the current resource allocation strategy with the benchmark parameters one by one, to obtain deviations corresponding to the number of physical resource blocks, the remaining bandwidth, the maximum schedulable modulation, and the coding strategy level, respectively;
[0034] Obtaining, by means of a quantization function of the network experience mapping table and according to the deviation, local perception scores corresponding to the number of physical resource blocks, the remaining bandwidth, the maximum schedulable modulation, and the coding strategy level, respectively;
[0035] The local perception scores are weighted and summed to obtain the perception score of the user.
[0036] Optionally, judging whether it is necessary to adjust the current resource allocation strategy of the user according to the perception score includes:
[0037] By comparing the perception score with a preset perception threshold, determining whether the current resource allocation strategy of the user needs to be adjusted;
[0038] If the perception score is lower than the preset perception threshold, determining that the current resource allocation strategy of the user needs to be adjusted;
[0039] If the perception score is not lower than the preset perception threshold, it is determined that there is no need to adjust the current resource allocation strategy of the user.
[0040] Optionally, obtaining a resource adjustment amount for the user in the next time period through a reinforcement learning strategy network based on the network usage pattern of the user, in combination with the remaining bandwidth of the user and the current resource allocation strategy, includes:
[0041] Using the network usage pattern, the remaining bandwidth, and the current resource allocation strategy as state inputs of the reinforcement learning strategy network;
[0042] Generating a set of candidate resource adjustment actions in a preset action space according to the state input through the reinforcement learning strategy network;
[0043] The candidate resource adjustment action set includes a physical resource block control action, a bandwidth control action, and a modulation and coding strategy level control action;
[0044] A value evaluation is performed on each control action in the candidate resource adjustment action set according to the reward function, and a joint optimization is performed with the user's perception score and the total throughput of the 5G network system as optimization targets to obtain the resource adjustment amount of the user in the next time period.
[0045] Optionally, determining a future resource allocation strategy for the user in the next time period based on the resource adjustment amount, and allocating resources to the user according to the future resource allocation strategy includes:
[0046] The resource adjustment amount is converted into a physical resource block allocation value, a bandwidth allocation value, and a modulation and coding strategy level allocation value in the next time period through a resource scheduling engine;
[0047] generating the future resource allocation strategy for the user according to the physical resource block allocation value, the bandwidth allocation value, and the modulation and coding strategy level allocation value;
[0048] The future resource allocation strategy is sent to the base station scheduler of the network slice where the user is located, and the base station scheduler allocates wireless resources to the user according to the future resource allocation strategy within the next time period.
[0049] In a second aspect, the present invention provides a 5G network resource allocation system, comprising:
[0050] An acquisition unit, configured to acquire traffic data of each user in the 5G network in a current time period and the user's current resource allocation strategy;
[0051] a feature extraction unit, configured to perform unsupervised physical layer-application layer joint feature extraction on the traffic data to obtain traffic features of each user in the current time period;
[0052] a prediction unit, configured to predict the network behavior of the user in the current time period based on the traffic characteristics in the current time period, and obtain a network usage pattern of the user;
[0053] a scoring unit, configured to determine a user experience score based on a network usage pattern and in combination with the current resource allocation strategy;
[0054] a judging unit, configured to judge whether it is necessary to adjust the current resource allocation strategy of the user according to the perception score;
[0055] an adjustment decision unit, configured to, when it is necessary to adjust the current resource allocation policy of the user, determine, based on the network usage pattern of the user, the remaining bandwidth of the user, and the current resource allocation policy, a resource adjustment amount for the user in the next time period through a reinforcement learning strategy network;
[0056] The resource allocation unit is configured to determine a future resource allocation strategy for the user in the next time period according to the resource adjustment amount, and allocate resources to the user according to the future resource allocation strategy.
[0057] In a third aspect, the present invention provides an electronic device comprising a memory and a processor;
[0058] The memory is used to store computer programs;
[0059] The processor is used to implement the 5G network resource allocation method as described above when executing the computer program.
[0060] The 5G network resource allocation method, system and electronic device of the present application can fully mine the deep features of the traffic data by performing unsupervised physical layer-application layer joint feature extraction after obtaining the traffic data and resource allocation strategy of the user, providing a more abundant and accurate data basis for subsequent analysis, making up for the shortcomings of traditional networks relying only on network-level data (wireless channel quality, base station buffer occupancy, etc.), and solving the problem of user experience evaluation deviation caused by limited data dimension. Based on the extracted traffic features, the network behavior of the user is predicted to obtain the network usage mode. The mode is combined with the current resource allocation strategy to determine the user experience score. In this process, the prediction of the network usage mode converts the traffic features into specific usage scenarios, making the calculation of the experience score more accurate and directly linking data processing and user experience evaluation. Then, whether the resource allocation strategy needs to be adjusted is determined according to the experience score. If adjustment is needed, the resource adjustment amount is determined by the reinforcement learning strategy network by comprehensively considering the user network usage mode, the remaining bandwidth and the current resource allocation strategy. The reinforcement learning strategy network plays a key role in this process. It learns the optimal resource adjustment strategy from historical data and real-time feedback, takes into account the user's remaining bandwidth to reasonably allocate resources, avoids resource waste or over-allocation, and solves the problem of unreasonable resource allocation caused by slow user experience data sampling and low dimensionality in the prior art. The present application fully links the experience evaluation, resource adjustment decision and network behavior prediction, etc., to form a complete closed loop, and realizes efficient dynamic allocation of resources.
[0061] In summary, the present application can significantly improve the user's network usage experience. In the traditional 5G network resource allocation method, due to the slow sampling and low dimensionality of user experience data, it is difficult to accurately capture the millisecond-level fluctuations of users during the process of performing real-time services such as short videos or games, such as changes in delay and code rate. This causes a deviation between the detection data on the network side and the real experience of the user, and part of the users in the process of pursuing the overall efficiency of the network have their network resources allocated unreasonably, resulting in poor user experience. The present application can more comprehensively and accurately obtain the traffic features of the user by obtaining the traffic data and the current resource allocation strategy of each user in the current time period, and performing unsupervised physical layer-application layer joint feature extraction on the traffic data. Based on these features, the network usage mode of the user is predicted, and then the current resource allocation strategy is combined to determine the user experience score, so that the real network usage experience of the user can be reflected in a timely and accurate manner.
[0062] Furthermore, a reinforcement learning strategy network is used to determine the user's resource adjustment amount for the next time period based on the user's network usage pattern, remaining bandwidth, and current resource allocation strategy. This is then used to determine the future resource allocation strategy, achieving efficient dynamic resource allocation. Compared to previous slicing-based multi-service carrying methods and simple deep reinforcement model output resource quota methods, this application takes into account the dynamic changes in user experience data and network resource conditions, and can better meet the personalized needs of users while ensuring overall network efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 This is a flowchart of a 5G network resource allocation method according to an embodiment of the present invention;
[0064] Figure 2 This is a structural block diagram of a 5G network resource allocation system according to an embodiment of the present invention;
[0065] Figure 3 Schematic diagram of the structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0066] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, specific embodiments of the present invention are described in detail below with reference to the accompanying drawings. Although certain embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as being limited to the embodiments described herein. Instead, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.
[0067] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0068] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to"; the term "based on" means "based at least in part on"; the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments"; the term "optionally" means "optional embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc. mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0069] It should be noted that the modifications of "one" and "multiple" mentioned in the present invention are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0070] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0071] In related technologies, user experience data is typically sampled at a low frequency, at intervals of only seconds or longer. This user experience data includes latency, lag rate, buffering time, image quality fluctuations, and interactive response time. However, many network activities with high real-time requirements, such as short video browsing and online gaming, generate significant millisecond-level network fluctuations, including latency, lag, and bitrate variations. These instantaneous fluctuations have a significant impact on the user's actual experience. For example, in online gaming, even a delay of just a few hundred milliseconds can cause player errors and compromise the gaming experience. Similarly, when watching high-definition short videos, momentary lags and bitrate drops can directly impact video smoothness and quality. Due to the limitations of low-frequency sampling, the network side is unable to capture these critical millisecond-level fluctuations, resulting in a significant discrepancy between network detection data and actual user experience. This discrepancy causes network resource allocation strategies to often lag behind actual user demand, failing to promptly respond to users' real-time experience needs.
[0072] Existing technologies often simplify complex user experience data into a single score, such as MOS (Monitoring Optimization) or lag rate, and then use a table lookup to map this score into a weighting factor for fine-tuning resource quotas. This simplified approach has the following drawbacks: First, user experience data is multidimensional. In addition to lag rate and MOS score, it also includes buffering time, image quality fluctuations, and interactive response time. A single-dimensional score fails to fully reflect the user's true experience. For example, in a video call, even if the lag rate is low, frequent image quality switching or audio-video desynchronization can still severely impact call quality. Second, the table lookup mapping approach lacks flexibility and adaptability. The relationship between user experience data and resource requirements varies across users, service types, and network environments. Fixed tables cannot dynamically adjust to these changes, resulting in reduced resource allocation accuracy. This simplified approach makes it difficult for the network to accurately quantify and analyze user experience, which in turn affects the accuracy of resource allocation decisions.
[0073] Existing resource allocation technologies often focus on improving overall network efficiency, which may come at the expense of network resources for some users. When network loads are high, to ensure a basic experience for most users, the network may restrict resources for individual users, such as reducing bandwidth allocation or the number of resource blocks. While this approach ensures overall network stability to a certain extent, it exacerbates user experience inconsistencies. For example, in high-density user scenarios, such as large-scale sporting events or concerts, some users may experience significant network latency and lag, while others enjoy relatively smooth service. This inconsistency not only impacts individual user experiences but may also raise questions about the fairness of network services. Furthermore, when network resources are limited, user experience varies significantly across different service types. However, existing technologies struggle to dynamically adjust to these real-time changes in user experience, further exacerbating user experience inconsistencies.
[0074] With the widespread adoption of 5G networks, services are becoming increasingly diverse, including enhanced mobile broadband (eMBB), ultra-reliable low-latency communications (URLLC), and massive machine-type communications (mMTC). Each service type has distinct requirements for network resources and user sensitivity. For example, eMBB services have extremely high bandwidth requirements, while URLLC services are extremely sensitive to latency. Existing resource allocation technologies struggle to accurately match resources to these diverse services. A unified resource allocation strategy cannot meet the differentiated needs of different service types, resulting in poor user experience for some services. Furthermore, user behavior on 5G networks is more complex and dynamic. Users may frequently switch between service types within a short period of time, such as switching from video streaming to online gaming, or from web browsing to video calling. Existing technologies struggle to quickly respond to these changes and adjust resource allocation strategies in a timely manner, failing to meet users' expectations for high-quality network services in diverse service scenarios.
[0075] In summary, existing 5G network resource allocation technologies have significant shortcomings in the collection, quantification, and application of user experience data, making it difficult to meet the demands of real-time, refined resource allocation. These issues not only impact the user experience but also restrict the performance of 5G networks in diverse business scenarios.
[0076] In response to the problems existing in the above-mentioned related technologies, this embodiment provides a 5G network resource allocation method, system and device.
[0077] Combine Figure 1 As shown, an embodiment of the present invention provides a 5G network resource allocation method, including:
[0078] Obtain the traffic data of each user in the 5G network during the current time period and the user's current resource allocation strategy.
[0079] Specifically, traffic monitoring modules deployed at 5G base stations (gNBs, the next generation Node Bs) collect traffic data for each user in the 5G network during the current time period. These modules can capture user data packets in real time and analyze information such as their size and transmission frequency. For example, when a user uses a video application, the data packets monitored by the base station will show high downlink traffic and high continuity. Current resource allocation strategies are primarily derived from pre-configured configuration files or database records in the network management system. These strategies include parameters such as the bandwidth allocated to each user and the number of resource blocks (RBs) used. For example, a 5G network operator allocates specific bandwidth ranges based on user plans or service requirements, and this information is clearly reflected in the resource allocation policy data. For example, when a user watches an HD video, the core network database will record that the user has been allocated a larger amount of spectrum resources to ensure smooth video playback. Alternatively, for smartphone users, detailed records are kept of download traffic generated by web browsing and upload traffic generated by operations such as sending images or files during a specific time period (e.g., 8:00 AM to 9:00 AM). The data format can be a traffic statistics series arranged by timestamp, such as traffic values per second or per minute, to facilitate subsequent analysis of traffic change trends.
[0080] Unsupervised physical layer-application layer joint feature extraction is performed on the traffic data to obtain traffic features of each of the users in the current time period.
[0081] Specifically, unsupervised learning algorithms, such as autoencoders or principal component analysis (PCA), are used to process traffic data. For example, an autoencoder is a neural network structure that automatically learns feature representations in data without the need for manual data labeling. Physical layer features include information such as signal strength and signal-to-noise ratio, while application layer features include information such as the type of service used by users (such as video, voice, and web browsing) and packet size distribution. To jointly extract features from these two layers, the physical and application layer data can be preprocessed separately, such as normalization, before being input into the autoencoder. The autoencoder learns a compressed representation of the data through its hidden layers, namely a joint feature vector. The value of each dimension represents a traffic feature extracted from both the physical and application layers. For example, a dimension might reflect the correlation between signal strength and packet size for video services.
[0082] The network behavior of the user in the current time period is predicted based on the traffic characteristics in the current time period to obtain the network usage pattern of the user.
[0083] Specifically, classification or clustering algorithms in machine learning are used to predict user network behavior patterns. For example, clustering algorithms (such as K-means clustering) are used to analyze user traffic characteristics and group users with similar traffic characteristics into the same category, thereby deriving different network usage pattern categories, such as "high-traffic video viewing mode" and "low-traffic web browsing mode."
[0084] The extracted traffic feature vectors are fed into the prediction model, which uses trained parameters (for supervised learning models) or cluster centers (for unsupervised clustering) to determine the user's current network behavior. For example, using K-means clustering, the distance between the traffic feature vector and each cluster center is calculated, and the vector is assigned to the category corresponding to the closest cluster center, thereby determining the user's network usage pattern.
[0085] In a preferred embodiment of the present invention, the network behavior of the user in the current time period is predicted based on the traffic characteristics in the current time period, and a support vector machine (SVM) model in a supervised learning model can be used.
[0086] Specifically, a large amount of traffic feature data labeled with network usage patterns is collected from base stations and the core network. These labels include common network usage patterns such as online gaming, video playback, web browsing, and file downloads. The collected data is normalized, mapping the traffic feature values to the range [0, 1] to improve model training efficiency and accuracy. Traffic characteristics closely related to network usage patterns are identified, such as uplink and downlink packet size, transmission frequency, and latency sensitivity. For example, online gaming typically has small uplink and downlink packet sizes but high transmission frequency and latency sensitivity; whereas video playback typically has large downlink packets with a certain degree of continuity.
[0087] The SVM model is trained using selected traffic features and corresponding network usage pattern labels. During training, the SVM classifies data based on different network usage patterns by finding the optimal hyperplane to maximize the classification margin. Selecting an appropriate kernel function, such as the Gaussian Radial Basis Function (RBF), can effectively process nonlinearly separable data and improve the model's classification performance. The feature extraction module in the core network processes real-time user traffic data, extracts the same traffic features as those used in the training phase, and normalizes them for input into the trained SVM model. The processed traffic features are input into the SVM model, which determines the user's current network usage pattern by calculating the distance between the input features and the support vector. The model outputs results, such as online gaming or video playback, and feeds the results back to the resource management control unit.
[0088] For example, consider a user playing a multiplayer online competitive game. During the game, the user's uplink and downlink data packets are small in size, sent very frequently, and are highly sensitive to latency. The feature extraction module in the core network extracts these traffic features and inputs them into a trained SVM model. Based on the mapping relationship between previously learned features and network usage patterns, the SVM model determines that the user's network usage pattern is an online gaming mode. This result is fed back to the resource management control unit for subsequent resource allocation and optimization based on the user's network usage pattern.
[0089] The user's perception score is determined based on the network usage pattern and in combination with the current resource allocation strategy.
[0090] Specifically, the user experience score comprehensively considers the user's service experience and the degree of matching of resource allocation. For example, for users in video viewing mode, the experience score may be related to factors such as the number of video pauses, loading time, and whether the allocated bandwidth meets the requirements for smooth video playback. For users in web browsing mode, the experience score is related to factors such as page loading speed and waiting time during browsing.
[0091] Based on different network usage patterns, a corresponding calculation model is established to determine the user experience score. Taking video services as an example, by analyzing the ratio of allocated bandwidth to required bandwidth, the number of freezes, and other parameters, a user experience score is calculated using a certain weighting formula (for example, freezes are given a higher weight). The score range can be set from 0 to 10, with a higher score indicating a better user experience.
[0092] According to the perception score, it is determined whether the current resource allocation strategy of the user needs to be adjusted.
[0093] Specifically, a threshold for the user experience score is first set. For example, if the experience score is below 6, the resource allocation policy is considered to be adjusted to improve the user's network experience. If the score is above or equal to 6, no adjustment is made and the existing resource allocation policy is maintained. The calculated experience score for each user is compared with the set threshold. If the score is below the threshold, the resource allocation policy adjustment mechanism is triggered; otherwise, the original policy remains unchanged. For example, in a 5G network cell with multiple users, each user's experience score is determined to meet the threshold, and the number of users requiring resource allocation policy adjustment and their specific user information are counted.
[0094] When the current resource allocation policy of the user needs to be adjusted, the resource adjustment amount for the user in the next time period is obtained through a reinforcement learning strategy network based on the user's network usage pattern, combined with the user's remaining bandwidth and the current resource allocation policy.
[0095] Specifically, a multi-layer neural network is constructed as a reinforcement learning policy network. The number of input layer nodes is determined according to the characteristic dimensions of the user network usage pattern, the remaining bandwidth, and the number of parameters of the current resource allocation strategy. For example, if the network usage pattern feature vector has 10 dimensions, and the remaining bandwidth and the current resource allocation strategy are each represented by one parameter, then the input layer has 12 nodes. The hidden layer can be designed to be 2-3 layers, with the number of nodes in each layer gradually decreasing, such as 20 nodes in the first hidden layer and 15 nodes in the second hidden layer. The number of output layer nodes is determined according to the possible value range of the resource adjustment amount. If the resource adjustment amount is a continuous value, then one node in the output layer is sufficient. If it is a discrete value, then each possible adjustment amount corresponds to one node.
[0096] The network's weights and biases are randomly initialized, and hyperparameters such as the learning rate and discount factor are set. The learning rate controls the step size of the network parameter update and is generally initially set to around 0.01. The discount factor measures the present value of future rewards and is typically between 0.9 and 0.99.
[0097] Based on the previously determined network usage pattern, the corresponding traffic characteristics are extracted as a pattern feature vector. Taking online gaming as an example, features such as uplink and downlink packet size, transmission frequency, and delay sensitivity are extracted and normalized to form a feature vector. The network usage pattern feature vector, the user's remaining bandwidth, and the current resource allocation policy parameters are concatenated to form a complete input state vector. For example, if the network usage pattern feature vector has 10 elements, the remaining bandwidth is represented by a normalized value, and the current resource allocation policy parameters (such as the allocated bandwidth size and number of resource blocks) are also normalized and represented by a single value, the input state vector will have a total of 12 elements.
[0098] Based on the resource allocation granularity of the 5G network and the user's maximum adjustable resource limit, define the possible range of resource adjustment amounts. For example, the bandwidth adjustment amount can range from -20MHz to +20MHz, with a step size of 5MHz; the resource block number adjustment amount can range from -10 to +10, with a step size of 1. If the resource adjustment amount is a continuous value, it can be discretized into multiple gears as an action space; if it is a discrete value, then each possible adjustment amount is directly determined as an action. For example, the bandwidth adjustment amount is quantized into 9 actions: -20MHz, -15MHz, -10MHz, -5MHz, 0, +5MHz, +10MHz, +15MHz, and +20MHz.
[0099] The reward function should comprehensively consider multiple factors, including improving user experience scores, resource utilization efficiency, and overall network performance. Improving user experience scores is the core goal, while high resource utilization efficiency can avoid resource waste, and good overall network performance can ensure service quality for multiple users.
[0100] For example, the reward function in the embodiment of the present invention is It includes the following parts:
[0101] ;
[0102] in: Indicates the change in user experience score, which is the core indicator for measuring whether the user's network experience has improved or deteriorated; Indicates resource utilization efficiency, which is used to measure whether the allocated resources are effectively utilized to avoid resource waste; Indicates changes in overall network performance, taking into account the impact of resource adjustments on other users in the network to ensure overall network performance stability; is the weight coefficient, which is used to balance the importance of each factor. In practical applications, it can be adjusted according to actual needs, for example , , .
[0103] Specifically, the change in user experience score , can be calculated through service quality indicators (QoS, Quality of Service) such as delay, packet loss rate, etc. Assume that the user experience score before adjustment is , after adjustment ,but Resource efficiency , which can be measured by the ratio of the amount of allocated resources to the amount of resources actually used. For example, if the allocated bandwidth is , the actual bandwidth used is ,but Overall network performance changes It can be measured by the average change in the perception scores of other users. Assume that the average perception score of other users in the network is and ,but .
[0104] For example, suppose that within a certain period of time, a user's resource allocation strategy needs to be adjusted. The specific steps include: In the initial state, the user's current perception score is (out of 10 points), the allocated bandwidth is MHz, the actual bandwidth used is MHz, the average perception score of other users in the network is Resource adjustment, through the reinforcement learning strategy network, it is decided to increase the bandwidth of 5MHz for this user. The new bandwidth allocation is MHz. The adjusted status is the user's new perception score is , the new actual bandwidth used is MHz, the average perception score of other users in the network is .
[0105] Calculation of various indicators includes the change in user experience score: , resource utilization efficiency: , the overall performance of the network changes: Calculating the reward value includes: assuming the weight coefficient is , , , reward value .
[0106] In this example, the user experience score improved significantly (by 1.0), resource utilization efficiency was high (0.8), and despite a slight decrease in overall network performance (-0.1), the overall reward value remained positive (1.35). This indicates that this resource adjustment was generally effective and should be rewarded positively to encourage the policy network to continue selecting similar actions. This reward function design effectively balances user experience, resource utilization efficiency, and overall network performance, thereby achieving intelligent 5G network resource allocation.
[0107] After resource adjustment, if the user's perception score significantly improves and resource utilization efficiency is high, a positive reward is given, for example, a reward value of +10; if the user's perception score improves but the resource utilization efficiency is low, a smaller positive reward is given, such as +5; if the user's perception score decreases, a negative reward is given, such as -10; at the same time, if the adjusted resource allocation has a negative impact on the overall network performance, such as causing a significant decrease in the perception scores of other users, an additional larger negative reward is given, such as -20.
[0108] During actual 5G network operation, data is continuously collected, including user network usage patterns, remaining bandwidth, current resource allocation strategies, corresponding resource adjustments, and changes in user experience scores. This collected data is fed as samples into a reinforcement learning policy network. Using optimization methods such as gradient descent, the network's weights and biases are updated based on reward signals, enabling the network to learn the optimal action strategy for selecting resource adjustments under different input conditions. For example, if the reward value obtained after a resource adjustment is large, the network parameter value for selecting that action in that state is increased; otherwise, the value is decreased.
[0109] When a user's resource allocation policy needs to be adjusted, the input state vector, consisting of the user's network usage pattern feature vector, remaining bandwidth, and current resource allocation policy parameters, is fed into the trained reinforcement learning policy network. Based on the input state and the learned policy, the policy network outputs the corresponding resource adjustment amount. For example, the network outputs a bandwidth adjustment of +5 MHz, indicating that the user should be allocated an additional 5 MHz of bandwidth in the next time period.
[0110] A future resource allocation strategy for the user in the next time period is determined according to the resource adjustment amount, and resources are allocated to the user according to the future resource allocation strategy.
[0111] Specifically, the user's current resource allocation is read from the network management system, including the currently allocated bandwidth, number of resource blocks, scheduling priority, etc. For example, the user's currently allocated bandwidth is , the number of resource blocks is . At the same time, obtain resource adjustments, such as bandwidth adjustments , Resource block adjustment amount These adjustments are calculated using reinforcement learning policy networks or other resource adjustment algorithms.
[0112] Update the user's resource allocation information based on the resource adjustment amount. For bandwidth resources, the bandwidth allocated in the future Similarly, the number of resource blocks allocated in the future It is necessary to ensure that the calculated future resource allocation is within a reasonable range, not exceeding the upper limit of available network resources, nor falling below the lower limit of resources required to ensure basic user services.
[0113] Based on the calculated future resource allocation, a detailed resource allocation strategy is developed. This includes determining the time period for resource allocation, the priority of resource allocation, etc. For example, Bandwidth and Resource blocks are allocated to meet their network usage needs. Resource allocation instructions are generated based on the resource allocation strategy. These instructions contain information such as user ID, type of allocated resources (such as bandwidth, resource blocks, etc.), amount of resources, and time period for resource allocation. For example, an instruction is generated: "Allocate user A in the next time period." Bandwidth and resource blocks".
[0114] The resource allocation instruction is sent to the base station and other network devices, which then make actual resource allocations to users based on the instruction. The base station allocates the specified resources to users by adjusting the resource allocation table, updating the scheduling information, etc. For example, based on the instruction received, the base station updates the resource allocation information of user A to Bandwidth and The system allocates a resource block and provides services to user A at the beginning of the next time period according to the new resource allocation information. During the resource allocation process, it is necessary to monitor the execution of resource allocation in real time to ensure its accuracy. If any anomalies in resource allocation are found, such as resource allocation failure or the allocated resources not matching the instructions, timely troubleshooting and resolution are carried out.
[0115] After resource allocation is complete, the network device provides feedback to the resource management control unit. This feedback includes information such as the actual amount of resources allocated and whether the resource allocation was successful. The resource management control unit records and updates the resource allocation status based on this feedback. If the resource allocation is successful, the new resource allocation information is recorded in the user's resource allocation file, serving as a basis for subsequent resource adjustments. If the resource allocation fails, the resource adjustment plan must be reconsidered and resource allocation must be repeated.
[0116] For example, suppose that at the end of a certain time period, according to the reinforcement learning policy network's calculations, user A's bandwidth resources are adjusted by 5 MHz, and the currently allocated bandwidth is 20 MHz. Then, the future resource allocation policy determines that user A will be allocated 25 MHz of bandwidth in the next time period. Based on this policy, the resource management control unit generates a resource allocation instruction and sends it to the base station. Upon receiving the instruction, the base station adjusts user A's bandwidth resources to 25 MHz at the beginning of the next time period and provides the corresponding network services to user A. Simultaneously, the base station sends feedback to the resource management control unit regarding the successful resource allocation, which then updates user A's resource allocation record, completing the entire resource allocation process.
[0117] The 5G network resource allocation method of the present invention obtains user traffic data and resource allocation policies, then performs unsupervised physical-layer and application-layer joint feature extraction. This method fully exploits the deep features of traffic data, providing a richer and more accurate data foundation for subsequent analysis. This overcomes the shortcomings of traditional networks that rely solely on network-level data (such as wireless channel quality and base station buffer occupancy), and addresses the bias in user experience evaluation caused by limited data dimensionality. Based on the extracted traffic features, user network behavior is predicted to obtain a network usage pattern. This pattern is then combined with the current resource allocation policy to determine a user experience score. This process transforms traffic features into specific usage scenarios, making the calculation of the experience score more accurate and directly linking data processing with user experience evaluation. The experience score is then used to determine whether the resource allocation policy needs to be adjusted. If adjustment is necessary, the amount of resource adjustment is determined using a reinforcement learning strategy network, taking into account the user's network usage pattern, remaining bandwidth, and the current resource allocation policy. The reinforcement learning policy network plays a key role here. Based on historical data and real-time feedback, it learns the optimal resource adjustment strategy. This strategy, which takes into account the user's remaining bandwidth, rationally allocates resources, avoiding waste or over-allocation. This addresses the existing problem of irrational resource allocation caused by slow sampling and limited dimensionality of user experience data. By fully integrating experience evaluation, resource adjustment decisions, and network behavior prediction, this invention forms a complete closed loop, achieving efficient dynamic resource allocation.
[0118] In general, the present invention can significantly improve the user's network usage experience. In the traditional 5G network resource allocation method, due to the slow sampling and small dimensions of user perception data, it is impossible to accurately capture the millisecond-level fluctuations of users in the process of performing services with high real-time requirements such as short videos or games, such as changes in delay and bit rate. This causes the detection data on the network side to deviate from the user's actual feelings. In the process of pursuing overall network efficiency, some users' network resources are unreasonably allocated, resulting in a poor user experience. However, this application obtains the traffic data and current resource allocation strategy of each user in the current time period, and performs unsupervised physical layer-application layer joint feature extraction on the traffic data, which can obtain the user's traffic characteristics more comprehensively and accurately. Based on these characteristics, the user's network usage pattern is predicted, and then the user perception score is determined in combination with the current resource allocation strategy, so that the user's real network usage experience can be reflected in a timely and accurate manner.
[0119] Further, the reinforcement learning strategy network is adopted to obtain the resource adjustment amount of the user in the next time period according to the user network usage mode, the residual bandwidth and the current resource allocation strategy, and to determine the future resource allocation strategy according to the resource adjustment amount, so as to realize efficient dynamic allocation of resources. Compared with the previous slice-based multi-service bearing mode and the method of outputting resource quotas by a simple deep reinforcement model, the application considers the dynamic changes of user experience data and network resource status, and can better meet the individual needs of users while ensuring the overall efficiency of the network.
[0120] Optionally, the traffic data of each user in the 5G network in the current time period and the current resource allocation strategy of the user are obtained, including:
[0121] The I / Q (In-phase / Quadrature) sample stream, control layer scheduling log, protocol data unit timestamp, packet arrival interval and instantaneous throughput value of each user are continuously collected according to a preset granularity, and the I / Q sample stream, control layer scheduling log, protocol data unit timestamp, packet arrival interval and instantaneous throughput value are used to form raw traffic data;
[0122] The raw traffic data is timestamp aligned and indexed at the slice level to obtain the traffic data of the user in the current time period;
[0123] The number of physical resource blocks, residual bandwidth, maximum schedulable modulation and coding policy level of the network slice in which the user is located in the current time period are obtained by a slice orchestrator;
[0124] The current resource allocation strategy is determined according to the number of physical resource blocks, residual bandwidth, maximum schedulable modulation and coding policy level.
[0125] Specifically, first, the I / Q (In-phase / Quadrature) sample stream, control layer scheduling log, protocol data unit timestamp, packet arrival interval and instantaneous throughput value of each user are continuously collected according to a preset granularity. For example, the I / Q sample stream can be collected once every time slot (such as 1ms) to record its amplitude and phase information; at the same time, scheduling decision information is extracted from the control layer scheduling log, including the scheduled user, the allocated resource block, etc.; the protocol data unit timestamp records the time when each protocol data unit enters and leaves the network layer; the packet arrival interval is determined by measuring the time difference between adjacent packet arrivals; the instantaneous throughput value is calculated according to the amount of data transmitted in each time period. These collected data are integrated together to form raw traffic data.
[0126] Next, the raw traffic data undergoes timestamp alignment and slice-level indexing. Timestamp alignment arranges data from different sources in chronological order, ensuring that each data point has an accurate time stamp. Slice-level indexing indexes data belonging to the same network slice based on the network slice's identifier, facilitating subsequent processing. For example, timestamp alignment arranges data such as I / Q sampling streams and control layer scheduling logs in chronological order, ensuring that each data point corresponds to a precise time point. Slice-level indexing assigns a unique index number to each network slice, marking and storing data belonging to the same network slice. Through these two steps, the user's traffic data for the current time period is obtained.
[0127] Next, the slice orchestrator obtains the number of physical resource blocks (PRBs), remaining bandwidth, and maximum schedulable modulation and coding strategy level for the user's network slice in the current time period. The slice orchestrator manages resource allocation and scheduling for the network slice. Based on the network slice's needs and current network conditions, it allocates PRBs, determines remaining bandwidth, and selects the appropriate modulation and coding strategy level. For example, during each time period, the slice orchestrator counts the number of allocated PRBs in the network slice, calculates remaining bandwidth, and determines the maximum schedulable modulation and coding strategy level based on channel conditions and service requirements. This information is fed back to the resource allocation system.
[0128] Finally, the user's current resource allocation strategy is determined based on the obtained number of physical resource blocks, remaining bandwidth, and the maximum schedulable modulation and coding strategy level. For example, the number of resource blocks that can be allocated to a user is determined based on the remaining bandwidth and number of physical resource blocks; the modulation method and coding rate selected for the user are determined based on the maximum schedulable modulation and coding strategy level. This information is combined to formulate the user's current resource allocation strategy, including parameters such as the number of allocated resource blocks and the modulation and coding strategy level.
[0129] In this optional embodiment, by collecting various detailed data at a preset granularity and performing timestamp alignment and slice-level indexing, users' traffic data and resource allocation can be accurately and comprehensively obtained. At the same time, the slice orchestrator obtains key resource information about network slices, providing a basis for determining precise resource allocation strategies. This improves the rationality and adaptability of resource allocation, helps improve the resource utilization efficiency of 5G networks and the user's network experience, enhances the overall performance and management capabilities of the network, and meets the diverse needs of different users.
[0130] Optionally, performing unsupervised physical layer-application layer joint feature extraction on the traffic data to obtain traffic features of each user in the current time period includes:
[0131] Dividing the traffic data into multiple continuous subsequences according to time windows;
[0132] Performing short-time Fourier transform on the I / Q sampling stream in each subsequence to obtain an in-phase orthogonal time-frequency diagram;
[0133] Time-aligning and concatenating the control layer scheduling log, the protocol data unit timestamp, the message arrival interval, and the instantaneous throughput value in the subsequence to obtain a cross-layer feature vector;
[0134] Inputting the in-phase orthogonal time-frequency map and the cross-layer feature vector into an unsupervised convolutional hybrid encoder to obtain an embedding vector of fixed dimension;
[0135] Time series pooling and normalization are performed on the embedding vectors corresponding to all the subsequences in the current time period to obtain a parameter-free traffic fingerprint of the user, and the parameter-free traffic fingerprint is used as the traffic feature.
[0136] Specifically, first, divide the traffic data into multiple continuous subsequences based on time windows. For example, if the current time period is one hour, it can be divided into multiple 10-second time windows, with the traffic data within each window forming a subsequence. This division can cut the data into segments that are easy to process and have temporal locality.
[0137] Next, a short-time Fourier transform (STFT) is performed on the I / Q sample stream in each subsequence to produce an in-phase orthogonal time-frequency representation (IQ-TFR, TFR, Time-Frequency Representation). The STFT reveals the frequency components of a signal at different time points and is very effective for extracting time-frequency information from I / Q sample streams. For example, using wireless signals, the STFT can convert I / Q sample data in the time domain into a representation in the time-frequency domain, enabling analysis of the signal's frequency characteristics and energy distribution at different time points.
[0138] Next, the control layer scheduling logs, protocol data unit timestamps, packet arrival intervals, and instantaneous throughput values in the subsequences are time-aligned and concatenated to produce a cross-layer feature vector. Time alignment involves arranging data from different layers or types in chronological order, ensuring that each data point corresponds to the same point in time. These aligned data are then concatenated to form a comprehensive cross-layer feature vector. For example, the scheduling priority in the control layer scheduling logs, the latency information in the protocol data unit timestamps, the statistical values of the packet arrival intervals, and the instantaneous throughput values are combined to form a vector of length N, where N is the sum of the feature dimensions.
[0139] The in-phase quadrature time-frequency map (IQ-TFR) and the cross-layer feature vector are then fed into an unsupervised convolutional hybrid encoder to produce a fixed-dimensional embedding vector. A convolutional hybrid encoder is a neural network architecture used to automatically learn feature representations. It extracts spatial features from the time-frequency map using convolutional layers, processes the cross-layer feature vector using fully connected layers, then fuses the two and uses an encoder to map the input data into a fixed-dimensional embedding space. For example, the convolutional hybrid encoder can convert the time-frequency map and cross-layer feature vector into a length-128 embedding vector that captures the key features of the original data.
[0140] Finally, time series pooling and normalization are performed on the embedding vectors corresponding to all subsequences in the current time period to obtain the user's parameter-free traffic fingerprint, which is used as the traffic feature. Time series pooling compresses information along the time dimension. For example, using average pooling or max pooling, the embedding vectors of multiple time steps are compressed into a fixed-length vector. Normalization scales the vector values to a specific range (such as [0, 1]) to improve data stability and comparability. The resulting parameter-free traffic fingerprint is a fixed-length vector that uniquely represents the user's traffic characteristics in the current time period.
[0141] In this optional embodiment, by dividing traffic data into time window subsequences and performing multi-level feature extraction and fusion, the joint characteristics of user traffic data at the physical layer and application layer can be effectively captured. Using an unsupervised convolutional hybrid encoder to extract embedding vectors and generate parameter-free traffic fingerprints eliminates the need for manual data annotation and reduces annotation costs. At the same time, this method can adaptively learn the inherent characteristic patterns in the data, improving the representation and discrimination of traffic characteristics, providing a more accurate basis for subsequent network behavior prediction and resource allocation, and enhancing the system's intelligence and performance.
[0142] Optionally, predicting the network behavior of the user in the current time period based on the traffic characteristics in the current time period to obtain the network usage pattern of the user includes:
[0143] Inputting the parameter-free traffic fingerprint into a network behavior prediction model;
[0144] Outputting a probability distribution of the network usage pattern of the user in the current time period through the network behavior prediction model;
[0145] Screening is performed according to the probability distribution of the network usage pattern, and the pattern with the highest probability value in the probability distribution is used as the network usage pattern of the user.
[0146] Specifically, before inputting the parameter-free traffic fingerprint into the network behavior prediction model, the data needs to be preprocessed to ensure that the model can effectively process the data. This includes normalizing the data and scaling each eigenvalue to a specific range (such as [0, 1]) to improve model training efficiency and accuracy. The preprocessed parameter-free traffic fingerprint is then fed into the network behavior prediction model as the input feature vector. For example, if the parameter-free traffic fingerprint is a vector of length 128, the model's input layer would be designed to have 128 neurons, with each neuron corresponding to one eigenvalue in the vector.
[0147] Network behavior prediction models can employ deep learning architectures such as multilayer perceptrons (MLPs), convolutional neural networks (CNNs), or recurrent neural networks (RNNs). Taking an MLP as an example, the model consists of an input layer, two hidden layers, and an output layer. The hidden layers use the ReLU activation function, while the output layer uses the softmax activation function to generate a probability distribution. During the model training phase, the model is trained using historical traffic data labeled with network usage patterns. Optimization algorithms (such as Adam, Adaptive Moment Estimation, and the Adam optimizer) and loss functions (such as cross-entropy loss) are used to adjust the model parameters, enabling the model to accurately predict the probability distribution of network usage patterns. Each neuron in the model's output layer corresponds to a network usage pattern (such as video playback, web browsing, and file downloading), and the output value represents the probability of that pattern. For example, the model output might be [0.1, 0.7, 0.2], corresponding to the probabilities of video playback, web browsing, and file downloading, respectively, with web browsing having the highest probability.
[0148] Analyze the probability distribution output by the model, compare the probability values corresponding to each network usage pattern, and identify the pattern with the highest probability. The pattern with the highest probability is determined as the user's network usage pattern. For example, in the output above, web browsing has a probability of 0.7, the highest of all patterns, so web browsing is determined as the user's current network usage pattern.
[0149] For example, suppose that during a certain time period, a user's parameter-free traffic fingerprint is preprocessed and then input into a trained MLP network behavior prediction model. The model's output layer has three neurons, corresponding to video playback, web browsing, and file downloading. The model outputs a probability distribution of [0.1, 0.7, 0.2]. Based on this probability distribution, web browsing has the highest probability, so web browsing is determined to be the user's network usage mode.
[0150] In this optional embodiment, accurate prediction of user network behavior can be achieved by inputting parameter-free traffic fingerprints into a network behavior prediction model and outputting a probability distribution. This method uses a deep learning model to automatically learn complex feature patterns in the data, eliminating the need for manual feature extraction, thereby improving the accuracy and efficiency of the prediction. At the same time, screening through probability distribution can provide an objective and quantitative basis for determining users' network usage patterns, enhancing the reliability and stability of the system. This probability distribution-based screening method can also provide detailed pattern matching information for subsequent resource allocation decisions, helping to further optimize resource allocation strategies and improve user experience and network performance.
[0151] Optionally, determining the user's perception score based on the network usage pattern in combination with the current resource allocation strategy includes:
[0152] According to the network usage pattern, obtaining a network experience mapping table corresponding to the network usage pattern;
[0153] Determining a standard resource allocation strategy corresponding to the network usage pattern according to the network experience mapping table;
[0154] Mapping the number of physical resource blocks, the remaining bandwidth, the maximum schedulable modulation, and the coding strategy level in the standard resource allocation strategy as benchmark parameters;
[0155] Comparing the number of physical resource blocks, the remaining bandwidth, the maximum schedulable modulation, and the coding strategy level in the current resource allocation strategy with the benchmark parameters one by one, to obtain deviations corresponding to the number of physical resource blocks, the remaining bandwidth, the maximum schedulable modulation, and the coding strategy level, respectively;
[0156] Obtaining, by means of a quantization function of the network experience mapping table and according to the deviation, local perception scores corresponding to the number of physical resource blocks, the remaining bandwidth, the maximum schedulable modulation, and the coding strategy level, respectively;
[0157] The local perception scores are weighted and summed to obtain the perception score of the user.
[0158] Specifically, a network experience mapping table is pre-stored in the system database. It records the ideal standard resource allocation strategy and corresponding quantization function for each network usage mode (such as video playback, web browsing, online gaming, etc.). After determining the user's network usage mode, the system calls the corresponding network experience mapping table from the database. For example, if the user's network usage mode is online gaming, the system calls the network experience mapping table for online gaming.
[0159] In the obtained network experience map, find the standard resource allocation strategy corresponding to the network usage model, including the standard number of physical resource blocks, remaining bandwidth, and the maximum schedulable modulation and coding strategy level. For example, the standard resource allocation strategy for online gaming might be: 15 physical resource blocks, 20 MHz remaining bandwidth, and a maximum schedulable modulation and coding strategy level of 6.
[0160] For example, a network experience mapping table is a data structure for associating network usage patterns with corresponding resource allocation strategies and user experience quantitative indicators. It is usually in the form of a table or database and specifically includes the following content, as shown in Table 1:
[0161] Table 1 Network experience mapping table
[0162]
[0163] As shown in Table 1, network usage models encompass common network application types, such as video playback, web browsing, online gaming, and file downloading. For each network usage model, the resource parameters required for optimal operation are defined, including the number of physical resource blocks, remaining bandwidth, and the maximum schedulable modulation and coding strategy level. These standard values are determined through network performance testing, user satisfaction surveys, and business needs analysis.
[0164] The quantization function is a mathematical function used to convert the deviation of resource allocation parameters into a local perception score. Different network usage patterns may correspond to different quantization functions. For example, video playback uses a linear function, with the local perception score decreasing linearly with increasing deviation. Web browsing uses a quadratic function, with the local perception score decreasing quadratically with increasing deviation. This function setting reflects the moderate sensitivity of web browsing to network resources. When resource allocation is close to ideal, the user experience is good, and the local perception score decreases slowly. However, as the deviation increases, the score decreases more rapidly, reflecting the significant impact of large deviations on the user experience. For example, when the deviation of the number of physical resource blocks is 0.2, the local perception score is 100×(1-0.2)²=64 points; when the deviation increases to 0.5, the score drops to 100×(1-0.5)²=25 points. Online gaming uses a Gaussian function, and the local perception score is more sensitive to changes in deviation. It decreases slowly when the deviation is small and decreases more rapidly as the deviation increases, reflecting the high sensitivity of online gaming to network resources. File downloading uses a quadratic function, emphasizing the serious impact of large deviations on user experience.
[0165] In a practical system, a network experience mapping table can be stored in a relational database, with each row representing a network usage pattern and its corresponding standard resource allocation strategy and quantification function. The network experience mapping table can be regularly updated based on the development of network technology, changes in user preferences, and new business needs. For example, with the evolution of 5G networks, the values in the standard resource allocation strategy may need to be adjusted to accommodate higher network performance. In the resource allocation system, when a user's network usage pattern is detected, the corresponding standard resource allocation strategy and quantification function are retrieved by querying the network experience mapping table. These are used to calculate the user experience score and guide resource adjustments. Through this network experience mapping table, the system can link abstract network usage patterns with specific resource allocation parameters and user experience quantitative indicators, providing a basis for refined network resource management and optimization.
[0166] The number of physical resource blocks, remaining bandwidth, and maximum schedulable modulation and coding strategy level in the current resource allocation strategy are compared with the corresponding parameters in the standard resource allocation strategy, and the respective deviations are calculated. The deviation can be expressed as the ratio of the difference between the current value and the standard value to the standard value. For example, if the number of physical resource blocks in the current resource allocation strategy is 12 and the standard value is 15, the deviation of the number of physical resource blocks is (12-15) / 15=-0.2; if the current remaining bandwidth is 18MHz and the standard value is 20MHz, the deviation is (18-20) / 20=-0.1; if the maximum schedulable modulation and coding strategy level is 5 and the standard value is 6, the deviation is (5-6) / 6≈-0.167.
[0167] Using the quantization function in the network experience mapping table, the corresponding local perception score is calculated based on the deviation of each parameter. The quantization function can be a predefined functional relationship. For example, for the number of physical resource blocks, the quantization function can be a linear function, and the deviation is inversely proportional to the local perception score. Assuming the quantization function is: local perception score = 100×(1-|deviation|), then in the above example, the local perception score of the number of physical resource blocks is 100×(1-0.2)=80; the local perception score of the remaining bandwidth is 100×(1-0.1)=90; and the local perception score of the maximum schedulable modulation and coding strategy level is 100×(1-0.167)=83.3.
[0168] A corresponding weight is set for each local perception score, and the sum of the weights is 1. For example, if the weight of the number of physical resource blocks is set to 0.4, the weight of the remaining bandwidth is set to 0.3, and the weight of the maximum schedulable modulation and coding strategy level is set to 0.3, the user perception score is: 80 × 0.4 + 90 × 0.3 + 83.3 × 0.3 = 83.99.
[0169] In this optional embodiment, network usage patterns are closely integrated with resource allocation strategies through a network experience mapping table and quantification function, enabling a quantitative assessment of user experience. This method accurately reflects the impact of the difference between the current resource allocation strategy and the ideal strategy on user experience, providing an intuitive quantitative basis for subsequent resource adjustment decisions. A weighted summation approach comprehensively considers the varying degrees of impact of different resource allocation parameters on user experience, making the experience score more representative and instructive, helping to improve user satisfaction and network service quality.
[0170] Optionally, judging whether it is necessary to adjust the current resource allocation strategy of the user according to the perception score includes:
[0171] By comparing the perception score with a preset perception threshold, determining whether the current resource allocation strategy of the user needs to be adjusted;
[0172] If the perception score is lower than the preset perception threshold, determining that the current resource allocation strategy of the user needs to be adjusted;
[0173] If the perception score is not lower than the preset perception threshold, it is determined that there is no need to adjust the current resource allocation strategy of the user.
[0174] Specifically, the preset experience threshold is a key indicator for determining whether resource allocation strategies need adjustment. This threshold is set by comprehensively considering factors such as user experience expectations, network performance requirements, and service type. For example, for services with high network quality requirements, such as online gaming, the preset experience threshold might be set at 80 points; while for services with relatively low network quality requirements, such as web browsing, the preset experience threshold might be set at 70 points.
[0175] Using the previously described method for determining user perception scores, calculate the current user's perception score. For example, assume the calculated user perception score is 75. Compare the calculated perception score with a preset perception threshold. If the perception score is lower than the preset threshold, determine that the current resource allocation policy needs to be adjusted. Conversely, if the perception score is not lower than the preset threshold, determine that no adjustment is required.
[0176] Assuming the preset perception threshold is 80 points, if the user's perception score is 75 points (lower than the threshold), it is determined that the resource allocation strategy needs to be adjusted; if the user's perception score is 85 points (not lower than the threshold), it is determined that no adjustment is required.
[0177] In this optional embodiment, by comparing the preset feeling threshold with the user feeling score, a quick and objective judgment is realized on whether the resource allocation strategy needs to be adjusted. This method can timely identify the situation that the user experience is poor, thereby triggering the resource adjustment mechanism to optimize the network resource allocation. At the same time, it avoids unnecessary resource adjustment, improves the stability and resource utilization efficiency of the system, and ensures that the user can obtain better network experience in most cases.
[0178] Optionally, the resource adjustment amount of the user in the next time period is obtained by a reinforcement learning strategy network according to the network usage mode of the user, in combination with the residual bandwidth of the user and the current resource allocation strategy, and includes:
[0179] The network usage mode, the residual bandwidth, and the current resource allocation strategy are taken as state inputs of the reinforcement learning strategy network;
[0180] A candidate resource adjustment action set is generated in a preset action space by the reinforcement learning strategy network according to the state inputs;
[0181] The candidate resource adjustment action set includes physical resource block control actions, bandwidth control actions, modulation and coding strategy level control actions, etc.
[0182] Each control action in the candidate resource adjustment action set is evaluated in value according to a reward function, and the feeling score of the user and the total throughput of the 5G network system are taken as optimization objectives for joint optimization to obtain the resource adjustment amount of the user in the next time period.
[0183] Specifically, the network usage mode, the residual bandwidth, and the current resource allocation strategy are taken as state inputs and provided to the reinforcement learning strategy network. For example, assuming that the network usage mode of the user is online gaming, the current residual bandwidth is 20MHz, and the current resource allocation strategy is 15 physical resource blocks and modulation and coding strategy level 6, these information will be encoded and input into the reinforcement learning strategy network.
[0184] The reinforcement learning strategy network generates a candidate resource adjustment action set in a preset action space according to the input state. The action space includes physical resource block control actions, bandwidth control actions, and modulation and coding strategy level control actions. For example, the physical resource block control actions may include increasing or decreasing 1-5 physical resource blocks; the bandwidth control actions may include increasing or decreasing 5-10MHz; and the modulation and coding strategy level control actions may include increasing or decreasing 1-2 levels.
[0185] A reward function is used to assess the value of each control action. The reward function considers the user experience score and the total throughput of the 5G network system as optimization objectives. For example, if a control action significantly improves the user experience score while having a minimal impact on the total throughput, the action is considered highly valuable. Specifically, the reward function can be defined as the weighted sum of the improvement in the user experience score and the change in the total throughput.
[0186] Using a reinforcement learning algorithm (such as Q-learning or deep reinforcement learning), the control action with the highest value is selected as the final resource adjustment. For example, if an evaluation shows that the combination of adding two physical resource blocks, 5MHz bandwidth, and increasing the modulation and coding strategy level by one maximizes the reward function value, then this combination will be selected as the resource adjustment for the next time period.
[0187] This optional embodiment dynamically adjusts resource allocation based on user needs and network status, effectively improving the user experience while also balancing the overall performance of the network system. By optimizing the network through reinforcement learning strategies, intelligent and automated resource allocation is achieved, improving resource utilization efficiency and enhancing the adaptability and flexibility of the network.
[0188] Optionally, determining a future resource allocation strategy for the user in the next time period based on the resource adjustment amount, and allocating resources to the user according to the future resource allocation strategy includes:
[0189] The resource adjustment amount is converted into a physical resource block allocation value, a bandwidth allocation value, and a modulation and coding strategy level allocation value in the next time period through a resource scheduling engine;
[0190] generating the future resource allocation strategy for the user according to the physical resource block allocation value, the bandwidth allocation value, and the modulation and coding strategy level allocation value;
[0191] The future resource allocation strategy is sent to the base station scheduler of the network slice where the user is located, and the base station scheduler allocates wireless resources to the user according to the future resource allocation strategy within the next time period.
[0192] Specifically, after receiving the resource adjustment from the reinforcement learning strategy network, the resource orchestration engine converts it into specific allocation values for physical resource blocks, bandwidth, and modulation and coding strategy levels. For example, assuming the resource adjustment is to increase by 2 physical resource blocks, increase by 5MHz bandwidth, and increase the modulation and coding strategy level by 1, the current allocation is 15 physical resource blocks, 20MHz bandwidth, and modulation and coding strategy level 6. The resource orchestration engine will calculate the specific allocation values for the next time period based on the adjustment: 17 physical resource blocks (15+2), 25MHz bandwidth (20+5), and modulation and coding strategy level 7 (6+1).
[0193] Based on the converted allocation values, the user's future resource allocation policy is generated. This policy specifies the number of physical resource blocks, bandwidth, and modulation and coding strategy level to be allocated to the user in the next time period. For example, the generated policy might read: "User A will be allocated 17 physical resource blocks, 25 MHz bandwidth, and modulation and coding strategy level 7 in the next time period."
[0194] The generated future resource allocation policy is sent to the base station scheduler in the user's network slice. The base station scheduler receives and interprets the policy, updating its internal resource allocation table. For example, after receiving user A's resource allocation policy, the base station scheduler stores it in the resource allocation table and prepares it for execution in the next time period.
[0195] At the beginning of the next time period, the base station scheduler allocates appropriate wireless resources to the user based on the updated resource allocation policy. For example, the base station scheduler allocates 17 physical resource blocks and 25 MHz bandwidth to user A according to the policy, and sets the modulation and coding strategy level to 7, ensuring that user A can obtain the appropriate resources to meet their network usage needs during the time period.
[0196] In this optional embodiment, the resource orchestration engine and base station scheduler work together to accurately convert and execute resource adjustments into specific resource allocations, ensuring timely updates and application of resource allocation policies. This process improves the flexibility and adaptability of resource allocation, enabling rapid adjustments based on dynamic user needs, thereby optimizing the user experience, improving network resource utilization efficiency, and enhancing the overall performance and service quality of the 5G network.
[0197] Combine Figure 2 As shown, the present invention provides a 5G network resource allocation system, including:
[0198] An acquisition unit, configured to acquire traffic data of each user in the 5G network in a current time period and the user's current resource allocation strategy;
[0199] a feature extraction unit, configured to perform unsupervised physical layer-application layer joint feature extraction on the traffic data to obtain traffic features of each user in the current time period;
[0200] a prediction unit, configured to predict the network behavior of the user in the current time period based on the traffic characteristics in the current time period, and obtain a network usage pattern of the user;
[0201] a scoring unit, configured to determine a user experience score based on a network usage pattern and in combination with the current resource allocation strategy;
[0202] a judging unit, configured to judge whether it is necessary to adjust the current resource allocation strategy of the user according to the perception score;
[0203] an adjustment decision unit, configured to, when it is necessary to adjust the current resource allocation policy of the user, determine, based on the network usage pattern of the user, the remaining bandwidth of the user, and the current resource allocation policy, a resource adjustment amount for the user in the next time period through a reinforcement learning strategy network;
[0204] The resource allocation unit is configured to determine a future resource allocation strategy for the user in the next time period according to the resource adjustment amount, and allocate resources to the user according to the future resource allocation strategy.
[0205] The advantages of the 5G network resource allocation system of the present invention over the prior art are the same as the advantages of the above-mentioned 5G network resource allocation method over the prior art, and will not be repeated here.
[0206] Combine Figure 3 As shown, the present invention provides an electronic device, including a memory and a processor;
[0207] The memory is used to store computer programs;
[0208] The processor is used to implement the 5G network resource allocation method as described above when executing the computer program.
[0209] The advantages of the electronic device of the present invention over the prior art are the same as the advantages of the above-mentioned 5G network resource allocation method over the prior art, and will not be repeated here.
[0210] Although the present invention is disclosed as above, the scope of protection disclosed by the present invention is not limited thereto. Those skilled in the art may make various changes and modifications without departing from the spirit and scope of the present invention, and these changes and modifications will fall within the scope of protection of the present invention.
Claims
1. A 5G network resource allocation method, characterized in that: include: Obtain the traffic data of each user in the 5G network during the current time period and the user's current resource allocation strategy; Performing unsupervised physical layer-application layer joint feature extraction on the traffic data to obtain traffic characteristics of each user in the current time period; Predicting the network behavior of the user in the current time period based on the traffic characteristics in the current time period to obtain a network usage pattern of the user; Determining the user's experience score based on the network usage pattern and the current resource allocation strategy; Determining whether it is necessary to adjust the current resource allocation strategy of the user according to the perception score; When the current resource allocation policy of the user needs to be adjusted, the resource adjustment amount for the user in the next time period is obtained through a reinforcement learning strategy network based on the user's network usage pattern, the user's remaining bandwidth, and the current resource allocation policy; A future resource allocation strategy for the user in the next time period is determined according to the resource adjustment amount, and resources are allocated to the user according to the future resource allocation strategy.
2. The 5G network resource allocation method according to claim 1, characterized in that: The obtaining of traffic data of each user in the 5G network in the current time period and the user's current resource allocation strategy includes: Continuously collecting the I / Q sampling stream, control layer scheduling log, protocol data unit timestamp, message arrival interval, and instantaneous throughput value of each user according to a preset granularity, and forming raw traffic data based on the I / Q sampling stream, the control layer scheduling log, the protocol data unit timestamp, the message arrival interval, and the instantaneous throughput value; Performing timestamp alignment and slice-level indexing on the original traffic data to obtain the traffic data of the user in the current time period; Obtaining, through the slice scheduler, the number of physical resource blocks, the remaining bandwidth, and the maximum schedulable modulation and coding strategy level of the network slice where the user is located in the current time period; The current resource allocation strategy is determined according to the number of physical resource blocks, the remaining bandwidth, the maximum schedulable modulation and the coding strategy level.
3. The 5G network resource allocation method according to claim 2, characterized in that: The performing unsupervised physical layer-application layer joint feature extraction on the traffic data to obtain traffic features of each user in the current time period includes: Dividing the traffic data into a plurality of continuous subsequences according to time windows; Performing short-time Fourier transform on the I / Q sampling stream in each subsequence to obtain an in-phase orthogonal time-frequency diagram; Time-aligning and concatenating the control layer scheduling log, the protocol data unit timestamp, the message arrival interval, and the instantaneous throughput value in the subsequence to obtain a cross-layer feature vector; Inputting the in-phase orthogonal time-frequency map and the cross-layer feature vector into an unsupervised convolutional hybrid encoder to obtain an embedding vector of fixed dimension; Time series pooling and normalization are performed on the embedding vectors corresponding to all the subsequences in the current time period to obtain a parameter-free traffic fingerprint of the user, and the parameter-free traffic fingerprint is used as the traffic feature.
4. The 5G network resource allocation method according to claim 3, characterized in that: The predicting the network behavior of the user in the current time period according to the traffic characteristics in the current time period to obtain the network usage pattern of the user includes: Inputting the parameter-free traffic fingerprint into a network behavior prediction model; Outputting a probability distribution of the network usage pattern of the user in the current time period through the network behavior prediction model; Screening is performed according to the probability distribution of the network usage pattern, and the pattern with the highest probability value in the probability distribution is used as the network usage pattern of the user.
5. The 5G network resource allocation method according to claim 2, wherein: The determining the user's perception score based on the network usage pattern and in combination with the current resource allocation strategy includes: According to the network usage pattern, obtaining a network experience mapping table corresponding to the network usage pattern; Determining a standard resource allocation strategy corresponding to the network usage pattern according to the network experience mapping table; Mapping the number of physical resource blocks, the remaining bandwidth, the maximum schedulable modulation, and the coding strategy level in the standard resource allocation strategy as benchmark parameters; Comparing the number of physical resource blocks, the remaining bandwidth, the maximum schedulable modulation, and the coding strategy level in the current resource allocation strategy with the benchmark parameters one by one, to obtain deviations corresponding to the number of physical resource blocks, the remaining bandwidth, the maximum schedulable modulation, and the coding strategy level, respectively; Obtaining, by means of a quantization function of the network experience mapping table and according to the deviation, local perception scores corresponding to the number of physical resource blocks, the remaining bandwidth, the maximum schedulable modulation, and the coding strategy level, respectively; The local perception scores are weighted and summed to obtain the perception score of the user.
6. The 5G network resource allocation method according to claim 1, wherein: The determining, based on the perception score, whether it is necessary to adjust the current resource allocation strategy of the user includes: By comparing the perception score with a preset perception threshold, determining whether the current resource allocation strategy of the user needs to be adjusted; If the perception score is lower than the preset perception threshold, determining that the current resource allocation strategy of the user needs to be adjusted; If the perception score is not lower than the preset perception threshold, it is determined that there is no need to adjust the current resource allocation strategy of the user.
7. The 5G network resource allocation method according to claim 1, characterized in that: The step of obtaining a resource adjustment amount for the user in the next time period through a reinforcement learning strategy network based on the network usage pattern of the user, in combination with the remaining bandwidth of the user and the current resource allocation strategy, includes: Using the network usage pattern, the remaining bandwidth, and the current resource allocation strategy as state inputs of the reinforcement learning strategy network; Generating a set of candidate resource adjustment actions in a preset action space according to the state input through the reinforcement learning strategy network; The candidate resource adjustment action set includes a physical resource block control action, a bandwidth control action, and a modulation and coding strategy level control action; A value evaluation is performed on each control action in the candidate resource adjustment action set according to the reward function, and a joint optimization is performed with the user's perception score and the total throughput of the 5G network system as optimization targets to obtain the resource adjustment amount of the user in the next time period.
8. The 5G network resource allocation method according to claim 1, wherein: The determining, based on the resource adjustment amount, a future resource allocation strategy for the user in the next time period, and allocating resources to the user according to the future resource allocation strategy, includes: The resource adjustment amount is converted into a physical resource block allocation value, a bandwidth allocation value, and a modulation and coding strategy level allocation value in the next time period through a resource scheduling engine; generating the future resource allocation strategy for the user according to the physical resource block allocation value, the bandwidth allocation value, and the modulation and coding strategy level allocation value; The future resource allocation strategy is sent to the base station scheduler of the network slice where the user is located, and the base station scheduler allocates wireless resources to the user according to the future resource allocation strategy within the next time period.
9. A 5G network resource allocation system, characterized in that: include: An acquisition unit, configured to acquire traffic data of each user in the 5G network in a current time period and the user's current resource allocation strategy; a feature extraction unit, configured to perform unsupervised physical layer-application layer joint feature extraction on the traffic data to obtain traffic features of each user in the current time period; a prediction unit, configured to predict the network behavior of the user in the current time period based on the traffic characteristics in the current time period, and obtain a network usage pattern of the user; a scoring unit, configured to determine a user experience score based on a network usage pattern and in combination with the current resource allocation strategy; a judging unit, configured to judge whether it is necessary to adjust the current resource allocation strategy of the user according to the perception score; an adjustment decision unit, configured to, when it is necessary to adjust the current resource allocation policy of the user, determine, based on the network usage pattern of the user, the remaining bandwidth of the user, and the current resource allocation policy, a resource adjustment amount for the user in the next time period through a reinforcement learning strategy network; The resource allocation unit is configured to determine a future resource allocation strategy for the user in the next time period according to the resource adjustment amount, and allocate resources to the user according to the future resource allocation strategy.
10. An electronic device, characterized in that: including memory and processor; The memory is used to store computer programs; The processor is configured to implement the 5G network resource allocation method according to any one of claims 1 to 8 when executing the computer program.
Citation Information
Patent Citations
Network resource allocation method and device, computer program product and electronic equipment
CN119996341A
Training reinforcement learning agents to perform multiple tasks across diverse domains
WO2024149747A1