Mobile internet access IP address intelligent distribution method and system based on dynamic strategy
By combining the masking autoencoder model and the near-end strategy optimization algorithm, the problem of inefficient IP address allocation in the mobile Internet is solved, dynamic perception and optimization control are realized, and the adaptability and stability of IP address allocation are improved.
Patent Information
- Application Number
- CN202510837221.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-23
AI Technical Summary
The prior art IP address allocation method in the mobile Internet is difficult to adapt to complex and changeable network environments, resulting in low allocation efficiency, increased probability of address conflict and waste of resources, and cannot meet the needs of mobile terminals for stable connections and low-latency access.
The masked autoencoder model and near-end strategy optimization algorithm are used to build an IP address intelligent allocation system through multi-source feature data modeling, combining policy trusted interval adjustment and external condition guidance, and realize dynamic perception and optimization control.
It significantly improves the adaptability and stability of IP address allocation, reduces the probability of conflict allocation, improves resource utilization efficiency and terminal access success rate, and adapts to changes in complex network conditions.
Smart Images

Figure CN120358218A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer networks, and in particular to a method and system for intelligently allocating mobile Internet access IP addresses based on dynamic strategies. Background Art
[0002] In the context of the rapid development of mobile Internet, the dynamic management and allocation of IP address resources has become one of the key issues to be solved in wireless communication systems. Traditional IP address allocation methods are mostly based on static allocation strategies or preset rules, such as allocation based on first-come, first-served (FIFO) or MAC address mapping by DHCP servers, which are difficult to adapt to complex and changing network environments. Especially in scenarios with frequent user access, drastic network load fluctuations, and limited address pool resources, there are often problems such as low allocation efficiency, increased probability of address conflicts, and waste of resources, which leads to a decline in the overall performance of the system and cannot meet the needs of mobile terminals for stable connections and low-latency access.
[0003] In existing research, some methods attempt to introduce rule engines, heuristic searches, or simple machine learning models to optimize the IP address allocation process, but these methods often have weak generalization capabilities and insufficient state modeling capabilities. Especially when dealing with high-dimensional, nonlinear, multi-source heterogeneous access behavior and network state data, existing methods find it difficult to construct high-quality state representations, and thus cannot effectively support the adaptive generation of allocation strategies. In addition, these methods generally lack effective use of network state feedback information after policy execution, resulting in untimely policy updates and difficulty in achieving long-term stable allocation performance.
[0004] Although deep reinforcement learning has been gradually applied to the field of network resource scheduling in recent years and has good policy self-learning capabilities, it still faces challenges in terms of the accuracy of state input and the controllability of policy output. Existing deep policy models generally rely on direct input of raw data to model network states, without processing missing, incomplete or noisy features in the data, which affects the model's perception accuracy. At the same time, there is often a lack of effective prior guidance mechanisms and policy stability constraints in the policy optimization process, which can easily lead to policy shocks and allocation anomalies.
[0005] Therefore, the prior art urgently needs a new method that integrates state representation enhancement, self-supervised learning and enhanced strategy optimization, which can extract robust network state features from multi-source data, introduce feedback loops and structural constraints in the process of strategy generation and update, and improve the intelligence, stability and resource utilization efficiency of IP address allocation. The present invention proposes innovative solutions to the above problems and makes up for the shortcomings of the prior art.
[0006] Therefore, how to provide an intelligent allocation method and system for mobile Internet access IP addresses based on dynamic policies is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0007] An object of the present invention is to propose an intelligent allocation method and system for mobile Internet access IP addresses based on dynamic policies. The present invention integrates a masked autoencoder model and a proximal policy optimization algorithm to construct an intelligent control process that perceives multi-source feature data and adaptively generates and optimizes IP allocation policies. Specifically, by introducing a multi-scale masking mechanism to process access behavior and network status data, the robustness of state modeling is improved, and combined with policy confidence interval adjustment, external condition guidance, and constraint regularization mechanisms, stable generation and efficient optimization of allocation policies are achieved. This method has the advantages of strong state perception ability, high policy execution flexibility, excellent resource utilization efficiency, and strong ability to adapt to dynamic network environments, and is suitable for dynamic scheduling and refined management of IP addresses in complex scenarios.
[0008] According to the intelligent allocation method for mobile Internet access IP addresses based on dynamic policies in an embodiment of the present invention, the method includes the following steps:
[0009] S1. Collect access behavior data of mobile terminals, current network status information, and IP address pool usage data, construct a multi-source feature data set, and perform masking processing on the multi-source feature data set to form masked input data;
[0010] S2. Construct a masked autoencoder model, input the masked input data into the masked autoencoder model, perform encoding and decoding operations, and generate a reconstructed network state representation vector;
[0011] S3. Use the network state representation vector as input, apply the proximal policy optimization algorithm to generate and optimize the IP address allocation policy, control the address scheduler to allocate IP addresses from the target address pool to the corresponding mobile terminals according to the IP address allocation policy, and record the allocation execution results and network status feedback information;
[0012] S4. Calculate the policy execution reward value based on the allocation execution results, and use the policy execution reward value to update the policy parameters in the proximal policy optimization algorithm;
[0013] S5. Input the network status feedback information into the masked autoencoder model, adjust the encoder and decoder parameters, and loop through steps S1 to S4 within a preset allocation period to achieve continuous optimization and allocation control of the IP address dynamic policy.
[0014] Optionally, the masking process on the multi-source feature dataset means that in the constructed multi-source feature dataset, a masking operation is performed on some feature dimensions in the multi-source feature dataset according to a masking ratio of 30%; the masking operation includes setting the position of the selected feature to zero, marking it as an invalid value, and replacing it with a mean value; the masking strategy adopted is a hybrid strategy combining random masking and feature importance-guided masking, where 70% of the masking positions are determined by randomly selecting feature dimensions, and the remaining 30% of the masking positions are determined according to the influence weight ranking of the features in the prior evaluation, and the high-weight key features are masked first.
[0015] Optionally, S2 specifically includes:
[0016] S21. Construct a masked input dataset for multi-source features , where is the number of samples, is the feature dimension, is the real number set, and multiple masked versions are generated for the masked input dataset , where the masking ratio ;
[0017] S22. Calculate the feature correlation matrix based on the co-variability between feature dimensions in the masked input dataset , where represents the statistical correlation coefficient between feature and feature . Based on the feature pairs in the feature correlation matrix that satisfy the correlation threshold , construct a heterogeneous graph structure , where the node set represents the feature dimension, and the edge set represents the connection relationship that satisfies . Use the graph attention mechanism to calculate the structural importance score for each node;
[0018] S23. Perform a mixed masking operation on each masked version according to the structural importance score, which includes generating the masking positions through random selection and graph-guided selection respectively according to the ratio to form a masked input matrix ;
[0019] S24. Construct a dual-encoder structure, including a main encoder and an auxiliary encoder, extract the network state feedback information and policy execution records in the previous allocation cycle, and construct a historical state dataset. The main encoder receives the masked input matrix , and the auxiliary encoder receives the historical state data ;
[0020] S25. Respectively input the masked input matrix and the historical state data into the main encoder and the auxiliary encoder, and use a fusion mechanism to calculate the latent state vector ;
[0021] S26. Construct a conditional decoder , collect external constraint information related to the current access policy, including network topology status, service level requirements, address pool tightness, and terminal device categories, construct an external conditional variable , concatenate the latent state vector and the external conditional variable , and input them into the decoder for conditional reconstruction;
[0022] S27. Construct multiple structurally heterogeneous parallel paths in the decoder, including a fully connected decoding path, a residual skip connection path, and an attention enhancement path, and respectively output the preliminary reconstruction results ;
[0023] S28. Construct a policy feedback-guided fusion controller module , where represents the feedback information of the previous round of policy execution, including the policy execution reward value , the IP allocation success rate and the remaining rate of the address pool . The fusion controller jointly generates the decoding path fusion weight based on the latent state vector and the feedback information , and introduces the task attention compensation output and the historical residual compensation vector , and calculates the final reconstruction output ;
[0024] S29. Based on the final reconstruction output and the previous round of policy allocation execution data, construct a feature contribution analysis module, calculate the contribution score of each feature dimension in the final reconstruction output in policy generation, form a confidence vector , and use a confidence threshold to screen the features that meet to form a network state representation vector , where d is the total number of feature dimensions.
[0025] Optionally, the specific steps of S3 include:
[0026] S31. Based on the network state representation vector output by the masked autoencoder model, and combine the external conditional variable corresponding to the current mobile terminal , construct a combined input ;
[0027] S32. Input the combined input into the policy network , and output the probability distribution of IP address allocation actions in the current state. The actions represent candidate IP address allocation schemes;
[0028] S33. Input the combined input into the value network , estimate the expected value in the current state , and evaluate the quality of the policy in combination with the advantage function;
[0029] S34. Calculate the policy update ratio based on importance sampling ;
[0030] S35. Use generalized advantage estimation to construct an advantage function ;
[0031] S36. Construct a policy objective function containing a dynamic clipping coefficient ;
[0032] S37. Introduce a policy confidence interval adjustment module based on the masking auto - encoder perception mechanism, introduce upper and lower bound confidence constraints for each allocation action . The upper and lower bounds are jointly determined according to the feature masking entropy value of the network state representation vector and the key constraint indicators in the external condition variable . Construct a set of legal actions , and perform validity screening on the policy output actions, only retaining those that satisfy: ;
[0033] where is a dynamic threshold, jointly adjusted according to the masking reconstruction error and the policy stability is the finally selected optimal allocation action is the set of legal actions represents the probability of selecting the allocation action a under the given conditions of the network state representation vector and the external condition variable C;
[0034] S38. Use the backpropagation method to update the policy network parameters and the value network parameters respectively, minimize the policy objective function and maintain the stability of the policy confidence interval;
[0035] S39. Use the updated conditional policy network Generate the optimal action, that is, the optimal allocation scheme in the target IP address pool;
[0036] S310. Control the address scheduler to allocate the IP addresses in the target address pool to the mobile terminals associated with the current external condition variable according to the generated optimal action ;
[0037] S311. Record the allocation execution result, including the allocated IP address, terminal identifier, allocation delay, and allocation status information;
[0038] S312. Collect the network status feedback information during the execution of this allocation. The network status feedback information includes connection success rate, address pool occupancy rate, and allocation failure rate, and is used for the modeling and policy update of the autoencoder in the next allocation cycle.
[0039] Optionally, the S4 specifically includes:
[0040] S41. Collect the feedback metrics after the execution of the allocation action and construct an allocation execution result vector , where represents the IP allocation success flag, represents the normalized delay time required to complete the allocation, represents the quality of service matching degree, represents the improvement value of the current allocation behavior on the availability rate of the target address pool;
[0041] S42. Calculate the policy execution reward value based on the feedback metrics ;
[0042] S43. Introduce a sliding window mechanism to calculate the difference between the current policy execution reward value and the average policy execution reward value within the historical window , which is used to determine the stability of the policy behavior. When the absolute value of the difference exceeds the set threshold, trigger the steady-state adjustment mechanism to suppress drastic parameter changes;
[0043] S44. Integrate the current policy execution reward value with the estimated result of the advantage function to construct a weighted advantage function, and improve the response sensitivity and adaptive ability of the policy to the allocation result feedback during the optimization process;
[0044] S45. Use the weighted advantage function as the core input variable of the policy optimization objective function to guide the parameter update of the policy network during the policy optimization process, and limit the policy update amplitude through the confidence interval adjustment mechanism;
[0045] S46. Store the updated policy network parameter set, and write the current policy execution reward value, state representation vector, and feedback metrics into the experience replay pool for sampling during the training iteration process.
[0046] An intelligent IP address allocation system for mobile Internet access based on dynamic policies according to an embodiment of the present invention includes the following modules:
[0047] A data collection module, configured to collect access behavior data of mobile terminals, current network status information, and IP address pool usage data, and construct a multi-source feature dataset;
[0048] A masking modeling module, configured to perform masking processing on the multi-source feature dataset and generate a network status representation vector through a masked autoencoder;
[0049] A policy structure construction module, configured to initialize a policy structure based on the network status representation vector and external condition variables;
[0050] A policy optimization module, configured to receive the policy structure and the network status representation vector, and use the proximal policy optimization algorithm to generate and optimize an IP address allocation policy;
[0051] An allocation execution module, configured to perform IP allocation, and record the allocation execution result and network status feedback information;
[0052] A reward update module, configured to construct a reward function based on the allocation execution result, calculate a policy execution reward value, and use it to update the parameters of the policy network and the value network;
[0053] A model adaptation module, configured to receive network status feedback information and update the parameters of the masked autoencoder model.
[0054] The beneficial effects of the present invention are:
[0055] By constructing an intelligent allocation mechanism that integrates a masked autoencoder model and a proximal policy optimization algorithm, the present invention significantly improves the adaptability and robustness of the IP address allocation policy in a mobile Internet environment. Compared with the existing traditional methods that rely on static rules or single policy models, the multi-scale masking processing mechanism introduced in the present invention effectively alleviates the modeling deviation caused by missing, incomplete, and noise interference in multi-source feature data, and ensures the accuracy and generalization ability of network status representation; at the same time, the improved model based on the conditional decoder and graph structure embedding makes the encoding and decoding process more context-dependent and structure-aware, providing a more refined and dynamic state input basis for policy generation.
[0056] In terms of policy optimization, the present invention combines the guidance mechanism of external condition variables with the dynamic adjustment method of the policy confidence interval, enhancing the response ability of the allocation policy to environmental changes and the execution stability; the introduced allocation constraint regularization term further ensures the legality of the policy output and the controllability of resource scheduling, effectively reducing the occurrence probability of conflict allocation and abnormal decision-making. In addition, after the system executes the policy, it synchronously collects the allocation results and network feedback information, which are used for parameter updates of the masking model and the policy network, constructing a complete perception-decision-feedback-update closed loop.
[0057] Through the above collaborative optimization mechanism, the present invention realizes the full-process dynamic control of intelligent IP address allocation, not only significantly improving the accuracy and real-time performance of policy decision-making, but also enhancing the adaptability of the system to complex network conditions such as high-concurrency access and frequent state changes, with significant engineering application value and promotion prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification, and are used to explain the present invention together with the embodiments of the present invention, and do not constitute a limitation to the present invention. In the drawings:
[0059] Figure 1 is a flowchart of the intelligent IP address allocation method for mobile Internet access based on dynamic policies proposed by the present invention;
[0060] Figure 2 is a schematic structural diagram of the intelligent IP address allocation system for mobile Internet access based on dynamic policies proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] Now, the present invention will be further described in detail with reference to the drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0062] Refer to Figure 1 , the intelligent IP address allocation method for mobile Internet access based on dynamic policies includes the following steps:
[0063] S1. Collect the access behavior data of the mobile terminal, the current network status information, and the IP address pool usage data, construct a multi-source feature data set, and perform masking processing on the multi-source feature data set to form masked input data;
[0064] S2. Construct a masked autoencoder model, input the masked input data into the masked autoencoder model, perform encoding and decoding operations, and generate a reconstructed network state representation vector;
[0065] S3. Use the network status representation vector as input, apply the Proximal Policy Optimization (PPO) algorithm to generate and optimize the IP address allocation policy, control the address scheduler to allocate IP addresses from the target address pool to the corresponding mobile terminals according to the IP address allocation policy, and record the allocation execution results and network status feedback information;
[0066] S4. Calculate the policy execution reward value based on the allocation execution results, and use the policy execution reward value to update the policy parameters in the Proximal Policy Optimization algorithm;
[0067] S5. Input the network status feedback information into the masked autoencoder model, adjust the encoder and decoder parameters, and loop through steps S1 to S4 within a preset allocation period to achieve continuous optimization and allocation control of the IP address dynamic policy.
[0068] Through the construction of a closed-loop optimization process that includes a masked autoencoder model and a Proximal Policy Optimization algorithm, the present invention realizes the intelligent allocation control of mobile terminal IP addresses in a dynamic network environment, with significant practical application effects. First, three types of key data, namely access behavior, network status, and address pool usage, are collected and fused to form a high-dimensional multi-source feature set, and a masking mechanism is introduced to handle data missing and noise problems, effectively enhancing the robustness and generalization ability of the model. Second, by constructing an improved masked autoencoder model, the network status representation vector is accurately generated, improving the state perception accuracy of subsequent policy optimization. In the policy generation stage, the Proximal Policy Optimization algorithm is introduced to generate high-quality allocation policies driven by access conditions and network status, and through policy credibility adjustment, external condition guidance, and allocation constraint mechanisms, the stability and legality of policy output are ensured. At the same time, by recording the allocation execution results and feedback information, and calculating the reward signal to dynamically update the policy parameters, the iterative evolution of policy optimization is realized. During the operation of the system, the masked autoencoder model is adjusted in reverse through the feedback information to construct a collaborative optimization closed-loop of policy learning and state modeling. The overall process has the advantages of data-driven, accurate perception, efficient policy, and strong adaptability, significantly improving the utilization rate of IP address resources and the success rate of terminal access.
[0069] In this embodiment, the masking process of the multi-source feature dataset refers to performing a masking operation on some feature dimensions in the constructed multi-source feature dataset according to a masking ratio of 30%; the masking operation includes setting the position of the selected feature to zero, marking it as an invalid value, and replacing it with a mean substitution value; the adopted masking strategy is a hybrid strategy combining random masking and feature importance-guided masking, where 70% of the masking positions are determined by randomly selecting feature dimensions, and the remaining 30% of the masking positions are determined according to the influence weight ranking of the features in the prior evaluation, and the high-weight key features are preferentially masked.
[0070] In this embodiment, S2 specifically includes:
[0071] S21. Construct a masked input dataset for multi-source feature data , where is the number of samples, is the feature dimension, is the set of real numbers, and generate multiple masked versions for the masked input dataset , where the masking ratio ;
[0072] S22. Calculate the feature correlation matrix based on the co-variability between feature dimensions in the masked input dataset , where represents the statistical correlation coefficient between feature and feature . Construct a heterogeneous graph structure based on the feature pairs that meet the correlation threshold in the feature correlation matrix, where the node set represents the feature dimensions, and the edge set represents the connection relationship that meets . Calculate the structural importance score for each node using the graph attention mechanism;
[0073] S23. Perform a mixed masking operation on each masked version according to the structural importance score, including generating the masked positions respectively through random selection and graph-guided selection according to the ratio to form a masked input matrix ;
[0074] S24. Construct a dual-encoder structure, including a main encoder and an auxiliary encoder, extract the network state feedback information and policy execution records in the previous allocation cycle, and construct a historical state dataset. The main encoder receives the masked input matrix , and the auxiliary encoder receives the historical state data ;
[0075] S25. Input the masked input matrix and the historical state data into the main encoder and the auxiliary encoder respectively, and calculate the latent state vector using a fusion mechanism: ;
[0076] where is the fusion weight coefficient, is the main encoder function, is the auxiliary encoder function;
[0077] Latent state vector The practical significance of the formula lies in constructing a latent state representation method that integrates the current masked input information and historical feedback data to enhance the comprehensive modeling ability of the network state. In the actual IP address allocation scenario, relying solely on the network observation data at the current moment is likely to overlook the behavioral characteristics and policy adjustment impacts accumulated during the system evolution process. Therefore, the formula introduces a dual-encoder mechanism, which inputs the masked input data of the current cycle and the historical state information such as the allocation records and network feedback of the previous cycle into the main encoder and the auxiliary encoder respectively. Through the fusion mechanism, the output results of the two encoders are integrated to generate a latent state representation with stronger temporal continuity and semantic integrity, thereby providing a more accurate and context-sensitive state input for subsequent policy optimization and address allocation. This mechanism not only retains the important features of the current network state but also incorporates the potential impacts of historical behaviors on the system state evolution, significantly enhancing the model's perception ability and robustness in dynamic scenarios and facilitating the construction of a more reasonable and stable IP address allocation strategy.
[0078] S26. Construct a conditional decoder Collect external constraint information related to the current access policy, including network topology status, service level requirements, address pool tightness, and terminal device categories, and construct external conditional variables Concatenate the latent state vector with the external conditional variables and input them into the decoder for conditional reconstruction;
[0079] S27. Construct multiple structurally heterogeneous parallel paths in the decoder, including a fully connected decoding path, a residual skip connection path, and an attention-enhanced path, and output preliminary reconstruction results The construction of multiple structurally heterogeneous parallel paths means inputting the latent state vector into three decoding paths with different structures to enhance the model's expressive ability and reconstruction accuracy. First, the fully connected decoding path performs layer-by-layer non-linear transformation on the latent state vector through a multi-layer perceptron to achieve the basic reconstruction function. Second, the residual skip connection path introduces the residual information of the previous layer's output in each layer to alleviate the gradient vanishing problem in the deep decoding process and enhance the feature transfer efficiency. Finally, the attention-enhanced path uses the multi-head attention mechanism to assign different weights to different dimensions in the latent state vector, highlighting the feature dimensions important for the current task and outputting a more focused reconstruction result. The three paths are processed in parallel and finally converge their respective preliminary reconstruction results;
[0080] S28. Construct a policy feedback-guided fusion controller module where represents the feedback information of the previous round of policy execution, including the policy execution reward value and the IP allocation success rate and the remaining rate of the address pool , the fusion controller is based on the potential state vector and the feedback information to jointly generate the decoding path fusion weight , and introduce the task attention compensation output and the historical residual compensation vector to calculate the final reconstruction output : ;
[0081] wherein, is the output of the task compensation decoding channel, is the v-th conditional decoder, and are weight coefficients;
[0082] The final reconstruction output The practical significance of the formula is to dynamically regulate the output performance of the masked autoencoder by fusing the policy feedback information and the multi-path decoding results, so as to improve the accuracy and adaptability of the reconstructed network state representation. In the scenario of intelligent IP address allocation, the network state is complex and changeable, and the external conditions are dynamically changing. A single-structured decoding path often fails to fully capture the reconstruction requirements of multi-dimensional features. Therefore, this step designs a fusion control mechanism composed of a task compensation channel, a conditional decoding path, and a historical residual compensation vector. According to the historical feedback information of policy execution in the decoding stage, it dynamically generates the fusion weights of each path and guides the information complementarity of different paths. By introducing task attention compensation and historical residual information, it not only enhances the perception ability of the decoding result to the policy behavior change, but also improves the model's ability to recover key features. The finally output reconstructed vector can more comprehensively reflect the current network state and its evolution trend, providing a high-confidence input for subsequent policy generation and allocation control. This mechanism enhances the representation ability and task alignment performance of the masked autoencoder, is an important bridge connecting modeling and decision-making, and has good generalization ability and practical value.
[0083] S29. Based on the final reconstruction output and the data of the previous round of policy allocation execution, construct a feature contribution analysis module to calculate the contribution score of each feature dimension in the final reconstruction output to the policy generation , form a confidence vector , and use a confidence threshold to screen the features that meet to form the network state representation vector , where d is the total number of feature dimensions. The construction of the feature contribution analysis module to calculate the contribution score of each feature dimension in the final reconstruction output during policy generation specifically means that the feature contribution analysis module matches the final reconstruction output with the data of the previous round of policy allocation and execution, adopts the sensitivity analysis method or the gradient-based backpropagation mechanism to evaluate the influence degree of each feature dimension in generating policy actions, quantifies its influence value on policy selection by perturbing each feature dimension one by one and observing the change range of the policy output, and then assigns a contribution score to each feature. Finally, all the contribution scores are composed into a confidence vector.
[0084] The present invention realizes the deep modeling and enhanced expression of complex multi-source feature data by constructing a masked autoencoder model integrating multi-scale masking mechanism, graph structure embedding, dual encoder structure, conditional decoder and policy feedback guidance mechanism, significantly improving the accuracy and generalization ability of network state representation. First, a multi-version masking and graph structure-guided masking strategy is adopted, and a heterogeneous graph structure is constructed by combining the correlation between features, enhancing the model's ability to understand the dependence of feature structures and effectively suppressing the interference of redundant features. Secondly, a main and auxiliary dual encoder structure is introduced to parallelly model the masked input of the current cycle and the historical feedback information, and a more temporally aware latent state representation is generated through a fusion mechanism. In terms of decoder design, external conditions of the network are collected to construct conditional variables, driving parallel reconstruction of multi-structure heterogeneous paths and comprehensively improving the model's adaptability to the context. At the same time, indicators such as policy execution rewards and resource status are introduced through a feedback fusion controller to realize the dynamic adjustment of the fusion weights of the decoding paths, making the output more in line with the actual network scheduling requirements. Finally, the key features of the policy are screened through the feature contribution analysis module to ensure that the output network state representation vector has strong interpretability and policy drivability. This improved scheme enables the masked autoencoder model to have the advantages of structural adaptability, feedback perception and refined state expression, providing stable and high-quality input support for the policy module.
[0085] In this embodiment, S3 specifically includes:
[0086] S31. Based on the network state representation vector output by the masked autoencoder model , and combined with the external conditional variables corresponding to the current mobile terminal , construct a joint input ;
[0087] S32. Input the joint input into the policy network , and output the probability distribution of the IP address allocation action in the current state. The action represents a candidate IP address allocation scheme;
[0088] S33. The joint input Input to the value network , estimate the expected value in the current state , and evaluate the quality of the policy in combination with the advantage function. The expected value in the current state is obtained by forward propagation calculation by inputting a combined feature formed by concatenating a network state representation vector generated by a masked autoencoder and a corresponding external conditional variable into the value network. The value network predicts the long-term reward that the state may obtain under the current policy based on this input, so as to provide a numerical reference for policy evaluation and update, and further conduct quantitative evaluation and optimization guidance for the quality of the policy in combination with the advantage function;
[0089] S34. Calculate the policy update ratio based on importance sampling : ;
[0090] Among them, is the IP address allocation action output by the policy network at time step t, is the network state representation vector generated by the masked autoencoder model at time step t, is the external conditional variable at time step t, is the current policy network with parameters When, for the network state representation vector And external conditional variable Under the output action Probability, Is the parameter before update;
[0091] Policy update ratio The actual meaning of the formula is to introduce an importance sampling mechanism to dynamically measure the probability difference between the current policy and the old policy in specific allocation actions, so as to provide a robust estimation basis for the policy update process. In the actual IP address intelligent allocation task, the policy network will continuously iterate and update, and the action distribution generated by each policy may change slightly or significantly. If not controlled, it may lead to instability or performance degradation in the policy training process. By constructing the probability ratio of the current policy and the old policy to output a certain action under the same state and external conditions, this formula effectively captures the drift degree in the policy update process and is used as an importance correction term in policy optimization. This correction ratio not only helps to adjust the gradient weights of each sample in the objective function, but also ensures that the update direction is more in line with the old policy trajectory, improving the stability of policy iteration and the sample utilization efficiency. This mechanism is the core policy convergence control means in the proximal policy optimization method, especially suitable for communication resource allocation scenarios with sensitive distribution and changing states, effectively avoiding unreasonable IP allocation behaviors caused by too fast policy deviation, and improving the reliability and adaptability of model decision-making.
[0092] S35. Construct the advantage function using Generalized Advantage Estimation : ;
[0093] where is the discount factor, is the Generalized Advantage Estimation control factor, is the temporal difference error at time step t + l, is the temporal difference error at time step t, is the expected value under the next state and conditional variables, is the network state representation vector generated by the masked autoencoder model at time step t + 1, is the external conditional variable at time step t + 1;
[0094] The advantage function The practical significance of the formula is to more smoothly and stably evaluate the quality of the IP address intelligent allocation strategy in a specific state by introducing the Generalized Advantage Estimation mechanism, thereby improving the optimization effect of the policy network. In practical applications, the single-step temporal difference error is often easily affected by short-term feedback fluctuations, resulting in unstable policy updates. The Generalized Advantage Estimation effectively integrates short-term and long-term information by introducing the weighted accumulation of multi-step temporal difference errors and combining the value prediction of future states, providing a more robust advantage function. This mechanism not only reduces the jitter caused by high-variance estimation in policy training but also enables the policy model to more accurately judge the value difference of a certain IP allocation action relative to the current state, thereby more specifically optimizing the selection probability of high-value actions. Combining the network state representation and external conditional variables, the Generalized Advantage Estimation realizes policy review and evaluation in the time dimension, which helps to improve the policy learning efficiency and allocation effect in a dynamic environment and is an important technical link to support the stable convergence of the policy reinforcement learning module.
[0095] S36. Construct the policy objective function containing a dynamic clipping coefficient : ;
[0096] where represents the expectation operation for all samples at time step t, is the minimum operation, is the clipping operation, where is the dynamically adjusted policy trust region coefficient: ;
[0097] where and are coefficient hyperparameters, is the element variance of the network state representation vector , is the metric function related to the quality of service in the external condition variable C;
[0098] Policy objective function The practical significance of the formula lies in constructing an objective function that integrates the policy update objective and the dynamic confidence interval adjustment mechanism, which is used to guide the stable optimization of the IP address allocation policy. In policy reinforcement learning, the policy network is prone to large policy deviations due to gradient fluctuations or drastic state changes during update, which affects the model performance and allocation stability. By setting a dynamic clipping coefficient, this formula restricts the policy probability ratio, thereby ensuring that the policy update amplitude of each step remains within a controllable range and avoiding policy jump changes. At the same time, the clipping threshold is adaptively adjusted according to the variance of the network state representation and the quality of service requirements in the external condition variable, enabling the confidence interval to have the ability to perceive the state complexity and external constraint changes. While optimizing the policy reward, the overall objective function embeds a confidence adjustment term, making the update process take into account the security and robustness of the policy while pursuing performance improvement. This design makes the policy optimization process more flexible, refined and adaptable, especially suitable for the mobile Internet access scenario with frequent dynamic changes and complex constraint conditions, and is an important mechanism to improve the intelligence and reliability of the system's IP address allocation.
[0099] S37. Introduce a policy confidence interval adjustment module based on the masked autoencoder perception mechanism for each allocation action Introduce upper and lower bound confidence constraints, and the upper and lower bounds are jointly determined according to the feature masking entropy value of the network state representation vector and the key constraint indicators in the external condition variable to construct a set of legal actions , and perform validity screening on the policy output actions, only retaining those that satisfy: ;
[0100] where is the dynamic threshold, jointly adjusted according to the masking reconstruction error and the policy stability is the finally selected optimal allocation action is the set of legal actions represents the probability of selecting the allocation action a under the given conditions of the network state representation vector and the external condition variable C;
[0101] S38. Use the backpropagation method to update the policy network parameters and the value network parameters respectively, minimize the policy objective function and maintain the stability of the policy confidence interval;
[0102] S39. Use the updated conditional policy network Generate the optimal action, which is the optimal allocation plan in the target IP address pool;
[0103] S310. Control the address scheduler to allocate the IP addresses in the target address pool to the mobile terminals associated with the current external condition variable The allocation of the IP addresses in the target address pool to the mobile terminals associated with the current external condition variable The mobile terminals associated with the current external condition variable refer to that after the address scheduler receives the optimal action output by the policy network, according to the allocation plan indicated by this action, select the corresponding available IP addresses from the target address pool, and complete the exact matching according to the mobile terminal characteristics identified in the external condition variable;
[0104] S311. Record the allocation execution results, including the allocated IP addresses, terminal identifiers, allocation delays, and allocation status information;
[0105] S312. Collect the network status feedback information during this allocation execution process. The network status feedback information includes connection success rate, address pool occupancy rate, and allocation failure rate, which are used for the masking autoencoder modeling and policy update in the next allocation cycle.
[0106] The present invention constructs a policy generation and execution process that integrates the masking autoencoder perception mechanism and the proximal policy optimization structure, significantly improving the intelligent level of IP address dynamic allocation and resource scheduling efficiency. First, the joint modeling of the network state representation vector and the external condition variable is realized, improving the context awareness ability of policy generation; the collaborative evaluation of the policy network and the value network is combined to effectively enhance the fine discrimination ability of the quality of actions. Second, the importance sampling and generalized advantage estimation strategies are used to accurately calculate the update ratio and the advantage function, enhancing the stability and convergence efficiency of policy optimization. In terms of constructing the objective function, a dynamic clipping boundary based on state variance and quality of service constraints is introduced to achieve adaptive control of the policy update amplitude and avoid policy oscillation. A credible interval adjustment mechanism is constructed, integrating the masking reconstruction error and the policy confidence index, further imposing constraints on the action output space, effectively eliminating illegal or unstable allocation plans, and enhancing the system robustness. Finally, a closed-loop system of allocation - feedback - update is constructed to achieve efficient response to mobile terminal access requests and dynamic balance management of resources. The overall method takes into account accuracy, flexibility, and scalability, and can significantly improve the IP resource utilization rate and terminal access success rate in complex access scenarios.
[0107] In this embodiment, the S4 specifically includes:
[0108] S41. Collect the feedback metrics after the allocation action is executed, and construct the allocation execution result vector , where represents the IP allocation success flag, represents the normalized delay time required to complete the allocation, represents the quality of service matching degree, represents the improvement value of the current allocation behavior on the availability rate of the target address pool;
[0109] S42. Calculate the policy execution reward value based on the feedback metrics , which is defined as follows: ;
[0110] where , , , are preset weight coefficients;
[0111] The policy execution reward value The practical significance of the formula is to establish a comprehensive reward evaluation function for the immediate feedback in the policy optimization process, so as to quantify the execution effect of each IP address allocation behavior. By introducing feedback metrics in multiple dimensions, including whether the allocation is successful, the required delay, the quality of service matching degree, and the improvement value of the address pool resource utilization efficiency, this formula constructs an immediate reward value reflecting the overall performance of the policy execution. The indicators in different dimensions are weighted and fused through preset weight coefficients, enabling the reward function to flexibly adjust the importance of each indicator according to application requirements. This mechanism not only considers the direct effect of the allocation action itself, but also takes into account the comprehensive balance of system resource optimization and service experience, thus providing a clear and targeted feedback basis for the update of the subsequent policy network parameters. In a dynamic environment, this reward design with multi-index fusion helps to enhance the sensitivity and adaptability of the policy, improve the comprehensive performance of the intelligent allocation system under multi-task constraints, and is the key technical foundation for realizing the application of reinforcement learning in complex scenarios.
[0112] S43. Introduce a sliding window mechanism to calculate the difference between the current policy execution reward value and the average policy execution reward value within the historical window , which is used to judge the stability of the policy behavior. When the absolute value of the difference exceeds the set threshold, trigger the steady-state adjustment mechanism to suppress drastic parameter changes;
[0113] S44. Linearly combine the current policy execution reward value and the estimated result of the advantage function according to the preset weight coefficient to construct a weighted advantage function, so as to enhance the response sensitivity and adaptive ability of the policy to the allocation result feedback during the optimization process;
[0114] S45. Use the weighted advantage function as the core input variable of the policy optimization objective function to guide the parameter update of the policy network during the policy optimization process, and limit the policy update range through the credible interval adjustment mechanism;
[0115] S46. Store the updated policy network parameter set, and write the current policy execution reward value, state representation vector, and feedback index into the experience replay pool for sampling during the training iteration process.
[0116] The present invention realizes the efficient linkage between IP address allocation strategy and actual allocation effect by introducing a feedback-driven strategy reward evaluation and update mechanism, and effectively enhances the dynamic adaptability and stability of strategy optimization. First, the allocation execution result vector is constructed to comprehensively quantify key indicators such as allocation success rate, delay, service quality matching and resource utilization improvement, providing a multi-dimensional basis for reward generation. The real-time strategy execution reward function is designed based on the weighted summation method, and multiple feedback dimensions are integrated to ensure that the reward signal has accuracy and controllability. The sliding window mechanism is introduced to analyze the strategy fluctuation trend by using the difference between the current and historical reward values, and the steady-state control is automatically triggered when the strategy changes drastically to avoid overfitting and model oscillation. At the same time, the real-time reward and advantage function are integrated to construct a weighted advantage signal to improve the sensitivity and responsiveness of the strategy to feedback. The weighted advantage signal is further introduced into the main process of strategy optimization, and the credible interval restriction mechanism is cooperated to ensure that the update process is both efficient and robust. Finally, a feedback memory mechanism is constructed through the experience replay pool to continuously accumulate allocation experience samples to support long-term strategy learning and generalization ability improvement. The overall mechanism ensures that the strategy can continuously evolve and be robustly optimized, and achieves efficient and reliable IP resource scheduling in complex access scenarios.
[0117] In this implementation manner, S5 specifically includes:
[0118] After each round of IP address allocation is completed, the real-time network status feedback information, including connection success rate, address pool occupancy rate and allocation failure rate, is input into the masked autoencoder model as the basis for model update. By adjusting the encoder and decoder parameters, the model can more accurately reconstruct the representation vector reflecting the current network status. The update process is performed periodically within the set allocation cycle, so that the model can continue to adapt to the dynamic changes of the network environment. At the same time, relying on the updated state representation vector, the generated IP address allocation strategy will be more timely and accurate, thereby realizing dynamic optimization and closed-loop control of the IP allocation strategy, and improving the resource allocation efficiency and stability of the system in complex mobile access environments.
[0119] refer to Figure 2 , the mobile Internet access IP address intelligent allocation system based on dynamic strategy includes the following modules:
[0120] The data collection module is used to collect the access behavior data of the mobile terminal, the current network status information and the IP address pool usage data, and construct a multi-source feature data set;
[0121] The mask modeling module is used to perform mask processing on multi-source feature datasets and generate network state representation vectors through mask autoencoders;
[0122] A policy structure building module is used to initialize the policy structure based on the network state representation vector and external condition variables;
[0123] A policy optimization module is used to receive the policy structure and the network state representation vector, and generate and optimize the IP address allocation policy using a proximal policy optimization algorithm;
[0124] Allocation execution module, used to perform IP allocation, record allocation execution results and network status feedback information;
[0125] The reward update module is used to construct a reward function based on the distribution execution results, calculate the policy execution reward value, and update the policy network and value network parameters;
[0126] The model adaptation module is used to receive network status feedback information and update the masked autoencoder model parameters.
[0127] Embodiment 1:
[0128] In order to verify the feasibility of the present invention in implementation, the present invention is applied to an IPv6 campus wireless network environment deployed in a certain university campus. Student terminals (such as mobile phones, tablets, and notebooks) frequently access in peak hours (8 am to 10 am, 8 pm to 10 pm), resulting in congestion in IP address allocation. Some devices have failed to reconnect multiple times, seriously affecting the use of the teaching platform and the online examination experience. The traditional static IP address allocation strategy cannot adapt to the real-time changes in terminal behavior, resulting in resource waste and reduced service quality.
[0129] To this end, the "intelligent allocation method for mobile Internet access IP addresses based on dynamic strategies" proposed in this paper is deployed, focusing on modeling the access state in combination with the masked autoencoder, and introducing the proximal strategy optimization algorithm to achieve dynamic allocation adjustment. In the experimental scenario, the system collects multi-source data including access behavior logs, network topology changes, terminal category labels, and address pool usage, uses a 30% masking ratio to process feature inputs, and builds a graph to guide the masking mechanism to mine the structural association between features, thereby enhancing the accuracy of state modeling.
[0130] During the policy optimization process, the model automatically reconstructs the state representation of the autoencoder every 30 minutes, and combines external conditional variables such as the service level of the current terminal and the occupancy rate of the address pool to output the optimal allocation action that meets the dynamic confidence interval. The system records the changes in key indicators before and after implementation, and selects 10 typical sets of terminal access data for comparative analysis.
[0131] Table 1 Comparison table of IP address allocation performance between traditional policy and the dynamic policy of the present invention
[0132] It can be seen from the data in Table 1 that in terms of the success rate of IP address allocation, the method of the present invention reached an average of 98.8% in 10 samples, which is significantly better than the average of 93.9% of the traditional method. Especially in sample numbers 1, 2, and 6, the present invention achieved a 100% success rate, showing extremely high stability. This shows that the present invention enhances the state perception ability through the autoencoder and effectively adjusts the IP allocation decision by combining the proximal policy optimization mechanism, which can significantly reduce the probability of allocation failure.
[0133] In terms of the average allocation delay, the method of the present invention took an average of 37.1 milliseconds, which is nearly 32% shorter than the 54.6 milliseconds of the traditional method. For example, in sample 5, the present invention completed the allocation in only 33 ms, while the traditional method took 59 ms. This reflects that the present invention improves the allocation efficiency through the optimization of the policy network structure and the dynamic action screening mechanism, and has strong real-time response capabilities.
[0134] In terms of service quality matching degree, the method of the present invention reached an average of 91.5% in all samples, higher than 81.3% of the traditional scheme. In samples 8 and 10, the matching degrees of the present invention were 94% and 96% respectively, far exceeding 75% and 79% of the traditional method, indicating that the present invention can better identify and match the service requirements of terminal devices, reflecting its adaptability advantage for heterogeneous devices.
[0135] In terms of the improvement value of the available rate of the address pool, the average improvement value of the method of the present invention is 7.1%, while the traditional method is only 2.4% on average. In sample 9, the improvement value is as high as 9.5%, far higher than 3.2% of the traditional method, which shows that the present invention can effectively alleviate the shortage of IP address resources and improve resource utilization efficiency through the dynamic allocation policy and the state feedback mechanism.
[0136] In summary, from the multi-dimensional performance comparison results, it can be seen that the present invention is superior to the traditional method in terms of allocation success rate, allocation delay, service matching, and resource utilization efficiency, verifying its significant practical value and promotion prospects in complex mobile network environments.
[0137] As described above, it is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.
Claims
1. An intelligent allocation method for mobile Internet access IP addresses based on dynamic policies, characterized in that It includes the following steps: S1. Collect the access behavior data of the mobile terminal, the current network status information, and the IP address pool usage data, construct a multi-source feature dataset, and perform a masking process on the multi-source feature dataset to form masked input data; S2. Construct a masked autoencoder model, input the masked input data into the masked autoencoder model, perform encoding and decoding operations, and generate a reconstructed network status representation vector; S3. Use the network status representation vector as the input, apply the proximal policy optimization algorithm to generate and optimize the IP address allocation policy, control the address scheduler to allocate IP addresses from the target address pool to the corresponding mobile terminals according to the IP address allocation policy, and record the allocation execution result and the network status feedback information; S4. Calculate the policy execution reward value based on the allocation execution result, and use the policy execution reward value to update the policy parameters in the proximal policy optimization algorithm; S5. Input the network status feedback information into the masked autoencoder model, adjust the encoder and decoder parameters, and loop through steps S1 to S4 within a preset allocation period to achieve continuous optimization and allocation control of the IP address dynamic policy.
2. The intelligent allocation method of mobile Internet access IP addresses based on dynamic policies according to claim 1, characterized in that The masking process on the multi-source feature dataset means that in the constructed multi-source feature dataset, according to a masking ratio of 30%, masking operations are performed on some feature dimensions in the multi-source feature dataset; where the masking operations include setting the positions of the selected features to zero, marking them as invalid values, and replacing them with mean substitution values; the masking strategy adopted is a hybrid strategy combining random masking and feature importance-guided masking, where 70% of the masking positions are determined by randomly selecting feature dimensions, and the remaining 30% of the masking positions are determined according to the influence weight ranking of the features in the prior evaluation, and the high-weight key features are preferentially masked.
3. The intelligent allocation method of mobile Internet access IP addresses based on dynamic policies according to claim 1, characterized in that The specific content of S2 includes: S21. Construct a masked input data set for multi-source feature data , where is the number of samples, is the feature dimension, is the set of real numbers, and generate multiple masked versions for the masked input data set , where the masking ratio ; S22. Calculate the feature correlation matrix based on the co-variability between each feature dimension in the masked input dataset , where represents the feature and the feature 's statistical correlation coefficient. Construct a heterogeneous graph structure based on the feature pairs that meet the correlation threshold in the feature correlation matrix, where the node set represents the feature dimension, and the edge set represents the connection relationship that meets . Use the graph attention mechanism to calculate the structural importance score for each node; S23. For each masked version perform a mixed masking operation according to the structural importance score, including in proportion generate the masked positions respectively through random selection and graph-guided selection to form a masked input matrix ; S24. Construct a dual-encoder structure, including a main encoder and an auxiliary encoder, extract the network state feedback information and policy execution records in the previous allocation cycle, and construct a historical state data set. The main encoder receives the masked input matrix , and the auxiliary encoder receives the historical state data ; S25. Respectively input the masked input matrix and the historical state data into the main encoder and the auxiliary encoder, and use the fusion mechanism to calculate the latent state vector ; S26. Construct a conditional decoder , collect external constraint information related to the current access policy, including network topology status, service level requirements, address pool tightness, and terminal device categories, and construct an external conditional variable , and concatenate the potential state vector with the external conditional variable , and input it into the decoder for conditional reconstruction; S27. Build multiple structurally heterogeneous parallel paths in the decoder, including a fully connected decoding path, a residual skip connection path, and an attention enhancement path, and respectively output preliminary reconstruction results ; S28. Construct a policy feedback-guided fusion controller module , where represents the feedback information of the previous round of policy execution, including the policy execution reward value , the IP allocation success rate and the remaining rate of the address pool . The fusion controller jointly generates the decoding path fusion weight based on the latent state vector and the feedback information , and introduces the task attention compensation output and the historical residual compensation vector to calculate the final reconstruction output ; S29. Based on the final reconstruction output Construct a feature contribution analysis module with the execution data of the previous round of policy allocation, and calculate the contribution score of each feature dimension in the final reconstruction output during policy generation to form a confidence vector , and use a confidence threshold to filter out features that satisfy to form a network state representation vector , where d is the total number of feature dimensions.
4. The intelligent allocation method of mobile Internet access IP address based on dynamic policy according to claim 3, characterized in that, The specific content of S3 includes: S31. Based on the network state representation vector output by the masked autoencoder model , and combining with the external condition variables corresponding to the current mobile terminal , construct a joint input ; S32. Input the combined input into the policy network , and output the probability distribution of the IP address allocation action in the current state. The action represents a candidate IP address allocation scheme; S33. Input the combined input into the value network , estimate the expected value in the current state , and evaluate the quality of the strategy in combination with the advantage function; S34. Calculate the policy update ratio based on importance sampling ; S35. Construct the advantage function using generalized advantage estimation ; S36. Construct a policy objective function including a dynamic cropping coefficient ; S37. Introduce a policy confidence interval adjustment module based on the masked autoencoder perception mechanism for each allocation action Introduce upper and lower bound confidence constraints, where the upper and lower bounds are jointly determined according to the feature masked entropy value of the network state representation vector and the key constraint indicators in the external condition variable to construct a set of legal actions , and perform validity screening on the policy output actions, only retaining those that satisfy: ; Among them, is the dynamic threshold, which is jointly adjusted according to the masking reconstruction error and the policy stability, is the finally selected optimal allocation action, is the set of legal actions, represents the probability of selecting the allocation action a under the given conditions of the network state representation vector and the external condition variable C; S38. Update the policy network parameters and the value network parameters respectively using the backpropagation method, minimize the policy objective function, and maintain the stability of the policy trust interval; and the value network parameters , minimize the policy objective function and maintain the stability of the policy trust interval; S39. Use the updated conditional policy network Generate the optimal action, which is the optimal allocation plan in the target IP address pool; S310. Control the address scheduler to allocate the IP addresses in the target address pool to the mobile terminal associated with the current external condition variable according to the generated optimal actions. associated with; S311. Record the allocation execution result, including the allocated IP address, terminal identifier, allocation delay, and allocation status information; S312. Collect the network status feedback information during this allocation execution process, where the network status feedback information includes the connection success rate, address pool occupancy rate, and allocation failure rate, and is used for masked autoencoder modeling and policy update in the next allocation cycle.
5. The intelligent allocation method of mobile Internet access IP addresses based on dynamic policies according to claim 1, characterized in that The specific content of S4 includes: S41. Collect the feedback metrics after the allocation action is executed, and construct an allocation execution result vector , where represents the IP allocation success flag, represents the normalized delay time required to complete the allocation, represents the service quality matching degree, represents the improvement value of the current allocation behavior on the availability rate of the target address pool; S42. Calculate the policy execution reward value based on the feedback metrics ; S43. Introduce a sliding window mechanism to calculate the difference between the current policy execution reward value and the average policy execution reward value within the historical window, which is used to determine the stability of policy behavior. When the absolute value of the difference exceeds the set threshold, trigger the steady-state adjustment mechanism to suppress drastic parameter changes; S44. Integrate the current policy execution reward value with the estimated result of the advantage function to construct a weighted advantage function, and enhance the response sensitivity and adaptive ability of the policy to the allocation result feedback during the optimization process; S45. Use the weighted advantage function as the core input variable of the policy optimization objective function, guide the parameter update of the policy network during the policy optimization process, and limit the policy update amplitude through a confidence interval adjustment mechanism; S46. Store the updated policy network parameter set, and write the current policy execution reward value, status representation vector, and feedback metrics into the experience replay pool for sampling during the training iteration process.
6. The intelligent IP address allocation system for mobile Internet access based on dynamic policies is applied to the intelligent IP address allocation method for mobile Internet access based on dynamic policies according to any one of claims 1 to 5, and is characterized in that, It includes the following modules: A data collection module, which is used to collect the access behavior data of the mobile terminal, the current network status information, and the IP address pool usage data, and construct a multi-source feature dataset; The masking modeling module is used to perform masking processing on the multi-source feature dataset and generate a network state representation vector through a masked autoencoder; The policy structure construction module is used to initialize the policy structure based on the network state representation vector and external conditional variables; The policy optimization module is used to receive the policy structure and the network state representation vector, and generate and optimize the IP address allocation policy by using the proximal policy optimization algorithm; The allocation execution module is used to perform IP allocation and record the allocation execution result and network state feedback information; The reward update module is used to construct a reward function based on the allocation execution result, calculate the policy execution reward value, and use it to update the policy network and value network parameters; The model adaptation module is used to receive the network state feedback information and update the masked autoencoder model parameters.
Citation Information
Patent Citations
Multi-type terminal random access competition solution based on near-end strategy optimization
CN117241409A
Acid gas capture system by heat exchange optimization having absorbent recovery line
KR1020250001989A
Method for allocating transmission resources using reinforcement learning
US20190124667A1