6G multi-access intelligent QoS unified framework and dynamic strategy execution method
By building a unified QoS mapping engine and an AI policy decision-maker, the problem of unified cross-domain QoS in the 6G multi-access environment is solved, realizing intelligent, rapid scheduling and adaptability of cross-domain resources, and improving the service reliability and efficiency of the 6G network.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-30
- Publication Date
- 2026-03-27
AI Technical Summary
The existing QoS guarantee mechanism lacks a unified mapping standard in the 6G multi-access environment, which leads to a "cliff-like" drop in service quality when services switch across domains. Traditional strategies cannot respond to dynamic changes in real time, and the distributed decision-making architecture results in complex signaling interaction and slow response, making it difficult to achieve unified and efficient scheduling of cross-domain resources.
A unified QoS mapping engine is built to normalize heterogeneous parameters. Combined with an AI policy decision-maker, semantic parsing and deep reinforcement learning are performed to generate dynamic QoS policy instructions. Enhanced user plane functional entities are used to achieve unified scheduling of cross-domain resources and programmable data plane technology for scheduling and forwarding of service flows.
It achieves unified and comparable cross-domain QoS, enhances the intelligence and adaptability of policy generation, improves the efficiency and reliability of cross-domain resource scheduling, and ensures the overall resilience and service reliability of 6G networks in complex multi-access scenarios.
Smart Images

Figure CN121751371A_ABST
Abstract
Description
[0001] TECHNICAL FIELD The present application relates to the technical field of information data, in particular to a 6G multi-access intelligent QoS unified framework and dynamic policy execution method.
[0002] BACKGROUND With the evolution of 5G-Advanced to 6G, network access forms are changing from single cellular networks to deep integration of ground cellular, non-3GPP (Wi-Fi) and non-ground networks (satellites); Under this background, emerging businesses such as immersive XR, holographic communication and industrial Internet of Things have made unprecedentedly high demands on bandwidth, latency and reliability. However, the existing QoS guarantee mechanism is facing severe challenges. Different access domains (cellular, Wi-Fi, satellite) use independent QoS parameter systems (5QI, EDCA parameters), lack of unified mapping standards, resulting in a "cliff-like" decline in service quality when switching across domains; Secondly, traditional QoS strategies rely on manual predefinition and static rules, which cannot real-time perceive and respond to dynamically changing business intentions and network states; The distributed decision-making architecture of the control plane leads to complex signaling interaction and slow response, making it difficult to achieve unified and efficient scheduling of cross-domain resources. Although existing technologies attempt to solve some problems, they do not address the complete scenario of 6G multi-access integration, nor do they achieve intention analysis and closed-loop optimization based on artificial intelligence. Therefore, they cannot meet the demand for intelligent, continuous and end-to-end QoS guarantee in the 6G era.
[0003] SUMMARY In order to solve the problems existing in the prior art, the present application aims to provide a 6G multi-access intelligent QoS unified framework and dynamic policy execution method.
[0004] The 6G multi-access intelligent QoS unified framework and dynamic policy execution method described in the present application comprises: S101, the user terminal reports the business demand and terminal context information, at the same time, the cellular access network element, the non-3GPP interworking function entity and the non-ground network gateway respectively collect and report the network state parameters of the access domain to which they belong, forming a cross-domain heterogeneous data set; S102, through the unified QoS mapping engine, the multi-dimensional heterogeneous QoS parameters in the cross-domain heterogeneous data set are normalized and dynamically weighted, generating a globally unified QoS level descriptor; S103, the AI policy decision maker performs semantic analysis on the business request of the user terminal to obtain the QoS constraint condition, combines the globally unified QoS level descriptor, calculates the optimal resource allocation action through a deep reinforcement learning model, and generates a structured QoS policy instruction; S104, the enhanced user plane function entity parses the QoS policy instruction, converts it into configuration commands for each access network element, and classifies, schedules and programs the forwarding path of the service flow based on programmable data plane technology, for unified scheduling and QoS guarantee across domains; S105, the probe device deployed in each access network element and the enhanced user plane function entity collects end-to-end key performance indicator data of the service flow in real time and reports it through the feedback channel; S106, the AI policy decision maker uses the reported key performance indicator data to perform online incremental learning and dynamic reward function correction on the deep reinforcement learning model, and updates and outputs the optimized QoS policy instruction.
[0005] Preferably, in step S101, the terminal collects service demand and context information, and structures the information by means of a preset data template to form a standardized terminal data set. Based on this, network status and parameters are obtained from cellular access network elements, interworking function entities and non-terrestrial network gateways, integrated into a cross-domain network status set through a unified interface, further classified by a support vector machine algorithm, and the priority of each domain data is identified. Key network parameters are extracted, and domains exceeding the threshold are marked. The source and scope of abnormal state are located by cross-domain correlation analysis, and resources are dynamically adjusted according to the preset scheduling rules to form an optimized allocation scheme. Whether the service demand is met is determined according to the real-time state evaluation, and the network operation state is determined.
[0006] Preferably, in step S102, the service quality parameters are obtained from multiple heterogeneous network domains by a cross-domain data collection module, and a standardized template is used for preliminary arrangement to form a structured parameter set. Based on this, a mapping engine is used to normalize each domain parameter to generate a uniform scale parameter value set. According to a dynamic weight distribution mechanism, the parameters are assigned weights, and parameters exceeding the threshold are marked with priority to obtain a weighted parameter combination. Based on the combination, the influence degree of each parameter is evaluated according to the global service quality standard, the priority is sorted, and the parameters are classified by a support vector machine to determine their service quality level, generating a global unified quality service level identifier. The identifier is used to monitor the cross-domain service quality in real time, and an automatic adjustment mechanism is triggered for domains that do not meet the standard to update the service state.
[0007] Preferably, in the step S103, first, the policy decision maker performs semantic analysis on the service request of the user terminal, extracts the core content and implicit constraints, and determines the basic demand range. Based on the analysis result, the quality of service requirement is classified by combining the preset level descriptor, the quality demand category is determined, the resource allocation scheme is calculated by using the deep reinforcement learning model, the deployment priority is determined, the scheme is compared with the threshold value, the resources exceeding the limit are marked, the marked resource list is formed, the list is converted into a formatted instruction, the execution order is adjusted according to the demand category, the final policy instruction set is generated, the inconsistent instructions are reordered by matching the terminal request type, the optimized instruction list is obtained, and finally the automatic distribution tool responds in real time and delivers to the terminal to complete the request processing whole process.
[0008] Preferably, in the step S104, the enhanced user plane function entity first analyzes the policy instruction set layer by layer, decomposes the specific operation for the access network element, forms a preliminary configuration command group, refines the service flow classification based on the command group by using the programmable data plane technology, determines the priority order of each service flow, plans the resource occupation scheme by using the scheduling algorithm according to the priority, dynamically adjusts the forwarding path, generates the path guidance information, implements flow sharing for the resource nodes exceeding the load based on the cross-domain resource pool data, obtains the optimized resource distribution, verifies the configuration command according to the quality of service guarantee requirement, and automatically delivers to each access network element to complete the scheduling and path programming of the service flow.
[0009] Preferably, in the step S105, the probe is deployed in the access network element and the user plane function entity to collect the end-to-end indicators of the service flow, form a preliminary data record, then clean the data according to the preset rule, remove the abnormal and redundant data, obtain the clean service flow data set, extract the key performance indicators from the data set, classify the high-priority data by using the support vector machine, and transmit the marked data through the feedback channel. Further analyze the indicator fluctuation of the marked data, locate the service flow segment with significant fluctuation, and track the distribution of the service flow segment in the network element and the function entity to determine whether there is local congestion or performance bottleneck. If it is confirmed that there is, the related access network element flow is adjusted through the preset scheduling rule, and finally the network operation state is optimized.
[0010] Preferably, in step S106, based on the collected performance data, first, the key indicators are preliminarily processed, the core data set is extracted, and a preliminarily sorted indicator set is formed, a deep reinforcement learning model is used to analyze the current service quality state, if the state is lower than the preset threshold, an online updating mechanism is triggered, the model parameter range to be adjusted is determined, on this basis, incremental training is implemented, the model weight is updated, and an optimized model version is obtained, a reward function adjustment basis is further extracted, if the output deviates from the expected target, the reward calculation rule is dynamically corrected, a service quality policy instruction is generated according to the new rule, a specific scheduling scheme is formed after matching the current system state, after the scheme is executed, the results are combined with the feedback process, the execution effect is evaluated through real-time closed-loop data, at the same time, the abnormal points in the analysis feedback are recorded, and the optimization record for the strategy decision maker to refer to is generated.
[0011] The 6G multi-access intelligent QoS unified framework and dynamic policy execution method described in the application has the advantages that the unity and comparability of cross-domain QoS are realized, the QoS parameters of heterogeneous access domains are normalized into global QoS level descriptors (G-QoS Index) by constructing a unified QoS mapping engine, the problem of parameter incomparability between cellular, Wi-Fi, and satellite networks is fundamentally solved, and a solid foundation is laid for seamless continuous experience of services in a multi-access environment; The intelligence and adaptability of the strategy generation are improved by introducing an AI strategy decision maker that integrates semantic analysis and deep reinforcement learning, which can automatically and accurately convert user's natural language service intention into QoS strategy in real time, overcoming the rigidity of traditional static configuration and realizing accurate and dynamic matching between service demand and network resources; The efficiency and reliability of the strategy execution are enhanced by the centralized architecture of the enhanced user plane function (UPF+) and the programmable data plane (SRv6), which realizes unified scheduling and millisecond-level fast path switching of cross-domain resources, avoids the decision delay and inconsistency problem of distributed architecture, and significantly improves the response speed and execution efficiency of the system; The system can dynamically optimize the AI decision model based on real-time network performance data through incremental learning, ensuring that the QoS guarantee strategy continuously improves itself with changes in network environment and service demand, and comprehensively improving the overall resilience and service reliability of 6G networks in complex multi-access scenarios. BRIEF DESCRIPTION OF DRAWINGS
[0012] Figure 1 is the flow of the 6G multi-access intelligent QoS unified framework and dynamic policy execution method described in the application Figure 1 ; Figure 2 is the flow of the 6G multi-access intelligent QoS unified framework and dynamic policy execution method described in the application Figure 2 . Detailed Implementation
[0013] like Figures 1-2 As shown in the figure, this application describes a unified intelligent QoS framework and dynamic policy execution method for 6G multi-access.
[0014] like Figures 1-2 As shown, in step S101, the user terminal reports service requirements and terminal context information. At the same time, cellular access network elements, non-3GPP interoperability function entities, and non-terrestrial network gateways collect and report network status parameters of their respective access domains to form a cross-domain heterogeneous dataset.
[0015] Furthermore, in step S101, service requirements and context information are collected through the user terminal, and the collected information is processed in a structured manner using a pre-established data template to obtain a standardized terminal dataset. Based on the standardized terminal dataset, the corresponding network status and network parameters are obtained from cellular access network elements, interconnection function entities and non-terrestrial network gateways, and integrated using a unified data interface to determine the cross-domain network status set. For cross-domain network state sets, heterogeneous datasets of each access domain are obtained, and the support vector machine algorithm is used to classify the data of different domains and determine the priority distribution of the data of each domain. Based on the priority distribution, key network parameters are extracted from the heterogeneous dataset. If the network parameters of a certain domain exceed the preset threshold, the data in that domain are labeled to obtain the labeled dataset. For the labeled dataset, cross-domain data correlation analysis is used to obtain the distribution of abnormal network states and determine the source domain and scope of influence of abnormal data; Based on the source domain and impact range of the abnormal data, network resources are dynamically adjusted using preset scheduling rules to obtain an optimized resource allocation scheme. By optimizing the resource allocation scheme, we can obtain real-time network status updates for each access domain, determine whether the standards for business requirements are met, and ultimately determine the network operating status.
[0016] Specifically, in step S101, during the process of a user terminal reporting service requirements and terminal context information, it is assumed that a user terminal requests a high-definition video stream service through a 5G network. Its service requirement is a minimum bandwidth guarantee of 2Mbps. Meanwhile, the context information includes the terminal location (latitude and longitude of 39.9042, 116.4074), device type (smartphone supporting 5G NR), and current battery level (60% remaining). The terminal uploads these data in JSON format to the unified data management function (UDM) of the core network through the API interface, and the data transmission adopts the AES-128 encryption algorithm to ensure security; Next, the cellular access network element (gNB base station) collects network state parameters, such as the current cell load is 75%, the signal-to-noise ratio (SNR) is 20 dB, and through analysis it is found that the high load may affect user experience, so it is necessary to dynamically adjust resource allocation, and a load balancing-based scheduling algorithm is used to offload part of the user data flow to the adjacent cell (the load is only 40%); Subsequently, the non-3GPP interworking function entity (N3IWF) collects data of the Wi-Fi access domain, such as the number of users in the access point coverage is 50, and the average throughput is 15 Mbps, and through regression analysis it is predicted that the throughput may drop to 10 Mbps in the next 10 minutes, so part of the users need to be switched to the cellular network in advance to ensure service quality; Finally, the non-terrestrial network gateway (NTN Gateway) collects satellite network parameters, such as signal delay of 250 ms and bandwidth utilization of 60%, and through time series analysis algorithm it is predicted that the delay may increase to 300 ms during peak period, so the system automatically adjusts the data routing strategy to preferentially transmit low-delay-demand business through the ground network; After these cross-domain data are summarized, a heterogeneous data set is formed, stored in the edge computing node, and the K-means clustering algorithm is used for classification processing of the data, the user demand is matched with the network state, and an optimized resource allocation scheme is generated, for example, 2.5 Mbps bandwidth is allocated to the user to ensure smooth video playback, and delay-sensitive business is routed to the cellular network with higher SNR, forming a complete closed-loop logic from data collection to optimization.
[0017] As shown in Figures 1-2 Step S102, through the unified QoS mapping engine, the multi-dimensional heterogeneous QoS parameters in the cross-domain heterogeneous data set are normalized and dynamically weighted to generate a globally unified QoS level descriptor.
[0018] Further, in step S102, through the cross-domain data collection module, quality service related parameters are obtained from multiple heterogeneous network domains, and the collected parameters are preliminarily sorted using a pre-established standardized template to obtain a structured parameter set; According to the structured parameter set, the mapping engine is used to normalize the parameters of each domain, and for different domain heterogeneous parameters, a unified scale conversion rule is used to determine the normalized parameter value set; Through the normalized parameter value set, for the dynamic weight distribution mechanism, each parameter is given a corresponding weight value, and if the weight value of a certain parameter exceeds a preset threshold, it is marked with a priority, and a weighted parameter combination is obtained; According to the weighted parameter combination, a global unified service quality evaluation standard is obtained, the parameters marked are prioritized, the influence of each parameter on the overall quality of service is judged, and a sorted parameter list is determined; Through the sorted parameter list, the parameters are classified by using a support vector machine algorithm, and according to the classification result, the service quality levels of different categories of parameters are obtained, and a classified level distribution is obtained; According to the classified level distribution, the quality service identifier of each level is mapped, and the corresponding global unified quality service level identifier is generated through the preset level identifier rule; Through the generated global unified quality service level identifier, the service quality in the cross-domain data is monitored in real time, and if the service quality level of a certain domain is lower than the preset standard, an automatic adjustment mechanism is triggered, and an updated service quality state is obtained.
[0019] Specifically, in step S102, when the multi-dimensional heterogeneous QoS parameters are processed by the unified QoS mapping engine in the cross-domain heterogeneous data set, it is assumed that the system receives QoS parameter data from different access domains, such as a delay of 50 ms and a throughput of 10 Mbps for a cellular network, a delay of 80 ms and a throughput of 8 Mbps for a non-3GPP Wi-Fi network, and a delay of 200 ms and a throughput of 5 Mbps for a satellite network; Firstly, the system uses a normalization algorithm to map these parameters to a unified range of 0 to 1, and calculates that the delay of the cellular network is normalized to 0.25 and the throughput is 1.0, the delay of the Wi-Fi network is normalized to 0.4 and the throughput is 0.8, and the delay of the satellite network is normalized to 1.0 and the throughput is 0.5; Next, the system dynamically allocates weights according to the business type, for example, for real-time communication business, the delay weight is set to 0.7 and the throughput weight is set to 0.3, and through weighted calculation, the comprehensive QoS score of the cellular network is 0.475, the Wi-Fi network is 0.52, and the satellite network is 0.85; Subsequently, the system analyzes these scores by using a decision tree algorithm, and generates a global unified QoS level descriptor in combination with the business demand (the delay of real-time communication needs to be less than 60 ms), and marks the cellular network as "high priority", the Wi-Fi network as "medium priority", and the satellite network as "low priority"; To ensure logical coherence, the system is further associated with an edge computing task distribution scenario, if the edge node needs to process real-time data flow, the "high priority" network is preferentially selected for transmission, and for batch data upload, the "medium priority" network can be selected, forming a complete link from parameter processing to business adaptation.
[0020] AsFigures 1-2 As shown, in step S103, the AI policy decision maker performs semantic analysis on the service request of the user terminal to obtain the QoS constraint condition, combines the globally unified QoS level descriptor, calculates the optimal resource allocation action through the deep reinforcement learning model, and generates a structured QoS policy instruction.
[0021] Further, in step S103, the service request issued by the user terminal is preliminarily analyzed by the policy decision maker, the core content in the request is decomposed by using semantic disassembly technology, the constraint condition implied therein is obtained, and the basic demand range of the request is determined; According to the constraint condition obtained by analysis, the quality requirement in the request is classified and processed in combination with the pre-established level descriptor, to obtain the classified quality demand category; For the classified quality demand category, a deep reinforcement learning model is used to calculate the resource allocation to generate a preliminary allocation scheme and determine the priority order of resource allocation; By comparing the generated allocation scheme with the preset threshold range, if the allocation proportion of a certain resource exceeds the threshold, the resource is marked to obtain a marked resource list; According to the marked resource list, the formatted instruction content is generated, the execution order of the instruction is adjusted for different categories of quality demand, and the final policy instruction set is determined; The final policy instruction set is obtained, and the request type of the user terminal is matched and processed, if a certain instruction is inconsistent with the terminal request, the instruction is reordered to obtain an optimized instruction list; Through the optimized instruction list, the terminal request is processed in real time, and the instruction content is distributed to the corresponding terminal by using an automatic distribution tool to complete the request processing flow.
[0022] Specifically, in step S103, in the cross-domain heterogeneous network environment, the AI policy decision maker first performs semantic analysis on the service request initiated by the user terminal, extracts the QoS constraint condition, for example, a certain video conference application request requires that the delay is not more than 100 ms and the bandwidth is not less than 5 Mbps, the system performs word segmentation and intent recognition on the request text by using a natural language processing algorithm, and generates a structured constraint condition vector [delay: 100 ms, bandwidth: 5 Mbps]; The system combines the existing globally unified QoS level descriptor database, which contains priority labels of different network types, such as "preferred level" for fiber network and "standard level" for 4G network, and matches the user request with the network labels in the database through a vector matching algorithm to calculate the matching score, the fiber network score is 0.9 and the 4G network score is 0.6; The system inputs the matching result into a deep reinforcement learning model, taking user satisfaction and resource utilization as the reward function, and the model optimizes the resource allocation action through multiple iterations. Assuming that the available bandwidth of the fiber network in the current resource pool is 8 Mbps and that of the 4G network is 3 Mbps, the model calculates that the action probability of allocating the fiber network is 0.85 and that of the 4G network is 0.15, and finally selects the fiber network as the optimal resource; The system generates a structured QoS policy instruction containing parameters such as network type allocation "fiber", bandwidth limit 6 Mbps, and delay guarantee 80 ms, and issues the instruction to the network scheduling module. To form a complete logical chain, the system also associates the policy instruction with subsequent business monitoring scenarios. If it is detected that the delay during a video conference exceeds 90 ms, the bandwidth adjustment mechanism is automatically triggered to ensure business continuity. Through the above process, the system realizes full automation from request analysis to resource allocation.
[0023] As shown in Figures 1-2 Step S104, the enhanced user plane function entity parses the QoS policy instruction, converts it into configuration commands for each access network element, and classifies, schedules, and programs the forwarding path of the business flow based on programmable data plane technology for unified scheduling and QoS guarantee across domains.
[0024] Further, in step S104, the enhanced user plane function entity parses the policy instruction set layer by layer, decomposes the specific operation content for the access network element, and obtains the preliminary configuration command group; According to the preliminary configuration command group, combined with the programmable data plane technology, the business flow classification is refined to determine the priority order of different business flows; Using the results of the priority order, the scheduling arrangement method is used to allocate and plan the business flow, and the corresponding resource occupation scheme is obtained; Through the resource occupation scheme, the forwarding path is dynamically adjusted, and the path guidance information for each access network element is generated; According to the path guidance information, combined with the data of the cross-domain resource pool, unified coordination is implemented to determine whether the load of a certain resource node exceeds the preset threshold, and if so, traffic sharing processing is performed to obtain the adjusted resource distribution state; Get the adjusted resource distribution state, and according to the requirements of the quality of service, verify the configuration command group to determine the command content issued to the access network element; Through the final verification of the command content, it is automatically issued to each access network element to complete the scheduling and path programming operation of the business flow.
[0025] Specifically, in step S104, in the cross-domain network resource management, after the enhanced user plane function entity first receives the QoS policy instruction, it is decomposed into specific configuration parameters by the built-in instruction parsing engine, for example, the bandwidth requirement of a certain service flow in the instruction is 4.5 Mbps, and the upper limit of the delay is 50 ms. The system converts the instruction text into a structured data matrix [bandwidth: 4.5 Mbps, delay: 50 ms, priority: high] using a syntax analysis algorithm, and generates configuration commands for different access network elements according to the matrix content. Assuming that the core network element needs to allocate 3.2 Mbps bandwidth and the edge network element needs to allocate 1.3 Mbps, the system calculates the core network load proportion to be 70% and the edge network element to be 30% by using a weight distribution algorithm; Based on programmable data plane technology, the system classifies and processes service flows, identifies service flow types using deep packet inspection technology, and marks them as "real-time flows". The system uses a priority queue algorithm for scheduling, and real-time flows are allocated to a high-priority queue. The queue buffer size is set to 200KB to avoid packet loss. At the same time, the system analyzes the current network congestion index, which is 0.75, and automatically adjusts the forwarding path. The delay of path A is calculated to be 40ms and the delay of path B is calculated to be 60ms. Path A is selected as the optimal forwarding path to ensure that the delay is less than 50ms. The system uses a software-defined network controller to uniformly schedule cross-domain resources and sends configuration commands to each network element. The core network element feedbacks that the bandwidth occupancy rate is 68% and the edge network element feedbacks that the bandwidth occupancy rate is 29%. The system dynamically adjusts the allocation proportion based on the feedback data to ensure that the resource utilization rate is above 90%. To form a complete logical chain, the system associates the service flow scheduling result with the subsequent network performance monitoring module. If the delay of path A rises to 48ms, the preloading mechanism of backup path B is automatically triggered to ensure the continuity of the service flow and the QoS guarantee. Through the above automatic process, the system realizes the whole-process optimization from instruction parsing to resource scheduling.
[0026] As shown in Figures 1-2 Step S105, probe devices deployed in each access network element and enhanced user plane function entity collect end-to-end key performance indicator data of service flows in real time and report them through a feedback channel.
[0027] Further, in step S105, probe devices are deployed in access network elements and user plane function entities to obtain end-to-end indicator information of service flow data and obtain preliminary collected data records. For the preliminary collected data records, a preset filtering rule is used for cleaning to eliminate abnormal values and redundant information, and a cleaned service flow data set is determined. According to the cleaned service flow data set, key performance related index data is extracted, and support vector machine algorithm is used for classification processing of the index data to judge the priority category of each index; If there is a high priority category in the classified index data, the part of the data is transmitted after being marked with priority through a feedback channel, and a marked high priority data stream is obtained; According to the marked high priority data stream, the fluctuation of its end-to-end index is analyzed, and the service flow segment corresponding to the index item with large fluctuation is determined; For the service flow segment corresponding to the index item with large fluctuation, its distribution in the access network element and user plane function entity is obtained, and whether there is local network congestion or performance bottleneck is judged; If it is judged that there is local network congestion or performance bottleneck, the related access network element is adjusted by the preset scheduling rule, and the optimized network running state is obtained.
[0028] Specifically, in step S105, in the probe equipment deployed in each access network element and enhanced user plane function entity, the system first installs probe software on each network element node through an automatic script, covering at least 95% of the key nodes to ensure the comprehensiveness of data collection; The probe equipment collects end-to-end key performance index data of service flow in real time, such as delay, packet loss rate and throughput. Assuming that a service flow collects average delay of 12.5 milliseconds, packet loss rate of 0.3%, and throughput of 800Mbps within 5 minutes, the data is preliminarily processed by the built-in sampling algorithm (100 times sampling per second to take the average value) to reduce the noise influence; Then, the system uses time series analysis algorithm to detect anomalies in the collected data, sets the delay threshold to 15 milliseconds, and marks it as an anomaly point if it exceeds the threshold, and calculates the duration of the anomaly, for example, a certain anomaly lasts for 30 seconds, the system automatically records and generates an alarm log; Subsequently, the probe equipment reports the processed data and alarm information to the central analysis platform through the preconfigured feedback channel (dedicated port based on UDP protocol), and the reporting frequency is set to once per minute to ensure that the data transmission bandwidth occupancy does not exceed 50Kbps, and the data packet size is controlled within 1KB to reduce network burden; During the reporting process, the system uses AES-128 encryption algorithm to ensure data security and prevent information leakage, and verifies data integrity through checksum mechanism. If the data loss rate is found to exceed 0.1%, the retransmission mechanism is triggered to ensure data reliability; To form a logical chain, if the delay of a certain service flow is continuously abnormal, the system will automatically associate the upstream node data of the service flow to analyze whether it is caused by upstream congestion. For example, when the upstream node throughput drops to below 500 Mbps, it is inferred to be a bottleneck point, and an optimization suggestion is generated and pushed to the operation and maintenance system to form a closed-loop processing.
[0029] As shown in Figures 1-2 Step S106, the AI policy decision maker uses the reported key performance indicator data to perform online incremental learning and reward function dynamic correction on the deep reinforcement learning model, updates and outputs the optimized QoS policy instruction.
[0030] Further, in step S106, the key indicators are preliminarily processed based on the collected performance data, and the core data set for subsequent analysis is extracted to obtain the preliminarily sorted indicator set. According to the preliminarily sorted indicator set, a deep reinforcement learning model is used to analyze the data and determine the state of the current service quality. If the state is below the preset threshold, an online update mechanism is triggered to determine the range of model parameters that need to be adjusted. For the determined range of model parameters, incremental training operation is implemented to obtain the latest training result, update the internal weights of the deep reinforcement learning model, and obtain the optimized model version. From the optimized model version, the adjustment basis of the reward function is extracted. If the output of the reward function deviates from the expected target, it is dynamically adjusted to determine the new reward calculation rule. According to the new reward calculation rule, the corresponding service quality policy instruction is generated, matched with the current system state, and the specific scheduling scheme is obtained. Through the execution of the scheduling scheme, the execution result is combined with the feedback process to obtain real-time closed-loop optimization data, and whether the execution effect meets the expected target is determined to determine whether the policy instruction needs to be further adjusted. After obtaining the closed-loop optimization data, the abnormal points in the feedback process are recorded and analyzed to generate improvement basis for the policy decision maker, and the final optimization record is obtained.
[0031] Specifically, in step S106, the AI policy decision maker realizes closed-loop optimization through a series of technical means, and the specific implementation method is as follows: The system collects key performance indicator data in real time, such as network delay of 50 ms, throughput of 10 Gbps, and packet loss rate of 0.5%. These data are automatically uploaded to the decision maker database through API interface, preprocessed by time series analysis algorithm, and 5-minute moving average value is calculated to smooth fluctuations to obtain stable delay mean value of 48 ms, which provides reliable basis for subsequent analysis. Based on these data, the system performs online incremental learning on the deep reinforcement learning model, adopts the DQN algorithm, updates the Q value table through the current state (delay, throughput) and historical data, sets the learning rate to 0.01 and the discount factor to 0.9, and after 1000 iterations, the model converges to a loss value of 0.02, outputting more optimal policy parameters; The system dynamically corrects the reward function, sets the delay weight to 0.6 and the throughput weight to 0.4 according to business needs, calculates the comprehensive reward value, for example, the current reward is 85 points, an increase of 5 points from the previous time, indicating that the strategy improvement is effective, and the corrected function is automatically updated to the model; The system outputs the optimized QoS strategy instructions according to the learning results, for example, adjusts the bandwidth allocation ratio from 60% to 70%, and issues them to network devices through the SDN controller, realizing automatic scheduling; The feedback link monitors the performance data after execution, for example, the delay is reduced to 45 ms, the system automatically compares the predicted value with the actual value, calculates the error rate of 6.25%, and if the error exceeds the threshold of 10%, a new round of learning is triggered, forming a closed loop; To ensure the integrity of the logical chain, the system also adjusts the business priority, if the delay continues to be lower than 50 ms, the proportion of high-priority business traffic is automatically increased to 80%, to optimize resource utilization; The above processes are all completed automatically by the system, with data-driven and algorithm coordination, ensuring continuous optimization of the strategy.
[0032] For those skilled in the art, various corresponding changes and modifications can be made to the above-described technical solutions and concepts, and all such changes and modifications should be within the scope of protection of the claims of the present application.
Claims
1. A unified intelligent QoS framework and dynamic policy execution method for 6G multi-access, characterized in that, include: S101. The user terminal reports service requirements and terminal context information. At the same time, cellular access network elements, non-3GPP interoperability function entities and non-terrestrial network gateways collect and report network status parameters of their respective access domains to form a cross-domain heterogeneous dataset. S102. Through a unified QoS mapping engine, the multidimensional heterogeneous QoS parameters in the cross-domain heterogeneous dataset are normalized and dynamically weighted to generate a globally unified QoS level descriptor. S103. The AI policy decision-maker performs semantic parsing on the user terminal's service requests to obtain QoS constraints. Combined with the globally unified QoS level descriptor, it calculates the optimal resource allocation action through a deep reinforcement learning model and generates structured QoS policy instructions. S104. The enhanced user plane function entity parses the QoS policy instructions, converts them into configuration commands for each access network element, and classifies, schedules and programs the forwarding paths of service flows based on programmable data plane technology for unified scheduling and QoS guarantee of cross-domain resources. S105. Probe devices deployed on each access network element and enhanced user plane functional entity collect end-to-end key performance indicator data of service flow in real time and report them through the feedback channel. S106. The AI policy decision-maker uses the reported key performance indicator data to perform online incremental learning and dynamic correction of the reward function on the deep reinforcement learning model, and updates and outputs the optimized QoS policy instructions.
2. The intelligent QoS unified framework and dynamic policy execution method for 6G multi-access as described in claim 1, characterized in that, In step S101: Business requirements and context information are collected through user terminals, and structured processing is performed using preset data templates to form a standardized terminal dataset. Based on the standardized terminal dataset, network status and parameters are obtained from cellular access network elements, non-3GPP interoperability functional entities and non-terrestrial network gateways, and integrated into a cross-domain network status set through a unified interface. The support vector machine algorithm is used to classify the heterogeneous data in the cross-domain network state set and identify the priority of data in each domain; Extract key network parameters, label domains that exceed a preset threshold, and obtain the labeled dataset; Cross-domain correlation analysis is used to locate the source and scope of impact of abnormal network states; network resources are dynamically adjusted according to preset scheduling rules to form an optimized resource allocation scheme. The final network operating status is determined by assessing whether the real-time network status meets business requirements.
3. The intelligent QoS unified framework and dynamic policy execution method for 6G multi-access as described in claim 1, characterized in that, In step S102: QoS parameters are obtained from multiple heterogeneous network domains through the cross-domain data acquisition module, and are initially organized using standardized templates to form a structured parameter set; The mapping engine is used to normalize the parameters of each domain, generating a set of parameter values with a uniform scale. Each parameter is assigned a weight based on a dynamic weight allocation mechanism, and parameters that exceed the threshold are marked with priority to obtain a weighted parameter combination. Based on the weighted parameter combination, the impact of each parameter is evaluated according to the global service quality standard, and the parameters are prioritized. The support vector machine algorithm is used to classify the parameters, determine their service quality level, and generate a globally unified QoS level descriptor. Based on the QoS level descriptor, cross-domain service quality is monitored in real time, and an automatic adjustment mechanism is triggered for domains that do not meet the standards to update the service quality status.
4. The intelligent QoS unified framework and dynamic policy execution method for 6G multi-access as described in claim 1, characterized in that, In step S103: The AI strategy decision-maker performs semantic parsing on the user terminal's business requests, extracts core content and implicit constraints, and obtains the parsing results. Based on the parsing results, the service quality requirements are classified according to the preset QoS level descriptors to determine the quality requirement categories; A deep reinforcement learning model is used to deduce a resource allocation scheme and determine the allocation priority. Resources that exceed the threshold are marked to obtain a list of marked resources; The marked resource list is converted into formatted instructions, and the execution order is adjusted according to the demand category to obtain the final strategy instruction set; Match the user terminal request type, reorder inconsistent instructions, and obtain an optimized instruction list; The automatic distribution tool sends instructions to the terminal in real time to complete the request processing.
5. The intelligent QoS unified framework and dynamic policy execution method for 6G multi-access as described in claim 1, characterized in that, In step S104: The enhanced user plane function entity parses the QoS policy instruction set layer by layer, decomposes it into specific operations for each access network element, and forms a preliminary configuration command group; Based on the initial configuration command group, programmable data plane technology is used to refine the business flow classification and clarify the priority order of each business flow; Based on priority order, a resource allocation scheme is planned through a scheduling algorithm, and forwarding paths are dynamically adjusted to generate path guidance information; By combining cross-domain resource pool data, traffic sharing is implemented on resource nodes that have exceeded their load limits to obtain an optimized resource distribution. The configuration commands are verified according to the service quality assurance requirements and automatically sent to each access network element to complete the scheduling and path programming of the service flow.
6. The intelligent QoS unified framework and dynamic policy execution method for 6G multi-access as described in claim 1, characterized in that, In step S105: By deploying probe devices on access network elements and user plane functional entities, end-to-end key performance indicator data of service flows are collected to obtain preliminary data records; The data is cleaned according to preset rules, outliers and redundant information are removed, and a clean business flow dataset is obtained. Key performance indicators are extracted from the cleanroom workflow dataset, and the support vector machine algorithm is used for classification to identify high-priority data. High-priority data is marked and transmitted via a feedback channel; Analyze the fluctuations in the tagged data, locate the service flow segments with significant fluctuations, and track their distribution in network elements and functional entities; Determine if there is local network congestion or performance bottleneck. If so, adjust the traffic of relevant access network elements through preset scheduling rules to optimize network operation.
7. The intelligent QoS unified framework and dynamic policy execution method for 6G multi-access as described in claim 1, characterized in that, In step S106: Based on the reported key performance indicator data, the key indicators are preliminarily processed, the core dataset is extracted, and a preliminary set of indicators is formed. A deep reinforcement learning model is used to analyze the current service quality status. If the status is lower than a preset threshold, an online update mechanism is triggered to determine the range of model parameters that need to be adjusted. Perform incremental training, update the model weights, and obtain an optimized model version; Extract the basis for adjusting the reward function; if the output deviates from the expected target, dynamically correct the reward calculation rules. QoS policy instructions are generated based on the new rules, and a specific scheduling scheme is formed after matching the current system status. After executing the scheduling plan, the results are combined with the feedback process, and the execution effect is evaluated through real-time closed-loop data. Record and analyze the anomalies in the feedback, generate optimization records that can be used as a reference for the strategy decision-maker, and form a complete iterative optimization closed loop.
8. The intelligent QoS unified framework and dynamic policy execution method for 6G multi-access as described in claim 1, characterized in that, The unified QoS mapping engine generates a globally unified QoS level descriptor called G-QoS Index, which is used for the comparison and seamless mapping of QoS parameters between cellular, Wi-Fi and satellite networks in a multi-access environment.
9. The intelligent QoS unified framework and dynamic policy execution method for 6G multi-access as described in claim 1, characterized in that, The programmable data plane technology includes the SRv6 protocol, which is used for unified scheduling of cross-domain resources and millisecond-level fast path switching.
10. The intelligent QoS unified framework and dynamic policy execution method for 6G multi-access as described in claim 1, characterized in that, The deep reinforcement learning model employs the DQN algorithm.