Traffic scheduling congestion control method based on transmission rate and available bearer capacity of vehicle networking access point
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-15
- Publication Date
- 2026-08-11
AI Technical Summary
[0006]针对上述问题,本发明提供一种基于传输速率与车联网接入点可用承载量的业务调度拥塞控制方法,旨在解决高速移动环境下5G与RSU异构网络中,因业务需求与实时容量失配导致的频繁切换、资源分配低下及网络拥塞问题
[0033](1) In step one, the present invention performs a refined three-class classification of the service flow by means of latency and bandwidth sensitivity, and combines a binarized indicator function to overcome the defect of the existing scheme that directly uses multi-source high-dimensional original data for fusion, which leads to the expansion of the state space. It realizes the discretized structure dimensionality reduction input based on the prior characteristics of the service, accurately perceives the spatial location and network layer state changes without increasing the feature dimension, and greatly improves the algorithm convergence speed and decision real-time performance in high dynamic scenarios.
Smart Images

Figure CN122554889A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of wireless communication and intelligent transportation technology, specifically involving network congestion control and service scheduling technology in the environment of heterogeneous vehicular networks (VANETs), and more specifically, a service scheduling congestion control method based on transmission rate and available capacity of VANET access points. Background Technology
[0002] With the development of intelligent transportation and autonomous driving, multi-layered heterogeneous vehicle-to-everything (V2X) networks composed of 5G base stations and RSUs are gradually evolving into a key architecture supporting intelligent mobility. In high-speed mobile environments, the network faces dual challenges: First, vehicles frequently traverse overlapping areas formed by 5G and multiple roadside nodes during high-speed travel. The spatial heterogeneity and dynamism of these areas lead to frequent vertical handovers and channel fluctuations, severely impacting service continuity. Second, service types are significantly differentiated, encompassing diverse needs ranging from low-rate text control (capacity-first) to high-rate multimedia video (rate-first). Due to the lack of coordinated optimization between transmission rate requirements and remaining node capacity, resource competition easily induces network congestion, leading to increased service blocking rates and difficulty in guaranteeing the quality of service required for ultra-reliable, low-latency communication. Therefore, how to achieve joint optimization of service scheduling and congestion control in heterogeneous and dynamic environments has become a pressing technical challenge.
[0003] The application CN202210221079.5 discloses "A Joint Computation Offloading and Resource Allocation Method Based on Multi-Agent DDQN". This method effectively alleviates resource contention and computational latency issues in wireless edge computing networks by utilizing multi-agent dual-deep Q-networks to collaboratively optimize computation offloading decisions and wireless channel resource allocation, thereby improving system resource utilization and energy efficiency. However, this method mainly focuses on the joint balance of general computing and channel resources, lacking proactive and forward-looking quantification of the available capacity of network nodes for high, medium, and low-rate service types in heterogeneous vehicular network scenarios. Furthermore, it fails to effectively address the "dimensional explosion" problem in the reinforcement learning state space that inevitably occurs when the physical layer spatial dynamic features and the multi-dimensional resource states of the network layer are directly spliced and stacked in high-speed mobile environments. This can easily lead to slower convergence or excessive computational overhead in the complex spatial locations and network layer states of vehicular networks, making it difficult to guarantee ultra-reliable and low-latency vehicular network service quality from the source.
[0004] The document with application number "CN202511806756.X" discloses "A Distributed Wireless Network User Handover Control Method Based on Deep Reinforcement Learning." This method uses deep reinforcement learning algorithms to guide users' vertical handover decisions based on dynamic channel quality or basic handover overhead, effectively solving the problem of frequent handovers and service interruptions in wireless networks during mobility, and improving the stability and continuity of mobile connections. However, this method mainly focuses on passively tracking the channel state. The algorithm does not explicitly use the real-time remaining available capacity of each access point for different rate service types as the core decision criterion. It also ignores the hard blocking penalty caused by rigidly executing the original action in high-speed mobile scenarios with sudden high traffic, which leads to the saturation of the specific resource pool of the preferred target network node. This can easily lead to a mismatch between resource supply and rate demand during network handover or access, causing widespread hard blocking.
[0005] As can be seen from the aforementioned publicly available documents, existing technologies still have limitations in congestion control for heterogeneous vehicle-to-everything (V2X) networks: 1. Existing solutions mostly rely on a single channel state or resource offloading, making it easy for sudden high-bandwidth services to cause source overload and frequent handovers; 2. When dealing with the dynamic nature of the physical layer space and the resource state of the network layer, existing reinforcement learning algorithms often use direct feature splicing without dimensionality reduction to simply merge high-dimensional spatiotemporal data with multi-layer network loads, lacking refined dimensionality reduction representation of service features, which can easily lead to a "dimensionality explosion" in the state space of deep reinforcement learning networks and slow down real-time decision-making; 3. Existing allocation and handover logic usually performs rigid blocking (such as denying access or forcibly dropping calls) when the preferred node is saturated, which is difficult to effectively resolve the hard blocking problem in sudden high-traffic scenarios. Summary of the Invention
[0006] To address the aforementioned problems, this invention provides a service scheduling congestion control method based on transmission rate and available capacity of vehicle network access points. It aims to solve the problems of frequent handover, low resource allocation, and network congestion caused by the mismatch between service demand and real-time capacity in heterogeneous 5G and RSU networks under high-speed mobile environments.
[0007] To achieve the above objectives, the technical solution provided by this invention is: a service scheduling congestion control method based on transmission rate and available capacity of vehicle network access points, comprising the following steps:
[0008] Step 1: Constructing environmental perception input: Obtain vehicle location, network coverage of each access point, and real-time available capacity for different service types in heterogeneous vehicle networking scenarios. Then, construct a multi-dimensional extended state space that includes physical layer location perception features and network layer resource status, and use it as the input for subsequent intelligent service scheduling environmental perception.
[0009] Step 2: Define the decision output: Define a discrete action space that includes network access actions and cross-node vertical handover actions, and use it as the output of the decision;
[0010] Step 3: Select actions from the discrete action space, design a joint utility function and access matching strategy, generate reward values to evaluate the merits of access or handover actions, and use these rewards as the basis for updating network parameters in the online optimization of the deep Q network.
[0011] First, for the candidate actions in step two, by setting differentiated weights for service transmission rate guarantee and available capacity utilization, actions are selected from the discrete action space based on the current state to obtain the service available capacity.
[0012] Then, by combining the transmission rate and the utilization rate of available capacity, a utility function is constructed as an evaluation criterion.
[0013] Finally, an access matching strategy is introduced when the carrying capacity of a specific service is saturated, and a reward value is generated to evaluate the merits of access or switching actions.
[0014] Step 4: Model building, model optimization, and execution scheduling:
[0015] First, the multidimensional extended state space constructed in step one is used as input, the discrete action space defined in step two is used as the decision output set, and the result of the joint utility function designed in step three is used as the reward signal to construct the DQN learning model.
[0016] Secondly, through the experience replay mechanism and network parameter updates, the deep Q network agent is driven to learn the optimal business scheduling strategy online until convergence.
[0017] Finally, the DQN learning model outputs and executes scheduling instructions to implement real-time network congestion control.
[0018] Furthermore, in step one above, the business is refined and dimensionality reduced into three categories: low-speed, medium-speed, and high-speed businesses.
[0019] The method for calculating real-time available capacity is as follows: when a service of a certain rate successfully accesses the corresponding network node, the capacity status value of the corresponding network node is incremented by 1.
[0020] Furthermore, in step two above, the discrete action space includes access actions and handover actions, and the action instructions within the handover action space are determined by a binarization factor. It means that among them :
[0021] when When this occurs, it means the system has accepted the initial access request from the target vehicle, or allowed it to perform a cross-node switch from the original node to the target node. At this time, the available service capacity of the corresponding node will be reduced accordingly.
[0022] when At this time, it means the system rejects the access request or maintains the vehicle's current network connection.
[0023] Furthermore, in step three above, the specific steps for constructing the utility function are as follows:
[0024] (1) Calculate the maximum transmission rate that the target vehicle can obtain if it connects to the corresponding network node;
[0025] (2) Real-time statistics of the remaining available capacity of each network node for low, medium and high speed services;
[0026] Based on (1) and (2) above, the formula for constructing the joint utility function is as follows:
[0027] in, For business type average transmission rate For business type Available carrying capacity and These are the rate weighting coefficient and the capacity weighting coefficient, respectively set for services with different rates. This is the redistribution penalty coefficient triggered based on the access matching policy.
[0028] Furthermore, in step three above, the specific access matching strategy when the service capacity is saturated is as follows:
[0029] When the capacity of low-speed services is saturated, overflow allocation is sequentially made to the medium-speed and high-speed service resource pools.
[0030] When the capacity of medium-speed services is saturated, the excess resources are allocated to the high-speed service resource pool.
[0031] A blocking penalty occurs when no resource pool can be matched.
[0032] Compared with existing methods, the beneficial effects of the present invention are:
[0033] (1) In step one, the present invention performs a refined three-class classification of the service flow by means of latency and bandwidth sensitivity, and combines a binarized indicator function to overcome the defect of the existing scheme that directly uses multi-source high-dimensional original data for fusion, which leads to the expansion of the state space. It realizes the discretized structure dimensionality reduction input based on the prior characteristics of the service, accurately perceives the spatial location and network layer state changes without increasing the feature dimension, and greatly improves the algorithm convergence speed and decision real-time performance in high dynamic scenarios.
[0034] (2) This invention overcomes the technical barrier of existing wireless network scheduling schemes that optimize "rate guarantee" and "available capacity" in isolation. This invention breaks through the limitation of existing technologies that pursue high throughput or low handover overhead in isolation. It constructs a joint utility reward function from transmission rate requirements and node available capacity, and guides the deep Q network parameters to actively align with the conflicting relationship between service rate guarantee and node resource hard boundary through differentiated weights, thereby realizing the coordinated optimization of vehicle service quality and heterogeneous vehicle network available capacity.
[0035] (3) This invention avoids the technical limitation of existing algorithms that use "saturation rigid hard blocking" when nodes are fully loaded, which leads to large-scale hard blocking. In step three, this invention innovatively embeds an "access matching strategy" when the carrying capacity of a specific service is saturated. When the preferred resource pool is fully loaded, it uses conditional branching logic to perform cross-rate dynamic overflow diversion buffering, which effectively avoids rigid hard blocking.
[0036] (4) This invention first obtains a low-dimensional extended state space perception input through the refined dimensionality reduction classification and binarization update logic of the multi-source service characteristics of the Internet of Vehicles. Such state perception input can be further enhanced by setting service differentiation weights and cross-rate dynamic overflow access matching strategies, so that the action selection boundary and elastic benchmarking ability of the agent in the network node saturation environment are enhanced, and the reward signal oscillation caused by rigid saturation hard blocking is smoothed. Then, it is fed into the designed deep Q network model containing online Q network and target Q network for online optimization, so as to obtain the optimal adaptive scheduling strategy that takes into account both transmission rate guarantee and load balancing. Compared with the joint resource allocation method and user switching control method, it realizes the synergistic maximization of vehicle service quality and heterogeneous resource utilization, effectively reduces the request blocking rate, and is conducive to realizing real-time congestion control of the Internet of Vehicles, realizing fast and robust real-time network congestion control. Attached Figure Description
[0037] Figure 1 This is a flowchart of the present invention;
[0038] Figure 2 This is a schematic diagram illustrating an applicable scenario for an example of the present invention;
[0039] Figure 3 Design a flowchart for the network congestion control algorithm of DQN;
[0040] Figure 4 Average reward during the training phase;
[0041] Figure 5 This is a comparison chart of the blocking rates of the DQN mechanism of this invention with three other strategies at different vehicle speeds. Detailed Implementation
[0042] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. Obviously, the described embodiments are merely some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0043] This invention is further described in detail with reference to specific implementation methods and accompanying drawings. The specific steps are as follows:
[0044] like Figure 2 As shown, the heterogeneous vehicle-to-everything (V2X) unit scenario constructed in this embodiment consists of 5G base stations and roadside units, forming a multi-layer network architecture with single coverage areas and overlapping coverage areas. During each transmission time interval, the system monitors vehicles on the road in real time. Based on the vehicle's location coordinates and movement trends, they are divided into three dynamic sets: a pre-access vehicle set (vehicles in areas without network coverage that are about to enter a coverage area), a switching vehicle set (vehicles in overlapping heterogeneous networks that need to perform network switching), and a stable operating vehicle set (vehicles that have established connections and do not currently require switching).
[0045] See Figure 1 The present invention provides a service scheduling congestion control method based on service transmission rate and available capacity of access network nodes, which specifically includes the following steps:
[0046] Step 1: Constructing Environment-Aware Input: This step constructs a heterogeneous network unit scenario and initializes a multi-dimensional extended state space (as the environment-aware input). Vehicle services are categorized into high, medium, and low sensitivity based on latency and bandwidth. A binarization function is used to convert the available capacity of each node into a scalar, thereby achieving dimensionality reduction and quantization of the vehicle-to-everything (V2X) environment state. This serves as the input feature for the DQN model. Specifically, it includes:
[0047] (1) Based on the sensitivity of typical Internet of Vehicles services to transmission rate and the requirements of network carrying capacity, the complex services are refined and reduced to three categories: low-speed services (text): low speed requirements, but continuous connection is required; medium-speed services (voice): a balance between speed and carrying capacity requirements; high-speed services (multimedia video): extremely high instantaneous throughput requirements.
[0048] In this embodiment, based on the sensitivity of typical vehicle-to-everything (V2X) services to transmission rate... and the degree of dependence on the available network capacity. Weights were assigned to different business types, as shown in Table 1. Representative business type
[0049] Table 1 Weights of Different Business Types
[0050]
[0051] (2) Real-time acquisition of the vehicle's current location coordinates The system includes physical coverage boundary indication information for 5G base stations and each roadside unit (RSU), as well as real-time carried data for the three types of services simultaneously extracted from each node.
[0052] By combining the vehicle's current location coordinates with the physical coverage boundaries of each network node, a multidimensional state space vector is formed. Specifically, the physical layer location perception features obtained in (1) and (2) above are aligned and fused with the network layer resource state data to construct a multidimensional extended state space vector, and the reduced-dimensional state space is output as the environmental perception state for the next step.
[0053] Within each transmission time interval, when the service management module detects that a service request has been accepted by the system and a communication bearer for the corresponding network node has been established, it determines that an access event has occurred and assigns a value of 1 to the binarization indicator function for the corresponding network node and service type. If the service request is rejected or the original connection is maintained, the binarization indicator function is assigned a value of 0. This construction method enables the model to accurately perceive the dynamic spatial location and network layer state changes of heterogeneous vehicle networks without adding useless dimensions.
[0054] Step 2: Define the decision output:
[0055] This step defines the action space as the decision output based on vehicle mobility and coverage characteristics.
[0056] Based on whether the vehicle is in the overlapping coverage area of 5G and RSU and its movement trend in the environmental perception state output in step one (i.e. the vehicle location and network coverage status identified in step one), a discrete action space covering network access actions and cross-node vertical handover actions is set up, and it is used as the output of the decision, so that the perceived dynamic environmental state is transformed into specific executable scheduling options.
[0057] Different network conditions will generate different access requirements. Based on the vehicle's access or handover requirements, discrete action spaces are defined. It includes two categories: access actions. and switching actions Regarding access actions The action set within the switching action space Action instructions are binarized by a factor It means that among them .
[0058] when When the system accepts the initial access request from the target vehicle, or allows it to perform a cross-node switch from the original node to the target node, the available service capacity of the corresponding target node is reduced accordingly; when At this time, it means the system rejects the access request or maintains the vehicle's current network connection.
[0059] Step 3: Select actions from the discrete action space, design a joint utility function and access matching strategy, generate reward values to evaluate the merits of access or handover actions, and use these rewards as the basis for updating network parameters in the online optimization of the deep Q network.
[0060] (1) For the candidate actions in step two, by setting differentiated weights for service transmission rate guarantee and available capacity utilization, an action is selected from the discrete action space based on the current state. Subsequently, in order to comprehensively evaluate and correct the merits of the selected action, the available service capacity is obtained. The following processing logic is executed:
[0061] The service transmission rate guarantee is set as follows:
[0062] This embodiment utilizes a conventional discrete-time channel model in the art to obtain the expected maximum service transmission rate of the target vehicle, which is then used as the basic quantization parameter for constructing the joint utility function. Specifically, it is assumed that in each time interval... With the internal channel state information remaining unchanged, the maximum service transmission rate of the target vehicle is calculated as follows:
[0063] (1)
[0064] in For channel bandwidth, For transmission power, For channel gain, Indicates noise power.
[0065] The available capacity utilization is set as follows:
[0066] In heterogeneous vehicular networks, the physical characteristics and service capabilities of each access point determine its carrying capacity. Specifically, 5G base stations provide a larger node carrying capacity than roadside units, while roadside units offer limited localized access capacity. To quantify the maximum number of services that can be supported, the maximum carrying capacity is defined from the perspective of the entire network unit as: ,in This indicates the carrying capacity of medium-speed services. This indicates the capacity for low-speed services. This indicates the capacity for high-speed services. The current usage of different business types is recorded as follows: , , .
[0067] (2)
[0068] (3)
[0069] (4)
[0070] Therefore, the available carrying capacity of the service can be expressed as:
[0071] (5)
[0072] (6)
[0073] (7)
[0074] (2) Combine transmission rate and available capacity to construct an optimized joint utility function as an evaluation criterion. The specific steps are as follows:
[0075] Calculate the maximum transmission rate that the target vehicle can obtain if it connects to the corresponding network node;
[0076] Real-time statistics on the remaining available capacity of each network node for low, medium, and high-speed services;
[0077] Based on the aforementioned maximum transmission rate and remaining available capacity, the joint utility function of this invention comprehensively considers two key factors: service transmission rate guarantee and available capacity of access network nodes.
[0078] (8)
[0079] in, and The rate weighting coefficient and capacity weighting coefficient are set for different rate services, and the specific values are shown in Table 1. Indicates business type average transmission rate Indicates business type Available carrying capacity This is the redistribution penalty coefficient triggered based on the access matching policy.
[0080] exist Service selection decisions are influenced by both service transmission rate and the available capacity of access network nodes. The optimization objective of this invention is to maximize a utility function that simultaneously considers the utilization rate of both transmission rate and available capacity of access network nodes. The optimization factor is defined based on a heterogeneous vehicular network environment. In this way, the proposed framework achieves coordinated optimization of transmission rate assurance and capacity utilization. Accordingly, the problem can be formulated as follows:
[0081] (9)
[0082] (10)
[0083] (11)
[0084] The optimization objective means maximizing the total utility function of the vehicle in the heterogeneous network units. This refers to the minimum transmission rate requirements for different service types. This indicates the maximum capacity limit for each service type. Ensure the minimum QoS transmission rate for different service types. Limit the capacity of each network node.
[0085] (3) Introduce an access matching strategy when the capacity of a specific service is saturated to achieve quantitative benchmarking of service requirements and network capabilities, thereby generating a reward value for evaluating the merits of access or handover actions:
[0086] When the capacity is not saturated, the system prioritizes allocating services to resource types that match the rate requirements to ensure optimal adaptation. Once a resource type becomes saturated, an access matching strategy is adopted. This strategy ensures orderly and robust access and handover in heterogeneous networks, enabling the system to dynamically adapt to diverse service demands under multi-mode capacity conditions.
[0087] In this embodiment, when the low-speed service capacity is saturated, overflow allocation is sequentially performed to the medium-speed and high-speed service resource pools; when the medium-speed service capacity is saturated, overflow allocation is performed to the high-speed service resource pool; a blocking penalty is only incurred when no resource pool is available for matching, thereby ensuring the system's access matching degree under sudden high traffic. See Table 2 for details.
[0088] Table 2 Access Matching Strategy
[0089]
[0090] Step 4: Model building, model optimization, and execution scheduling:
[0091] This step aims to comprehensively integrate the environmental perception inputs, decision outputs, and evaluation criteria sequentially output from steps one, two, and three into a deep reinforcement learning framework. By reformulating the congestion control optimization objective of vehicular heterogeneous network service scheduling using a Markov decision process, and running a DQN-based service scheduling congestion control algorithm, the deep neural network is updated with the goal of maximizing long-term expected rewards, achieving online optimization and service access scheduling. See also... Figure 3 The specific execution logic is strictly divided into the following three merged sub-steps:
[0092] First, a DQN learning model is constructed, comprising an Online Q-Network and a Target Q-Network. In terms of data mapping and loop closure construction, this step directly follows the previous sequential association steps: the multidimensional extended state space, constructed and output in step one to prevent dimensionality explosion, is used as the real-time state input (i.e., the perceived state input) of the Online Q-Network. The discrete action space defined in step two, encompassing network access and cross-node vertical handover, is used as the decision output set (i.e., the candidate action space) of the online Q-network. Simultaneously, the joint utility function result calculated in step three by setting differentiated weights and the access matching strategy is used as the instantaneous reward signal (i.e., action value reward value) fed back to the model. Thus, the discrete physical outputs of steps one, two, and three are completely transformed into the core training elements of the reinforcement learning model in step four.
[0093] Secondly, through experience replay mechanisms and network parameter updates, the deep Q-network agent is driven to learn the optimal service scheduling strategy online until the strategy converges. During the sample collection and training phase, the agent first starts from a randomly selected current environment state. Perform initialization and based on the preset exploration factors. use A greedy strategy selects an action from the action space. After taking this action, the agent interacts with the vehicle's wireless communication environment, and the environment then returns to its evolution state for the next moment. and the instantaneous reward value determined in step three. The system will transfer the complete tuple. The samples are stored in real time in the Experience Replay Buffer. Once the number of samples in the Experience Replay Buffer reaches a preset size, the Q-network training and optimization procedure is triggered: the system periodically randomly selects mini-batch interactive samples from the buffer, calculates the target Q-value using the target Q-network, and updates and optimizes the network parameters of the online Q-network by gradient, with the optimization objective being to minimize the loss function between the target Q-value and the predicted value of the online Q-network. After a fixed number of gradient updates, the latest parameters of the online Q-network are periodically copied to the target Q-network for stable model training. Through the continuous iteration of the above "interaction-collection-parameter update" closed loop, until the reward function curve gradually stabilizes and eventually converges, it indicates that the deep Q-network has successfully learned the optimal traffic allocation and congestion control strategy in the current heterogeneous vehicle network scenario.
[0094] Finally, the DQN learning model outputs and executes scheduling instructions to implement real-time network congestion control. During the actual online scheduling operation of the system, the decision center directly inputs the vehicle dynamic spatial location and network layer state characteristics collected and updated in real time in step one into the optimized online Q-network. The deep learning model then calculates and outputs the optimal service scheduling action (i.e., the best network access point selection or 5G and RSU cross-node handover decision) for different high, medium, and low rate service types in the current state. The system then automatically issues and executes the scheduling action instruction, thereby avoiding overloaded nodes at the source of the heterogeneous multi-layer network and resolving the hard congestion problem under sudden high traffic, ultimately achieving the goal of reducing the system request blocking rate and significantly optimizing service quality.
[0095] The following specific simulation examples will be provided to illustrate the invention in detail:
[0096] In the experiment, to reflect the movement of vehicles during peak hours, the vehicle density was set to 200 vehicles per kilometer, which represents a typical high-load scenario on urban roads during peak hours. The specific parameter settings such as the coverage and bandwidth of the base station and RSU are shown in Table 3.
[0097] Table 3 Simulation Parameters
[0098]
[0099] In the experiment, the vehicle arrival process followed a Poisson distribution, with the rate varying with vehicle speed. Specifically, for each vehicle speed, the arrival rate was taken from the dataset. ,in Indicates the corresponding vehicle speed The Poisson arrival rate. This dataset is shown in Table 4. In each transmission interval... Within the system, each vehicle can generate at most one service request. It is assumed that the probability of a service request occurring in the Roadside Unit (RSU) network (dual coverage area) and the 5G network (single coverage area) is equal, both being 50%. Service types follow the high, medium, and low rate classifications defined in the system model. The duration of all services follows an exponential distribution with a mean of 2 seconds. For each ongoing service, the probability of leaving the network within a given time is 50%. By studying the congestion probabilities generated at different vehicle speeds, this simulation further validates the effectiveness of the proposed mechanism in heterogeneous vehicular networks.
[0100] Table 4. Correspondence between vehicle speed and Poisson's arrival rate
[0101]
[0102] Figure 4 The figure shows the convergence of the reward function of the DQN-based service scheduling congestion control mechanism proposed in this invention during the model training phase. As can be seen from the figure, the reward function curve fluctuates significantly in the early stages of training. This is because the network parameters in the Q-Network are randomly initialized at the beginning of training, and the experience replay pool is not yet full, resulting in the agent's actions being largely random. As training continues, the experience pool gradually accumulates and fills, and the model begins to sample from the experience pool and update its parameters, causing the reward function to show an upward trend. After 8000 training rounds, the reward curve gradually stabilizes and eventually converges, indicating that the DQN agent has learned the optimal service allocation strategy.
[0103] Figure 5 The paper compares the blocking rates of the proposed method with various congestion control strategies at different vehicle speeds. Results show that within a speed range of 1 m / s to 30 m / s (covering major driving scenarios such as urban roads and main arteries), the proposed DQN-based service scheduling congestion control mechanism maintains the lowest blocking rate compared to other comparison algorithms. The blocking rate is reduced by at least 12%, demonstrating its significant advantages in urban and suburban vehicular network environments. Throughout the dynamic evolution range of vehicle speeds, the proposed method exhibits strong spatiotemporal environmental adaptability. By jointly optimizing the service transmission rate and the available capacity of the access point, robust service scheduling congestion control is achieved, demonstrating significant technological advancements across all heterogeneous vehicular network scenarios.
[0104] The above description is a specific illustration of the present invention, and not a limitation thereof. Those skilled in the art can make various equivalent technical solutions without departing from the scope of the present invention; therefore, all equivalent technical solutions should be included within the protection scope of the present invention.
Claims
1. A service scheduling congestion control method based on transmission rate and available capacity of vehicle network access points, characterized in that, Includes the following steps: Step 1: Constructing environmental perception input: Obtain vehicle location, network coverage of each access point, and real-time available capacity for different service types in heterogeneous vehicle networking scenarios. Then, construct a multi-dimensional extended state space that includes physical layer location perception features and network layer resource status, and use it as the input for subsequent intelligent service scheduling environmental perception. Step 2: Define the decision output: Define a discrete action space that includes network access actions and cross-node vertical handover actions, and use it as the output of the decision; Step 3: Select actions from the discrete action space, design a joint utility function and access matching strategy, generate reward values to evaluate the merits of access or handover actions, and use these rewards as the basis for updating network parameters in the online optimization of the deep Q network. First, for the candidate actions in step two, by setting differentiated weights for service transmission rate guarantee and available capacity utilization, actions are selected from the discrete action space based on the current state to obtain the service available capacity. Then, by combining the transmission rate and the utilization rate of available capacity, a utility function is constructed as an evaluation criterion. Finally, an access matching strategy is introduced when the carrying capacity of a specific service is saturated, and a reward value is generated to evaluate the merits of access or switching actions. Step 4: Model building, model optimization, and execution scheduling: First, the multidimensional extended state space constructed in step one is used as input, the discrete action space defined in step two is used as the decision output set, and the result of the joint utility function designed in step three is used as the reward signal to construct the DQN learning model. Secondly, through the experience replay mechanism and network parameter updates, the deep Q network agent is driven to learn the optimal business scheduling strategy online until convergence. Finally, the DQN learning model outputs and executes scheduling instructions to implement real-time network congestion control.
2. The service scheduling congestion control method based on transmission rate and available capacity of vehicle network access points according to claim 1, characterized in that, In step one, the business is refined and dimensionality reduced into three categories: low-speed, medium-speed, and high-speed businesses. The method for calculating real-time available capacity is as follows: when a service of a certain rate successfully accesses the corresponding network node, the capacity status value of the corresponding network node is incremented by 1.
3. The service scheduling congestion control method based on transmission rate and available capacity of vehicle network access points according to claim 1, characterized in that, In step two, the discrete action space includes access actions and switching actions, and the action instructions in the switching action space are determined by a binarization factor. It means that, among them : when When this occurs, it means the system has accepted the initial access request from the target vehicle, or allowed it to perform a cross-node switch from the original node to the target node. At this time, the available service capacity of the corresponding node will be reduced accordingly. when At this time, it means the system rejects the access request or maintains the vehicle's current network connection.
4. The service scheduling congestion control method based on transmission rate and available capacity of vehicle network access points according to claim 1, characterized in that, In step three, the construction steps of the utility function are specifically as follows: (1) Calculate the maximum transmission rate that the target vehicle can obtain if it connects to the corresponding network node; (2) Real-time statistics of the remaining available capacity of each network node for low, medium and high speed services; Based on (1) and (2) above, the formula for constructing the joint utility function is as follows: in, For business type average transmission rate For business type Available carrying capacity and These are the rate weighting coefficient and the capacity weighting coefficient, respectively set for services with different rates. This is the redistribution penalty coefficient triggered based on the access matching policy.
5. The service scheduling congestion control method based on transmission rate and available capacity of vehicle network access points according to claim 1, characterized in that, In step three, the specific access matching strategy when the service capacity is saturated is as follows: When the capacity of low-speed services is saturated, overflow allocation is sequentially made to the medium-speed and high-speed service resource pools. When the capacity of medium-speed services is saturated, the excess resources are allocated to the high-speed service resource pool. A blocking penalty occurs when no resource pool can be matched.
Citation Information
Patent Citations
Joint computing unloading and resource allocation method based on multi-agent DDQN
CN114584951A
Distributed wireless network user switching control method based on deep reinforcement learning
CN121619625A