Urban road intelligent height limit control method and system based on Internet of Vehicles and Internet of Things

By applying Internet of Vehicles and Internet of Things technology in height-limiting devices and optimizing height-limiting control strategies using reinforcement learning algorithms, the problem that existing height-limiting devices cannot be dynamically adjusted is solved, and more efficient and safe traffic management is achieved.

CN120108174APending Publication Date: 2025-06-06POWER CHINA KUNMING ENG CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510155585.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Existing height limit devices cannot dynamically adjust based on real-time traffic flow and vehicle height distribution, resulting in inefficient traffic efficiency and increased risk of traffic accidents, and fail to make full use of vehicle and Internet of Things technologies to achieve intelligent height limit control.

Method used

The intelligent height limit control method based on the Internet of Vehicles and the Internet of Things is adopted, and by obtaining historical status information and using attention network, actor network, reward critic network and cost critic network for training, the height limit control strategy is optimized, and the height limit value is dynamically adjusted to respond to real-time traffic flow and vehicle height distribution.

Benefits of technology

Dynamic adjustment of urban road height limit devices has been achieved, road traffic efficiency has been improved, traffic congestion has been reduced, and traffic accident risks have been reduced due to improper height limits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108174A_ABST
    Figure CN120108174A_ABST
Patent Text Reader

Abstract

The invention discloses an urban road intelligent height limit control method and system based on the Internet of Vehicles and the Internet of Things, and the method comprises the steps: obtaining historical state information, including observation information, actions, reward values and cost values; inputting the observation information into an attention network, obtaining features, inputting the features into an actor network and a criticer network, and respectively obtaining a probability value, an award value and a cost value; obtaining an evaluation value through the advantage evaluation function and the cost advantage evaluation function; inputting the evaluation value and the probability value into a target function to optimize the actor network, and inputting the reward value and the value into a loss function to optimize the criticer network; repeating the above steps until a preset number of times is exceeded, and obtaining a trained actor network; and acquiring current observation information, and inputting the current observation information into the trained actor network to obtain actions so as to control the height limiting device. According to the invention, the intelligent control of the urban road height limiting device can be realized, the road passing efficiency is improved, the traffic jam is reduced, and the traffic accident risk caused by improper height limiting is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent transportation systems, and in particular to an urban road intelligent height limit control method and system based on vehicle networking and the Internet of Things. Background Art

[0002] In urban traffic management, the installation of road height limit devices is of great significance for ensuring traffic safety and improving traffic efficiency. Traditional height limit devices usually adopt fixed height settings and cannot be dynamically adjusted according to real-time traffic flow and vehicle height distribution. This fixed height limit method often leads to low traffic efficiency and even traffic congestion when facing different time periods and different vehicle types. In addition, fixed height limit devices may cause unnecessary restrictions on overheight vehicles in some cases, increasing the risk of traffic accidents.

[0003] Existing height limit control technologies mainly rely on manual adjustment or simple sensor triggering mechanisms. These methods are not only inefficient, but also unable to respond to changes in traffic flow in real time. With the development of Internet of Vehicles and Internet of Things technologies, intelligent transportation systems have gradually become a research hotspot. Internet of Vehicles technology enables real-time communication between vehicles and road infrastructure, while Internet of Things technology enables comprehensive perception of the road environment. These technologies make it possible to dynamically adjust height limit devices, but there is currently no mature technical solution that can fully utilize these technologies to achieve intelligent height limit control.

[0004] In the process of implementing the embodiments of the present invention, the inventors found that there are at least the following problems or defects in the prior art: the existing height limiting device cannot be dynamically adjusted according to the real-time traffic flow and vehicle height distribution, resulting in low traffic efficiency and increased risk of traffic accidents; the prior art fails to fully utilize the Internet of Vehicles and Internet of Things technologies to realize intelligent height limit control, and lacks an effective dynamic adjustment mechanism. Summary of the invention

[0005] The embodiments of the present invention aim to provide a method and system for intelligent height limit control of urban roads based on Internet of Vehicles and Internet of Things, so as to solve the technical problems raised in the prior art.

[0006] The embodiment of the present invention solves the technical problem by adopting the following technical solution:

[0007] In a first aspect, a method for intelligent height limit control of urban roads based on Internet of Vehicles and Internet of Things is provided, comprising:

[0008] Step 1: Acquire multiple historical state information, each of which includes first observation information, second observation information, action, reward value and cost value; the first observation information and the second observation information correspond to adjacent moments; the reward value is calculated by a height limit reward function, which is constructed based on the traffic flow of each road section and the vehicle height distribution of the corresponding road section; the cost value is calculated by a height limit cost function, which is constructed by the traffic flow of the road section and the preset maximum height limit ratio;

[0009] Step 2: Input the first observation information and the second observation information of each historical state information into the attention network respectively to obtain the first feature and the second feature; input the first feature into the actor network to obtain the first probability value; input the first feature and the second feature into the reward critic network respectively to obtain the first reward value and the second reward value; input the first feature and the second feature into the cost critic network respectively to obtain the first cost value and the second cost value; input the reward value, the first reward value and the second reward value into the reward advantage evaluation function to obtain the advantage evaluation value; input the cost value, the first cost value and the second cost value into the cost advantage evaluation function to obtain the cost evaluation value;

[0010] Step 3: Input the advantage evaluation value, the cost evaluation value and the first probability value into the objective function and optimize to obtain the optimized actor network; input the reward value and the first reward value into the loss function of the reward critic network and optimize to obtain the optimized reward critic network; input the cost value and the first cost value into the loss function of the cost critic network and optimize to obtain the optimized cost critic network;

[0011] Step 4: Based on the optimized actor network, reward critic network and cost critic network, repeat steps 1 to 3 until the preset number of times is exceeded to obtain the trained actor network;

[0012] Step 5: Obtain the observation information at the current moment and input it into the trained actor network to obtain the current action to control the height limit device.

[0013] Furthermore, it also includes: executing the current action to obtain observation information at the next moment;

[0014] The current reward value is calculated based on the limit reward function;

[0015] The current cost value is calculated according to the height-limited cost function;

[0016] Construct the current state information based on the observation information at the current moment, the observation information at the next moment, the current action, the current reward value, and the current cost value;

[0017] Based on the current state information, the actor network, reward critic network and cost critic network are optimized again, and the observation information at the next moment is processed using the re-optimized actor network.

[0018] Furthermore, it also includes: using the mean square error to construct the loss function of the reward critic network and the cost critic network, minimizing the loss function through the Adam gradient descent algorithm, updating the parameters of the reward critic network and the parameters of the cost critic network, and obtaining the optimized reward critic network and the cost critic network.

[0019] Furthermore, the expression of the height limit reward function is:

[0020]

[0021] Among them, R i,t represents the height limit reward value of road section i at time t, i represents the identification number of the road section, t represents the time, l represents the identification number of the lane, L i represents the lane set of road section i, H l,t represents the average height of vehicles in lane l of section i at time t, H max Indicates the preset maximum height limit, Q l,t represents the traffic flow of lane l of road section i at time t.

[0022] Furthermore, the expression of the height limit cost function is:

[0023]

[0024] Among them, C i,t represents the height restriction cost of road section i at time t, H l,t represents the average height of vehicles in lane l of section i at time t, ratiol represents the preset maximum height limit ratio, Q l,t represents the traffic flow of lane l of road section i at time t.

[0025] Furthermore, the expression of the reward advantage evaluation function is:

[0026]

[0027] in, represents the reward advantage evaluation value of road segment i at time t, k represents the stage of advantage evaluation, γ represents the discount factor, m represents the mth stage in the advantage evaluation process, R i,t represents the height limit reward value of road section i at time t, V R,i,t Indicates the second reward value V corresponding to the second observation information R,i,t-1 represents the first reward value corresponding to the first observation information, λ GAErepresents the advantage value parameter, which is used to control the degree of averaging of the advantage value, and Y represents the length of the sampling trajectory.

[0028] Furthermore, the expression of the cost advantage evaluation function is:

[0029]

[0030] in, represents the cost advantage evaluation value of road section i at time t, k represents the stage of advantage evaluation, γ represents the discount factor, m represents the mth stage in the advantage evaluation process, C i,t represents the height restriction cost of road section i at time t, V C,i,t Represents the second cost value corresponding to the second observation information, V C,i,t-1 represents the first cost value corresponding to the first observation information, λ GAE represents the advantage value parameter, which is used to control the degree of averaging of the advantage value, and Y represents the length of the sampling trajectory.

[0031] Furthermore, the objective function is expressed as:

[0032]

[0033] Among them, L(θ i ,λ i ) represents the target value, Represented in the actor strategy network The expected experience return value under γ represents the discount factor, R i,t represents the height limit reward value of road section i at time t, C i,t represents the height restriction cost of road section i at time t, λ i represents the Lagrange multiplier of road segment i.

[0034] In a second aspect of the present invention, an intelligent height limit control system for urban roads based on Internet of Vehicles and Internet of Things is provided, comprising:

[0035] The training module is used to perform the following steps, including:

[0036] Step 1: Obtain multiple historical state information, each historical state information includes first observation information, second observation information, action, reward value and cost value; the moments corresponding to the first observation information and the second observation information are adjacent moments; the reward value is calculated by a height limit reward function, and the height limit reward function is constructed based on the traffic flow of each road section and the vehicle height distribution of the corresponding road section; the cost value is calculated by a height limit cost function, and the height limit cost function is constructed by the traffic flow of the road section and the preset maximum height limit ratio; Step 2: Input the first observation information and the second observation information of each historical state information into the attention network respectively to obtain the first feature and the second feature; input the first feature into the actor network to obtain the first probability value; input the first feature and the second feature into the reward critic network respectively to obtain the first reward value and the second reward value; input the first feature and the second feature into the reward critic network respectively to obtain the first reward value and the second reward value; input the first feature and the second feature into the reward critic network respectively Input into the cost critic network to obtain the first cost value and the second cost value; input the reward value, the first reward value and the second reward value into the reward advantage evaluation function to obtain the advantage evaluation value; input the cost value, the first cost value and the second cost value into the cost advantage evaluation function to obtain the cost evaluation value; step three: input the advantage evaluation value, the cost evaluation value and the first probability value into the objective function and optimize to obtain the optimized actor network; input the reward value and the first reward value into the loss function of the reward critic network and optimize to obtain the optimized reward critic network; input the cost value and the first cost value into the loss function of the cost critic network and optimize to obtain the optimized cost critic network; step four: based on the optimized actor network, reward critic network and cost critic network, repeat steps one to three until the preset number of times is exceeded to obtain the trained actor network;

[0037] The execution module obtains the observation information at the current moment and inputs it into the trained actor network to obtain the current action to control the height limit device.

[0038] Furthermore, the training module further includes: executing the current action to obtain observation information at the next moment;

[0039] The current reward value is calculated based on the limit reward function;

[0040] The current cost value is calculated according to the height-limited cost function;

[0041] Construct the current state information based on the observation information at the current moment, the observation information at the next moment, the current action, the current reward value, and the current cost value;

[0042] Based on the current state information, the actor network, reward critic network and cost critic network are optimized again, and the observation information at the next moment is processed using the re-optimized actor network.

[0043] The above-mentioned embodiments of the present invention have at least the following beneficial effects: the present invention can realize the dynamic adjustment of the height limit device of urban roads, and intelligently control the height limit value according to the real-time traffic flow and vehicle height distribution, thereby improving the road traffic efficiency and reducing traffic congestion. In addition, through the Internet of Vehicles and Internet of Things technology, the present invention can obtain road and vehicle information in real time, ensure the accuracy and timeliness of height limit control, and reduce the risk of traffic accidents caused by improper height limit.

[0044] The present invention can also optimize the height limit control strategy through reinforcement learning algorithm to further improve the adaptability and robustness of the system. This method can not only adapt to changes in different time periods and vehicle types, but also maintain stable performance in complex traffic environments, providing a more intelligent solution for urban traffic management. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] One or more embodiments are exemplarily described by the figures in the corresponding drawings, and these exemplifications do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, and unless otherwise stated, the figures in the drawings do not constitute proportional limitations.

[0046] Figure 1 A schematic diagram of a flow chart of an intelligent height limit control method for urban roads based on Internet of Vehicles and Internet of Things provided by one embodiment of the present invention;

[0047] Figure 2 A schematic diagram of the structure of an intelligent height limit control system for urban roads based on the Internet of Vehicles and the Internet of Things provided by one embodiment of the present invention;

[0048] Figure 3 The schematic diagram schematically shows the structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0049] In order to facilitate the understanding of the present invention, the present invention is described in more detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that when an element is described as "connecting" another element, it can be directly on another element, or there can be one or more centered elements therebetween. The orientation or positional relationship indicated by the terms "upper", "lower", "left", "right", "upper end", "lower end", "top" and "bottom" used in this specification is based on the orientation or positional relationship shown in the accompanying drawings, only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0050] Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0051] Combine the following Figure 1 , the urban road intelligent height limit control method 100 based on the Internet of Vehicles and the Internet of Things provided in the embodiment of the present application is described in detail through specific embodiments.

[0052] Figure 1 The figure is a flow chart of the intelligent height limit control method for urban roads based on the Internet of Vehicles and the Internet of Things provided by the present invention. The intelligent height limit control method for urban roads based on the Internet of Vehicles and the Internet of Things provided by one embodiment of the present invention comprises:

[0053] Step 1: Acquire multiple historical state information, each of which includes first observation information, second observation information, action, reward value and cost value; the first observation information and the second observation information correspond to adjacent moments; the reward value is calculated by a height limit reward function, which is constructed based on the traffic flow of each road section and the vehicle height distribution of the corresponding road section; the cost value is calculated by a height limit cost function, which is constructed by the traffic flow of the road section and the preset maximum height limit ratio;

[0054] Step 2: Input the first observation information and the second observation information of each historical state information into the attention network respectively to obtain the first feature and the second feature; input the first feature into the actor network to obtain the first probability value; input the first feature and the second feature into the reward critic network respectively to obtain the first reward value and the second reward value; input the first feature and the second feature into the cost critic network respectively to obtain the first cost value and the second cost value; input the reward value, the first reward value and the second reward value into the reward advantage evaluation function to obtain the advantage evaluation value; input the cost value, the first cost value and the second cost value into the cost advantage evaluation function to obtain the cost evaluation value;

[0055] Step 3: Input the advantage evaluation value, the cost evaluation value and the first probability value into the objective function and optimize to obtain the optimized actor network; input the reward value and the first reward value into the loss function of the reward critic network and optimize to obtain the optimized reward critic network; input the cost value and the first cost value into the loss function of the cost critic network and optimize to obtain the optimized cost critic network;

[0056] Step 4: Based on the optimized actor network, reward critic network and cost critic network, repeat steps 1 to 3 until the preset number of times is exceeded to obtain the trained actor network;

[0057] Step 5: Obtain the observation information at the current moment and input it into the trained actor network to obtain the current action to control the height limit device.

[0058] It should be noted that the method includes obtaining multiple historical state information, each of which includes first observation information, second observation information, action, reward value and cost value. The historical state information here refers to the state data recorded by the system at a certain moment in the past, and the first observation information and the second observation information respectively represent the traffic conditions observed at two adjacent moments. Action refers to the height limit control measures taken by the system at a certain moment. The reward value and cost value are calculated by the height limit reward function and the height limit cost function respectively, and are used to evaluate the effectiveness of the height limit control measures.

[0059] Specifically, the first observation information and the second observation information may include parameters such as the height, speed, and traffic flow of the vehicle. The height limit reward function is constructed based on the traffic flow and vehicle height distribution of each road section, with specific parameters such as the average height of the vehicle and the preset maximum height limit value. The height limit cost function is constructed by the traffic flow of the road section and the preset maximum height limit ratio, with specific parameters such as the average height of the vehicle and the preset maximum height limit ratio. The attention network is used to process observation information, the actor network is used to generate control actions, and the reward critic network and the cost critic network are used to evaluate the reward value and the cost value, respectively.

[0060] Preferably, the first observation information and the second observation information are respectively input into the attention network to obtain the first feature and the second feature; the first feature is input into the actor network to obtain the first probability value; the first feature and the second feature are respectively input into the reward critic network and the cost critic network to obtain the reward value and the cost value. Based on these features and evaluation values, by optimizing the objective function and the loss function, the actor network, the reward critic network and the cost critic network are gradually optimized until the preset number of times is exceeded to obtain the trained actor network. Finally, the observation information at the current moment is obtained and input into the trained actor network to obtain the current action to control the height limit device.

[0061] In some embodiments, the method further includes: executing the current action to obtain observation information at the next moment;

[0062] The current reward value is calculated based on the limit reward function;

[0063] The current cost value is calculated according to the height-limited cost function;

[0064] Construct the current state information based on the observation information at the current moment, the observation information at the next moment, the current action, the current reward value, and the current cost value;

[0065] Based on the current state information, the actor network, reward critic network and cost critic network are optimized again, and the observation information at the next moment is processed using the re-optimized actor network.

[0066] It should be noted that the method also includes executing the current action to obtain the observation information at the next moment. The current action here refers to the height limit control instruction generated by the trained actor network, and the observation information at the next moment refers to the new traffic condition data observed by the system after executing the action. These data will be used to further optimize the height limit control strategy and ensure the dynamic adaptability of the system.

[0067] Specifically, after executing the current action, the system will calculate the current reward value based on the height limit reward function, which evaluates the effect of height limit control based on the traffic flow and vehicle height distribution of each road section. At the same time, the system will also calculate the current cost value based on the height limit cost function, which evaluates the cost of height limit control by comparing the traffic flow of the road section with the preset maximum height limit ratio. These reward values ​​and cost values ​​will be used together with the observation information at the current moment, the observation information at the next moment, and the current action to construct the current state information, which will be used to optimize the actor network, reward critic network, and cost critic network again.

[0068] Preferably, the system can use the current state information to re-optimize the actor network, the reward critic network and the cost critic network. For example, the performance of the network can be improved by adjusting the parameters of the network, such as the learning rate, the discount factor, etc.

[0069] Furthermore, more observation information can be introduced, such as vehicle type, weather conditions, etc., to enrich the status information, so that the system can more comprehensively evaluate the effect of height limit control. In this way, the system can continuously learn and adapt to new traffic conditions, thereby achieving more accurate height limit control.

[0070] In some embodiments, it also includes: using the mean square error to construct the loss function of the reward critic network and the cost critic network, minimizing the loss function through the Adam gradient descent algorithm, updating the parameters of the reward critic network and the parameters of the cost critic network, and obtaining the optimized reward critic network and cost critic network.

[0071] It should be noted that the method also includes constructing the loss function of the reward critic network and the cost critic network using the mean square error, minimizing the loss function through the Adam gradient descent algorithm, updating the parameters of the reward critic network and the parameters of the cost critic network, and obtaining the optimized reward critic network and the cost critic network. The mean square error here refers to the average of the squares of the differences between the predicted value and the true value, which is used to measure the accuracy of the model prediction. The Adam gradient descent algorithm is an optimization algorithm used to adjust network parameters to minimize the loss function.

[0072] Specifically, the calculation formula of the mean square error is:

[0073]

[0074] Among them, y i is the true value, is the predicted value, and n is the number of samples. In the reward critic network and the cost critic network, the loss function can be expressed as and Where R and C are the real reward value and cost value respectively. and is the reward and cost value predicted by the network. The Adam gradient descent algorithm adjusts the learning rate by calculating the exponential weighted average of the gradient. The specific formula is m t =β 1 m t-1 +(1-β 1 ) t and Where m t and v t They are the first-order moment estimate and the second-order moment estimate of the gradient, β 1 and β 2 is the decay rate, g t is the current gradient.

[0075] Preferably, the parameters of the Adam algorithm can be set, such as the learning rate α, the decay rate β 1 and β 2 , and ∈ is used for numerical stability. For example, we can set α = 0.001, β 1 =0.9,β 2 =0.999,∈=10 -8 .

[0076] Furthermore, regularization terms, such as L2 regularization, can be introduced to prevent overfitting. The parameter λ of the regularization term can be set to 0.01. By adjusting these parameters, the performance of the reward critic network and the cost critic network can be further optimized, and the accuracy and stability of the height limit control can be improved.

[0077] In some embodiments, the expression of the height limit reward function is:

[0078]

[0079] Among them, R i,t represents the height limit reward value of road section i at time t, i represents the identification number of the road section, t represents the time, l represents the identification number of the lane, L i represents the lane set of road section i, H l,t represents the average height of vehicles in lane l of section i at time t, H max Indicates the preset maximum height limit, Q l,t represents the traffic flow of lane l of road section i at time t.

[0080] It should be noted that the expression of the limit reward function is

[0081]

[0082] Among them, R s,a represents the height limit reward value of section a at time s, L s represents the lane set of road section s, h s,l represents the average height of vehicles in lane l of section a at time s, H max Indicates the preset maximum height limit, f s,l Represents the traffic flow of lane l of section a at time s. This function evaluates the effect of height control by calculating the ratio of the average height of vehicles in each lane to the preset maximum height limit value and combining it with the traffic flow.

[0083] Specifically, the parameters in the height limit reward function are set as follows: H max It can be set according to the road design specifications and safety requirements. For example, for ordinary urban roads, it can be set to 4.5 meters; for highways, it can be set to 5.0 meters. Traffic flow f s,l It can be obtained in real time through traffic sensors in units of vehicles per hour. The average height of vehicles h s,l It can be obtained through vehicle height sensors or vehicle networking data in meters. The concept of this function is that when the average height of the vehicle approaches the preset maximum height limit, the reward value will decrease, thereby encouraging the system to take height limit measures to avoid the passage of overheight vehicles, while considering the impact of traffic flow to balance traffic efficiency and safety.

[0084] Preferably, the height limit reward function can be further refined. For example, a weight factor α can be introduced to adjust the influence of the ratio of the average height of the vehicle to the preset maximum height limit value, that is, The weight factor α can be set according to actual needs, for example, α=0.8.

[0085] Furthermore, we can also consider introducing the time factor and weighting the traffic flow and vehicle height at different times to better adapt to the dynamic changes of traffic flow. For example, for peak hours, we can increase the weight of traffic flow, while for non-peak hours, we can increase the weight of vehicle height. Through these refinements and alternatives, we can further improve the adaptability and effectiveness of the height limit reward function.

[0086] In some embodiments, the expression of the height limit cost function is:

[0087]

[0088] Among them, C i,t represents the height restriction cost of road section i at time t, H l,t represents the average height of vehicles in lane l of section i at time t, ratiol represents the preset maximum height limit ratio, Q l,t represents the traffic flow of lane l of road section i at time t.

[0089] It should be noted that the expression of the height limit cost function is

[0090]

[0091] Among them, C s,a represents the height restriction cost of section a at time s, L s represents the lane set of road section s, h s,l represents the average height of vehicles in lane l of section a at time s, H max Indicates the preset maximum height limit, H ratio Indicates the preset maximum height ratio, f s,l represents the traffic flow of lane l in section a at time s. This function evaluates the cost of height control by calculating the ratio of the average vehicle height to the preset maximum height limit value and combining it with the traffic flow.

[0092] Specifically, the parameters in the height limit cost function are set as follows: The preset maximum height limit value H max It can be set according to the road design specifications and safety requirements. For example, for ordinary urban roads, it can be set to 4.5 meters; for expressways, it can be set to 5.0 meters. ratio It can be set according to actual needs, for example, it can be set to 0.9, which means that the average height of the vehicle should not exceed 90% of the preset maximum height limit. s,l It can be obtained in real time through traffic sensors in units of vehicles per hour. The average height of vehicles h s,lIt can be obtained through vehicle height sensors or vehicle networking data in meters. The concept of this function is that when the average height of the vehicle exceeds the preset maximum height limit ratio, the cost value will increase, thereby encouraging the system to take height limit measures to avoid the passage of overheight vehicles, while considering the impact of traffic flow to balance traffic efficiency and safety.

[0093] Preferably, the height limit cost function can be further refined. For example, a weight factor β can be introduced to adjust the influence of the difference between the average height of the vehicle and the preset maximum height limit ratio, that is, The weight factor β can be set according to actual needs, for example, β=1.2.

[0094] Furthermore, we can also consider introducing the time factor and weighting the traffic flow and vehicle height at different times to better adapt to the dynamic changes of traffic flow. For example, for peak hours, we can increase the weight of traffic flow, while for non-peak hours, we can increase the weight of vehicle height. Through these refinements and alternatives, we can further improve the adaptability and effectiveness of the height limit cost function.

[0095] In some embodiments, the reward advantage evaluation function is expressed as:

[0096]

[0097] in, represents the reward advantage evaluation value of road segment i at time t, k represents the stage of advantage evaluation, γ represents the discount factor, m represents the mth stage in the advantage evaluation process, R i,t represents the height limit reward value of road section i at time t, V R,i,t Indicates the second reward value V corresponding to the second observation information R,i,t-1 represents the first reward value corresponding to the first observation information, λ GAE represents the advantage value parameter, which is used to control the degree of averaging of the advantage value, and Y represents the length of the sampling trajectory.

[0098] It should be noted that the expression of the reward advantage evaluation function is

[0099]

[0100] Among them, A s,a,t(k) represents the reward advantage evaluation value of segment a at time s, k represents the stage of advantage evaluation, γ represents the discount factor, T represents the length of the sampled trajectory, R s,a represents the height limit reward value of section a at time s, V s,arepresents the reward value of road section a at time s, and λ represents the advantage value parameter, which is used to control the average degree of the advantage value. This function evaluates the advantage of height control by calculating the difference between the reward value and the reward value, and combining the discount factor and the advantage value parameter.

[0101] Specifically, the parameters in the reward advantage evaluation function are set as follows: The discount factor γ is usually set between 0 and 1, for example, it can be set to 0.99, indicating the degree of discount on future rewards. The length T of the sampling trajectory can be set according to actual needs, for example, it can be set to 10, indicating that the rewards of the next 10 time steps are considered when evaluating the advantage. The advantage value parameter λ is usually set between 0 and 1, for example, it can be set to 0.95, indicating the control of the average degree of reward value when calculating the advantage value. The concept of this function is to evaluate the advantage of height control by calculating the difference between the current reward value and the future reward value, and combining the discount factor and advantage value parameters, so as to provide guidance for the optimization of the actor network.

[0102] Preferably, the reward advantage evaluation function can be further refined. For example, a smoothing parameter α can be introduced to adjust the smoothness of the reward value, that is, The smoothing parameter α can be set according to actual needs, for example, α=0.9.

[0103] Furthermore, it is also possible to consider introducing a time factor and weighting the reward value and reward value of different time periods to better adapt to the dynamic changes in traffic flow. For example, for peak hours, the weight of the reward value can be increased, while for non-peak hours, the weight of the reward value can be increased. Through these refinements and alternatives, the adaptability and effectiveness of the reward advantage evaluation function can be further improved.

[0104] In some embodiments, the cost advantage evaluation function is expressed as:

[0105]

[0106] in, represents the cost advantage evaluation value of road section i at time t, k represents the stage of advantage evaluation, γ represents the discount factor, m represents the mth stage in the advantage evaluation process, C i,t represents the height restriction cost of road section i at time t, V C,i,t Represents the second cost value corresponding to the second observation information, V C,i,t-1 represents the first cost value corresponding to the first observation information, λ GAE represents the advantage value parameter, which is used to control the degree of averaging of the advantage value, and Y represents the length of the sampling trajectory.

[0107] It should be noted that the expression of the cost advantage evaluation function is

[0108]

[0109] Among them, A s,a,t(k) represents the cost advantage evaluation value of section a at time s, k represents the stage of advantage evaluation, γ represents the discount factor, T represents the length of the sampling trajectory, C s,a represents the height restriction cost of section a at time s, V s,a represents the cost value of road section a at time s, and λ represents the advantage value parameter, which is used to control the average degree of the advantage value. This function evaluates the cost advantage of height control by calculating the difference between the cost value and the cost value, and combining the discount factor and the advantage value parameter.

[0110] Specifically, the parameters in the cost advantage evaluation function are set as follows: The discount factor γ is usually set between 0 and 1, for example, it can be set to 0.99, indicating the degree of discount for future costs. The length T of the sampling trajectory can be set according to actual needs, for example, it can be set to 10, indicating that the costs of the next 10 time steps are considered when evaluating the cost advantage. The advantage value parameter λ is usually set between 0 and 1, for example, it can be set to 0.95, indicating the control of the average degree of cost value when calculating the cost advantage value. The concept of this function is to evaluate the cost advantage of height limit control by calculating the difference between the current generation value and the future cost value, and combining the discount factor and advantage value parameters, so as to provide guidance for the optimization of the actor network.

[0111] Preferably, the cost advantage evaluation function can be further refined. For example, a smoothing parameter α can be introduced to adjust the smoothness of the cost value, that is, The smoothing parameter α can be set according to actual needs, for example, α=0.9.

[0112] Furthermore, we can also consider introducing the time factor and weighting the cost value and the cost value in different time periods to better adapt to the dynamic changes of traffic flow. For example, for peak hours, the weight of the cost value can be increased, while for non-peak hours, the weight of the cost value can be increased. Through these refinements and alternatives, the adaptability and effectiveness of the cost advantage evaluation function can be further improved.

[0113] In some embodiments, the objective function is expressed as:

[0114]

[0115] Among them, L(θ i ,λ i ) represents the target value, Represented in the actor strategy network The expected experience return value under γ represents the discount factor, R i,t represents the height limit reward value of road section i at time t, C i,t represents the height restriction cost of road section i at time t, λ i represents the Lagrange multiplier of road segment i.

[0116] It should be noted that the expression of the objective function is

[0117]

[0118] Among them, J(π,λ) represents the target value, π represents the actor strategy network, λ represents the Lagrange multiplier of the road segment, γ represents the discount factor, and R s,a represents the height limit reward value of section a at time s, C s,a represents the height limit cost value of road section a at time s, and T represents the length of the sampled trajectory. This function optimizes the actor network by calculating the difference between the reward value and the cost value, and combines the discount factor and the Lagrange multiplier to achieve the optimal strategy for height limit control.

[0119] Specifically, the parameters in the objective function are set as follows: The discount factor γ is usually set between 0 and 1, for example, it can be set to 0.99, indicating the degree of discount for future rewards and costs. The Lagrange multiplier λ is used to balance the impact of rewards and costs. For example, it can be set to 0.1, indicating the degree of importance attached to costs during the optimization process. The length T of the sampling trajectory can be set according to actual needs. For example, it can be set to 10, indicating that the rewards and costs of the next 10 time steps are considered during the optimization process. The concept of this function is to optimize the actor network by calculating the difference between the current reward value and the cost value, and combining the discount factor and the Lagrange multiplier, so as to achieve the optimal strategy for height control to improve road traffic efficiency and safety.

[0120] Preferably, the objective function can be further refined. For example, a regularization term λ can be introduced reg To prevent overfitting, Regularization term λ reg It can be set according to actual needs, such as λ reg =0.01.

[0121] Furthermore, we can also consider introducing the time factor and weighting the reward and cost values ​​in different time periods to better adapt to the dynamic changes of traffic flow. For example, for peak hours, we can increase the weight of the reward value, while for non-peak hours, we can increase the weight of the cost value. Through these refinements and alternatives, we can further improve the adaptability and effectiveness of the objective function.

[0122] The above-mentioned embodiments of the present invention have the following beneficial effects: the present invention can realize the dynamic adjustment of the height limit device of urban roads, and intelligently control the height limit value according to the real-time traffic flow and vehicle height distribution, thereby improving the road traffic efficiency and reducing traffic congestion. In addition, through the Internet of Vehicles and Internet of Things technology, the present invention can obtain road and vehicle information in real time, ensure the accuracy and timeliness of height limit control, and reduce the risk of traffic accidents caused by improper height limit. The present invention can also optimize the height limit control strategy through reinforcement learning algorithm, further improve the adaptability and robustness of the system, adapt to changes in different time periods and vehicle types, maintain stable performance in complex traffic environments, and provide a more intelligent solution for urban traffic management.

[0123] like Figure 2 As shown, in some embodiments, an intelligent height limit control system 200 for urban roads based on the Internet of Vehicles and the Internet of Things includes:

[0124] The training module 201 is used to perform the following steps, including:

[0125] Step 1: Obtain multiple historical state information, each historical state information includes first observation information, second observation information, action, reward value and cost value; the moments corresponding to the first observation information and the second observation information are adjacent moments; the reward value is calculated by a height limit reward function, and the height limit reward function is constructed based on the traffic flow of each road section and the vehicle height distribution of the corresponding road section; the cost value is calculated by a height limit cost function, and the height limit cost function is constructed by the traffic flow of the road section and the preset maximum height limit ratio; Step 2: Input the first observation information and the second observation information of each historical state information into the attention network respectively to obtain the first feature and the second feature; input the first feature into the actor network to obtain the first probability value; input the first feature and the second feature into the reward critic network respectively to obtain the first reward value and the second reward value; input the first feature and the second feature into the reward critic network respectively to obtain the first reward value and the second reward value; input the first feature and the second feature into the reward critic network respectively Input into the cost critic network to obtain the first cost value and the second cost value; input the reward value, the first reward value and the second reward value into the reward advantage evaluation function to obtain the advantage evaluation value; input the cost value, the first cost value and the second cost value into the cost advantage evaluation function to obtain the cost evaluation value; step three: input the advantage evaluation value, the cost evaluation value and the first probability value into the objective function and optimize to obtain the optimized actor network; input the reward value and the first reward value into the loss function of the reward critic network and optimize to obtain the optimized reward critic network; input the cost value and the first cost value into the loss function of the cost critic network and optimize to obtain the optimized cost critic network; step four: based on the optimized actor network, reward critic network and cost critic network, repeat steps one to three until the preset number of times is exceeded to obtain the trained actor network;

[0126] The execution module 202 obtains the observation information at the current moment and inputs it into the trained actor network to obtain the current action to control the height limiting device.

[0127] It is understandable that the modules recorded in the urban road intelligent height limit control system 200 based on the Internet of Vehicles and the Internet of Things are the same as those in the reference Figure 1 Therefore, the operations, features and beneficial effects described above for the urban road intelligent height limit control method based on the Internet of Vehicles and the Internet of Things are also applicable to the urban road intelligent height limit control system 200 based on the Internet of Vehicles and the Internet of Things and the modules contained therein, and will not be repeated here.

[0128] In some embodiments, the training module further includes:

[0129] Execute the current action and obtain the observation information of the next moment;

[0130] The current reward value is calculated based on the limit reward function;

[0131] The current cost value is calculated according to the height-limited cost function;

[0132] Construct the current state information based on the observation information at the current moment, the observation information at the next moment, the current action, the current reward value, and the current cost value;

[0133] Based on the current state information, the actor network, reward critic network and cost critic network are optimized again, and the observation information at the next moment is processed using the re-optimized actor network.

[0134] It should be noted that the system includes a training module and an execution module. The training module is responsible for obtaining historical state information, training through the attention network, actor network, reward critic network and cost critic network, and optimizing the actor network. The execution module is responsible for obtaining the observation information at the current moment, inputting the trained actor network to obtain the current action to control the height limit device. The historical state information here refers to the state data recorded by the system at a certain moment in the past, including observation information, action, reward value and cost value. The attention network is used to process observation information, the actor network is used to generate control actions, and the reward critic network and cost critic network are used to evaluate the reward value and cost value respectively.

[0135] Specifically, the workflow of the training module is as follows: First, obtain multiple historical state information, each of which includes first observation information, second observation information, action, reward value and cost value. Then, the first observation information and the second observation information of each historical state information are respectively input into the attention network to obtain the first feature and the second feature. Next, the first feature is input into the actor network to obtain the first probability value; the first feature and the second feature are respectively input into the reward critic network and the cost critic network to obtain the reward value and the cost value. Finally, by optimizing the objective function and the loss function, the actor network, the reward critic network and the cost critic network are gradually optimized until the preset number of times is exceeded to obtain the trained actor network. The workflow of the execution module is as follows: obtain the observation information at the current moment, and input it into the trained actor network to obtain the current action to control the height limit device.

[0136] Preferably, the operating steps of the training module and the execution module can be further refined. For example, in the training module, more observation information, such as vehicle type, weather conditions, etc., can be introduced to enrich the state information so that the system can more comprehensively evaluate the effect of height limit control. In the execution module, a real-time feedback mechanism can be introduced to dynamically adjust the control action according to the actual operation of the height limit device to improve the adaptability and stability of the system. In addition, it is also possible to consider introducing a multi-agent collaboration mechanism so that multiple height limit devices can cooperate with each other and jointly optimize the height limit control strategy. Through these refinements and alternatives, the performance and reliability of the system can be further improved.

[0137] Reference below Figure 3 , which shows a schematic diagram of a structure 300 of an electronic device suitable for implementing some embodiments of the present invention. The electronic devices in some embodiments of the present invention may include, but are not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 3 The terminal device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0138] like Figure 3As shown, the electronic device 300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0139] Typically, the following devices may be connected to the I / O interface 305: input devices 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 308 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 309. The communication devices 309 may allow the electronic device 300 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 3 The electronic device 300 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead. Figure 3 Each block shown in the figure may represent one device, or may represent multiple devices as required.

[0140] Furthermore, the storage medium of the embodiment of the present application stores program instructions that can implement all the above methods, wherein the program instructions can be stored in the above storage medium in the form of a software product, including several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, or terminal devices such as a computer, a server, a mobile phone, and a tablet.

[0141] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them. Under the concept of the present invention, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other changes in different aspects of the present invention as above, which are not provided in detail for the sake of simplicity. Although the present invention is described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An intelligent height limit control method for urban roads based on Internet of Vehicles and Internet of Things, characterized in that: include: Step 1: Acquire multiple historical state information, each of which includes first observation information, second observation information, action, reward value and cost value; the first observation information and the second observation information correspond to adjacent moments; the reward value is calculated by a height limit reward function, which is constructed based on the traffic flow of each road section and the vehicle height distribution of the corresponding road section; the cost value is calculated by a height limit cost function, which is constructed by the traffic flow of the road section and the preset maximum height limit ratio; Step 2: Input the first observation information and the second observation information of each historical state information into the attention network respectively to obtain the first feature and the second feature; input the first feature into the actor network to obtain the first probability value; input the first feature and the second feature into the reward critic network respectively to obtain the first reward value and the second reward value; input the first feature and the second feature into the cost critic network respectively to obtain the first cost value and the second cost value; input the reward value, the first reward value and the second reward value into the reward advantage evaluation function to obtain the advantage evaluation value; input the cost value, the first cost value and the second cost value into the cost advantage evaluation function to obtain the cost evaluation value; Step 3: Input the advantage evaluation value, the cost evaluation value and the first probability value into the objective function and optimize to obtain the optimized actor network; input the reward value and the first reward value into the loss function of the reward critic network and optimize to obtain the optimized reward critic network; input the cost value and the first cost value into the loss function of the cost critic network and optimize to obtain the optimized cost critic network; Step 4: Based on the optimized actor network, reward critic network and cost critic network, repeat steps 1 to 3 until the preset number of times is exceeded to obtain the trained actor network; Step 5: Obtain the observation information at the current moment and input it into the trained actor network to obtain the current action to control the height limit device.

2. According to claim 1, a method for controlling the height limit of urban roads based on the Internet of Vehicles and the Internet of Things is characterized in that: Also includes: Execute the current action and obtain the observation information of the next moment; The current reward value is calculated based on the limit reward function; The current cost value is calculated according to the height-limited cost function; Construct the current state information based on the observation information at the current moment, the observation information at the next moment, the current action, the current reward value and the current cost value; Based on the current state information, the actor network, reward critic network and cost critic network are optimized again, and the observation information at the next moment is processed using the re-optimized actor network.

3. According to claim 1, a method for controlling the height limit of urban roads based on the Internet of Vehicles and the Internet of Things is characterized in that: Also includes: The mean square error is used to construct the loss function of the reward critic network and the cost critic network. The loss function is minimized by the Adam gradient descent algorithm, and the parameters of the reward critic network and the cost critic network are updated to obtain the optimized reward critic network and the cost critic network.

4. According to claim 1, a method for controlling the height limit of urban roads based on the Internet of Vehicles and the Internet of Things is characterized in that: The expression of the height limit reward function is: Among them, R i,t represents the height limit reward value of road section i at time t, i represents the identification number of the road section, t represents the time, l represents the identification number of the lane, L i represents the lane set of road section i, H l,t represents the average height of vehicles in lane l of section i at time t, H max Indicates the preset maximum height limit, Q l,t represents the traffic flow of lane l of road section i at time t.

5. According to claim 1, a method for controlling the intelligent height limit of urban roads based on the Internet of Vehicles and the Internet of Things is characterized in that: The expression of the height limit cost function is: Among them, C i,t represents the height restriction cost of road section i at time t, H l,t represents the average height of vehicles in lane l of section i at time t, ratiol represents the preset maximum height limit ratio, Q l,t represents the traffic flow of lane l of road section i at time t.

6. The method for controlling the intelligent height limit of urban roads based on the Internet of Vehicles and the Internet of Things according to claim 1 is characterized in that: The expression of the reward advantage evaluation function is: in, represents the reward advantage evaluation value of road segment i at time t, k represents the stage of advantage evaluation, γ represents the discount factor, m represents the mth stage in the advantage evaluation process, R i,t represents the height limit reward value of road section i at time t, V R,i,t Indicates the second reward value V corresponding to the second observation information R,i,t-1 represents the first reward value corresponding to the first observation information, λ GAE represents the advantage value parameter, which is used to control the degree of averaging of the advantage value, and Y represents the length of the sampling trajectory.

7. The method for controlling the intelligent height limit of urban roads based on the Internet of Vehicles and the Internet of Things according to claim 1 is characterized in that: The expression of the cost advantage evaluation function is: in, represents the cost advantage evaluation value of road section i at time t, k represents the stage of advantage evaluation, γ represents the discount factor, m represents the mth stage in the advantage evaluation process, C i,t represents the height restriction cost of road section i at time t, V C,i,t Represents the second cost value corresponding to the second observation information, V C,i,t-1 represents the first cost value corresponding to the first observation information, λ GAE represents the advantage value parameter, which is used to control the degree of averaging of the advantage value, and Y represents the length of the sampling trajectory.

8. The method for controlling the intelligent height limit of urban roads based on the Internet of Vehicles and the Internet of Things according to claim 1 is characterized in that: The expression of the objective function is: Among them, L(θ i ,λ i ) represents the target value, Represented in the actor strategy network The expected experience return value under γ represents the discount factor, R i,t represents the height limit reward value of road section i at time t, C i,t represents the height restriction cost of road section i at time t, λ i represents the Lagrange multiplier of road segment i.

9. An intelligent height limit control system for urban roads based on Internet of Vehicles and Internet of Things, characterized in that: include: The training module is used to perform the following steps, including: Step 1: Acquire multiple historical state information, each of which includes first observation information, second observation information, action, reward value and cost value; the first observation information and the second observation information correspond to adjacent moments; the reward value is calculated by a height limit reward function, which is constructed based on the traffic flow of each road section and the vehicle height distribution of the corresponding road section; the cost value is calculated by a height limit cost function, which is constructed by the traffic flow of the road section and the preset maximum height limit ratio; Step 2: Input the first observation information and the second observation information of each historical state information into the attention network respectively to obtain the first feature and the second feature; input the first feature into the actor network to obtain the first probability value; input the first feature and the second feature into the reward critic network respectively to obtain the first reward value and the second reward value; input the first feature and the second feature into the cost critic network respectively to obtain the first cost value and the second cost value; input the reward value, the first reward value and the second reward value into the reward advantage evaluation function to obtain the advantage evaluation value; input the cost value, the first cost value and the second cost value into the cost advantage evaluation function to obtain the cost evaluation value; Step 3: Input the advantage evaluation value, the cost evaluation value and the first probability value into the objective function and optimize to obtain the optimized actor network; input the reward value and the first reward value into the loss function of the reward critic network and optimize to obtain the optimized reward critic network; input the cost value and the first cost value into the loss function of the cost critic network and optimize to obtain the optimized cost critic network; Step 4: Based on the optimized actor network, reward critic network and cost critic network, repeat steps 1 to 3 until the preset number of times is exceeded to obtain the trained actor network; The execution module obtains the observation information at the current moment and inputs it into the trained actor network to obtain the current action to control the height limit device.

10. The intelligent height limit control system for urban roads based on vehicle networking and the Internet of Things according to claim 9 is characterized in that: The training module also includes: executing the current action to obtain observation information at the next moment; The current reward value is calculated based on the limit reward function; The current cost value is calculated according to the height-limited cost function; Construct the current state information based on the observation information at the current moment, the observation information at the next moment, the current action, the current reward value and the current cost value; Based on the current state information, the actor network, reward critic network and cost critic network are optimized again, and the observation information at the next moment is processed using the re-optimized actor network.