RSU auxiliary video data routing method based on deep Q learning

By adopting the RSU-assisted video data routing method based on deep Q learning in VANET network, the problem of insufficient real-time and effectiveness of video data routing algorithms in the prior art in dynamic network environments is solved, and more efficient and reliable video data transmission is achieved.

CN119946766AActive Publication Date: 2025-05-06HENAN UNIV OF SCI & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411908494.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-06
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

When the existing video data routing algorithm faces a highly dynamic VANET network environment, the meta-heuristic algorithm needs more time to converge to the optimal solution, and may not be able to adapt to the rapid changes in network topology in time, affecting the real-time and effectiveness of the algorithm.

Method used

The RSU assisted video data routing method based on deep Q learning is adopted. The next RSU between adjacent road segments is selected as the forwarding direction for SVC video data of different layers based on deep Q learning on the RSU, and the best relay vehicle is selected according to the frame type in the road segment.

Benefits of technology

It improves the transmission efficiency and quality of video data in VANET, ensures reliable segment forwarding of video data, reduces transmission delay, enhances adaptability to congested environments, and improves the possibility of video decoding on the receiver.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119946766A_ABST
    Figure CN119946766A_ABST
Patent Text Reader

Abstract

The invention provides an RSU auxiliary video data routing method based on deep Q learning, and the method comprises the steps: building a system model, collecting a real-time road condition video through a moving source vehicle, carrying out the SVC coding of the real-time road condition video, and forwarding the real-time road condition video to an RSU at an intersection in wired connection with a traffic management department; the method comprises the following steps of: performing problem description, selecting proper paths for videos of different layers for forwarding, maximizing the MOS (Metal Oxide Semiconductor) measurement of each GOP (Group of Pictures) under the constraint of time delay, and describing the problem as follows: a source vehicle depends on decision information of RSUs to select a first RSU to which to-be-forwarded data is to arrive; selecting a next RSU between adjacent road sections as a forwarding direction for SVC video data of different layers on the RSU based on deep Q learning; in the forwarding in the road section, the optimal relay vehicle is selected based on the neutrosophic set analytic hierarchy process according to the frame type, and the data is subjected to multiple decisions among the road sections and the forwarding in the road section until the data reaches the destination RSU. According to the invention, better overall performance can be effectively obtained under the scenes of different vehicle densities and different network loads.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of vehicle networking, and in particular to an RSU-assisted video data routing method based on deep Q learning. Background Art

[0002] In recent years, intelligent transportation systems have gradually become a widely discussed topic. Among them, vehicle networking technology is a key part of intelligent transportation and is also the focus of the current wireless communication field. Vehicle self-organizing network VANET is a special type of mobile self-organizing network MANET that provides communication between vehicles (V2V) and vehicles and infrastructure (V2I), and also forwards information in the form of opportunistic routing. In VANET, vehicles not only act as communication nodes, but also as routing nodes to forward messages from other vehicles, thereby creating a dynamic, self-organizing communication network. This network supports a variety of traffic management and safety applications, such as collision warning, traffic flow control, real-time traffic information updates, and emergency messaging. Sharing information through VANET helps support safer driving decisions and provide location-related services, thereby enhancing road safety and optimizing the driving experience.

[0003] Among the numerous VANET research areas, one of the most valuable VANET applications is video streaming, which can provide drivers and passengers with more understandable and attractive information services. It can not only provide rich traffic information, such as real-time traffic conditions and accident alerts, but also support more diverse entertainment and commercial applications, such as Figure 1 As shown in the figure, people's in-vehicle experience is significantly improved. However, the key issues faced in VANET video transmission research include highly dynamic topology changes of vehicles, limited bandwidth of wireless network environments, and packet loss caused by unstable links, all of which make high-quality video transmission more challenging.

[0004] As the demand for video data continues to increase, existing work has also conducted some research on the transmission of video data in VANET. Almotairi et al. proposed an improved multipath video data routing algorithm for vehicle-mounted delay-tolerant networks, which evaluates routing paths based on link stability, available bandwidth and transmission delay and is used for the transmission of video frames of different priorities. Vafaei et al. proposed a QoS-aware video stream VANET routing algorithm based on ACO, which uses the ant colony optimization algorithm ACO to find primary and secondary paths based on global QoS, delay, packet delivery rate and other parameters, and transmits video frames of different priorities through TCP and UDP respectively. Bouzid et al. proposed LEQRV, which uses Q learning to evaluate link efficiency and estimate MOS scores, and improves the stability and reliability of video stream transmission links through distributed single-hop reinforcement learning. Since most of the existing video data routing algorithms use meta-heuristic algorithms for routing, these algorithms perform well when the network scale is small and the node mobility is not high, but in the face of highly dynamic VANET network environments, meta-heuristic algorithms need more time to converge to the optimal solution and may not be able to adapt to the rapid changes in network topology in time, thus affecting the real-time and effectiveness of the algorithm. Traditional RL methods are usually suitable for problems with smaller state spaces, but have poor generalization capabilities when faced with complex or high-dimensional state spaces. Summary of the invention

[0005] The purpose of the present invention is to provide an RSU-assisted video data routing method based on deep Q learning, which can effectively achieve better overall performance in scenarios with different vehicle densities and different network loads.

[0006] In order to achieve the above object, the technical solution adopted by the present invention is: an RSU-assisted video data routing method based on deep Q learning, comprising the following steps: Step 1: Establish a system model. The mobile source vehicle collects real-time traffic video, encodes it with SVC and forwards it to the RSU at the intersection that is wired to the traffic management department. Step 2: Describe the problem. Select appropriate paths for forwarding videos at different layers. Maximize the MOS metric for each GOP under the delay constraint. Describe the problem as P1: Step 3: The source vehicle selects the first RSU to which the forwarded data is to be sent based on the decision information of the RSU; Step 4: Based on deep Q learning, the next RSU between adjacent sections is selected as the forwarding direction for SVC video data of different layers on the RSU; Step 5: Forwarding within the road section selects the best relay vehicle based on the frame type based on the CIHI analytic hierarchy process, and makes multiple decisions between road sections and forwarding within the road section until the data reaches the destination RSU.

[0007] Preferably, in step 1, advanced video coding SVC is used to perform video encoding on the data.

[0008] Preferably, the SVC video stream includes a base layer BL1 and N-1 enhancement layers {EL2, EL3, L, EL N}.

[0009] Preferably, the problem is described as P1:

[0010] P1:max{MOS GOP},stD≤D MAX ..

[0011] Preferably, the step 5 comprises the following steps: Step 51, determine the goal, decompose the problem level to represent the goal, standard and possibility of alternative solutions; Step 52, listing the three-dimensional neutrosophic set and specifying the relative preference of the standard layer metric; Step 53: When the neutral intelligence set is used in the relay selection scheme, the three-dimensional neutral intelligence number is converted into a clear value to obtain a clear value matrix; Step 54: Calculate the weight of each metric according to the clarity value matrix, and calculate the average value of the row sums of the obtained pairwise comparison matrix; Step 55: consistency test of the pairwise comparison matrix; Step 56: Obtain the alternative scores of neighbor vehicles by multiplying each alternative by its corresponding weight relative to the corresponding standard, and select the vehicle with the largest score as the relay.

[0012] The beneficial effects of the present invention are:

[0013] 1. This solution ensures that the lower-layer video data is forwarded through reliable sections, thereby reducing the loss of basic layer data due to local optimization, while reducing transmission delay and improving adaptability to congested environments.

[0014] 2. The problem that the fixed scoring mechanism cannot meet the specific transmission requirements of different types of frames is solved, focusing on ensuring the reliable transmission of I frames and improving the decoding possibility of the video at the receiving end.

[0015] 3. Solve the problem that the existing video opportunistic routing algorithm has low adaptability to dynamic environments and poor generalization ability. Ensure the video quality in low congestion environments and the basic watchability of videos in high congestion environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0017] Figure 1 This is the application background diagram of video streaming transmission in existing VANET.

[0018] Figure 2 It is a frame structure diagram of the AVC and SVC video streams of the present invention.

[0019] Figure 3 This is a diagram of the urban VANET network model of the present invention.

[0020] Figure 4 A neural network architecture diagram of the road segment selection process based on deep Q learning of the present invention.

[0021] Figure 5 This is a hierarchical structure model diagram of the present invention.

[0022] Figure 6 It is a convergence performance diagram of MOS and delay metrics in a single group of GOPs of the present invention.

[0023] Figure 7 The figure is a performance comparison chart of the average frame delivery rate, end-to-end delay, peak signal-to-noise ratio, and mean opinion score of the present invention and the existing algorithm under different numbers of vehicles.

[0024] Figure 8 The figure is a performance comparison chart of the average frame delivery rate, end-to-end delay, peak signal-to-noise ratio, and mean opinion score of the present invention and the existing algorithm under different numbers of video streams. DETAILED DESCRIPTION

[0025] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0026] The present invention discloses a RSU-assisted video data routing method based on deep Q learning. The embodiment includes the following steps:

[0027] Step 1: Establish a system model. The mobile source vehicle collects real-time traffic video, encodes it with SVC and forwards it to the RSU at the intersection that is wired to the traffic management department. Step 2: Describe the problem. Select appropriate paths for forwarding videos at different layers. Maximize the MOS metric for each GOP under the delay constraint. Describe the problem as P1: Step 3: The source vehicle selects the first RSU to which the forwarded data is to be sent based on the decision information of the RSU; Step 4: Based on deep Q learning, the next RSU between adjacent sections is selected as the forwarding direction for SVC video data of different layers on the RSU; Step 5: Forwarding within the road section selects the best relay vehicle based on the frame type based on the CIHI analytic hierarchy process, and makes multiple decisions between road sections and forwarding within the road section until the data reaches the destination RSU.

[0028] Video Encoding

[0029] Video data routing usually requires video encoding of the data, which reduces the required transmission bandwidth by compressing the video data. In the VANET environment, due to the large changes in bandwidth-constrained network conditions, efficient encoding can significantly improve the efficiency and quality of data transmission. The H.264 standard, also known as Advanced Video Coding AVC, is the dominant video compression standard in the industry. This standard promotes lossy compression by exploiting the similarity between adjacent pixels and the movement of objects between consecutive frames. In AVC, a dependent frame structure called a group of pictures GOP is used. AVC includes three types of frames: intra-frames, forward prediction frames, and bidirectional prediction frames. Scalable video coding SVC is an extension of AVC and provides scalability of video content. SVC achieves temporal, spatial, and quality scalability. Through layered coding, the base layer supports the lowest quality video, while the upper layer can provide enhanced video quality on this basis. This layered structure greatly enhances the adaptability and flexibility of video streams, helps optimize the use of network resources, and improves user experience. The encoded video consists of a base layer BL and one or more enhancement layers EL. The data of the upper layer needs to rely on the lower layer for decoding. Although SVC maintains the same GOP hierarchy as AVC, it introduces a parallel structure based on BL+EL, such as Figure 2 , which is very suitable for dynamic scenarios such as VANET. The system allows the removal of higher layers while still retaining the ability to decode video.

[0030] Network Model

[0031] In the scenario of this embodiment, the mobile source vehicle needs to collect real-time traffic video, encode it with SVC, and forward it to an RSU at an intersection that is wired to the traffic management department through V2V, V2R, or R2V. The two-way road section is adjacent to the intersection, and the RSU connected to the cloud server is deployed at the intersection, such as Figure 3 Rsu at the two intersections P and Rsu Q , the section in between is Road PQ, RSU communicates with vehicles and assists in forwarding video data packets. Vehicles are equipped with GPS systems and digital maps, so that they can obtain the real-time geographic location of themselves and RSUs. The cloud server uses historical traffic flow information, network load, etc. from the data sections uploaded by RSUs to train the deep Q learning model. Vehicles and RSUs broadcast beacon packets at regular intervals, including information such as location and speed. When video data packets of different layers arrive at the intersection, RSU determines the next RSU between adjacent sections as the forwarding direction of the data packet through the maximum Q value output by the deep Q network, and at the same time notifies the corresponding section selection decision to the vehicle driving away from itself, so as to help the source vehicle and potential source vehicles that may participate in the video capture task in the future to better coordinate the forwarding of different layers of video data. The relay process in the section between two RSUs adopts a greedy forwarding strategy based on NS-AHP to select appropriate vehicle relays for I frames, P frames, and B frames. After several intersection RSU section selections and vehicle relays in the section, it finally reaches the destination RSU.

[0032] Video Quality Metrics

[0033] The SVC video stream consists of a base layer BL1 and N-1 enhancement layers {EL2, EL3, L, EL N}, in the scenario of this embodiment, there are three layers {BL1, EL2, EL3}, and in this embodiment, the adaptive selection of the video layer is not explored, so it is necessary to make the receiving end receive the video data of all layers as much as possible. The source vehicle converts the captured video into an SVC video stream. At the receiving end, MOS is used as a video quality metric. The video quality score ranges from 1 to 5. MOS measures MOS νideo Expressed in equation (1):

[0034] The model allows the MOS estimation of the video to be obtained by the frame loss rate flr and the video bit rate vbr (Mbps), where v1, v2, v3, and v4 are model-specific parameters to approximate the impact on the MOS value. In order to adapt to the scenario of this embodiment, they are fine-tuned: v1 = 4.0, 2 = 0.1, 3 = 3.6, and 4 = r. Since the base layer BL1 can be decoded independently, while the enhancement layer EL2 needs to rely on BL1, and the enhancement layer EL3 needs to rely on BL1 and EL2, if the receiving end receives the base layer BL1 very poorly, even if the enhancement layers EL2 and EL3 are perfectly received, good video quality cannot be obtained. Therefore, flr is expressed by equation (2):

[0035] in, is the delivery success rate of different layers. This formula can reflect the impact of the reception of different layers of SVC video on the quality of video reconstruction. Since the basic layer contains the key information necessary for decoding the entire video, its transmission path usually requires higher reliability.

[0036] Problem Description

[0037] In the scenario of this embodiment, the encoded SVC video needs to be delivered from the source vehicle to the destination RSU. By selecting a path with better quality or more stable to transmit the basic layer, the packet loss rate of key data can be significantly reduced to ensure the basic watchability of the video. The enhancement layer is used to improve the video quality. In the case of poor network conditions, it can be transmitted through other suboptimal or less used paths, which helps to reduce the load on the basic layer forwarding section. In addition, a suitable relay vehicle is selected for the key frame within the section to ensure its reception to improve the decoding effectiveness of the reference frame. The goal of this embodiment is to select the appropriate path for forwarding videos of different layers, and to ensure that the delay constraint D MAX Under this condition, the MOS metric of each GOP is maximized, and the problem is described as P1: P1:max{MOS GOP}, st D≤D MAX . (3)

[0038] The routing algorithm of the present invention is referred to as QRAVDR

[0039] The algorithm includes the following three key points: 1. Adaptive selection of First_Rsu by the source vehicle; 2. Selection of Next_Rsu between road sections based on deep Q learning; 3. Selection of relay vehicles within the road section based on NS-AHP.

[0040] Adaptive selection of source vehicle for First_Rsu

[0041] In the scenario of this embodiment, the first RSU that the video data generated by the source vehicle passes through during the forwarding process is defined as First_Rsu. Since the source vehicle switches the driving section during the movement, it is necessary to constantly adjust the source vehicle's selection of First_Rsu. An unreasonable selection will lead to invalid data transmission, increase unnecessary forwarding hops, and increase transmission delay. Therefore, it is very important to coordinate the forwarding direction of different layers of video data by the source vehicle using the RSU's section forwarding decision information for different video layer data.

[0042] To solve the above problem, if the source vehicle v s In Rsu i Within the communication range and v s Towards Rsu i If driving, select Rsu directly. iAs the First_Rsu for BL1, EL2, EL3 video layer data forwarding, this is temporary, and the source vehicle will switch the driving section soon. i When receiving the source vehicle group, the deep Q network model selects the appropriate road section for different video layers and forwards it to Next_Rsu. s The driving section is from Road ij Transformed to Road in , v s Divergence from Rsu i Driving, then according to the Rsu i The received Next_Rsu decision information for BL1, EL2, EL3 is used to adjust First_Rsu. If Next_Rsu = Rsu exists in BL1, EL2, EL3 n If the decision information is received, the source vehicle will synchronize the First_Rsu of the corresponding video layer transmission to Rsu n , in order to reduce certain invalid forwarding. In addition, the source vehicle v s The interaction flag Mark with the RSU is set to 1. If the source vehicle does not update the interaction information with other RSUs within the interaction validity period, Mark is set to 0. This ensures that the source vehicle leaves the RSU. i After the communication range is reached, the subsequent selection of First_Rsu can rely on the DQL decision from each RSU to ensure the rationality of the selection of First_Rsu.

[0043] If the source vehicle has not interacted with any RSU in the near term before it starts forwarding video data, and the source vehicle is not within the communication range of any RSU, it cannot rely on the decision-making knowledge of the RSU. In this case, the source vehicle needs to make a subjective selection of the First_Rsu. The road topology is obtained from the digital map, and the positions of the vehicle itself, the RSUs at both ends of the road section, and the destination RSU are obtained through GPS. The distance between the RSUs at the intersections at both ends and the destination RSU and the cosine of the source vehicle's motion angle are compared. The RSUs with the larger weight are selected from the RSUs at the intersections at both ends as the First_Rsu for forwarding BL1, EL2, and EL3 video layer data. The weight calculation equation is shown in equation (4):

[0044] in is the source vehicle v s Rsu at the intersections at both ends of the road section i and destination roadside unit Rsu dThe Manhattan distance between them is because the Manhattan distance is more suitable than the Euclidean distance to characterize the path of the routing process in urban traffic; S is the path of two Rsu i with Rsu d The sum of the Manhattan distances between them is used to normalize the metric; is the source vehicle v at time slot t s The displacement vector of Yes s To the intersections at both ends i The position vector of . Reflects v s The driving direction and Rsu i The relationship between the positions, if the vehicle is moving towards Rsu i If the value of the cosine in the equation is positive, then it is negative if the value of the cosine in the equation is positive; otherwise, it is negative. ω is a weighting factor between 0 and 1.

[0045] After the source vehicle selects First_Rsu through the above process, it directly forwards the video data of different layers BL1, EL2, and EL3 to BL1_First_Rsu, EL2_First_Rsu, and EL3_First_Rsu respectively or selects a suitable relay vehicle to forward it to them.

[0046] Selection of the next RSU between road segments based on deep Q-learning

[0047] When a data packet arrives at an RSU, if the RSU is not the destination RSU of the video data packet, the deep Q network trained based on the traffic flow information and resource utilization status of the road segment estimates the Q value of the adjacent road segment, and then determines the forwarding sections of the video data at different layers, and forwards the video frame to the corresponding Next_Rsu. The reinforcement learning model and the Next_Rsu selection algorithm will be described.

[0048] Reinforcement Learning Model

[0049] This process can be modeled as a reinforcement learning problem of Markov process MDP, which learns the road segment selection strategy by observing the network and traffic status of the road segment and constantly interacting with neighboring vehicles. The definition of each element in the DRL agent is described below.

[0050] Denote the state space as S:{s1,s2,...,s N}, N represents the total number of RSUs. Due to the dynamic changes in network traffic at different times, the state changes over time. M The decision of is determined by the information of adjacent segments and is defined as follows:

[0051] Layer M Indicates that the currentM The video layer to which the data to be forwarded belongs, Vehicles MN _Ava is the road section MN In Rsu M The number of vehicles within the communication range directly determines whether the data can be forwarded to the road section immediately. M→N It is Rsu M To adjacent intersection Rsu N Number of vehicles in the direction, Vehicle N→M is the number of vehicles in the opposite direction, ρ MN Indicates road segment MN The vehicle density on the network is due to the fact that the link connection of vehicles traveling in the opposite direction is more unstable than that of vehicles in the same packet forwarding direction

[16] . Therefore, ρ MN The definition of is shown in equation (10):

[0052] Where Length MN For Road MN Length.

[0053] When Rsu M Received from the adjacent road segment Road MN When grouping, Road MN _out M Increase; and when Rsu M Forward the packet to the road MN When MN _in M Increase, when the RSU will forward a packet along one of the roads, the average road load The calculation of is shown in equation (11):

[0054] Among them, Road MN _in M and Road MN _in N From Rsu M and Rsu N Enter the road MN Number of groups, Road MN _out M and Road MN _out N It is Rsu M and Rsu N From Road MN The number of packets received. M→N +VehicleN→M Representative Road MN The total number of vehicles in the road. The road load can reflect the network congestion status on the road.

[0055] The action space is defined as A:{a1,a2,...,a M}, when Rsu M When receiving a packet, it is necessary to select the RSU of the adjacent intersection and forward it to it, so RSU M The optional action is represented by a M :{x|x∈Nbr(Rsu M )}. Therefore, for each RSU, the size of its action space is equal to the number of RSUs in its neighboring sections. Once the data packet arrives at RSU at time step t M , according to the state and in a t ∈a M Select an action and send the current data packet to the selected RSU through V2V transmission within the road section.

[0056] The reward function is used to guide the agent to formulate an effective strategy for the goal: to improve the MOS metric of the video as much as possible under the condition of delay constraints. Due to the dependencies between different video layers, the reward is related to the reception of different video layers, the delay, and the distance from the target RSU. The definition of the reward function is shown in equation (12): If reaching Rsu d ,ψ=1,otherwise,ψ=0

[0057] Among them, MOS GOP (D MAX ) is the maximum tolerable delay D for a group of GOPs using equation (1). MAX The quality estimate obtained under the scenario D MAX =3s,Conrti Layer is the contribution coefficient of different video layers to the estimated quality based on the forwarding success rate, Frame_num Layer Represents the number of video frames received by a group of GOP corresponding to the video layer, dist t is the Manhattan distance between the current RSU and the destination RSU, dist t+1 is the Manhattan distance between the RSU at the next intersection and the destination RSU, dist SD is the Manhattan distance between the source vehicle and the destination RSU. D is the delay from the current RSU to the next RSU, Size GOPIt is the number of frames contained in a group of GOPs. This means that if you want to maximize the reward, it depends not only on the reception of a single video frame, but also on the overall reception performance of each video layer and the dependencies between them. When the overall reception is good, a higher MOS score will be obtained. In addition, if you want to get a higher reward, it also depends on the reception of the lower video layer. Therefore, driven by the reward, the agent will pay more attention to the reception of the lower video layer and reduce the transmission delay to achieve the goal.

[0058] The neural network architecture of the Next_Rsu selection process based on deep Q learning is as follows Figure 4 As shown. The agent uses two neural networks, namely the online network and the target network. The online network estimates the Q value Q of taking action a in the current state s online (s, a), and the target network outputs the Q value Q of taking action a' in the next state s' target (s′, a′). The online network is trained at each learning step to reduce the loss function. The target Q value and loss function are calculated as shown in Equation (16) and Equation (17): Loss=(T DQN -Q online (s,a)) 2 (17)

[0059] T DQN is the target Q value, which is determined by the immediate reward R t , the target network predicts the discounted value of the maximum Q value of the next state s′, and the loss function Loss reflects the difference between the target Q value and the current Q value predicted by the online network.

[0060] The discount factor γ determines the importance that the agent attaches to immediate rewards and future rewards. A lower discount factor means that the agent attaches more importance to immediate rewards, while a higher discount factor means that the agent attaches more importance to long-term rewards. In the traditional Q-learning algorithm, the discount factor γ is a fixed value between 0 and 1. In the scenario, in order to make the agent make more reliable decisions, the discount factor γ is defined as a dynamic parameter related to the distance to the destination RSU, as shown in equation (18):

[0061] where dist t is the Manhattan distance between the current RSU and the target RSU, dist t+1 is the distance between the next RSU at the intersection and the destination RSU. If the next RSU is farther from the destination RSU than the current RSU, then γ tThis dynamic discount factor makes the reliability of future rewards poor when the data packet is far away from the destination due to the dynamic network topology, and the reliability of future rewards gradually increases as the data packet approaches the destination.

[0062] At the beginning of the learning process, the weights of the target network and the online network are the same. The weight values ​​of the target network are not updated immediately to enhance learning stability. During the training phase, the weights of the target network are regularly updated to match the online network after a predetermined number of learning steps. The target network is used to provide a relatively stable target value for calculating the loss function and guiding the update of the online network, which helps avoid fluctuations in the learning process caused by constant changes between the target and the estimated value.

[0063] Experience replay is a key technology in DQN to improve the effect and stability of training. At each time step, the agent performs an action, observes the new state and reward, and (S t ,A t ,R t ,S t+1 ) is stored as an experience tuple in the replay buffer. The replay buffer has a certain size limit. When it is full, the new experience will replace the oldest experience. During the training process, the algorithm randomly extracts a batch of experience from the replay buffer, which helps to break the temporal correlation between the data. Then, this batch of experience is used to update the parameters of the Q network through the gradient descent method, reducing the variance of the loss function and improving the stability of training.

[0064] Exploration refers to the agent trying actions that have not been tried before or have been tried rarely to understand the potential rewards of these actions. Exploration is very important for discovering the optimal strategy, especially in the case of incomplete or dynamically changing environmental information, and in the scenario of being able to adapt to dynamic vehicle self-organizing networks. However, too much exploration may cause the agent to spend too much time trying low-return actions before finding an effective strategy. Exploitation refers to the agent using its existing knowledge to select actions that are believed to bring the greatest reward. The ε-Greedy strategy is adopted in the scenario, which uses an adjustable parameter ε∈[0,1] to determine the probability of the agent exploiting and exploring. For each learning step, a random value τ∈[0,1] is generated. If τ<ε, the agent will explore, otherwise the agent will exploit. The DQN agent selects action A according to equation (19):

[0065] Gradually reducing ε in the ε-Greedy strategy can increase the agent's exploration opportunities in the early stages of learning, and gradually transition to using existing experience to make reasonable decisions over time. Use a fixed decay rate decr to control the decrease of ε, and the ε value decreases from the maximum value ε max It starts with , and decreases linearly with the decay rate decr at the end of each training round until it reaches the minimum value ε min , ε is updated at the end of each round using equation (20): ε=ε max -round×decr(20)

[0066] Next_Rsu selection algorithm based on DQN

[0067] The algorithm implements a learning process to find the best forwarding segments for different layers of video data, as shown in Algorithm 1. First, the weights θ and θ of the online and target networks are initialized. And replay the experience pool, and then learn through the loop from line 2 to line 21. In each cycle, the DQN agent completes a cycle and the state S t To state S t+1 , where the inner loop executes as follows: First, the DQN agent updates ε with a decay rate of decr and selects action A from the action space based on Eq. (19) t , and then calculate the reward R in the environment through Eq. (12) t , the DQN agent also obtains the next state S t+1 , experience exp t =(S t ,A t, S t+1 ,R t ) is saved in the replay experience pool. If the current step number has reached the replay start threshold, the DQN agent randomly selects a mini-batch of data from the replay experience pool to train the online network and estimate Q(S t ,A t ), and use the target network to calculate T DQN (S t ,A t ), adjust the weights and biases of the online network through back-propagation and gradient descent to minimize the loss calculated by Eq. (17). The current step number reaches the target network update frequency f up , the DQN agent replaces the target network’s weights with the online network’s weights θ The DQN agent transitions to the next state.

[0068] Relay vehicle selection within a road section based on the analytic hierarchy process of CIGI

[0069] When the video frame does not reach the destination RSU and needs to be forwarded to First_Rsu or Next_Rsu, the vehicles in the road section need to relay. Selecting the most suitable vehicle relay for I-frames, P-frames and B-frames is the key to improving transmission efficiency and ensuring video quality. Since I-frames carry complete picture information and are the basis for decoding other frames, the accuracy and reliability of their transmission are crucial. P-frames and B-frames contain differential information and rely on previous or subsequent frames for decoding. The proposed method uses four routing metrics to help the vehicle or RSU carrying the packet to select the next-hop vehicle. The NS-AHP process will use these routing metrics to calculate the weights of candidate relay vehicles in the road section. Finally, if the node carrying the video frame does not have a suitable neighbor vehicle to relay, it will continue to carry the data until a neighbor vehicle with a larger weight is identified.

[0070] Definition of relay selection indicators

[0071] The distance DM from neighbor vehicle n to Frist_Rsu or Next_Rsu on the same road segment as the current vehicle nr :

[0072] Among them, ED nr is the Euclidean distance between n and Frist_Rsu or Next_Rsu, ED current is the Euclidean distance from the current vehicle to Frist_Rsu or Next_Rsu, ED nr_min is the minimum value among the distances defined above.

[0073] Link expiration time LET mn :

[0074] The high-speed movement of vehicles in VANET is the source of rapid changes in topology, resulting in frequent link disconnections. To calculate the link expiration time LET, the proposed scheme uses the equation in. LET is considered to be the time that two vehicles remain connected. Let (x m ,y m ) and (x n ,y n ) is in the direction θ m ,θ n The coordinates of the two vehicles m and n moving on the m ,θ n ≤2π, respectively with speed v m and v n , r is the communication range. mn Equation (22) is used for estimation:

[0075] Where a = v m cosθ m -v n cosθ n ,b=x m -x n ,c=v m sinθ m -v n sinθ n ,d=y m -y n .

[0076] Vehicle available buffer metric BM:

[0077] Buffer n is the available buffer size for vehicle n, Buffer max and Buffer min are the maximum and minimum values ​​of the available buffer between the current vehicle and n.

[0078] Available Bandwidth Estimation ABE:

[0079] The effective transmission of video streams depends heavily on the available bandwidth in the communication channel. In an urban scenario, the collision probability of forwarding vp-bit video packets at a given average speed sp and number of vehicles vn is expressed by equation (24) according to the research of Tripp-Barba et al.: prob(vp,vn,sp)=f(vp)×prob hello (vp,vn,sp)(24)

[0080] where f(vp) is the Lagrange interpolation polynomial obtained from the simulation offline. hello (vp,vn,sp) is the probability of Hello message collision. The additional overhead introduced by the binary exponential backoff mechanism (K) is calculated as follows:

[0081] Where D diff is the time difference between the transmission of two video frames, DIFS is the distributed inter-frame space, which represents the fixed time interval for a vehicle node to access the channel medium when the channel medium is idle for no longer than DIFS. is the average backoff for sending a single frame.

[0082] The sending vehicle or RSU calculates the ABE of the wireless link of each neighboring vehicle in the same road segment n(0≤ABE n ≤1), as shown in equation (26): ABE n =(1-K)×(1-prob(vp,vn,sp))×ID rv ×ID sn ×C(26)

[0083] Where ID rv and ID sn They represent the idle duration ratios of the receiving vehicle and the sending vehicle estimated by monitoring the channel, and C is the link capacity.

[0084] Relay vehicle selection based on NS-AHP

[0085] Step 1: Determine the goal; break down the problem hierarchy to represent the goals, criteria, and possibilities of alternative solutions, such as Figure 5 shown. Table 4.1 Logical terms and corresponding three-dimensional neutrosophic numbers

[0086] Step 2: The three-dimensional neutrosophic set is shown in Table 4.1, which is used to specify the relative preference of the standard layer metric, where the pairwise comparison matrix is ​​expressed as:

[0087] The matrix It represents the result of the k-th decision maker scoring the relative importance of the criteria according to his own judgment. The form of the three-dimensional neutron number is as follows in are the lower bound, median value and upper bound of the neutrosophic number, which represent the possible range of the decision maker's assessment of the importance of criterion i relative to criterion j. is the minimum possible value, is the maximum possible value, and is the most likely value. They are respectively the truth, uncertainty and falsity of the neutron number in three dimensions.

[0088] In the scenario, since the I frame contains a complete data image and is the most important frame in the video stream, link stability is extremely important. In addition, the amount of data in the I frame is large, so a higher bandwidth and available buffer are required to ensure its successful transmission. Although the P frame size is smaller than the I frame, link stability is still an important factor in the transmission of P frames. It also requires a certain amount of bandwidth to support it, and the requirements for the available buffer are relatively loose. For B frames, the importance of link stability and distance measurement is obviously greater than bandwidth and available buffer. Selecting a farther neighbor to forward the B frame each time can make full use of vehicles with weaker link stability to participate in forwarding, and can also reach the destination with fewer hops. Based on the above reasons, the pairwise comparison matrix is ​​defined as follows:

[0089] Convert the comparison matrix into a three-dimensional neutrosophic set:

[0090] Step 3: In the case of using the neutrality set in the relay selection scheme, the three-dimensional neutrality number is converted into a crisp value based on equation (28).

[0091] This gives the clarity value matrix:

[0092] Step 4: According to the clarity value matrix, the weight of each metric is calculated according to the following steps, and the average of the row sums of the obtained pairwise comparison matrix is ​​calculated, as shown in equation (29):

[0093] And the weight w i Normalization:

[0094] The weights of different metrics for I, P, and B frame transmission can be obtained by calculation, which are

[0095] Step 5: Consistency test of the pairwise comparison matrix. The consistency ratio CR is calculated as shown in equation (31):

[0096] Among them, CI is the consistency index and RI is the random consistency index. Table 4.2 Random consistency index (RI) of various standards N 3 4 5 RI 0.58 0.90 1.12

[0097] Where n is the number of judgment criteria. RI depends on the number of judgment criteria. In the scenario, there are 4 judgment criteria. According to Table 4.2, RI = 0.9. If CR ≤ 0.1, the pairwise comparison matrix is ​​consistent; otherwise, it must be redefined. By calculation, the pairwise comparison matrix A can be obtained respectively. I ,A P ,A B The consistency ratio is:

[0098] All passed the consistency test of the pairwise comparison matrix

[0099] Step 6: Obtain the alternative score of neighbor vehicle n by multiplying each alternative by its corresponding weight relative to the corresponding criterion, calculated as shown in Equation (34), where δ = I, P, B, and select the vehicle with the largest score as the relay.

[0100] QRAVDR routing process

[0101] Assume that the source vehicle v s The captured video data has been SVC encoded, and the video frame to be sent is encapsulated and prepared for forwarding. The complete routing process will be described, and the specific process is shown in Algorithm 2. First, the source vehicle determines frame_first_Rsu for the frame to guide the forwarding direction, and directly delivers it or selects a suitable relay vehicle for the frame through NS-AHP to indirectly deliver it to frame_first_Rsu. When the frame arrives at the RSU, if the current RSU is not the destination RSU, the trained deep Q learning model is used to determine frame_Next_Rsu to guide the forwarding direction, and then a suitable relay vehicle is selected based on NS-AHP to forward it until it reaches frame_Next_Rsu. The frame undergoes several frame_Next_Rsu selections and corresponding intra-section forwardings until it reaches the destination RSU.

[0102] Simulation and performance analysis

[0103] The performance of the proposed scheme is evaluated through computer simulation. The proposed routing algorithm QRAVDR is implemented on the network simulator platform NS-2.34. The urban transport tool SUMO is used to simulate vehicle movement, and the reconstructed video quality at the receiving end is evaluated by Evalvid.

[0104] Experimental environment

[0105] For the simulation area, the electronic map topology and data were obtained from OpenStreetMap contributors. The selected area is a 1900m×2100m urban area, which includes 19 intersections and 29 two-way lanes. The maximum speed of vehicles in each section is 100km / h. The communication range between the vehicle and the RSU is set to 250 meters, and SUMO is used to generate the vehicle mobility model. The highway.yuv from the Xiph video repository is used as the video for simulation transmission, and the JSVM 9.19.15 encoder is used to encode the video into a three-layer structure, where the bit rates of BL1, EL2, and EL3 are 77.97kbps, 116.71kbps, and 203.64kbps, respectively. More detailed simulation parameter information is provided in Table 5.1, which includes the configuration of key parameters such as urban simulation area, number of vehicles, vehicle speed, media access control MAC protocol, and video information. Table 5.1 Simulation parameter configuration Parameter Values Urban simulation area <![CDATA[1900×2100m 2 ]]> Number of vehicles 100-500 Transmission range 250m Maxvehiclespeed 100km / h Mobility generator SUMO MAC protocol 802.11p Videofile highway.yuv Videoresolution 352×288 Length 2000frames Video encoding JSVM Number of Layers 3 TransmissionProtocol UDP Simulation time 100s

[0106] Performance Evaluation

[0107] In order to compare the performance of the proposed algorithm with the existing algorithms RLOR-AOMDV, EPSO-MS and MNH-FGR, the following metrics are considered.

[0108] Frame delivery rate FDR: the ratio of the number of video frames successfully received by the destination RSU to the total number of video frames sent from the source vehicle.

[0109] Average end-to-end delay: The average time it takes for a video frame to be forwarded from the source vehicle to the destination RSU.

[0110] Peak signal-to-noise ratio (PSNR): One of the most widely used indicators for measuring the quality of video after transmission, reflecting how close the reconstructed video image at the receiving end is to the original video image.

[0111] Mean Opinion Score (MOS): is a subjective video quality assessment method that uses equation (1) to estimate the entire video segment, which is determined by model parameters, encoding bit rate, and equivalent frame loss rate.

[0112] The results are averaged through multiple simulations as the performance of these algorithms in multiple indicators.

[0113] Convergence analysis of the algorithm

[0114] First, this embodiment Figure 6 The figure shows the convergence performance of QRAVDR in the MOS metric of a single group of GOPs, where the dynamic discount factor related to the distance to the destination RSU reflects the model's preference for current and future rewards during training. For example, 500 rounds of training are used in a scenario with 200 vehicles to demonstrate the training effect.

[0115] from Figure 6 It can be seen that the MOS value gradually converges with the increase of training rounds, which reflects the effectiveness of the proposed method. It can be observed that in the early stage of training, the MOS metric has a large oscillation amplitude. This is because in the early stage of training, the algorithm tends to explore more possible behaviors to understand the environment, resulting in large fluctuations in decision quality. As learning progresses, MOS grows rapidly from a low value, and the model performance is significantly improved. And as the number of cycles increases further, it gradually stabilizes and fluctuates around 3.0 overall. Figure 6 b reflects the convergence performance in terms of latency. Compared with the convergence process of MOS values, the MOS metric has entered the convergence stage when the number of iterations is 200, while it is still in the latency optimization stage. The latency did not enter the convergence stage until the number of iterations was about 250, and the overall fluctuation was between 0.6 and 0.8 seconds. At this time, a near-optimal solution has been found. In addition, the architecture of the online network and the target network helps to reduce the fluctuation of the learning target value, thereby making the entire learning process more stable.

[0116] Comparative analysis of performance under different vehicle densities

[0117] By increasing the number of simulated vehicles from 100 to 500, the Average FDR, Average Delay, Average PSNR and Average MOS of the proposed algorithm as well as RLOR-AOMDV, EPSO-MS and MNH-FGR are measured under different vehicle densities through multiple simulations, and the simulation results are discussed and analyzed. Figure 7 7a, 7b, 7c, and 7d respectively show the performance of these algorithms in multiple indicators.

[0118] Figure 7a shows the comparison results of the average frame delivery rate. As the number of vehicles increases, QRAVDR has the best overall performance. When the number of vehicles is 500, the frame delivery rate is 85.94%, which is 6.38%, 7.96%, and 8.95% higher than EPSO-MS, RLOR-AOMDV, and MNH-FGR, respectively. This shows that the decision accuracy of QRAVDR is not only improved with the increase in the number of vehicles, but also can adapt to the dynamically changing network environment through the decision-making mechanism based on reinforcement learning. In addition, based on NS-AHP, the appropriate relay vehicle is selected according to different types of video frames, which further ensures the successful delivery of video frames. EPSO-MS performs second best, but its performance is better than QRAVDR when the number of vehicles is 100. Their delivery rates are 73.43 and 71.01, respectively. This is because EPSO-MS uses the particle swarm optimization algorithm to better adapt to the environment with fewer vehicles and thus improve the delivery rate of data packets. Compared with QRAVDR, RLOR-AOMDV, which is also based on reinforcement learning decision-making mechanism, performs poorly. This is because QRAVDR uses relevant parameters such as video quality and delay as reward functions to guide learning, while RLOR-AOMDV only uses delay parameters for learning, which directly affects the performance of RLOR-AOMDV in terms of delivery rate. Since MNH-FGR mainly evaluates candidate relay vehicles through set rules and membership functions, this relatively static decision-making mechanism is prone to fall into local optimal solutions in the face of rapidly changing network conditions and complex decision-making environments, which leads to its low delivery rate.

[0119] Figure 7 b describes the average end-to-end delay. It can be seen that when the number of vehicles is 100, the average delays of QRAVDR, RLOR-AOMDV and EPSO-MS are all between 1.6 and 1.8 seconds, and the average delay of MNH-FGR is 2.73 seconds, which are relatively high in general. This is because the vehicle density is low, which requires more vehicles to carry video frames during the forwarding process of the road section until a suitable relay vehicle is encountered. However, EPSO-MS performs best at this time, because when the number of nodes is small, the particle swarm optimization algorithm can solve the routing strategy faster and use it for forwarding video frames. As the number of vehicles increases, RLOR-AOMDV performs best in terms of average delay, which is related to its use of a delay-based reward function. QRAVDR performs second best. The average delay of EPSO-MS reaches the lowest value of 0.75 seconds when the number of vehicles is 300, but as the number of vehicles continues to increase, the average delay increases instead. As the number of particles increases, the calculation of the information update of each particle in each iteration becomes more complex and time-consuming, and the increase in the solution space prolongs the time to converge to the optimal solution. The latency of MNH-FGR's relatively static routing strategy gradually decreases as the number of vehicles increases, but its overall performance is still poor.

[0120] Figure 7 c shows the results of average PSNR. Since QRAVDR and EPSP-MS are more focused on the reception of base layer video frames in their design concepts, their average PSNR performance is better than RLOR-AOMDV and MNH-FGR overall. Since EPSP-MS has a higher frame delivery rate than QRAVDR when the number of vehicles is 100, it also successfully obtains a higher average PSNR (28.81). Although the frame delivery rate cannot completely determine the video PSNR at the receiving end, it still shows that EPSP-MS is better than QRAVDR in this case. However, as the number of vehicles increases, QRAVDR can adapt to the environment with a large number of vehicles by adaptively adjusting the routing decision through reinforcement learning, and achieve the reception of base layer video frames as much as possible to receive more enhanced layer video frames to improve the quality, and use NS-AHP to select more suitable relay vehicles for different types of frames to ensure the transmission of key frames, which makes the performance of QRAVDR steadily improve, and the average PSNR reaches 32.65 when the number of vehicles is 500. However, EPSP-MS still shows the phenomenon that the average PSNR increases first and then decreases as the number of vehicles increases. RLOR-AOMDV and MNH-FGR obtain lower average PSNR because they lose more key frames, resulting in the inability to successfully decode the related video frames.

[0121] Figure 7 d describes the average MOS. Since MOS is largely affected by the reception of the base layer, this shows that as the number of vehicles increases, QRAVDR can adaptively rely on the decision of reinforcement learning to stably maintain the reception of the base layer video frames, and improve MOS by receiving as many enhanced layer video frames as possible. When the number of vehicles is 500, it is 0.46, 0.87, and 0.93 higher than EPSO-MS, RLOR-AOMDV, and MNH-FGR, respectively. EPSP-MS performs second best, but it can still maintain MOS around 3.0 by maintaining a relatively stable reception of the base layer video frames when the vehicle density is high. Although RLOR-AOMDV and MNH-FGR show a continuous increasing trend in MOS, they perform poorly in the reception of the base layer video frames, which naturally leads to a lower MOS.

[0122] Performance comparison analysis under different network loads

[0123] In order to compare the performance of each algorithm under different network loads, the number of source vehicles is gradually increased and the Average FDR, Average Delay, Average PSNR and Average MOS of these algorithms are measured through multiple simulations. The simulation results are discussed and analyzed. In this scenario, when the number of vehicles is fixed at 300, Figure 8 The performance of these algorithms in terms of multiple indicators as the network load increases is shown in Figure 2.

[0124] from Figure 8 As can be seen in a, as the number of video streams increases, the average frame delivery rate of all algorithms shows a certain downward trend. This is because the transmission of multiple video streams will compete for limited bandwidth resources at the same time, and the limited buffer capacity will make it impossible for the vehicle to store all the data packets waiting to be forwarded, resulting in some data packets being discarded, thereby reducing the delivery rate. However, the proposed QRAVDR performs best, and when the number of video streams is 10, it is 9.38%, 12.4% and 16.18% higher than EPSO-MS, RLOR-AOMDV and MNH-FGR respectively. This is because it can effectively identify congested sections through reinforcement learning, and consider the importance of available buffers and bandwidth for video frame forwarding in the process of relay vehicle selection. EPSO-MS performs suboptimally, because as the network complexity increases, the particle swarm optimization algorithm easily converges to a suboptimal solution and it is difficult to find the global optimal path in time. Although RLOR-AOMDV also performs routing based on reinforcement learning, it only focuses on optimizing latency, resulting in a lower frame delivery rate. MNH-FGR uses fuzzy logic to make decisions based on local information and empirical rules, which makes it difficult to find the global optimal path in a complex network environment, resulting in poor delivery rate.

[0125] Figure 8 b shows the relationship between the average end-to-end delay and the network load, where RLOR-AOMDV has the lowest average delay because its learning strategy only optimizes the delay parameter and ignores other parameters. QRAVDR performs second best. The penalty parameter based on the maximum tolerant delay helps to focus on delay control while optimizing the video quality at the receiving end. EPSO-MS further reduces its performance in terms of delay due to the increase in solution time. MNH-FGR cannot consider the comprehensive network status, resulting in higher transmission delay.

[0126] from Figure 8As can be seen in c, as the network load increases, QRAVDR can avoid more congested sections to ensure the reliable transmission of the base layer video frames, so that the reconstructed video at the receiving end maintains a high PSNR. When the number of video streams is 10, it is 1.02, 5.99 and 6.51 higher than EPSO-MS, RLOR-AOMDV and MNH-FGR respectively. EPSO-MS also transmits base layer video frames by selecting a more reliable transmission path, but its adaptability to congested networks is worse than QRAVDR. RLOR-AOMDV and MNH-FGR lose more key frames, so the video frames that need to be decoded cannot function, resulting in a lower PSNR.

[0127] Figure 8 d describes the average MOS. Despite the increase in the number of video streams and the increase in network congestion, QRAVDR is still able to ensure the basic watchability of the video by maintaining stable reception of basic layer video frames, controlling the rapid decrease of MOS. When the number of video streams is 10, it is 0.11, 0.53, and 0.61 higher than EPSO-MS, RLOR-AOMDV, and MNH-FGR, respectively. EPSP-MS performs second best. Although it can also focus on the reception of basic layer video frames under network congestion conditions, its adaptability to complex environments is not as good as QRAVDR. RLOR-AOMDV and MNH-FG obtained lower MOS, which is related to the fact that they lost more basic layer video frames.

[0128] It should be noted that the parts not described in detail in the above embodiments are all prior art.

[0129] The above description is only a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed as a preferred embodiment as above, it is not intended to limit the present invention. Any technician familiar with the profession can make some changes or modifications to equivalent embodiments of equivalent changes using the technical contents disclosed above without departing from the scope of the technical solution of the present invention. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention should be covered within the protection scope of the present invention.

Claims

1. A RSU-assisted video data routing method based on deep Q learning, characterized in that: The following steps are involved: Step 1: Establish a system model. The mobile source vehicle collects real-time traffic video, encodes it with SVC and forwards it to the RSU at the intersection that is wired to the traffic management department. Step 2: Describe the problem. Select appropriate paths for forwarding videos at different layers. Maximize the MOS metric for each GOP under the delay constraint. Describe the problem as P1: Step 3: The source vehicle selects the first RSU to which the forwarded data is to be sent based on the decision information of the RSU; Step 4: Based on deep Q learning, the next RSU between adjacent sections is selected as the forwarding direction for SVC video data of different layers on the RSU; Step 5: Forwarding within the road section selects the best relay vehicle based on the frame type based on the CIHI analytic hierarchy process, and makes multiple decisions between road sections and forwarding within the road section until the data reaches the destination RSU.

2. According to claim 1, a RSU-assisted video data routing method based on deep Q learning is characterized in that: In the step 1, advanced video coding (SVC) is used to perform video encoding on the data.

3. The RSU-assisted video data routing method based on deep Q learning according to claim 1, characterized in that: The SVC video stream includes a base layer BL1 and N-1 enhancement layers {EL2, EL3, L, EL N }.

4. The RSU-assisted video data routing method based on deep Q learning according to claim 1, characterized in that: The problem is described as P1: P1:max{MOS GOP },s.t.D≤D MAX .。 5. The RSU-assisted video data routing method based on deep Q learning according to claim 1, characterized in that: The step 5 comprises the following steps: Step 51, determine the goal, decompose the problem level to represent the goal, standard and possibility of alternative solutions; Step 52, listing the three-dimensional neutrosophic set and specifying the relative preference of the standard layer metric; Step 53: When the neutral intelligence set is used in the relay selection scheme, the three-dimensional neutral intelligence number is converted into a clear value to obtain a clear value matrix; Step 54: Calculate the weight of each metric according to the clarity value matrix, and calculate the average value of the row sums of the obtained pairwise comparison matrix; Step 55: consistency test of the pairwise comparison matrix; Step 56: Obtain the alternative scores of neighbor vehicles by multiplying each alternative by its corresponding weight relative to the corresponding standard, and select the vehicle with the largest score as the relay.

Citation Information

Patent Citations

  • Method for transmitting SVC video in Internet of Vehicles

    CN110418143A

  • Unmanned aerial vehicle assisted Internet of Vehicles real-time video transmission method based on deep reinforcement learning

    CN117857737A

  • A Guide to drum

    KR1020220153444A