Optimization method, device and system based on communication-aware integrated system and readable storage medium
By combining terahertz communication and hybrid precoding with deep reinforcement learning and meta-learning, the analog precoding and power allocation strategies of vehicles are optimized, solving the problem of communication and perception synchronization optimization in vehicle-to-everything (V2X) systems under dynamic environments, and improving spectrum utilization and system performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- UNIV OF SCI & TECH BEIJING
- Filing Date
- 2023-03-30
- Publication Date
- 2026-05-01
AI Technical Summary
Existing integrated communication and sensing systems struggle to achieve efficient communication and synchronous optimization between vehicles and base stations in dynamic environments, failing to meet the ultra-low latency requirements of 6G. Furthermore, traditional resource optimization methods are static and alternating, making it impossible to effectively integrate communication and sensing system metrics.
By employing terahertz communication and hybrid precoding technology, combined with deep reinforcement learning and meta-learning, the TD3 neural network optimizes the vehicle's analog precoding and power allocation strategies, thereby achieving intelligent joint resource optimization in dynamic environments.
This study improved the spectrum utilization of the vehicle-to-everything (V2X) system, reduced latency, optimized the overall performance of the communication and sensing systems, solved the resource allocation challenges in dynamic environments, and analyzed the signal interference and power constraints between vehicles and between vehicles and facilities.
Smart Images

Figure CN116801285B_ABST
Abstract
Description
Optimization methods, devices, systems, and readable storage media based on integrated communication and sensing systems. Technical Field
[0001] This invention relates to the field of integrated communication and sensing technology, and in particular to an optimization method, apparatus, system and readable storage medium based on an integrated communication and sensing system. Specifically, it relates to an optimization method, apparatus, system and readable storage medium based on an integrated communication and sensing system for intelligent joint resource optimization in dynamic environments. Background Technology
[0002] With the development of smart cities, smart transportation, and intelligent industry, connected and autonomous driving will dominate. In the future vehicle-to-everything (V2X) architecture, the network must not only possess high-speed data transmission capabilities but also have perception capabilities to provide accurate positioning services for vehicles, thus developing safer and more intelligent autonomous driving. The evolution from radar communication, joint radar communication, and dual-function radar communication to the more extensive integrated communication and perception technology will further support the development of future mobile communication networks. Integrated communication and perception highly integrates communication and perception, increasing spectrum utilization on the basis of shared spectrum, and simultaneously improving overall system performance through intelligent resource optimization using artificial intelligence. Current research mostly focuses on the base station scheduling mode, where the base station communicates and perceives with the vehicle. This approach cannot meet the ultra-low latency communication requirements of 6G, and because the vehicle communicates with the base station but needs to perceive other vehicles and facilities, it cannot achieve a high degree of integration between perception and communication. The terminal scheduling resource mode in V2X, where the vehicle directly and simultaneously communicates and perceives the target while optimizing its own resources, can reduce V2X latency. Furthermore, using communication signals as radar detection waveforms allows for a high degree of integration between communication and perception. However, current resource optimization methods are static and involve alternating solutions, making it difficult for traditional methods to jointly and synchronously optimize the communication and sensing system metrics for integrated communication and sensing systems.
[0003] Therefore, it is necessary to study an optimization method, device, system, and readable storage medium based on a communication-sensing integrated system for intelligent joint resource optimization in dynamic environments to address the shortcomings of existing technologies and solve or mitigate one or more of the aforementioned problems. Summary of the Invention
[0004] In view of this, the present invention provides an optimization method, apparatus, system and readable storage medium for intelligent joint resource optimization based on an integrated communication and sensing system in dynamic environments. The proposed intelligent joint synchronous resource optimization in vehicle networks has guiding significance for the development of future optimization systems based on integrated communication and sensing systems in dynamic environments.
[0005] On one hand, the present invention provides an optimization method based on an integrated communication and sensing system. The vehicle-to-everything (V2X) system includes an integrated communication and sensing system and other subsystems. The integrated communication and sensing system is a system formed by the fusion of the communication system and the sensing system in the V2X system. The optimization method based on the integrated communication and sensing system includes the following steps:
[0006] S1: Initialize the integrated communication and sensing system based on terahertz communication;
[0007] S2: After dimensionality reduction of other subsystems in the vehicle networking system, it is fused with the communication and perception integrated system initialized in S1 to obtain the original current state of the agent in deep reinforcement learning.
[0008] S3: Input the original current state into the TD3 neural network for training, and output the optimized current state;
[0009] S4: Optimize the current state input into the TD3 neural network and output the optimal policy action of the agent in a time-varying dynamic environment, which combines hybrid precoding and power allocation.
[0010] In addition to the aspects and any possible implementations described above, an implementation is further provided, wherein S1 specifically involves: initializing the channel state at different time points in the integrated communication and sensing system, the analog precoding matrix of all vehicles and the power allocation strategy actions, the number of training rounds of deep reinforcement learning, and the structure and parameters of the TD3 neural network and meta-learning network;
[0011] Channel state is a fundamental parameter of the integrated communication and sensing system in the Internet of Vehicles (IoV). Simulated precoding matrix and power allocation are optimization strategies and actions. Deep reinforcement learning, TD3 neural network and meta-learning are methods for optimizing strategies and actions.
[0012] In addition to the aspects and any possible implementations described above, an implementation is further provided, wherein S2 specifically includes:
[0013] S21: Divide the channel state at different times in the vehicle network system into multiple tasks to form a task set, and start the meta-learning after initialization. The meta-learning generalizes the learning ability by learning the task set.
[0014] S22: Determine whether meta-learning in S31 has ended. If meta-learning has ended, proceed to S4. If meta-learning has not ended, randomly sample multiple tasks from the task set and proceed to S23. The meta-learning generalizes the learning ability through the learning of the task set in S21.
[0015] S23: Reduce the channel state matrix of the system under a single task in multiple tasks to one dimension and combine it with the initialized one-dimensional policy action, and then output the original current state. The initialized one-dimensional policy action is a random hybrid precoding matrix and power allocation factor.
[0016] In addition to the aspects and any possible implementations described above, a further implementation is provided, wherein S3 specifically includes:
[0017] S31: Determine if the preset number of training rounds for meta-learning has been reached. If the preset number of training rounds has been reached, optimize and update the parameters of the meta-learning network, and output the original current state as the optimized current state; if the preset number of training rounds has not been reached, proceed to step S32:
[0018] S32: Input the original current state into the TD3 neural network, and output the analog precoding and power allocation strategy actions of all vehicles in the system through the TD3 neural network;
[0019] S33: The analog precoding matrix obtained by the TD3 neural network is used to calculate the digital precoding matrix through the zero-forcing method and together they form a hybrid precoding scheme. The TD3 neural network will also output the power allocation scheme and, combined with the current channel state, perform reward calculation through the established mathematical model. The reward calculation objects are the average transmission rate of the communication system and the beam error of the sensing system. The reward calculation is a summation calculation.
[0020] S34: Store the four elements—hybrid precoding scheme, power allocation scheme, reward, and current channel state—into the experience pool;
[0021] S35: Determine whether the number of experiences in the experience pool has reached the preset maximum capacity of the experience pool. If it has not reached the maximum capacity, proceed to S36. If the experience pool is full of experiences, randomly extract experiences from the experience pool for learning and training. During the learning and training process, continuously update the parameters of the TD3 neural network and use these parameters to optimize the parameters of the meta-learning network.
[0022] S36: Recombine the channel state and the policy action output by the neural network into the original current state, and repeat S31.
[0023] In addition to the aspects and any possible implementations described above, a further implementation is provided in which the number of training rounds preset in S31 is set based on the convergence results of S31-S36, and the upper limit of the number of training rounds is set as the number of training rounds when the reward value no longer increases.
[0024] In addition to the aspects and any possible implementations described above, a further implementation is provided, wherein the method for determining the preset number of training rounds is as follows: first, a preliminary judgment is made on the initial range of the number of training rounds; if the preliminary judgment result is inconsistent, the initial range is increased by an equal amount; then, a preliminary judgment is made again until the reward value no longer increases. The initial range is 400-600 rounds, and the equal increase range is 400-600 rounds.
[0025] In addition to the aspects and any possible implementations described above, a further implementation is provided, wherein the method for determining whether meta-learning needs to end in S22 is as follows:
[0026] Meta-learning updates parameters using gradient descent, and the update method is as follows:
[0027] ω′=ω-αG;
[0028] Where ω′ represents the parameters updated by meta-learning, ω represents the parameters before meta-learning, G is the calculated task loss gradient, and α is the learning rate for updating the meta-learning parameters. Meta-learning ends after all tasks have been iterated with meta-learning.
[0029] In accordance with the aspects described above and any possible implementation, an optimization device based on a communication-sensing integrated system is further provided. This optimization device is an apparatus for resource optimization within the communication-sensing integrated system, and includes:
[0030] An initialization module is used to initialize the integrated communication and sensing system based on terahertz communication.
[0031] The dimension reduction and fusion processing module is used to reduce the dimension of the vehicle network system and then fuse it with the optimized system based on the initial communication and perception integrated system to obtain the original current state.
[0032] The training and optimization module is used to input the current state into the TD3 neural network for training and output an optimized current state.
[0033] The policy output module is used to input the optimized current state into the TD3 neural network and output the optimal policy action in a time-varying dynamic environment.
[0034] In accordance with the aspects described above and any possible implementation, an optimization system based on a communication-sensing integrated system is further provided, the system comprising: one or more processors; a memory; and one or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the processor to implement any of the optimization methods based on the communication-sensing integrated system described above.
[0035] In accordance with the aspects described above and any possible implementation thereof, a computer-readable storage medium is further provided having a computer program stored thereon, the computer program being loaded by a processor to perform the steps in the optimization method based on any of the described communication-sensing integrated systems.
[0036] Compared with the prior art, the present invention can achieve the following technical effects:
[0037] 1) This invention uses terahertz and hybrid precoding techniques to propose a suitable modeling method. Taking the Internet of Vehicles as an example, it is used to analyze the transmission rate in the communication system and the beam error in the sensing system under dynamic environment, and to characterize the optimized performance index of the system based on the integrated communication and sensing system.
[0038] 2): This invention analyzes the performance indicators of the entire regional system. The vehicle will communicate and sense other vehicles and facilities, and the combined communication and sensing system indicators will be used as the final optimization indicators of the vehicle network system based on the integrated communication and sensing system.
[0039] 3): This invention proposes an intelligent joint synchronous resource optimization method for dynamic environments, which combines meta-learning and Twin Delayed Deep Deterministic policy gradient algorithm (TD3) deep reinforcement learning. Meta-learning will output the parameters of the TD3 neural network for the policy actions required in the dynamic environment, and TD3 deep reinforcement learning will complete the policy optimization of hybrid precoding and power allocation under a single task.
[0040] 4) This invention combines the vehicle simulation precoding matrix and power allocation strategy actions in the system into a whole action under a single task, and optimizes the resources through joint synchronization.
[0041] 5): This invention analyzes vehicle-to-vehicle communication and sensing, as well as vehicle-to-road communication and sensing, in the vehicle network, and analyzes the signal interference and power constraint problems in the vehicle network system.
[0042] 6) In this invention, the effects of intelligent joint synchronous resource optimization in dynamic environments are achieved through the interaction of TD3 neural network parameters and meta-learning network parameters.
[0043] Of course, any product implementing this invention does not necessarily need to achieve all of the technical effects described above at the same time. Figure Description
[0044] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0045] Figure 1 is a flowchart of an intelligent joint synchronization method for an optimization system based on a communication and sensing integrated system in a dynamic environment, according to an embodiment of the present invention.
[0046] Figure 2 is a cross-sectional view of a method provided in an embodiment of the present invention. Detailed Implementation Methods
[0047] To better understand the technical solution of the present invention, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0048] It should be understood that the described embodiments are merely some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0049] The terminology used in the embodiments of this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. The singular forms “a,” “the,” and “the” as used in the embodiments of this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0050] This invention provides an optimization method based on an integrated communication and sensing system for intelligent joint resource optimization in dynamic environments. The optimization method provided in this invention involves hybrid precoding and power allocation optimization of the integrated communication and sensing system. It is an optimization technology for the application of integrated communication and sensing systems in the Internet of Vehicles (IoV). The object of this method is the integrated communication and sensing system in the IoV, which includes the integrated communication and sensing system and other subsystems. The integrated communication and sensing system is the system after the fusion of the communication system and the sensing system in the IoV. The optimization method based on the integrated communication and sensing system includes the following steps:
[0051] S1: Initialize the integrated communication and sensing system based on terahertz communication;
[0052] S2: After reducing the dimensionality of other subsystems in the vehicle networking system, it is fused with the communication and perception integrated system initialized in S1 to obtain the original current state of the agent in deep reinforcement learning; the agent includes the vehicle and the intelligent working units on the vehicle.
[0053] S3: Input the original current state into the TD3 neural network for training, and output the optimized current state;
[0054] S4: Optimize the current state input into the TD3 neural network and output the optimal policy action of the agent in a time-varying dynamic environment, which combines hybrid precoding and power allocation.
[0055] Specifically, S1 involves: initializing the channel state at different time points in the integrated communication and sensing system, the simulated precoding matrix and power allocation strategy actions of all vehicles, the number of training rounds for deep reinforcement learning, and the structure and parameters of the TD3 neural network and meta-learning network.
[0056] Channel state is a fundamental parameter of the integrated communication and sensing system in the Internet of Vehicles (IoV). Simulated precoding matrix and power allocation are optimization strategies and actions. Deep reinforcement learning, TD3 neural network and meta-learning are methods for optimizing strategies and actions.
[0057] S2 specifically includes:
[0058] S21: Divide the channel state at different times in the vehicle network system into multiple tasks to form a task set, and start the meta-learning after initialization. The meta-learning generalizes the learning ability by learning the task set.
[0059] S22: Determine whether meta-learning in S31 has ended. If meta-learning has ended, proceed to S4. If meta-learning has not ended, randomly sample multiple tasks from the task set and proceed to S23. The meta-learning generalizes the learning ability through the learning of the task set in S21.
[0060] S23: Reduce the channel state matrix of the system under a single task in multiple tasks to one dimension and combine it with the initialized one-dimensional policy action, and then output the original current state. The initialized one-dimensional policy action is a random hybrid precoding matrix and power allocation factor.
[0061] S3 specifically includes:
[0062] S31: Determine if the preset number of training rounds for meta-learning has been reached. If the preset number of training rounds has been reached, optimize and update the parameters of the meta-learning network, and output the original current state as the optimized current state; if the preset number of training rounds has not been reached, proceed to step S32:
[0063] S32: Input the original current state into the TD3 neural network, and output the analog precoding and power allocation strategy actions of all vehicles in the system through the TD3 neural network;
[0064] S33: The analog precoding matrix obtained by the TD3 neural network is used to calculate the digital precoding matrix through the zero-forcing method and together they form a hybrid precoding scheme. The TD3 neural network will also output the power allocation scheme and, combined with the current channel state, perform reward calculation through the established mathematical model. The reward calculation objects are the average transmission rate of the communication system and the beam error of the sensing system. The reward calculation is a summation calculation.
[0065] S34: Store the four elements—hybrid precoding scheme, power allocation scheme, reward, and current channel state—into the experience pool;
[0066] S35: Determine whether the number of experiences in the experience pool has reached the preset maximum capacity of the experience pool. If it has not reached the maximum capacity, proceed to S36. If the experience pool is full of experiences, randomly extract experiences from the experience pool for learning and training. During the learning and training process, continuously update the parameters of the TD3 neural network and use these parameters to optimize the parameters of the meta-learning network.
[0067] S36: Recombine the channel state and the policy action output by the neural network into the original current state, and repeat S31.
[0068] The number of training rounds preset in S31 is set based on the convergence results of S31-S36, and the upper limit of the number of training rounds is set as the number of training sessions when the reward value no longer increases.
[0069] The method for determining the preset number of training rounds is as follows: First, a preliminary judgment is made on the initial range of the number of training rounds. If the preliminary judgment result is inconsistent, the initial range is increased by the same amount. Then, the preliminary judgment is made again until the reward value no longer increases. The initial range is 400-600 rounds, and the range of the equal increase is 400-600 rounds.
[0070] Preferably, it is necessary to first explore a wide range of training rounds, starting with 500 rounds and then adding 500 rounds each time, until the performance of the optimized system based on the integrated communication and sensing system no longer shows a significant increase. This number will then be used as the number of training rounds for the next period.
[0071] The specific method for determining whether meta-learning needs to end is as follows:
[0072] Meta-learning uses gradient descent to update parameters, with the update method being ω′=ω-αG, where ω′ represents the updated parameters, ω represents the parameters before the update, G is the calculated task loss gradient, and α is the learning rate for updating the meta-learning parameters. If the number of iterations in meta-learning is too large, it can lead to abnormal parameters. Therefore, meta-learning needs to set a small number of tasks, which is consistent with the characteristic of meta-learning to improve the learning generalization ability with a small number of learnings. When all tasks have been iterated with meta-learning, meta-learning can end.
[0073] The method for determining whether meta-learning needs to end in S22 is as follows:
[0074] Meta-learning updates parameters using gradient descent, and the update method is as follows:
[0075] ω′=ω-αG;
[0076] Where ω′ represents the parameters updated by meta-learning, ω represents the parameters before meta-learning, G is the calculated task loss gradient, and α is the learning rate for updating the meta-learning parameters. Meta-learning ends after all tasks have been iterated with meta-learning.
[0077] The present invention also provides an optimization device based on a communication-sensing integrated system. The optimization device is a means for resource optimization within the communication-sensing integrated system, and includes:
[0078] An initialization module is used to initialize the integrated communication and sensing system based on terahertz communication.
[0079] The dimension reduction and fusion processing module is used to reduce the dimension of the vehicle network system and then fuse it with the optimized system based on the initial communication and perception integrated system to obtain the original current state.
[0080] The training and optimization module is used to input the current state into the TD3 neural network for training and output an optimized current state.
[0081] The policy output module is used to input the optimized current state into the TD3 neural network and output the optimal policy action in a time-varying dynamic environment.
[0082] The present invention also provides an optimization system based on a communication-sensing integrated system, the system comprising: one or more processors; a memory; and one or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the processor to implement the optimization method based on any one of the communication-sensing integrated systems.
[0083] The present invention also provides a computer-readable storage medium having a computer program stored thereon, the computer program being loaded by a processor to execute the steps in the optimization method based on any of the described communication-sensing integrated systems.
[0084] Example 1:
[0085] This invention proposes an intelligent joint synchronous resource optimization method for dynamic environments, used to achieve intelligent resource management and control while improving the performance of an optimization system based on an integrated communication and sensing system. The highly integrated communication and sensing systems, along with the dynamic environment surrounding the vehicle, present both temporal and strategic challenges to intelligent resource management and optimization. This invention constructs a terahertz-based channel model for the vehicle. The time-varying channel states at different moments in a dynamic environment are used as different tasks for meta-learning. In each task, the current channel state, analog precoding, and power allocation strategy actions are combined to form the current state, which is then input into a TD3 neural network for outputting strategy actions. The output analog precoding, power allocation strategy actions, and the vehicle's channel state are used to calculate a reward value. This reward value represents the optimization based on the integrated communication and sensing system. The reward value is jointly calculated using the transmission rate in the communication system (analyzing signal interference between the vehicle and roadside equipment) and the beam error in the sensing system (meeting power constraints) as the reward value for deep reinforcement learning and synchronous optimization. By combining meta-learning and TD3 deep reinforcement learning, while solving resource allocation schemes under individual identical channel conditions, the neural network parameters for the allocation of policy actions required by the channel at different time points in a dynamic environment can be predicted and optimized using the meta-learning network. This enables intelligent joint synchronous resource optimization under dynamic conditions, thereby improving the performance of the optimization system based on the integrated communication and perception system in the Internet of Vehicles.
[0086] As shown in Figure 1, the specific steps of this invention are as follows:
[0087] Step 1: Initialize the time-varying channel state at different moments based on the terahertz communication system, the analog precoding matrix of all vehicles and the power allocation policy actions, the number of training rounds of deep reinforcement learning, and the structure and parameters of the TD3 neural network and meta-learning network;
[0088] Step 2: Divide the channel state at different times in the vehicle-to-everything (V2X) system into multiple tasks to form a task set;
[0089] Step 3: Determine whether meta-learning needs to end. If meta-learning ends, proceed to step 11; if meta-learning does not end, randomly sample multiple tasks from the task set and continue to step 4.
[0090] Step 4: Reduce the channel state matrix of the system under a single task in multiple tasks to one dimension and combine it with the initialized one-dimensional policy action to form the current state, and then proceed to step 5;
[0091] Step 5: Determine if the required number of training epochs has been reached. If the required number of training epochs has been reached, it indicates that the TD3 neural network parameters for a single task have been trained, and the parameters of the meta-learning network will be optimized and updated accordingly. If the required number of training epochs has not been reached, proceed to Step 6.
[0092] Step 6: Input the current state into the TD3 neural network, and output the analog precoding and power allocation strategy actions of all vehicles in the system through the TD3 neural network;
[0093] Step 7: The analog precoding matrix obtained by the TD3 neural network is used to calculate the digital precoding matrix through the zero-forcing method and together they form a hybrid precoding scheme. The TD3 neural network will also output the power allocation scheme and, combined with the current channel state, perform reward calculation through the established mathematical model.
[0094] Step 8: Store the four elements mentioned above—the current state, the policy action output by the neural network, the calculated reward value, and the next state—into the experience pool;
[0095] Step 9: Determine if the amount of experience in the experience pool has reached the set maximum capacity. If the maximum capacity has not been reached, proceed directly to Step 10. If the experience pool is full, the system will randomly extract experience from the experience pool for learning and training. During the learning and training process, the parameters of the TD3 neural network will be continuously updated, and these parameters will be used to optimize the parameters of the meta-learning network.
[0096] Step 10: Recombine the channel state and the policy action output by the neural network into the current state, and then jump to step 5;
[0097] Step 11: Use a meta-learning network to predict the policy network parameters under this environment model, and intelligently and quickly obtain the optimal policy action in a time-varying dynamic environment.
[0098] In step 1, the channel state modeling of the terahertz-based communication system is performed, and the channel response formula from vehicle k to the receiving terminal (i.e., other vehicles and roadside equipment around vehicle k) u is given by h. k,u (f,d)
[0099]
[0100] The set of vehicles is K = {1, 2, ..., k, ..., K}, and there exists L around vehicle k. kThe number of roadside devices is L, and the set of roadside devices is denoted as L. k ={1,2,…l k ,…,L k}, u∈U k U k =(K\{k})∩L k N T This represents the number of transmitting antennas per vehicle, PL(f,d) represents the transmission path attenuation, and Ω is the antenna's transmit gain. Let be the antenna array response vector from vehicle k to receiving device u.
[0101] The formula for calculating the path loss PL(f,d) is as follows:
[0102]
[0103] Where L sl (f,d) represents free-space transmission attenuation, L mal (f,d) represents molecular absorption attenuation, where f represents the frequency in THz, d represents the transmission distance, k(f) represents the molecular absorption coefficient, and e is a natural constant in mathematics.
[0104] In step 1, the initialization actions of all vehicles, namely the simulated pre-coded phase shift angle θ value and power allocation, are all 0.
[0105] The TD3 neural network is composed of both Actor and Critic networks. The Actor neural network is a three-layer linear network. The input state passes through the first and second layers of the neural network and is activated by the ReLU function. When it passes through the third layer, the Tanh activation function is used, and a linear transformation is applied to convert it to the [0,1] interval, providing a basis for subsequent joint synchronization resource optimization. The Critic neural network is a six-layer linear network. The first three layers are used to output Q1, and the last three layers are used to output Q2. The Q-value is used for action value estimation.
[0106] In step 4, the channel state matrix in the individual task is reduced to one dimension and combined with the one-dimensional analog precoding and power allocation strategy actions to form the current state, which serves as the input to the TD3 neural network. The channel state matrix structure here has two variations depending on the neural network:
[0107] 1) When the neural network is a real neural network, the original complex matrix is split into its real and imaginary parts and combined to form a new real matrix;
[0108] 2) When the neural network is a complex neural network, the complex form is used directly. Then, it is reduced to one dimension by rows. Then, it is combined with the one-dimensional policy action to form the state input into the neural network.
[0109] In step 6, the TD3 neural network outputs the simulated precoding and power allocation strategy actions, and the simulated precoding matrix F k for
[0110]
[0111] in q = 1, 2, ... N RF , express Find the m-th element and use the breaking-zero method to find the digital precoding matrix:
[0112] in, N T N represents the number of transmitting antennas. RF θ represents the number of radio frequency links in hybrid precoding, and θ represents the phase shift parameter of the phase shifter in analog precoding.
[0113]
[0114] Where K is the number of vehicles in the system, L k Let N be the number of roadside facilities around vehicle k. RF This represents the number of radio frequency links in hybrid precoding.
[0115] The reward value is calculated in step 7. The reward is composed of the system's communication metrics and perception metrics, where η c η is the scaling factor for the communication system. s This is the scaling factor for the sensing system. The average transmission rate of the vehicle-to-everything (V2X) system is the sum of the transmission rates between vehicles and between vehicles and other devices in the system. It is calculated as follows:
[0116]
[0117] Where ζ is the proportional factor for adjusting the weights of vehicles and roadside equipment. The beam error is the beam error of the vehicle-to-everything (V2X) system, which is the average beam error of the entire system.
[0118]
[0119] in, This represents the beam error of vehicle k relative to receiving device i.
[0120] Vehicle beam error:
[0121]
[0122] in Indicates the direction of the target The desired beam This represents the beam in the target direction of the pre-coded matrix strategy obtained through a neural network.
[0123] Target response matrix of the transmitted signal:
[0124]
[0125] Where, x k This indicates the signal transmitted by vehicle k. Let V represent the transmission power of vehicle k, ρ represent the power allocation factor, and V represent the transmission power of vehicle k. k V represents the hybrid precoding matrix. k =F k D k .
[0126] And power constraints need to be met.
[0127] Step 7 includes the calculation and analysis of the transmission rate of the communication system and the beam error in the sensing system, which are included in the reward.
[0128] Transmission rate per vehicle:
[0129] R k′ =Blog2(1+SINR) k′ );
[0130] Where B represents the vehicle's transmission bandwidth and SINR represents the signal-to-interference-plus-noise ratio.
[0131] The transmission rate of each roadside device is:
[0132]
[0133] The meanings of the characters are the same as in the formula above.
[0134] The optimization model based on the integrated communication and sensing system involves interference between vehicle-to-vehicle communication and between vehicle and roadside equipment communication signals. The signal-to-interference-plus-noise ratio (SIR) for each vehicle is:
[0135]
[0136] Among them, V k (:,k′) represents matrix V k The k′th column.
[0137] The signal-to-interference-to-noise ratio of the roadside equipment is:
[0138]
[0139] Vehicle beam error in the perception system:
[0140]
[0141] Target response matrix of the transmitted signal:
[0142]
[0143] And power constraints need to be met.
[0144] Specifically, the model in this invention mainly focuses on an optimization model based on an integrated communication and sensing system under the terminal resource scheduling mode in the Internet of Vehicles (IoV), as shown in Figure 2. The system includes K vehicles, with a vehicle set K = {1,2,…,k,…,K}, and L vehicles surrounding the k-th vehicle. k There are L roadside devices. k ={1,2,…l k ,…,L k The signal sent by the kth vehicle is... The first K-1 signals are signals sent to the vehicle, and the following L... k Each signal is sent to its own roadside equipment. The average power of the signal transmitted by each vehicle is... Let u represent the transmit power factor of the signal allocated from the k-th vehicle to the u-th device, where u = {1, 2, ..., K-1 + L}. k}
[0145] The simulated precoding matrix of the k-th user in the system is as follows:
[0146]
[0147] middle express Find the m-th element and use the breaking-zero method to obtain the digital precoding matrix. And the power constraint must be satisfied as follows: ||[D k (:,u)]|| 2 ≤1.
[0148] Terahertz communication The channel model is as follows:
[0149]
[0150]
[0151] The original channel matrix is a complex matrix. Its real and imaginary parts are split and combined to form a new real channel matrix, which is reduced to one dimension and combined with one-dimensional actions to form the current state input into the Actor network.
[0152] The calculation of the reward value in step 7. The reward is composed of the system's communication metrics and perception metrics, where η c η is the scaling factor for the communication system. s This is the scaling factor for the sensing system. The average transmission rate of the vehicle-to-everything (V2X) system is the sum of the transmission rates between vehicles and between vehicles and other devices within the system.
[0153] The calculation method is as follows:
[0154]
[0155] Where μ is the scaling factor for adjusting the weights of vehicles and roadside equipment. The beam error is the beam error of the vehicle-to-everything (V2X) system, while the beam error is the average beam error of the entire system.
[0156] The reward includes the calculation and analysis of beam error in the communication system's transmission rate sensing system, with a transmission rate R for each vehicle. k′ =Blog2(1+SINR) k′ The transmission rate of each roadside device is:
[0157]
[0158] The optimization model based on the integrated communication and sensing system involves interference between vehicle-to-vehicle communication and between vehicle and roadside equipment communication signals. The signal-to-interference-plus-noise ratio (SIR) for each vehicle is:
[0159]
[0160] The signal-to-interference-to-noise ratio of the roadside equipment is:
[0161]
[0162] Vehicle beam error in the perception system Target response matrix of transmitted signal And power constraints must be met:
[0163]
[0164] This invention employs terahertz and hybrid precoding techniques to propose a suitable modeling approach. Taking vehicle-to-everything (V2X) systems as an example, it analyzes the transmission rate in the communication system and the beam error in the sensing system under dynamic conditions, and characterizes the system's optimized performance indicators based on an integrated communication and sensing system. This invention analyzes the performance indicators of the entire regional system, where vehicles communicate and sense with other vehicles and facilities, and combines the communication and sensing system indicators as the final optimized indicators for the V2X system based on the integrated communication and sensing system.
[0165] This invention proposes an intelligent joint synchronization resource optimization method for dynamic environments. It combines meta-learning with Twin Delayed Deep Deterministic policy gradient (TD3) deep reinforcement learning. Meta-learning outputs the parameters of the TD3 neural network for the policy actions required in the dynamic environment, while TD3 deep reinforcement learning completes the policy optimization of hybrid precoding and power allocation under a single task. Under a single task, the vehicle simulation precoding matrix and power allocation policy actions in the system are combined into a unified action for joint synchronization resource optimization. The invention analyzes vehicle-to-vehicle communication and perception, as well as vehicle-to-roadside equipment communication and perception in the vehicular network, and also analyzes signal interference and power constraints in the vehicular network system. The interaction of TD3 neural network parameters and meta-learning network parameters achieves intelligent joint synchronization resource optimization under dynamic environments.
[0166] The intelligent joint synchronous resource optimization proposed in this invention has guiding significance for the development of optimization systems based on integrated communication and sensing systems in future dynamic environments.
[0167] The foregoing has provided a detailed description of an optimization method, apparatus, system, and readable storage medium for intelligent joint resource optimization in dynamic environments based on an integrated communication and sensing system, as provided in the embodiments of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and core ideas of this application; furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
[0168] Certain terms are used in the specification and claims to refer to specific components. Those skilled in the art will understand that hardware manufacturers may use different names to refer to the same component. This specification and claims do not distinguish components based on differences in name, but rather on differences in function. The terms "comprising" and "including" used throughout the specification and claims are open-ended and should be interpreted as "comprising / including but not limited to". "Approximately" means that within an acceptable margin of error, those skilled in the art can solve the technical problem and substantially achieve the technical effect within a certain margin of error. The following descriptions in the specification are preferred embodiments for carrying out this application; however, these descriptions are for the purpose of illustrating the general principles of this application and are not intended to limit the scope of this application. The scope of protection of this application shall be determined by the appended claims.
[0169] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a product or system comprising a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a product or system. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the product or system that includes said element.
[0170] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0171] The foregoing description illustrates and describes several preferred embodiments of this application. However, as previously stated, it should be understood that this application is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and environments, and can be altered within the scope of the application concept described herein through the foregoing teachings or techniques or knowledge in related fields. Any modifications and variations made by those skilled in the art that do not depart from the spirit and scope of this application should be within the protection scope of the appended claims.
Claims
1. An optimization method based on an integrated communication and sensing system, wherein the vehicle-to-everything (V2X) system includes an integrated communication and sensing system and other subsystems, wherein the integrated communication and sensing system is a system formed by merging the communication system and the sensing system within the V2X system, characterized in that, The optimization method based on the integrated communication and sensing system Includes the following steps: S1: Initialize the integrated communication and sensing system based on terahertz communication; S2: After dimensionality reduction of other subsystems in the vehicle networking system, it is fused with the communication and perception integrated system initialized in S1 to obtain the original current state of the agent in deep reinforcement learning. S3: Input the original current state into the TD3 neural network for training, and output the optimized current state; S4: Optimize the current state by inputting it into the TD3 neural network, and output the optimal policy action of the agent in a time-varying dynamic environment using hybrid precoding and power allocation; S1 specifically involves: initializing the channel state at different time points in the integrated communication and sensing system, the simulated precoding matrix and power allocation policy actions of all vehicles, the number of training rounds for deep reinforcement learning, and the structure and parameters of the TD3 neural network and meta-learning network; where the channel state is a basic parameter of the integrated communication and sensing system under the vehicle network, the simulated precoding matrix and power allocation are the optimized policy actions, and deep reinforcement learning, the TD3 neural network, and meta-learning are the methods for optimizing the policy actions; S2 specifically includes: S21: Divide the channel state at different time points in the vehicle network system into... The system is divided into multiple tasks forming a task set. Meta-learning is initiated after initialization, and this meta-learning generalizes learning capabilities through learning from the task set. S22: Determine if meta-learning in S21 has ended. If it has, proceed to S4; if not, randomly sample multiple tasks from the task set and proceed to S23. Meta-learning generalizes learning capabilities through learning from the task set in S21. S23: Reduce the system's channel state matrix for a single task in the multiple tasks to one dimension and combine it with the initialized one-dimensional policy action. Then, output the original current state. The initialized one-dimensional policy action is a random mixture of a precoding matrix and a power allocation factor. Specifically, S3 includes: S31: Determine if the meta-learning training... If the preset number of training rounds has been reached, the parameters of the meta-learning network are optimized and updated, and the original current state is output as the optimized current state. If the preset number of training rounds has not been reached, proceed to step S32: S32: Input the original current state into the TD3 neural network, and output the analog precoding and power allocation strategy actions of all vehicles in the system through the TD3 neural network; S33: Calculate the digital precoding matrix from the analog precoding matrix obtained by the TD3 neural network using the zero-forcing method, and combine them to form a hybrid precoding scheme. The TD3 neural network will also output the power allocation scheme, and combined with the current channel state, perform reward calculation through the established mathematical model to obtain the reward value. S34: The calculation objects are the average transmission rate of the communication system and the beam error of the sensing system. The reward calculation is a summation calculation. S35: Store the four elements of the hybrid precoding scheme, power allocation scheme, reward value and current channel state into the experience pool. S36: Determine whether the number of experiences in the experience pool has reached the preset maximum capacity of the experience pool. If it has not reached the maximum capacity, proceed to S37. If the experience pool is full, randomly select experiences from the experience pool for learning and training. During the learning and training process, continuously update the parameters of the TD3 neural network and use these parameters to optimize the parameters of the meta-learning network. S37: Recombine the channel state and the policy action output by the neural network into the original current state and repeat S38.
2. The optimization method based on the integrated communication and sensing system according to claim 1, characterized in that, The number of training rounds preset in S31 is set based on the convergence results of S31-S36, and the upper limit of the number of training rounds is set as the number of training sessions when the reward value no longer increases.
3. The optimization method based on the integrated communication and sensing system according to claim 2, characterized in that, The method for determining the preset number of training rounds is as follows: First, a preliminary judgment is made on the initial range of the number of training rounds. If the preliminary judgment result is inconsistent, the initial range is increased by the same amount. Then, the preliminary judgment is made again until the reward value no longer increases. The initial range is 400-600 rounds, and the range of the equal increase is 400-600 rounds.
4. The optimization method based on the integrated communication and sensing system according to claim 1, characterized in that, The method for determining whether meta-learning needs to end in S22 is as follows: Meta-learning updates parameters using gradient descent, and the update method is as follows: ;in This represents the parameters updated after meta-learning. The parameters before meta-learning update. To calculate the task loss gradient, The learning rate is used to update the meta-learning parameters. Meta-learning ends after all tasks have been iterated with the meta-learning.
5. An optimization device based on a communication-sensing integrated system, implemented based on the optimization method for a communication-sensing integrated system according to any one of claims 1 to 4, wherein the optimization device is a device for resource optimization in a communication-sensing integrated system, characterized in that, The optimization device includes: an initialization module for initializing a terahertz-based integrated communication and sensing system; a dimensionality reduction and fusion processing module for reducing the dimensionality of the vehicle network system and then fusing it with the initialized optimized system based on the integrated communication and sensing system to obtain the original current state; a training and optimization module for inputting the current state into a TD3 neural network for training and outputting an optimized current state; and a policy output module for inputting the optimized current state into the TD3 neural network and outputting the optimal policy action under time-varying dynamic conditions.
6. An optimized system based on an integrated communication and sensing system, characterized in that, The system includes: one or more processors; a memory; and one or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor to implement the optimization method based on the communication-sensing integrated system as described in any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that, It stores a computer program, which is loaded by a processor to execute the steps in the optimization method based on the integrated communication and sensing system as described in any one of claims 1 to 4.