World model-based information value perception vehicle networking spectrum sharing method and system
Patent Information
- Application Number
- CN202610705519.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-21
- Publication Date
- 2026-09-29
AI Technical Summary
[0011]本发明目的在于解决现有车联网频谱共享技术在动态城市环境下存在的样本利用率低下、决策视野短视以及难以平衡信息新鲜度与干扰抑制等技术问题
[0031]1、本发明通过构建世界模型内化隐式的缓存队列演化动力学,使调度器能够通过想象在潜在空间中预判未来的缓存溢出和数据包过期风险,有效降低信息价值损失,确保在动态信道条件下的高信息新鲜度。
Smart Images

Figure CN122846136A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method and system for information value perception in vehicle-to-everything (V2X) networks based on a world model, and belongs to the field of wireless communication. Background Technology
[0002] Against the backdrop of the deep integration of 6G mobile communication technology and vehicle-to-everything (V2X) communication, building intelligent and efficient communication networks has become a core foundation for supporting smart transportation and autonomous driving. Due to the highly dynamic nature of urban road environments, the high-speed movement of vehicles causes wireless channels to exhibit rapidly changing characteristics, exacerbating the contradiction between the diverse communication needs of massive amounts of onboard devices and limited spectrum resources. Although spectrum sharing technology effectively improves spectrum utilization by allowing vehicle links to reuse frequency resources, in real-world scenarios, vehicle mobility, resource scheduling coupling, and complex signal obstruction and interference control remain major obstacles to maintaining high levels of information utility.
[0003] For vehicle-to-everything (V2X) safety-sensitive services such as collision warning and collaborative driving, the freshness and timeliness of information directly impact the safety of decision-making. Traditional quality of service (QoS) metrics (such as latency and throughput) are insufficient to accurately represent the stringent real-time value requirements of these services. If the data received by the receiver is outdated, it not only leads to inefficient use of bandwidth resources but may also induce incorrect driving decisions, increasing the risk of traffic accidents. Therefore, introducing "information value" as the core metric for measuring the effectiveness and timeliness of information, and using this as a guide for resource allocation, is crucial for ensuring traffic safety and system efficiency.
[0004] However, achieving efficient spectrum resource scheduling in dense urban environments faces multiple challenges. When a large number of vehicles, roadside units, and public network users access the network concurrently, dedicated spectrum is prone to instantaneous saturation. In spectrum sharing mode, vehicles, as secondary users, need to reuse resources while ensuring the communication quality of primary users. This leads to an inherent conflict between high-frequency state updates and stringent interference control. On the one hand, increasing information value requires vehicle nodes to increase transmission power and frequency; on the other hand, interference to primary users and other vehicles must be strictly limited. Furthermore, the vehicle-to-everything (V2X) environment exhibits high non-stationarity and partial observability. Changes in channel state, signal fluctuations caused by non-line-of-sight propagation, and the evolution of buffer queues affected by scheduling strategies collectively create a complex environment that makes spectrum sharing a highly challenging fundamental scientific problem.
[0005] To address the aforementioned challenges, existing research approaches mainly fall into two categories: model-based traditional optimization methods and model-free deep reinforcement learning methods. The first is model-based traditional optimization methods. These methods typically utilize theories such as convex optimization, Lyapunov optimization, or game theory to construct resource allocation models. For example, patent CN114884595B discloses a cognitive UAV spectrum sensing scheme that optimizes flight trajectories through mathematical modeling to obtain idle bandwidth; related literature, such as "Energy-efficient resource allocation for V2X communications," explores the application of traditional optimization algorithms in energy efficiency allocation. While these methods possess theoretical completeness, they heavily rely on accurate global channel state information and environmental dynamics models. In complex urban electromagnetic environments, obtaining real-time and accurate channel information often involves significant signaling overhead, and the computational complexity of the algorithms increases exponentially with network expansion, making it difficult to meet the real-time constraints of vehicle-to-everything (V2X) services.
[0006] Second, model-free deep reinforcement learning methods. To reduce reliance on precise physical models, algorithms such as dual-delay deep deterministic policy gradient (TD3) and proximal policy optimization (PPO) have been introduced into the field of vehicle-to-everything (V2X) communication. For example, patent CN111818625B discloses a V2X resource allocation method based on deep reinforcement learning, which learns the optimal allocation strategy through iterative interaction between the agent and the environment; related literature such as "Deep deterministic policy gradient to minimize the age of information in cellular V2X communications" utilizes this type of algorithm to optimize information freshness. These algorithms acquire decision-making strategies through a "trial and error" mechanism, exhibiting strong environmental adaptability and the ability to handle non-convex and nonlinear scheduling problems.
[0007] However, standard model-free reinforcement learning has significant drawbacks when dealing with highly dynamic and resource-constrained tasks: Firstly, it suffers from low sample efficiency, often requiring millions of physical interactions for convergence. Due to the lack of internal representations in the connected vehicle environment, agents cannot simulate experience, leading to numerous suboptimal decisions during the training phase before convergence, severely diminishing information value. Secondly, it lacks foresight in decision-making. These algorithms tend to make greedy decisions based on the current state, lacking the ability to predict the long-term evolution of the environment. When handling tasks with buffer queues and deadlines, agents cannot effectively foresee the impact of current decisions on future queue stability, resulting in performance oscillations during traffic load fluctuations and difficulty in maintaining long-term stable information value.
[0008] Therefore, reinforcement learning pathways based on world models have also been proposed. To balance sample efficiency and long-term planning, model-based reinforcement learning techniques, represented by world models, are gradually emerging. The paper "Dream to control: Learning behaviors by latent imagination" proposes the core idea of simulating environmental dynamics and guiding decision-making by constructing internal representations. Agents can perform "imaginary" deductions in the latent space, reducing their reliance on interactions with the real environment.
[0009] Despite the excellent performance of world models in visual tasks, their application in the field of vehicle-to-everything (V2X) spectrum sharing still faces a technological vacuum. Firstly, there is a data modality mismatch: existing world models are mostly designed for structured visual data, while wireless communication data is characterized by high noise, sparsity, and high randomness, making it difficult to capture the intrinsic features of fast channel fading and buffer queues. Secondly, existing research often focuses on maximizing explicit rewards, failing to deeply integrate the composite indicator of information value. This results in models struggling to internalize the balance between information freshness and interference suppression during the visualization phase.
[0010] In summary, existing technologies struggle to balance information efficiency, learning efficiency, and long-term stability when dealing with highly dynamic environments and information value constraints. Therefore, developing a novel intelligent spectrum sharing method capable of deeply understanding and predicting the implicit dynamic characteristics of the vehicle-to-everything (V2X) environment is of significant technological importance for realizing smart communication in the 6G era. Summary of the Invention
[0011] The purpose of this invention is to address the technical problems of existing vehicle-to-everything (V2X) spectrum sharing technologies in dynamic urban environments, such as low sample utilization, short-sighted decision-making perspective, and difficulty in balancing information freshness and interference suppression. To this end, this invention proposes a world-model-based information value perception V2X spectrum sharing method and system. By introducing a world model capable of learning implicit environmental dynamics, predictive planning is achieved within the potential space. Under the premise of strictly satisfying spectrum interference constraints, the system's accumulated information value is maximized, ensuring the robustness and efficiency of spectrum resource scheduling.
[0012] The technical solution adopted by this invention to solve its technical problem is: a method for information value perception in vehicle-to-everything (V2X) spectrum sharing based on a world model, which includes the following steps:
[0013] Step S1: Construct a vehicle-to-everything (V2X) communication environment model and define reinforcement learning elements. Initialize the system state space, action space, and reward function. The state space is composed of an explicit spectrum observation dimension and an implicit cache dynamic dimension; the reward function uses information value as the core indicator, aiming to quantify the timeliness and effectiveness of data. The implicit cache dynamic state is driven by historical arrival data and scheduling decisions, exhibiting inherent uncertainty. The reward function is configured to be positively correlated with the overall network information value and incorporates penalties for spectrum interference exceeding limits and cache queue instability.
[0014] Step S2: Establish a spectrum-based cognitive world model for dynamic environments. This world model aims to map high-dimensional environmental observations to a low-dimensional latent state space and establish a state transition mechanism. The latent state representation adopts a hybrid architecture, including a deterministic component for remembering long-term historical features and a random component for characterizing channel random fluctuations. Specifically, the model includes: a temporal memory unit for recursively evolving deterministic states, maintaining internal memory of channel evolution and cached trajectories; a state inference unit for fusing current real-time observations to generate posterior latent states to calibrate cognitive biases; a state prediction unit, detached from real-world feedback, inferring future states solely based on historical prior information to support virtual planning; and an observation reconstruction and reward prediction unit inversely maps latent features to the observation and reward space, ensuring the physical interpretability of the representation vectors.
[0015] Step S3: Learn the latent dynamic characteristics of the environment based on the interaction trajectory. Utilize historical trajectory data accumulated from the interaction between the scheduling agent and the environment to perform end-to-end optimization of the spectrum-based cognitive world model. By constructing a general objective function, the model can internalize the evolutionary laws of the cache queue and accurately predict rewards. The objective function includes: an environment representation consistency loss to constrain the deviation between reconstructed observations and true values; a task objective prediction loss to align predicted rewards with environmental feedback rewards; and a state distribution constraint loss to minimize the divergence between prior and posterior distributions, improving prediction stability.
[0016] Step S4: Perform predictive planning based on the imagined potential space. Using the optimized world model, perform multi-step deductions within the potential space to generate an imagined trajectory. Evaluate the long-term effectiveness of the candidate action sequence based on this trajectory. Specifically, starting from the current potential state, recursively predict the state sequence for several future steps, calculate the long-term cumulative return implied by future packet expiration and spillover risks; and generate a spectrum access strategy through a policy generation unit, guided by maximizing the aforementioned long-term returns.
[0017] Step S5 involves implementing online scheduling control and closed-loop iterative optimization. The generated strategy is translated into specific transmit power control and bandwidth allocation commands. Simultaneously, environmental feedback observation and reward data are collected in real time to update the experience replay unit, and the world model and strategy parameters are iterated periodically to achieve adaptive evolution in a dynamic environment.
[0018] Secondly, the present invention provides a world-model-based information value perception vehicle-to-everything (V2X) spectrum sharing system for performing the above-described method, the system comprising:
[0019] The environment modeling and state perception module is responsible for defining the vehicle's state space, which includes explicit and implicit dimensions, and calculating the information value reward that integrates disturbance penalties and stability constraints.
[0020] The spectrum-based cognitive world model construction module is responsible for maintaining a hybrid latent space based on a recursive state space architecture, realizing the mapping of high-dimensional observations to deterministic and random components;
[0021] The Environmental Potential Dynamics Learning Module is responsible for training the model by maximizing the lower bound of variational evidence, enabling it to grasp the endogenous laws of cache evolution and channel changes.
[0022] The Information Value Perception Predictive Planning module is responsible for performing multi-step deductions in the potential space, using an actor-critic architecture to assess long-term risks and output the optimal access strategy;
[0023] The online execution and closed-loop optimization module is responsible for issuing scheduling instructions and performing rolling optimization of the system model based on real-time feedback data.
[0024] The implicit cache dynamic state described in this invention is the evolutionary state of the vehicle node cache queue. This state is formed by the arrival of historical data packets and scheduling decisions, and has partial observability for the scheduler.
[0025] The spectral cognitive world model of the present invention is configured to map environmental observations into latent state representations, wherein the latent state representations include at least a first latent component for characterizing historical correlation and a second latent component for characterizing randomness.
[0026] The model update module of the present invention is configured to train the world model based on historical interaction data between the scheduler and the environment, so that it has the ability to predict the evolution of future potential states and corresponding environmental feedback.
[0027] The decision planning module of this invention is configured to perform multi-step predictive planning in the potential state space constructed by the world model, so that potential future caching risks can influence current spectrum scheduling decisions.
[0028] The decision planning module of the present invention includes a value assessment unit and a strategy generation unit. The value assessment unit is used to assess the long-term value of the predicted trajectory, and the strategy generation unit is used to generate a spectrum access control strategy based on the long-term value.
[0029] The spectrum access execution module of the present invention is configured to perform at least one of transmit power control and bandwidth allocation, and the model update module is configured to continuously update empirical data based on real-time environmental feedback and iteratively optimize the world model and spectrum access control strategy.
[0030] Beneficial effects:
[0031] 1. This invention internalizes implicit buffer queue evolution dynamics by constructing a world model, enabling the scheduler to predict future buffer overflow and data packet expiration risks in the potential space through imagination, effectively reducing information value loss and ensuring high information freshness under dynamic channel conditions.
[0032] 2. This invention utilizes a world model to perform multi-step virtual simulations in the latent space, avoiding the need for expensive large-scale data sampling in the real environment. It quickly converges to the optimal policy within a shorter training step, significantly improving sample efficiency compared to model-free reinforcement learning methods.
[0033] 3. This invention deeply aligns spectrum resource allocation with information value, maximizing the total information value of the system while ensuring spectrum sharing interference constraints. This invention increases the average information value by approximately 26.7%, effectively preventing critical information from becoming outdated, and improves spectrum efficiency by an order of magnitude, making it suitable for highly dynamic and high-density vehicle-to-everything (V2X) communication scenarios. Attached Figure Description
[0034] Figure 1 This is a schematic diagram of the vehicle-to-everything (V2X) communication network environment model based on the world model of this invention.
[0035] Figure 2 This is a comparison chart of the total information value curve of the system and the baseline algorithm during the model training process of this invention.
[0036] Figure 3 This is a performance comparison chart of the delay-constrained value guarantee efficiency of the method of this invention and the baseline algorithm under different signal-to-noise ratio conditions.
[0037] Figure 4 This is a performance comparison chart of the throughput-value conversion efficiency of the method of this invention and the baseline algorithm under different signal-to-noise ratio conditions.
[0038] Figure 5 This is a performance comparison chart of the delay-constrained value guarantee efficiency of the method of this invention and the baseline algorithm under different transmission loads.
[0039] Figure 6 This is a performance comparison chart of the throughput-value conversion efficiency of the method of this invention and the baseline algorithm under different transmission loads. Detailed Implementation
[0040] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings.
[0041] It should be understood that these descriptions are merely exemplary and not intended to limit the scope of the invention. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concept of the invention.
[0042] like Figure 1 As shown, this invention provides a method for information value perception in vehicle-to-everything (V2X) spectrum sharing based on a world model; furthermore, this invention also provides a system for executing the above method, wherein each functional module in the system is used to execute the corresponding steps in the following method. The method specifically includes the following steps:
[0043] Step S1: Modeling the vehicle-to-everything (V2X) communication network environment.
[0044] This invention first defines a vehicle-to-everything (V2X) communication scenario, which includes... A mobile vehicle node as a secondary user, One relay drone, one ground primary user, and multiple disaster-stricken areas. Considering distance-dependent path loss, channel fluctuations caused by Rician fading, and mutual interference between vehicles, the first... The achievable communication link speed of a vehicle node in time slot t Represented as:
[0045]
[0046] in, and These represent the bandwidth and transmission power allocated to the vehicle node, respectively. This represents the channel gain of the vehicle-to-relay link. The power spectral density of Gaussian white noise. This indicates that the vehicle originated from another vehicle node. The interference channel gain.
[0047] Step S2, Information Value (VoI) Measurement System.
[0048] To accurately characterize the timeliness and effectiveness of connected vehicle business data, this invention introduces an information value index. Its function expression is defined as follows:
[0049]
[0050] In this definition, Characterizes the inherent initial utility of the data packet; This indicates the latency of a data packet waiting to be transmitted in the buffer queue; This is the preset time-related decay coefficient; This is an indicator function that takes the value 1 only when the transmission is successful.
[0051] Step S3: Construct the system's total reward function.
[0052] This invention rewards the total information value of the system. Set as a multi-objective optimization function to guide resource allocation strategies:
[0053]
[0054] in, This item is used to penalize spectral interference caused to the primary user's receiver; This item is used to characterize the stability risk of the cache queue; and These are the weighted adjustment coefficients for the corresponding items.
[0055] Step S4: Implicit cache dynamics evolution.
[0056] This invention treats implicit cached dynamic states as endogenous variables in the environment model. (Vehicle node) Cache queue length It follows the following dynamic evolution law:
[0057]
[0058] in, Representing time slots The number of newly arrived data packets, Representing time slots The number of data packets successfully sent.
[0059] In the model described in this invention, if the length of the buffer queue exceeds a preset stability threshold, or the waiting time of a single data packet exceeds its relative deadline... If the packet expires, it is determined to be expired and discarded. Given that the evolution of the buffer queue is driven by the coupling of external random data arrivals and internal scheduling decisions, this state exhibits significant endogenous uncertainty for the scheduler.
[0060] Step S5: Construct and learn a spectral cognitive world model of the potential dynamics of the environment.
[0061] In step S5, in order to achieve continuous adaptation to the dynamic environment, the present invention further performs model architecture maintenance and parameter iterative learning through the following sub-steps:
[0062] Step S51: Maintain and run the spectrum cognitive world model for the dynamic environment of the Internet of Vehicles.
[0063] Based on the communication environment model constructed in step S1, this study addresses the issues of fast time-varying channels and buffer queues. Leveraging the characteristics of endogenous evolution, this invention runs a spectrum-based cognitive world model within a scheduling agent, based on a recursive state-space model (RSSM), internalizing environmental features through the mixing of latent state spaces. This model architecture specifically comprises the following four core units:
[0064] Sequential memory unit: based on the deterministic latent state of the previous moment. Random latent states and the joint action vector executed by the scheduler The action vector is derived from step S1. and The current deterministic state is formed according to the following formula. :
[0065]
[0066] in, For parameters Nonlinear recurrent activation units, such as GRUs, are used. These units iteratively apply historical channel gains. With the evolution of caching Implicit encoding to This effectively alleviates the partial observability problem of dynamic caching.
[0067] Representation unit: During the online learning phase, based on real-world environmental observations at the current moment. The observation includes the channel gain defined in step S1. Cache length And interference measurements, extract random states This is used to correct perceptions of the environment, and its posterior distribution inference process is expressed as:
[0068]
[0069] in, Indicates parameters The posterior probability distribution function.
[0070] Transition Unit: During the potential space deduction phase, the transition unit relies solely on its internal deterministic state. Predicting the prior distribution of the random state in the next time step This supports virtual imagination detached from real observation:
[0071]
[0072] Decoding unit: Maps the latent state back to the original observation space and reward space. Among them, the observation reconstruction value... Ensure the potential state has physical meaning, and reward the predicted value. Corresponding to the total information value reward of the system defined in step S1 The expression is as follows:
[0073]
[0074] Step S52: Perform dynamic parameter learning based on maximizing the lower bound of variational evidence.
[0075] To enable the spectrum-aware world model to accurately grasp the channel evolution patterns and buffer dynamics in the vehicle-to-everything (V2X) scenario, this invention utilizes the historical interaction trajectory data accumulated by the scheduler in step S5, and minimizes the total dynamic loss function. For model parameters Perform joint optimization:
[0076]
[0077] in, To observe the reconstruction loss and ensure that the model can reconstruct the complete communication environment characteristics based on finite spectrum feedback, including channel dynamics and buffer state, ; To reward the prediction loss, the model is trained to accurately predict different resource allocation schemes. Impact on the overall information value of the system; The KL divergence loss is used to constrain prior predictions to approximate posterior inferences, ensuring that virtual simulations conform to physical reality. For regularization weights. Furthermore, to enhance the model's stability in multi-step prediction planning, as described in step S4, this invention further introduces a potential overshoot loss. :
[0078]
[0079] in, The preset number of advance prediction steps, superscript Indicates the future The prediction term for each step. By co-optimizing the above loss function, the world model establishes a strong causal mapping between the scheduling decision history and the implicit cache dynamic state in the latent space, thus providing accurate dynamic support for efficient spectrum sharing.
[0080] Step S6: Perform predictive spectrum scheduling and online policy optimization based on potential inference.
[0081] This step aims to utilize the world model trained in step S5 for forward-looking planning and to achieve dynamic evolution of the strategy during the actual issuance of instructions. The specific sub-steps are as follows:
[0082] Step S61: Perform predictive planning and strategy learning based on perceived information value.
[0083] Based on the transfer unit described in step S51, the process is carried out within the potential space. Step-by-step virtual "imagination" deduction to generate an imagination trajectory Based on this trajectory, the policy network is updated using an actor-critic framework. The critic network minimizes the evaluation loss. This includes consideration of long-term cache overflow risks. -Target return value. Performer Network: Optimization strategy aimed at maximizing the long-term accumulated information value. This enables it to anticipate and mitigate the risk of future data packet expiration.
[0084] Step S62: Online spectrum scheduling and distribution, and the closed loop of perception-decision-learning.
[0085] The scheduler executes a real-time online inference stream: first, it maps the real-time observations... Encoding as latent states Next, decision-making is made by generating actions through performer network sampling. Includes the steps defined in step S1 and Next, feedback is provided, actions are performed in a real-world environment, and results are obtained. and The data is then stored in the experience playback unit. Finally, the system evolves by periodically sampling experience data and repeatedly triggering steps S52 and S61 to achieve continuous adaptive optimization of the vehicle-to-everything (V2X) dynamic environment.
[0086] The effects of this invention will be further illustrated below with simulation experiments. Specifically, this includes:
[0087] 1) Simulation conditions and parameter settings
[0088] The simulation experiments of this invention were conducted on a simulation platform using Python 3.11 and PyTorch 2.0. The experiments were carried out in a 600m × 600m Manhattan grid urban environment. The path loss exponent n = 3.0, the Rician K-factor = 3.0 dB, and the Gaussian white noise power spectral density were used. = -174 dBm / Hz. World model configuration: GRU latent state dimension is 200, random state dimension is 64; imagined trajectory length H = 15.
[0089] The comparison algorithms include the TD3 algorithm and the PPO algorithm.
[0090] 2) Simulation content
[0091] Figure 2 The figure shows a comparison of the convergence performance of the technical solution of the present invention and the baseline algorithm during the training process. Figure 2 The horizontal axis represents the number of training steps, and the vertical axis represents the total information value of the system. It can be seen that the method of this invention exhibits a rapid upward trend in the early stages of training, and after approximately 7.5 × 10^4 steps, the total information value of the system converges to a stable value with small variance. This indicates that the invention has extremely high sample efficiency and effectively ensures the timeliness of information through latent programming, thereby enhancing the value of information freshness for decision-making.
[0092] Figure 3 This is a performance comparison chart of the delay-constrained value guarantee efficiency of the method of this invention and the baseline algorithm under different signal-to-noise ratio conditions. The experimental signal-to-noise ratio is set from... Increment to As shown in the bar chart, the value assurance efficiency of each algorithm increases with the improvement of the signal-to-noise ratio (SNR). At all SNR levels, the performance of the method described in this invention significantly outperforms the two baseline algorithms. Because this invention incorporates a spectrum-based cognitive world model, it can accurately capture the fast fading characteristics of the channel, thus maintaining high information freshness even in low SNR environments. In contrast, the baseline algorithms (such as TD3) show a significant performance degradation under severe interference, demonstrating the robustness of this invention in complex communication environments.
[0093] Figure 4 This chart compares the throughput-to-value conversion efficiency of the proposed method and the baseline algorithm under different signal-to-noise ratio (SNR) conditions. Experimental results show that the proposed method maintains the highest conversion efficiency across the entire SNR range. This indicates that the proposed method does not blindly pursue maximizing physical layer throughput, but rather predicts future cache dynamics through a world model, prioritizing the allocation of limited spectrum resources to data packets with higher "information value." The baseline algorithm, lacking forward-looking planning, is prone to ineffective transmissions during SNR fluctuations, resulting in high throughput but low information conversion value. This verifies the superiority of the proposed method in optimizing resource utilization efficiency.
[0094] Figure 5 This graph compares the performance of the proposed method and the baseline algorithm under different transmission loads, highlighting the performance differences in latency-constrained value guarantee efficiency. As shown in the graph, the transmission load varies from... Increase to During the process, the performance of each algorithm in latency-constrained scenarios was observed. As the transmission load increased, the risk of network congestion increased, and the guarantee efficiency of all algorithms decreased. However, the curve of the method in this invention remained at the top and was far above the "data buffer boundary" dotted line. When the load exceeded... At that time, the performance of the baseline algorithms (TD3, PPO) experienced a significant decline, gradually approaching or falling below the buffer boundary, meaning that they could not stably cache under strict latency constraints. The method of this invention, with its multi-step extrapolation capability of the potential space, can predict the risk of buffer overflow under high load and adjust the scheduling strategy in advance, thereby effectively delaying performance degradation.
[0095] Figure 6 This graph compares the throughput-value conversion efficiency of the method of this invention and the baseline algorithm under different transmission loads, demonstrating the impact of different transmission loads on the system's value conversion efficiency. Under low load conditions, the differences between the algorithms are small; however, as the load continues to increase, the leading advantage of the method of this invention becomes increasingly apparent. Figure 6 The magnified view shows that during load fluctuations, the curve of the method described in this invention exhibits smaller fluctuations and remains consistently above the "resource utilization baseline." This demonstrates that the predictive planning based on latent imagination described in this invention enables scheduling decisions to maintain greater stability when traffic load intensifies, avoiding frequent oscillations between optimal states in the baseline algorithm.
[0096] It should be understood that the specific embodiments described above are merely illustrative or explanatory of the principles of the invention and do not constitute a limitation thereof. Therefore, any modifications, equivalent substitutions, improvements, etc., made without departing from the spirit and scope of the invention should be included within the protection scope of the invention. Furthermore, the appended claims are intended to cover all variations and modifications falling within the scope and boundaries of the appended claims, or equivalent forms of such scope and boundaries.
Claims
1. A method for information value perception-based vehicle-to-everything (V2X) spectrum sharing, characterized in that: The method includes the following steps: Step S1: Construct a vehicle-to-everything (V2X) communication environment model and define reinforcement learning elements; The system state space, action space, and reward function are initialized. The state space is composed of an explicit spectrum observation dimension and an implicit cache dynamic dimension. The reward function takes information value as the core indicator and aims to quantify the timeliness and effectiveness of the data. Step S2: Establish a spectrum-based cognitive world model for dynamic environments; The world model aims to map high-dimensional environmental observations to a low-dimensional latent state space and establish a state transition mechanism. The latent state representation adopts a hybrid architecture, including a deterministic component for remembering long-term historical features and a random component for characterizing channel random fluctuations. Step S3: Learn the potential dynamic characteristics of the environment based on the interactive trajectory; By utilizing historical trajectory data accumulated from the interaction between the scheduling agent and the environment, the spectrum-based cognitive world model is optimized end-to-end; by constructing a total objective function, the model can internalize the evolution law of the cache queue and accurately predict rewards. Step S4: Perform predictive planning based on potential space imagination; based on the optimized world model, conduct multi-step deduction in the potential space to generate an imagined trajectory; evaluate the long-term effectiveness of the candidate action sequence based on the trajectory, and generate a spectrum access strategy through the strategy generation unit with the goal of maximizing long-term cumulative returns. Step S5: Implement online scheduling control and closed-loop iterative optimization; The generated strategy is transformed into specific transmit power control and bandwidth allocation instructions. The observation and reward data from environmental feedback are collected in real time and the experience replay unit is updated. The world model and strategy parameters are iterated periodically.
2. The method for information value perception-based vehicle-to-everything (V2X) spectrum sharing according to claim 1, characterized in that, In step S1, the implicit cache dynamic state is the evolution state of the vehicle node cache queue, which is formed by the arrival of historical data packets and scheduling decisions, and has partial observability for the scheduler; the reward function is configured to be positively correlated with the cumulative information value of the entire network, and includes a penalty term for violating the preset interference threshold and a penalty term for cache queue instability.
3. The method for information value perception-based vehicle-to-everything (V2X) spectrum sharing according to claim 1, characterized in that, In step S2, the spectrum-based cognitive world model includes: a temporal memory unit, used to evolve the deterministic state characteristics of the current moment based on the potential states and actions of historical moments, in order to maintain the memory of historical channel and buffer dynamics; a state inference unit, used to combine the real environment observations of the current moment to generate the posterior potential state of the current moment, in order to correct the cognition of the current environment; a state prediction unit, used to predict the prior potential state of the current or future moments based only on historical information, in order to support virtual inference detached from the real environment; and an observation reconstruction and reward prediction unit, used to map the potential state back to the observation space and reward space, in order to ensure that the potential state has physical meaning and can guide task optimization.
4. The method for information value perception-based vehicle-to-everything (V2X) spectrum sharing according to claim 1, characterized in that, In step S3, the process of learning the latent dynamics of the environment includes: constructing a total objective function that includes an environment representation consistency loss, a task objective prediction loss, and a state distribution constraint loss; wherein, the environment representation consistency loss is used to constrain the reconstructed observations generated by the model to be consistent with the true observations, the task objective prediction loss is used to align the predicted reward with the environmental feedback reward, and the state distribution constraint loss is used to constrain the distribution of the prior latent state to approximate the distribution of the posterior latent state.
5. The information value perception vehicle-to-everything (V2X) spectrum sharing method based on a world model according to claim 1, characterized in that, In step S4, the specific steps of performing predictive planning based on latent space imagination include: starting from the potential state at the current moment, recursively predicting the sequence of potential states for several future time steps in the potential space using the spectrum cognitive world model to form the imagined trajectory; estimating the long-term cumulative return of the imagined trajectory using a value assessment unit, wherein the long-term cumulative return implies a prediction of possible future data packet expiration and buffer overflow; and generating the spectrum access control policy using a policy generation unit with the goal of maximizing the long-term cumulative return output by the value assessment unit.
6. The method for information value perception in vehicle-to-everything (V2X) networks based on a world model, as described in claim 5, is characterized in that... The input of the value assessment unit includes the potential states in the imagined trajectory, and the output is the state value estimate; the input of the policy generation unit includes the potential states, and the output is the probability distribution or deterministic action of the spectrum access action.
7. The method for information value perception-based vehicle-to-everything (V2X) spectrum sharing according to claim 1, characterized in that, In step S5, the online scheduling control and closed-loop iterative optimization process includes: vehicle nodes performing transmit power control and bandwidth allocation according to the actions output by the spectrum access control strategy; receiving the next moment's real observation and real reward from environmental feedback, and storing the current moment's observation, action, reward, and next moment's observation as experience data in the experience playback unit; periodically sampling experience data from the experience playback unit, repeating steps S3 and S4, and updating the spectrum cognitive world model and spectrum access control strategy.
8. A world-model-based information value perception vehicle-to-everything (V2X) spectrum sharing system, characterized in that, The system includes: The environment modeling and state perception module is used to acquire the environmental state of the vehicle-to-everything (V2X) communication network. The environmental state includes at least spectrum observation information and implicit states that reflect the dynamic evolution of the buffer queue, and generates information value reward evaluation indicators for spectrum access decisions. The spectrum-based cognitive world model construction module is used to construct a spectrum-based cognitive world model based on the environmental state, and to model and predict the evolution process of the environmental state. The decision planning module is used to generate a spectrum access control strategy based on the prediction results of the world model, under the condition of satisfying spectrum interference constraints. The spectrum access execution module is used to control the vehicle node to perform spectrum access operations according to the spectrum access control strategy. The model update module is used to update and optimize the spectrum cognitive world model and spectrum access control strategy based on communication feedback.
9. A world-model-based information value perception vehicle-to-everything (V2X) spectrum sharing system according to claim 8, characterized in that, The implicit cache dynamic state is the evolution state of the vehicle node cache queue. This state is formed by the arrival of historical data packets and scheduling decisions, and it has partial observability for the scheduler. The latent state representation includes at least a first latent component for characterizing historical correlation and a second latent component for characterizing randomness; The model update module is configured to train the world model based on historical interaction data between the scheduler and the environment, enabling it to predict the evolution of future potential states and corresponding environmental feedback. The decision planning module is configured to perform multi-step predictive planning in the potential state space constructed by the world model, so that potential future caching risks can affect the current spectrum scheduling decision. The decision planning module includes a value assessment unit and a policy generation unit. The value assessment unit is used to evaluate the long-term value of the predicted trajectory, and the policy generation unit is used to generate a spectrum access control policy based on the long-term value.
10. A vehicle-to-everything (V2X) spectrum sharing system based on a world model for perceiving information value, as described in claim 8, is characterized in that... The spectrum access execution module is configured to perform at least one of transmit power control and bandwidth allocation, and the model update module is configured to continuously update empirical data based on real-time environmental feedback and iteratively optimize the world model and spectrum access control strategy.
Citation Information
Patent Citations
Power consumption control methods, devices, storage media and electronic equipment
CN111818625B