Vehicle-mounted edge element universe optimization method and system capable of being energized by common inductance calculation

By introducing the synaesthesia-enabled method of active reasoning and deep reinforcement learning in the Internet of Vehicles, problems such as weak generalization ability and uneven coverage in the Internet of Vehicles are solved, efficient and stable resource allocation and decision optimization are achieved, and system performance and adaptability are improved.

CN120692598APending Publication Date: 2025-09-23XIAN UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510268400.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing technologies in the Internet of Vehicles have problems such as weak generalization ability, imbalance between exploration and utilization, high sample complexity, and uneven coverage of the Internet of Vehicles, which limit the combined application of multi-access edge computing and the Internet of Vehicles.

Method used

The in-vehicle edge metaverse optimization method empowered by synaesthesia is adopted, combined with active reasoning technology and deep reinforcement learning. Through the interaction between the intelligent agent and the environment, the free energy principle and reward function are used to optimize resource allocation and decision-making, build a multi-objective optimization model, realize dynamic channel perception and modeling, and support edge collaborative computing and security protection.

Benefits of technology

It significantly improves the efficiency and adaptability of the Internet of Vehicles system, reduces network latency, improves resource utilization, enhances system stability and reliability, and supports efficient and intelligent Internet of Vehicles communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120692598A_ABST
    Figure CN120692598A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of wireless communication, and discloses a vehicle-mounted edge element universe optimization method and system capable of being energized by a general inductance algorithm. The method combines multi-access edge computing and time-frequency diversity technologies, and aims to optimize the allocation and utilization efficiency of resources in the Internet of Vehicles. By introducing an active reasoning mechanism and a deep reinforcement learning algorithm, the method can dynamically adjust a resource allocation strategy in real time according to changes in an Internet of Vehicles environment. Firstly, the system processes communication requirements of the Internet of Vehicles equipment based on an MEC platform; and secondly, by utilizing deep reinforcement learning, the system continuously improves the accuracy and efficiency of resource allocation through autonomous learning and strategy optimization according to the current traffic and network conditions. According to the method, the delay of the system is effectively reduced, the communication quality between the vehicles is improved, the overall performance of the Internet of Vehicles is enhanced, and the method has higher adaptability and robustness especially in a high-density dynamic environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of wireless communication technology, and in particular relates to a vehicle-mounted edge metaverse optimization method and system enabled by telepathic computing. Background Art

[0002] Multi-access edge computing (MEC) is an emerging distributed computing architecture designed to deliver low-latency, high-bandwidth, and highly reliable services by deploying computing and storage resources at the edge of the network, close to access points (such as base stations or Wi-Fi access points). Compared to traditional cloud computing, MEC enables real-time processing of data locally or at nearby nodes, significantly reducing data transmission paths. This significantly meets the needs of scenarios requiring high real-time performance and resource efficiency, such as augmented reality, virtual reality, autonomous driving, and smart manufacturing. With the widespread adoption of 5G communication technology, the technical advantages of MEC are becoming increasingly prominent. Its deep integration with ultra-high bandwidth and ultra-low latency network characteristics provides a solid foundation for supporting the massive access of IoT devices and the processing of complex tasks. Research and application of MEC are currently experiencing rapid development. MEC not only supports edge processing of high-performance computing tasks, but also, through integration with artificial intelligence technologies, enables edge nodes to run lightweight deep learning models and achieve localized intelligent decision-making. Furthermore, MEC uses technologies such as network slicing to achieve dynamic resource allocation and on-demand scheduling, thereby improving overall system efficiency. More importantly, MEC provides an end-to-end low-latency solution that significantly optimizes user experience, such as achieving millisecond-level responses in real-time gaming, video streaming, and industrial control systems. Furthermore, with the increasing complexity of distributed architectures, MEC also faces challenges in data privacy and system security. Current research is focusing on multi-edge node collaboration, privacy-preserving algorithm design, and the improvement of network security defense mechanisms. Multi-access edge computing, as a crucial component driving 5G and subsequent network technologies, not only demonstrates broad application prospects in areas such as intelligent transportation and telemedicine, but also provides new directions and technical support for the future development of information technology.

[0003] The Internet of Vehicles (IoV) is an intelligent transportation network that integrates vehicles, road infrastructure, pedestrians, and cloud systems. It achieves efficient and safe traffic management and services through multi-dimensional information exchange and coordinated control. With the rapid development of key technologies such as autonomous driving, 5G communications, and edge computing, the IoV has become a crucial technical support for intelligent transportation and smart cities. Its core applications include traffic flow optimization, real-time navigation, vehicle status monitoring, and autonomous driving assistance, aiming to improve the safety, efficiency, and sustainability of transportation systems. 5G technology has significantly improved the low latency and high reliability of vehicle-to-vehicle, vehicle-to-road, and vehicle-to-cloud communications, providing technical support for scenarios with extremely high real-time requirements, such as autonomous driving. At the same time, the deep integration of edge computing and artificial intelligence has driven the data processing and decision-making capabilities of the IoV down to the lowest levels, enhancing the system's real-time performance and adaptability to complex and dynamic traffic environments.

[0004] In practical applications, Hangzhou's smart traffic management system has achieved a deep integration of vehicle-to-vehicle (IoV) technology, edge computing, and 5G communications by integrating real-time traffic flow data, road infrastructure data, and dynamic environmental perception data. This demonstrates the potential of intelligent transportation technologies in urban governance. Furthermore, the deployment of a collaborative autonomous driving test platform on highways demonstrates the significant impact these technologies have on improving traffic flow efficiency and driving safety.

[0005] When integrating MEC into the Internet of Vehicles architecture, several challenges need to be addressed:

[0006] (1) Collaboration and resource management of heterogeneous networks: The Internet of Vehicles involves multiple communication protocols and network topologies. MEC nodes need to simultaneously handle the diverse needs of vehicle-to-vehicle communication, vehicle-to-road communication, and vehicle-to-cloud communication. To achieve efficient collaboration between MEC and IoV, it is necessary to develop optimized resource management algorithms to dynamically allocate edge computing resources and wireless resources to ensure the service quality of different communication tasks.

[0007] (2) Guarantee of low latency and highly reliable communication: Internet of Vehicles applications such as collaborative autonomous driving and real-time navigation have extremely high requirements for communication latency and reliability. The combination of MEC and IoV requires maintaining millisecond-level end-to-end latency control in complex traffic scenarios, while also addressing issues such as frequent handoffs and channel fading caused by high-speed vehicle movement. This requires incorporating efficient scheduling mechanisms and fault tolerance into the architecture design.

[0008] (4) Edge intelligence and communication collaborative optimization: MEC nodes need to run deep learning algorithms for environmental perception and decision-making, but the complexity of these algorithms may affect real-time performance. The combination of IoV and MEC requires the design of lightweight, real-time edge intelligence algorithms, while solving the dynamic collaborative optimization problem between computing tasks and communication resources to improve overall system performance.

[0009] (5) Security and privacy protection: The distributed nature of the Internet of Vehicles (IoV) creates more potential attack surfaces. The data processing of MEC nodes and the channel dynamics of IoV communication may lead to security vulnerabilities such as signal tampering or malicious attacks. To achieve the combination of the two, it is necessary to strengthen edge security protection mechanisms and design privacy protection algorithms for distributed architectures to ensure the security and integrity of data during transmission and processing.

[0010] Secondly, in the in-vehicle edge metaverse system powered by synaesthesia, deep reinforcement learning (DRL) has been proven to be an effective optimization solution. DRL is a key branch of machine learning that combines the advantages of deep learning and reinforcement learning. Reinforcement learning learns optimal decision-making strategies through the interaction between an agent and its environment. The agent obtains reward signals through trial and error and continuously adjusts its behavior to maximize long-term returns. Deep learning uses deep neural networks to efficiently represent and process complex data. DRL combines these two approaches, leveraging deep neural networks to process high-dimensional and complex state spaces, enabling agents to make decisions in dynamic environments with incomplete knowledge. DRL is widely used in fields such as autonomous driving, robotic control, and game AI. In these applications, DRL can help agents achieve self-learning and optimization in complex environments without relying on manual programming. Through large-scale data training, DRL can continuously improve performance and cope with challenges in unknown environments and changing conditions. It is one of the key technologies driving the development of adaptive decision-making capabilities in artificial intelligence. However, traditional DRL methods have some drawbacks:

[0011] (1) Instability and convergence issues: Because the DRL algorithm relies on the training of deep neural networks, the algorithm is prone to instability, especially in the early stages of training, when the neural network parameter updates may cause severe fluctuations in the strategy. In addition, in some cases, traditional DRL algorithms have difficulty ensuring stable convergence in complex environments.

[0012] (2) Low exploration efficiency: When faced with large-scale and high-dimensional state spaces, traditional DRL methods may have inefficient exploration strategies during the learning process, causing the intelligent agent to fall into a local optimal solution during the learning process and making it difficult to reach the global optimal solution.

[0013] (3) Dependence on environment modeling: Traditional DRL algorithms usually assume that the environment is fully observable, but in practical applications, many environments are partially observable, which requires DRL algorithms to be able to handle partially observable Markov decision processes, which is more complicated in practice.

[0014] (4) Sample imbalance and reward delay: In some tasks, the rewards obtained by the agent may be sparse and delayed. Traditional DRL algorithms find it difficult to effectively handle rewards with long time spans, which may lead to slow convergence or poor performance during the learning process.

[0015] The technical problems existing in the industrial application of existing technologies are mainly reflected in the following aspects:

[0016] 1. Weak generalization ability:

[0017] The poor generalization ability of the DRL algorithm means that the algorithm performs well in the training environment, but performs poorly in unknown environments. This is mainly due to the overfitting problem, where the algorithm relies too much on the specific characteristics of the training environment, resulting in an inability to adapt to new situations. In addition, DRL usually requires a large number of samples for training, and the complexity and diversity of the training environment make it difficult for the algorithm to maintain good performance in the dynamically changing real environment. Therefore, the lack of effective policy transfer and adaptability is the main reason for its poor generalization ability. These problems limit the widespread promotion of DRL in practical applications, prompting researchers to seek to improve its generalization ability through transfer learning, multi-task learning, and other methods.

[0018] 2. Exploration and Exploitation Imbalance:

[0019] In DRL, the imbalance between exploration and exploitation refers to the trade-off faced by intelligent agents during the learning process: exploration involves trying new actions to gain more information about the environment, while exploitation involves selecting the optimal action based on a known strategy. Excessive exploration can lead to wasted time on suboptimal strategies, resulting in inefficient training. Excessive exploitation can trap agents in local optima, making them unable to adapt to environmental changes and affecting their generalization capabilities.

[0020] 3. High sample complexity:

[0021] Traditional DRL algorithms typically require a large number of interaction samples to converge, especially in complex environments. High sample complexity leads to long training times and significant computational resource consumption. Acquiring sufficient data is often costly and time-consuming, especially in real-world applications.

[0022] 4. Uneven coverage of the Internet of Vehicles:

[0023] The Internet of Vehicles (IoV) requires stable communication services across large geographic areas. However, significant differences in network infrastructure between urban and rural areas, and between roads and non-road areas, result in limited network signal coverage in remote areas, tunnels, and overpasses, hindering the widespread adoption and reliability of IoV services.

[0024] In summary, the main technical challenges facing existing technologies in industrial applications are weak generalization, an imbalance between exploration and exploitation, high sample complexity, and uneven coverage of the Internet of Vehicles. These issues limit the effectiveness and scope of existing technologies in practical applications, necessitating the introduction of new mechanisms and methods for improvement and optimization. Summary of the Invention

[0025] In response to the problems existing in the prior art, the present invention provides a vehicle-mounted edge metaverse optimization method and system enabled by synaesthesia computing.

[0026] The present invention is implemented as follows: a vehicle-mounted edge metaverse optimization method and system enabled by synaesthesia computing includes:

[0027] S101, initialize optimization algorithm parameters; initialize initial state s t , transfer model parameters θ1, reward model parameters θ2 (the union of θ1 and θ2 is θ), number of rounds M, maximum number of training steps T, number of alternative strategies J, number of optimal candidate strategies k, initialization of global network learning rate α, discount factor γ, initialization of replay buffer size D, initialization of strategy distribution q(π), initialization of transition probability distribution p(s τ |s τ-1 ,θ,π); initialize hyper parameters, such as system bandwidth B m , channel gain A, etc.; initialize random parameters, such as drone transmission power The number of vehicles N, etc.

[0028] S102, except for the first round, the initial state s t Due to the state transition, the strategy distribution is randomly sampled to obtain J alternative strategies;

[0029] S103. Randomly sample J actions from each of these strategies, then obtain J conditional transition probability distributions based on the J alternative strategies, and calculate the corresponding current rewards from the corresponding J actions;

[0030] S104, the state s t Input reward model to get reward r t According to the active reasoning and free energy principle, the free energy of each alternative strategy is obtained by using the conditional transition probability distribution and the transition model output, and the indicative index is obtained by adding the two together.

[0031] S105, for the first k smallest Find the average and get the current strategy distribution based on the average;

[0032] S106: Sampling a strategy based on the obtained strategy distribution, and then sampling an action as the current agent's behavior, and interacting with the environment to obtain the next state;

[0033] S107, storing the experience tuple in the replay buffer; if the replay buffer is full, deleting the oldest experience and storing the latest experience;

[0034] S108, calculating the error function between the predicted value of the global network output free energy and the actual free energy, and updating the global network parameters using the back propagation algorithm and gradient descent;

[0035] S109. Repeat the training until the algorithm converges and finally obtains the strategy distribution, so that the strategy of each step can be randomly extracted from the distribution to control the action of the intelligent agent and obtain the optimal resource control allocation.

[0036] Furthermore, in S102, at the beginning of each round, the initial state is set; J candidate strategies are randomly sampled from the strategy distribution, and then J actions are randomly sampled from each of the J strategies; the state of the agent is represented as:

[0037]

[0038] Where D(t)={D n (t)}, D n (t) is the data volume requirement for vehicle n in time slot t; is the mAP demand of vehicle n for parking spaces i in time period t; η(t) = {η n (t)}, η n (t) is the direction of vehicle n at time slot t; X(t) = {X n (t)}, X n (t) is the x-axis position of vehicle n at time slot t; Y(t) = {Y n (t)}, Y n (t) is the y-axis position of vehicle n at time slot t; is the wireless resource price of edge server m in time slot t; is the transmission power price of edge server m in time slot t. t The sampled action is represented as:

[0039] a(t)={s(t),P(t),fser (t)}.,

[0040] in, represents the resolution transmitted from the edge server to the vehicle; P(t) = {P m,n (t)}, Indicates the transmit power allocation; Indicates the allocation of computing resources.

[0041] Furthermore, in step S103, J conditional transition probability distributions are obtained based on the J alternative strategies, and the corresponding current rewards are calculated from the corresponding J actions. The calculation formula of the immediate reward r(t) is as follows:

[0042]

[0043] Among them, γ1 and γ2 are weight coefficients used to balance the relative size of vehicle QoE and total delay, ensuring that the two are on the same scale and unifying the units of the two. In addition, ∈1 and ∈2 are used to balance the importance of the two items, satisfying 1,2 ∈[0,1], and 1+2=1. Obviously, the larger the value, the better.

[0044] Furthermore, the S104: according to active inference and free energy principle, using the conditional transition probability distribution p(s τ |s τ-1 ,θ,π) and the output of the transfer model to obtain the indicators of J alternative strategies in The calculation formula is:

[0045]

[0046] The above S105: averaging the first k smallest free energies and obtaining the current strategy distribution based on the average value is equivalent to taking the first k maximum values ​​of the opposite values ​​of the free energies, sorting the values ​​from large to small, and then obtaining the average value:

[0047]

[0048] And get the policy distribution:

[0049]

[0050] where σ(·) represents a continuous distribution related to the exponential of the natural number e, such as the exponential distribution and the gamma distribution;

[0051] S106: Sampling strategy π based on the obtained strategy distribution q(π) t , and then sample action a tAs the current agent's behavior, it interacts with the environment to obtain the next state; the extraction process is summarized as follows:

[0052] π t ~q(π),a t ~π t

[0053] Furthermore, the S108: transfer model network output free energy prediction value Q(s t ,a t ; θ1), and the target (ie, actual free energy) The error function L(θ1) is used to update the transfer model network parameters using the back propagation algorithm and gradient descent; the loss function is given by the following formula:

[0054]

[0055] The gradient of the loss function is given by:

[0056]

[0057] Then update the parameters of the transfer model network through gradient descent as follows:

[0058]

[0059] Where θ1 is the internal parameter of the transfer model network, and α is the learning rate. Similarly, the reward model network parameter θ2 is updated using this method.

[0060] Another object of the present invention is to propose a system that uses an active inference-based optimization method to assist in the in-vehicle edge metaverse system enabled by synaesthesia. The system consists of the following key modules:

[0061] System Initialization Module: This module is responsible for initializing key parameters required for the backpropagation algorithm, including the network's learning rate, discount factor, and replay buffer size. It also builds the network architecture and sets parameters related to network layout, such as the number of drones and the location of ground edge servers.

[0062] Agent Decision Module: At the beginning of each cycle, this module generates appropriate action decisions based on the current network state. It combines the active inference mechanism and the free energy principle to optimize the policy distribution by selecting the mean of the previous stage that minimizes the free energy.

[0063] Action execution module: This module is responsible for actually executing resource allocation operations and putting decisions into practice.

[0064] Reward Evaluation Module: After an action is executed, this module calculates the immediate reward based on the inverse of the weighted average of latency, energy consumption, and operator cost of all devices in the system, and manages the transition process from the current state to the next state.

[0065] Experience storage module: This module stores the system state, executed actions, rewards obtained, accumulated rewards, and next state after each operation, forming experience tuples for subsequent learning and training.

[0066] Sample sampling module: This module is responsible for randomly extracting samples from the policy distribution and self-sampling from the policy itself for the agent to optimize its decision.

[0067] Global model update module: This module uses the back-propagation algorithm to update the global neural network, and uses a parameter optimization unit to adjust the network parameters using the gradient descent method to achieve optimization.

[0068] Another object of the present invention is to provide a computer device comprising a memory and a processor, wherein the memory stores a computer program. When the processor executes the program, the method and system for optimizing the in-vehicle edge metaverse enabled by synaesthesia computing are implemented according to specified steps.

[0069] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, prompts the processor to execute a vehicle-mounted edge metaverse optimization process enabled by synaesthesia.

[0070] Another object of the present invention is to provide an information data processing terminal specifically used to realize the system function of the above-mentioned synaesthesia-enabled in-vehicle edge metaverse optimization.

[0071] In combination with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are as follows:

[0072] First, this invention develops an innovative system architecture that significantly improves system performance and adaptability by integrating MEC technology into the IoV. This invention develops an innovative system architecture that significantly improves system performance and adaptability by integrating mobile edge computing technology into the IoV. In this architecture, MEC nodes are designed as distributed computing and storage units capable of processing large amounts of data generated by vehicles in real time, thereby reducing the load on central servers and significantly reducing network latency. Furthermore, this system fully considers the dynamic nature of the IoV environment, achieving efficient resource allocation and load balancing through intelligent scheduling algorithms, ensuring service quality in different scenarios. Furthermore, to further enhance the system's adaptability, this architecture incorporates deep learning and edge intelligence technologies, enabling it to dynamically perceive environmental changes and automatically adjust strategies to meet the ever-changing communication needs between vehicles and road infrastructure. By combining the efficient data processing capabilities of MEC technology, the overall communication performance of the IoV is further improved, providing a new technical path for building efficient and intelligent IoV communication systems in the future.

[0073] This paper proposes a synaesthesia-enabled in-vehicle edge metaverse optimization method and system. This method, based on active inference technology, achieves the following key technological advancements:

[0074] (1) Dynamic channel perception and modeling: This paper constructs a model that can accurately capture the characteristics of dynamic vehicle network channels, improving communication reliability in complex multipath propagation and high-speed movement scenarios.

[0075] (2) Intelligent resource allocation: The present invention uses active inference technology to achieve adaptive optimization allocation of computing resources, communication resources, and storage resources in the Internet of Vehicles, significantly improving system resource utilization and reducing network latency and communication overhead.

[0076] (3) Multi-objective optimization strategy: The present invention adopts a multi-objective optimization algorithm, taking into account the requirements of delay, energy consumption and service quality, and realizes the global optimization of resource allocation in a highly dynamic environment.

[0077] (4) Scalable system architecture: This paper proposes an architecture design with good scalability, which supports seamless integration with other Internet of Vehicles technologies (such as V2X communication protocol and 5G network), and provides a solid foundation for the development of future intelligent transportation systems.

[0078] The core of the proposed in-vehicle edge metaverse optimization method and system, powered by synaesthesia, lies in the use of mathematical models to guide the system's behavior and learning process. The characteristics of these mathematical models and their technical effects are as follows:

[0079] (1) Resource Allocation Optimization Model: Using a multi-objective optimization approach, a resource allocation mathematical model was established that simultaneously considers latency, energy consumption, and service quality. This achieved the global optimal allocation of system resources, balanced various performance indicators, and significantly improved the overall system performance.

[0080] (2) Edge Collaborative Computing Model: Based on the MEC architecture, an edge node collaborative optimization model is designed to support task allocation and collaborative processing among multiple nodes. This significantly improves the utilization efficiency of edge computing resources and reduces the load and communication delay of the central server.

[0081] (3) Real-time reasoning and adaptive optimization model: An adaptive optimization model is built through active reasoning technology, which supports dynamic adjustment of system strategies to adapt to environmental changes. This ensures that the system can still respond quickly in complex and dynamic environments, improving real-time performance and intelligence.

[0082] (4) Innovative incentive mechanism: Taking both free energy and reward into consideration, and combining machine learning technology with the basic principles of biological brain neural activity, an incentive strategy based on a simple additive form is designed. This mechanism captures the uncertainty of the system and the changes in information entropy through dynamic adjustment of free energy, reflecting the real-time characteristics of environmental complexity; at the same time, it strengthens goal-oriented behavior optimization through the reward function to ensure the effectiveness and controllability of resource allocation and decision-making processes. The integration of free energy and reward not only improves the model's adaptability to complex Internet of Vehicles environments, but also draws on the generation mechanism of reward signals in biological neural systems, providing theoretical support for brain-like intelligence for incentive strategies. Compared with traditional incentive methods, this mechanism shows higher stability and learning efficiency in complex dynamic environments, promoting the further development of intelligent resource allocation optimization in Internet of Vehicles.

[0083] The application of the mathematical model provided by this invention not only improves the operational efficiency and decision-making quality of the synaesthesia-enabled in-vehicle edge metaverse optimization system, but also enhances its adaptability to environmental changes and long-term stability. These technical effects are crucial for modern edge computing environments that process large amounts of data and high-frequency interactions.

[0084] This paper proposes a synaesthesia-enabled in-vehicle edge metaverse optimization method and system, which improves network performance through the interaction between intelligent agents and the environment. Initializing the agent's state involves multiple parameters, including the number of vehicles and base stations, the distance between the vehicles and base stations, the channel gain between the edge server and the vehicle, the vehicle's two-dimensional position coordinates, and the vehicle's bandwidth. These parameters collectively constitute the agent's environmental state at a given moment and influence its decision-making.

[0085] The algorithm is still based on DRL, and the process mainly includes four stages: environmental interaction, policy update, value evaluation and model optimization. The agent interacts with the environment (Environment), selects an action (Action) based on the current state (State), and obtains a reward (Reward) and the next state based on environmental feedback. The agent uses a deep neural network to approximate the policy function, and guides behavior optimization through an indicator obtained by summing the maximum reward (whose calculation takes into account the long-term delay of all devices in the system and the quality of vehicle experience, aiming to improve the user experience in many aspects) and free energy. During the training process, the agent uses a combination of exploration and utilization to collect experience, stores it in the experience replay pool, and updates the model through small batch sampling to improve learning efficiency and stability. The entire process is continuously optimized in a cyclic iterative manner until the agent shows sufficient performance on the target task or reaches a convergence state. These steps and the application of mathematical models have brought significant technological progress:

[0086] Improved multi-objective optimization capabilities: Free energy is introduced as an additional metric to help intelligent agents identify and cope with uncertainties within the system, improving robustness and stability in dynamic environments.

[0087] Free-energy-guided uncertainty handling: This approach leverages deep neural networks to dynamically approximate the policy function, enabling the agent to rapidly adapt to changes in the complex connected vehicle environment and optimize decision-making. The combination of small-batch sampling and deep network training significantly accelerates the algorithm's training convergence speed while ensuring the stability and reliability of policy optimization.

[0088] Enhanced real-time dynamic optimization capabilities: By combining exploration and exploitation strategies, the intelligent agent can achieve real-time strategy optimization during task execution, improving the response speed and flexibility in the Internet of Vehicles environment.

[0089] These advances demonstrate the potential of deep reinforcement learning with active inference in complex network systems, especially in realizing intelligent and efficient vehicle-to-vehicle communication networks.

[0090] Second, traditional deep reinforcement learning algorithms usually rely on the design of several fixed reward functions, which limits their applicability to specific scenarios. Once the demand or application scenario changes, the training results are often difficult to meet expectations and show poor generalization ability. Although studies have attempted to integrate theoretical methods (such as Lyapunov optimization theory and attention mechanism) into the DRL framework to enhance its adaptability, the problem of insufficient generalization ability has not been completely solved. The present invention significantly improves the generalization performance of the DRL algorithm by introducing the active reasoning mechanism in brain neuroscience and using the free energy principle to guide algorithm design. Specifically, the free energy is combined with the reward function, and the DRL technology and biological neuroscience theory are integrated in an additive form, guiding the traditional "black box" network in a direction that is more in line with biological instincts. This method not only overcomes the limitations of traditional algorithms, but also enables the algorithm to show higher adaptability and stability in a variety of dynamic environments.

[0091] Third, this paper proposes an innovative solution to the difficulty faced by traditional deep reinforcement learning algorithms in effectively balancing the exploration of new strategies with the exploitation of existing knowledge. By simultaneously considering the free energy and reward values ​​of candidate strategies and selecting the optimal action based on the weighted result, this method significantly improves the balance between exploration and exploitation, enhancing the overall efficiency and performance of the algorithm.

[0092] This invention further optimizes the resource allocation mechanism in the Internet of Vehicles. By introducing policy distribution and conditional transition probability distribution, the algorithm can dynamically adjust the resource utilization strategy of vehicles and other devices. This not only significantly improves resource utilization but also effectively enhances system stability and reliability.

[0093] This invention has also demonstrated remarkable technical value in practical applications. By optimizing vehicle transmission power and computing resource allocation, it has successfully achieved a significant reduction in system latency and significantly improved the service quality of vehicle users, providing strong technical support for the in-depth application of MEC in the Internet of Vehicles. BRIEF DESCRIPTION OF THE DRAWINGS

[0094] Figure 1 This is a flowchart of the vehicle-mounted edge metaverse optimization method and system enabled by synaesthesia computing provided in an embodiment of the present invention.

[0095] Figure 2 This is an applicable scenario diagram provided by an embodiment of the present invention.

[0096] Figure 3 This is a flowchart of the implementation of the in-vehicle edge metaverse optimization method and system enabled by synaesthesia computing provided in an embodiment of the present invention.

[0097] Figure 4This is a diagram showing the relationship between the accumulated system reward and the system bandwidth after simulating the convergence of the algorithm, provided by an embodiment of the present invention.

[0098] Figure 5 This is a graph comparing the convergence performance of the embodiment of the present invention with several baseline algorithms.

[0099] Figure 6 This is a structural block diagram of the vehicle-mounted edge metaverse optimization system enabled by synaesthesia computing provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0100] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0101] The technical solution of this invention proposes an optimization method for a vehicle-to-vehicle communication system that supports multi-access edge computing (MEC). The method is based on active inference and deep reinforcement learning (DRL) framework. The key components of the solution include:

[0102] (1) Deep reinforcement learning framework: A deep neural network is constructed as the global decision model, and an active reasoning-enhanced DRL method based on the free energy principle is used to perform strategy learning and decision optimization.

[0103] (2) Network configuration: Initialize the parameters of the global model network and set key hyperparameters in the training process, such as learning rate, discount factor, etc., to ensure the efficiency of the training process.

[0104] (3) Environmental interaction and status update: The intelligent agent interacts with multiple alternative strategies in a virtual simulation environment, selects the optimal strategy based on a comprehensive evaluation of free energy and reward, and updates the system status information based on the execution results.

[0105] (4) Experience replay mechanism: Use the experience replay pool to store historical interaction data, and use cumulative rewards rather than immediate rewards as the basis for learning. When the buffer reaches the capacity limit, the system will update to ensure the timeliness of the training data.

[0106] (5) Network parameter optimization: By calculating the loss function and using the gradient descent method to adjust the parameters of the global model, the strategy selection and resource allocation in the decision-making process are continuously optimized.

[0107] (6) Strategy convergence and implementation: After ensuring that the strategy reaches a stable state through continuous training, the strategy is applied to optimize the path planning and task offloading of the UAV, thereby improving the overall efficiency and stability of the system.

[0108] The following are two specific embodiments and their implementation schemes provided by the present invention:

[0109] Example 1: Intelligent Traffic Flow Management System

[0110] The system optimizes traffic flow management by leveraging MEC and IoV technologies through the collaborative work of vehicles and road infrastructure.

[0111] 1) Network configuration: Configure the communication parameters of vehicles and base stations based on the road network scale, vehicle speed, and traffic demand, and optimize base station selection and network topology.

[0112] 2) Strategy interaction: Vehicles use deep reinforcement learning algorithms to select the optimal driving route and speed based on traffic flow monitoring data and environmental changes, and optimize traffic flow based on the principle of free energy minimization.

[0113] 3) Data collection and analysis: Vehicles collect traffic conditions and vehicle information through on-board sensors and edge computing nodes, and transmit them to the MEC server in real time for data processing and analysis.

[0114] 4) Task optimization and resource management: Use deep reinforcement learning strategies to optimize the distribution of computing tasks, reduce communication delays between vehicles, and offload tasks on the MEC platform to improve system efficiency.

[0115] 5) Network Adjustment and Performance Improvement: Dynamically adjust the base station's load distribution and routing strategies based on real-time vehicle data and network resource usage to ensure optimal system response speed and overall performance.

[0116] Example 2: Autonomous Vehicle Collaborative Control System

[0117] The system supports collaborative control and path planning between autonomous vehicles through MEC and IoV technologies.

[0118] 1) Network configuration: Optimize the communication connection between the MEC server and vehicles based on the distribution, movement speed, and communication requirements of the vehicles, and select appropriate signal processing methods to cope with the dynamically changing environment.

[0119] 2) Strategy Interaction: Autonomous vehicles use deep reinforcement learning to select the optimal driving strategy based on real-time traffic data and pre-set safety policies. By summing free energy and reward functions, they optimize coordinated control actions between vehicles, enabling autonomous driving in a platoon.

[0120] 3) Data collection and analysis: Vehicles and road facilities collect road information through sensors and edge computing nodes, and send the data to the MEC platform for real-time processing and decision support.

[0121] 4) Task Optimization and Resource Management: Utilize deep reinforcement learning strategies to optimize the allocation of computing tasks and communication resources in autonomous driving systems, reducing latency and improving the accuracy of collaborative control.

[0122] 5) Network adjustment and performance improvement: Dynamically adjust signal processing parameters based on vehicle location and communication conditions to ensure efficient and stable communication connections, while optimizing task scheduling and resource utilization of the MEC platform.

[0123] Example 3

[0124] In response to the challenges in current technology, this embodiment provides a method and system for optimizing the edge metaverse of a vehicle enabled by synaesthesia. Figure 2 As shown, the method of the present invention is applicable to specific application scenarios. The vehicle network of interest is a multi-access vehicle-mounted system consisting of M base stations and N vehicles, such as Figure 2 As shown. Let the base station set be As a database service provider. There are N vehicles randomly distributed on the road, denoted as Since the presented services are owned by different providers, each vehicle must purchase them from the base station. For each vehicle, the service request content of vehicle n can be expressed as V n (t) = {D n (t),mAP n (t)},D n (t) represents the data size of the service, which is determined by the needs of the vehicle. In the system of the present invention, the service of each vehicle is divided into 1 3D block with different resolutions. represents the set of slice resolutions for vehicle n, where represents the resolution of tile i for vehicle n. Since the contribution of each tile to vehicle QoE depends on its position and orientation, the resolution of each tile is adjusted according to its importance. The mAPn(t) score plays a key role in influencing these adjustments to ensure the optimal resolution of the tile. Notation Represents a collection of time slots.

[0125] In this scenario, at the beginning of each time slot, a user in a car, located anywhere, submits a virtual service request to the operator. This request is forwarded to the base station, which then forwards the task to the edge server and charges the user for the data. Ultimately, all data is processed and returned by the edge server, alleviating the heavy load on the base station and allowing the operator to collect the fee.

[0126] Secondly, the present invention provides an in-vehicle edge metaverse optimization system powered by telepresence computing. This system is networked with several vehicles, base stations, edge servers, and other IoT devices. During each fixed time slot, in-vehicle user devices generate tasks, which are then processed by the edge computing platform within the specified time slot. By rationally allocating limited resources in the communication system, the overall user experience quality of the system is significantly improved.

[0127] Example 4

[0128] Given the dynamic nature of the network and the uncertainty of information acquisition, the problem becomes quite challenging. To effectively address it, we reformulate the problem as a Markov decision process (MDP). Because different environments have different orientations, we use an improved reinforcement learning algorithm based on active inference. This algorithm supports both continuous and discrete action spaces, enabling real-time online decision making.

[0129] like Figure 1 As shown, the in-vehicle edge metaverse optimization method and system enabled by synaesthesia computing provided in this embodiment is characterized by comprising the following steps:

[0130] S101, initialize optimization algorithm parameters; initialize initial state s t , transfer model parameters θ1, reward model parameters θ2 (the union of θ1 and θ2 is θ), number of rounds M, maximum number of training steps T, number of alternative strategies J, number of optimal candidate strategies k, initialization of global network learning rate α, discount factor γ, initialization of replay buffer size D, initialization of strategy distribution q(π), initialization of transition probability distribution p(s τ |s τ-1 ,θ,π); initialize hyper parameters, such as system bandwidth B m , channel gain A, etc.; initialize random parameters, such as drone transmission power P m trans (t), the number of vehicles N, etc.

[0131] S102, except for the first round, the initial state s t Due to the state transition, the strategy distribution is randomly sampled to obtain J alternative strategies;

[0132] S103. Randomly sample J actions from each of these strategies, then obtain J conditional transition probability distributions based on the J alternative strategies, and calculate the corresponding current rewards from the corresponding J actions;

[0133] S104, the state s t Input reward model to get reward r tAccording to the active reasoning and free energy principle, the free energy of each alternative strategy is obtained by using the conditional transition probability distribution and the transition model output, and the indicative index is obtained by adding the two together.

[0134] S105, for the first k smallest Find the average and get the current strategy distribution based on the average;

[0135] S106: Sampling a strategy based on the obtained strategy distribution, and then sampling an action as the current agent's behavior, and interacting with the environment to obtain the next state;

[0136] S107, storing the experience tuple in the replay buffer; if the replay buffer is full, deleting the oldest experience and storing the latest experience;

[0137] S108, calculating the error function between the predicted value of the global network output free energy and the actual free energy, and updating the global network parameters using the back propagation algorithm and gradient descent;

[0138] S109. Repeat the training until the algorithm converges and finally obtains the strategy distribution, so that the strategy of each step can be randomly extracted from the distribution to control the action of the intelligent agent and obtain the optimal resource control allocation.

[0139] At the beginning of each round, the initial state is set in S102. J candidate strategies are randomly sampled from the strategy distribution, and then J actions are randomly sampled from each of the J strategies. The state of the agent is represented as:

[0140]

[0141] Where D(t)={D n (t)}, D n (t) is the data volume requirement of vehicle n in time slot t; is the mAP demand of vehicle n for parking spaces i in time period t; η(t) = {η n (t)}, η n (t) is the x-axis position of vehicle n at time slot t; X(t) = {X n (t)}, X n (t) is the x-axis position of vehicle n at time slot t; Y(t) = {Y n (t)}, Y n (t) is the y-axis position of vehicle n at time slot t; is the wireless resource price of edge server m in time slot t; is the transmission power price of edge server m in time slot t. t The sampled action is represented as:

[0142] a(t)={s(t),P(t),f ser (t)}.,

[0143] in, represents the resolution transmitted from the edge server to the vehicle; P(t) = {P m,n (t)}, Indicates the transmit power allocation; Indicates the allocation of computing resources.

[0144] S103: Based on the J alternative strategies, J conditional transition probability distributions are obtained, and the corresponding current rewards are calculated from the corresponding J actions. The calculation formula of the immediate reward r(t) is as follows:

[0145]

[0146] Among them, γ1 and γ2 are weight coefficients used to balance the relative size of vehicle QoE and total delay, ensuring that the two are on the same scale and unifying the units of the two. In addition, ∈1 and ∈2 are used to balance the importance of the two items, satisfying 1,2 ∈[0,1], and 1+2=1. Obviously, the larger the value, the better.

[0147] S104: According to active inference and free energy principle, using conditional transition probability distribution p(s τ |s τ-1 ,θ,π) and the output of the transfer model to obtain the indicators of J alternative strategies in The calculation formula is:

[0148]

[0149] The above S105: averaging the first k smallest free energies and obtaining the current strategy distribution based on the average value is equivalent to taking the first k maximum values ​​of the opposite values ​​of the free energies, sorting the values ​​from large to small, and then obtaining the average value:

[0150]

[0151] And get the policy distribution:

[0152]

[0153] where σ(·) represents a continuous distribution related to the exponential of the natural number e, such as the exponential distribution and the gamma distribution;

[0154] S106: Sampling strategy π based on the obtained strategy distribution q(π) t , and then sample action a t As the current agent's behavior, it interacts with the environment to obtain the next state; the extraction process is summarized as follows:

[0155] π t ~q(π),a t ~π t

[0156] S108: Transfer model network outputs predicted value Q(s) of free energy t ,a t ; θ1), and the target (ie, actual free energy) The error function L(θ1) is used to update the transfer model network parameters using the back propagation algorithm and gradient descent; the loss function is given by the following formula:

[0157]

[0158] The gradient of the loss function is given by:

[0159]

[0160] Then update the parameters of the transfer model network through gradient descent as follows:

[0161]

[0162] Where θ1 is the internal parameter of the transfer model network, and α is the learning rate. Similarly, the reward model network parameter θ2 is updated using this method.

[0163] In order to elaborate on the in-vehicle edge metaverse optimization method and system enabled by synaesthesia, the present invention provides two specific application embodiments, including key details of the implementation scheme.

[0164] In order to prove the creativity and technical value of the technical solution of the present invention, this section provides application examples of the claimed technical solution on specific products or related technologies.

[0165] Case Application 1: Intelligent Traffic Flow Management Solution

[0166] 1) System Startup and Configuration: In the IoV traffic management center, the global model network is deployed and started, and network parameters are initialized, including key hyperparameters such as the critic network's learning rate, the replay buffer capacity, and the time slot length. Simultaneously, vehicles are deployed to collaborate with traffic infrastructure nodes (such as smart traffic lights and road sensors), leveraging the MEC platform to ensure efficient traffic monitoring and real-time adjustments.

[0167] 2) Strategy Interaction and Environmental Adaptation: The vehicle adjusts its route and speed based on active inference, based on traffic data and environmental changes. The system supports multi-threaded parallel processing and efficient computing to ensure real-time decision-making and rapid response.

[0168] 3) Data collection and analysis: Vehicles and infrastructure nodes collect traffic data through sensors and upload it to the MEC platform for processing in real time to generate traffic flow predictions and optimization suggestions.

[0169] 4) Task allocation and offloading: Deep reinforcement learning-based strategies optimize traffic flow control and task offloading, and intelligently allocate tasks through the MEC platform to ensure optimal resource utilization and minimize latency.

[0170] 5) System Iteration and Performance Improvement: Strategies are regularly optimized and updated based on traffic management effectiveness, resource usage, and system feedback. Continuous reinforcement learning iterations improve the system's responsiveness and adaptability, ensuring the efficient and stable operation of the connected vehicle system in complex and dynamic traffic environments.

[0171] Implementation Application Case 2: Autonomous Driving Vehicle Collaborative Control System

[0172] 1) System Startup and Configuration: Deploy and start the global model network on the autonomous vehicle management platform, and initialize relevant network parameters, such as the critic network's learning rate and replay buffer capacity. Configure communication between the vehicle and road infrastructure, and implement efficient data exchange through edge computing nodes.

[0173] 2) Strategy Interaction and Environmental Adaptation: Autonomous vehicles execute path planning and task offloading operations generated by active inference mechanisms based on real-time environmental data and pre-defined collaborative control strategies. The system utilizes multi-threaded parallel processing to improve computational efficiency in decision-making and support efficient collaboration between vehicles.

[0174] 3) Data collection and analysis: Vehicles and road facilities collect road information through on-board sensors and road equipment, and transmit the data to the MEC platform for real-time processing and decision support.

[0175] 4) Task Allocation and Offloading: A deep reinforcement learning-based strategy is used to optimize the allocation and offloading of vehicle computing tasks, ensuring the timely processing of various tasks while improving the accuracy and stability of collaborative control between vehicles.

[0176] 5) System Iteration and Performance Improvement: Strategies are regularly updated and optimized based on the autonomous driving system's operational performance, inter-vehicle collaboration, and resource efficiency. Through continuous iterations of reinforcement learning, the system achieves rapid adaptation and efficient collaboration in complex and dynamic road environments, improving the overall performance and safety of autonomous vehicles.

[0177] In these two embodiments, an efficient, reliable and energy-saving solution is provided for the vehicle network communication system supporting MEC, which is suitable for different application scenarios, from intelligent traffic flow management to cooperative control of autonomous driving vehicles, demonstrating its wide application potential and technical advantages.

[0178] In order to more clearly demonstrate the positive effects achieved by the embodiments of the present invention during the research process, the advantages of the embodiments of the present invention in research and development compared with the prior art are described below.

[0179] exist Figure 4 In this paper, we simulated the relationship between the average system reward and system bandwidth after algorithm convergence. It can be seen that when the bandwidth is very small, the total reward obtained by all algorithms is negative. As the bandwidth increases, the total reward gradually increases. From the vehicle's perspective, the more bandwidth available, the lower the latency achieved. Therefore, as the available bandwidth increases, the total reward increases, indicating that the algorithm proposed in this paper always performs best. This shows that, given a given bandwidth, system performance can be improved by jointly optimizing resolution, transmit power allocation, and computational resource allocation.

[0180] exist Figure 5 Figure 2 shows the convergence performance of all algorithms using reward as the performance metric under default parameter settings. It is clear from the figure that all three algorithms perform well in terms of convergence. Furthermore, the algorithm proposed in this paper outperforms the other two algorithms in terms of reward interval and convergence speed.

[0181] like Figure 6 As shown, another object of the present invention is to provide an optimization system based on active reasoning in low-altitude communication supporting MEC, including:

[0182] Agent Decision Module: At the beginning of each cycle, this module generates appropriate action decisions based on the current network state. It combines the active inference mechanism and the free energy principle to optimize the policy distribution by selecting the mean of the previous stage that minimizes the free energy.

[0183] Action execution module: This module is responsible for actually executing resource allocation operations and putting decisions into practice.

[0184] Reward Evaluation Module: After an action is executed, this module calculates the immediate reward based on the inverse of the weighted average of latency, energy consumption, and operator cost of all devices in the system, and manages the transition process from the current state to the next state.

[0185] Experience storage module: This module stores the system state, executed actions, rewards obtained, accumulated rewards, and next state after each operation, forming experience tuples for subsequent learning and training.

[0186] Sample sampling module: This module is responsible for randomly extracting samples from the policy distribution and self-sampling from the policy itself for the agent to optimize its decision.

[0187] Global model update module: This module uses the back-propagation algorithm to update the global neural network, and uses a parameter optimization unit to adjust the network parameters using the gradient descent method to achieve optimization.

[0188] At the beginning of each system cycle, the agent's decision-making module generates appropriate action decisions based on the current network state. Using an active inference mechanism, this module analyzes the likelihood of state transitions within the network and, incorporating the free energy principle, calculates the optimal policy distribution. The free energy calculation sorts the metrics of candidate policies and selects the policy with the lowest free energy, thereby generating an optimized policy distribution. The agent then samples actions from this policy distribution, ensuring optimal resource allocation and laying the foundation for improved system performance.

[0189] The action execution module allocates communication resources, computing resources, and power allocation based on the actions generated by the agent, and implements these decisions. The reward evaluation module then calculates the immediate reward based on the results of the action execution. This immediate reward is calculated by taking a weighted average of latency, energy consumption, and operator costs, and the inverse of these averages reflects the degree of improvement in system performance. This module is also responsible for recording the transition between the current state and the next state resulting from the action execution, providing accurate feedback signals to the agent.

[0190] The experience storage module records the state, action, immediate reward, cumulative reward, and the entire process of transitioning to the next state after each action is executed, forming an experience tuple. This stored experience data provides data support for the agent's learning and training. The sample sampling module randomly samples training samples from the policy distribution or self-samples based on the policy, providing input data for further decision optimization. Through efficient sample sampling, this module ensures the diversity of training data and the robustness of system decisions.

[0191] The global model update module optimizes the parameters of the global neural network using a backpropagation algorithm. The module first calculates the error between the predicted and target values ​​and uses gradient descent to adjust the network parameters, ensuring a better fit between the reward function and the free energy distribution. The parameter optimization unit dynamically adjusts the learning rate to ensure rapid model convergence and optimal performance. After multiple iterations, the optimization of global network parameters significantly improves the decision-making efficiency of the intelligent agents and the global optimization of system resources, thereby enabling continuous improvement in MEC system performance in low-altitude communication scenarios.

[0192] Case Application 1: Intelligent Traffic Flow Management Solution

[0193] 1) System Configuration and Initialization: Deploy the global model network in the IoV traffic management center and initialize key parameters, including hyperparameters such as the critic network's learning rate, replay buffer capacity, and time slot length. Configure vehicles and road infrastructure (such as smart traffic lights and onboard sensors) to work together, using the MEC platform to monitor and optimize real-time traffic flow.

[0194] 2) Strategy Execution and Environmental Adaptation: The vehicle uses dynamic path planning and speed adjustment strategies generated by active inference mechanisms, combined with real-time traffic data, to execute precise driving decisions. This system is powered by an efficient computing platform and leverages multi-threaded parallel processing capabilities to improve the speed and accuracy of decision responses.

[0195] 3) Data Collection and Processing: Vehicles use sensors to collect traffic information such as road conditions and speed, and transmit this data in real time to the MEC platform for rapid processing. The platform, combined with advanced traffic flow prediction and optimization models, provides accurate data support for traffic scheduling.

[0196] 4) Task Allocation and Offloading: Using a deep reinforcement learning-based strategy, the system intelligently optimizes the allocation and offloading of computing tasks, ensuring efficient use of network and computing resources. Through continuous reinforcement learning, the system continuously improves task execution efficiency and enhances decision-making accuracy.

[0197] 5) System Updates and Optimization: Strategies are regularly optimized based on traffic flow management effectiveness and resource usage. Through continuous learning and strategy iteration, the system dynamically adapts to varying traffic environments, providing smooth and stable traffic management support, improving overall traffic efficiency, and ensuring the accuracy and responsiveness of intelligent traffic management.

[0198] Implementation Application Case 2: Autonomous Driving Vehicle Collaborative Control System

[0199] 1) System Deployment and Initialization: The global model network is deployed in the control center of the autonomous vehicle, and network parameters such as the critic network's learning rate, replay buffer capacity, and time slot length are initialized. The vehicle works in conjunction with the edge computing platform to implement path planning, task scheduling, and inter-vehicle collaboration through MEC.

[0200] 2) Strategy Execution and Environmental Adaptation: Based on task offloading, path planning, and collaborative driving strategies generated by active inference mechanisms, vehicles dynamically adjust driving decisions based on the surrounding environment and real-time data. Powered by efficient computing, vehicles can respond to changing traffic conditions in real time and collaborate with other vehicles to optimize driving paths.

[0201] 3) Data Collection and Real-Time Processing: Vehicles use sensors to collect surrounding traffic information, including vehicle location, speed, obstacles, and other data, and transmit it to the MEC platform for real-time processing. The platform uses this data to schedule tasks and dynamically plan routes, providing real-time decision support for vehicles.

[0202] 4) Task Allocation and Collaborative Offloading: A deep reinforcement learning-based strategy is used to optimize task allocation and offloading, coordinating the driving paths and resource sharing of multiple autonomous vehicles. Through continuous reinforcement learning iterations, the system improves collaborative efficiency and vehicle safety.

[0203] 5) System Optimization and Performance Improvement: Based on the effectiveness of vehicle collaborative control and resource usage, the system regularly updates and optimizes its strategies to ensure that the intelligent driving system can efficiently handle various complex traffic scenarios, improve the driving experience and system stability, and achieve more efficient collaborative control and autonomous driving.

[0204] In these two embodiments, an efficient, reliable and energy-saving solution is provided for the communication system supporting MEC, which is suitable for different application scenarios, from intelligent traffic flow management to cooperative control of autonomous driving vehicles, demonstrating its wide application potential and technical advantages.

[0205] exist Figure 4 In the simulation, the relationship between the average system reward and system bandwidth after algorithm convergence is simulated. It can be seen that when the bandwidth is very small, the total reward obtained by all algorithms is negative. As the bandwidth increases, the total reward gradually increases. From the vehicle's perspective, the more bandwidth available, the lower the latency achieved. Therefore, as the available bandwidth increases, the total reward increases, indicating that the algorithm proposed in this paper always performs best. This shows that, given a given bandwidth, system performance can be improved by jointly optimizing resolution, transmit power allocation, and computational resource allocation.

[0206] exist Figure 5Figure 2 shows the convergence performance of all algorithms using reward as the performance metric under default parameter settings. It is clear from the figure that all three algorithms perform well in terms of convergence. Furthermore, the algorithm proposed in this paper outperforms the other two algorithms in terms of reward interval and convergence speed.

[0207] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware portion can be implemented using dedicated logic; the software portion can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. Those skilled in the art will appreciate that the above-mentioned devices and methods can be implemented using computer-executable instructions and / or contained in processor control code, for example, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above-mentioned hardware circuits and software, such as firmware.

[0208] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with this technical field within the technical scope disclosed by the present invention and within the spirit and principles of the present invention should be covered by the scope of protection of the present invention.

Claims

1. A synaesthesia-enabled vehicle-mounted edge metaverse optimization method, characterized in that: The following steps are involved: Initialize optimization algorithm parameters, including state variables, transition model parameters, reward model parameters, number of training rounds, maximum number of training steps, number of candidate policies, number of optimal candidate policies, learning rate, discount factor, replay buffer size, policy distribution, transition probability distribution, hyperparameters, and random parameters, to account for the dynamic characteristics of the in-vehicle edge computing environment; Based on the current state, an adaptive policy sampling mechanism is used to generate multiple candidate policies. This mechanism dynamically adjusts the exploration and exploitation trade-off of policy distribution based on the communication quality between vehicles, computing resource availability, task complexity, and service quality requirements to improve search efficiency and optimization performance. Calculate the conditional transition probability distribution and immediate reward for the actions generated by the alternative strategy; A multi-scale free energy calculation model is constructed, combining short-term immediate rewards with long-term cumulative rewards. Based on conditional transition probabilities and transition model outputs, the dynamic weight adjustment mechanism for free energy calculation is optimized. After ranking the free energy indicators of alternative strategies, an adaptive dynamic weighted average method is used to update the strategy distribution, enabling the intelligent agent to quickly adapt to different in-vehicle edge computing scenarios and improving the global optimality of decision-making. Based on the updated policy distribution, the current policy and action are generated in combination with environmental context information, and the next state is obtained through interaction with the environment. The environmental context information includes the load status of the on-board computing nodes, the quality of the communication link, the type of task requirements, etc., to improve the environmental adaptability of the decision-making; Self-supervised contrastive learning is used to optimize the transfer model and reward model. The empirical data is stored in the replay buffer, and the network parameters are dynamically adjusted by combining the backpropagation algorithm and gradient descent method to improve the training efficiency of the model. Repeat the above steps until the algorithm converges, and finally obtain the optimal policy distribution, which is applied to tasks such as resource scheduling, computing offloading, and intelligent service allocation in the vehicle-mounted edge computing environment.

2. The in-vehicle edge metaverse optimization method enabled by synaesthesia and computing as claimed in claim 1, characterized in that: The optimization algorithm parameter initialization includes: Initialize state parameters, including vehicle communication requirements, parking space requirements, vehicle location, and server resource prices; Initialize the policy distribution and transition probability distribution; Initialize hyperparameters, including system bandwidth, channel gain, transmit power, and number of vehicles.

3. The in-vehicle edge metaverse optimization method enabled by synaesthesia and computing as claimed in claim 1, characterized in that: The calculation of the instant reward is based on the weighted sum of the vehicle's quality experience and the total system delay, and specifically includes the following steps: Normalize the vehicle's quality experience and the total system delay separately; By setting weight parameters, the vehicle quality experience and the total system delay are balanced so that they have the same scale; The instant reward is calculated by weighted subtraction of the two indicators according to the weight parameter.

4. The in-vehicle edge metaverse optimization method enabled by synaesthesia and computing as claimed in claim 1, characterized in that: The calculation of the free energy is based on the conditional transition probability distribution and the output of the transition model, and specifically includes the following steps: Determine the conditional transition probability distribution of the current state and action; The free energy is obtained by weighted summing of the logarithms of the conditional transition probability distributions.

5. The in-vehicle edge metaverse optimization method enabled by synaesthesia and computing as claimed in claim 1, characterized in that: The update of the policy distribution is based on the sorting and weighting of free energy, which includes the following steps: Sort the calculated free energies and extract the first several smaller free energy values; Average the extracted free energy values ​​and update the policy distribution based on the average; Adjust the strategy distribution using exponentially related distribution methods.

6. The in-vehicle edge metaverse optimization method enabled by synaesthesia and computing as claimed in claim 1, characterized in that: The update of network parameters is completed by gradient descent of the error function, which includes the following steps: Calculate the error between the predicted value and the target value; Calculate the gradient of network parameters through error; The network parameters are iteratively updated according to the learning rate and the calculated gradient.

7. A synaesthesia-enabled in-vehicle edge metaverse optimization system that implements the synaesthesia-enabled in-vehicle edge metaverse optimization method and system as described in any one of claims 1-6, characterized in that: include: The initialization module is used to initialize various parameters of the optimization algorithm, including state, transition model parameters, reward model parameters, number of rounds, maximum number of training steps, number of alternative strategies, number of optimal candidate strategies, learning rate, discount factor, replay buffer size, strategy distribution, transition probability distribution, hyperparameters, and random parameters; The strategy generation module is used to generate multiple alternative strategies based on the current state by randomly sampling the strategy distribution, and obtain multiple actions by randomly sampling the alternative strategies; The free energy calculation module is used to calculate free energy and related indicators based on the conditional transition probability distribution and the output of the transition model, and then sort the indicators and extract the top several items to average them to update the strategy distribution; The interaction module is used to sample and generate the current strategy and action based on the updated strategy distribution, and interact with the environment to obtain the next state; Parameter update module, used to update the network parameters of the transfer model and reward model through backpropagation algorithm and gradient descent; The training module is used to repeat the above process until the algorithm converges to obtain the optimal policy distribution.

8. The in-vehicle edge metaverse optimization system enabled by synaesthesia and computing as claimed in claim 7, characterized in that: The initialization module includes: A state initialization unit, used to initialize the vehicle's communication requirements, parking space requirements, vehicle location, and server resource prices; Distribution initialization unit, used to initialize policy distribution and transition probability distribution; The parameter initialization unit is used to initialize hyperparameters, including system bandwidth, channel gain, transmit power, and number of vehicles.

9. The in-vehicle edge metaverse optimization system enabled by synaesthesia and computing as claimed in claim 7, characterized in that: The free energy calculation module includes: Conditional transfer calculation unit, used to calculate the conditional transfer probability distribution of the current state and action; a free energy index calculation unit, for processing data based on conditional transition probability distribution to generate a free energy value; A free energy update unit that sorts, extracts, and averages free energy values ​​to update the policy distribution.

10. The in-vehicle edge metaverse optimization system enabled by synaesthesia and computing as claimed in claim 7, characterized in that: The parameter updating module includes: An error calculation unit, used to calculate the error between the predicted value and the target value; A gradient calculation unit, for calculating a gradient value based on the error; The parameter adjustment unit is used to update the network parameters of the transfer model and reward model according to the gradient value and learning rate.