Semantic perception multi-mode resource allocation method for vehicle networking fleet system

By adopting the semantic-aware multimodal resource allocation method in the C-V2X fleet system and using the MADDPG distributed architecture for dynamic semantic resource allocation, the problems of low resource management efficiency, low multimodal data transmission efficiency and low semantic information transmission success rate in the existing technology are solved, and efficient user experience quality and semantic transmission success rate are achieved.

CN120201393APending Publication Date: 2025-06-24XINKONG (SUQIAN) TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510425190.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing C-V2X fleet system has low resource management efficiency in dynamic environments, low multimodal data transmission efficiency, and low semantic information transmission success rate, making it difficult to meet the needs of efficient communication and real-time decision-making.

Method used

The semantic perceived multimodal resource allocation method for the Internet of Vehicles fleet system is adopted. By establishing a semantic communication model, optimizing resource allocation strategies, and dynamic semantic resource allocation using MADDPG distributed architecture, improving the efficiency of channel allocation, power control and semantic symbol transmission.

Benefits of technology

It significantly improves the quality of user experience (QoE) and semantic transmission success rate (SRS), realizes distributed resource management, adapts to dynamic environments, optimizes multimodal semantic information transmission, solves the resource allocation problem in fleet systems, and is suitable for real-time applications of intelligent transportation systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120201393A_ABST
    Figure CN120201393A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of Internet of Vehicles edge calculation reinforcement learning, and particularly relates to an Internet of Vehicles fleet system-oriented semantic perception multi-modal resource allocation method, which specifically comprises the following steps of: 1, establishing a semantic communication model of a C-V2X fleet multi-modal task scene, 2, defining an optimization target of a semantic data packet transmission success probability SRS of QoE and V2V links, the QoE is decomposed into semantic speed and semantic accuracy so as to adapt to the requirements of different tasks; setting configurable weight parameters according to task requirements, and solving an optimal resource allocation strategy; step 3, optimizing by taking maximization of QoE and SRS of all vehicles as a target function according to the main decision variable; and step 4, establishing an MADDPG distributed architecture, and carrying out dynamic semantic resource allocation. According to the method, the quality of experience (QoE) and the semantic transmission success rate (SRS) are improved, distributed resource management is realized, the method adapts to a dynamic environment, multi-modal semantic information transmission is optimized, and the resource allocation problem in a formation system is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of edge computing reinforcement learning in the vehicle networking, and particularly relates to a semantic-aware multimodal resource allocation method for a vehicle networking platoon system. Background Art

[0002] With the rapid development of the Intelligent Transportation System (ITS), ensuring the safety, efficiency, and reliability of transportation has become increasingly important. As an important part of ITS, the platoon system improves traffic flow and safety by having multiple autonomous vehicles move forward closely in coordination. In the platoon system, usually a platoon leader (PL) guides the movement of the platoon, and other platoon members (PM) ensure the stability of the platoon by maintaining coordinated speeds and inter-vehicle distances.

[0003] In the platoon system, efficient communication within the platoon (intra-platoon) and between the platoon and the infrastructure (inter-platoon) is crucial for improving the overall efficiency and safety of the platoon. To achieve the dual goals of intra-platoon and inter-platoon communication, the integration of Cellular Vehicle-to-Everything (C-V2X) communication is particularly critical. C-V2X communication includes Vehicle-to-Vehicle (V2V) communication, which is used for sharing cooperative perception information among vehicles within the platoon to ensure synchronized movement; and Vehicle-to-Infrastructure (V2I) communication, which is used for information exchange between vehicles and base stations to ensure the sharing of more extensive traffic and safety information. These communication modes enable the platoon system to respond promptly to dynamic traffic situations, thereby improving the safety and operating efficiency of the platoon system.

[0004] However, with the increasing connectivity and communication requirements in the C-V2X platoon system, how to manage network resources has become an important challenge. Efficient resource management is crucial for maintaining communication reliability and reducing latency, especially in real-time decision-making scenarios of autonomous driving. In addition, the frequency of information exchange among platoon members is directly related to the reaction speed of platoon members to potential obstacles, so efficient resource management is required to ensure timely and reliable communication.

[0005] Existing C-V2X platoon systems face multiple technical challenges in resource management:

[0006] Difficulty in resource allocation in a dynamic environment: Traditional centralized resource management methods need to rely on global information for resource allocation. However, due to the rapid and unpredictable changes in network status, the performance of centralized methods in dynamic channel conditions is often unsatisfactory, and it is prone to problems such as large signaling overhead and slow response speed.

[0007] Low efficiency of multi-modal data transmission: As the task complexity increases, the tasks in the platoon system gradually shift from single-mode (such as visual and auditory information) tasks to multi-modal tasks that integrate multiple data types. However, existing research usually fails to combine multi-modal data processing with resource management, resulting in low transmission efficiency in complex environments.

[0008] Low success rate of semantic information transmission (SRS): Traditional communication methods mainly transmit raw data and do not fully utilize the semantic information of the data. In the platoon system, semantic communication can improve the efficiency of information transmission, ensure that more important information is transmitted first, and thus enhance the cooperation ability between vehicles and the overall decision-making efficiency of the system. However, how to effectively improve the success rate of semantic information transmission during the resource allocation process remains an urgent problem to be solved. Summary of the Invention

[0009] To solve the limitations of centralized resource management under dynamic conditions, the present invention has developed a semantic-aware multi-modal resource allocation method for the platoon system of the vehicle-to-everything (V2X) network, aiming to improve the quality of experience (QoE).

[0010] To achieve the above technical objectives, the present invention adopts the following technical means:

[0011] A semantic-aware multi-modal resource allocation method for the platoon system of the vehicle-to-everything (V2X) network, where the platoon system of the vehicle-to-everything (V2X) network includes a base station BS and a platoon composed of several autonomous vehicles, each platoon contains several vehicles, and the first vehicle of the platoon serves as the platoon leader PL, responsible for communicating with the BS and coordinating communication within the platoon. The allocation method specifically includes the following steps:

[0012] Step 1: Establish a semantic communication model for the C-V2X platoon multi-modal task scenario, supporting two communication methods: vehicle-to-vehicle (V2V) communication within the platoon and vehicle-to-infrastructure (V2I) communication between platoons;

[0013] Step 2: Define the QoE and the optimization objective of the success probability SRS of semantic data packet transmission on the V2V link, and decompose the QoE into semantic speed and semantic accuracy to adapt to the requirements of different tasks; set configurable weight parameters according to the task requirements, and solve the optimal resource allocation strategy;

[0014] Step 3: Optimize with the objective of maximizing the QoE and SRS of all vehicles according to the main decision variables including channel allocation binary variables, communication category binary variables, power control, average semantic symbol number, and constraint semantic data packet transmission success probability threshold, exclusivity of sub-channel allocation, semantic similarity, average semantic symbol number, and quality of service constraints;

[0015] Step 4: Establish a MADDPG (Multi-Agent Deep Deterministic Policy Gradient) distributed architecture for dynamic semantic resource allocation; design a two-level critic network architecture where the global critic is responsible for multi-workshop channel allocation and power control, and the local critic optimizes the semantic coding strategy based on individual observations, including channel allocation, power control, semantic symbol length, etc., to improve individual transmission efficiency. Combine the delayed update mechanism to enhance the convergence speed of training.

[0016] Preferably, the specific method of the said Step 1 is as follows:

[0017] (1) Number the vehicles in each convoy as M n , n , where n ∈ 1, 2,..., N, N is the number of convoys, and the number of vehicles in each convoy n is M n ; the set of vehicles within a convoy is defined as M = 1,..., m,..., M; determine the communication mode based on communication requirements and task types; when performing a unimodal task, use the DeepSC (Deep Semantic Conmmunication) model for semantic encoding and decoding of text data; when performing a multimodal task, use the MU-DeepSC (Multi-user Deep Semantic Conmmunication) model to support the transmission of text and image data; among them, unimodal communication is mainly used for communication between the convoy leader PL and the base station BS, mainly transmitting text semantic information; multimodal communication is carried out within the convoy; when the number of vehicles is even, the vehicles within the convoy use the MU-DeepSC model in pairs for multimodal communication, and when the number of vehicles is odd, multimodal communication is used for vehicle pairs (two vehicles as a vehicle pair), and the remaining vehicles perform unimodal communication;

[0018] (2) Unimodal communication of the DeepSC model: When a vehicle conducts unimodal communication within a convoy, the vehicle generates the input text S T , and realizes semantic encoding through a bidirectional long short-term memory network Bi-LSTM to extract semantic features, denoted as Then perform channel encoding to generate symbols X suitable for wireless channel transmission T , where α T and β T are the network parameters of semantic encoding and channel encoding, and T indicates that the transmitted is text; within the convoy, the vehicle sends the symbol X T through the wireless channel, and the receiving vehicle decodes the received symbol Y T,n,m [k] for channel decoding, denoted as and restores the semantic information of the input text through semantic decoding ​ where Y T,n,m [k] represents the symbol received on the k-th subchannel, and α T -1 and β T -1 represent the network parameters for semantic decoding and channel decoding;

[0019] Multi-modal communication of the DeepSC model: The vehicle generates the input image S I , performs semantic encoding through a deep convolutional neural network, and extracts the image semantic features Then, through channel encoding generates the symbol X suitable for wireless channel transmission I ; The image symbol X I and the text symbol X T are sent through the wireless channel. At the receiving end, channel decoding and semantic decoding are respectively performed to obtain the image semantic information and the text semantic information and the two are fused to obtain the complete semantic information;

[0020]

[0021] The transmission symbols X T,n and X I,n generated by the convoy leader PL through channel encoding and semantic encoding are transmitted to the base station BS on the selected subchannel k; At the receiving end, the base station BS performs channel decoding and semantic decoding on the received symbols, calculates the signal-to-interference-plus-noise ratio (SINR) of the signal to evaluate the signal quality; According to the communication requirements within the convoy, determine the types of vehicle-to-vehicle (V2V) communication and vehicle-to-infrastructure (V2I) communication within the convoy;

[0022] Set the binary selection variable. When ρ n,k = 1, select the n-th convoy to perform V2V communication on the k-th subchannel; When ρ n,k = 0, select to perform V2I communication; The encoded text symbol and image symbol are respectively sent to the target vehicle or the base station through V2V or V2I communication; In V2V communication, the vehicles within the convoy exchange information through a direct communication link, and the received signal expression at the receiving vehicle is:

[0023] Y T,n,m [k] = ρ n,k h n,m [k]X T,n [k] + I T,n,m [k] + χ T,n,m , (3)

[0024] Y I,n,m [k] = ρ n,k hn,m [k]X I,n [k]+I I,n,m [k]+χ I,n,m (4)

[0025] Among them, h n,m [k] is the channel gain, representing the channel gain when the m-th vehicle of the n-th convoy communicates on the k-th subchannel. I T,n,m [k] and I I,n,m [k] are interference terms. I T,n,m [k] represents the interference generated when the m-th vehicle of the n-th convoy performs text communication on the k-th subchannel. I I,n,m [k] represents the interference generated when the m-th vehicle of the n-th convoy performs image communication on the k-th subchannel. χ T,n,m and χ I,n,m are noise. X T,n [k] and X I,n [k] represent the text and image transmission symbols generated through channel coding and semantic coding on the k-th subchannel;

[0026] For text and image transmission, the signal-to-noise ratio (SINR) is expressed as:

[0027]

[0028] Among them, p T,n [k] and p I,n [k] are the transmission powers of text and image transmission respectively; for text and image transmission, the interference is expressed as:

[0029]

[0030] β n′,k represents whether the n'-th vehicle selects the k-th subchannel to transmit information. The receiver recovers the transmitted symbols through channel decoding and semantic decoding; for single-modal text transmission, the received signal is decoded by the channel decoder and semantic decoder:

[0031]

[0032] For multi-modal transmission, the received signals of text and image are decoded independently and then fused to infer the final semantic message:

[0033]

[0034] Preferably, the specific method of step 2 is as follows:

[0035] The QoE model evaluates the overall performance of the semantic communication system by combining semantic accuracy and semantic transmission rate;

[0036]

[0037] Score R (φ q ) and Score A (ξ q ) represent the semantic rate φ q and the semantic accuracy ξ q scores, where φ target and ξ target are the target levels of semantic rate and accuracy considered optimal for the safe operation of the vehicle fleet, while γ q and δ μ are the sensitivity parameters that determine the deviation of the scoring function from these target values; where ξ q is used as a performance metric to evaluate the accuracy of semantic transmission, ω q represents the weight for balancing the semantic rate and semantic accuracy, and q represents the q-th vehicle; this metric quantifies the degree of similarity between the received message and the transmitted message in terms of its meaning, and the semantic similarity is expressed as:

[0038]

[0039] where 0 ≤ ξ q ≤ 1, ξ q = 1 indicates the highest similarity between two sentences, while ξ q = 0 indicates no similarity; the term represents the average number of semantic symbols transmitted by vehicle m, and the function Ψ q maps these inputs to a semantic similarity score;

[0040] The semantic rate φ q measures the amount of meaningful information transmitted by a vehicle per second, reflecting the effectiveness of data communication in the vehicle network; different from the traditional data rate that focuses on the transmission of raw data, the semantic rate emphasizes the transmission of information related to the understanding and decision-making processes of specific tasks, and is defined as:

[0041]

[0042] where W represents the bandwidth allocated to the communication channel, represents the semantic entropy of this task.

[0043] Preferably, the specific method of step 3 is as follows:

[0044] (1) Optimize including channel allocation, power allocation, and the selection of the number of semantic symbols for the transmission of text and image data; formally defined as follows:

[0045]

[0046] The objective function of Equation (14) aims to maximize the comprehensive QoE of all vehicles and the semantic transmission success rate SRS within the time budget ΔT. λ is a weight factor used to balance the quality of user services and the importance of timely delivery of safety-critical messages; the term Pr represents the probability of successful timely transmission of semantic data packets through the V2V link, and B s represents the transmitted semantic data packets, while represents the number of data packets of semantic information transmitted within the coherence time ΔT; the first constraint ensures the channel allocation variables, which means that each subchannel is either allocated to a vehicle fleet or not; the second constraint limits that each subchannel can be allocated to at most one vehicle fleet, and the third constraint ensures that each vehicle fleet can be allocated at most one subchannel; the fourth and fifth constraints define the number of semantic symbols for text data and image data transmission and the allowed range of cannot exceed the maximum number of semantic symbols of text and images and The sixth constraint limits the transmission power of each vehicle; the seventh constraint ensures that the semantic rate Score R,m and the accuracy score Score A,m meet the minimum threshold G th , and the last constraint ensures that the total size of semantic information transmitted through the V2V link is within the allowed limit and cannot exceed where, represents the semantic rate of the m-th vehicle on the k-th subchannel at each time period t, and d represents the distance;

[0047] (2) Define the state. For the current time step t, the instantaneous channel information between the n-th vehicle fleet and the base station the instantaneous channel information between the n-th vehicle fleet and its followers the previous interference from other vehicle fleets and as well as the residual information of the in-vehicle communication load Therefore, the state space is represented as:

[0048]

[0049] (3) Define the actions. The vehicle fleet includes various actions, among which is used to select the communication subchannel, is used to select whether to perform vehicle-to-vehicle communication or in-vehicle communication, and represent the communication power level and the sentence length related to vehicle-to-infrastructure (V2I) and vehicle-to-vehicle (V2V) communication, respectively and When the vehicle platoon chooses to perform inter-vehicle communication, only one sentence length needs to be selected; while when choosing to perform in-platoon communication, the sentence lengths of M followers need to be selected; therefore, the action space of agent k at time slot t is defined as:

[0050]

[0051] (4) Define the rewards. The goal of the rewards is twofold. The reward of each agent consists of two aspects: one is the global reward reflecting the cooperation between agents, and the other is the local reward to help each agent explore the optimal action; the local and global rewards of the p-th vehicle platoon at time slot $t$ are respectively:

[0052]

[0053] Preferably, the specific method of step 4 is as follows:

[0054] (1) Implement the SAMRAMARL algorithm (Semantic-Aware Multi-Modal Resource Allocation Method based on Multi-Agent Reinforcement Learning (MARL)): Each agent is equipped with its own actor and local evaluator network for decision-making, while the optimization of the entire system is guided by a pair of global evaluators; Consider a vehicle environment consisting of P vehicle platoons. The policies of all agents are π = π1, π2,..., π N , and the policy π N of the n-th agent, the Q function and the double global evaluator Q function are represented by parameters φ n , ψ1, ψ2 respectively, and g1, g2 represent the abbreviations of two global networks for evaluating the joint state-action pairs of all agents;

[0055] The training process starts from initializing the environment simulator. The simulator is configured with P vehicle platoons to reflect the real multi-agent communication scenario; the double global evaluator network and are initialized. These global evaluators are responsible for evaluating the performance of the entire system according to the global reward, i.e., QoE; at the same time, each agent n initializes its policy network φ n for action selection and initializes a local evaluator network for evaluating the quality of its actions according to local observations;

[0056] Each agent n observes its local state s t at time step t and selects an action according to the policy π N ; the actions of all agents form an action vector The states of all agents form a global state vector According to the global state st and action a t , calculate the global reward This reward is used to measure the performance of the entire system; meanwhile, each agent n calculates a local reward based on its local state and action This reward reflects its specific semantic communication requirements; store the experience of agent at time step t into the shared replay buffer; when the buffer size reaches a preset threshold, randomly sample a mini - batch of transition samples for training;

[0057] Update the global evaluator by minimizing the following loss function and

[0058]

[0059] E represents the mathematical expectation; represents the double - global evaluator network under the global state s and action a;

[0060] where the target value y g is defined as:

[0061]

[0062] γ is the discount factor, π′ is the target policy, and its parameters are θ′ = θ 1′ , θ 2′ ,..., θ N′ ;

[0063] Update the local evaluator of each agent n by minimizing the following loss function

[0064]

[0065] where the target value is defined as:

[0066]

[0067] Combining the feedback from the global and local evaluators, update the policy network of each agent n, and its policy gradient is:

[0068]

[0069] ▽J(θ n ) represents the gradient of the objective function J with respect to the policy parameter θ n of agent n, which is used for policy optimization; E[·] represents the expected value, which means taking the average over the trajectories (state and action distributions) generated by the policy; The policy π of agent n n Selects action a n at state s n with probability (or the output of a deterministic policy) with respect to parameter θ n gradient; Denotes the gradient of the global Q function with respect to agent n's action a n where s = (s1, s2,..., s P ) and a = (a1, a2,..., a P ) are the global state and action vectors. Denotes the gradient of agent n's local Q function with respect to its own action a n gradient;

[0070] Update the target networks of the global evaluator and local evaluator through the following soft update mechanism:

[0071] θ n′ ← τθ n +(1 - τ)θ n′ , ψ n′ ← τψ n +(1 - τ)ψ n′ , (24)

[0072] where τ ∈ (0, 1) is the smoothing factor used to control the update speed of the target network and the main network, θ n′ and ψ n′ represent the policy parameters and network parameters of the dual global evaluator Q function after update.

[0073] Compared with the prior art, the present invention has the following beneficial effects:

[0074] (1) Improve the quality of user experience (QoE) and the success rate of semantic transmission (SRS): By introducing semantic-aware communication, the quality of user experience (QoE) and the success rate of inter-vehicle semantic transmission (SRS) are significantly improved, enhancing the accuracy and reliability of information transmission.

[0075] (2) Achieve distributed resource management and adapt to dynamic environments: By adopting a distributed multi-agent framework, each vehicle can make autonomous decisions and resource allocations, better adapting to complex dynamic network environments and improving the efficiency and flexibility of resource allocation.

[0076] (3) Optimize multi-modal semantic information transmission: Redefine the metrics applicable to semantic and multi-modal data, optimize channel allocation, power allocation, and the transmission length of semantic symbols, thereby improving the transmission efficiency of multi-modal information.

[0077] (4) Solve the resource allocation problem in the formation system: By introducing semantic communication into resource allocation optimization, the challenges that traditional methods are difficult to handle in the formation system are solved, and the overall resource utilization rate is significantly improved.

[0078] (5) Apply to real-time applications in intelligent transportation systems: This method has high robustness and stability, can achieve real-time resource adjustment and semantic optimization in intelligent transportation systems, and provides a reliable solution for future vehicle networking systems. Description of the Drawings

[0079] Figure 1 It is the reward evolution relationship of different algorithms of the present invention's solution.

[0080] Figure 2 It is the relationship between the intra-fleet distance and QoE of the present invention's solution.

[0081] Figure 3 It is the relationship between the intra-fleet distance and the semantic transmission success rate of the present invention's solution.

[0082] Figure 4 It is the relationship between the size of semantic demand and the quality of user experience of the present invention's solution.

[0083] Figure 5 It is the relationship between the size of semantic demand and the semantic transmission success rate of the present invention's solution.

[0084] Figure 6 It is the relationship between the conversion factor and the quality of user experience of the present invention's solution.

[0085] Figure 7 It is the relationship between the conversion factor and the semantic transmission success rate of the present invention's solution. Detailed Implementation Manner

[0086] The following further elaborates in detail the semantic perception resource management method of the vehicle networking communication system based on multi-agent reinforcement learning of the present invention in conjunction with the specification drawings and embodiments. The implementation manners of the present invention include but are not limited to the following embodiments.

[0087] A semantic perception multi-modal resource allocation method for a vehicle networking fleet system of the present invention, the specific steps are as follows:

[0088] Step (1): Number the vehicles in each fleet as M n , where n ∈ 1, 2,..., N, N is the number of fleets, and the number of vehicles in each fleet n is M n; The vehicle set within the convoy is defined as M = 1, …, m, …, M; Two communication methods are supported: vehicle-to-vehicle (V2V) communication within the convoy and vehicle-to-infrastructure (V2I) communication between convoys. Based on communication requirements and task types, the communication mode is determined; When performing single-modal tasks, the DeepSC model is used for semantic encoding and decoding of text data; When performing multi-modal tasks, the MU-DeepSC model is used to support the transmission of text and image data; Among them, single-modal communication is mainly used for communication between the convoy leader PL and the base station BS, mainly transmitting text semantic information; Multi-modal communication is carried out within the convoy. When the number of vehicles is even, the vehicles within the convoy use the MU-DeepSC model in pairs for multi-modal communication. When the number of vehicles is odd, multi-modal communication is used for vehicle pairs, and the remaining vehicles perform single-modal communication.

[0089] Step (2): Single-modal communication of the DeepSC model: When a vehicle performs single-modal communication within the convoy, the vehicle generates the input text S T , and realizes semantic encoding through a Bi-LSTM network, extracting semantic features, denoted as Then perform channel encoding to generate symbols X suitable for wireless channel transmission T , where α T and β T are the network parameters of semantic encoding and channel encoding, and T indicates that the transmitted content is text. Within the convoy, the vehicle sends the symbol X T through the wireless channel, and the receiving vehicle decodes the received symbol Y T,n,m [k] for channel decoding, denoted as and restores the semantic information of the input text through semantic decoding where Y T,n,m [k] represents the symbol received on the k-th sub-channel, and α T -1 and β T -1 represent the network parameters of semantic decoding and channel decoding.

[0090] Multi-modal communication of the DeepSC model: The vehicle generates the input image S I , and performs semantic encoding through a deep convolutional neural network (such as ResNet-101) to extract image semantic features Then through channel encoding generate symbols X suitable for wireless channel transmission I . Transmit the image symbol X I and the text symbol X T through the wireless channel, and the receiving end performs channel decoding and semantic decoding respectively to obtain the image semantic information and the text semantic information And fuse the two to obtain complete semantic information.

[0091]

[0092] The transmission symbol X generated by the convoy leader PL through channel coding and semantic coding T,n and X I,n , is transmitted to the base station BS on the selected subchannel k; at the receiving end, the base station BS performs channel decoding and semantic decoding on the received symbol, calculates the signal-to-interference-plus-noise ratio (SINR) of the signal to evaluate the signal quality. According to the in-convoy communication requirements, determine the types of in-convoy (V2V) communication and between-convoy (V2I) communication;

[0093] Set binary selection variables. When ρ n,k = 1, select the nth convoy to perform V2V communication on the kth subchannel; when ρ n,k = 0 = 0, select to perform V2I communication. Transmit the encoded text symbols and image symbols to the target vehicle or base station through V2V or V2I communication respectively; in V2V communication, the vehicles within the convoy exchange information through a direct communication link, and the received signal expression at the receiving vehicle is:

[0094] Y T,n,m [k] = ρ n,k h n,m [k]X T,n [k] + I T,n,m [k] + χ T,n,m , (3)

[0095] Y I,n,m [k] = ρ n,k h n,m [k]X I,n [k] + I I,n,m [k] + χ I,n,m (4)

[0096] Where h n,m [k] is the channel gain, representing the channel gain when the mth vehicle of the nth convoy communicates on the kth subchannel, I T,n,m [k] and I I,n,m [k] are interference terms, I T,n,m [k] represents the interference generated when the mth vehicle of the nth convoy performs text communication on the kth subchannel, I I,n,m [k] represents the interference generated when the mth vehicle of the nth convoy performs image communication on the kth subchannel, χ T,n,m and χ I,n,m are noise, X T,n [k] and X I,n[k] represents the text and image transmission symbols generated by channel coding and semantic coding on the k-th sub-channel.

[0097] For text and image transmission, the signal-to-noise ratio (SINR) can be expressed as:

[0098]

[0099] where p T,n [k] and p I,n [k] are the transmission powers of text and image transmission respectively. For text and image transmission, the interference can be expressed as:

[0100]

[0101] β n′,k indicates whether the n'-th vehicle selects the k-th sub-channel for information transmission, and the receiver recovers the transmitted symbols through channel decoding and semantic decoding. For single-modal text transmission, the received signal is decoded by the channel decoder and the semantic decoder:

[0102]

[0103] For multi-modal transmission, the received signals of text and image are decoded independently and then fused to infer the final semantic message:

[0104]

[0105] Step (3): In the vehicle platoon network, the effectiveness of resource management is crucial, especially for autonomous driving applications. We propose new metrics based on semantic entropy and quality of experience (QoE) to optimize resource allocation and improve system performance. These metrics aim to go beyond traditional metrics that mainly focus on bit-level accuracy by quantifying the meaningful information transmitted. The QoE model evaluates the overall performance of the semantic communication system by combining semantic accuracy and semantic transmission rate.

[0106]

[0107] Score R (φ q ) and Score A (ξ q ) represent the scores of the semantic rate φ q and the semantic accuracy ξ q , φ target and ξ target are the target levels of semantic rate and accuracy considered to be the best for the safe operation of the vehicle platoon, while γ q and δ μ are the sensitivity parameters that determine the deviation of the scoring function from these target values; where ξ is used as a performance metricq to evaluate the accuracy of semantic transmission, ω q represents the weight for balancing the semantic rate and semantic accuracy, q represents the q-th vehicle; this metric quantifies the similarity degree of the received message in its meaning to the transmitted message. The semantic similarity can be expressed as:

[0108]

[0109] where 0 ≤ ξ q ≤ 1, ξ q = 1 indicates the highest similarity between two sentences, while ξ q = 0 indicates no similarity. The term represents the average number of semantic symbols transmitted by vehicle m, and the function Ψ q maps these inputs to a semantic similarity score. For multimodal communication, it is similar, which is the fusion of text and image data.

[0110] The semantic rate φ q measures the amount of meaningful information transmitted by a vehicle per second, reflecting the effectiveness of data communication in the vehicle network. Different from the traditional data rate that focuses on the transmission of raw data, the semantic rate emphasizes the information transmission related to the understanding and decision-making processes of specific tasks, and its definition is:

[0111]

[0112] where W represents the bandwidth allocated to the communication channel, represents the semantic entropy of this task.

[0113] Step (4): We formulate a semantics-centered resource allocation problem aiming to maximize the quality of experience (QoE) of all vehicles in the vehicle fleet network while improving the successful semantic transmission rate (SRS). The optimization includes channel allocation, power allocation, and the selection of the number of semantic symbols for text and image data transmission. This problem can be formally defined as follows:

[0114]

[0115] The objective function of Equation (14) aims to maximize the comprehensive QoE of all vehicles and the SRS of semantic transmission within the time budget ΔT. λ is a weight factor used to balance the user service quality and the importance of the timely delivery of safety-critical messages; the term Pr represents the probability of the successful timely transmission of semantic data packets through the V2V link, B s represents the transmitted semantic data packets, and represents the number of data packets of semantic information transmitted within the coherence time ΔT; the first constraint ensures the channel allocation variables, which means that each sub-channel is either allocated to a vehicle fleet or not; the second constraint restricts that each sub-channel can be allocated to at most one vehicle fleet, and the third constraint ensures that each vehicle fleet can be allocated at most one sub-channel; the fourth and fifth constraints define the allowable ranges of the number of semantic symbols for text data and image data transmission, which cannot exceed the maximum number of semantic symbols of text and images and the allowable range of, which cannot exceed the maximum number of semantic symbols of text and images and The sixth constraint restricts the transmission power of each vehicle; the seventh constraint ensures the semantic rate Score R,m and the accuracy score Score A,m meet the minimum threshold G th Finally, the constraint ensures that the total size of semantic information transmitted through the V2V link is within the allowable limit and cannot exceed where represents the semantic rate of the m-th vehicle on the k-th sub-channel at each time period t, and d represents the distance

[0116] Step (5): We selected the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm to address the challenges of vehicular networks. This choice is based on the following considerations: Due to the complexity and non-convexity of the optimization problems in dynamic vehicle networks, traditional optimization techniques (such as convex methods) are insufficient. This complexity stems from the rapid changes in vehicle network topologies, interference between links, and changes in communication requirements. These factors make the problem highly dynamic, with non-linear dependencies between variables, resulting in a non-convex objective function and making it impossible to be easily solved by traditional methods. Therefore, DRL is more suitable for this environment because it can adaptively learn appropriate policies through interaction with the environment and is more effective in real-time scenarios. However, centralized DRL methods are impractical due to large communication overheads and lack of scalability. To address these challenges, we adopted the Multi-Agent Deep Deterministic Policy Gradient (MADDPG), which provides a decentralized learning framework. MADDPG allows each vehicle (agent) to independently learn the optimal policy while considering the actions of other vehicles, which is crucial in multi-agent systems such as vehicle fleets. In addition, MADDPG can handle continuous action spaces and is very suitable for complex resource management tasks such as power allocation and channel allocation, which conforms to the dynamic characteristics of V2X communication

[0117] Step (6): Before solving, we first need to simplify the optimization problem by converting the binary channel allocation variable β n,kRelax it into a continuous variable ranging from [0, 1], thus transforming the mixed-integer nonlinear programming (MINLP) problem into a more tractable nonlinear programming (NLP) problem. By approximating the success rate through a logical function, the continuity and differentiability of the calculation are improved, facilitating the use of gradient-based optimization methods. The state space, action space, and reward function are defined to construct the DRL framework. Each agent interacts at time t through state s t , action a t , and reward r t . and transfers to the next state s t+1 .

[0118] Define the state. For the current time step t, the instantaneous channel information between the nth vehicle fleet and the base station The instantaneous channel information between the nth vehicle fleet and its followers The previous interference from other vehicle fleets and And the residual information of the in-vehicle communication load Therefore, the state space can be expressed as:

[0119]

[0120] Step (7): Define the action. The vehicle fleet includes multiple actions, where is used to select the sub-channel for communication, is used to select whether to perform vehicle-to-vehicle communication or in-vehicle communication, and respectively represent the power level of communication, and the sentence lengths related to vehicle-to-infrastructure (V2I) and vehicle-to-vehicle (V2V) communication and When the vehicle fleet selects to perform vehicle-to-vehicle communication, only one sentence length needs to be selected; while when selecting to perform in-vehicle communication, the sentence lengths of M followers need to be selected. Therefore, the action space of agent k at time slot t is defined as:

[0121]

[0122] Step (8): Define the reward. The goal of the reward is twofold: The reward of each agent consists of two aspects: one is the global reward reflecting the cooperation between agents, and the other is the local reward helping each agent explore the optimal action. In deep reinforcement learning (DRL), the reward can be set flexibly, and a good reward mechanism can improve the system performance. Here, our goal is to optimize the quality of experience (QoE) level of agent p, as well as the corresponding payload transmission probability. Therefore, the local and global rewards of the pth vehicle fleet at time slot t are respectively:

[0123]

[0124] Step (9): Perform the SAMRAMARL algorithm. The MADDPG algorithm is adopted to optimize the global system objective and individual agent behaviors in a dynamic semantic communication environment through a dual-evaluator framework, which is particularly suitable for the convoy driving scenario. This algorithm uses a multi-agent actor-critic method, where each agent is equipped with its own actor and local evaluator network for decision-making, while the optimization of the entire system is guided by a pair of global evaluators. We consider a vehicle environment consisting of P convoys, and the policies of all agents are π = π1, π2,..., π N . The policy π of the nth agent N , Q function and the dual global evaluator Q function are represented by the parameters φ n , ψ1, ψ2 respectively, and g1, g2 represent the abbreviations of two global networks used to evaluate the joint state-action pairs of all agents.

[0125] The training process starts with initializing the environment simulator, which is configured with $P$ convoys to reflect the real multi-agent communication scenario. The dual global evaluator networks and are initialized to evaluate the joint state-action pairs of all agents. These global evaluators are responsible for evaluating the performance of the entire system based on the global reward (i.e., QoE). At the same time, each agent n initializes its policy network φ n for action selection and initializes a local evaluator network for evaluating the quality of its actions based on local observations.

[0126] Each agent n observes its local state s t at time step t and selects an action according to the policy π N ; the actions of all agents form an action vector The states of all agents form a global state vector Based on the global state s t and action a t , the global reward is calculated, which is used to measure the performance of the entire system; at the same time, each agent n calculates the local reward and action according to its local state This reward reflects its specific semantic communication requirements. The experience of the agent at time step t is stored in the shared replay buffer; when the buffer size reaches the preset threshold, a small batch of transition samples are randomly drawn from it for training.

[0127] Update the global evaluator by minimizing the following loss function and

[0128]

[0129] E represents the mathematical expectation; represents the dual global evaluator network under the global state s and action a;

[0130] where the target value y g is defined as:

[0131]

[0132] γ is the discount factor, π′ is the target policy, and its parameters are θ′ = θ 1′ , θ 2′ ,..., θ N′ .

[0133] Update the local evaluator of each agent n by minimizing the following loss function

[0134]

[0135] where the target value is defined as:

[0136]

[0137] Combining the feedback of the global and local evaluators, update the policy network of each agent n, and its policy gradient is:

[0138]

[0139] ▽J(θ n ) represents the gradient of the objective function J with respect to the policy parameter θ n of agent n, which is used for policy optimization; E[·] represents the expected value, which means taking the average over the trajectories (state and action distributions) generated by the policy; The policy π n of agent n selects the action a n under the state s n with probability (or the output of the deterministic policy) with respect to the parameter θ n ; represents the gradient of the global Q function with respect to the action a n of agent n, where s = (s1, s2,..., s P ) and a = (a1, a2,..., a P ) are the global state and action vectors. Denote the gradient of the local Q - function of agent n with respect to its own action a n 。

[0140] Update the target networks of the global evaluator and the local evaluator through the following soft - update mechanism:

[0141] θ n′ ←τθ n +(1 - τ)θ n′ ,ψ n′ ←τψ n +(1 - τ)ψ n′ ,(24)

[0142] where τ∈(0,1) is a smoothing factor used to control the update speed of the target network and the main network, θ n′ and ψ n′ represent the policy parameters after update and the network parameters of the double - global - evaluator Q - function.

[0143] Experimental comparison:

[0144] Figure 1 shows the evolution process of different algorithms in terms of the reward value, including algorithms such as SAMRAMARL, DDPG, and TD3 of the present invention. The comparison results show that with the support of the multi - agent structure, SAMRAMARL can converge quickly and exhibit higher stability in the dynamically changing vehicle - to - everything (V2X) environment. Compared with DDPG and TD3, SAMRAMARL can better adapt to environmental changes in complex scenarios and achieve better resource allocation, enabling each vehicle to utilize communication resources more effectively.

[0145] Figure 2 shows the impact of the intra - platoon distance (i.e., the physical distance between vehicles) on the quality of experience (QoE). As the intra - platoon distance increases, the QoE of traditional algorithms such as DDPG and TD3 gradually decreases because the increased path loss due to the increased distance reduces the communication quality. However, the SAMRAMARL algorithm of the present invention can dynamically adjust the semantic symbol length and power allocation when the distance increases, thus improving the QoE initially and maintaining the stability of the QoE after reaching a certain distance. This shows that SAMRAMARL can maintain good communication efficiency and stable user experience at different distances.

[0146] Figure 3Shows the relationship between the in-vehicle distance and the semantic transmission success rate (SRS). When the in-vehicle distance increases, the SRS of algorithms without semantic awareness (such as DDPG_NO_SC) shows an obvious downward trend, showing strong signal attenuation and low-efficiency transmission effects. While algorithms with semantic awareness (such as DDPG and TD3) can better adapt to distance changes and maintain a certain SRS. The SAMRAMARL of the present invention is further optimized on this basis, using multi-agent cooperation to dynamically adjust the transmission strategy, maintaining a high transmission success rate within a larger distance range, thus significantly improving the reliability of communication.

[0147] Figure 4 Shows the impact of the size of semantic requirements on the quality of experience (QoE) of users. As the size of semantic data packets increases, the QoE of all algorithms increases accordingly, reflecting that the transmission of richer semantic information helps to improve the user experience. However, due to its failure to effectively process semantic information, the QoE of DDPG_NO_SC is always low. After the semantic requirements reach a certain level, the QoE of DDPG and TD3 shows a bottleneck, while SAMRAMARL can still maintain a high QoE under this requirement, demonstrating its strong processing ability and high efficiency in resource allocation under high semantic loads.

[0148] Figure 5 Shows the impact of the size of semantic requirements on the semantic transmission success rate (SRS). When the semantic demand increases, the SRS of the DDPG_NO_SC algorithm without considering semantic awareness shows a linear decrease, indicating that it is prone to resource congestion when processing a large amount of semantic information. While other algorithms achieve the stability of SRS by using semantic information and can still maintain a high transmission success rate when the demand increases. Especially the SAMRAMARL of the present invention, through a multi-agent reinforcement learning framework, can adjust resource allocation according to real-time feedback, improve transmission efficiency, and make the SRS still maintain an advantage under high semantic demands.

[0149] Figure 6 Shows the impact of the conversion factor on the quality of experience (QoE) of users. For the $DDPG_NO_SC$ algorithm, as the conversion factor increases, its QoE gradually decreases, which is related to the fact that the conversion factor is used as the denominator in traditional QoE calculations, resulting in a decrease in QoE when the complexity increases. However, algorithms with semantic awareness evaluate performance through the semantic rate, weakening the impact of conversion factor changes on QoE. Especially the SAMRAMARL of the present invention can maintain the stability of QoE under different conversion factor conditions by adjusting the semantics of the transmitted content, thus ensuring the effectiveness of communication.

[0150] Figure 7Shows the relationship between the conversion factor and the semantic transmission success rate (SRS). The $DDPG\_NO\_SC$ algorithm without semantic awareness shows stable SRS values as the conversion factor increases because it does not utilize semantic information for optimization. Semantic awareness algorithms (such as DDPG and TD3), on the other hand, can achieve an increase in SRS by preferentially transmitting more semantically valuable data as the conversion factor increases. The SAMRAMARL of the present invention is particularly outstanding in this regard, being able to effectively coordinate resource allocation and adapt to different network conditions, thus maintaining a high transmission success rate over a wider range of conversion factor changes, demonstrating its strong adaptability and optimization effect in dynamic environments.

[0151] The above embodiments are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and equivalent replacements can be made. These technical solutions obtained by improving and equivalently replacing the claims of the present invention all fall within the protection scope of the present invention.

Claims

1. A semantically aware multimodal resource allocation method for an Internet of Vehicles fleet system. The Internet of Vehicles fleet system includes a base station BS and a fleet of several autonomous vehicles. Each fleet contains several vehicles. The first vehicle in the fleet serves as the fleet leader PL, which is responsible for communicating with the BS and coordinating communications within the fleet. The method is characterized by: The allocation method specifically includes the following steps: Step 1: Establish a semantic communication model for the C-V2X fleet multimodal mission scenario, supporting two communication modes: intra-fleet communication V2V and inter-fleet communication V2I; Step 2: Define the optimization objectives of QoE and the semantic packet transmission success probability SRS of the V2V link, and decompose QoE into semantic speed and semantic accuracy to adapt to the needs of different tasks; set configurable weight parameters according to task requirements and solve the optimal resource allocation strategy; Step 3: According to the main decision variables including channel allocation binary variable, communication category binary variable, power control, average number of semantic symbols and constraint semantic data packet transmission success probability threshold, sub-channel allocation exclusivity, semantic similarity, average number of semantic symbols and service quality constraint, the optimization is performed with maximizing the QoE and SRS of all vehicles as the objective function; Step 4: Establish a MADDPG distributed architecture for dynamic semantic resource allocation; design a two-level critic network architecture, in which the global critic is responsible for multi-workshop channel allocation and power control, and the local critic optimizes the semantic coding strategy based on individual observations, including channel allocation, power control, semantic symbol length, etc., to improve individual transmission efficiency, and combine the delayed update mechanism to improve the convergence speed of training.

2. The semantically aware multimodal resource allocation method for a vehicle network fleet system according to claim 1, characterized in that: The specific method of step 1 is as follows: (1) Number the vehicles in each fleet as M n , where n∈1,2,...,N, N is the number of fleets, and the number of vehicles in each fleet n is M n ; The set of vehicles in the convoy is defined as M = 1, ..., m, ..., M; The communication mode is determined based on the communication requirements and task type; When performing unimodal tasks, the DeepSC model is used for semantic encoding and decoding of text data; When performing multimodal tasks, the MU-DeepSC model is used to support the transmission of text and image data; Among them, unimodal communication is mainly used for communication between the convoy leader PL and the base station BS, mainly transmitting text semantic information; multimodal communication is carried out within the convoy; When the number of vehicles is even, the vehicles in the convoy use the MU-DeepSC model in pairs for multimodal communication, and when the number of vehicles is odd, multimodal communication is used for vehicle pairs, and the remaining vehicles perform unimodal communication; (2) Unimodal communication of DeepSC model: When vehicles perform unimodal communication within a fleet, the vehicle generates input text S T , semantic encoding is realized through the bidirectional long short-term memory network Bi-LSTM, and semantic features are extracted, which is expressed as Then channel coding is performed Generate symbol X suitable for wireless channel transmission T , where α T and β T is the network parameter for semantic coding and channel coding, T indicates that the transmission is text; in the convoy, the vehicle sends the symbol X through the wireless channel T , the receiving vehicle receives the symbol Y T,n,m [k] performs channel decoding, expressed as And through semantic decoding Restore the semantic information of the input text where Y T,n,m [k] represents the symbol received on the kth subchannel, α T -1 and β T -1 Represents the network parameters for semantic decoding and channel decoding; Multimodal communication of DeepSC model: vehicle generation input image S I , semantic encoding is performed through deep convolutional neural networks to extract image semantic features Then, channel coding Generate symbol X suitable for wireless channel transmission I ; Change the image symbol X I and text symbol X T The image is sent through a wireless channel, and the receiving end performs channel decoding and semantic decoding to obtain the image semantic information. and text semantic information And fuse the two to obtain complete semantic information; The transmission symbol X generated by the team leader PL through channel coding and semantic coding T,n and X I,n , and transmit it to the base station BS on the selected subchannel k; at the receiving end, the base station BS performs channel decoding and semantic decoding on the received symbols, calculates the signal interference and noise ratio of the signal to evaluate the signal quality; according to the communication requirements within the fleet, determines the type of internal communication V2V and inter-fleet communication V2I; Set the binary choice variable when ρ n,k = 1, the nth convoy is selected to perform V2V communication on the kth subchannel; when ρ n,k =0, V2I communication is selected; the encoded text symbols and image symbols are sent to the target vehicle or base station through V2V or V2I communication respectively; in V2V communication, vehicles in the fleet exchange information through a direct communication link, and the signal expression received by the receiving vehicle is: Y T,n,m [k]=ρ n,k h n,m [k]X T,n [k]+I T,n,m [k]+x T,n,m , (3) Y I,n,m [k]=ρ n,k h n,m [k]X I,n [k]+I I,n,m [k]+x I,n,m (4) Among them, h n,m [k] is the channel gain, which means the channel gain when the mth vehicle in the nth convoy communicates on the kth subchannel. T,n,m [k] and I I,n,m [k] is the interference term, I T,n,m [k] represents the interference generated when the mth vehicle in the nth convoy communicates on the kth subchannel, I I,n,m [k] represents the interference generated when the mth vehicle in the nth convoy performs image communication on the kth subchannel, T,n,m and χ I,n,m is the noise, X T,n [k] and X I,n [k] represents the text and image transmission symbols generated by channel coding and semantic coding on the kth subchannel; For text and image transmission, the signal-to-noise ratio is expressed as: where p T,n [k] and p I,n [k] is the transmission power of text and image transmission respectively; for text and image transmission, the interference is expressed as: β n′,k Indicates whether the n'th vehicle selects the kth subchannel for information transmission. The receiver recovers the transmitted symbols through channel decoding and semantic decoding. For single-mode text transmission, the received signal is decoded by the channel decoder and semantic decoder: For multimodal transmission, the received signals of text and image are decoded independently and fused to infer the final semantic message:

3. The semantically aware multimodal resource allocation method for a vehicle network fleet system according to claim 2, characterized in that: The specific method of step 2 is as follows: The QoE model evaluates the overall performance of the semantic communication system by combining semantic accuracy and semantic transmission rate; Score R (φ q ) and Score A (ξ q ) represents the semantic rate φ q and semantic accuracyξ q The fraction, φ target and target is the semantic rate and accuracy target level considered optimal for safe fleet operation, and γ q and δ μ is a parameter that determines the sensitivity of the scoring function to deviations from these target values; As the performance indicator ξ q To evaluate the accuracy of semantic transfer, ω q represents the weight of the balance between semantic rate and semantic accuracy, and q represents the qth vehicle. This indicator quantifies the degree of similarity between the received message and the transmitted message in terms of its meaning. The semantic similarity is expressed as: Where 0≤ξ q ≤1,ξ q =1 means the similarity between the two sentences is the highest, and ξ q = 0 means no similarity; term represents the average number of semantic symbols transmitted by vehicle m, function Ψ q Map these inputs to semantic similarity scores; Semantic rate φ q The semantic rate measures the amount of meaningful information transmitted per second by vehicles, reflecting the effectiveness of data communication in vehicle networks. Unlike traditional data rates that focus on raw data transmission, the semantic rate emphasizes the transmission of information related to the understanding and decision-making process of specific tasks, which is defined as: Where W represents the bandwidth allocated to the communication channel, Represents the semantic entropy of this task.

4. The semantically aware multimodal resource allocation method for a vehicle network fleet system according to claim 3, characterized in that: The specific method of step 3 is as follows: (1) Optimization includes channel allocation, power allocation, and the selection of the number of transmission semantic symbols for text and image data; formally defined as follows: The objective function of formula (14) aims to maximize the comprehensive QoE of all vehicles and the semantic transmission success rate SRS within the time budget ΔT. λ is a weight factor used to balance the user service quality The importance of timely delivery of safety-critical messages; the term Pr represents the probability of timely successful transmission of semantic data packets over the V2V link, B s represents the semantic data packet transmitted, while The first constraint ensures that the channel allocation is variable, which means that each subchannel is either allocated to a fleet or not. The second constraint restricts each subchannel to be allocated to at most one fleet, while the third constraint ensures that each fleet can be allocated to at most one subchannel. The fourth and fifth constraints define the number of semantic symbols for text data and image data transmission. and The allowed range cannot exceed the maximum number of semantic symbols for text and images. and The sixth constraint limits the transmission power of each vehicle; the seventh constraint ensures the semantic rate Score R,m and accuracy score Score A,m Meet the minimum threshold G th ,The last constraint ensures that the total semantic information size transmitted through the V2V link is within the allowed limit and cannot exceed in, represents the semantic rate of the mth vehicle on the kth subchannel in each time period t, and d represents the distance; (2) Define the state. For the current time step t, the instantaneous channel information between the nth convoy and the base station is The instantaneous channel information between the nth convoy and its followers Previous interference from other teams and and residual information about the fleet's internal communication load Therefore, the state space is represented as: (3) Definition of actions. The convoy includes multiple actions, including Used to select the subchannel for communication, Used to select whether to conduct inter-vehicle communication or intra-fleet communication. and The power level of the communication and the length of the sentences related to vehicle-to-infrastructure (V2I) and vehicle-to-vehicle (V2V) communications, respectively. and When the fleet chooses to communicate with each other, it only needs to choose a sentence length; when it chooses to communicate within the fleet, it needs to choose the sentence length of M followers; therefore, the action space of agent k at time slot t is defined as: (4) Define rewards. The goal of rewards is twofold. The reward for each agent includes two aspects: one is a global reward that reflects the cooperation between agents, and the other is a local reward that helps each agent explore the optimal action. The local and global rewards of the p-th team at time slot $t$ are:

5. The semantically aware multimodal resource allocation method for a vehicle network fleet system according to claim 4, characterized in that: The specific method of step 4 is as follows: (1) Perform the SAMRAMARL algorithm: Each agent is equipped with its own actor and local evaluator network for decision making, while a pair of global evaluators guides the optimization of the entire system. Consider an environment consisting of a fleet of P vehicles, and the strategies of all agents are π = π1, π2, ..., π N , the strategy π of the nth agent N , Q function and the dual global evaluator Q function The parameters φ n , ψ1, ψ2 denote, g1, g2 denote the abbreviations of two global networks used to evaluate the joint state-action pairs of all agents; The training process starts with initializing the environment simulator, which is configured with P fleets to reflect the real multi-agent communication scenario; the dual global evaluator network and are initialized, these global evaluators are responsible for evaluating the performance of the entire system based on the global reward, i.e., QoE; at the same time, each agent n initializes its policy network φ n Used to select actions and initialize a local evaluator network Used to evaluate the quality of its actions based on local observations; Each agent n observes its local state s at time step t t , and according to the strategy π N Select actions; all the agents’ actions form an action vector The states of all agents constitute the global state vector According to the global state s t and action a t , calculate the global reward This reward is used to measure the performance of the entire system; at the same time, each agent n is rewarded according to its local state. and actions Calculating local rewards This reward reflects its specific semantic communication needs; the agent's experience at time step t Store them in a shared replay buffer; when the buffer size reaches a preset threshold, randomly extract a small batch of transfer samples for training; The global evaluator is updated by minimizing the following loss function and E represents mathematical expectation; Represents a dual global evaluator network under the global state s and action a; Among them, the target value y g Defined as: γ is the discount factor, π′ is the target strategy, and its parameter is θ′=θ 1′ ,θ 2′ ,...,θ N′ ; The local evaluator of each agent n is updated by minimizing the following loss function Among them, the target value Defined as: Combining the feedback from the global and local evaluators, the policy network of each agent n is updated, and its policy gradient is: represents the policy parameter θ of the objective function J for agent n n The gradient of is used for policy optimization; E[·] represents the expected value, which means averaging the trajectories (state and action distributions) generated by the policy; The strategy π of agent n n In status n Next select action a n The probability of (or the output of a deterministic policy) is about the parameter θ n The gradient of Represents the global Q function for agent n action a n The gradient of s = (s1, s2, ..., s P ) and a=(a1,a2,...,a P ) are the global state and action vector, Represents the local Q function of agent n for its own action a n The gradient of The target network of the global evaluator and the local evaluator is updated through the following soft update mechanism: i n′ ←tth n +(1-τ)θ n′ ,ψ n′ ←ts n +(1-t)ψ n′ , (24) Where τ∈(0,1) is a smoothing factor used to control the update speed of the target network and the main network, θ n′ and ψ n′ Represents the updated policy parameters and network parameters of the dual global evaluator Q function.

Citation Information

Cited By

  • State synchronization method and device for agent cluster and computer equipment

    CN121078060A