A semantic-driven holographic content adaptive transmission method
By combining the semantic-driven communication framework and the DRL algorithm assisted by the gradient projection method in holographic content transmission, the adaptability and reliability problems of holographic content transmission are solved, efficient resource allocation and task utility improvement are achieved, and the scope of application of the transmitted content is expanded.
Patent Information
- Application Number
- CN202510061506.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-01-15
AI Technical Summary
Existing technologies find it difficult to achieve adaptive semantic-driven holographic content transmission under dynamic wireless channel conditions. Especially during the holographic content transmission process, existing research has failed to effectively combine the advantages of numerical optimization and deep reinforcement learning, and the computational complexity is high, making it difficult to meet the unique challenges of holographic content.
A semantic-driven holographic content adaptive transmission method is adopted. By establishing a communication framework, base station association, semantic compression ratio and power allocation are optimized. A deep reinforcement learning algorithm assisted by the gradient projection method is combined to optimize BS association and power allocation. Convex optimization technology is used to solve non-convex problems. A GPM-assisted DRL algorithm is designed to improve transmission efficiency.
It has achieved the adaptability and reliability of holographic content transmission, improved task utility by up to 78%, expanded the scope of application of transmitted content, and improved the efficiency and accuracy of resource allocation.
Smart Images

Figure CN119967593B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of communication technology and relates to a semantic-driven holographic content adaptive transmission method. Background Art
[0002] As Internet of Things (IoT) deployments continue to scale, thousands of IoT devices will engage in task-oriented communication to enable smart services. To enable more immersive and data-rich interactions in IoT applications, holographic content transmission is emerging as a promising technology. Specifically, IoT devices equipped with depth sensors and cameras capture real-world objects and environments as 3D holographic data. This data is then transmitted to a receiver (e.g., an edge server or cloud platform) for performing various computational tasks such as object classification, scene segmentation, and 3D reconstruction, enabling applications such as augmented reality, remote medical diagnosis, and industrial quality inspection. However, holographic content typically involves large amounts of data, making direct transmission over wireless networks resource-intensive and challenging. Semantic communication has received considerable attention as a potential solution, focusing on transmitting task-related information rather than ensuring accurate bit recovery. This approach is particularly promising for task-oriented holographic content transmission. However, implementing adaptive semantic-driven transmission and resource allocation under dynamic wireless channel conditions presents significant challenges.
[0003] In recent years, research on semantically driven wireless network transmission optimization has been increasing. This includes wireless resource optimization methods for text transmission, improving task performance by maximizing the semantic rate, and schemes that balance communication resources and task performance by introducing successful transmission probability and semantic compression ratio (SCR). However, these numerical optimization-based methods have high computational complexity and limited adaptability in dynamic wireless environments. To this end, researchers have begun exploring deep reinforcement learning (DRL) techniques to achieve adaptive transmission and resource optimization. These include DRL-based resource allocation algorithms for downlink text recovery, semantically driven downlink image transmission latency minimization, and DRL bitrate adaptation algorithms for video segmentation tasks. Despite this, existing research still has two major limitations: first, the modeling approach primarily focuses on text, image, or video transmission, failing to address the unique challenges of holographic content transmission; second, the algorithmic approach primarily relies on numerical optimization or DRL, failing to effectively combine the advantages of both to further improve performance. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a semantic-driven holographic content adaptive transmission method, providing a design idea for the practice of transmitting holographic content.
[0005] In order to achieve the above object, the present invention provides the following technical solutions:
[0006] A semantic-driven holographic content adaptive transmission method comprises the following steps:
[0007] S1: Build a semantic-driven communication framework for transmitting holographic content;
[0008] S2: Establish a mathematical model for the optimization problem of joint base station (BS) association, SCR selection and power allocation;
[0009] S3: Convert the non-convex mixed-integer nonlinear programming (MINLP) problem into a sub-problem that satisfies the minimum probability of successful transmission and determine the feasibility conditions for transmit power allocation;
[0010] S4: Design a gradient projection method (GPM)-assisted DRL algorithm to obtain an effective solution for optimizing variable base station association, semantic compression ratio selection, and power allocation.
[0011] Furthermore, the S1 specifically includes:
[0012] The semantic-driven holographic content transmission framework established includes N base stations, K IoT devices and semantic models. All base stations use Indicates that all IoT devices use The device uses non-orthogonal multiple access (NOMA) to uplink the extracted holographic content semantic information to the corresponding base station, which then performs specific task reasoning. The specific steps are as follows:
[0013] B1. Establish a semantic communication model consisting of a semantic encoder, a channel encoder, a channel decoder, and a semantic decoder. Represents the holographic data generated by the kth IoT device, which is encoded by the semantic encoder Processed In addition, define η k For SCR selection, 0<η k <1 indicates the ratio of extracted semantic information to original information.
[0014] B2.Transmission Model
[0015] αk,n represents the BS association indicator, α k,n =1 means the kth IoT device is connected to the nth BS, otherwise α k,n = 0. Since each IoT device can only be associated with one base station, the BS association indicator should meet the following constraints:
[0016]
[0017] In addition, the large-scale channel gain g between the kth IoT device and the nth base station is k,n Including path loss effect, small-scale fading coefficient h k,n The total channel gain of device k on the nth BS is given by H k,n =g k,n h k,n Without loss of generality, it is assumed that the channel gains of IoT devices on each BS are arranged in ascending order:
[0018] p k Indicates the transmission power of the kth device, that is, the power distribution. Considering the heterogeneity of IoT devices, the maximum transmission power of each device is Therefore, there is a transmission power constraint:
[0019]
[0020] At the receiving BS, successive interference cancellation is used to recover the signal. The base station first decodes the signal from the IoT device with the highest channel power gain and treats the signals from other devices as interference. The decoded signal is then subtracted from the superimposed signal. The Signal-to-Interference-Plus-Noise-Ratio (SINR) of device k can be given by
[0021]
[0022] Among them, W n represents the bandwidth of the nth BS, and N0 represents the noise power spectrum density. represents the bandwidth of device k. Therefore, the transmission rate achievable by device k can be expressed by R k =W k log2(1+γ k )calculate.
[0023] After semantic compression of holographic data, the actual data stream transmitted through the wireless channel still exists in the form of data packets. After basic feature extraction and channel coding, the size of the data packet containing semantic information is represented as d0. The SCR selection η of device kk Only from the pre-trained SCR set Therefore, the following constraints apply:
[0024]
[0025] Therefore, the amount of data transmitted by device k is d k =η k d0, the transmission delay of device k can be expressed as Next, we derive the formula for the probability of successful transmission based on semantics. For the sake of simplicity of notation, let and in represents the set of devices associated with the nth BS, represents an indicator function. In order to ensure the reliability of semantic transmission, the delay t k Less than the threshold t max The probability can be calculated as:
[0026]
[0027] in Assume h k It obeys the Rayleigh distribution with scale parameter δ. When x≥0, its cumulative distribution function (CDF) is given by Therefore, we can get P(t k ≤t max )'s explicit expression:
[0028]
[0029] In order to ensure the reliability of semantic-driven transmission, the minimum success probability requirement P must be met. th ,Right now
[0030]
[0031] Task utility is used to evaluate the performance of task execution. For example, in the holographic classification task, task utility is measured by classification accuracy, while in the holographic content reconstruction task, it is evaluated using Mean Squared Error (MSE). Task utility is a function of SINR and SCR. The task utility of device k can be expressed as β k =β k (γ k ,η k ). It should be noted that the function β k (γ k ,η k) has no closed-form expression and can only be obtained by fitting a pre-trained semantic model. A logistic regression function is used to model the task utility and (γ k ,η k ) can be written as:
[0032] here, and It is a parameter related to SCR.
[0033] Furthermore, the S2 specifically includes the following steps:
[0034] make Indicates BS association, Indicates SCR selection, represents power allocation. The goal is to jointly optimize BS association, SCR selection, and power allocation. Therefore, mathematically, the optimization problem can be expressed as
[0035]
[0036] Among them, C1-C2 give the BS association indicator constraints, C3 is the transmission power constraint, C4 is the SCR selection constraint, and C5 gives the transmission delay constraint to ensure the reliability of semantic transmission.
[0037] Problem 1 is a MINLP problem, whose high computational complexity poses a significant challenge to traditional optimization techniques. In addition, directly applying DRL methods usually fails to produce accurate results.
[0038] Furthermore, the S3 specifically includes the following steps:
[0039] The DRL method derives the optimal policy from the reward signal. A direct way to solve Problem 1 is to model all optimization variables as actions. However, this simplified approach ignores the intrinsic relationship between variables. Therefore, it may be difficult for the agent to learn an effective state-action mapping, thereby increasing the risk of converging to a suboptimal solution or even failing to converge. Based on the structure of Problem 1, two key features are obtained. First, the joint BS association α and SCR selection η are discrete optimization variables, making DRL particularly suitable for processing these discrete actions. Second, once α and η are determined, the original problem is simplified to a power allocation subproblem, which can be efficiently solved using convex optimization techniques. It is worth noting that compared with the exploration process in DRL, numerical optimization methods provide more accurate and efficient solutions. Taking these factors into account, the present invention adopts the DRL method to determine the optimal BS association α and SCR selection η. Convex optimization technology is used to assist DRL in efficiently solving the power allocation subproblem.
[0040] B1. Given the BS association α and SCR selection η, Problem 1 can be transformed into Problem 2, which is the power allocation sub-problem:
[0041]
[0042] Due to the transmit power limitation in C3, the minimum successful transmission probability requirement in C5 may not be feasible, resulting in an inability to solve Problem 1. To solve this problem, we must first determine the feasibility condition. C5 can be equivalently restated as follows:
[0043]
[0044] Therefore, the following lemma is proposed to establish the feasibility conditions for Problem 2.
[0045] Lemma 1 (Feasibility Conditions for Problem 2): Problem 2 is feasible if and only if the following constraints are satisfied:
[0046]
[0047] in Indicates the minimum transmit power requirement of the kth device.
[0048] Due to the complex coupling between transmit power and mission utility, Problem 2 is non-convex. However, by exploiting the continuity of the objective function, GPM provides an efficient method to obtain a solution, namely Algorithm 1, whose specific steps are as follows:
[0049] First, at each iteration x, the objective function ▽f(p k (x)) to determine the direction of ascent; secondly, take a step in this direction, and the step size κ controls the update amplitude, so for IoT device k, calculate Finally, step 4 projects the power value into the above feasible range to calculate p k (x+1) and calculate the task utility β k (x+1); Update x and repeat the above steps until convergence.
[0050]
[0051] B2. The controller deployed on the BS acts as an intelligent agent, leveraging channel information from IoT devices to determine the optimal BS association α and SCR selection η. Algorithm 1 is used to compute the optimal power allocation. The BS then broadcasts these decisions to the devices. To avoid excessive dimensionality in the action space, the scheduling process for all devices is modeled as a sequential decision process. In each round, a decision is made for one device at each time step t. A round consists of at most T time steps, where T = K. At the end of each round, all solutions are determined. The key definitions are as follows:
[0052] 1) State: At time step t, the state includes the channel state information (CSI) of the kth device on all BSs, as well as the potential interference from the previously associated k-1 devices, i.e. in
[0053] 2) Action: The agent outputs the BS association α and the SCR selects η as the action of device k at time step t, expressed as
[0054] 3) Reward: The reward is used to evaluate the quality of the action. To be consistent with the optimization goal, the reward is set to the task utility of device k, i.e.
[0055] Compared with the traditional DRL framework, Proximal Policy Optimization (PPO) shows excellent exploration ability and convergence characteristics. This paper uses PPO as a training framework, and the agent learns the policy function that maps the state to the action. The parameter θ is adjustable, and the critic network uses the parameter φ to estimate the state To evaluate the relative value of taking a particular action in a particular state, the advantage function
[0056]
[0057] where γ∈[0,1] is a discount factor that balances the weight between current rewards and future rewards. To ensure training stability, a clipping mechanism is introduced to limit the amplitude of policy updates. Let ∈ be the clipping parameter. The objective function of the policy is given by:
[0058]
[0059] in, is the ratio between the old and new policies, the clipping function
[0060]
[0061] Furthermore, the S4 specifically includes the following steps:
[0062] B1. In each training set, make decisions for IoT devices in ascending order starting from k=1.
[0063] B2. At each time step t, obtain the state from the environment and generate the action from the actor network Calculate the minimum transmit power of device k And get the next state if The current round is stopped. After a round ends, all actions constitute α and η.
[0064] B3. Algorithm 1 calculates the optimal transmission power p of all devices based on the actions of the intelligent agent k and task utility β k and will reward Set to β k . Data obtained in the current round Stored in the replay buffer for network training.
[0065] B4. Update network parameters, The above process is iterated until the training model converges.
[0066] The beneficial effects of the present invention are that: the method of the present invention takes into account that the transmission content is not limited to text, images and videos but is extended to holographic content, and combines numerical optimization and DRL to formulate resource allocation strategies, which is more realistic; in addition, compared with the representative baseline, the proposed method can improve the task utility by up to 78%.
[0067] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:
[0069] Figure 1 A diagram of the framework model of the semantically driven holographic content transmission system provided by the present invention;
[0070] Figure 2 In the present invention, different delay requirements t max The relationship between the probability of successful transmission and SCR is shown below;
[0071] Figure 3 This is a graph showing the relationship between different numbers of IoT devices associated with a BS and task utilities under Algorithm 1 of the present invention;
[0072] Figure 4 This is a graph showing the relationship between different discount factors and task utility under Algorithm 2 of the present invention;
[0073] Figure 5 The task utility and maximum delay limit t of all solutions in this invention aremax The relationship diagram between
[0074] Figure 6 The mission utility and minimum successful transmission probability P of all schemes in this invention are th The relationship diagram between . DETAILED DESCRIPTION
[0075] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.
[0076] Among them, the accompanying drawings are only for illustrative purposes and represent only schematic diagrams rather than actual pictures, and should not be understood as limiting the present invention. In order to better illustrate the embodiments of the present invention, some parts of the accompanying drawings may be omitted, enlarged or reduced, and do not represent the dimensions of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions may be omitted in the accompanying drawings.
[0077] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "back", etc. indicating directions or positional relationships, they are based on the directions or positional relationships shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, the terms describing the positional relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.
[0078] A semantic-driven holographic content adaptive transmission method comprises the following steps:
[0079] S1: Build a semantic-driven communication framework for transmitting holographic content;
[0080] S2: Establish a mathematical model for the optimization problem of joint BS association, SCR selection and power allocation;
[0081] S3: Convert the MINLP problem into a sub-problem that satisfies the minimum probability of successful transmission and determine the feasibility conditions for transmission power allocation;
[0082] S4: Design a GPM-assisted DRL algorithm to obtain effective solutions for the three optimization variables: base station association, semantic compression ratio selection, and power allocation.
[0083] A semantic-driven holographic content transmission framework in S1 is as follows Figure 1 As shown in the figure, the framework includes N base stations, K IoT devices and semantic models. All base stations are Indicates that all IoT devices use The device uses NOMA to transmit the extracted holographic content semantic information uplink to the corresponding base station, which then performs specific task reasoning. The specific steps are as follows:
[0084] B1. Establish a semantic communication model consisting of a semantic encoder, a channel encoder, a channel decoder, and a semantic decoder. Represents the holographic data generated by the kth IoT device, which is encoded by the semantic encoder Processed In addition, define η k For SCR selection, 0<η k <1 indicates the ratio of extracted semantic information to original information.
[0085] B2.Transmission Model
[0086] α k,n represents the BS association indicator, α k,n =1 means the kth IoT device is connected to the nth BS, otherwise α k,n = 0. Since the present invention defines that each IoT device can only be associated with one base station, the BS association indicator should meet the following constraints:
[0087]
[0088] The present invention considers that the large-scale channel gain g between the kth IoT device and the nth base station is k,n Including path loss effect, small-scale fading coefficient h k,n The total channel gain of device k on the nth BS is given by H k,n =g k,n h k,n Without loss of generality, it is assumed that the channel gains of IoT devices on each BS are arranged in ascending order:
[0089] p k Indicates the transmission power of the kth device, that is, the power distribution. Considering the heterogeneity of IoT devices, the maximum transmission power of each device is Therefore, there is a transmission power constraint:
[0090]
[0091] At the receiving BS, successive interference cancellation is used to recover the signal. The base station first decodes the signal from the IoT device with the highest channel power gain, treating the signals from other devices as interference. The decoded signal is then subtracted from the superimposed signal. The SINR of device k is given by
[0092]
[0093] Among them, W n represents the bandwidth of the nth BS, and N0 represents the noise power spectrum density. represents the bandwidth of device k. Therefore, the transmission rate achievable by device k can be expressed by R k =W k log2(1+γ k )calculate.
[0094] After semantic compression of holographic data, the actual data stream transmitted through the wireless channel still exists in the form of data packets. After basic feature extraction and channel coding, the size of the data packet containing semantic information is represented as d0. The SCR selection η of device k k Only from the pre-trained SCR set Therefore, the following constraints apply:
[0095]
[0096] Therefore, the amount of data transmitted by device k is d k =η k d0, the transmission delay of device k can be expressed as Next, we derive the formula for the probability of successful transmission based on semantics. For the sake of simplicity of notation, let and in represents the set of devices associated with the nth BS, Represents an indicator function. In order to ensure the reliability of semantic transmission, the delay t k Less than the threshold t max The probability can be calculated as:
[0097]
[0098] in Assume h k It obeys the Rayleigh distribution with scale parameter δ. When x≥0, its CDF is given by Therefore, we can get P(t k ≤t max )'s explicit expression:
[0099]
[0100] Under different delay requirements t max The relationship between the success transmission probability and SCR is as follows: Figure 2 As shown, in order to ensure the reliability of semantic-driven transmission, the minimum success probability requirement P needs to be met. th ,Right now
[0101]
[0102] Task utility is used to evaluate the performance of task execution. For example, in the holographic classification task, task utility is measured by classification accuracy, while in the holographic content reconstruction task, MSE is used to evaluate. Task utility is a function of SINR and SCR. The task utility of device k can be expressed as β k =β k (γ k ,η k ). It should be noted that the function β k (γ k ,η k ) has no closed-form expression and can only be obtained by fitting a pre-trained semantic model. This paper uses a logistic regression function to model the task utility and (γ k ,η k ) can be written as:
[0103] here, and It is a parameter related to SCR.
[0104] In S2, a mathematical model for the optimization problem of joint BS association, SCR selection, and power allocation is established under the constraints of BS association symbol and maximum transmission power. The model includes the following steps:
[0105] make Indicates BS association, Indicates SCR selection, The goal of this invention is to jointly optimize BS association, SCR selection and power allocation, so mathematically, the optimization problem can be expressed as
[0106]
[0107] Among them, C1-C2 give the BS association indicator constraints, C3 is the transmission power constraint, C4 is the SCR selection constraint, and C5 gives the transmission delay constraint to ensure the reliability of semantic transmission.
[0108] Problem 1 is a MINLP problem, whose high computational complexity poses a significant challenge to traditional optimization techniques. In addition, directly applying DRL methods usually fails to produce accurate results.
[0109] S3 transforms the MINLP problem into a subproblem satisfying the minimum probability of successful transmission and determines the feasibility conditions for transmit power allocation. First, the joint BS association α and SCR selection η are discrete optimization variables, making DRL particularly well-suited for handling these discrete actions. Second, once α and η are determined, the original problem is simplified to the power allocation subproblem, which can be efficiently solved using convex optimization techniques. It includes the following steps:
[0110] B1. Given the BS association α and SCR selection η, Problem 1 can be transformed into Problem 2, which is the power allocation sub-problem:
[0111]
[0112] Due to the transmit power limitation in C3, the minimum successful transmission probability requirement in C5 may not be feasible, resulting in an inability to solve Problem 1. To solve this problem, we must first determine the feasibility condition. C5 can be equivalently restated as follows:
[0113]
[0114] Therefore, the following lemma is proposed to establish the feasibility conditions for Problem 2.
[0115] Lemma 1 (Feasibility Conditions for Problem 2): Problem 2 is feasible if and only if the following constraints are satisfied:
[0116]
[0117] in Indicates the minimum transmit power requirement of the kth device.
[0118] Due to the complex coupling between transmit power and mission utility, Problem 2 is non-convex. However, by exploiting the continuity of the objective function, GPM provides an efficient way to obtain a solution.
[0119] Algorithm 1 outlines the detailed steps of this method, which are as follows:
[0120] First, at each iteration x, the objective function ▽f(p k (x)) to determine the direction of ascent; secondly, take a step in this direction, and the step size κ controls the update amplitude, so for IoT device k, calculate Finally, step 4 projects the power value into the above feasible range to calculate p k (x+1) and calculate the task utility βk (x+1); Update x and repeat the above steps until convergence.
[0121]
[0122] B2. The controller deployed on the BS acts as an intelligent agent, leveraging channel information from IoT devices to determine the optimal BS association α and SCR selection η. Algorithm 1 is used to compute the optimal power allocation. The BS then broadcasts these decisions to the devices. To avoid excessive dimensionality in the action space, the scheduling process for all devices is modeled as a sequential decision process. In each round, a decision is made for one device at each time step t. A round consists of at most T time steps, where T = K. At the end of each round, all solutions are determined. The key definitions are as follows:
[0123] State: At time step t, the state includes the CSI of the kth device on all BSs, as well as the potential interference from the previously associated k-1 devices, i.e. in
[0124] Action: The agent outputs the BS association α and the SCR selects η as the action of device k at time step t, expressed as
[0125] Reward: The reward is used to evaluate the quality of the action. To be consistent with the optimization goal, the reward is set to the task utility of device k, i.e.
[0126] Compared with the traditional DRL framework, PPO shows superior exploration ability and convergence characteristics. Using PPO as a training framework. The agent learns a policy function that maps states to actions. The parameter θ is adjustable. The critic network estimates the state with the parameter φ To evaluate the relative value of taking a particular action in a particular state, the advantage function
[0127]
[0128] where γ∈[0,1] is a discount factor that balances the weights between current rewards and future rewards. To ensure training stability, a clipping mechanism is introduced to limit the policy update amplitude. Let ∈ be the clipping parameter. The objective function of the policy is given by:
[0129]
[0130] in, is the ratio between the old and new policies, the clipping function
[0131] S4 designed a GPM-assisted DRL algorithm to obtain effective power allocation and training data. The specific steps include the following:
[0132] B1. In each training set, make decisions for IoT devices in ascending order starting from k=1.
[0133] B2. At each time step t, obtain the state from the environment and generate actions from the actor network Calculate the minimum transmit power of device k And get the next state if The current round is stopped. After a round ends, all actions constitute α and η.
[0134] B3. Algorithm 1 calculates the optimal transmission power p of all devices based on the actions of the intelligent agent k and task utility β k and will reward Set to β k . Data obtained in the current round Stored in the replay buffer for network training.
[0135] B4. Update network parameters, The above process is iterated until the training model converges.
[0136] like Figure 3 As shown in Figure 2, the present invention presents the relationship between the number of IoT devices associated with a BS and the task utility based on Algorithm 1. Compared with the initial value, Algorithm 1 improves the task utility by 5%. However, as the number of devices increases, the average task utility decreases due to the increased interference caused by multiple devices sharing the same frequency band.
[0137] like Figure 4 As shown in the figure, the present invention presents the relationship between different discount factors γ and task utility based on Algorithm 2. As can be seen from the figure, the performance of Algorithm 2 starts from a low initial level and gradually converges to a satisfactory accuracy with small fluctuations. In addition, the convergence performance under different γ remains consistent, highlighting the reliability of Algorithm 2.
[0138] like Figure 5 and Figure 6 As shown, the present invention shows the task utility and maximum delay limit t of all schemes max and the minimum successful transmission probability P th relationship. Figure 5 As can be seen from the figure, the proposed scheme (using Algorithm 1) shows excellent performance, with a 78% improvement in task utility compared to the baseline scheme. In contrast, the proposed scheme (without Algorithm 1) The transmission power of each IoT device is allocated in order to illustrate the effectiveness of Algorithm 1. The obvious performance gap between the proposed scheme using Algorithm 1 and that not using Algorithm 1 is mainly due to the significant enhancement brought by the power allocation based on GPM. Figure 5 In, with t max With the increase of , the task utility of all schemes will be improved due to the relaxation of the delay requirement. Figure 6 In the case of P th From the range of 0.95 to 0.99, the proposed scheme always maintains high utility. In contrast, Baseline 1, Baseline 2 and 3 show that the utility increases with P th Because higher P th Requires stricter transmission reliability. Overall, the advantages of the proposed scheme are attributed to the intelligent decision-making through DRL and precise power allocation through GPM.
[0139] The present invention proposes a semantic-driven holographic content adaptive transmission method, which takes into account that the transmitted content is not limited to text, images and videos but is extended to holographic content, and combines numerical optimization and DRL to formulate resource allocation strategies, which is more realistic; in addition, compared with the representative baseline, the proposed method can improve the task utility by up to 78%.
[0140] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.
Claims
1. A semantically driven holographic content adaptive transmission method, characterized by: The following steps are involved: S1: Establish a semantically driven communication framework for transmitting holographic content; specifically, The semantic-driven holographic content transmission framework established includes N BSs, K IoT devices and semantic models. All base stations use Indicates that all IoT devices use The device uses non-orthogonal multiple access (NOMA) to uplink the extracted holographic content semantic information to the corresponding base station, which then performs task reasoning. B1. Establish a semantic communication model consisting of a semantic encoder, a channel encoder, a channel decoder, and a semantic decoder. Represents the holographic data generated by the kth IoT device, which is encoded by the semantic encoder Processed Define η k For SCR selection, 0<η k <1 indicates the ratio of extracted semantic information to original information; B2.Transmission Model α k,n Represents the BS association indicator, the kth IoT device is connected to the nth BS with α k,n =1, otherwise α k,n =0; each IoT device can only be associated with one base station. The BS association indicator meets the following constraints: The large-scale channel gain g between the kth IoT device and the nth BS k,n Including path loss effect, small-scale fading coefficient h k,n Following the Rayleigh distribution, the total channel gain of device k on the nth BS is given by H k,n =g k,n h k,n Given; suppose the channel gains of IoT devices on each BS are arranged in ascending order: p k represents the transmission power of the kth device, that is, the power distribution; considering the heterogeneity of IoT devices, the maximum transmission power of each device is different; there is a transmission power constraint of: At the receiving BS, successive interference cancellation is used to recover the signal. The base station first decodes the signal from the IoT device with the highest channel power gain, treating the signals from other devices as interference. The decoded signal is subtracted from the superimposed signal. The signal-to-interference-plus-noise ratio (SINR) of device k is given by: Among them, W n represents the bandwidth of the nth BS, N0 represents the noise power spectrum density; let represents the bandwidth of device k; the transmission rate achieved by device k is expressed by R k =W k log2(1+γ k )calculate; After semantic compression of holographic data, the actual data stream transmitted through the wireless channel still exists in the form of data packets; after basic feature extraction and channel coding, the size of the data packet containing semantic information is expressed as d0; the SCR selection of device k is η k Only from the pre-trained SCR set The following constraints apply to the selection: The amount of data transmitted by device k is d k =η k d0, the transmission delay of device k is expressed as Derive the formula for the probability of successful transmission based on semantics; let and in represents the set of devices associated with the nth BS, represents an indicator function; to ensure the reliability of semantic transmission, the delay t k Less than the threshold t max The probability of is calculated as: in Assume h k It obeys the Rayleigh distribution with scale parameter δ; when x ≥ 0, its cumulative distribution function CDF is given by Given; get P(t k ≤t max )'s explicit expression: To ensure the reliability of semantic-driven transmission, the minimum success probability requirement P th ,Right now Task utility is used to evaluate the performance of task execution; in the holographic classification task, task utility is measured by classification accuracy; in the holographic content reconstruction task, mean square error (MSE) is used for evaluation; task utility is a function of SINR and SCR, and the task utility of device k is expressed as β k =β k (γ k ,η k ); function β k (γ k ,η k ) has no closed-form expression and can only be obtained by fitting a pre-trained semantic model; a logistic regression function is used to model the task utility and (γ k ,η k ) can be written as: and It is a parameter related to SCR; S2: Establish a mathematical model for the optimization problem of joint base station BS association, semantic compression ratio SCR selection and power allocation; specifically, the following steps are included: make Indicates BS association, Indicates SCR selection, represents power allocation; the goal is to jointly optimize BS association, SCR selection, and power allocation; the optimization problem, i.e., Problem 1, is formulated as: Among them, C1-C2 give the BS association indicator constraints, C3 is the transmission power constraint, C4 is the SCR selection constraint, and C5 gives the transmission delay constraint to ensure the reliability of semantic transmission; S3: Convert the non-convex mixed integer nonlinear programming (MINLP) problem into a sub-problem that satisfies the minimum probability of successful transmission and determine the feasibility condition of transmit power allocation; S4: Design a deep reinforcement learning (DRL) algorithm assisted by the gradient projection method (GPM) to obtain effective solutions for the three optimization variables: base station association, semantic compression ratio selection, and power allocation.
2. The semantically driven holographic content adaptive transmission method according to claim 1, characterized in that: The S3 specifically includes: The DRL method derives the optimal policy from the reward signal. Problem 1 is addressed by modeling all optimization variables as actions. First, the joint BS association α and SCR selection η are discrete optimization variables, making DRL particularly well-suited for handling these discrete actions. Second, once α and η are determined, the original problem is reduced to a power allocation subproblem, which is solved using convex optimization techniques. The DRL method is used to determine the optimal BS association α and SCR selection η. Convex optimization techniques are used to assist DRL in efficiently solving the power allocation subproblem. The specific steps are as follows: B1. Given the BS association α and SCR selection η, transform Problem 1 into Problem 2, which is the power allocation sub-problem: Determine feasibility conditions; C5 is equivalently restated as follows: The following lemma is proposed to establish the feasibility conditions of Problem 2; Lemma 1, feasibility condition for Problem 2: Problem 2 is feasible if and only if the following constraints are satisfied: in Indicates the minimum transmit power requirement of the kth device; Problem 2 is non-convex; B2. The controller deployed on the BS acts as an intelligent agent and uses the channel information from the IoT devices to determine the optimal BS association α and SCR selection η. Algorithm 1 is used to calculate the optimal power allocation. The BS broadcasts the decision to the devices. Algorithm 1 is as follows: First, in each iteration x, the objective function ▽f(p k (x)) to determine the ascending direction: Secondly, taking a step in this direction, the step size κ controls the update amplitude. For IoT device k, we calculate Finally, the power value is projected into the feasible range to calculate p k (x+1) and calculate the task utility β k (x+1); Update x and repeat the above steps until convergence; To avoid excessive dimensionality in the action space, the scheduling process for all devices is modeled as a sequential decision process. In each round, a decision is made for one device at each time step t. A round consists of at most T time steps, where T = K. At the end of each round, a solution is determined. The definition is as follows: 1) State: At time step t, the state includes the channel state information CSI of the kth device on all BSs, as well as the potential interference from the previously associated k-1 devices, i.e. in 2) Action: The agent outputs the BS association α and the SCR selects η as the action of device k at time step t, expressed as 3) Reward: The reward is used to evaluate the quality of the action; to be consistent with the optimization goal, the reward is set to the task utility of device k, that is, Using proximal policy optimization (PPO) as a training framework, the agent learns a policy function that maps states to actions. The parameter θ is adjustable, and the critic network uses the parameter φ to estimate the state To evaluate the relative value of taking a specific action in a specific state, we use the advantage function Where γ∈[0,1] is a discount factor used to balance the weight between current rewards and future rewards. To ensure training stability, a clipping mechanism is introduced to limit the policy update amplitude. Let ∈ be the clipping parameter. The objective function of the policy is given by the following formula: in, is the ratio between the old and new policies, the clipping function 3. The semantically driven holographic content adaptive transmission method according to claim 1, characterized in that: The S4 specifically includes: B1. In each training set, make decisions about IoT devices in ascending order starting from k=1; B2. At each time step t, obtain the state from the environment and generate the action from the actor network Calculate the minimum transmit power of device k And get the next state if The current round is stopped; after a round, all actions constitute α and η; B3. Algorithm 1 calculates the optimal transmission power p of all devices based on the actions of the intelligent agent k and task utility β k and will reward Set to β k ; Data obtained in the current round Stored in the replay buffer for network training; B4. Update network parameters, The above process is iterated until the training model converges.
Citation Information
Patent Citations
Thermal power unit denitration control system optimization method fusing deep reinforcement learning and linear active disturbance rejection control
CN117311162A
Electric power semantic short packet communication method, device and system based on communication and sensing integration
CN118338321A