Semantic communication aided multi-beam satellite anti-eavesdropping secure transmission method and system
Patent Information
- Application Number
- CN202611264456.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-20
- Publication Date
- 2026-09-22
AI Technical Summary
[0005]为了解决现有技术中卫星通信因星地信道增益差异小导致物理层安全保密速率难以保障,以及语义通信在知识库可能被窃取后安全优势丧失、缺乏与传统通信协同资源管理方案的问题,本发明提供了一种语义通信辅助的多波束卫星抗窃听安全传输方法及系统,通过构建包含用户和窃听者的多波束卫星通信系统模型并计算信道增益,为差异化安全防护提供基础;通过以系统等效保密速率与用户公平性为联合优化目标并采用深度强化学习算法进行子信道和功率资源的联合分配,实现资源利用与安全性能的协同优化;通过根据各链路信干噪比与预设阈值的比较结果为各用户分配语义通信范式或传统比特通信范式,利用两类通信范式在不同信干噪比区间的性能互补特性弥补低信干噪比下物理层安全的短板;通过将分配结果转化为星载装备可执行的指令格式并下发至各执行单元完成数据传输,构建了具备工程闭环的抗窃听安全传输体系
在安全性能方面,本发明通过构建语义通信与传统比特通信的共信道异构传输架构,并引入基于信干噪比门限的自适应范式切换机制,即使在窃听者已获取完整语义知识库的极端条件下,仍能利用语义通信在低信干噪比区域与传统通信在高信干噪比区域的性能互补特性,弥补传统物理层安全在合法信道质量较差时的性能短板,有效保障安全用户的正保密速率。
Smart Images

Figure CN122802023A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multi-beam satellite communication technology, and in particular to a semantic communication-assisted multi-beam satellite anti-eavesdropping secure transmission method and system. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] Satellite communication, as a crucial component of the 6G integrated space-air-ground network, aims to expand the coverage of existing terrestrial networks and provide communication services to areas with weak infrastructure or complex geographical conditions. It has become a core infrastructure of the national space information network. With the development of satellite communication, corresponding eavesdropping technologies are also constantly evolving. The inherent characteristics of satellite communication—wide-area coverage and line-of-sight transmission—exacerbate the risks of eavesdropping. Physical layer security technology utilizes the physical characteristics of wireless channels to achieve anti-eavesdropping transmission and is one of the important technical approaches to improve the security of satellite communication. However, compared to terrestrial networks, satellite transmission distances are much longer, and the distances from different users to the satellite are relatively small. Therefore, similar path losses result in small differences in channel gain between users, making it difficult to meet the fundamental prerequisite of physical layer security—that the quality of the legitimate channel must be superior to that of the eavesdropping channel—directly limiting the effectiveness of the confidentiality rate.
[0004] Semantic communication, as an emerging 6G technology, leverages the asymmetry between legitimate knowledge bases and eavesdropping decoding methods, making it difficult for eavesdropping terminals to reconstruct the initial information. Since semantic communication codecs primarily rely on local knowledge bases rather than instantaneous channel state information, they can reduce the need for inter-channel differences. However, semantic communication faces unique challenges in security applications: the knowledge base is at risk of being intercepted by third parties during distribution. Once an eavesdropper obtains the same semantic knowledge base as a legitimate user, the inherent security advantages of semantic communication will be significantly weakened, and in extreme cases, even completely lost. How to achieve reliable and secure communication using the transmission characteristics of semantic communication even with the potential leakage of the knowledge base is a critical problem that urgently needs to be solved in current technological development. Furthermore, the complementary performance relationship between semantic communication and traditional communication under different channel conditions has not yet been systematically applied to resource management and paradigm scheduling for satellite anti-eavesdropping transmissions, lacking a dynamic configuration mechanism for security performance optimization. Summary of the Invention
[0005] To address the challenges in existing satellite communication technologies, such as the difficulty in guaranteeing physical layer security rates due to the small difference in channel gain between satellite and ground, the loss of security advantages of semantic communication after knowledge base theft, and the lack of resource management solutions in conjunction with traditional communication, this invention provides a semantic communication-assisted multi-beam satellite anti-eavesdropping security transmission method and system. By constructing a multi-beam satellite communication system model incorporating both users and eavesdroppers and calculating channel gain, a foundation for differentiated security protection is provided. By using the system's equivalent security rate and user fairness as joint optimization objectives and employing deep reinforcement learning algorithms for joint allocation of sub-channels and power resources, synergistic optimization of resource utilization and security performance is achieved. By assigning semantic communication paradigms or traditional bit communication paradigms to each user based on the comparison results of the signal-to-interference-plus-noise ratio (SNR) of each link with a preset threshold, the complementary performance characteristics of the two communication paradigms in different SNR ranges are utilized to compensate for the shortcomings of physical layer security at low SNR. Finally, by converting the allocation results into an executable instruction format for onboard equipment and distributing it to each execution unit to complete data transmission, an anti-eavesdropping security transmission system with an engineering closed loop is constructed.
[0006] On the one hand, a semantic communication-assisted multi-beam satellite anti-eavesdropping secure transmission method is provided, including: A multi-beam satellite communication system model is constructed, and the channel gain between the user and the satellite is calculated based on the multi-beam satellite communication system model. The multi-beam satellite communication system model includes users and eavesdroppers, and the users include secure users and ordinary users. With the system's equivalent security rate and user fairness as joint optimization objectives, a deep reinforcement learning algorithm is used to jointly allocate the sub-channel resources and power resources in the multi-beam satellite communication system model, and output the sub-channel allocation results and power allocation results for each user. Based on the sub-channel allocation results and power allocation results, the signal-to-interference-plus-noise ratio (SINR) of each user's corresponding link is calculated, and a communication paradigm is assigned to each user based on the comparison result of the SINR with a preset threshold. The sub-channel allocation results, power allocation results, and communication paradigm allocation results are converted into an instruction format executable by the spaceborne equipment and sent to each execution unit to complete downlink data transmission.
[0007] Furthermore, before constructing the multi-beam satellite communication system model, it also includes: constructing a beam bandwidth occupancy model based on tri-color multiplexing technology, dividing the total bandwidth into three parts on average, each part further divided into a finite number of sub-channels, with each beam occupying a portion of the three parts of bandwidth, and no mutual interference between adjacent beams.
[0008] Furthermore, the secure user is a user who needs to combat eavesdroppers intercepting downlink information. The eavesdropper is used to intercept downlink information sent by the satellite to the secure user, and the eavesdropper has the same semantic knowledge base as the secure user and has semantic decoding capabilities.
[0009] Furthermore, the system equivalent security rate is the lower bound of the system total equivalent security semantic rate, and the user fairness is the system total exponentially accumulated equivalent semantic rate.
[0010] Furthermore, the deep reinforcement learning algorithm employs the SAC algorithm, which includes a decision network and an evaluation network. The input of the decision network is the system state, and the output is the action policy. The system state includes the accumulated equivalent semantic rate of all ordinary users at the previous time step, the global channel gain, the eavesdropper's location information, and the tri-color multiplexing grouping of the beam. The action policy includes global power allocation variables and global sub-channel allocation variables. The evaluation network is used to evaluate the expected reward of the state-action pair and uses a dual-Q network architecture to suppress overestimation of the reward. The SAC algorithm only outputs the sub-channel allocation results and power allocation results.
[0011] Furthermore, the reward function of the SAC algorithm includes: a first reward component, representing the lower bound of the total equivalent confidentiality semantic rate of secure users; a second reward component, representing the total exponential accumulation equivalent semantic rate of the system; and a third reward component, used to penalize state-action pairs that violate the maximum satellite transmit power constraint and the sub-channel exclusive constraint; the reward function is a weighted sum of the first reward component, the second reward component, and the third reward component.
[0012] Furthermore, the step of assigning a communication paradigm to each user based on the comparison result of the signal-to-interference-plus-noise ratio (SIR) and a preset threshold specifically includes: for any user, when the SIR of the link corresponding to that user is greater than the preset threshold, the traditional bit communication paradigm is adopted; when the SIR of the link corresponding to that user is less than or equal to the preset threshold, the semantic communication paradigm is adopted; the preset threshold is a single threshold determined based on the complementary performance characteristics of semantic communication and traditional bit communication in different SIR ranges.
[0013] On the other hand, a semantic communication-assisted multi-beam satellite anti-eavesdropping secure transmission system is provided, including: The system modeling module is configured to: construct a multi-beam satellite communication system model, and calculate the channel gain between the user and the satellite based on the multi-beam satellite communication system model. The system model includes users and eavesdroppers, and the users include secure users and ordinary users. The resource allocation module is configured to: use a deep reinforcement learning algorithm to jointly allocate sub-channel resources and power resources in the multi-beam satellite communication system model with system equivalent security rate and user fairness as joint optimization objectives, and output the sub-channel allocation results and power allocation results for each user. The communication paradigm switching module is configured to: calculate the signal-to-interference-plus-noise ratio (SINR) of each user's corresponding link based on the sub-channel allocation result and power allocation result, and allocate a communication paradigm to each user based on the comparison result of the SINR and a preset threshold. The execution module is configured to convert the sub-channel allocation results, power allocation results, and communication paradigm allocation results into an instruction format executable by the spaceborne equipment, and then send them to each execution unit to complete downlink data transmission.
[0014] In another aspect, a computer device is also provided, including a computer-readable storage medium, a processor, and a computer program stored on the computer-readable storage medium and executable on the processor, wherein when the processor executes the program, it performs the method described in the first aspect.
[0015] In another aspect, a computer-readable storage medium is also provided, on which a computer program is stored, which, when executed by a processor, performs the method described in the first aspect.
[0016] The above technical solution has the following advantages or beneficial effects: In terms of security performance, this invention constructs a co-channel heterogeneous transmission architecture for semantic communication and traditional bit communication, and introduces an adaptive paradigm switching mechanism based on the signal-to-interference-plus-noise ratio (SINR) threshold. Even under extreme conditions where the eavesdropper has obtained the complete semantic knowledge base, it can still utilize the complementary performance characteristics of semantic communication in the low SINR region and traditional communication in the high SINR region to make up for the performance shortcomings of traditional physical layer security when the quality of legitimate channels is poor, and effectively guarantee the positive confidentiality rate of secure users.
[0017] Regarding resource utilization and user fairness, this invention takes the system's equivalent secure rate and user fairness as joint optimization objectives, and uses a deep reinforcement learning algorithm to jointly allocate sub-channels and power resources. This overcomes the technical limitations of traditional resource management schemes, which struggle to balance secure transmission efficiency and multi-user service fairness, and achieves synergistic optimization of system security performance and resource utilization efficiency.
[0018] In terms of algorithm complexity and engineering feasibility, this invention uses the SAC algorithm for macro-level resource allocation and combines it with a low-complexity paradigm scheduling algorithm based on the signal-to-interference-plus-noise ratio threshold to form a two-stage decision framework, which significantly reduces the occupation of onboard computing resources. At the same time, by directly converting the resource allocation and paradigm switching results into executable instruction formats for onboard equipment and distributing them to each execution unit, an engineering closed loop of "observation-decision-execution-feedback" is formed, which has high engineering feasibility. Attached Figure Description
[0019] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0020] Figure 1 This is a schematic diagram of a scenario in Embodiment 1 of the present invention; Figure 2 This is a flowchart of the method in Embodiment 1 of the present invention; Figure 3 This is a flowchart of the algorithm in Embodiment 1 of the present invention; Figure 4 This is a diagram of the co-channel heterogeneous transmission architecture of Embodiment 1 of the present invention; Figure 5 This is an approximate regression graph of the semantic similarity function in Embodiment 1 of the present invention; Figure 6 This is a diagram of the SAC algorithm framework in Embodiment 1 of the present invention; Figure 7 This is a flowchart of the control command issuance and execution in Embodiment 1 of the present invention. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings. Those skilled in the art should understand that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0022] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0023] Example 1 This embodiment considers the user downlink secure communication scenario in a multi-beam GEO satellite system, such as... Figure 1As shown in the diagram, in this scenario, the satellite transmits data to various user terminals within the ground coverage area via a multi-beam antenna. These user terminals include mobile terminals, vehicle terminals, and multimedia terminals. An equal number of users are randomly distributed within each beam; these users are categorized as either ordinary users or secure users, with varying proportions in different beams. Considering that the user's movement speed is much smaller than the satellite-to-ground distance and the beam coverage radius, this embodiment models users as fixed-location users. Furthermore, a different number of eavesdroppers are also distributed within each beam, including fixed eavesdroppers and eavesdropping vehicles. The links between these eavesdroppers and the satellite constitute eavesdropping links, used to intercept downlink information received by secure users within their defined range. This invention employs a three-color multiplexing method, dividing the total system bandwidth into three equal parts, labeled B1, B2, and B3. Each part of the bandwidth is further divided into a finite number of sub-channels, with each beam occupying a portion of the three bandwidth parts, and adjacent beams do not interfere with each other. The satellite connects to the ground network through a gateway station, which assists in the allocation and management of communication resources.
[0024] In the above scenario, the core objective of the eavesdropper is to intercept downlink information transmitted from the satellite to the secure user. However, when the secure user and the eavesdropper are close to each other, the inter-satellite channel states are similar, and the channel gain difference is often insufficient to meet the basic premise of traditional physical layer security that the legitimate channel is superior to the eavesdropping channel, leading to the risk of failure of the physical layer security mechanism. To overcome the above difficulties, this embodiment constructs a semantic communication-assisted satellite transmission architecture. This architecture integrates semantic encoding and decoding, channel encoding and decoding, and source encoding and decoding, and configures a switching procedure to support the heterogeneous coexistence and adaptive switching of the traditional bit communication paradigm and the semantic communication paradigm on the same physical channel. Since the performance of the two communication paradigms differs under different channel quality conditions, this embodiment considers the extreme case where the secure user and the eavesdropper share the same semantic knowledge base and both have semantic decoding capabilities. Therefore, it is necessary to dynamically configure the transmission paradigm according to the user's service type and the channel state of the link it is on, thereby achieving multi-objective joint optimization of system security rate and user fairness.
[0025] Figure 2 This is a flowchart of a method according to an embodiment of the present invention. The method includes the following steps: S101: Construct a multi-beam satellite communication system model, and calculate the channel gain between the user and the satellite based on the multi-beam satellite communication system model. The multi-beam satellite communication system model includes users and eavesdroppers. Users include secure users and ordinary users. S102: Taking the system's equivalent security rate and user fairness as joint optimization objectives, a deep reinforcement learning algorithm is used to jointly allocate sub-channel resources and power resources in the multi-beam satellite communication system model, and output the sub-channel allocation results and power allocation results for each user. S103: Based on the sub-channel allocation results and power allocation results, calculate the signal-to-interference-plus-noise ratio (SINR) of the corresponding link for each user, and allocate a communication paradigm to each user based on the comparison result of the SINR and the preset threshold. S104: Convert the sub-channel allocation results, power allocation results, and communication paradigm allocation results into an instruction format executable by the spaceborne equipment, and send them to each execution unit to complete downlink data transmission.
[0026] like Figure 3 As shown, this invention employs the SAC algorithm architecture to implement the macro-resource allocation stage in a two-stage decision-making framework. First, a beam bandwidth occupancy model is constructed based on tri-color multiplexing technology, and sub-channel partitioning is completed. Then, a co-channel heterogeneous transmission architecture is designed to support adaptive switching between traditional bit communication paradigms and semantic communication paradigms. The satellite collects the location information and security attributes of each user and transmits this information to the resource management module. This module, through the SAC decision network, outputs the power allocation value and sub-channel allocation result for each user based on the current system state (including channel gain for each user, eavesdropper location information, accumulated equivalent semantic rate at the previous moment, and tri-color multiplexing grouping status), and transmits the resource allocation action to the performance evaluation module.
[0027] The performance evaluation module evaluates the long-term performance expectation of the current resource allocation decision based on the action-value function through the SAC evaluation network, quantifies the merits of the decision through the reward function, and feeds the evaluation results back to the resource management module for updating the decision network parameters. The above process of "resource allocation-performance evaluation-parameter update" is iterated until the optimization objective converges, thereby achieving the joint maximization of system security rate and user fairness.
[0028] Based on the output power of the SAC algorithm and the sub-channel resource allocation results, the communication paradigm scheduling algorithm further allocates traditional bit communication paradigm or semantic communication paradigm to each user according to the comparison results of the signal-to-interference-plus-noise ratio of each link with the preset threshold. Finally, the resource allocation results and paradigm scheduling results are converted into an instruction format that can be executed by the spaceborne equipment and sent to each execution unit to complete the downlink data transmission.
[0029] Specifically, in step S101, a multi-beam satellite communication system model is constructed and the channel gain between the user and the satellite is calculated. This embodiment considers a single multi-beam GEO satellite, including... The beam, and the first beam There are within each beam A secure user, A regular user and There are several eavesdroppers. The secure user is the user who needs to counter the eavesdropper's interception of downlink information. The eavesdropper is used to intercept downlink information sent by the satellite to the secure user, and the eavesdropper has the same semantic knowledge base as the secure user and has semantic decoding capabilities.
[0030] The system employs tri-color multiplexing technology, with a total bandwidth of It is divided into three equal parts, and each part is further divided into indexes. of Each sub-channel. The beam is also divided into three parts according to the bandwidth occupied, with indices as follows: , and ,in .
[0031] Based on this, users With satellite beam The channel power gain between can be expressed as:
[0032] in, Indicates the transmission gain of the satellite antenna. Indicates user With beam The angle between the axes, A positive value indicates path loss. For users Distance to satellite, Indicates the receiving antenna gain. The angle by which the direction of the received signal is offset from the axis of the receiving antenna.
[0033] exist At any moment, the user Transmission signals in sub-channels The signal-to-interference-plus-noise ratio (SIR) in the equation is expressed as:
[0034] in, The 0-1 variables represent the matching relationship between the sub-channel and the user. Indicates user exist Sub-channels can be occupied at any time ,otherwise ; Indicates the satellite in Sub-channels in each beam Allocated transmission power, Indicates the satellite in Sub-channels in each beam Allocated transmission power, Indicates user With satellite beam Channel power gain between This indicates the power of the AWGN.
[0035] In each beam, each sub-channel at time It can be assigned to at most one user, thus satisfying the following constraint:
[0036] Therefore, ordinary users The reachable rate can be expressed as:
[0037] Due to the presence of eavesdroppers, security users The security rate is affected by both the reachability rate and the eavesdropping rate, and can be expressed as:
[0038] in, , Indicates beam The collection of eavesdroppers in the middle, Indicates eavesdropper For security users In sub-channel The interference-to-noise ratio (IRR) of the eavesdropping signal can be expressed as:
[0039] in, Indicates eavesdropper With satellite beam Channel power gain between Indicates eavesdropper With satellite beam Channel power gain between.
[0040] Based on formulas (2)-(6), the achievable rate may suffer from severe co-channel interference, resulting in a low signal-to-interference-plus-noise ratio; at the same time, due to the presence of eavesdroppers, the confidential rate is further compressed.
[0041] While adjusting resources such as power and sub-channels can improve overall resource utilization to some extent, the similarity of channel state information may limit the gains in resource management. To further enhance the system's secure transmission capability under low signal-to-interference-plus-noise ratio (SINR) conditions, this embodiment introduces a semantic communication auxiliary mechanism into the system model. This mechanism leverages the performance advantages of semantic communication in low SINR regions to improve the overall security rate of the system.
[0042] like Figure 4The diagram illustrates the heterogeneous transmission architecture using a shared channel in this invention. During downlink communication, the satellite transmitter is equipped with a semantic communication coding module and a traditional bit communication coding module, and has a switching mechanism. The semantic communication coding module includes semantic coding and channel coding, while the traditional bit communication coding module includes source coding and channel coding. The user receiver is equipped with a semantic communication decoding module and a traditional bit communication decoding module, and also has a switching mechanism. The semantic communication decoding module includes semantic decoding and channel decoding, while the traditional bit communication decoding module includes source decoding and channel decoding. At the satellite transmitter, the switching mechanism selects the appropriate coding link based on the current link's channel conditions. Source coding can compress the original information, or semantic coding can extract relevant semantic features; both require further channel coding. The encoded signal is transmitted to the user receiver via the physical channel. During transmission, the signal is affected by the superposition of channel noise and semantic noise. At the user receiver, the switching mechanism selects the corresponding decoding link based on the transmitter's communication paradigm. After signal reception, if the traditional bit communication paradigm is used, the original information is recovered sequentially through channel decoding and source decoding; if the semantic communication paradigm is used, the original information is recovered sequentially through channel decoding and semantic decoding. Satellite background knowledge is used for semantic encoding at the transmitting end, while user background knowledge is used for semantic decoding at the receiving end. Both parties ensure consistency in the encoding and decoding process through a knowledge-sharing mechanism, enabling the receiving end to accurately reconstruct the semantic information transmitted by the transmitting end.
[0043] This embodiment is designed for applications such as satellite short message communication, emergency information broadcasting, and remote command issuance, and considers the data payload in the system to be text-based. At any given moment, if a semantic communication paradigm is used to transmit signals, the user Transmission signals in sub-channels The semantic rate in the text is represented as:
[0044] in, 0-1 variables Time satellite via sub-channel To users Communication transmission paradigm Represents a semantic communication paradigm. This represents the traditional bit communication paradigm; This indicates the expected word count for each sentence. This represents the average number of semantic symbols per word. This represents the expected number of semantic information items in each sentence. This is a semantic similarity function, which represents the similarity between the original sentence sent from the sender and the sentence recovered from the receiver. Using the DeepSC tool and data regression methods, an approximate calculation of the semantic similarity function can be obtained, such as... Figure 5The figure shown is an approximate regression diagram of the semantic similarity function in this invention. Therefore, formula (7) is approximately expressed as:
[0045] in, This represents a semantic similarity approximation function.
[0046] If signals are transmitted using the traditional bit communication paradigm, the user Transmission signals in sub-channels The equivalent semantic rate in the equation is represented as:
[0047] in, This represents the average number of bits per word. This represents the semantic similarity of traditional bit communication. This value is affected by bit errors and is usually set to a constant of 1.
[0048] Therefore, ordinary users The equivalent semantic rate can be derived as the semantic rate under the semantic communication paradigm. Equivalent semantic rate under the traditional bit communication paradigm The summation, and in the sub-channel Summing the above is expressed as:
[0049] Considering the extreme case that an eavesdropper might acquire the complete background knowledge base of a legitimate user, thereby gaining semantic decoding capabilities, this embodiment further provides an upper bound on the semantic rate and an upper bound on the equivalent semantic rate for the eavesdropper within the semantic communication paradigm. Specifically, the eavesdropper... For security users In sub-channel The upper bounds of the semantic rate and the equivalent semantic rate in the equation can be expressed as follows:
[0050] Therefore, secure users The equivalent confidential semantic rate lower bound can be expressed as:
[0051] The lower bound of the system's total equivalent confidential semantic rate can be expressed as:
[0052] Furthermore, while considering security rates, this embodiment also optimizes fairness for ordinary users by constructing a long-term proportional fairness mechanism to guarantee the basic communication capabilities of each user. Ordinary users The cumulative equivalent semantic rate can be expressed as:
[0053] in, Indicates the time-scale influence factor. Indicates a regular user exist The cumulative equivalent semantic rate over time.
[0054] The system's total exponential accumulation equivalent semantic rate can be expressed as:
[0055] In step S102, based on the system model constructed in step S101 and the defined lower bound of the equivalent confidentiality semantic rate for secure users and the accumulated equivalent semantic rate for ordinary users, this step further coordinates power allocation, sub-channel allocation, and paradigm scheduling to construct a joint optimization problem to achieve a joint improvement in system confidentiality rate and user fairness. The mathematical model of this optimization problem is shown in formula (17):
[0056] in, C 1 indicates that the total power of all users must not exceed the satellite's maximum transmission power. P max , C 2 indicates that each subchannel within each beam can be allocated to at most one user at a single moment, i.e., subchannel exclusive constraint; C 3 -C 4 represents 0-1 variable constraints, indicating the value ranges of the sub-channel allocation variable and the communication paradigm variable, respectively. C 3. Each user is limited to a sub-channel allocation state of 0 or 1 in each time slot. C 4. Each user is limited to either the semantic communication paradigm or the traditional bit communication paradigm in each sub-channel of each time slot.
[0057] The optimization problem described above belongs to nonconvex optimization and sequential decision problems, and its optimization objective is... Since resource decisions depend on all previous time slots, theoretically, obtaining the globally optimal solution requires joint optimization of all decision variables over the entire time window, resulting in excessive complexity. However, breaking the problem down into individual time slots for successive solutions makes it difficult to guarantee optimal long-term cumulative performance. Therefore, this invention models the time-dependent continuous resource allocation process as a Markov decision process and utilizes a deep reinforcement learning strategy incorporating Bellman equations to construct resource decision networks and resource evaluation networks, achieving a balance between online real-time decision-making and optimal long-term cumulative performance.
[0058] Meanwhile, considering the strict constraints on computing resources, power budget, and real-time response capabilities faced by satellite equipment during on-orbit operation, its core tasks such as communication and attitude control already consume the majority of energy, leaving extremely limited power margin for computing tasks, making it difficult to support highly complex iterative optimization algorithms. Based on this, this invention further adopts a two-stage decision-making framework: the deep reinforcement learning algorithm is only responsible for the joint allocation of sub-channels and power resources, while the selection of the communication paradigm is independently completed by a subsequent paradigm scheduling algorithm based on the signal-to-interference-plus-noise ratio (SINR) threshold. By decoupling the paradigm switching decision from the reinforcement learning action space, this framework effectively compresses the dimensionality of the decision space while ensuring basic performance, significantly reducing the complexity of onboard online solutions and the dependence on real-time computing power.
[0059] In the first stage of the two-stage framework described above, this invention employs the SAC algorithm as the core solution framework for deep reinforcement learning. For example... Figure 6 As shown, the algorithm framework consists of a decision network and an evaluation network, with an experience pool used to store historical interaction samples. The decision network uses policy gradients to make resource decisions, and its output is the action policy. The network parameters are updated by calculating the policy gradient through the optimizer; the network output action-value function (i.e., Q-value function) is evaluated. , The algorithm evaluates the long-term performance of decisions and updates network parameters by calculating temporal difference error (TD error) through an optimizer. The evaluation network further passes the parameters to the corresponding Q-target networks (Q1 and Q2 target networks) via soft updates to stabilize the training process. The algorithm also utilizes an entropy maximization mechanism to maintain the randomness of the policy, preventing it from getting trapped in local optima. During training, an experience pool continuously stores state transition samples generated by the agent's interactions with the environment. When the network is updated, a mini-batch of samples is randomly sampled from the experience pool for batch learning. An experience replay mechanism breaks the temporal correlation between samples to improve training efficiency.
[0060] In this embodiment, a multi-beam satellite secure communication system is used as the environment, and the satellite itself acts as an intelligent agent, continuously sampling the system state and interacting with the SAC algorithm network. The network selects power and sub-channel resources based on the current environment state, and approximates the optimal resource allocation strategy through iterative learning. The network structure, state, actions, and rewards are designed in detail below.
[0061] Decision Network: This network is responsible for executing the action selection strategy. Its input is the system state vector, and its output is the action mean and log-standard deviation. Based on this, a Gaussian distribution function of the actions is constructed, and then the current action is generated through reparameterized sampling. The decision network adjusts its parameters based on the action-value function fed back from the evaluation network.
[0062] State design: The system state includes all useful information in the current environment to guide the strategy in selecting resource actions. Based on the optimization problem in formula (17), the state includes the accumulated equivalent semantic rate of all ordinary users in the previous time step. This parameter changes over time; in addition, static information, including global channel gain, must also be considered. Location information of the eavesdropper Three-color multiplexing grouping of beams Therefore, the state vector can be represented as:
[0063] Action design: The action describes the resource selection currently being performed, including the global power variable and the global sub-channel variable in the optimization problem (17). The action space is a mixed type of continuous and discrete, which can be represented as:
[0064] Evaluation Network: This network is responsible for evaluating the expected reward of the state-action pair. It employs a dual-Q network architecture to suppress the overestimation problem of the reward. The input is the current state and action, and the output is a single scalar Q-value. Since deep reinforcement learning algorithms only optimize power and sub-channel parameters, the current rate value is not the final performance. Therefore, the evaluation network needs to wait for feedback from subsequent communication paradigm resource scheduling, use this feedback value to calculate the target Q-value, and correct the network parameters.
[0065] Reward Design: The overall reward of the system is the cumulative value over time, and a discount factor is used to control the impact of future rewards. However, the design of immediate rewards needs to consider the feasibility and optimality of each state-action pair. Since both the decision network and the evaluation network need to update parameters using rewards, the rewards must align with the resource allocation objectives. Based on the optimization problem in formula (17), the reward consists of three parts: Part 1: Considering the Optimization Objective Using the lower bound of the sum of secure users' equivalent confidentiality semantic rate as the first reward component, it can be expressed as:
[0066] Part Two: Considering the Optimization Objective Using the system's total exponential accumulation equivalent semantic rate as the second reward component, it can be expressed as:
[0067] Part Three: Considering Resource Constraints C 1 and C 2. A negative reward needs to be set to penalize state-action pairs that violate this constraint, in order to reduce the probability of the policy making an incorrect action selection. Therefore, the third reward component can be represented as:
[0068] in, Indicates a power constraint indicator. Indicates corresponding to beam Neutron Channel The sub-channel constraint indicators are represented as follows:
[0069] in, and All are negative constants; in this embodiment, they are taken as -1 and -0.1, respectively.
[0070] The total reward is a weighted sum of the three components mentioned above. By adjusting the weight parameters to determine the relative importance of each optimization objective in the decision-making process, it can be expressed as:
[0071] in, , , The weighting parameters are 0.4, 0.4, and 0.2 respectively in this embodiment.
[0072] Regarding the hyperparameter settings for network training, the decision network has two hidden layers, each containing 256 neurons, using the ReLU activation function. The output layer uses the Sigmoid activation function to construct the action mean and the Softplus activation function to construct the action variance, and is optimized using the Adam optimizer with a learning rate of 1e-4. The evaluation network has three hidden layers, containing 256, 256, and 128 neurons respectively, using the ReLU activation function. The output layer uses a linear activation function to ensure the availability of positive and negative rewards, and is optimized using the Adam optimizer with a learning rate of 1e-3. The discount factor is set to 0.9, the experience pool can hold 10,000 sample tuples, the mini-batch size for each sampling is 256, the initial value for automatic entropy adjustment is 0.1, and the soft update coefficient is 0.005.
[0073] Through the processing of the SAC algorithm described above, this step outputs the sub-channel allocation results and power allocation results for each user, providing a resource allocation basis for the communication paradigm switching based on the signal-to-interference-plus-noise ratio threshold in the subsequent step S103.
[0074] In step S103, based on the sub-channel allocation results and power allocation results of each user output in step S102, this step further determines the actual communication paradigm adopted by each user. It should be noted that the deep reinforcement learning algorithm in step S102 only completed the joint allocation of sub-channels and power resources, and did not involve the selection of the communication paradigm. This selection is completed independently by this step, thus constituting the second stage of the two-stage decoupled decision framework.
[0075] like Figure 5 As shown, semantic communication and traditional bit communication exhibit different performance characteristics under different signal-to-interference-plus-noise ratios (SNR). In the low SNR region, semantic codecs can recover the core meaning of information using background knowledge bases even under poor signal quality conditions, thereby maintaining a high semantic rate. Therefore, the rate advantage of semantic communication is more obvious. In the high SNR region, traditional bit communication is not constrained by the accuracy of semantic similarity recovery and can achieve a high rate performance close to the theoretical limit based on Shannon capacity. Moreover, its encoding and decoding complexity is significantly lower than that of semantic communication, combining the dual advantages of low complexity and high rate.
[0076] Leveraging the complementary performance characteristics of the two communication paradigms mentioned above, the global signal-to-interference-plus-noise ratio (SIR) set under the resource allocation scheme output by the deep reinforcement learning algorithm in step S102 is calculated, and a preset threshold is established. To schedule communication paradigms.
[0077] Specifically, for any user (including security users and ordinary users), when the signal-to-interference-plus-noise ratio (SIR) of the link corresponding to that user is greater than a preset threshold... When the channel quality is sufficient to support high-fidelity transmission of traditional bit communication, the traditional bit communication paradigm is allocated; however, when the signal-to-interference-plus-noise ratio (SIR) of the link corresponding to the user is less than or equal to a preset threshold... At this time, the semantic rate corresponding to the channel quality is better than the equivalent semantic rate under the traditional bit communication paradigm, so the semantic communication paradigm is allocated.
[0078] The solution for the paradigm scheduling can be expressed as:
[0079] in, This represents the unit step function.
[0080] The preset threshold is a single threshold value determined based on the complementary performance characteristics of the two communication paradigms. Its value can be pre-calibrated according to the system operating point and performance requirements. In this embodiment, it is set to 10dB.
[0081] Through the aforementioned signal-to-interference-plus-noise ratio (SINR) threshold decision mechanism, this step achieves the adaptive switching of communication paradigms between semantic communication and traditional bit communication with extremely low computational overhead. This provides a complete resource scheduling scheme for instruction conversion and downlink data transmission in the subsequent step S104, including three decision outputs: sub-channel allocation result, power allocation result, and communication paradigm allocation result.
[0082] In step S104, based on the sub-channel allocation results, power allocation results, and communication paradigm allocation results output in steps S101 to S103, this step further transforms the above algorithm decisions into physical instructions that the spaceborne equipment can actually execute, so as to complete the actual transmission of downlink data.
[0083] Figure 7 This is a flowchart illustrating the control command issuance and execution process of this invention. First, after the resource management module completes the resource scheduling decision for the current time slot, it forms a complete resource configuration scheme, including power allocation values for each user, sub-channel allocation results, and communication paradigm mode labels. Then, the command generator converts floating-point numbers into fixed-width integers according to preset quantization rules and packages them into binary control command frames according to the spaceborne communication protocol format. Finally, the command frames are sent to each execution unit via the spaceborne bus. These execution units, in addition to the traditional baseband resource mapper, also include the paradigm switching module and semantic encoding / decoding module proposed in this invention: the baseband resource mapper maps baseband signals to corresponding physical resource blocks according to the sub-channel allocation results and configures the transmit power according to the power allocation values; the paradigm switching module routes signals to the corresponding processing links according to the communication paradigm mode labels; and the semantic encoding / decoding module performs semantic encoding or decoding processing on the signals in semantic communication mode. All the above modules work together to complete the baseband processing and physical transmission of downlink data.
[0084] In addition, to form a continuously optimized control loop, after the satellite equipment executes the command, its operating status (including actual transmission power, signal-to-interference-plus-noise ratio and other information fed back by the user) can be transmitted back to the on-board base station through the telemetry link, triggering a new round of resource decision-making, forming a continuously optimized control loop, and realizing the closed loop of "algorithm decision-making, command generation, and equipment execution".
[0085] Example 2 This embodiment provides a semantic communication-assisted multi-beam satellite anti-eavesdropping secure transmission system, including: The system modeling module is configured to: construct a multi-beam satellite communication system model, calculate the channel gain between the user and the satellite based on the multi-beam satellite communication system model, and include users and eavesdroppers in the multi-beam satellite communication system model, including secure users and ordinary users; The resource allocation module is configured to: use a deep reinforcement learning algorithm to jointly allocate sub-channel resources and power resources in the multi-beam satellite communication system model with the system equivalent security rate and user fairness as joint optimization objectives, and output the sub-channel allocation results and power allocation results for each user. The communication paradigm switching module is configured to: calculate the signal-to-interference-plus-noise ratio (SINR) of each user's corresponding link based on the sub-channel allocation results and power allocation results, and allocate a communication paradigm to each user based on the comparison result of the SINR and a preset threshold. The execution module is configured to convert the sub-channel allocation results, power allocation results, and communication paradigm allocation results into an instruction format executable by the spaceborne equipment, and then distribute them to each execution unit to complete downlink data transmission.
[0086] It should be noted that each module in this embodiment corresponds one-to-one with each step in Embodiment 1, and their specific implementation process is the same, so it will not be repeated here.
[0087] Example 3 This embodiment also provides a computer device, including a computer-readable storage medium, a processor, and a computer program stored on the computer-readable storage medium and executable on the processor. When the processor executes the program, it completes the method described in Embodiment 1.
[0088] It should be understood that in this embodiment, the processor can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0089] Memory may include read-only memory and random access memory, and provides instructions and data to the processor. A portion of the memory may also include non-volatile random access memory. For example, memory may also store information about the device type.
[0090] In the implementation process, each step of the above method can be completed by the integrated logic circuits in the processor hardware or by software instructions.
[0091] The method in Embodiment 1 can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor. The software modules can reside in readily available storage media in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, a detailed description is not provided here.
[0092] Those skilled in the art will recognize that the units and algorithm steps described in connection with the various examples of this embodiment can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this invention.
[0093] Example 4 This embodiment also provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, complete the method described in Embodiment 1.
[0094] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A semantic communication-assisted multi-beam satellite anti-eavesdropping secure transmission method, characterized in that, include: A multi-beam satellite communication system model is constructed, and the channel gain between the user and the satellite is calculated based on the multi-beam satellite communication system model. The multi-beam satellite communication system model includes users and eavesdroppers, and the users include secure users and ordinary users. With the system's equivalent security rate and user fairness as joint optimization objectives, a deep reinforcement learning algorithm is used to jointly allocate the sub-channel resources and power resources in the multi-beam satellite communication system model, and output the sub-channel allocation results and power allocation results for each user. Based on the sub-channel allocation results and power allocation results, the signal-to-interference-plus-noise ratio (SINR) of each user's corresponding link is calculated, and a communication paradigm is assigned to each user based on the comparison result of the SINR with a preset threshold. The sub-channel allocation results, power allocation results, and communication paradigm allocation results are converted into an instruction format executable by the spaceborne equipment and sent to each execution unit to complete downlink data transmission.
2. The semantic communication-assisted multi-beam satellite anti-eavesdropping secure transmission method according to claim 1, characterized in that, Before constructing the multi-beam satellite communication system model, the following steps are also taken: constructing a beam bandwidth occupancy model based on tri-color multiplexing technology, dividing the total bandwidth into three parts on average, and further dividing each part into a finite number of sub-channels. Each beam occupies a portion of the bandwidth of the three parts, and there is no mutual interference between adjacent beams.
3. The semantic communication-assisted multi-beam satellite anti-eavesdropping secure transmission method according to claim 1, characterized in that, The secure user is a user who needs to counter eavesdroppers intercepting downlink information. The eavesdropper is used to intercept downlink information sent by the satellite to the secure user, and the eavesdropper has the same semantic knowledge base as the secure user and has semantic decoding capabilities.
4. The semantic communication-assisted multi-beam satellite anti-eavesdropping secure transmission method according to claim 1, characterized in that, The system equivalent security rate is the lower bound of the system's total equivalent security semantic rate, and the user fairness is the system's total exponentially accumulated equivalent semantic rate.
5. The semantic communication-assisted multi-beam satellite anti-eavesdropping secure transmission method according to claim 1, characterized in that, The deep reinforcement learning algorithm employs the SAC algorithm, which includes a decision network and an evaluation network. The input of the decision network is the system state, and the output is the action policy. The system state includes the accumulated equivalent semantic rate of all ordinary users at the previous time step, the global channel gain, the eavesdropper's location information, and the tri-color multiplexing grouping of the beam. The action policy includes global power allocation variables and global sub-channel allocation variables. The evaluation network is used to evaluate the expected reward of the state-action pair and uses a dual-Q network architecture to suppress overestimation of the reward. The SAC algorithm only outputs the sub-channel allocation results and power allocation results.
6. The semantic communication-assisted multi-beam satellite anti-eavesdropping secure transmission method according to claim 5, characterized in that, The reward function of the SAC algorithm includes: a first reward component, representing the lower bound of the total equivalent confidentiality semantic rate of secure users; a second reward component, representing the total exponential accumulation equivalent semantic rate of the system; and a third reward component, used to penalize state-action pairs that violate the maximum satellite transmit power constraint and the sub-channel exclusive constraint; the reward function is a weighted sum of the first reward component, the second reward component, and the third reward component.
7. The semantic communication-assisted multi-beam satellite anti-eavesdropping secure transmission method according to claim 1, characterized in that, The step of assigning a communication paradigm to each user based on the comparison result of the signal-to-interference-plus-noise ratio (SIR) and a preset threshold specifically includes: for any user, when the SIR of the link corresponding to the user is greater than the preset threshold, the traditional bit communication paradigm is adopted; when the SIR of the link corresponding to the user is less than or equal to the preset threshold, the semantic communication paradigm is adopted; the preset threshold is a single threshold determined based on the complementary performance characteristics of semantic communication and traditional bit communication in different SIR ranges.
8. A semantic communication-assisted multi-beam satellite anti-eavesdropping secure transmission system, characterized in that, include: The system modeling module is configured to: construct a multi-beam satellite communication system model, and calculate the channel gain between the user and the satellite based on the multi-beam satellite communication system model. The multi-beam satellite communication system model includes users and eavesdroppers, and the users include secure users and ordinary users. The resource allocation module is configured to: use a deep reinforcement learning algorithm to jointly allocate sub-channel resources and power resources in the multi-beam satellite communication system model with system equivalent security rate and user fairness as joint optimization objectives, and output the sub-channel allocation results and power allocation results for each user. The communication paradigm switching module is configured to: calculate the signal-to-interference-plus-noise ratio (SINR) of each user's corresponding link based on the sub-channel allocation result and power allocation result, and allocate a communication paradigm to each user based on the comparison result of the SINR and a preset threshold. The execution module is configured to convert the sub-channel allocation results, power allocation results, and communication paradigm allocation results into an instruction format executable by the spaceborne equipment, and then send them to each execution unit to complete downlink data transmission.
9. A computer device comprising a computer-readable storage medium, a processor, and a computer program stored on the computer-readable storage medium and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the semantic communication-assisted multi-beam satellite anti-eavesdropping secure transmission method as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the semantic communication-assisted multi-beam satellite anti-eavesdropping secure transmission method as described in any one of claims 1-7.