Security beamforming method, device, medium and equipment based on reinforcement learning
Through the adaptive design of beamforming matrix of neural networks trained by reinforcement learning, the problems of small coverage and multi-object interference in the communication and perception integrated system are solved, the system's secure transmission efficiency and reliability are improved, and the interference resistance to eavesdropping targets is enhanced.
Patent Information
- Application Number
- CN202211651558.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-21
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2042-12-21
AI Technical Summary
In the integrated communication perception system, the prior art has problems such as small coverage, large path loss, and interference between users in multi-user and multi-target scenarios that affect communication performance, resulting in insufficient secure transmission efficiency and reliability.
Using a secure beamforming method based on reinforcement learning, the intelligent control of the base station beamforming matrix, RIS phase shift matrix and switch control vector is adopted, and the beamforming matrix is adaptively designed to improve system security and reliability through intelligent control of the base station beamforming matrix, combined with reinforcement learning, and training neural networks.
In multi-user and multi-object scenarios, the secure transmission efficiency and reliability of the communication and perception integrated system are improved, the interference resistance to eavesdropping targets is enhanced, and the spectrum efficiency and system robustness are improved.
Smart Images

Figure CN116054896B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer and communication technology, and more specifically, to a reinforcement learning-based secure beamforming method, apparatus, medium, and device. Background Art
[0002] With the advent of the 6G era, the communication spectrum is moving towards millimeter waves and terahertz frequencies. To address spectrum congestion, Integrated Sensing and Communication (ISAC) technology has been proposed. Among these, the Dual-Functional Radar-Communication (DFRC) system uses a dual-function hardware platform to transmit beam signals, simultaneously implementing both sensing and communication functions. This significantly reduces complexity and cost, and has garnered widespread attention from academia and industry. To improve radar sensing performance, base stations aim to focus radiated power in the direction of the target when transmitting beams. However, since the transmitted signal carries the user's information, if the target is a potential eavesdropper, there is a risk that confidential information could be intercepted and decoded. Therefore, new physical layer security solutions are needed to achieve secure DFRC information transmission.
[0003] Among current technical solutions, a reconfigurable intelligent surface (RIS) is a uniform planar array composed of many low-cost passive reflective elements. Each element can adaptively adjust its reflection amplitude and phase to control the intensity and direction of electromagnetic waves. This allows the RIS to enhance or weaken the reflected signal for different users. However, single RIS assistance still suffers from issues such as limited coverage and path loss. In practical applications, with multiple users and multiple targets, inter-user interference can affect communication performance. Therefore, improving the efficiency and reliability of secure transmission in integrated communication and perception systems has become a pressing technical challenge. Summary of the Invention
[0004] The embodiments of the present application provide a reinforcement learning-based secure beamforming method, apparatus, medium, and device, which can improve the efficiency and reliability of secure transmission of a communication-awareness integrated system, at least to a certain extent.
[0005] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by practice of the present application.
[0006] According to one aspect of an embodiment of the present application, a reinforcement learning-based secure beamforming method is provided, which is applied to a communication and perception integrated system. The communication and perception integrated system includes a dual-function radar communication base station and several RISs, wherein the base station is connected to each of the RISs via an intelligent controller.
[0007] The method comprises:
[0008] At the beginning of each time slot, based on the current base station beamforming matrix, RIS phase shift matrix and switch control vector, the current communication state of the communication and perception integrated system is determined, wherein the communication state includes the channel state and achievable rate corresponding to the base station and each communication user and each eavesdropping target;
[0009] Selecting a target action from the action set and determining an action reward corresponding to executing the target action, wherein the target action includes a changed base station beamforming matrix, a RIS phase shift matrix, and a switch control vector;
[0010] After the communication-awareness integrated system performs the target action, determining a communication state of the communication-awareness integrated system at the beginning of a next time slot;
[0011] Associating the current communication state, target action, action reward, and communication state at the start of the next time slot corresponding to the communication-perception integrated system, and storing them in an experience pool as training data;
[0012] When the amount of training data in the experience pool reaches a predetermined condition, a number of training data are obtained from the experience pool to train the pre-constructed neural network to be trained, so as to obtain a target neural network for secure beamforming.
[0013] According to one aspect of an embodiment of the present application, a reinforcement learning-based secure beamforming device is provided, which is applied to a communication and perception integrated system. The communication and perception integrated system includes a dual-function radar communication base station and several RISs, wherein the base station is connected to each of the RISs via an intelligent controller.
[0014] The device comprises:
[0015] A first determination module is configured to determine, at the beginning of each time slot, a current communication state of the communication-awareness integrated system based on a current base station beamforming matrix, a RIS phase shift matrix, and a switch control vector, wherein the communication state includes a channel state and a achievable rate corresponding to the base station and each communication user and each eavesdropping target;
[0016] A second determination module is configured to select a target action from the action set and determine an action reward corresponding to executing the target action, wherein the target action includes the changed base station beamforming matrix, RIS phase shift matrix, and switch control vector;
[0017] A third determination module is configured to determine a communication state of the communication perception integrated system at the beginning of a next time slot after the communication perception integrated system performs the target action;
[0018] A storage module is used to associate the current communication state, target action, action reward and communication state at the beginning of the next time slot corresponding to the communication perception integrated system, and store them in an experience pool as training data;
[0019] The processing module is used to obtain a number of training data from the experience pool to train the pre-built neural network to be trained when the amount of training data in the experience pool reaches a predetermined condition, so as to obtain a target neural network for security beamforming.
[0020] According to one aspect of an embodiment of the present application, a computer-readable medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the reinforcement learning-based secure beamforming method as described in the above embodiment is implemented.
[0021] According to one aspect of an embodiment of the present application, an electronic device is provided, comprising: one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the reinforcement learning-based secure beamforming method as described in the above embodiments.
[0022] According to one aspect of an embodiment of the present application, a computer program product or computer program is provided. The computer program product or computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the reinforcement learning-based secure beamforming method provided in the above-described embodiments.
[0023] In the technical solutions provided in some embodiments of the present application, at the beginning of each time slot, based on the current base station beamforming matrix, RIS phase shift matrix and switch control vector, the current communication state of the communication perception integrated system is determined, the communication state including the channel state and achievable rate corresponding to the base station and each communication user and each eavesdropping target, and a target action is selected from the action set, and the action reward corresponding to the target action is determined after the target action is executed, the target action including the changed base station beamforming matrix, RIS phase shift matrix and switch control vector, and then after the communication perception integrated system executes the target action, the communication state of the communication perception integrated system at the beginning of the next time slot is determined, so as to associate the current communication state, target action, action reward and communication state corresponding to the communication perception integrated system at the beginning of the next time slot and store them as training data in an experience pool, and when the number of training data in the experience pool reaches a predetermined condition, a number of training data are obtained from the experience pool to train the pre-constructed neural network to be trained, thereby obtaining a target neural network. Thus, the target neural network obtained by the above method enables the base station to adaptively design the beamforming matrix using the reinforcement learning method in multiple RIS-assisted, multi-user, and multi-target scenarios, thereby improving the efficiency and reliability of secure transmission of the communication perception integrated system.
[0024] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The accompanying drawings are incorporated into and constitute a part of the specification, illustrating embodiments consistent with the present application and, together with the specification, explaining the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and those skilled in the art can derive other drawings based on these drawings without inventive effort. In the drawings:
[0026] Figure 1 A schematic diagram showing an exemplary system architecture to which the technical solutions of the embodiments of the present application can be applied;
[0027] Figure 2 A flowchart of a secure beamforming method based on reinforcement learning according to an embodiment of the present application is shown;
[0028] Figure 3 A block diagram of a security beamforming device based on reinforcement learning according to an embodiment of the present application is shown;
[0029] Figure 4 A schematic diagram of the structure of a computer system suitable for implementing an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0030] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art.
[0031] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner.In the following description, many specific details are provided so as to provide a full understanding of the embodiments of the present application. However, it will be appreciated by those skilled in the art that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps etc. can be adopted. In other cases, known methods, devices, implementations or operations are not shown or described in detail to avoid blurring the various aspects of the application.
[0032] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically separate entities. That is, these functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0033] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.
[0034] Figure 1 A schematic diagram shows an exemplary system architecture to which the technical solutions of the embodiments of the present application can be applied.
[0035] like Figure 1 As shown, the system architecture may include a DFRC base station, several intelligent controllers and several RIS, wherein the DFRC base station is connected to each RIS via the intelligent controller. T A uniform linear array of 10 antennas communicates with K single-antenna users and simultaneously senses M eavesdropping targets. L RISs are deployed between the DFRC base station and the users for auxiliary communication. The base station and RISs are connected via an intelligent controller, and multiple RISs adaptively control the number of accesses using switches.
[0036] It should be noted that the reinforcement learning-based secure beamforming method provided in the embodiments of the present application is generally executed by a base station. Accordingly, the reinforcement learning-based secure beamforming device is generally provided in the base station. However, in other embodiments of the present application, the terminal device may also have similar functions to the server, thereby implementing the reinforcement learning-based secure beamforming method provided in the embodiments of the present application.
[0037] The following is a detailed description of the implementation details of the technical solution of the embodiment of the present application:
[0038] Figure 2 A flowchart of a security beamforming method based on reinforcement learning according to an embodiment of the present application is shown. The security beamforming method based on reinforcement learning for a communication perception integrated system can be executed by a base station, which can be Figure 1 The DFRC base station shown in . Figure 2 As shown, the reinforcement learning-based secure beamforming method includes at least steps S210 to S250, which are described in detail as follows:
[0039] In step S210, at the beginning of each time slot, the current communication state of the communication perception integrated system is determined based on the current base station beamforming matrix, RIS phase shift matrix and switch control vector. The communication state includes the channel state and achievable rate corresponding to the base station and each communication user and each eavesdropping target.
[0040] In this embodiment, the base station can observe the system to determine the current communication state of the communication perception integrated system at the beginning of each time slot based on the base station beamforming matrix, RIS phase shift matrix and switch control vector executed by the current base station. The communication state may include but is not limited to the channel state and achievable rate corresponding to the base station and each communication user and each eavesdropping target.
[0041] In particular, at the initial time slot (ie, time t=0), the initial base station beamforming matrix w, RIS phase shift matrix Ψ, and switch control vector β may be randomly generated to facilitate subsequent observations.
[0042] In one embodiment, at the beginning of each time slot, the current communication state of the communication-awareness integrated system is determined based on the current base station beamforming matrix, the RIS phase shift matrix, and the switch control vector, including:
[0043] Determine, based on the DFRC pilot signal transmitted by the base station, the channel vector between the base station and each communication user, the channel matrix between the base station and each RIS, and the channel vector between each RIS and the communication user; and determine, based on the DFRC radar reflection signal transmitted by the base station, the angle between the base station and each eavesdropping target;
[0044] At the beginning of each time slot, the current communication state of the communication perception integrated system is determined according to the first channel vector, the channel matrix, the second channel vector, the angle, the current base station beamforming matrix, the RIS phase shift matrix and the switch control vector.
[0045] In this embodiment, the base station can obtain the channel vector h between itself and the kth communication user by sending a pilot signal. bu,k , the channel matrix H between it and the lth RIS br,l , and the channel vector h between the lth RIS and the kth communication user ru,l,k And obtain the angle θ of the mth eavesdropping target through the DRFC radar reflection signal m Based on the above information and the current base station beamforming matrix, RIS phase shift matrix and switch control vector, the current communication state of the communication-awareness integrated system can be determined.
[0046] Specifically, the base station can determine the channel vector (i.e., channel state) of the communication user according to the following formula:
[0047]
[0048] Among them, h k is the channel vector of the kth communication user, h bu,k is the channel vector between the base station and the kth communication user, h ru,l,k represents the channel vector between the lth RIS and the kth communication user, β l represents the switch control variable of the lth RIS predicted in the action, Ψ l represents the phase shift matrix of the lth RIS predicted in the action, H br,l is the channel matrix between the base station and the lth RIS.
[0049] The channel vector calculation formula of the mth eavesdropping target is as follows:
[0050]
[0051] Among them, h m is the channel vector of the mth eavesdropping target, κ represents the path loss coefficient, a(θ) represents the array response vector, θ m represents the angle of the mth eavesdropping target, N T is the number of antennas of the base station.
[0052] According to the above channel status, two achievable rates can be calculated. The achievable rate of the kth communication user is calculated as follows:
[0053]
[0054]
[0055] Among them, R j k and SINR j k They represent the achievable rate and signal-to-noise ratio of the jth communication user with respect to the kth communication user, w k represents the predicted beamforming vector of the kth communication user, represents the variance of the Gaussian noise in the j-th user received signal.
[0056] Similarly, the calculation formula for the reachable rate of the eavesdropping target with respect to the kth communication user is:
[0057]
[0058]
[0059] in, and They represent the achievable rate and signal-to-noise ratio of the mth eavesdropping target with respect to the kth communication user, The two channel vectors and two achievable rates mentioned above constitute the state set in
[0060] Please continue to refer to Figure 2 In step S220, a target action is selected from the action set, and an action reward corresponding to the execution of the target action is determined. The target action includes the changed base station beamforming matrix, RIS phase shift matrix, and switch control vector.
[0061] In this embodiment, the action set may be pre-set by a person skilled in the art, and may include several executable actions, each of which includes a corresponding base station beamforming matrix, RIS phase shift matrix, and switch control vector, i.e. in In one example, the base station may randomly select an action from the action set as the target action; in another example, the base station may select from the action set sequentially (eg, from front to back or from back to front, etc.).
[0062] It should be understood that the target actions selected by the base station in different time slots may be the same or different, and this application does not impose any special limitation on this.
[0063] After determining the target action, the base station can execute the target action, namely transmitting signals using the beamforming matrix and communicating with the intelligent controller to control the access and phase shift of the RIS. The base station then determines an action reward after executing the target action, and the communication quality after executing the target action is determined based on the action reward.
[0064] In one embodiment of the present application, when determining the action reward corresponding to the execution of the target action, the base station can determine the corresponding action reward based on the confidentiality interruption probability after the execution of the target action, the matching degree of the reachable rate of the communication user, the matching degree of the eavesdropping target with respect to the reachable rate of the communication user, and the matching degree of the transmission beam pattern.
[0065] Specifically, the calculation formula for action rewards is as follows:
[0066]
[0067] Among them, p out represents the probability of confidentiality interruption, and p d represents the matching degree of the achievable rate of the kth communication user, the matching degree of the eavesdropping target with respect to the achievable rate of the kth communication user, and the matching degree of the transmission beam pattern. μ1, μ2, and μ3 represent the weight coefficients of the above three matching degrees, respectively.
[0068] The calculation formula for the confidentiality interruption probability is as follows:
[0069]
[0070]
[0071] Where [x] + =max(0,x), represents the confidentiality rate of the kth communication user, is the confidentiality rate threshold.
[0072] The calculation formulas for the three matching degrees are as follows:
[0073]
[0074]
[0075]
[0076] in, and γ p Both are matching thresholds. X represents the covariance matrix of the transmitted signal, R d represents the covariance matrix of the desired generated beam pattern.
[0077] Please continue to refer to Figure 2 In step S230, after the communication perception integration system executes the target action, the communication state of the communication perception integration system at the beginning of the next time slot is determined.
[0078] In this embodiment, the base station may observe the communication state at the beginning of the next time slot, thereby obtaining the state set s'.
[0079] In step S240, the current communication state, target action, action reward and communication state at the beginning of the next time slot corresponding to the communication-awareness integrated system are associated and stored in the experience pool as training data.
[0080] In this embodiment, the base station can associate the current communication state s, target action a, action reward r and communication state s' at the beginning of the next time slot corresponding to the communication perception integrated system to form a vector set {s, a, r, s'}, and store it in the experience pool as training data.
[0081] In step S250, when the amount of training data in the experience pool reaches a predetermined condition, a number of training data are obtained from the experience pool to train the pre-constructed neural network to be trained, so as to obtain a target neural network for security beamforming.
[0082] In this embodiment, the predetermined condition can be pre-set by a person skilled in the art based on prior experience, such as when the amount of training data in the experience pool reaches 80% of the storable amount, or when the experience pool is full, etc. This application does not impose any special restrictions on this.
[0083] At this time, the base station can obtain some training data from the training data stored in the experience pool to train the pre-built neural network to be trained, wherein the amount of training data used for training can be determined according to actual implementation needs and is not specifically limited.
[0084] In one example, the D3QN method can be used for training. At the beginning of the initial time slot, the reinforcement learning agent can be initialized. The base station can randomly extract N from the experience pool. batch The neural network is trained using data samples and the target value of the neural network to be trained is calculated using the following formula:
[0085]
[0086] Where γ is the discount factor, and are the parameters of the evaluation network and the target network, respectively. Q and Q' are the Q values calculated by the evaluation network and the target network, respectively.
[0087] Based on the target value, the loss function is updated, and the evaluation network parameters are updated according to the updated loss function.
[0088] Specifically, the loss function calculation formula is as follows:
[0089]
[0090] The network parameters are updated and evaluated according to the loss function. The calculation formula is as follows:
[0091]
[0092] Among them, α is the learning rate.
[0093] Then, based on the updated evaluation network parameters, the parameters of the target network are soft-updated. The calculation formula is as follows:
[0094]
[0095] Where τ is the soft update rate.
[0096] The base station can repeat the above training steps to complete the training, thereby obtaining the target neural network.
[0097] Therefore, the reinforcement learning-based secure beamforming method provided by the aforementioned embodiments outperforms traditional synaesthesia-integrated secure beamforming design methods in terms of complexity and system performance. Furthermore, the reinforcement learning agent can interact with the environment and adaptively design the beamforming matrix and phase shift matrix, making it highly robust to channel variations and achieving high spectral efficiency.
[0098] In one embodiment of the present application, the method further comprises:
[0099] Determine, based on the dual-function radar communication pilot signal transmitted by the base station, the channel vector between the base station and each communication user, the channel matrix between the base station and each RIS, and the channel vector between each RIS and the communication user; and determine, based on the DFRC radar reflection signal transmitted by the base station, the angle between the base station and each eavesdropping target;
[0100] Determine the current channel state and achievable rate corresponding to the base station and each communication user and each eavesdropping target based on the channel vector between the base station and each communication user, the channel matrix between the base station and each RIS, the channel vector between each RIS and the communication user, and the angle between the base station and each eavesdropping target;
[0101] Inputting the channel states and achievable rates corresponding to the current base station and each communication user and each eavesdropping target into the target neural network, so that the target neural network outputs a target beamforming matrix, a target RIS phase shift matrix, and a target switch control vector;
[0102] Signal transmission is performed according to the target beamforming matrix, the target RIS phase shift matrix, and the target switch control vector.
[0103] In this embodiment, the base station determines the channel vector between the base station and each communication user, the channel matrix between the base station and each RIS, and the channel vector between each RIS and the communication user by transmitting a DFRC pilot signal. The angle between the base station and each eavesdropping target is determined based on the DFRC radar reflection signal transmitted by the base station.
[0104] Based on the above information, the communication state of the system is calculated, that is, the channel state and achievable rate corresponding to the base station and each communication user and each eavesdropping target, which is input as environmental data into the trained target neural network. The target neural network can be trained using the method provided in the above embodiment. The target neural network can output a relatively optimal target beamforming matrix, a target RIS phase shift matrix, and a target switch control vector based on the input environmental data. The base station transmits a signal based on the target beamforming matrix and controls the intelligent controller to control the phase shift and access of the RIS.
[0105] Therefore, when a DFRC base station communicates with multiple users and detects multiple targets around it at the same time, there is a risk of information leakage. This application uses reinforcement learning to perform secure beamforming design in multiple RIS-assisted scenarios to improve information transmission security, increase spectrum efficiency, and increase system robustness.
[0106] The following describes an embodiment of the device of the present application, which can be used to perform the method in the above embodiment of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the method in the above embodiment of the present application.
[0107] Figure 3 A block diagram of a reinforcement learning-based secure beamforming device according to an embodiment of the present application is shown. The device is applied to a communication and perception integrated system, comprising a dual-function radar communication base station and several RISs, each of which is connected to the RISs via an intelligent controller.
[0108] Reference Figure 3 As shown, according to an embodiment of the present application, a security beamforming device based on reinforcement learning includes:
[0109] A first determination module 310 is configured to determine, at the beginning of each time slot, a current communication state of the communication-awareness integrated system based on a current base station beamforming matrix, a RIS phase shift matrix, and a switch control vector, wherein the communication state includes a channel state and a achievable rate corresponding to the base station and each communication user and each eavesdropping target;
[0110] A second determination module 320 is configured to select a target action from the action set and determine an action reward corresponding to executing the target action, wherein the target action includes the changed base station beamforming matrix, RIS phase shift matrix, and switch control vector;
[0111] A third determination module 330 is configured to determine a communication state of the communication perception integrated system at the beginning of a next time slot after the communication perception integrated system performs the target action;
[0112] The storage module 340 is used to associate the current communication state, target action, action reward and communication state at the beginning of the next time slot corresponding to the communication perception integrated system, and store them in the experience pool as training data;
[0113] The processing module 350 is configured to obtain a plurality of training data from the experience pool to train the pre-built neural network to be trained when the amount of training data in the experience pool reaches a predetermined condition, so as to obtain a target neural network for security beamforming.
[0114] In one embodiment of the present application, the processing module 350 is used to: obtain training data from the experience pool to determine the target value of the current neural network to be trained based on the training data; update the loss function according to the target value and update the evaluation network parameters according to the updated loss function; and soft-update the parameters of the target network based on the updated evaluation network parameters.
[0115] In one embodiment of the present application, the first determination module 310 is used to: determine the channel vector between the base station and each communication user, the channel matrix between the base station and each RIS, and the channel vector between the RIS and the communication user based on the dual-function radar communication pilot signal sent by the base station, and determine the angle between the base station and each eavesdropping target based on the dual-function radar reflection signal transmitted by the base station; at the beginning of each time slot, determine the current communication status of the communication perception integrated system based on the channel vector between the base station and each communication user, the channel matrix between the base station and each RIS, the channel vector between the RIS and the communication user, the angle between the base station and each eavesdropping target, the current base station beamforming matrix, the RIS phase shift matrix, and the switch control vector.
[0116] In one embodiment of the present application, the second determination module 320 is used to determine the corresponding action reward based on the probability of confidentiality interruption after executing the target action, the matching degree of the reachable rate of the communication user, the matching degree of the eavesdropping target with respect to the reachable rate of the communication user, and the matching degree of the transmission beam pattern.
[0117] In one embodiment of the present application, the processing module 350 is further used to: determine the channel vector between the base station and each communication user, the channel matrix between the base station and each RIS, and the channel vector between the RIS and the communication user according to the dual-function radar communication pilot signal transmitted by the base station; determine the angle between the base station and each eavesdropping target according to the dual-function radar reflection signal transmitted by the base station; determine the current channel state and achievable rate corresponding to the base station and each communication user and each eavesdropping target according to the channel vector between the base station and each communication user, the channel matrix between the base station and each RIS, the channel vector between the RIS and the communication user, and the angle between the base station and each eavesdropping target; input the current channel state and achievable rate corresponding to the base station and each communication user and each eavesdropping target into the target neural network, so that the target neural network outputs a target beamforming matrix, a target RIS phase shift matrix, and a target switch control vector; and transmit a signal according to the target beamforming matrix, the target RIS phase shift matrix, and the target switch control vector.
[0118] Figure 4 A schematic diagram of the structure of a computer system suitable for implementing an electronic device according to an embodiment of the present application is shown.
[0119] It should be noted that Figure 4 The computer system of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0120] like Figure 4 As shown, the computer system includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 402 or the program loaded from the storage part 408 into the random access memory (RAM) 403, such as executing the method described in the above embodiment. Various programs and data required for system operation are also stored in the RAM 403. The CPU 401, ROM 402 and RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0121] The following components are connected to the I / O interface 405: an input section 406 including a keyboard, a mouse, and the like; an output section 407 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 408 including a hard disk and the like; and a communication section 409 including a network interface card such as a LAN (Local Area Network) card or a modem. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as needed. Removable media 411, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 410 as needed, so that computer programs read therefrom can be installed into the storage section 408 as needed.
[0122] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 409, and / or installed from a removable medium 411. When the computer program is executed by the central processing unit (CPU) 401, the various functions defined in the system of the present application are executed.
[0123] It should be noted that the computer-readable medium shown in the embodiments of the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable computer program. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. A computer program embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0124] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. Among them, each box in the flowchart or block diagram can represent a module, program segment, or part of the code, and the above-mentioned module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0125] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. In some cases, the names of these units do not constitute limitations on the units themselves.
[0126] As another aspect, the present application further provides a computer-readable medium, which may be included in the electronic device described in the above embodiments, or may exist independently without being incorporated into the electronic device. The computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device implements the method described in the above embodiments.
[0127] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiment of the application, the features and functions of two or more modules or units described above can be concretized in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.
[0128] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0129] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of this application and include common knowledge or customary techniques in the art that are not disclosed herein.
[0130] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. A secure beamforming method based on reinforcement learning, characterized in that: Applied to a communication and perception integrated system, the communication and perception integrated system includes a dual-function radar communication base station and several RIS, the base station is connected to each of the RIS through an intelligent controller; The method comprises: At the beginning of each time slot, based on the current base station beamforming matrix, RIS phase shift matrix and switch control vector, the current communication state of the communication and perception integrated system is determined, wherein the communication state includes the channel state and achievable rate corresponding to the base station and each communication user and each eavesdropping target; Selecting a target action from the action set and determining an action reward corresponding to executing the target action, wherein the target action includes a changed base station beamforming matrix, a RIS phase shift matrix, and a switch control vector; After the communication-awareness integrated system performs the target action, determining a communication state of the communication-awareness integrated system at the beginning of a next time slot; Associating the current communication state, target action, action reward, and communication state at the start of the next time slot corresponding to the communication-perception integrated system, and storing them in an experience pool as training data; When the amount of training data in the experience pool reaches a predetermined condition, a plurality of training data are obtained from the experience pool to train a pre-constructed neural network to be trained, so as to obtain a target neural network for secure beamforming; At the beginning of each time slot, the current communication state of the communication-awareness integrated system is determined based on the current base station beamforming matrix, RIS phase shift matrix, and switch control vector, including: Determine, based on the dual-function radar communication pilot signal transmitted by the base station, the channel vector between the base station and each communication user, the channel matrix between the base station and each RIS, and the channel vector between each RIS and the communication user; and determine, based on the dual-function radar reflection signal transmitted by the base station, the angle between the base station and each eavesdropping target; At the beginning of each time slot, the current communication state of the communication-awareness integrated system is determined based on the channel vectors between the base station and each communication user, the channel matrix between the base station and each RIS, the channel vectors between each RIS and the communication user, the angle between the base station and each eavesdropping target, the current base station beamforming matrix, the RIS phase shift matrix, and the switch control vector. Determining the action reward corresponding to executing the target action includes: The corresponding action reward is determined according to the confidentiality interruption probability after executing the target action, the matching degree of the communication user's achievable rate, the matching degree of the eavesdropping target with respect to the communication user's achievable rate, and the matching degree of the transmission beam pattern.
2. The method according to claim 1, characterized in that Acquiring a plurality of training data from the experience pool to train the pre-built neural network to be trained includes: Acquiring training data from the experience pool to determine a target value of the current neural network to be trained based on the training data; According to the target value, updating the loss function and updating the evaluation network parameters according to the updated loss function; Based on the updated evaluation network parameters, soft update is performed on the parameters of the target network.
3. The method according to claim 2, characterized in that The channel vector of the communication user is determined according to the following formula: Among them, h k is the channel vector of the kth communication user, h bu,k is the channel vector between the base station and the kth communication user, h ru,l,k represents the channel vector between the lth RIS and the kth communication user, β l represents the switch control variable of the lth RIS predicted in the action, Ψ l represents the phase shift matrix of the lth RIS predicted in the action, H br,l is the channel matrix between the base station and the lth RIS; The channel vector of the eavesdropping target is determined according to the following formula: Among them, h m is the channel vector of the mth eavesdropping target, κ represents the path loss coefficient, a(θ) represents the array response vector, θ m represents the angle of the mth eavesdropping target, N T is the number of antennas of the base station; The achievable rate of the kth communication user is calculated according to the following formula: in, and They represent the achievable rate and signal-to-noise ratio of the jth communication user with respect to the kth communication user, w k represents the predicted beamforming vector of the kth communication user, represents the variance of Gaussian noise in the jth user received signal; The achievable rate of the eavesdropping target with respect to the kth communication user is calculated according to the following formula: in, and They represent the achievable rate and signal-to-noise ratio of the mth eavesdropping target with respect to the kth communication user, represents the variance of the Gaussian noise in the received signal of the m-th eavesdropping target.
4. The method according to claim 1, wherein The method further comprises: Determine, based on the dual-function radar communication pilot signal transmitted by the base station, the channel vector between the base station and each communication user, the channel matrix between the base station and each RIS, and the channel vector between each RIS and the communication user; and determine, based on the dual-function radar reflection signal transmitted by the base station, the angle between the base station and each eavesdropping target; Determine the current channel state and achievable rate corresponding to the base station and each communication user and each eavesdropping target based on the channel vector between the base station and each communication user, the channel matrix between the base station and each RIS, the channel vector between each RIS and the communication user, and the angle between the base station and each eavesdropping target; Inputting the channel states and achievable rates corresponding to the current base station and each communication user and each eavesdropping target into the target neural network, so that the target neural network outputs a target beamforming matrix, a target RIS phase shift matrix, and a target switch control vector; Signal transmission is performed according to the target beamforming matrix, the target RIS phase shift matrix, and the target switch control vector.
5. A secure beamforming device based on reinforcement learning, characterized in that: Applied to a communication and perception integrated system, the communication and perception integrated system includes a dual-function radar communication base station and several RIS, the base station is connected to each of the RIS through an intelligent controller; The device comprises: A first determination module is configured to determine, at the beginning of each time slot, a current communication state of the communication-awareness integrated system based on a current base station beamforming matrix, a RIS phase shift matrix, and a switch control vector, wherein the communication state includes a channel state and a achievable rate corresponding to the base station and each communication user and each eavesdropping target; A second determination module is configured to select a target action from the action set and determine an action reward corresponding to executing the target action, wherein the target action includes the changed base station beamforming matrix, RIS phase shift matrix, and switch control vector; A third determination module is configured to determine a communication state of the communication perception integrated system at the beginning of a next time slot after the communication perception integrated system performs the target action; A storage module is used to associate the current communication state, target action, action reward and communication state at the beginning of the next time slot corresponding to the communication perception integrated system, and store them in an experience pool as training data; a processing module, configured to, when the amount of training data in the experience pool reaches a predetermined condition, obtain a plurality of training data from the experience pool to train a pre-built neural network to be trained, so as to obtain a target neural network for secure beamforming; The first determination module is further configured to determine, based on the dual-function radar communication pilot signal transmitted by the base station, a channel vector between the base station and each communication user, a channel matrix between the base station and each RIS, and a channel vector between each RIS and the communication user, and determine an angle between the base station and each eavesdropping target based on the dual-function radar reflection signal transmitted by the base station; At the beginning of each time slot, the current communication state of the communication-awareness integrated system is determined based on the channel vectors between the base station and each communication user, the channel matrix between the base station and each RIS, the channel vectors between each RIS and the communication user, the angle between the base station and each eavesdropping target, the current base station beamforming matrix, the RIS phase shift matrix, and the switch control vector. Among them, the second determination module is also used to determine the corresponding action reward based on the confidentiality interruption probability after executing the target action, the matching degree of the communication user's reachable rate, the matching degree of the eavesdropping target with respect to the communication user's reachable rate and the matching degree of the transmission beam pattern.
6. The device according to claim 5, characterized in that The processing module is used for: Acquiring training data from the experience pool to determine a target value of the current neural network to be trained based on the training data; According to the target value, updating the loss function and updating the evaluation network parameters according to the updated loss function; Based on the updated evaluation network parameters, soft update is performed on the parameters of the target network.
7. A computer-readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.
8. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, causes the one or more processors to implement the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Secure transmission method and system based on air-based reconfigurable intelligent surface
CN113472419A
Wireless communication transmission method and device based on IRS assistance, terminal and storage medium
CN113747442A