Game theory based access authentication and communication resource control method, device and equipment

By employing a game theory-based access authentication and communication resource control method, channel policy and transmit power are optimized using channel feature data and a deep deterministic policy gradient algorithm. This solves the problems of unauthorized user access and low resource utilization efficiency in wireless networks, achieving higher security and resource utilization efficiency.

CN119155686BActive Publication Date: 2025-11-21XIAMEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410951569.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-16
Publication Date
2025-11-21
Estimated Expiration
2044-07-16

AI Technical Summary

Technical Problem

Wireless networks pose a security threat of unauthorized users impersonating legitimate users to access the network and steal data. Furthermore, traditional optimization methods cannot effectively utilize communication resources when faced with a large number of users or complex environments.

Method used

A game theory-based access authentication and communication resource control method is adopted. By acquiring channel feature data for legitimacy authentication, a dynamic adversarial game model is established, and a deep deterministic policy gradient algorithm is combined to optimize channel policy and transmit power to improve security and resource utilization efficiency.

Benefits of technology

It improves the security of users accessing the network and effectively utilizes communication resources in the event of malicious interference, thereby enhancing communication capacity and system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119155686B_ABST
    Figure CN119155686B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a game theory-based access authentication and communication resource control method, device and equipment. The method comprises: obtaining and preprocessing channel feature data corresponding to a target user; performing legality authentication on the target user according to the preprocessed channel feature data; in the case that the result of the legality authentication is a legal user, establishing a dynamic confrontation game model to form a dynamic confrontation game equilibrium solution, i.e. an optimal channel strategy set, according to different needs of the target user and an interferer; using a deep deterministic policy gradient algorithm, taking the optimal channel strategy set as the state input of an agent, and taking the transmission power of the target user as the action of the agent, to determine a target channel strategy and a target transmission power corresponding to the target user, which make the communication capacity under the condition of maximum malicious interference maximum. The technical solution of the embodiments of the present application can improve the security of user access to the network, and ensure the effective use of communication resources in the presence of communication interference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of computer and communication technology, and more specifically, to a method, apparatus, and device for access authentication and communication resource control based on game theory. Background Technology

[0002] As 5G applications expand from communication between people and information services between people and machines to communication and control between people and objects, and between objects themselves, large-scale access has become a prominent demand. However, compared to wired networks, wireless network communication is more susceptible to interference and disruption by unauthorized users. Some unauthorized users impersonate legitimate users to access the network and steal data, thus threatening network security. Furthermore, facing increasingly fierce competition for communication resources, dynamic spectrum optimization has become an important method to improve resource utilization. Traditional optimization methods, however, are insufficient when dealing with large numbers of users or complex environments because they do not consider the potential impact of malicious interference. Therefore, improving the security of user access to the network and ensuring the effective utilization of communication resources in the presence of communication interference have become urgent technical problems to be solved. Summary of the Invention

[0003] The embodiments of this application provide a game theory-based access authentication and communication resource control method, apparatus, and device, which can at least to some extent improve the security of user access to the network and ensure the effective utilization of communication resources in the presence of communication interference.

[0004] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.

[0005] According to one aspect of the embodiments of this application, a game theory-based access authentication and communication resource control method is provided, including:

[0006] Acquire and preprocess channel feature data corresponding to the target user;

[0007] Based on the preprocessed channel feature data, the target user is authenticated.

[0008] If the result of the legitimacy authentication is that the user is legitimate, a dynamic adversarial game model is established according to the different needs of the target user and the jammer. In this model, the jammer is the active party and the target user is the passive party. During the dynamic adversarial game, the active party first selects a channel strategy and calculates the jamming payoff. The passive party adjusts its own channel according to the active party's channel strategy, obtains the user payoff under different actions, and updates the channel selection probability. The two parties iterate continuously until a dynamic adversarial game equilibrium solution, i.e., the optimal channel strategy set, is formed.

[0009] A deep deterministic policy gradient algorithm is used, with the optimal channel policy set as the state input of the agent and the transmit power of the target user as the action of the agent, to determine the target channel policy and target transmit power that maximize the communication capacity under the maximum malicious interference condition for the target user.

[0010] According to one aspect of the embodiments of this application, a game theory-based access authentication and communication resource control device is provided, comprising:

[0011] The acquisition module is used to acquire and preprocess the channel feature data corresponding to the target user;

[0012] The authentication module is used to authenticate the legitimacy of the target user based on the preprocessed channel feature data;

[0013] The first processing module is used to establish a dynamic adversarial game model based on the different needs of the target user and the jammer when the result of the legality authentication is a legitimate user. In this model, the jammer is the active party and the target user is the passive party. During the dynamic adversarial game, the active party first selects a channel strategy and calculates the jamming payoff. The passive party adjusts its own channel according to the active party's channel strategy, obtains the user payoff under different actions, and updates the channel selection probability. The two parties iterate continuously until a dynamic adversarial game equilibrium solution, i.e., the optimal channel strategy set, is formed.

[0014] The second processing module is used to employ a deep deterministic policy gradient algorithm, taking the optimal channel policy set as the state input of the agent and the transmit power of the target user as the action of the agent, to determine the target channel policy and target transmit power corresponding to the target user that maximizes the communication capacity under the maximum malicious interference condition.

[0015] According to one aspect of the embodiments of this application, a computer-readable medium is provided having a computer program stored thereon, which, when executed by a processor, implements the game theory-based access authentication and communication resource control method as described in the above embodiments.

[0016] According to one aspect of the embodiments of this application, an electronic device is provided, including: one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the game theory-based access authentication and communication resource control method as described in the above embodiments.

[0017] According to one aspect of the embodiments of this application, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the game theory-based access authentication and communication resource control method provided in the above embodiments.

[0018] In some embodiments of this application, the technical solutions provide that by acquiring and preprocessing channel feature data corresponding to the target user, and performing legitimacy authentication on the target user based on the preprocessed channel feature data, the security of user access to the network is improved. Furthermore, when the legitimacy authentication result is that the user is legitimate, a dynamic adversarial game model is established according to the different needs of the target user and the jammer. In this model, the jammer is the active party and the target user is the passive party. During the adversarial game, the active party first selects a channel strategy and calculates the interference payoff. The passive party adjusts its own channel according to the active party's channel strategy, obtains the user payoff under different actions, and updates the channel selection probability. The two parties iterate continuously until a dynamic adversarial game equilibrium solution, i.e., the optimal channel strategy set, is formed. A deep deterministic policy gradient algorithm is adopted, using the optimal channel strategy set as the state input of the agent and the transmit power of the target user as the action of the agent. This determines the target channel strategy and target transmit power corresponding to the target user that maximizes the communication capacity under the maximum malicious interference condition, thereby ensuring the effective utilization of communication resources even in the presence of malicious interference.

[0019] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0020] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:

[0021] Figure 1 The diagram illustrates an application scenario where the technical solutions of the embodiments of this application can be applied.

[0022] Figure 2 A flowchart illustrating a game theory-based access authentication and communication resource control method according to an embodiment of this application is shown.

[0023] Figure 3A block diagram of a game theory-based access authentication and communication resource control device according to an embodiment of this application is shown;

[0024] Figure 4 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown. Detailed Implementation

[0025] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art.

[0026] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.

[0027] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0028] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0029] Figure 1 The illustration shows an application scenario diagram where the technical solutions of the embodiments of this application can be applied.

[0030] like Figure 1As shown, this application scenario can include legitimate users, illegal users, base stations (BS), and jammers. In this scenario, illegal users may impersonate legitimate users to access the network and steal data, thus threatening network security. Furthermore, jammers can interfere with base station communications, affecting the interaction between the base station and legitimate users.

[0031] The technical solutions of this application embodiment can be applied to base stations, thereby improving the security of user access to the network and ensuring the effective utilization of communication resources when communication interference exists.

[0032] Specifically, Figure 2 A flowchart illustrating a game theory-based access authentication and communication resource control method according to an embodiment of this application is shown.

[0033] like Figure 2 As shown, the method includes at least steps S210 to S240, which are described in detail below:

[0034] In step S210, channel feature data corresponding to the target user is acquired and preprocessed.

[0035] In this embodiment, when the base station receives a signal sent by the target user, it can collect channel feature data corresponding to the received signal. This channel feature data can be used to describe the characteristic information of the transmission channel from the target user to the base station. It should be noted that the target user can be a new user or a previously disconnected user; there are no special limitations on this.

[0036] After acquiring the corresponding channel feature data, the base station can perform corresponding preprocessing on the channel feature data. Specifically, the base station can first perform data denoising on the collected channel feature data, and then extract features from it to facilitate subsequent legality authentication.

[0037] In one example, the channel characteristic data could be channel state information (CSI). It should be understood that CSI is a series of parameters used in a wireless communication system to describe the characteristics of a wireless channel. These parameters reflect the path loss, multipath effects, shadowing effects, interference, and other influences experienced by the signal as it travels from the transmitter to the receiver. Using CSI for authentication can effectively identify legitimate and illegitimate users, thereby improving the security and reliability of wireless communication networks.

[0038] In step S220, the target user is authenticated based on the preprocessed channel feature data.

[0039] In one embodiment, after the channel feature data is preprocessed, the base station can compare it with the channel feature data of known legitimate users to determine the similarity between the two. When the similarity reaches a certain threshold, the target user can be confirmed as a legitimate user. If the similarity is low, it indicates that the target user is an illegitimate user, and the base station can refuse its access.

[0040] In one embodiment of this application, the legitimacy authentication of the target user is performed based on the preprocessed channel feature data, including:

[0041] The preprocessed channel state data is input into a pre-trained convolutional neural network, so that the convolutional neural network outputs a judgment result on whether the target user is a legitimate user based on the preprocessed channel state data. The convolutional neural network is trained from a sample dataset containing channel state data of legitimate users and illegitimate users.

[0042] In this embodiment, the base station can pre-build and train a convolutional neural network (CNN) for authenticating the channel state data of users. It should be understood that a CNN is a machine learning algorithm that includes three common structures: convolutional layers, pooling layers, and fully connected layers. It excels at extracting data features and has strong representational capabilities. During the training of this CNN, a sample dataset containing channel state data from both legitimate and illegitimate users can be used for training.

[0043] In one example, CSI samples of 4000 combined users and 4000 illegal users can be collected separately. These samples are then divided into two parts: "Sample Set A" and "Sample Set B". Each sample set contains CSI samples of 2000 combined users and CSI samples of 2000 illegal users. Each CSI sample is 128 in length.

[0044] In the first stage, 75% of the samples from "Sample Set A" are selected as the training set to train the parameters of the convolutional neural network, and the remaining 25% of the samples are used as the test set. In the second stage, the network trained in the first stage is used to identify all samples in "Sample Set B". In this application, the classification threshold η is controlled to be 5 × 10. -1 This allows for better training of convolutional neural networks, ensuring their accuracy in recognition.

[0045] In one embodiment, to better evaluate the effectiveness of the convolutional neural network channel authentication, this application further defines the false alarm probability and the false alarm probability. Specifically, the false alarm probability is the probability that the system incorrectly authenticates a legitimate user as an illegitimate user; the false alarm probability is the probability that the system fails to correctly detect an illegitimate user. The lower these two probabilities are, the better the algorithm's recognition performance. Therefore, these two probabilities can be used to evaluate the convolutional neural network, ensuring its reliability when deployed online.

[0046] Once the convolutional neural network (CNN) is trained, the base station can use it to identify the preprocessed channel state data corresponding to the target user. Specifically, the preprocessed channel state data can be input into the CNN, causing it to output a judgment result on whether the target user is legitimate (i.e., to perform legitimacy authentication on the target user). Through CNN authentication technology based on channel feature information, legitimate and illegitimate users can be effectively identified, ensuring the effectiveness and reliability of communication.

[0047] Please continue to refer to this. Figure 2 In step S230, if the result of the legitimacy authentication is a legitimate user, a dynamic adversarial game model is established according to the different needs of the target user and the jammer. In this model, the jammer is the active party and the target user is the passive party. During the dynamic adversarial game, the active party first selects a channel strategy and calculates the jamming payoff. The passive party adjusts its own channel according to the active party's channel strategy, obtains the user payoff under different actions, and updates the channel selection probability. The two parties iterate continuously until a dynamic adversarial game equilibrium solution, i.e., the optimal channel strategy set, is formed.

[0048] In this embodiment, when the target user is determined to be a legitimate user, the base station can establish a dynamic adversarial game model based on the different needs of the target user and the jammer, thereby optimizing the channel strategy for the target user. It should be understood that a channel strategy can be a series of rules and methods adopted in a wireless communication system for the use and management of the channel in order to optimize communication performance and resource utilization. These strategies involve how to effectively allocate and use wireless spectrum resources between users and base stations, and how to adjust transmission parameters under different communication conditions to improve signal quality and system capacity.

[0049] In dynamic adversarial game, the jammer acts as the active party and the target user acts as the passive party. The active party first selects a channel strategy and calculates the jamming payoff. The passive party adjusts its own channel according to the active party's channel strategy, obtains the user payoff under different actions, and updates the channel selection probability. The two iterate continuously until a dynamic adversarial game equilibrium solution is formed, which is the optimal channel strategy set.

[0050] In one embodiment, the user set in the system is initialized. Base station set jammer set and user channel set Channel set of the jammer Given the information, the user's maximum transmit power and the jammer's maximum transmit power Set the user's initial channel selection probability and the initial channel selection probability of the jammer

[0051] In one example, suppose the number of time slots per round for the jammer in the system, L, and the number of time slots per round for the user, T, are both 100. Then, time slots 1 ≤ l ≤ 100 and 1 ≤ t ≤ 100. The communication system has N = 20 users, randomly distributed on a 1000 × 1000 plane. Some users are legitimate, and some are illegitimate. The maximum transmit power for each user is... The number of jammer transmitters is J=2 and their positions are fixed. The jammer transmission power is... The number of base stations K = 3, and each base station provides communication services to users in the terrestrial cell. Assume the number of user channels and the number of jammer channels in the system are equal, i.e., Z = C = 6. Initialize the user channel selection probability. and jammer channel selection probability

[0052] In the system, users are randomly distributed across a plane, while the locations of the jammer and base station are fixed. In time slot t, user n passes through the channel... Data is sent to base station k. This is because neighboring user m has also accessed the channel. Therefore, co-channel interference is generated at base station k. Where P m Let ω be the transmit power of user m. m,k Let δ(x,y) be the channel gain between user m and base station k, and let δ(x,y) be the Kronecker function. Let d be the set of neighboring users of user n. th It is the farthest distance at which co-channel interference can occur.

[0053] Jammer j continuously transmits interference power P to base station k j To generate malicious interference attacks Where, ω j,k It is the channel gain between the jammer and base station k;

[0054] The total interference received by base station k is Where σ 2 It is Gaussian white noise;

[0055] Then the communication capacity of target user n for:

[0056]

[0057] Among them, P n Let ω be the transmit power of the target user n. n,k B is the channel gain between target user n and base station k, and B is the system bandwidth.

[0058] At this point, the optimization objective for target user n is to obtain the optimal channel strategy with the lowest co-channel interference and malicious interference. This improves communication capacity. Therefore, the user optimization problem can be formulated as follows:

[0059]

[0060] in, The channel selected by target user n in time slot t; The channel selected for users other than the target user n in time slot t.

[0061] The goal of a jammer is to maximize the jamming attack effect; therefore, the optimization problem of a jammer can be described as follows:

[0062]

[0063] In time slot l of the dynamic adversarial game, the jammer, as the active party, first determines the initial probability... Select Channel Calculate the jammer's benefit in, The channel selected by target user n in time slot t; The channel selected for jammer j; Let n be the communication capacity for the target user n; P is the transmission cost factor of the jammer j; j Let be the interference power of jammer j.

[0064] After the jammer selects a channel, the passive party updates the channel selection probability based on the gains obtained under different actions. The user's benefit is maximized on the channel with probability 1 within time slot T, while the probability of the other channels is 0. Let n represent the user benefit for target user n, where n is the user benefit for target user n. The channel selected by target user n in time slot t; The channel selected by users other than the target user n in time slot t; It is the transmission cost factor for the target user n; P n Let n be the transmit power of the target user. Let n be the communication capacity for the target user n; Let n be the set of neighboring users. Let m be the communication capacity of the neighboring user.

[0065] Next, the jammer updates the channel selection probability. The user and the jammer iterate continuously until a dynamic adversarial game equilibrium solution is reached, which is the optimal channel strategy set.

[0066] Please continue to refer to this. Figure 2 In step S240, a deep deterministic policy gradient algorithm is used to determine the target channel policy and target transmission power corresponding to the target user that maximizes the communication capacity under the maximum malicious interference condition. The optimal channel policy set is used as the state input of the agent, and the transmission power of the target user is used as the action of the agent.

[0067] In this embodiment, a method of deep reinforcement learning guided by dynamic adversarial game theory is adopted, with the goal of optimizing the user's channel selection strategy and transmission power to improve the target user's communication capacity.

[0068] Specifically, in power control, the Deep Deterministic Policy Gradient Algorithm (DDPG algorithm) is employed, where the agent's state is derived from the optimal channel strategy obtained through dynamic adversarial game theory. The action is the target user's transmit power, with the aim of increasing their objective function value under maximum interference attack conditions.

[0069] Therefore, the power optimization problem for the target user can be expressed as:

[0070]

[0071] At the same time, power control conditions are specified:

[0072] Condition 1:

[0073] Condition 2:

[0074] Condition 3:

[0075] Condition 4:

[0076] Condition 5:

[0077] Where, γ min This is the minimum communication capacity requirement for users. E represents the energy consumed by the user during communication. max Indicates the upper limit of energy. This represents the communication time between user n and base station k, where Φ represents the size of the transmitted data in bits. This is the user's maximum transmit power.

[0078] In one example, in the DDPG algorithm, the Actor network and the Critic network each have two hidden layers with 128 and 200 nodes respectively. The learning rates for the Actor and Critic networks are 5e-4 and 1e-3, respectively. The reward discount factor is 0.95. Each network consists of two identical neural networks: an online Critic network, an online Actor network, a target Critic network, and a target Actor network, with network parameters θ. Q θ μ θ Q′ and θ μ′ The experience replay buffer size is 10. 5 .

[0079] The agent obtains its current state based on dynamic adversarial game theory. And input it into the network, by Obtain and execute the action in the current state. This indicates the exploration of noise to better facilitate strategy exploration.

[0080] Receive immediate rewards from the environment based on the actions performed.

[0081]

[0082] Where, γ min This is the minimum communication capacity requirement; Let n be the communication capacity for the target user n; It is the transmission cost factor for the target user n; P n Let n be the transmit power of the target user. It is the set of neighboring users of the target user n. Let m be the communication capacity of the neighboring user.

[0083] Therefore, based on the rewards corresponding to different actions of the intelligent agent, the target transmission power corresponding to the target user can be optimized to maximize the communication capacity under the condition of maximum malicious interference.

[0084] In one embodiment, the method further includes:

[0085] The state, action, reward, and next state of each execution of the intelligent agent are stored as experience data in the experience replay buffer.

[0086] Select a portion of the experience data from the experience playback buffer;

[0087] The target value is calculated using a target value network based on the selected empirical data.

[0088] Based on the target value, the weights of the Critic network are updated using the mean squared error loss function method, and the parameters of the Actor network are updated using the policy gradient algorithm. The target network parameters are then softly updated.

[0089] In this embodiment, the agent observes the reward returned by the environment. and the next state s t+1 and will The data is stored in the experience playback buffer for subsequent experience playback. The base station can randomly extract small batches of experience data of size x from the experience playback buffer during idle periods or predetermined time periods.

[0090] Based on the selected empirical data, the target value is calculated using a target value network:

[0091]

[0092] Where Q′ is the target Critic network, μ′ is the target Actor network, and λ is the discount factor.

[0093] Based on the calculated target value, the Critic network weights are updated using the mean squared error loss function method:

[0094]

[0095] After a certain number of updates, the Actor network parameters are updated using the policy gradient algorithm:

[0096]

[0097] Where Q is the online Critic network and μ is the online Actor network.

[0098] For the target network parameters θ Q′ and θ μ′ Perform a soft update:

[0099]

[0100] Where τ is the soft update factor, which is 0.99.

[0101] The above steps can be repeated until the algorithm converges.

[0102] Based on the above embodiments, the DDPG algorithm can help agents learn the optimal channel strategy and transmit power in a dynamically changing communication environment, so as to maximize the communication capacity of the target user and resist interference attacks.

[0103] The following describes an apparatus embodiment of this application, which can be used to execute the game theory-based access authentication and communication resource control method in the above embodiments of this application. For details not disclosed in the apparatus embodiments of this application, please refer to the above embodiments of the game theory-based access authentication and communication resource control method of this application.

[0104] Figure 3 A block diagram of a game theory-based access authentication and communication resource control device according to an embodiment of this application is shown.

[0105] Reference Figure 3 As shown, a game theory-based access authentication and communication resource control device according to an embodiment of this application includes:

[0106] The acquisition module is used to acquire and preprocess the channel feature data corresponding to the target user;

[0107] The authentication module is used to authenticate the legitimacy of the target user based on the preprocessed channel feature data;

[0108] The first processing module is used to establish a dynamic adversarial game model based on the different needs of the target user and the jammer when the result of the legality authentication is a legitimate user. In this model, the jammer is the active party and the target user is the passive party. During the dynamic adversarial game, the active party first selects a channel strategy and calculates the jamming payoff. The passive party adjusts its own channel according to the active party's channel strategy, obtains the user payoff under different actions, and updates the channel selection probability. The two parties iterate continuously until a dynamic adversarial game equilibrium solution, i.e., the optimal channel strategy set, is formed.

[0109] The second processing module is used to employ a deep deterministic policy gradient algorithm, taking the optimal channel policy set as the state input of the agent and the transmit power of the target user as the action of the agent, to determine the target channel policy and target transmit power corresponding to the target user that maximizes the communication capacity under the maximum malicious interference condition.

[0110] In one embodiment, the second processing module is further configured to: store the state, action, reward, and next state corresponding to each execution of the agent as experience data in an experience replay buffer; select a portion of experience data from the experience replay buffer; calculate a target value using a target value network based on the selected experience data; update the Critic network weights using a mean squared error loss function method and update the Actor network parameters using a policy gradient algorithm based on the target value, and perform a soft update on the target network parameters.

[0111] Figure 4 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown.

[0112] It should be noted that, Figure 4 The computer system of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0113] like Figure 4 As shown, the computer system includes a Central Processing Unit (CPU) 401, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 402 or programs loaded from storage portion 408 into Random Access Memory (RAM) 403, such as performing the methods described in the above embodiments. The RAM 403 also stores various programs and data required for system operation. The CPU 401, ROM 402, and RAM 403 are interconnected via a bus 404. An Input / Output (I / O) interface 405 is also connected to the bus 404.

[0114] The following components are connected to I / O interface 405: an input section 406 including a keyboard, mouse, etc.; an output section 407 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to I / O interface 405 as needed. A removable medium 411, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 410 as needed so that computer programs read from it can be installed into storage section 408 as needed.

[0115] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 409, and / or installed from removable medium 411. When the computer program is executed by central processing unit (CPU) 401, it performs various functions defined in the system of this application.

[0116] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0117] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0118] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.

[0119] In another aspect, this application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods described in the above embodiments.

[0120] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0121] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the method according to the embodiments of this application.

[0122] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.

[0123] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A game theory-based access authentication and communication resource control method, characterized in that, include: Acquire and preprocess channel feature data corresponding to the target user; Based on the preprocessed channel feature data, the target user is authenticated. If the result of the legitimacy authentication is that the user is legitimate, a dynamic adversarial game model is established according to the different needs of the target user and the jammer. In this model, the jammer is the active party and the target user is the passive party. During the dynamic adversarial game, the active party first selects a channel strategy and calculates the jamming payoff. The passive party adjusts its own channel according to the active party's channel strategy, obtains the user payoff under different actions, and updates the channel selection probability. The two parties iterate continuously until a dynamic adversarial game equilibrium solution, i.e., the optimal channel strategy set, is formed. A deep deterministic policy gradient algorithm is used, with the optimal channel policy set as the state input of the agent and the transmit power of the target user as the action of the agent, to determine the target channel policy and target transmit power corresponding to the target user that maximizes the communication capacity under the maximum malicious interference condition; The interference benefit of the jammer is calculated according to the following formula: in, The channel selected by target user n in time slot t; The channel selected for jammer j; Let n be the communication capacity for the target user n; P is the transmission cost factor of the jammer j; j Let j be the interference power of the jammer. The user benefit for the target user is calculated using the following formula: in, The channel selected by users other than the target user n in time slot t; It is the transmission cost factor for the target user n; P n Let n be the transmit power of the target user. Let n be the set of neighboring users. Let m be the communication capacity of the neighboring user. Wherein, assume that target user n passes through channel in time slot t. Data is sent to base station k, and neighboring user m also accesses the channel. The co-channel interference generated on base station k Among them, P m Let ω be the transmit power of neighboring user m. m,k δ(x,y) is the channel gain between neighboring user m and base station k, and δ(x,y) is the Kronecker function. d is the set of neighboring users of the target user n. th It is the farthest distance at which co-channel interference can occur; Jammer j continuously transmits interference power P to base station k j To generate malicious interference attacks Where, ω j,k It is the channel gain between the jammer and base station k; The total interference received by base station k is Where σ 2 It is Gaussian white noise; Then the communication capacity of target user n for: Among them, P n Let ω be the transmit power of the target user n. n,k B is the channel gain between target user n and base station k, and B is the system bandwidth.

2. The method according to claim 1, characterized in that, The channel feature data is channel state data; Based on the preprocessed channel feature data, the target user is authenticated, including: The preprocessed channel state data is input into a pre-trained convolutional neural network, so that the convolutional neural network outputs a judgment result on whether the target user is a legitimate user based on the preprocessed channel state data. The convolutional neural network is trained from a sample dataset containing channel state data of legitimate users and illegitimate users.

3. The method according to claim 1, characterized in that, The agent obtains its current state based on dynamic adversarial game theory. Input it into the network, by Obtain and execute the action in the current state, where, Indicates exploratory noise; Get instant rewards from the environment Where, γ min This is the minimum communication capacity requirement for the target user n; Let n be the communication capacity for the target user n; It is the transmission cost factor for the target user n; P n Let n be the transmit power of the target user. It is the set of neighboring users of the target user n. Let m be the communication capacity of the neighboring user. Based on the rewards corresponding to different actions of the intelligent agent, the target channel strategy and target transmit power corresponding to the target user that maximizes the communication capacity under the condition of maximum malicious interference are determined.

4. The method according to claim 3, characterized in that, The method further includes: The state, action, reward, and next state of each execution of the intelligent agent are stored as experience data in the experience replay buffer. Select a portion of the experience data from the experience playback buffer; The target value is calculated using a target value network based on the selected empirical data. Based on the target value, the weights of the Critic network are updated using the mean squared error loss function method, and the parameters of the Actor network are updated using the policy gradient algorithm. The target network parameters are then softly updated.

5. A game theory-based access authentication and communication resource control device, characterized in that, include: The acquisition module is used to acquire and preprocess the channel feature data corresponding to the target user; The authentication module is used to authenticate the legitimacy of the target user based on the preprocessed channel feature data; The first processing module is used to establish a dynamic adversarial game model based on the different needs of the target user and the jammer when the result of the legality authentication is a legitimate user. In this model, the jammer is the active party and the target user is the passive party. During the dynamic adversarial game, the active party first selects a channel strategy and calculates the jamming payoff. The passive party adjusts its own channel according to the active party's channel strategy, obtains the user payoff under different actions, and updates the channel selection probability. The two parties iterate continuously until a dynamic adversarial game equilibrium solution, i.e., the optimal channel strategy set, is formed. The second processing module is used to employ a deep deterministic policy gradient algorithm, taking the optimal channel policy set as the state input of the agent and the transmit power of the target user as the action of the agent, to determine the target channel policy and target transmit power corresponding to the target user that maximizes the communication capacity under the maximum malicious interference condition. The interference benefit of the jammer is calculated according to the following formula: in, The channel selected by target user n in time slot t; The channel selected for jammer j; Let n be the communication capacity for the target user n; P is the transmission cost factor of the jammer j; j Let j be the interference power of the jammer. The user benefit for the target user is calculated using the following formula: in, The channel selected by target user n in time slot t; The channel selected by users other than the target user n in time slot t; It is the transmission cost factor for the target user n; P n Let n be the transmit power of the target user. Let n be the communication capacity for the target user n; Let n be the set of neighboring users. Let m be the communication capacity of the neighboring user. Wherein, assume that target user n passes through channel in time slot t. Data is sent to base station k, and neighboring user m also accesses the channel. The co-channel interference generated on base station k Among them, P m Let ω be the transmit power of neighboring user m. m,k δ(x,y) is the channel gain between neighboring user m and base station k, and δ(x,y) is the Kronecker function. d is the set of neighboring users of the target user n. th It is the farthest distance at which co-channel interference can occur; Jammer j continuously transmits interference power P to base station k j To generate malicious interference attacks Where, ω j,k It is the channel gain between the jammer and base station k; The total interference received by base station k is Where σ 2 It is Gaussian white noise; Then the communication capacity of target user n for: Among them, P n Let ω be the transmit power of the target user n. n,k B is the channel gain between target user n and base station k, and B is the system bandwidth.

6. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the game theory-based access authentication and communication resource control method as described in any one of claims 1 to 4.

7. An electronic device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the game theory-based access authentication and communication resource control method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Communication power control method based on interference countering game

    CN111555838A

  • Anti-intelligent interference channel decision-making method based on interference consciousness learning

    CN116896422A