A fuzzy learning anti-interference method and system for dealing with incomplete channel information

By mapping incomplete channel information to fuzzy space, and dynamically adjusting communication power and interference power using Q-learning algorithm and Steinberg game theory, the anti-interference problem under dynamic uncertainty of channel state is solved, and stable solution convergence in a dynamic uncertain environment is achieved.

CN117528538BActive Publication Date: 2025-05-16PLA PEOPLES LIBERATION ARMY OF CHINA STRATEGIC SUPPORT FORCE AEROSPACE ENG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311551774.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-21
Publication Date
2025-05-16
Estimated Expiration
2043-11-21

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with the anti-interference problem caused by dynamic uncertainty of channel state, especially in scenarios where there is no obvious pattern or the channel state changes very quickly, the algorithm may not converge, and the statistical-based optimization method has the problems of large overhead and high time cost in obtaining statistical information.

Method used

By mapping incomplete channel information to fuzzy space, a confrontation process model between users and interference is established using Steinberg game theory, and a Q-learning algorithm is used to create a user Q table and interference Q table, and the communication power and interference power are dynamically adjusted to achieve fuzzy Steinberg equalization.

Benefits of technology

The problem that traditional methods cannot cope with dynamic uncertain environments is solved. The fuzzy Steinberg equilibrium is achieved through fuzzy learning algorithms, which reduces the algorithm's sensitivity to uncertain information and ensures that a stable solution can be obtained in dynamic uncertain environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117528538B_ABST
    Figure CN117528538B_ABST
Patent Text Reader

Abstract

The invention relates to the field of wireless communication technology, and specifically discloses a fuzzy learning anti-interference method and system for coping with incomplete channel information, comprising: mapping the incomplete channel information to a fuzzy space, and then obtaining the fuzzy benefits of a user and interference through a fuzzy benefit function; using the Steinberg game theory to establish a confrontation process model between the user and the interference; creating a Q table of the user and the interference, initializing the communication power of the user and the interference power of the interference; obtaining the fuzzy benefit of the interference, evaluating the fuzzy benefit of the interference by using a function, and obtaining an evaluation value; updating the Q table of the interference, and re-determining the interference power of the interference, and repeatedly calculating the interference power of the interference until the maximum number of iterations; obtaining the fuzzy benefit of the user, and evaluating the fuzzy benefit of the interference by using a satisfaction function, and obtaining an evaluation value; updating the Q table of the user, and re-determining the communication power of the user, and repeatedly calculating the communication power of the user until the maximum number of iterations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of wireless communication technology, and in particular to a fuzzy learning anti-interference method and system for coping with incomplete channel information. Background Art

[0002] With the development of wireless communication networks, malicious interference has become increasingly rampant, posing a huge challenge to the security and reliability of wireless communication systems. Although some achievements have been made in communication anti-interference technology, many of the research objects are non-intelligent jammers, which essentially still belong to the traditional anti-interference category, that is, the interference mode is relatively fixed and the channel environment is relatively stable. There is currently a lack of in-depth research on anti-interference methods under uncertain channel conditions.

[0003] In order to solve the difficulty of anti-interference problem caused by the dynamic uncertainty of channel state, the more common solutions are mainly the following two methods, namely (1) using model-free reinforcement learning (L. Xiao et al., “IoT Security Techniques Based on Machine Learning: How Do IoT Devices Use AI to EnhanceSecurity?,” IEEE Signal Process Mag, vol. 35, no. 5, pp. 41-49, Sept. 2018.), to explore the environment in an unknown environment by trial and error, and then continuously adjust the decision of the intelligent agent so that it finally converges to the optimal strategy. (2) Use statistical tools to collect statistics on the distribution of channel state information, such as channel gain, signal-to-interference-noise ratio, and other parameters, and then perform Bayesian optimization on user utility in a targeted manner (L. Jia et al., “Bayesian Stackelberg Game for Antijamming Transmission With Incomplete Information,” IEEE Commun. Lett., vol. 20, no. 10, pp. 1991-1994, Oct. 2016.). The above optimization process usually obtains clear rewards for actions in a steady-state environment, and then uses clear rewards to optimize action strategies.

[0004] Although these methods can solve the high dynamic and uncertainty problems in interference environments to a certain extent, they still have the following shortcomings:

[0005] (1) The reinforcement learning method performs “decision-feedback-adjustment” in real time. Its essence is to explore the law of interference changes. For scenarios with no obvious law or extremely fast changes in channel status, the algorithm is likely to not converge.

[0006] (2) Optimization based on statistics is essentially about optimizing the long-term returns of users. However, there are practical problems such as high overhead and time cost in obtaining statistical information, and the inability of the fluctuating returns of a single time slot to directly guide decision optimization.

[0007] Therefore, how to provide a fuzzy learning anti-interference method and system for dealing with incomplete channel information is a technical problem that needs to be solved urgently by those skilled in the art. Summary of the invention

[0008] A fuzzy learning anti-interference method for coping with incomplete channel information is provided to achieve the purpose of the present invention, comprising:

[0009] Step S1: Map the incomplete channel information to the fuzzy space and express the channel gain with a fuzzy number, and then obtain the user fuzzy gain and interference fuzzy gain through the fuzzy gain function;

[0010] Step S2: Using the Steinberg game theory, a confrontation process model between users and interference is established;

[0011] Step S3: According to the fuzzy benefit function and the confrontation process model, a user Q table and an interference Q table are respectively created using a Q-learning algorithm to initialize the user's current communication power and the current interference power;

[0012] Step S4: Obtain interference fuzzy benefits, evaluate the interference fuzzy benefits, and obtain an evaluation value;

[0013] Step S5: Update the interference Q table, and re-determine the interference power according to the updated interference Q table, jump to step S4, and repeat the process of steps S4 and S5 until the maximum number of iterations is reached;

[0014] Step S6: Obtain user fuzzy benefits, and evaluate the interference fuzzy benefits using a satisfaction function to obtain an evaluation value;

[0015] Step S7: Update the user Q table, and re-determine the user's communication power according to the updated user Q table, jump to step S4, and repeat the process of steps S4 to S7 until the maximum number of iterations is reached.

[0016] In some specific embodiments, in step S1, the membership function of the channel gain is:

[0017] ;

[0018] in,m The independent variable when the probability density of the channel gain takes the maximum value x The value of l The independent variable when the probability density of the channel gain takes a minimum value x A value of r The independent variable when the probability density of the channel gain takes a minimum value x Another value of l < r .

[0019] In some specific embodiments, in step S1, the fuzzy benefit function includes a user fuzzy benefit function and an interference fuzzy benefit function, and the user fuzzy benefit function is:

[0020] ;

[0021] The interference fuzzy benefit function is:

[0022] ;

[0023] in, represents the user communication power, represents the interference power; is the fuzzy number representation of user channel gain, is the fuzzy number representation of interference channel gain; represents the user's unit power cost, represents the interference unit power cost; is the background noise power.

[0024] In some specific embodiments, the evaluation formula of the fuzzy benefit is:

[0025] ;

[0026] in, is a satisfaction function representing the user's preference for live or disturbed benefits, including:

[0027] optimism: ;

[0028] neutral: ;

[0029] pessimistic: ;

[0030] is the membership function of the user or interference fuzzy benefit.

[0031] In some specific embodiments, the user Q table or interference Q table update formula is:

[0032] ;

[0033] in, Indicates the selection action in the Q table Q value; represents the learning rate; It is the fuzzy benefit evaluation value of the above-mentioned user or interference.

[0034] In some specific embodiments, in step S5 and step S7, the user communication power and the interference power are determined according to the Q table based on the following formula:

[0035] q t ( a ) = exp [ Q ( a ) / τ ] ∑ a ∈  or  exp [ Q ( a ) / τ ] ;

[0036] in, Select power for user or interference in time slot t probability; Select power for user or interference; Selecting actions for the Q table Q value; is the user's action space; is the action space for interference; is the temperature coefficient; as the learning time goes by, .

[0037] To achieve the purpose of the present invention, the present application also provides a fuzzy learning anti-interference system for dealing with incomplete channel information, comprising:

[0038] Information mapping module: used to map incomplete channel information to fuzzy space and express channel gain with fuzzy numbers, and then obtain user fuzzy gain and interference fuzzy gain through fuzzy gain function;

[0039] Model building module: used to build a user-interference confrontation process model using Steinberg game theory;

[0040] Q table creation module: used to create a user Q table and an interference Q table respectively according to the fuzzy benefit function and the confrontation process model using the Q-learning algorithm, and initialize the user's current communication power and the current interference power;

[0041] Interference evaluation module: used for obtaining interference fuzzy benefits, evaluating the interference fuzzy benefits, and obtaining an evaluation value;

[0042] Interference update module: used to update the interference Q table, and re-determine the interference power of the interference according to the updated interference Q table, jump to the interference evaluation module, and repeatedly execute the interference evaluation module and the interference update module until the maximum number of iterations is reached;

[0043] User evaluation module: used to obtain user fuzzy benefits, evaluate the interference fuzzy benefits using a satisfaction function, and obtain an evaluation value;

[0044] User update module: used to update the user Q table, and re-determine the user's communication power according to the updated user Q table, jump to the interference evaluation module, and repeatedly execute the interference evaluation module to the user update module until the maximum number of iterations is reached.

[0045] In some specific embodiments, the fuzzy utility function includes a fuzzy benefit function of the user and a fuzzy benefit function of interference, and the fuzzy benefit function of the user is:

[0046] ;

[0047] The fuzzy benefit function of the interference is:

[0048] ;

[0049] in, represents the user communication power, represents the interference power; is the fuzzy number representation of user channel gain, is the fuzzy number representation of interference channel gain; represents the user's unit power cost, represents the interference unit power cost; is the background noise power.

[0050] In some specific embodiments, the evaluation formula of the fuzzy benefit is:

[0051] ;

[0052] in, is a satisfaction function representing the user's preference for live or disturbed benefits, including:

[0053] optimism: ;

[0054] neutral: ;

[0055] pessimistic: ;

[0056] is the membership function of the user or interference fuzzy benefit.

[0057] In some specific embodiments, the Q table update formula of the user or interference is:

[0058] ;

[0059] in, Indicates the selection action in the Q table Q value; represents the learning rate; It is the fuzzy benefit evaluation value of the above-mentioned user or interference.

[0060] Beneficial effects of the above technical solution:

[0061] The present invention uses fuzzy representation of incomplete channel information to model the dynamic system as an anti-interference fuzzy Stackelberg game model, and then obtains the fuzzy Stackelberg equilibrium through the fuzzy learning algorithm, solving the problem that traditional methods cannot cope with dynamic uncertain environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0063] Figure 1 A flowchart of a fuzzy learning anti-interference method for coping with incomplete channel information provided by an embodiment of the present invention;

[0064] Figure 2 A schematic diagram of the structure of a fuzzy learning anti-interference system for coping with incomplete channel information provided by an embodiment of the present invention;

[0065] Figure 3 A fuzzy learning schematic diagram of a fuzzy learning anti-interference method and system for coping with incomplete channel information provided by an embodiment of the present invention;

[0066] Figure 4 A definite learning schematic diagram of a fuzzy learning anti-interference method and system for coping with incomplete channel information provided by an embodiment of the present invention;

[0067] Figure 5 A schematic diagram showing a comparison between the effects of a fuzzy learning anti-interference method and system for coping with incomplete channel information provided by an embodiment of the present invention and a prior art learning method. DETAILED DESCRIPTION

[0068] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0069] Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar symbols throughout represent the same or similar elements or elements with the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.

[0070] Embodiment 1

[0071] An embodiment of the present invention provides a fuzzy learning anti-interference method for dealing with incomplete channel information, referring to Figure 1 As shown, including:

[0072] Step S1: Map the incomplete channel information to the fuzzy space and express the channel gain with a fuzzy number, and then obtain the user fuzzy gain and interference fuzzy gain through the fuzzy gain function.

[0073] In a specific embodiment of the present invention, in step S1, the membership function of the channel gain is:

[0074] ;

[0075] in, m The independent variable when the probability density of the channel gain takes the maximum value x The value of l The independent variable when the probability density of the channel gain takes a minimum value x A value of r The independent variable when the probability density of the channel gain takes a minimum value x Another value of l < r .

[0076] In a specific embodiment of the present invention, in step S1, the fuzzy benefit function includes a user fuzzy benefit function and an interference fuzzy benefit function, and the user fuzzy benefit function is:

[0077] ;

[0078] The interference fuzzy benefit function is:

[0079] ;

[0080] in, represents the user communication power, represents the interference power; is the fuzzy number representation of user channel gain, is the fuzzy number representation of interference channel gain; represents the user's unit power cost, represents the interference unit power cost; is the background noise power.

[0081] Step S2: Using the Steinberg game theory, a model of the confrontation process between users and interference is established.

[0082] In a specific embodiment of the present invention, the adversarial process model is expressed as ,in Represent users and interference respectively.  = [ p min , p max ] and  = [ j min , j max ] Represent the power selection intervals of user and interference respectively.

[0083] Step S3: According to the fuzzy benefit function and the adversarial process model, a user Q table and an interference Q table are respectively created using a Q-learning algorithm to initialize the user's current communication power and current interference power.

[0084] Step S4: Obtain interference fuzzy gain, evaluate the interference fuzzy gain, and obtain an evaluation value.

[0085] Step S5: Update the interference Q table, and re-determine the interference power according to the updated interference Q table, jump to step S4, and repeat the process of steps S4 and S5 until the maximum number of iterations is reached.

[0086] Step S6: Obtain the user fuzzy benefit, and evaluate the interference fuzzy benefit using a satisfaction function to obtain an evaluation value.

[0087] Step S7: Update the user Q table, and re-determine the user's communication power according to the updated user Q table, jump to step S4, and repeat the process of steps S4 to S7 until the maximum number of iterations is reached.

[0088] In a specific embodiment of the present invention, the Q table update formula of the user or interference is:

[0089] In a specific embodiment of the present invention, the evaluation formula of the fuzzy benefit is:

[0090] ;

[0091] in, is a satisfaction function representing the user's preference for live or disturbed benefits, including:

[0092] optimism: ;

[0093] neutral: ;

[0094] pessimistic: ;

[0095] is the membership function of the user or interference fuzzy benefit.

[0096] ;

[0097] in, Indicates the selection action in the Q table Q value; represents the learning rate; It is the fuzzy benefit evaluation value of the above-mentioned user or interference.

[0098] In a specific embodiment of the present invention, in step S5 and step S7, the user communication power and the interference power are determined according to the Q table based on the following formula:

[0099] q t ( a ) = exp [ Q ( a ) / τ ] ∑ a ∈  or  exp [ Q ( a ) / τ ] ;

[0100] in, Select power for user or interference in time slot t probability; Select power for user or interference; Selecting actions for the Q table Q value; is the user's action space; is the action space for interference; is the temperature coefficient; as the learning time goes by, .

[0101] Reference Figure 3 and Figure 4 As shown in the figure, the user's convergence process is given. It can be seen that the fuzzy learning scheme proposed in the present invention can quickly converge to a stable solution. This is because the present invention reduces the sensitivity of the algorithm to uncertain information and can still obtain the proposed equilibrium in a dynamic uncertain environment. On the contrary, due to the turbulence of the channel gain, the traditional deterministic learning method cannot converge to a stable solution, and its application scenario is limited.

[0102] Reference Figure 5 As shown in Figure 2, the performance comparison of the three learning schemes is given. Compared with other schemes, the fuzzy learning method has the best anti-interference effect, while the effects of deterministic learning and random selection are similar. This is because they cannot converge to a stable solution and the optimization effect is poor.

[0103] Embodiment 2

[0104] An embodiment of the present invention provides a fuzzy learning anti-interference system for dealing with incomplete channel information, referring to Figure 2 As shown, including:

[0105] Information mapping module 10: used to map incomplete channel information to fuzzy space and express channel gain with fuzzy numbers, and then obtain user fuzzy gain and interference fuzzy gain through fuzzy gain function.

[0106] In a specific embodiment of the present invention, in the information mapping module 10, the membership function of the channel gain is:

[0107] ;

[0108] in, m The independent variable when the probability density of the channel gain takes the maximum value x The value of l The independent variable when the probability density of the channel gain takes a minimum value x A value of r The independent variable when the probability density of the channel gain takes a minimum value x Another value of l < r .

[0109] In a specific embodiment of the present invention, the fuzzy benefit function includes a user fuzzy benefit function and an interference fuzzy benefit function, and the user fuzzy benefit function is:

[0110]

[0111] The interference fuzzy benefit function is:

[0112]

[0113] in, represents the user communication power, represents the interference power; is the fuzzy number representation of user channel gain, is the fuzzy number representation of interference channel gain; represents the user's unit power cost, represents the interference unit power cost; is the background noise power.

[0114] Model building module 20: used to build a confrontation process model between users and interference by using Steinberg game theory.

[0115] In a specific embodiment of the present invention, the adversarial process model is expressed as ,in Represent users and interference respectively.  = [ p min , p max ] and  = [ j min , j max ] Represent the power selection intervals of user and interference respectively.

[0116] The Q table creation module 30 is used to create a user Q table and an interference Q table respectively according to the fuzzy benefit function and the adversarial process model using a Q-learning algorithm, and initialize the user's current communication power and current interference power.

[0117] The interference evaluation module 40 is used to obtain the interference fuzzy benefit, evaluate the interference fuzzy benefit, and obtain an evaluation value.

[0118] Interference updating module 50: used to update the interference Q table, and re-determine the interference power of the interference according to the updated interference Q table, jump to the interference evaluation module, and repeatedly execute the interference evaluation module and the interference updating module until the maximum number of iterations is reached.

[0119] The user evaluation module 60 is used to obtain the user fuzzy benefit and evaluate the interference fuzzy benefit using a satisfaction function to obtain an evaluation value.

[0120] User update module 70: used to update the user Q table, and re-determine the user's communication power according to the updated user Q table, jump to the interference evaluation module, and repeatedly execute the interference evaluation module to the user update module until the maximum number of iterations is reached.

[0121] In a specific embodiment of the present invention, the evaluation formula of fuzzy benefit is:

[0122] ;

[0123] in, is a satisfaction function representing the user's preference for live or disturbed benefits, including:

[0124] optimism: ;

[0125] neutral: ;

[0126] pessimistic: ;

[0127] is the membership function of the user or interference fuzzy benefit.

[0128] In a specific embodiment of the present invention, the Q table update formula of the user or interference is:

[0129] ;

[0130] in, Indicates the selection action in the Q table Q value; represents the learning rate; It is the fuzzy benefit evaluation value of the above-mentioned user or interference.

[0131] In a specific embodiment of the present invention, the user communication power and the interference power are determined according to the Q table based on the following formula:

[0132] q t ( a ) = exp [ Q ( a ) / τ ] ∑ a ∈  or  exp [ Q ( a ) / τ ] ;

[0133] in, Select power for user or interference in time slot t probability; Select power for user or interference; Selecting actions for the Q table Q value; is the user's action space; is the action space for interference; is the temperature coefficient; as the learning time goes by, .

[0134] Reference Figure 3 and Figure 4 As shown in the figure, the user's convergence process is given. It can be seen that the fuzzy learning scheme proposed in the present invention can quickly converge to a stable solution. This is because the present invention reduces the sensitivity of the algorithm to uncertain information and can still obtain the proposed equilibrium in a dynamic uncertain environment. On the contrary, due to the turbulence of the channel gain, the traditional deterministic learning method cannot converge to a stable solution, and its application scenario is limited.

[0135] Reference Figure 5 As shown in Figure 2, the performance comparison of the three learning schemes is given. Compared with other schemes, the fuzzy learning method has the best anti-interference effect, while the effects of deterministic learning and random selection are similar. This is because they cannot converge to a stable solution and the optimization effect is poor.

[0136] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

[0137] Each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referenced to each other. The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the functions in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing terminal device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which implements the functions specified in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable terminal device provide for implementing the process in the process. Figure 1 A process or multiple processes and / or boxes Figure 1The steps of the functions specified in one or more boxes. Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the attached claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention. Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or terminal device including a series of elements includes not only those elements, but also other elements that are not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of more restrictions, the elements defined by the sentence "including one..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.

[0138] The method and device provided by the present invention are introduced in detail above. Specific examples are used in this article to illustrate the principle and implementation mode of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation mode and application scope. In summary, the content of this specification should not be understood as limiting the present invention.

[0139] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", "one specific embodiment" or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0140] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A fuzzy learning anti-interference method for dealing with incomplete channel information, characterized in that: include: Step S1: Map the incomplete channel information to the fuzzy space and express the channel gain with a fuzzy number, and then obtain the user fuzzy gain and interference fuzzy gain through the fuzzy gain function; Step S2: Using the Steinberg game theory, a confrontation process model between users and interference is established; Step S3: According to the fuzzy benefit function and the confrontation process model, a user Q table and an interference Q table are respectively created using a Q-learning algorithm to initialize the user's current communication power and the current interference power; Step S4: Obtain interference fuzzy benefits, evaluate the interference fuzzy benefits, and obtain an evaluation value; Step S5: Update the interference Q table, and re-determine the interference power according to the updated interference Q table, jump to step S4, and repeat the process of steps S4 and S5 until the maximum number of iterations is reached; Step S6: Obtain user fuzzy benefits, and evaluate the interference fuzzy benefits using a satisfaction function to obtain an evaluation value; Step S7: Update the user Q table, and re-determine the user's communication power according to the updated user Q table, jump to step S4, and repeat the process of steps S4 to S7 until the maximum number of iterations is reached; In step S1, the fuzzy benefit function includes a user fuzzy benefit function and an interference fuzzy benefit function, and the user fuzzy benefit function is: The interference fuzzy benefit function is: Where p represents the user communication power, j represents the interference power; is the fuzzy number representation of user channel gain, is the fuzzy number representation of interference channel gain; c s represents the user unit power cost, c j represents the interference unit power cost; is the background noise power; In step S1, the membership function of the channel gain is: Wherein, m is the value of the independent variable x when the probability density of the channel gain takes the maximum value; l is a value of the independent variable x when the probability density of the channel gain takes the minimum value, and r is another value of the independent variable x when the probability density of the channel gain takes the minimum value, and l<r; The evaluation formula of the fuzzy benefit is: in, is a satisfaction function representing the user or interference benefit preference, including: optimism: neutral: pessimistic: is the membership function of the user or interference fuzzy benefit, x represents the independent variable of the satisfaction function of the user or interference benefit preference, and y represents the independent variable of the fuzzy benefit membership function.

2. The fuzzy learning anti-interference method for dealing with incomplete channel information according to claim 1, characterized in that: The update formula of user Q table or interference Q table is: Q(a)=(1-λ)Q(a)+λE v Where Q(a) represents the Q value of action a in the Q table; λ∈(0,1) represents the learning rate; E v Fuzzy benefit evaluation value for users or interference.

3. The fuzzy learning anti-interference method for dealing with incomplete channel information according to claim 1, characterized in that: In step S5 and step S7, the user communication power and the interference power are determined according to the Q table based on the following formula: Among them, q t (a) is the probability that the user or interference selects power a in time slot t; a is the selected power of the user or interference; Q(a) is the Q value of selecting action a in the Q table; P is the action space of the user, and J is the action space of the interference; τ is the temperature coefficient, and as the learning time goes by, τ→0.

4. A fuzzy learning anti-interference system for dealing with incomplete channel information, characterized in that: include: Information mapping module: used to map incomplete channel information to fuzzy space and express channel gain with fuzzy numbers, and then obtain user fuzzy gain and interference fuzzy gain through fuzzy gain function; Model building module: used to build a user-interference confrontation process model using Steinberg game theory; Q table creation module: used to create a user Q table and an interference Q table respectively according to the fuzzy benefit function and the confrontation process model using the Q-learning algorithm, and initialize the user's current communication power and the current interference power; Interference evaluation module: used for obtaining interference fuzzy benefits, evaluating the interference fuzzy benefits, and obtaining an evaluation value; Interference update module: used to update the interference Q table, and re-determine the interference power of the interference according to the updated interference Q table, jump to the interference evaluation module, and cyclically call the interference evaluation module and the interference update module until the maximum number of iterations is reached; User evaluation module: used to obtain user fuzzy benefits, evaluate the interference fuzzy benefits using a satisfaction function, and obtain an evaluation value; User update module: used to update the user Q table, and re-determine the user's communication power according to the updated user Q table, jump to the interference evaluation module, and cyclically call the interference evaluation module, interference update module, user evaluation module, and user update module until the maximum number of iterations is reached; The fuzzy benefit function includes a user fuzzy benefit function and an interference fuzzy benefit function. The user fuzzy benefit function is: The interference fuzzy benefit function is: Where p represents the user communication power, j represents the interference power; is the fuzzy number representation of user channel gain, is the fuzzy number representation of interference channel gain; c s represents the user unit power cost, c j represents the interference unit power cost; is the background noise power; The evaluation formula of the fuzzy benefit is: in, is a satisfaction function representing the user's benefit or interference benefit preference, including: optimism: neutral: pessimistic: is the membership function of the user or interference fuzzy benefit, x represents the independent variable of the satisfaction function of the user or interference benefit preference, and y represents the independent variable of the fuzzy benefit membership function.

5. The fuzzy learning anti-interference system for dealing with incomplete channel information according to claim 4, characterized in that: The update formula of user Q table or interference Q table is: Q(a)=(1-λ)Q(a)+λE v Where Q(a) represents the Q value of action a in the Q table; λ∈(0,1) represents the learning rate; E v Fuzzy benefit evaluation value for users or interference.