Security transmission scheme selection method and device, computer device, storage medium and program product

By constructing a multi-agent model in the smart grid CPS system and optimizing the selection of relay nodes, the problem of high computational complexity of traditional security measures is solved, and the security and stability of data transmission are improved.

CN116707959BActive Publication Date: 2026-05-22SHENZHEN POWER SUPPLY BUREAU
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN POWER SUPPLY BUREAU
Filing Date
2023-06-29
Publication Date
2026-05-22

Smart Images

  • Figure CN116707959B_ABST
    Figure CN116707959B_ABST
Patent Text Reader

Abstract

The application relates to a secure transmission scheme selection method and device, computer equipment and a storage medium. The method comprises the following steps: performing transmission rate calculation according to transmission state information of a transmission node, relay state information of a relay node and reception state information of a reception node, and obtaining a transmission rate; performing model training on a to-be-trained multi-agent model according to the transmission rate, and obtaining a target multi-agent model; and determining a target secure transmission scheme according to agent state information of a transmission process and the target multi-agent model. The method can realize strong connectivity of data transmission between a physical network and an information network, and improve the security of data transmission.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of smart grid information security technology, and in particular to a method, apparatus, computer equipment, and storage medium for selecting a secure transmission scheme. Background Technology

[0002] Smart grid cybersecurity systems (CPS) bear the important mission of safeguarding urban economic development and residents' lives. With the rapid development of smart grid CPS systems, network access terminals are becoming increasingly diversified, further exposing these systems to the open environment of the internet. The convergence of physical and information layers also provides new avenues for malicious attacks. For example, attacks on the information transmission process between the physical and information layers could lead to the leakage of confidential information from the smart grid CPS system, thereby threatening the system's operational security. Traditionally, wireless transmission security relies primarily on cryptography to encrypt information at the information layer to prevent leaks or on directly transplanting high-level security protocols from wired systems.

[0003] However, such security measures undoubtedly lead to high computational complexity and energy consumption, and the processes of key distribution, storage, and management also incur significant overhead. On the other hand, existing research on secure transmission schemes for power grids mainly focuses on traditional power grid systems. Traditional power grid security algorithms are not perfectly applicable to the high degree of coupling between physical and information networks in smart grid CPS (Cyber-Physical Systems). Summary of the Invention

[0004] Therefore, it is necessary to provide a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for selecting a secure transmission scheme that can effectively ensure the security of the wireless transmission process and effectively reduce the operating costs of the power grid system, in response to the above-mentioned technical problems.

[0005] Firstly, this application provides a method for selecting a secure transmission scheme. The method includes:

[0006] The transmission rate is calculated based on the transmission status information of the transmitting node, the relay status information of the relay node, and the reception status information of the receiving node.

[0007] The model is trained according to the transmission rate to obtain the target multi-agent model.

[0008] Based on the agent state information during the transmission process and the target multi-agent model, a secure transmission scheme for the target is determined.

[0009] In one embodiment, the transmission rate is calculated based on the transmission status information of the transmitting node, the relay status information of the relay node, and the reception status information of the receiving node to obtain the transmission rate, including:

[0010] Construct a channel gain model between the transmitter and receiver;

[0011] The signal-to-noise ratio (SNR) is calculated based on the transmission status information, relay status information, and reception status information to obtain the first SNR and the second SNR.

[0012] The transmission rate is obtained by performing transmission calculations based on the first signal-to-noise ratio, the second signal-to-noise ratio, and the channel gain model.

[0013] In one embodiment, the transmission rate is obtained by performing transmission calculations based on a first signal-to-noise ratio, a second signal-to-noise ratio, and a channel gain model, including:

[0014] The transmission efficiency is calculated based on the first signal-to-noise ratio and the second signal-to-noise ratio to obtain the codeword transmission rate.

[0015] The eavesdropping efficiency is calculated based on the channel gain model to obtain the eavesdropping transmission rate.

[0016] The secure transmission rate is determined based on the codeword transmission rate and the efficiency of eavesdropping transmission.

[0017] In one embodiment, redundancy calculation is performed based on the codeword transmission rate and the secure transmission rate to obtain the current link redundancy rate;

[0018] Based on the current link redundancy rate and the eavesdropping transmission rate, a transmission security assessment is performed to obtain the security assessment result.

[0019] In one embodiment, model training is performed based on the transmission rate, agent state information during the training process, and the multi-agent model to be trained to obtain a target multi-agent model, including:

[0020] When training the model at the current training moment, decisions are made based on the agent's state information and transmission rate at the current training moment to obtain the agent's action at the current training moment and its corresponding current reward value.

[0021] The multi-agent model to be trained at the current training time is trained based on the current agent's action to obtain the updated model parameters, and the multi-agent model to be trained at the next training time is determined based on the updated model parameters.

[0022] When training the model at the next training time, the agent state information for the next training time is determined based on the current reward value. The multi-agent model to be trained at the next training time is then trained based on the agent state information for the next training time until the model training is completed, and the target multi-agent model is obtained.

[0023] In one embodiment, a target secure transmission scheme is determined based on the agent state information during the transmission process and the target multi-agent model, including:

[0024] After training is completed, decisions are made based on the agent's state information at the current transmission time and the target multi-agent model to obtain the agent's action at the current transmission time and its corresponding reward value at the current transmission time.

[0025] The agent's state information at the current transmission time is updated based on the reward value at the current transmission time to obtain the agent's state information at the next transmission time.

[0026] The secure transmission scheme for the current transmission time is updated based on the agent's state information at the next transmission time to obtain the secure transmission scheme for the next transmission time, and then the secure transmission scheme for the next transmission time is set as the target secure transmission scheme.

[0027] Secondly, this application also provides a secure transmission scheme selection device. The device includes:

[0028] The transmission calculation module is used to calculate the transmission rate based on the transmission status information of the transmitting node, the relay status information of the relay node, and the reception status information of the receiving node.

[0029] The model training module is used to train the model based on the transmission rate, the agent state information during the training process, and the multi-agent model to be trained, so as to obtain the target multi-agent model.

[0030] The scheme selection module is used to make decisions based on the agent state information during the transmission process and the target multi-agent model to determine the target secure transmission scheme.

[0031] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:

[0032] The transmission rate is calculated based on the transmission status information of the transmitting node, the relay status information of the relay node, and the reception status information of the receiving node.

[0033] The model is trained according to the transmission rate to obtain the target multi-agent model.

[0034] Based on the agent state information during the transmission process and the target multi-agent model, a secure transmission scheme for the target is determined.

[0035] Fourthly, this application also provides a computer-readable storage medium. This computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:

[0036] The transmission rate is calculated based on the transmission status information of the transmitting node, the relay status information of the relay node, and the reception status information of the receiving node.

[0037] The model is trained according to the transmission rate to obtain the target multi-agent model.

[0038] Based on the agent state information during the transmission process and the target multi-agent model, a secure transmission scheme for the target is determined.

[0039] Fifthly, this application also provides a computer program product. This computer program product includes a computer program that, when executed by a processor, performs the following steps:

[0040] The transmission rate is calculated based on the transmission status information of the transmitting node, the relay status information of the relay node, and the reception status information of the receiving node.

[0041] The model is trained according to the transmission rate to obtain the target multi-agent model.

[0042] Based on the agent state information during the transmission process and the target multi-agent model, a secure transmission scheme for the target is determined.

[0043] The aforementioned secure transmission scheme selection method, apparatus, computer equipment, storage medium, and computer program products utilize the state information of transmitting nodes, relay nodes, and receiving nodes in the transmission environment. Based on the channel gain of the transmission channel between the transmitting and receiving ends, they calculate the transmission rates during data transmission. Then, they acquire agent state information to train a multi-agent model. After training, the agent state information at the time of transmission is input into the target agent model to obtain action decisions and corresponding reward values. Furthermore, the selection of relay nodes is continuously adjusted based on the agent state information at the time of transmission. This achieves strong connectivity for data transmission between physical and information networks and enhances the security of data transmission. Attached Figure Description

[0044] Figure 1 This is a diagram illustrating the application environment of a secure transmission scheme selection method in one embodiment.

[0045] Figure 2 This is a flowchart illustrating a secure transmission scheme selection method in one embodiment;

[0046] Figure 3 This is a flowchart illustrating the process of determining the transmission rate in one embodiment;

[0047] Figure 4 This is a schematic diagram of the multi-agent training process in one embodiment;

[0048] Figure 5 This is a structural block diagram of a secure transmission scheme selection device in one embodiment;

[0049] Figure 6 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0050] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0051] The secure transmission scheme selection method provided in this application embodiment can be applied to, for example, Figure 1 In the application environment shown, set up as follows: Figure 1 The smart grid shown is a Cyber-Physical System (CPS), which includes a physical network and an information network. Nodes in the physical network connect to physical devices (such as generators and sensors) to acquire real-time device operation data and transmit confidential data to the information network. However, malicious eavesdroppers exist within the system, intending to steal confidential information (such as control commands and device operation data) transmitted between the physical and information networks.

[0052] In this scenario, consider N physical network nodes. These physical network nodes cannot directly connect to the information processing system within the information network; they need to be relayed through intermediate information network nodes to transmit confidential data to the information system. Due to the presence of obstacles in the actual scenario, physical nodes can effectively avoid these obstacles by selecting suitable relay nodes. The scenario also considers M idle information network nodes as potential relay nodes.

[0053] In one embodiment, such as Figure 2 As shown, a method for selecting a secure transmission scheme is provided, which is applied to... Figure 1 Taking the physical network nodes in the example, the following steps are included:

[0054] Step 202: Calculate the transmission rate based on the transmission status information of the transmitting node, the relay status information of the relay node, and the reception status information of the receiving node to obtain the transmission rate.

[0055] Specifically, after determining the transmitting and receiving nodes of the data transmission link and selecting a relay node, the transmission rate from the transmitting node to the relay node and the transmission rate from the relay node to the receiving node are calculated. Furthermore, when the receiving node is an illegitimate intercepting node, the transmission rate from the transmitting node to the intercepting node is calculated.

[0056] Step 204: Train the multi-agent model to be trained according to the transmission rate to obtain the target multi-agent model.

[0057] Specifically, a multi-agent model to be trained is constructed, with physical network nodes acting as agents. These agents interact with the environment to obtain training samples for the multi-agent model. The multi-agent model is then trained using these training samples, and after a preset number of training iterations, the target multi-agent model is obtained.

[0058] Step 206: Determine the target secure transmission scheme based on the agent state information during the transmission process and the target multi-agent model.

[0059] Specifically, during transmission, the agent's state information is input into the target multi-agent model, and reasonable action decisions are made based on the agent's state information, thereby ensuring the security of data transmission.

[0060] In this embodiment, by using the state information of the transmitting node, relay node, and receiving node in the transmission environment, and calculating the transmission rate during data transmission based on the channel gain of the transmission channel between the transmitting and receiving ends, the state information of the agents is obtained to train a multi-agent model. After training is completed, the state information of the agents at the time of transmission is input into the target agent model to obtain action decisions and corresponding reward values. The selection of relay nodes is continuously adjusted according to the state information of the agents at the time of transmission, thereby achieving strong connectivity of data transmission between the physical network and the information network and improving the security of data transmission.

[0061] In one embodiment, such as Figure 3 As shown, the transmission rate is calculated based on the transmission status information of the transmitting node, the relay status information of the relay node, and the reception status information of the receiving node, and includes:

[0062] Step 302: Construct a channel gain model between the transmitter and receiver.

[0063] Specifically, the channel gain model between transmitter X and receiver Y can be expressed by the following formula:

[0064]

[0065] in, This indicates small-scale fading between the transmitter and receiver; This represents the path loss between the transmitter and receiver; The blocking coefficient between the transmitter and receiver is used to describe the characteristics of an obstacle; the antenna gain between the transmitter and receiver can be expressed as... and express.

[0066] Step 304: Calculate the signal-to-noise ratio (SNR) based on the transmission status information, relay status information, and reception status information to obtain the first SNR and the second SNR.

[0067] Specifically, the transmitting node and relay node are set up as the first stage, and the relay node and receiving node are set up as the second stage. After receiving the signal transmitted by the transmitting node, the relay node forwards the signal to the receiving node using an amplify-and-forward (AF) mechanism. The signal-to-noise ratio (SNR) of the transmission process in the first and second stages are then calculated respectively:

[0068]

[0069]

[0070] in, The transmit power of the transmitting node. This represents the amplification gain coefficient of the relay node. The noise figure of the channel. Let S be the channel gain between the transmitting node and the relay node, where S represents the transmitting node, R represents the relay node, and D represents the receiving node.

[0071] Step 306: Perform transmission calculations based on the first signal-to-noise ratio, the second signal-to-noise ratio, and the channel gain model to obtain the transmission rate.

[0072] The transmission rate includes codeword transmission rate, eavesdropping transmission rate, and secure transmission rate.

[0073] Specifically, after obtaining the signal-to-noise ratios (SNRs) of the first and second stages, the transmission efficiency is calculated based on the first and second SNRs to obtain the codeword transmission rate. :

[0074]

[0075] in, The 1 / 2 coefficient is due to the fact that there are two stages of data transmission in the relay-assisted transmission process.

[0076] The eavesdropping efficiency is calculated based on the channel gain model to obtain the eavesdropping transmission rate. :

[0077]

[0078] in, For the channel gain between the transmitting node and the eavesdropper, Antenna gain for the eavesdropper.

[0079] The secure transmission rate is determined based on the codeword transmission rate and the efficiency of eavesdropping transmission; that is, the secure transmission rate between the physical network and the information network is calculated based on Wyner's eavesdropping principle. :

[0080] .

[0081] The secure transmission rate between the k-th physical network node and the information network is then... .

[0082] By selecting appropriate relay nodes, to increase The codeword transmission rate is increased, and the eavesdropping rate of the eavesdropper is reduced. And by utilizing variables This represents the relationship between the nth relay node and the kth physical network node. When... The time indicates that the k-th physical network node selects the n-th relay node for relay forwarding; otherwise... .

[0083] Simultaneously, while calculating the transmission rate, the relay node selection scheme is optimized, taking into account the corresponding constraints, specifically:

[0084]

[0085]

[0086]

[0087]

[0088] Among them, constraint C1 means that at any given time, a relay node can provide data forwarding services to at most one physical network node; constraint C2 means that at any given time, a physical network node can select at most one relay node to provide data forwarding services; and C3 means that the decision variable is a binary indicator variable.

[0089] In this embodiment, a channel gain model is constructed for the data transmission process between the physical network and the information network. Then, the transmission rate between the transmitting node, the relay node, and the receiving node is calculated. Based on the transmission rate, the eavesdropping transmission rate is reduced accordingly, thereby optimizing the selection of relay nodes and improving the security of the transmitted signal.

[0090] In one embodiment, the secure transmission scheme selection method further includes:

[0091] Redundancy calculations are performed based on the codeword transmission rate and the secure transmission rate to obtain the current link redundancy rate; transmission security is assessed based on the current link redundancy rate and the eavesdropping transmission rate to obtain the security assessment result.

[0092] Specifically, two rates need to be determined before data transmission: the transmission codeword rate and the transmission rate. and the secure transmission rate of confidential information ,therefore Redundant coding rate to protect confidential information.

[0093] Then, the redundancy coding rate is compared with the eavesdropping transmission rate. If the redundancy coding rate is less than the eavesdropping transmission rate, it indicates that the signal transmission process may be intercepted, which in turn triggers a security interruption, that is, the secure transmission link between the physical network and the information network is interrupted.

[0094] In this embodiment, by calculating the redundancy coding rate during transmission and comparing it with the eavesdropping transmission rate, a security interruption is triggered when the redundancy coding rate is determined to be less than the eavesdropping transmission rate. This effectively determines that the relay node selection is incorrect, while the other way around indicates that the selection is reasonable, thus ensuring the security and stability of the transmitted signal.

[0095] In one embodiment, such as Figure 4 As shown, the target multi-agent model is obtained by training the model based on the transmission rate, the agent state information during the training process, and the multi-agent model to be trained, including:

[0096] Step 402: When training the model at the current training moment, make decisions based on the agent's state information and transmission rate at the current training moment to obtain the agent's action and its corresponding current reward value at the current training moment.

[0097] The agent state information includes the current node's device state information, channel state information, blocking information, and information about idle relay nodes.

[0098] Specifically, after acquiring the intelligent state information, the agent selects a relay node, which is equivalent to making an action. After the action acts on the environment, the environment changes. Subsequently, the environment provides a specific reward for the agent's action. The reward design for the k-th agent is as follows:

[0099]

[0100] in, Secure transmission rate, For reward design based on a secure transmission rate threshold, the agent's action decision is the selection of relay nodes, i.e., the decision variable. .

[0101] When the secure transmission rate is less than the threshold, that is... ≤ At this time, the secure transmission performance of the node cannot be guaranteed, so its decision-making behavior is subject to necessary penalties, namely:

[0102]

[0103] in, This is the penalty coefficient.

[0104] Then the secure connection probability (SCP) can be obtained as the instantaneous secure rate being greater than the secure rate threshold. The probability, i.e. .

[0105] Step 404: Train the multi-agent model to be trained at the current training time according to the current agent's action, obtain the updated model parameters, and determine the multi-agent model to be trained at the next training time according to the updated model parameters.

[0106] Specifically, a multi-agent model is constructed and trained based on the D3QN deep reinforcement learning algorithm. For the k-th agent, it obtains the state at the current training time. And make action decisions based on the state. When an action is applied to an environment, the environment provides a corresponding reward based on the action. And enter the state of the next training moment. The decision-making process is { , , }

[0107] Step 406: When training the model at the next training time, determine the agent state information at the next training time based on the current reward value, and train the multi-agent model to be trained at the next training time based on the agent state information at the next training time until the model training is completed and the target multi-agent model is obtained.

[0108] Specifically, based on the state at the next training moment Make action decisions When an action is applied to an environment, the environment provides a corresponding reward based on the action. If this process is repeated cyclically, the decision-making process can be represented as: { , ,…,}.

[0109] Based on steps 404 and 406, the objective of model training can be expressed as:

[0110]

[0111] in, The discount factor for model training. For the training network in the D3QN model, These are the parameters of the neural network. The target network in the D3QN model. These are the parameters of the neural network.

[0112] During training, gradient descent is used for model training and updating, and its loss function is... for:

[0113] .

[0114] In this embodiment, a multi-agent model is constructed and trained based on the D3QN deep reinforcement learning algorithm, and then the model parameters, such as the number of iterations and neural network parameters, are continuously adjusted to make the multi-agent model more realistic. This ensures that the decision-making is closer to reality in subsequent transmission moments, and ensures that the action decisions are continuously adjusted during the transmission process, thereby optimizing the performance of secure transmission.

[0115] In one embodiment, a target secure transmission scheme is determined based on the agent state information during transmission and the target multi-agent model, including:

[0116] After training is completed, decisions are made based on the agent's state information at the current transmission time and the target multi-agent model to obtain the agent's action at the current transmission time and its corresponding reward value at the current transmission time. The agent's state information at the current transmission time is updated based on the reward value at the current transmission time to obtain the agent's state information at the next transmission time. The secure transmission scheme at the current transmission time is updated based on the agent's state information at the next transmission time to obtain the secure transmission scheme at the next transmission time, and the secure transmission scheme at the next transmission time is set as the target secure transmission scheme.

[0117] Specifically, each agent observes the environmental state at the current transmission moment during the transmission process and inputs it into the target multi-agent model to make reasonable action decisions based on the environmental state at the current transmission moment. Simultaneously, the reward value corresponding to the action state at the current transmission moment is used to generate the environmental state for the next transmission moment, thereby updating the secure transmission scheme.

[0118] In this embodiment, action decisions are made based on the environmental state during transmission and corresponding reward values ​​are generated. The environmental state in the next transmission process is updated based on the reward values. This enables mutual cooperation without further training, jointly ensuring the security of data transmission between the physical network and the information network in the power grid CPS system.

[0119] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0120] Based on the same inventive concept, this application also provides a secure transmission scheme selection device for implementing the secure transmission scheme selection method described above. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations in one or more secure transmission scheme selection device embodiments provided below can be found in the limitations of the secure transmission scheme selection method described above, and will not be repeated here.

[0121] In one embodiment, such as Figure 5 As shown, a secure transmission scheme selection device is provided, comprising: a transmission calculation module 502, a model training module 504, and a scheme selection module 506, wherein:

[0122] The transmission calculation module 502 is used to calculate the transmission rate based on the transmission status information of the transmitting node, the relay status information of the relay node, and the receiving status information of the receiving node, and to obtain the transmission rate.

[0123] The model training module 504 is used to train the model based on the transmission rate, the agent state information during the training process, and the multi-agent model to be trained, so as to obtain the target multi-agent model.

[0124] The scheme selection module 506 is used to make decisions based on the agent state information during the transmission process and the target multi-agent model to determine the target secure transmission scheme.

[0125] In one embodiment, the transmission calculation module 502 is further configured to construct a channel gain model between the transmitter and the receiver; calculate the signal-to-noise ratio (SNR) based on the transmission status information, relay status information, and receiver status information to obtain a first SNR and a second SNR; and perform transmission calculations based on the first SNR, the second SNR, and the channel gain model to obtain the transmission rate.

[0126] In one embodiment, the transmission calculation module 502 is further configured to calculate the transmission efficiency based on the first signal-to-noise ratio and the second signal-to-noise ratio to obtain the codeword transmission rate; calculate the eavesdropping efficiency based on the channel gain model to obtain the eavesdropping transmission rate; and determine the confidential transmission rate based on the codeword transmission rate and the eavesdropping transmission efficiency.

[0127] In one embodiment, the secure transmission scheme selection device is further configured to perform redundancy calculation based on the codeword transmission rate and the confidential transmission rate to obtain the current link redundancy rate; and to perform transmission security judgment based on the current link redundancy rate and the eavesdropping transmission rate to obtain a security judgment result.

[0128] In one embodiment, the model training module 504 is further configured to, when training the model at the current training moment, make decisions based on the agent's state information and transmission rate at the current training moment to obtain the agent's actions and corresponding current reward values ​​at the current training moment; train the multi-agent model to be trained at the current training moment based on the agent's actions to obtain updated model parameters; determine the multi-agent model to be trained at the next training moment based on the updated model parameters; when training the model at the next training moment, determine the agent's state information at the next training moment based on the current reward value; train the multi-agent model to be trained at the next training moment based on the agent's state information at the next training moment, until the model training is completed and the target multi-agent model is obtained.

[0129] In one embodiment, the scheme selection module 506 is further configured to, after completing training, make decisions based on the agent state information at the current transmission time and the target multi-agent model to obtain the agent action at the current transmission time and its corresponding reward value at the current transmission time; update the agent state information at the current transmission time based on the reward value at the current transmission time to obtain the agent state information at the next transmission time; update the secure transmission scheme at the current transmission time based on the agent state information at the next transmission time to obtain the secure transmission scheme at the next transmission time, and set the secure transmission scheme at the next transmission time as the target secure transmission scheme.

[0130] Each module in the aforementioned secure transmission scheme selection device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0131] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 6As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores confidential data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a secure transmission scheme selection method.

[0132] Those skilled in the art will understand that Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0133] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0134] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0135] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0136] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0137] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0138] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0139] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for selecting a secure transmission scheme, characterized in that, The method includes: Construct a channel gain model between the transmitter and receiver; The signal-to-noise ratio (SNR) is calculated based on the transmission status information, relay status information, and reception status information to obtain the first SNR and the second SNR. The transmission efficiency is calculated based on the first signal-to-noise ratio and the second signal-to-noise ratio to obtain the codeword transmission rate. The eavesdropping efficiency is calculated based on the channel gain model to obtain the eavesdropping transmission rate. The secure transmission rate is determined based on the codeword transmission rate and the eavesdropping transmission rate. The multi-agent model to be trained is trained according to the transmission rate to obtain the target multi-agent model; the transmission rate includes the codeword transmission rate, the eavesdropping transmission rate and the confidential transmission rate; Based on the agent state information during the transmission process and the target multi-agent model, a target secure transmission scheme is determined; The method further includes: Redundancy calculations are performed based on the codeword transmission rate and the secure transmission rate to obtain the current link redundancy rate. Based on the current link redundancy rate and the eavesdropping transmission rate, a transmission security judgment is made to obtain a security judgment result.

2. The method according to claim 1, characterized in that, The step of training the multi-agent model to be trained according to the transmission rate to obtain the target multi-agent model includes: When training the model at the current training moment, a decision is made based on the agent's state information at the current training moment and the transmission rate to obtain the agent's action at the current training moment and its corresponding current reward value. The multi-agent model to be trained at the current training time is trained based on the agent's actions at the current training time to obtain updated model parameters, and the multi-agent model to be trained at the next training time is determined based on the updated model parameters. When training the model at the next training time, the agent state information at the next training time is determined based on the current reward value. The multi-agent model to be trained at the next training time is trained based on the agent state information at the next training time until the model training is completed, and the target multi-agent model is obtained.

3. The method according to claim 1, characterized in that, The step of determining a secure transmission scheme for the target based on the agent state information during transmission and the target multi-agent model includes: After training is completed, decisions are made based on the agent's state information at the current transmission time and the target multi-agent model to obtain the agent's action at the current transmission time and its corresponding reward value at the current transmission time. The agent's state information at the current transmission time is updated based on the reward value at the current transmission time to obtain the agent's state information at the next transmission time. The secure transmission scheme for the current transmission time is updated based on the agent's state information at the next transmission time to obtain the secure transmission scheme for the next transmission time, and then the secure transmission scheme for the next transmission time is set as the target secure transmission scheme.

4. A secure transmission scheme selection device, characterized in that, The device includes: The transmission calculation module is used to construct a channel gain model between the transmitter and receiver; calculate the signal-to-noise ratio (SNR) based on transmission status information, relay status information, and receiver status information to obtain a first SNR and a second SNR; calculate the transmission efficiency based on the first SNR and the second SNR to obtain the codeword transmission rate; calculate the eavesdropping efficiency based on the channel gain model to obtain the eavesdropping transmission rate; and determine the secure transmission rate based on the codeword transmission rate and the eavesdropping transmission rate. The model training module is used to train the model based on the transmission rate, the agent state information during the training process, and the multi-agent model to be trained, to obtain the target multi-agent model; the transmission rate includes the codeword transmission rate, the eavesdropping transmission rate, and the confidential transmission rate; The scheme selection module is used to make decisions based on the agent state information during the transmission process and the target multi-agent model to determine the target secure transmission scheme; The device is also used to perform redundancy calculation based on the codeword transmission rate and the confidential transmission rate to obtain the current link redundancy rate; and to perform transmission security judgment based on the current link redundancy rate and the eavesdropping transmission rate to obtain a security judgment result.

5. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 3.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 3.

7. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 3.