A network security method for supporting user-side resource participation demand response and related device

By leveraging blockchain to optimize demand response models and secure resource allocation, and dynamically adjusting the number of blocks and the sub-channel allocation of consensus nodes, the problem of the contradiction between consensus throughput and latency in traditional methods is solved, achieving efficient, secure, and stable operation of the demand response system.

CN118764166BActive Publication Date: 2026-01-13ELECTRIC POWER RES INST CHINA SOUTHERN POWER GRID CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410998055.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2026-01-13
Estimated Expiration
2044-07-24

AI Technical Summary

Technical Problem

Traditional blockchain-based demand response information exchange methods have failed to effectively coordinate and optimize consensus throughput and block consensus latency, and lack long-term decentralized constraints, resulting in high trust costs, information opacity, and low transaction efficiency, making it difficult to guarantee the security of demand response.

Method used

By leveraging blockchain to enable demand response models, combined with secure resource allocation optimization models and Lyapunov optimization theory, the number of blocks and the sub-channel allocation of consensus nodes are dynamically adjusted to maximize the weighted average of consensus throughput and block consensus latency. Long-term decentralized constraints are introduced, and adaptive deep learning is used to optimize resource allocation, thereby achieving dynamic balance and security in the consensus process.

Benefits of technology

It improves the processing speed and response efficiency of the demand response system, ensures that the blockchain network maintains a high degree of decentralization in long-term operation, enhances its risk resistance, and guarantees the security and efficiency of demand response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118764166B_ABST
    Figure CN118764166B_ABST
Patent Text Reader

Abstract

The application provides a network security method for supporting user-side resource participation in demand response and a related device, and belongs to the technical field of demand response in the power industry. The method provides a decentralized trust basis for the demand response process through the application of blockchain technology. The introduction of a secure resource allocation optimization model enables the system to more efficiently manage resources during the consensus process. By optimizing the allocation of the number of blocks and sub-channels of block consensus nodes in each time period, the consensus throughput is maximized and the block consensus delay is reduced, thereby improving the processing speed and response efficiency of the entire demand response system and achieving a dynamic balance between the two contradictory indicators. The setting of long-term decentralized constraints ensures that the blockchain network can still maintain a high degree of decentralization during scale expansion and long-term operation. The long-term risk resistance capability in the block consensus process is ensured, and the security of demand response is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of demand response technology in the power industry, specifically relating to a network security method and related apparatus that supports user-side resources in participating in demand response. Background Technology

[0002] With the accelerated construction of new power systems and the continuous integration of high-proportion renewable energy sources into the grid, the intermittent and fluctuating output of distributed renewable energy sources such as solar and wind power, as well as the unregulated charging of large-scale electric vehicles, seriously threaten the reliability of power distribution networks. Demand response incentivizes users to reduce load during peak hours or shift load to off-peak periods to achieve peak shaving and valley filling, ensuring a balance between supply and demand in the distribution network. However, the execution of demand response involves multiple stakeholders, including grid companies, aggregators, and electricity users. Due to the differences between these stakeholders, the exchange of demand response information among multiple stakeholders faces problems such as high trust costs, information opacity, and low transaction efficiency, posing a significant challenge to the secure exchange of demand response information.

[0003] Blockchain technology, with its advantages of distributed ledger, no need for a trusted third party, and immutability, provides a solution for trusted interaction in multi-entity demand response. Leveraging blockchain technology to empower demand response systems can effectively reduce trust costs among diverse, large-scale market participants with varying electricity consumption habits. However, traditional blockchain-based demand response information interaction methods still face the following challenges. First, traditional methods do not consider the inherent contradiction between consensus throughput and block consensus latency; optimizing one metric will worsen the other, making them unsuitable for synergistic optimization and hindering the achievement of a dynamic balance between conflicting metrics. Second, traditional methods do not consider long-term decentralization constraints, making it difficult to guarantee long-term resilience during the block consensus process and compromising the security of demand response. Therefore, a cybersecurity technology that supports user-side resource participation in demand response is urgently needed. Summary of the Invention

[0004] In view of this, the present invention aims to provide a network security method and related apparatus that supports user-side resources in participating in demand response, so as to solve the above-mentioned problems existing in the secure interaction of demand response information in traditional methods.

[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0006] In a first aspect, the present invention provides a network security method for supporting user-side resources to participate in demand response, comprising the following steps:

[0007] The system obtains user demand response behavior plans and achieves consensus based on the blockchain-enabled demand response model. The blockchain-enabled demand response model is used to reach consensus on user demand response behavior plans according to a preset consensus algorithm.

[0008] The security resource allocation optimization model is used to solve the number of blocks and sub-channel allocation decisions of block consensus nodes in each time period consensus process. Under the constraints of sub-channel allocation and long-term decentralization, the security resource allocation optimization model maximizes the weighted average of consensus throughput and the reciprocal of block consensus latency by optimizing the number of blocks and sub-channel allocation decisions of block consensus nodes in each time period consensus process. The long-term decentralization constraint is used to ensure that the blockchain-enabled demand response model maintains a high degree of decentralization that meets long-term requirements, so as to ensure the security of user-side resources participating in demand response.

[0009] The consensus process is optimized based on the number of blocks solved and the sub-channel allocation decision of the block consensus nodes to achieve consensus, thereby supporting user-side resources to participate in demand response based on the user demand response behavior scheme that has reached consensus.

[0010] Furthermore, the optimization model for security resource allocation is represented by the first optimization problem, as follows:

[0011]

[0012] In the formula, X(t) represents the block quantity decision vector for time period t, W(t) represents the sub-channel allocation vector for time period t, T is the number of time periods, t is the time period index, Y(t) represents the consensus throughput for time period t, and θ M T represents the weighting coefficient. total (t) represents the block consensus latency in time period t, C1 is the block quantity decision indicator variable, I is the number of block quantity levels, i is the block quantity level index, and x i (t) = 1 indicates the number of blocks of the i-th level selected in time period t, C2 is the sub-channel allocation constraint, and W represents the total system bandwidth. Represented as node S n The number of sub-channels allocated, W0 represents the average bandwidth of the sub-channels, and C3 is the long-term decentralization constraint. This represents the set of blocks up to time period t. The Gini coefficient, G max This represents the upper limit of the maximum tolerable Gini coefficient.

[0013] Furthermore, in solving the optimization model for secure resource allocation, a virtual queue is constructed based on Lyapunov optimization theory. This virtual queue transforms the long-term decentralization constraint into a queue stability constraint, resulting in the second optimization problem, as follows:

[0014]

[0015] In the formula, θ Z Let Z(t) be the weighting coefficient, and Z(t) be the actual Gini coefficient of the system at the end of time period t, and G be the upper limit of the maximum tolerable Gini coefficient.max The deviation.

[0016] Furthermore, the second optimization problem is constructed as a temporal Markov process for solution. The constructed temporal Markov process includes: state space, action space, and reward function.

[0017] The state space includes the signal-to-noise ratio between different edge servers, the computing resources of the edge servers, and the decentralized deficit virtual queue;

[0018] The action space includes decisions on the number of blocks and the allocation of sub-channels;

[0019] The reward function is used to penalize and reduce the reward value when the constraints of the optimization problem are not met, based on the degree of constraint non-compliance.

[0020] Furthermore, the temporal Markov process is solved using an adaptive DQN block and channel resource cooperative allocation optimization method driven by security boundary violation penalties. This cooperative allocation optimization method includes an evaluation network and a target network, and the solution process includes:

[0021] The state information of the current time slot edge layer is input into the evaluation network to obtain the state-action value function of different actions in the current time slot. Based on this, the action with the highest state-action value in the current time slot is selected according to the strategy obtained by the exploration and utilization adaptive adjustment algorithm. The exploration and utilization adaptive adjustment algorithm is used to reduce the penalty during the learning process by using the existing optimal resource management strategy when the security boundary violation penalty in the previous time slot meets the condition of a large value. When the security boundary violation penalty in the previous time slot does not meet the condition of a large value, other optimal resource management strategies are explored to avoid getting trapped in local optima.

[0022] Execute the selected action, calculate the reward value, move to the next state, and store the action chain in the experience pool. The action chain includes the current state, the corresponding action and reward value, and the next state.

[0023] A set of empirical data is randomly selected from the experience pool, and the temporal difference error of the evaluation network in the current time slot is calculated.

[0024] The parameters of the evaluation network are updated based on the temporal difference error, and the target network is synchronized with the evaluation network at set time intervals.

[0025] Repeat the above steps iteratively until the optimization cycle ends.

[0026] Furthermore, the consensus throughput Y(t) and block consensus latency T total (t) is determined using the consensus throughput model and the consensus latency model, respectively, as follows:

[0027] Consensus throughput model:

[0028]

[0029] In the formula, Ξ(k(t)) represents the total number of transactions recorded in a single block during time period t, k(t) represents the number of blocks that the master node needs to produce during time period t, and IB(B(t),k(t)) represents the master node B during time period t. p Transmitted to master node B in the next time period p′ At that time, the number of blocks ignored; k(t)B(t) is the sum of the sizes of all blocks generated by the master node in the t-th time period. Average size per transaction For B in the t-th time period p Able to send to B p′ The amount of data transmitted, B(t) is the size of a single block in the t-th time period;

[0030] Consensus latency model:

[0031]

[0032] In the formula, Let t be the transmission delay in time period t. To prepare for transmission delay during the preparation phase, To reduce transmission delay during the preparation phase, For the transmission delay during the submission phase, The response phase transmission delay is represented by k(t), which is the number of blocks that the master node needs to produce in the t-th time period. Let t be the calculation delay for the t-th time period. Calculate the delay for the request phase. Calculate the delay for the preparation phase. Calculate the delay during the preparation phase. The delay is calculated for the submission phase. Calculate the delay for the response phase.

[0033] Furthermore, the set of block counts up to time period t. Gini coefficient Determined according to the following formula:

[0034]

[0035] In the formula, k(i) and k(j) are the number of blocks that the master node needs to produce in the i-th and j-th time periods, respectively, and N represents the number of nodes that take turns serving as the master node.

[0036] Secondly, the present invention provides a network security system that supports user-side resources in participating in demand response, comprising: a cloud layer, an edge layer, and a physical layer;

[0037] The cloud layer is used to distribute user demand response behavior plans to the edge layer for consensus; it is also used to solve the number of blocks and sub-channel allocation decisions for block consensus nodes in each time period based on the security resource allocation optimization model. The security resource allocation optimization model is used to maximize the weighted average of consensus throughput and the reciprocal of block consensus latency by optimizing the number of blocks and sub-channel allocation decisions for block consensus nodes in each time period under constraints including sub-channel allocation and long-term decentralization. The long-term decentralization constraint is used to ensure that the blockchain-enabled demand response model maintains a high degree of decentralization that meets long-term requirements, so as to ensure the security of user-side resources participating in demand response; it also guides the consensus process based on the solved number of blocks and sub-channel allocation decisions for block consensus nodes.

[0038] The edge layer is used to reach consensus based on the blockchain-enabled demand response model, which is used to reach consensus on user demand response behavior plans based on a preset consensus algorithm; it is also used to distribute the consensus-reached user demand response behavior plans to the physical layer.

[0039] The physical layer is used to execute the response plan for user needs.

[0040] Accordingly, the present invention also provides a computer device, the device including a processor and a memory:

[0041] The memory is used to store computer programs and send the instructions of the computer programs to the processor;

[0042] The processor executes, according to instructions from a computer program, a network security method that supports user-side resource participation in demand response, as described in the first aspect.

[0043] Accordingly, the present invention also provides a computer-readable storage medium, characterized in that a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements a network security method for supporting user-side resource participation in demand response as described in the first aspect.

[0044] In summary, this invention provides a network security method and related apparatus that supports user-side resource participation in demand response. This method, through the application of blockchain technology, provides a decentralized trust foundation for the demand response process. The introduction of a secure resource allocation optimization model enables the system to manage resources more efficiently during the consensus process. By optimizing the allocation of the number of blocks and the sub-channels of block consensus nodes in each time period, consensus throughput is maximized while block consensus latency is reduced, thereby improving the processing speed and response efficiency of the entire demand response system and achieving a dynamic balance between these two conflicting indicators. The long-term decentralization constraint ensures that the blockchain network maintains a high degree of decentralization even during expansion and long-term operation. This guarantees long-term resilience in the block consensus process and also ensures the security of demand response. This invention, by combining blockchain technology and a secure resource allocation optimization model, improves the security and efficiency of the demand response process and supports network security for user-side resource participation in demand response. Attached Figure Description

[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0046] Figure 1 A flowchart illustrating a network security method for supporting user-side resources in demand response, provided by an embodiment of the present invention;

[0047] Figure 2 A diagram of a blockchain-enabled security requirement response system provided in this embodiment of the invention;

[0048] Figure 3 A schematic diagram of a security resource allocation software module based on security boundary violation penalty-driven adaptive deep learning provided in an embodiment of the present invention;

[0049] Figure 4 This is a flowchart of a network security technology that supports user-side resources in demand response, provided in an embodiment of the present invention.

[0050] Figure 5 This is a schematic diagram of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0051] To make the objectives, features, and advantages of this invention more apparent and understandable, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described below are only some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0052] Please see Figure 1 This embodiment provides a network security method for supporting user-side resources to participate in demand response, including the following steps:

[0053] S1: Obtain user demand response behavior plans and reach consensus based on the blockchain-enabled demand response model. The blockchain-enabled demand response model is used to reach consensus on user demand response behavior plans according to a preset consensus algorithm.

[0054] It should be noted that this step aims to enhance the consensus process for user demand response measures through blockchain technology. Specifically, all relevant user demand response measures (such as adjusting electricity usage time, changing electricity consumption, etc.) are submitted to the blockchain network and verified and agreed upon using a pre-defined consensus algorithm (such as Proof of Stake, Proof of Work, etc.).

[0055] This step leverages the immutability and transparency of blockchain to ensure the fairness and transparency of the demand response process, enabling all participants to have a shared understanding and trust in the response plan, thereby laying a solid foundation for subsequent resource allocation and execution.

[0056] S2: Based on the security resource allocation optimization model, solve for the number of blocks and the sub-channel allocation decision of block consensus nodes in each time period consensus process. The security resource allocation optimization model is used to maximize the weighted average of consensus throughput and the reciprocal of block consensus latency by optimizing the number of blocks and the sub-channel allocation decision of block consensus nodes in each time period consensus process under the constraints of sub-channel allocation and long-term decentralization. The long-term decentralization constraint is used to ensure that the blockchain-enabled demand response model maintains a high degree of decentralization that meets long-term requirements, so as to ensure the security of user-side resources participating in demand response.

[0057] It should be noted that, in order to further optimize the consensus process, especially to improve consensus efficiency and decentralization, this step adopts a secure resource allocation optimization model. This model dynamically adjusts the number of blocks and the sub-channel allocation of consensus nodes in each time period, taking into account both sub-channel allocation (ensuring data transmission efficiency) and long-term decentralization constraints (maintaining network decentralization).

[0058] This step maximizes consensus throughput (processing capacity) while minimizing latency. This model ensures the real-time and effective response to demands, while maintaining the stability of the blockchain network and avoiding the risks associated with centralization. It also achieves a dynamic balance between these two conflicting metrics, guaranteeing long-term resilience and security of demand response during the block consensus process.

[0059] S3: Optimize the consensus process based on the number of blocks solved and the sub-channel allocation decision of the block consensus nodes to achieve consensus, thereby supporting user-side resources to participate in demand response based on the user demand response behavior scheme that has reached consensus.

[0060] It should be noted that, based on the optimal number of blocks and sub-channel allocation decisions obtained in the first two steps, this step optimizes the consensus process, speeds up the consensus process, and ensures that the user demand response behavior plan can be effectively executed.

[0061] Through these optimization measures, user-side resources can be adjusted according to the agreed-upon demand response plan, such as reducing electricity consumption during peak hours or increasing consumption when there is a surplus of renewable energy, thereby achieving flexible adjustment of electricity supply and demand and improving the overall efficiency and reliability of the energy system.

[0062] Figure 2 This is a diagram of a blockchain-enabled security requirement response system. Figure 3 This is a schematic diagram of a security resource allocation software module based on adaptive deep learning driven by security boundary violation punishment. Figure 4 This is a flowchart illustrating the cybersecurity technologies that support user-side resources in demand response. The following is combined with... Figure 2-4 Further embodiments of the present invention will be described.

[0063] The method proposed in this invention can be applied to blockchain-enabled security requirement response systems, such as... Figure 1 As shown, the system comprises a physical layer, an edge layer, and a cloud layer. The cloud layer includes the power grid company's demand response cloud platform and the demand response aggregator cloud platform. Demand response operations are jointly performed by the power grid company's demand response cloud platform, the demand response aggregator cloud platform, and power users. The power grid company's demand response cloud platform is responsible for managing and publishing demand response signals. After receiving the demand response signals published by the power grid company's demand response cloud platform, the demand response aggregator cloud platform executes a smart contract to determine the user's demand response behavior plan and sends the response information to the edge layer. The edge layer consists of N edge servers and 5G base stations, and the set of edge servers is represented as... The base station is responsible for providing communication coverage within the area, supporting communication between the physical layer, edge layer, and cloud layer. The edge server is responsible for receiving demand response information from the demand response aggregator cloud platform, distributing demand response requests to electricity users, and maintaining the blockchain platform. The physical layer consists of electricity users with distributed power sources, loads, and storage devices, and is responsible for adjusting their corresponding electricity consumption behavior after receiving demand response requests from the edge layer.

[0064] This invention considers T time periods, whose set is represented as follows: The duration of each time period is τ. At the beginning of each time period, firstly, the power grid company's demand response cloud platform sends a demand response signal to the demand response aggregator cloud platform. Upon receiving the demand response signal, the demand response aggregator cloud platform executes a smart contract to determine the user's demand response behavior plan and distributes it to the edge layer base station. Secondly, each edge server reaches a consensus on the received user demand response behavior plan, and then the edge server acting as the master node packages the consensus-reaching user demand response behavior plan into a new block and uploads it to the blockchain network. The base station distributes the consensus-reaching user demand response behavior plan to the power users at the physical layer, initiating the demand response. Finally, the power users execute the demand response task according to the received demand response requirements, and the demand response aggregator cloud platform pays the participating power users a reward, completing the demand response. The demand response aggregator cloud platform maximizes the weighted sum of consensus throughput and the reciprocal of the total block consensus latency by optimizing the number of blocks and spectrum allocation strategies.

[0065] Practical Byzantine Fault Tolerance (PBFT) is a distributed consensus algorithm that does not require a trusted third party. It enables consistency verification of transactions and data in a decentralized environment, exhibiting high security and fault tolerance. Therefore, in the proposed architecture, each edge server uses PBFT to reach consensus on user demand response behavior schemes issued by the demand response aggregator cloud platform. This decentralized management approach supports the power system's demand response operations, ensuring the authenticity and consistency of information within the demand response system while improving system security and reliability. In time period t, the demand response aggregator cloud platform determines the number of blocks k(t) that the master node needs to produce in the current time period. Definition Define x i (t)∈{0,1} is the block quantity decision indicator variable. i (t) = 1, indicating the number of blocks selected by the demand response aggregator cloud platform for the i-th level during time period t, i.e., the number of blocks k(t) = K. iThe master node is formed by N edge servers taking turns serving as the master node. The more blocks a master node produces, the more demand response requests need to be completed within that time period. On the other hand, since block generation requires consensus among the edge servers, an excessively large number of blocks will lead to a significant increase in consensus latency.

[0066] The security resource allocation optimization model in step S2 is represented by the joint optimization problem of the number of blocks in the consensus process and the sub-channel allocation decision of the block consensus nodes, i.e., the first optimization problem. In a preferred embodiment of the present invention, a method for constructing a security resource allocation optimization model for blockchain-enabled demand response is proposed. This method uses a demand response consensus model, a consensus throughput model, a consensus latency model, a sub-channel bandwidth constraint model, and a decentralized constraint model to construct the security resource allocation optimization model. The steps of this method are as follows:

[0067] S21: A demand response consensus model is constructed based on the PBFT consensus process. Compared with the traditional demand response business process, in this embodiment, the edge server, in addition to processing the user demand response behavior schemes issued by the demand response aggregator cloud platform, also needs to verify signatures and Message Authentication Codes (MACs). This can be regarded as the computational overhead of executing consensus. Define α, β, and θ to represent the computational complexity of processing the demand response information issued by the demand response aggregator cloud platform, generating or verifying a signature, and generating or verifying a MAC, respectively. The demand response consensus process includes five steps: request, pre-preparation, preparation, submission, and response, as detailed below.

[0068] (1) Request: In time period t, define the master node as edge server S. p User demand response behavior schemes issued by the demand response aggregator cloud platform will be stored as transactions in the pending pool. For a transaction, the master node first verifies its signature. If valid, it verifies the transaction's MAC address. If still valid, the master node will process the demand response information in the transaction and obtain the processing result. At the end of the period, all processed transactions, processing results, and important transaction-related information will be packaged into a new block.

[0069] Assume the average size of each transaction is Let B(t) be the size of a single block in time period t. Then, the maximum number of transactions that a block can contain is: This is a floor function. Assume that g(t) is correct in the demand response information sent by the demand response aggregator cloud platform during time period t. At this stage, S... p Verification required The signature and MAC address of each transaction still need to be processed. Request phase master node S receives request response information from each transaction. p The data processing complexity is expressed as

[0070]

[0071] The computation latency of a single block during the request phase is expressed as:

[0072]

[0073] In the formula, Indicates the master node S p Its computing power.

[0074] (2) Pre-preparation: After generating a new block, the master node S p The pre-prepared message and signed block are multicast to all backup nodes via a 5G communication network. The transmission delay generated by this process depends on the node that last receives the block; therefore, the transmission delay of a single block during the preparation phase is expressed as...

[0075]

[0076] In the formula, Indicates from the master node S p To backup node S n To determine the data transmission rate, this paper considers multicast OFDMA, where one sender uses a subchannel to multicast the same message to multiple receivers. The entire spectrum bandwidth W is divided into E ≥ N subchannels, each with a bandwidth of W0. Represented as

[0077]

[0078] In the formula, For the demand response aggregation cloud platform in time period t, the main node is S. p Number of sub-channels allocated The signal-to-noise ratio is given. Simultaneously, the master node needs to generate one signature and N-1 MAC addresses, with a data processing complexity of O(n).

[0079]

[0080] Backup nodes need to verify the primary node's signature and MAC address, and also need to process... The backup node's data processing complexity is expressed as: (The text abruptly ends here, so the translation stops as well.)

[0081]

[0082] In the formula, S n ≠Sp The completion of the pre-preparation phase of the consensus process requires all nodes to have processed the information in the transaction and verified the signature. Therefore, the computation latency of a single block in the pre-preparation phase is expressed as...

[0083]

[0084] In the formula, Represents node S n Its computing power.

[0085] (3) Preparation: After verifying the signature and MAC, each backup node will send a preparation message to all other nodes. Simultaneously, each node checks whether the received preparation message matches the pre-preparation message received during the pre-preparation phase. According to the PBFT protocol, once 2f preparation messages are received from other nodes, the commit phase begins.

[0086] Similar to the preparation phase, the transmission delay of a single block in the preparation phase is expressed as...

[0087]

[0088] In the formula, Indicates backup node S n To backup node S n′ The transmission rate, its calculation method is the same as... resemblance.

[0089] Let f ≤ (N-1) / 3 be the number of dishonest nodes in PBFT. The master node needs to verify 2f signatures and MACs from the backup nodes, and its data processing complexity is O(n).

[0090]

[0091] Other backup nodes need to generate one signature and N-1 MACs for the prepared message, and need to verify 2f signatures and MACs from the master node. Therefore, the data processing complexity of the backup nodes is expressed as...

[0092]

[0093] In the formula, S n ≠S p Similar to the pre-preparation phase, the computation latency of a single block in the preparation phase is expressed as...

[0094]

[0095] (4) Commit: After each backup node receives 2f preparation confirmation messages consistent with the pre-preparation messages, it sends a commit message to all nodes. During the commit phase, the transmission delay of a single block is expressed as...

[0096]

[0097] During the commit phase, each node needs to generate one signature and N-1 MACs for the commit message. After receiving the commit message, each node also needs to verify 2f signatures and MACs. Therefore, the data processing complexity for each node is O(n^2).

[0098]

[0099] Therefore, the computation latency of a single block during the commit phase is expressed as:

[0100]

[0101] (5) Reply: After collecting 2f commit messages, each backup node confirms the new block as a valid block and generates a local copy of the block. Simultaneously, each backup node sends a reply message to the master node. Upon receiving the reply message, the master node updates the new block into the blockchain. During the reply phase, the transmission latency for a single block is...

[0102]

[0103] During the response phase, each backup node needs to generate... A signature is required, which the master node needs to generate. The computational complexity of the backup node is MAC addresses. For the master node, 2f signatures and MACs need to be verified, and its computational complexity is O(n). Therefore, the computation latency of a single block in the response phase is expressed as:

[0104]

[0105] S22: This embodiment constructs a consensus throughput model. During block generation, the number of transactions that successfully reach consensus is limited by the block size and the master node's data processing capabilities. In a preferred embodiment of the invention, the consensus throughput model is as follows:

[0106] Considering the block size, the total number of transactions recorded in a single block during time period t is:

[0107]

[0108] Since the master node needs to generate k(t) blocks consecutively, the last few blocks may be ignored due to excessive transmission delay between the current master node and the master node in the next time period. Let B be the master node in the (t+1)th time period. p′ Then the master node B in time period t p Transmitted to master node B in the next time period p′ The transmission rate is Therefore, the number of ignored blocks is expressed as

[0109]

[0110] In the formula, k(t)B(t) is the sum of the sizes of all blocks generated by the master node in the t-th time period. For B in the t-th time period p Able to send to B p′ The amount of data transmitted. If, in time period t, the sum of the block sizes generated by the master node exceeds B... p Able to send to B p′ The amount of data transmitted will affect the throughput, meaning some blocks will be ignored because they cannot be transmitted before the end of the time period. Therefore, consensus throughput is defined as the number of transactions successfully processed within a single time period, used to measure the transaction processing performance of the blockchain-enabled demand-response architecture in each time period. This can be expressed as...

[0111]

[0112] S23: This step constructs a consensus latency model, defining the block consensus latency as the time from when the edge server receives the demand response information from the demand response aggregator cloud platform to when the block is successfully uploaded to the blockchain. The block consensus latency specifically includes two parts: transmission latency and computation latency. In a preferred embodiment of this invention, the consensus latency model is as follows:

[0113] In time period t, the transmission delay is the sum of the transmission delays generated during the preparation, preparation, submission, and response phases of that time period, denoted as:

[0114]

[0115] Similarly, computation latency is the sum of the computation latency incurred during the request phase, pre-preparation phase, preparation phase, submission phase, and response phase within that time period, expressed as:

[0116]

[0117] Therefore, the block consensus latency in time period t is expressed as:

[0118]

[0119] S24: This embodiment constructs a decentralized constraint model. In the blockchain-enabled demand response model, the degree of decentralization not only reflects the overall risk resistance of the model but also serves as a measure of fairness. A highly decentralized architecture can more effectively prevent single points of failure and network attacks, while ensuring that the interests of all system participants are considered and maintained equally. The Gini coefficient is used to describe the degree of decentralization of the demand response resource trading system. As time progresses, the blockchain system will generate different numbers of blocks. The set of block counts up to time period t is defined as... Therefore, in a preferred embodiment of the present invention, The Gini coefficient is expressed as

[0120]

[0121] In the formula, A smaller Gini coefficient indicates a smaller difference in the number of blocks generated by the blockchain-enabled demand response model across different time periods, signifying higher decentralization and better security. To ensure the blockchain-enabled demand response model maintains a high degree of decentralization in the long term, this paper introduces a long-term decentralization constraint. Let G be defined. max If the maximum tolerable Gini coefficient for a blockchain system is defined as follows: The long-term stable decentralized constraint is represented as

[0122]

[0123] By quantifying the degree of decentralization using the Gini coefficient, we can ensure that power and resources in blockchain-enabled demand response models are not excessively concentrated on a few nodes, thereby maintaining the stability and sustainable development of the entire system.

[0124] S25: Constructing Sub-channel Allocation Constraints. Due to the limited total bandwidth resources available in the blockchain-enabled demand response resource trading system, the number of sub-channels allocated by the demand response aggregator cloud platform to each block consensus node must meet the following constraints.

[0125]

[0126] W represents the total system bandwidth. Represented as node S n The number of sub-channels allocated, W0 represents the average bandwidth of the sub-channels, and C2 represents that the bandwidth allocated to all nodes should not exceed the total bandwidth of the system.

[0127] S26: A security resource allocation optimization model for blockchain-enabled demand response, whose main performance evaluation indicators are consensus throughput and block consensus latency. Consensus throughput represents the number of demand response services completed in each time period, while block consensus latency represents the time it takes for each edge server to reach a consensus on the user demand response behavior plan issued by the demand response aggregator cloud platform, including the transmission and computation latency of five stages: request stage, pre-preparation stage, preparation stage, submission stage, and response stage. A longer block consensus latency will affect the progress of demand response services and increase the risk of demand response information being altered or revoked. Therefore, this paper optimizes the number of blocks in the consensus process of each time period and the sub-channel allocation decision of the block consensus nodes to maximize the weighted sum of consensus throughput and the reciprocal of block consensus latency. In a preferred embodiment of this invention, the first optimization problem is constructed as follows:

[0128]

[0129] In the formula, θ M This is a weighting coefficient used to balance the order of magnitude and to weigh consensus throughput against block consensus latency performance. Let represent the decision vector for the number of blocks in time period t. Let C1 and C2 represent the sub-channel allocation vector for time period t. C1 and C2 represent the block quantity decision indicator variable and the sub-channel allocation constraint, respectively, and C3 represents the long-term decentralization constraint.

[0130] In a preferred embodiment of the present invention, the optimization problem P1 cannot be directly solved due to the coupling between the long-term decentralized constraint C3 and short-term decision-making, as well as the unknowability of future global information. Therefore, this paper constructs a virtual queue based on Lyapunov optimization theory, and transforms the long-term decentralized constraint C3 into a queue stability constraint in the decentralized queue construction submodule.

[0131] Define a decentralized deficit virtual queue Z(t), and express its update formula as follows:

[0132]

[0133] In the formula, Z(t) represents the actual Gini coefficient of the system at the end of time period t, and the upper limit of the maximum tolerable Gini coefficient G. max The larger the deviation Z(t), the lower the degree of decentralization and the lower the security of the blockchain-enabled demand response architecture.

[0134] Based on the virtual queue and Lyapunov optimization theory, the virtual queue backlog reflects the satisfaction of long-term decentralization constraints, thus decoupling short-term block quantity decisions and sub-channel allocation decisions from long-term decentralization constraints. This transforms optimization problem P1 into optimization problem P2, the second optimization problem.

[0135]

[0136] In the formula, θ Z , which is a weighting coefficient used to balance the order of magnitude and to weigh the optimization objective against the stability of the queue. The security boundary violation penalty is calculated by the security boundary violation penalty calculation submodule. It represents the negative impact of failing to meet the decentralization constraint on the optimization objective. The larger the value, the lower the degree of decentralization, forcing future optimization decisions to be more inclined to reduce the backlog of the decentralized deficit virtual queue in order to meet the long-term decentralization constraint.

[0137] Based on the above embodiments, the proposed method for constructing a secure resource allocation optimization model for blockchain-enabled demand response is as follows: First, a blockchain-enabled secure demand response resource transaction model is constructed, including a demand response consensus model, a consensus throughput model, a consensus latency model, a sub-channel bandwidth constraint model, and a decentralized constraint model. Next, a joint optimization problem is constructed for the number of blocks in the consensus process and the sub-channel allocation decision of block consensus nodes. Finally, a secure resource allocation software module based on security boundary violation penalty-driven adaptive deep learning is used. In the decentralized queue construction sub-module, a virtual queue is constructed based on Lyapunov optimization theory, transforming long-term decentralized constraints into queue stability constraints for easier solution. This provides support for the subsequent secure resource allocation decision sub-module to implement an adaptive collaborative allocation method for block and channel resources, thereby ensuring the authenticity and consistency of information in the demand response system while improving the system's security and reliability.

[0138] In conjunction with the above embodiments, addressing the problem of low efficiency in demand response information interaction due to the lack of consideration for the joint optimization of block consensus throughput and consensus latency, this invention proposes a method for constructing a secure resource allocation optimization model for blockchain-enabled demand response. By constructing this model, the weighted sum of consensus throughput and the reciprocal of block consensus latency is used as the optimization objective. Through dynamic adjustment of the optimization weights, joint guarantees of consensus throughput and block consensus latency are achieved. Furthermore, based on Lyapunov optimization theory, the short-term joint guarantee problem is decoupled from long-term decentralization constraints, providing support for the collaborative allocation optimization of block and channel resources, thereby improving the security and efficiency of demand response information interaction.

[0139] In a preferred embodiment of the present invention, a block and channel resource adaptive allocation method based on security boundary violation penalty-driven adaptive deep learning is proposed. This method constructs a temporal Markov process based on the joint optimization problem P2, as follows:

[0140] The optimization problem P2 is described as an MDP process. The demand response aggregator cloud platform is defined as the decision-maker. At the beginning of each time period, the demand response aggregator cloud platform executes specific block quantity and sub-channel allocation decisions based on the current state and transitions to a new state. Since the channel states and decentralized deficit virtual queues between edge servers cannot be predicted in advance before the demand response aggregator cloud platform makes these decisions, i.e., global information is unavailable, the state identification submodule collects relevant information such as the historical channel states and historical virtual queue states of the previous time period as the basis for block quantity and bandwidth allocation decisions. The state space construction submodule constructs the state space Φ(t), which includes the signal-to-noise ratio between different edge servers, the computing resources of the edge servers, and the decentralized deficit virtual queue, specifically represented as follows:

[0141]

[0142] In the formula,

[0143] In time period t, the demand response aggregator cloud platform determines the action space A(t) for the current time period based on the state space, including block quantity decisions and sub-channel allocation decisions, denoted as...

[0144] A(t) = {X(t); W(t)}. (30)

[0145] It is worth noting that when the demand response aggregator cloud platform executes actions, it needs to satisfy constraints C1 and C2.

[0146] Based on the concept of penalty functions, when the constraints of the optimization problem are satisfied, the reward function is equivalent to the objective function. When the constraints are not satisfied, a penalty is imposed according to the degree of constraint non-compliance, reducing the reward value. This reward value calculation is performed by the reward calculation submodule. The reward function r(t) for the demand response aggregator cloud platform aggregator selection action A(t) in the t-th time period is defined as follows:

[0147]

[0148] In a further embodiment of the present invention, the block and channel resource adaptive allocation method based on security boundary violation penalty-driven adaptive deep learning further includes making action decisions driven by security boundary violation penalty; then, each edge server executes the action decided by the demand response aggregator cloud platform; finally, the demand response aggregator cloud platform performs experience learning and network updates to realize the optimization strategy for the number of blocks and channel resource allocation. The collaborative allocation optimization method (i.e., the block and channel resource adaptive allocation method based on security boundary violation penalty-driven adaptive deep learning) includes an evaluation network and a target network, and includes the following process:

[0149] S31: Input the state information of the current time slot edge layer into the evaluation network to obtain the state-action value function of different actions in the current time slot. Based on this, select the action with the highest state-action value in the current time slot according to the strategy obtained by the exploration and utilization adaptive adjustment algorithm. The exploration and utilization adaptive adjustment algorithm is used to reduce the penalty in the learning process by utilizing the existing optimal resource management strategy when the security boundary violation penalty in the previous time slot meets the condition of a large value. When the security boundary violation penalty in the previous time slot does not meet the condition of a large value, explore other optimal resource management strategies to avoid getting trapped in local optima.

[0150] S32: Execute the selected action, calculate the reward value, move to the next state, and store the action chain in the experience pool. The action chain includes the current state, the corresponding action and reward value, and the next state.

[0151] S33: Randomly select a set of empirical data from the experience pool and calculate the temporal difference error of the evaluation network in the current time slot;

[0152] S34: Update the evaluation network parameters based on the time-series difference error, and synchronize the target network with the evaluation network every set time interval;

[0153] S35: Iterate through the above steps until the optimization cycle ends.

[0154] The proposed adaptive block and channel resource allocation method based on security boundary violation penalty-driven adaptive deep learning can be implemented using a security resource allocation software module based on security boundary violation penalty-driven adaptive deep learning. This module is as follows: Figure 3 As shown. Deployed on the demand response aggregator cloud platform, it includes a state recognition submodule, a decentralized queue construction submodule, a security boundary violation penalty submodule, an exploration and exploitation adaptive adjustment submodule, a state space construction submodule, a reward calculation submodule, a security resource allocation decision submodule, an experience learning and network update submodule, and a security resource allocation decision execution submodule. Figure 3 As shown, the module's functions are described in detail below:

[0155] State recognition submodule: It perceives information on historical channel status, computing resources of each server, and historical virtual queue status to provide support for state space construction.

[0156] The decentralized queue construction submodule constructs a decentralized virtual queue, transforming long-term decentralization constraints into queue stability constraints for easier solution.

[0157] Security boundary violation penalty submodule: Calculates the value of the adaptive factor for security boundary violation penalty based on the state information of the edge layer.

[0158] The exploration and exploitation adaptive adjustment submodule adjusts exploration and exploitation based on the security boundary violation penalty adaptive factor.

[0159] State Space Construction Submodule: Constructs the action space of the MDP problem based on the information obtained from the state recognition module, including historical channel state, computing resources of each server, and historical virtual queue state.

[0160] Security resource allocation decision submodule: adopts the ε-greedy strategy driven by security boundary violation penalty to select the action with the highest action value in the current time slot state.

[0161] Reward Calculation Submodule: Calculates the rewards obtained from performing different actions.

[0162] Experience learning and network update submodule: After each learning session, data and experience are collected, and the system's decision-making strategy is continuously updated based on the DQN network.

[0163] The security resource allocation decision execution submodule outputs security resource allocation decisions, including decisions on the number of blocks and bandwidth resources. For ease of description, the implementation process of the block and channel resource adaptive allocation method based on security boundary violation penalty-driven adaptive deep learning is introduced in conjunction with the above module. It should be understood that this module is only for descriptive purposes and does not limit the specific implementation vehicle of the method.

[0164] Traditional deep reinforcement learning algorithms employ a fixed exploration-exploitation tradeoff when executing actions. This leads to a failure to promptly adjust block quantity and sub-channel allocation decisions when the security bound is breached, thereby reducing the system's decentralization and resilience. To address this issue, an adaptive DQN block and channel resource collaborative allocation optimization method based on security bound violation penalty is adopted. The security resource allocation decision submodule of the demand response aggregator cloud platform maintains two deep neural networks simultaneously, including one valuation network and one target network. The network parameter vectors of the valuation network and the target network in the cloud platform's security resource allocation decision submodule can be represented as ω. e (t) and ω t(t). The valuation network realizes the mapping from state-action combinations to the state-action value function Q[Φ(t), A(t)]. The state-action value function represents the expected total future reward that the demand response aggregator cloud platform can obtain when choosing action A(t) under state Φ(t). The cloud platform's security resource allocation decision submodule learns and executes a block number and sub-channel allocation decision in each time period through the valuation network; the edge server then feeds back the updated information to the demand response aggregator cloud platform. The algorithm reduces the correlation of sample data between consecutive time slots by setting a target network and an experience replay pool, thereby improving the optimization performance of the evaluation network. The proposed algorithm learns block and channel collaborative resource management strategies under non-global information based on the deep Q network, and adaptively adjusts the probability of exploration and utilization during the learning process through the exploration and utilization adaptive adjustment submodule. When the security boundary violation penalty in the previous time period is large, the algorithm will tend to utilize the existing optimal resource management strategy to reduce the penalty received during the learning process. Conversely, the algorithm will tend to explore other resource management strategies to avoid getting trapped in local optima. The specific process is described as follows:

[0165] (1) Action decision driven by security boundary violation punishment: The state space construction submodule of the demand response aggregator cloud platform inputs the state information Φ(t) of the current time slot edge layer into the evaluation network of the security resource allocation decision submodule to obtain the state-action value function Q[Φ(t),A(t)|ω] of different actions A(t) in the current time slot. e [t], and based on this, the ε-greedy policy based on the output of the exploration and utilization submodule selects the action with the highest action value in the current time slot state, i.e.

[0166]

[0167] Where Ω∈[0,1], is a random number generated by the security resource allocation decision submodule of the demand response aggregator cloud platform when selecting actions. Considering the dynamic adjustment of the agent's exploration and utilization probabilities in each time period based on the exploration and utilization submodule, As an adaptive exploration factor, that is, the agent uses... It randomly selects an action from the action space with a certain probability to execute; otherwise, it executes the action with the largest Q value.

[0168] Where ε0 is the initial exploration factor.

[0169] (2) Action Execution: The security resource allocation decision execution submodule of the demand response aggregator cloud platform executes action A. * (t), the reward calculation submodule calculates the reward value, updates relevant information, transitions to the next state Φ(t+1), and sets the action chain [Φ(t),A * [(t),r(t),Φ(t+1)] are stored in the experience pool.

[0170] (3) Experience Learning and Network Updates: The experience learning and network update sub-module of the demand response aggregator cloud platform starts from the experience pool. A set of empirical data was randomly selected from the data. The calculation formula is as follows: [Formula omitted for brevity]

[0171]

[0172] In the formula, This represents the number of action chains drawn from the experience pool. y(t) represents the state-action value updated by the demand response aggregator cloud platform through the target network after receiving the reward r(t), and its update formula is:

[0173]

[0174] In the formula, γ∈[0,1] is a discount factor used to measure the impact of future rewards on the current state-action value. The larger γ is, the greater the impact of future rewards on the current action strategy, and the more the algorithm focuses on long-term returns.

[0175] Subsequently, the parameters of the evaluation network are updated based on the temporal difference error. The target network synchronizes with the evaluation network every T0 time intervals and remains unchanged for the next T0-1 time intervals. The above three steps are iteratively executed until the entire optimization cycle ends.

[0176] Based on the above embodiments, this invention addresses the problem of low long-term resilience and low security of demand response caused by the long-term decentralized constraints of block consensus. It proposes an optimization algorithm for collaborative allocation of block and channel resources based on adaptive DQN driven by security boundary violation penalties. A security boundary penalty is constructed based on a decentralized virtual queue deficit and incorporated into the reward function of the MDP model. The weights of exploration and utilization during the learning process are adaptively adjusted according to the security boundary penalty to optimize the number of blocks and the channel allocation learning strategy. When the security boundary penalty increases, the algorithm tends to utilize the existing optimal resource management strategy to reduce the penalty received during the learning process. Conversely, the algorithm tends to explore other resource management strategies to avoid getting trapped in local optima. This provides an efficient, secure, and adaptive resource allocation method for demand response systems, effectively ensuring the stability and reliability of the system.

[0177] Based on the same inventive concept, this application also provides a network security system for supporting user-side resource participation in demand response, which is used to implement the network security method for supporting user-side resource participation in demand response as described above. The solution provided by this system is similar to the solution described in the above method. Therefore, the specific limitations in the network security system embodiments for supporting user-side resource participation in demand response provided below can be found in the limitations of the network security method for supporting user-side resource participation in demand response described above, and will not be repeated here.

[0178] This embodiment provides a network security system that supports user-side resources in demand response, including: cloud layer, edge layer and physical layer;

[0179] The cloud layer is used to distribute user demand response behavior plans to the edge layer for consensus; it is also used to solve the number of blocks and sub-channel allocation decisions for block consensus nodes in each time period based on the security resource allocation optimization model. The security resource allocation optimization model is used to maximize the weighted average of consensus throughput and the reciprocal of block consensus latency by optimizing the number of blocks and sub-channel allocation decisions for block consensus nodes in each time period under constraints including sub-channel allocation and long-term decentralization. The long-term decentralization constraint is used to ensure that the blockchain-enabled demand response model maintains a high degree of decentralization that meets long-term requirements, so as to ensure the security of user-side resources participating in demand response; it also guides the consensus process based on the solved number of blocks and sub-channel allocation decisions for block consensus nodes.

[0180] The edge layer is used to reach consensus based on the blockchain-enabled demand response model, which is used to reach consensus on user demand response behavior plans based on a preset consensus algorithm; it is also used to distribute the consensus-reached user demand response behavior plans to the physical layer.

[0181] The physical layer is used to execute the response plan for user needs.

[0182] The system comprises a physical layer, an edge layer, and a cloud layer. The cloud layer includes the power grid company's demand response cloud platform and the demand response aggregator cloud platform. Demand response operations are jointly performed by the power grid company's demand response cloud platform, the demand response aggregator cloud platform, and electricity users. The power grid company's demand response cloud platform manages and publishes demand response signals. Upon receiving a demand response signal from the power grid company's demand response cloud platform, the demand response aggregator cloud platform executes a smart contract to determine the user's demand response behavior plan and sends the response information to the edge layer. The edge layer consists of edge servers and 5G base stations. The base stations are responsible for providing communication coverage within the area and supporting communication between the physical layer, edge layer, and cloud layer. The edge servers are responsible for receiving demand response information from the demand response aggregator cloud platform, issuing demand response requests to electricity users, and maintaining the blockchain platform. The physical layer consists of electricity users with distributed power sources, loads, and storage devices, and is responsible for adjusting their corresponding electricity consumption behavior after receiving demand response requests from the edge layer. This provides platform support for subsequent software modules for secure resource allocation based on security boundary violation penalties.

[0183] This system can execute the block and channel resource collaborative allocation optimization algorithm based on security boundary violation penalty-driven adaptive DQN proposed in the aforementioned embodiments. This algorithm is deployed on the software module for allocating secure resources based on security boundary violation penalty. A security boundary penalty is constructed based on a decentralized virtual queue deficit and incorporated into the reward function of the MDP model. The weights of exploration and utilization during the learning process are adaptively adjusted according to the security boundary penalty to optimize the number of blocks and the channel allocation learning strategy. When the security boundary penalty increases, the algorithm tends to utilize the existing optimal resource management strategy to reduce the penalty received during the learning process. Conversely, the algorithm tends to explore other resource management strategies to avoid getting trapped in local optima. This provides an efficient, secure, and adaptive resource allocation method for demand response systems, effectively ensuring the stability and reliability of the system.

[0184] Furthermore, the software module for allocating security resources based on security boundary violation penalties can be deployed on the demand response aggregator cloud platform. This module includes submodules for state identification, decentralized queue construction, security boundary violation penalties, adaptive adjustment for exploration and exploitation, state space construction, reward calculation, security resource allocation decision-making, experience learning and network update, and security resource allocation decision execution. The state identification submodule senses information about historical channel states, server computing resources, and historical virtual queue states to support state space construction. The decentralized queue construction submodule constructs decentralized virtual queues, transforming long-term decentralization constraints into queue stability constraints for easier solution. The security boundary violation penalty submodule calculates the value of the security boundary violation penalty adaptive factor based on edge layer state information. The adaptive adjustment for exploration and exploitation submodule adjusts exploration and exploitation based on the security boundary violation penalty adaptive factor. The state space construction submodule constructs the action space of the MDP problem based on information obtained from the state identification module, including historical channel states, server computing resources, and historical virtual queue states. The security resource allocation decision-making submodule employs an ε-greedy strategy driven by security boundary violation penalties to select the action with the highest action value in the current time slot state. The reward calculation submodule calculates the rewards obtained from performing different actions. The experience learning and network update submodule collects data and experience after each learning cycle and continuously updates the system's decision-making strategy based on the DQN network. The security resource allocation decision execution submodule outputs the security resource allocation decision, including the allocation of block quantity and bandwidth resources. Through adaptive adjustment, real-time monitoring, and continuous learning mechanisms, the system's security, stability, and resource allocation efficiency are improved.

[0185] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0186] Reference Figure 5The present invention also provides a computer device 1, including: a memory 12 and a processor 11, and a computer program 13 stored on the memory 12. When the computer program 13 is executed on the processor 11, it implements a network security method for supporting user-side resource participation in demand response as described in any of the above methods.

[0187] The computer device 1 may be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device 1 may include, but is not limited to, a processor 11 and a memory 12. Those skilled in the art will understand that... Figure 5 The computer device 1 is merely an example and does not constitute a limitation on the computer device 1. It may include more or fewer components than shown in the figure, or combine certain components, or different components, such as input / output devices, network access devices, etc.

[0188] The processor 11 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0189] In some embodiments, the memory 12 may be an internal storage unit of the computer device 1, such as a hard disk or memory of the computer device 1. In other embodiments, the memory 12 may be an external storage device of the computer device 1, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the computer device 1. Furthermore, the memory 12 may include both internal and external storage units of the computer device 1. The memory 12 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory 12 can also be used to temporarily store data that has been output or will be output.

[0190] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a network security method for supporting user-side resource participation in demand response as described in any of the above methods.

[0191] In this embodiment, if the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0192] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0193] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0194] In the embodiments disclosed in this application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0195] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A cyber-security method to support user-side resource participation in demand response, characterized by, The method comprises the following steps: obtaining a user demand response behavior scheme and performing consensus on the user demand response behavior scheme according to a blockchain-enabled demand response model, wherein the blockchain-enabled demand response model is used to perform consensus on the user demand response behavior scheme according to a preset consensus algorithm; solving the number of blocks and sub-channel allocation decisions of block consensus nodes in each time period according to a security resource allocation optimization model, wherein the security resource allocation optimization model is used to maximize the weighted sum of consensus throughput and block consensus delay inverse by optimizing the number of blocks and sub-channel allocation decisions of block consensus nodes in each time period under the constraints of sub-channel allocation and long-term decentralization, and the long-term decentralization constraint is used to ensure that the blockchain-enabled demand response model maintains a long-term high requirement for decentralization degree to ensure the security of user-side resources participating in demand response; optimizing the consensus process based on the solved number of blocks and sub-channel allocation decisions of block consensus nodes to reach consensus, thereby supporting user-side resources to participate in demand response based on the user demand response behavior scheme that has reached consensus; the security resource allocation optimization model is represented by a first optimization problem as follows: ; In the formula, Indicates the first The decision vector for the number of blocks in a given time period. Indicates the first Time-segment sub-channel allocation vector, Number of time periods For time period index, Indicates the first Consensus throughput over a given period These are the weighting coefficients. For the first The block consensus latency for a given time period, and C1 is an indicator variable for the number of blocks. The number of blocks is the level of quantity. C2 is the block quantity level index, and C2 is the sub-channel allocation constraint. Indicates the total system bandwidth. Represented as nodes The number of sub-channels allocated. C3 represents the average bandwidth of the sub-channel, and C3 is the long-term decentralization constraint. Indicates up to the number Set of blocks for a given time period The Gini coefficient, This represents the upper limit of the maximum tolerable Gini coefficient.

2. The network security method of claim 1, wherein, when solving the security resource allocation optimization model, a virtual queue is constructed based on Lyapunov optimization theory, the long-term decentralization constraint is converted into a queue stability constraint based on the virtual queue, and a second optimization problem is obtained as follows: ; wherein is a weight coefficient, is the first the deviation of the actual Gini coefficient of the system at the end of the period from the upper limit of the maximum tolerable Gini coefficient the deviation of the actual Gini coefficient of the system at the end of the period from the upper limit of the maximum tolerable Gini coefficient 3. The cyber-security method of supporting user-side resource participation demand response according to claim 2, characterized in that, the second optimization problem is constructed as a time series Markov process for solving, and the constructed time series Markov process comprises a state space, an action space and a reward function; the state space comprises signal-to-noise ratios among different edge servers, computing resources of edge servers and decentralized deficit virtual queues; the action space comprises block quantity decisions and sub-channel allocation decisions; the reward function is used to punish according to the degree of constraint not being met when the constraint condition of the optimization problem is not met, thereby reducing the reward value.

4. The network security method of claim 3, wherein, the time series Markov process is solved by using a self-adaptive DQN block and channel resource collaborative allocation optimization method based on security boundary violation punishment driving, the collaborative allocation optimization method comprises one evaluation network and one target network, and the solving process comprises: inputting state information of the current time slot edge layer into the evaluation network to obtain state-action value functions of different actions in the current time slot, and selecting an action with the maximum state-action value based on a policy obtained by an exploration and utilization adaptive adjustment algorithm, wherein the exploration and utilization adaptive adjustment algorithm is used to utilize an existing optimal resource management strategy to reduce the punishment in the learning process when the security boundary violation punishment in the last time period meets a large condition, and to explore other optimal resource management strategies to avoid falling into a local optimum when the security boundary violation punishment in the last time period does not meet the large condition; performing the selected action, calculating the reward value, transferring to the next state and storing the action chain into an experience pool, wherein the action chain comprises the current time state, the corresponding action and the reward value, and the next time state; randomly sampling a set of experience data from the experience pool, calculating a timing difference error of the evaluation network in a current time slot; updating the evaluation network parameters according to the timing difference error, and synchronizing the target network with the evaluation network every set time period; iteratively performing the above steps until the optimization period ends.

5. The cyber-security method of supporting user-side resource participation demand response according to claim 1, characterized in that, the consensus throughput and the block consensus latency are determined by a consensus throughput model and a consensus latency model, respectively, as follows: Consensus throughput model: ; ; ; wherein, the number of blocks produced by the master node for the total number of transactions recorded in a single block, the number of blocks produced by the master node for the total number of blocks produced by the master node for the the master node for the the master node for the the number of blocks ignored by the master node for the the total size of all blocks produced by the master node for the the total size of all blocks produced by the master node for the the average size of each transaction, the total size of all blocks produced by the master node for the the total size of all blocks produced by the master node for the the total size of all blocks produced by the master node for the the total size of all blocks produced by the master node for the the total size of all blocks produced by the master node for the the size of a single block for the the size of a single block for the Consensus latency model: ; ; ; In the formula, is the number of blocks produced by the primary node in the time period; is the transmission latency of the pre-preparation phase, is the transmission latency of the preparation phase, is the transmission latency of the submission phase, is the transmission latency of the reply phase, is the number of blocks produced by the primary node in the time period; is the computation latency of the primary node in the time period; is the computation latency of the request phase, is the computation latency of the pre-preparation phase, is the computation latency of the preparation phase, is the computation latency of the submission phase, is the computation latency of the reply phase.

6. The cyber-security method of supporting user-side resource participation demand response according to claim 1, wherein, Up to the Set of blocks for a given time period Gini coefficient Determined according to the following formula: ; In the formula, and are the first and the number of blocks that the primary node needs to produce in the time period, represents the number of nodes that take turns to act as primary nodes.

7. A cyber-security system supporting user-side resource participation demand response, characterized by, comprising: a cloud layer, an edge layer and a physical layer; the cloud layer is configured to issue a user demand response behavior scheme to the edge layer for consensus; is also used to solve the block quantity of each time period consensus process and the subchannel allocation decision of the block consensus node according to a secure resource allocation optimization model, the secure resource allocation optimization model is used to maximize the weighted sum of consensus throughput and block consensus latency inverse by optimizing the block quantity of each time period consensus process and the subchannel allocation decision of the block consensus node under the constraints including subchannel allocation and long-term decentralization, and the long-term decentralization constraint is used to ensure that the blockchain-enabled demand response model maintains a long-term higher requirement for decentralization degree, so as to ensure the security of user-side resource participation in demand response; also based on the solved block quantity and subchannel allocation decision of the block consensus node to guide the consensus process; the edge layer is configured to perform consensus according to the blockchain-enabled demand response model, and the blockchain-enabled demand response model is used to perform consensus on the user demand response behavior scheme according to a preset consensus algorithm; also used to issue the user demand response behavior scheme reached consensus to the physical layer; the physical layer is configured to execute the user demand response behavior scheme; the secure resource allocation optimization model is represented by a first optimization problem as follows: ; In the formula, Indicates the first The decision vector for the number of blocks in a given time period. Indicates the first Time-segment sub-channel allocation vector, Number of time periods For time period index, Indicates the first Consensus throughput over a given period These are the weighting coefficients. For the first The block consensus latency for a given time period, and C1 is an indicator variable for the number of blocks. The number of blocks is the level of quantity. C2 is the block quantity level index, and C2 is the sub-channel allocation constraint. Indicates the total system bandwidth. Represented as nodes The number of sub-channels allocated. C3 represents the average bandwidth of the sub-channel, and C3 is the long-term decentralization constraint. Indicates up to the number Set of blocks for a given time period The Gini coefficient, This represents the upper limit of the maximum tolerable Gini coefficient.

8. A computer device, comprising: The device comprises a processor and a memory: The memory is configured to store a computer program and send instructions of the computer program to the processor; The processor executes the network security method for supporting user-side resource participation in demand response according to the instructions of the computer program.

9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to realize the network security method for supporting user-side resource participation in demand response. The computer readable storage medium stores a computer program, and the computer program is executed by the processor to realize the network security method for supporting user-side resource participation in demand response.