A routing and scheduling method, system, device and medium based on the service priority of a virtual power plant

By updating the virtual power plant business flow model and optimizing the routing strategy using Q-learning reinforcement learning method, the performance degradation of traditional routing scheduling methods under complex business needs and dynamic network conditions is solved, and efficient and reliable data transmission and service security are achieved.

CN119697105BActive Publication Date: 2025-06-17STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510197318.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-17
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

The existing traditional routing scheduling methods rely on fixed routing algorithms and are difficult to adapt to complex business needs and dynamic network conditions, resulting in low network resource utilization and unavailability of service security.

Method used

The routing scheduling method based on the service priority of virtual power plant is adopted, and the current network status of the virtual power plant business flow model is updated, and the updated network status is iteratively optimized by using Q-learning reinforcement learning method to output the optimal routing scheduling strategy.

Benefits of technology

Dynamically adjusting the routing scheduling strategy improves network resource utilization, ensures business security, and achieves more efficient and reliable data transmission in complex network environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119697105B_ABST
    Figure CN119697105B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of virtual power plants, and discloses a routing and scheduling method, system, device and medium based on the service priority of virtual power plants. This method updates the current network state of a pre-constructed virtual power plant service flow model according to the service priority of newly accessed service requests, and uses the Q-learning reinforcement learning method to iteratively optimize the updated network state until the termination condition is met, so as to output the optimal routing and scheduling strategy. This dynamic adjustment method overcomes the limitations of traditional routing and scheduling methods that rely on fixed routing algorithms and are difficult to adapt to complex and changeable network environments, effectively improves the utilization rate of network resources, and can achieve more efficient and reliable data transmission on the basis of ensuring service security; by comprehensively considering the virtual power plant service traffic and service priority, this method effectively solves the problems of routing and tight network resource allocation caused by multi-concurrent flow phenomena in virtual power plant networks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of virtual power plant services, and particularly relates to a routing and scheduling method, system, device and medium based on the service priority of virtual power plants. Background Art

[0002] A virtual power plant is one of the important technologies of an intelligent distribution network. Through advanced information communication and intelligent metering technologies, a large number of distributed clean energy sources, controllable loads and energy storage systems in the distribution network are efficiently aggregated to participate in the peak regulation, frequency modulation and demand response services of the large power grid. How to effectively manage the routing of data transmission according to the importance and real-time requirements of different services is an important research direction for the service scheduling of virtual power plants.

[0003] Traditional routing and scheduling methods usually rely on fixed routing algorithms to guide the transmission path of data packets in the network. Most traditional routing and scheduling methods are based on static network topologies and predefined routing policies, aiming to transmit data efficiently and reliably in the network; however, in the face of complex service requirements and dynamic network conditions, the performance of traditional routing and scheduling methods will decline sharply, resulting in low utilization of network resources and security risks for services.

[0004] It can be seen that existing traditional routing and scheduling methods usually rely on fixed routing algorithms. In the face of complex service requirements and dynamic network conditions, the performance will decline sharply, resulting in low utilization of network resources and inability to ensure the security of services. Summary of the Invention

[0005] The present invention provides a routing and scheduling method, system, device and medium based on the service priority of virtual power plants to solve the technical problem that existing traditional routing and scheduling methods usually rely on fixed routing algorithms, and in the face of complex service requirements and dynamic network conditions, the performance will decline sharply, resulting in low utilization of network resources and inability to ensure the security of services.

[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0007] In a first aspect, the present invention provides a routing and scheduling method based on the service priority of virtual power plants, including:

[0008] Updating the current network state of a pre-constructed virtual power plant service flow model based on the service priority of a newly received service request; wherein, the service priority is included in the virtual power plant service flow model;

[0009] Performing iteration on the updated network state based on the Q-learning reinforcement learning method until a termination condition is met, and outputting an optimal routing and scheduling strategy.

[0010] The virtual power plant service flow model is constructed based on service types, service priorities, service flow paths, service bandwidth requirements, and service latency requirements, and is specifically expressed as:

[0011]

[0012] Among them, is the service type; is the service priority; is the service bandwidth requirement; is the service latency requirement; is the service flow path, expressed as a sequence of a series of nodes and connections: ; among them; represents a node in the network, 1 ≤ i ≤ k; and k ≥ 2.

[0013] A further improvement of the present invention lies in that the current network state of the pre-constructed virtual power plant service flow model is updated based on the service priority of the new service request based on access, including:

[0014] When a new service request accesses, according to the service priority of the new service request and the parameters of the virtual power plant service flow model , the new service request is added to the corresponding service queue, and the current network state of the pre-constructed virtual power plant service flow model is updated to obtain the updated network state.

[0015] A further improvement of the present invention lies in that the updated network state is iterated based on the Q-learning reinforcement learning method until the termination condition is met, and the optimal routing and scheduling strategy is output, including:

[0016] Initialize the value table , and all values are set to 0; define the state space and action space; set the learning rate , the discount factor and the exploration rate ;

[0017] Observe the updated network state ; among them, is the current network topology; is the service queue length of each node;

[0018] Action selection process: Use the greedy strategy to select the action ;

[0019] Scheduling process: Perform service scheduling and path selection according to the selected action to meet the requirements of the virtual power plant service flow model;

[0020] State observation process: Action After the execution is completed, observe the next network state and the current reward obtained ;

[0021] Value update process: According to the current value update formula to update value;

[0022] Transfer the updated network state s to the next network state , repeat the action selection process, scheduling process, state observation process and value update process until the termination condition is met, and output the optimal routing and scheduling strategy;

[0023] Among them, the value update formula is as follows:

[0024]

[0025] Among them, is the updated network state; is the current action; is the learning rate; is the current reward; discount factor; is the next network state; is the action optional in the next network state; is the maximum of the next state value.

[0026] A further improvement of the present invention is that the current reward is expressed as: ; Among them, : respectively represent different weight coefficients for adjusting the priorities and the weights of the delay in the current reward in.

[0027] A further improvement of the present invention is that the virtual power plant service flow model interacts with virtual power plant data streams; among them, the virtual power plant data streams include acquisition streams and control streams.

[0028] A further improvement of the present invention is that the virtual power plant services are divided into 4 levels for management, including real-time services, quasi-real-time services, non-real-time services and default services; the service priorities of the service requests corresponding to the 4-level virtual power plant services are levels 1 to 4.

[0029] In a second aspect, the present invention provides a routing and scheduling system based on the service priority of a virtual power plant, including:

[0030] A service request access module, configured to update the current network state of a pre-constructed virtual power plant service flow model based on the service priority of an accessed new service request; wherein, the virtual power plant service flow model includes service priorities.

[0031] A scheduling policy output module, configured to perform iteration on the updated network state based on the Q-learning reinforcement learning method until a termination condition is met, and output an optimal routing and scheduling policy.

[0032] The virtual power plant service flow model is constructed based on service types, service priorities, service flow paths, service bandwidth requirements, and service latency requirements, and is specifically expressed as:

[0033]

[0034] Wherein, is the service type; is the service priority; is the service bandwidth requirement; is the service latency requirement; is the service flow path, expressed as a sequence of a series of nodes and connections: ; wherein; represents a node in the network, 1 ≤ i ≤ k; and k ≥ 2.

[0035] In a third aspect, the present invention provides a device, including:

[0036] A memory, configured to store a computer program;

[0037] A processor, configured to implement the steps of the above-mentioned routing and scheduling method based on the service priority of a virtual power plant when executing the computer program.

[0038] In a fourth aspect, the present invention provides a computer-readable storage medium, which stores a computer program, and the computer program is used to implement the steps of the above-mentioned routing and scheduling method based on the service priority of a virtual power plant when executed by a processor.

[0039] Compared with the prior art, the present invention has the following beneficial effects:

[0040] The present invention also provides a routing and scheduling method based on the service priority of a virtual power plant. This method updates the current network state of a pre-constructed virtual power plant service flow model according to the service priority of newly accessed service requests, and uses the Q-learning reinforcement learning method to iteratively optimize the updated network state until the termination condition is met, thereby outputting an optimal routing and scheduling strategy. This dynamic adjustment method overcomes the limitation of traditional routing and scheduling methods that rely on fixed routing algorithms and are difficult to adapt to complex and changeable network environments, effectively improving the network resource utilization rate, and enabling more efficient and reliable data transmission on the basis of ensuring service security; by comprehensively considering the virtual power plant service traffic and service priority, this method effectively solves the problems of routing and tight network resource allocation caused by multi-concurrent flow phenomena in the virtual power plant network.

[0041] Preferably, in the present invention, the virtual power plant service flow model is defined, including service types, priorities, paths, bandwidth requirements, and delay requirements, providing a more comprehensive information basis for routing and scheduling; by defining the service flow path as a sequence of a series of nodes and connections, the flexibility of routing and scheduling is increased, and it can better adapt to complex network environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 It is a system architecture diagram of the virtual power plant service flow provided by an embodiment of the present invention;

[0043] Figure 2 It is a flowchart of a routing and scheduling method based on the service priority of a virtual power plant provided by an embodiment of the present invention;

[0044] Figure 3 It is a flowchart of a routing and scheduling method based on the service priority of a virtual power plant provided by the present invention;

[0045] Figure 4 It is a schematic structural diagram of a routing and scheduling system based on the service priority of a virtual power plant provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] Embodiment 1

[0047] The present invention provides a routing and scheduling method based on the service priority of a virtual power plant, as Figure 3 shown, including the following steps:

[0048] S1: Update the current network state of the pre-constructed virtual power plant service flow model based on the service priority of the newly accessed service request; wherein, the service priority is included in the virtual power plant service flow model;

[0049] S2: Iterate the updated network state based on the Q-learning reinforcement learning method until the termination condition is met, and output the optimal routing and scheduling strategy;

[0050] Among them, the virtual power plant service flow model interacts with the virtual power plant data flow. Among them, the above virtual power plant data flow includes a collection flow and a control flow, which are used to support scheduling, marketing services, and source-network-load-storage interaction. The virtual power plant service flow model provides strong support for the scheduling, marketing services, and source-network-load-storage interaction of the virtual power plant by interacting with the virtual power plant data flow; the virtual power plant data flow includes a collection flow and a control flow, which helps to achieve more efficient data interaction and service processing.

[0051] In this method, the virtual power plant services are managed in 4 levels, including real-time services, quasi-real-time services, non-real-time services, and default services; the service priorities of the service requests corresponding to the 4-level virtual power plant services are from level 1 to level 4. In this embodiment, the virtual power plant services are managed in 4 levels and corresponding service priorities are set, which helps to give priority to ensuring the transmission requirements of key services when network resources are limited; through reasonable priority settings, the allocation and utilization of network resources can be optimized, and the operation efficiency and stability of the entire network system can be improved.

[0052] Specifically, the above virtual power plant service flow model includes service type, service priority, service flow path, service bandwidth requirement, and service delay requirement, which are specifically expressed as:

[0053]

[0054] is the service type; is the service priority; is the service bandwidth requirement; is the service delay requirement; is the service flow path, which is expressed as a sequence of a series of nodes and connections: ; where; represents a node in the network, 1 ≤ i ≤ k; and k ≥ 2.

[0055] In this embodiment, the virtual power plant service flow model is defined, including service type, priority, path, bandwidth requirement, and delay requirement, providing a more comprehensive information basis for routing and scheduling; by defining the service flow path as a sequence of a series of nodes and connections, the flexibility of routing and scheduling is increased, which can better adapt to complex network environments, and the service flow path is the network transmission path in the virtual power plant service flow model.

[0056] Here, based on the service priority of the newly received service request, update the current network state of the pre-constructed virtual power plant service flow model, including:

[0057] When a new service request is to be connected, according to the service priority of the new service request and the parameters of the virtual power plant service flow model , add the new service request to the corresponding service queue, update the current network state of the pre-built virtual power plant service flow model, and obtain the updated network state.

[0058] Specifically, perform iteration on the updated network state based on the Q-learning reinforcement learning method until the termination condition is met, and output the optimal routing and scheduling strategy, including:

[0059] Initialization value table , set all values to 0; define the state space and action space; set the learning rate , discount factor and exploration rate ;

[0060] Observe the updated network state ; where is the current network topology structure; is the service queue length of each node;

[0061] Action selection process: Use the greedy policy to select an action ;

[0062] Scheduling process: Perform service scheduling and path selection according to the selected action to meet the requirements of the virtual power plant service flow model;

[0063] State observation process: After the action is executed, observe the next network state and the current reward obtained ;

[0064] Value update process: Update the value according to the current value update formula;

[0065] Transfer the updated network state s to the next network state , repeat the action selection process, scheduling process, state observation process and value update process until the termination condition is met, and output the optimal routing and scheduling strategy;

[0066] Among them, the value update formula is as follows:

[0067]

[0068] Among them, is the updated network state; is the current action; is the learning rate; is the current reward; Discount factor; is the next network state; is the action optional in the next network state; is the maximum value of the next state.

[0069] Here, the updated network state is expressed as: ;

[0070] The current reward is expressed as: ; Among them, : represents different weight coefficients, used to adjust the weights of priority and delay in the current reward.

[0071] In this embodiment, by initializing the Q-value table and learning parameters, a good starting point is provided for the operation of the Q-learning algorithm; the ε-greedy greedy strategy is adopted to select actions, realizing intelligent service scheduling and path selection, and improving the intelligent level of network scheduling; by continuously iteratively updating the Q-value, the optimal routing scheduling strategy is gradually approximated, enhancing network performance and resource utilization. At the same time, the representation of the updated network state contains rich network information, which helps to more accurately evaluate the network state and formulate routing scheduling strategies; by representing the flexible current reward, the weights of factors such as priority and delay in the reward can be adjusted according to actual needs, realizing more refined scheduling control.

[0072] As Figure 4 shown, the present invention also provides a routing scheduling system based on the service priority of a virtual power plant, including: a service request access module, configured to update the current network state of a pre-constructed virtual power plant service flow model based on the service priority of an accessed new service request; wherein, the virtual power plant service flow model contains service priorities; a scheduling policy output module, configured to perform iteration on the updated network state based on the Q-learning reinforcement learning method until a termination condition is met, and output an optimal routing scheduling strategy; the virtual power plant service flow model is constructed based on service types, service priorities, service flow paths, service bandwidth requirements, and service delay requirements, and is specifically expressed as:

[0073]

[0074] Among them, is the service type; is the business priority; is the business bandwidth requirement; is the business latency requirement; is the business flow path, represented as a sequence of a series of nodes and connections: ; where; represents a node in the network, 1 ≤ i ≤ k; and k ≥ 2.

[0075] The present invention also provides a device, including: a memory for storing a computer program; a processor for implementing the steps of the above-mentioned routing and scheduling method based on the business priority of the virtual power plant when executing the computer program.

[0076] When the processor executes the computer program, it implements the steps of the above-mentioned routing and scheduling based on the business priority of the virtual power plant, for example: updating the current network state of the pre-constructed virtual power plant business flow model based on the business priority of the newly accessed service request; wherein, the virtual power plant business flow model contains business priorities; performing iteration on the updated network state based on the Q-learning reinforcement learning method until the termination condition is met, and outputting an optimal routing and scheduling strategy; the virtual power plant business flow model is constructed based on the service type, business priority, business flow path, business bandwidth requirement, and business latency requirement, and is specifically represented as:

[0077]

[0078] wherein, is the service type; is the business priority; is the business bandwidth requirement; is the business latency requirement; is the business flow path, represented as a sequence of a series of nodes and connections: ; where; represents a node in the network, 1 ≤ i ≤ k; and k ≥ 2.

[0079] Alternatively, when the processor executes the computer program, it implements the functions of each module in the above-mentioned system, for example: a service request access module for updating the current network state of the pre-constructed virtual power plant business flow model based on the business priority of the newly accessed service request; wherein, the virtual power plant business flow model contains business priorities; a scheduling strategy output module for performing iteration on the updated network state based on the Q-learning reinforcement learning method until the termination condition is met, and outputting an optimal routing and scheduling strategy; the virtual power plant business flow model is constructed based on the service type, business priority, business flow path, business bandwidth requirement, and business latency requirement, and is specifically represented as:

[0080]

[0081] Among them, is the service type; is the service priority; is the service bandwidth requirement; is the service latency requirement; is the service flow path, expressed as a sequence of a series of nodes and connections: ; among them; represents a node in the network, 1 ≤ i ≤ k; and k ≥ 2.

[0082] Exemplarily, the computer program can be divided into one or more modules / units, and the one or more modules / units are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of completing preset functions, and the instruction segments are used to describe the execution process of the computer program in the routing and scheduling device based on the service priority of the virtual power plant. For example, the computer program can be divided into a service request access module and a scheduling policy output module; the specific functions of each module are as follows: The service request access module is used to update the current network state of the pre-constructed virtual power plant service flow model based on the service priority of the newly accessed service request; among them, the virtual power plant service flow model contains service priorities; the scheduling policy output module is used to perform iterations on the updated network state based on the Q-learning reinforcement learning method until the termination condition is met, and output the optimal routing and scheduling policy; the virtual power plant service flow model is constructed based on the service type, service priority, service flow path, service bandwidth requirement, and service latency requirement, and is specifically expressed as:

[0083]

[0084] Among them, is the service type; is the service priority; is the service bandwidth requirement; is the service latency requirement; is the service flow path, expressed as a sequence of a series of nodes and connections: ; among them; represents a node in the network, 1 ≤ i ≤ k; and k ≥ 2.

[0085] The routing and scheduling device based on the business priority of the virtual power plant may be a computing device such as a desktop computer, a notebook, a palm computer, or a cloud server. The routing and scheduling device based on the business priority of the virtual power plant may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the above are examples of the routing and scheduling device based on the business priority of the virtual power plant, which do not constitute a limitation on the routing and scheduling device based on the business priority of the virtual power plant. It may include more components than the above, or combine some components, or different components. For example, the routing and scheduling device based on the business priority of the virtual power plant may also include input / output devices, network access devices, buses, etc.

[0086] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc. The processor is the control center of the routing and scheduling based on the business priority of the virtual power plant, and uses various interfaces and lines to connect all parts of the routing and scheduling device based on the business priority of the virtual power plant.

[0087] The memory can be used to store the computer programs and / or modules. The processor realizes various functions of the routing and scheduling device based on the business priority of the virtual power plant by running or executing the computer programs and / or modules stored in the memory, and by calling the data stored in the memory.

[0088] The memory may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0089] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the routing and scheduling method based on the service priority of a virtual power plant.

[0090] If the modules / units integrated in the routing and scheduling system based on the service priority of a virtual power plant are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.

[0091] Based on such an understanding, all or part of the processes in the routing and scheduling method based on the service priority of a virtual power plant of the present invention can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the routing and scheduling method based on the service priority of a virtual power plant can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or a preset intermediate form, etc.

[0092] The computer-readable storage medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0093] It should be noted that the content included in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable storage medium does not include electrical carrier signals and telecommunication signals.

[0094] The present invention will be further described below with reference to embodiments and the accompanying drawings:

[0095] Embodiment 2

[0096] As described in the background art, traditional routing and scheduling methods usually rely on fixed routing algorithms to guide the transmission path of data packets in the network. Most traditional routing and scheduling methods are based on static network topologies and predefined routing policies, aiming to efficiently and reliably transmit data in the network; however, in the face of complex service requirements and dynamic network conditions, the performance of traditional routing and scheduling methods will decline sharply, resulting in low utilization rate of network resources and security risks for services.

[0097] To solve the above problems, the present invention provides a routing and scheduling method based on the service priority of a virtual power plant. This method relies on a reinforcement learning optimization algorithm to achieve priority scheduling of virtual power plant service routing, improve the allocation efficiency of network resources and data transmission performance, and support the implementation and efficient operation of virtual power plants.

[0098] This embodiment provides a routing and scheduling method based on the service priority of a virtual power plant, which is oriented to virtual power plant services and applied under the virtual power plant service flow model. The construction of the virtual power plant service flow model is as follows:

[0099] As Figure 1 shown, Figure 1 is the system architecture of the virtual power plant service flow, including the virtual power plant service master station, the virtual power plant service terminal aggregation device, and the virtual power plant distributed resources; among them, the virtual power plant service aggregation node is connected to the distributed resources of the virtual power plant, such as distributed photovoltaics, distributed energy storage, flexible adjustable loads, etc., through a wired or wireless network downward to realize the access of devices; at the same time, the virtual power plant service aggregation node is connected to the virtual power plant service master station through the network infrastructure upward to support the virtual power plant service.

[0100] The virtual power plant data stream includes a collection stream and a control stream, which are used to support scheduling, marketing services, source-network-load-storage interaction, etc. The collection stream and the control stream have both periodic collection and non-periodic execution service characteristics, and usually adopt the 104 protocol and DL / T1867-2018. In this embodiment, based on the above system architecture of the virtual power plant service flow, a virtual power plant service flow model is constructed, which is specifically represented as:

[0101]

[0102] Among them, is the service type; is the service priority; is the service bandwidth requirement; is the service delay requirement; is the service flow path, which is represented as a sequence of a series of nodes and connections: ; among them; represents a node in the network, 1≤i≤k; and k≥2.

[0103] In this embodiment, in order to improve the allocation efficiency of network resources and data transmission performance, the service priority is defined as follows:

[0104] In this embodiment, the virtual power plant services are managed at four levels. Among them, the real-time service with the highest service priority is the first-level service; the quasi-real-time service is the second-level service; the non-real-time service is the third-level service; the default service has the lowest service priority, which is the fourth-level service, and the service priorities of the corresponding service requests are 1 to 4 levels respectively. When scheduling each service in the power communication network, it is necessary to consider the service priority of the newly added service request, and first schedule the service with the highest service priority to ensure the safe and stable operation of the power grid, and then schedule the service with a lower service priority, as shown in Table 1 specifically:

[0105] Table 1 is the classification of service priorities

[0106]

[0107] Among them, QoS (Quality of Service) is a network term used to describe the level of service quality provided by the network for specific services or users when transmitting data packets through the network. The goal of QoS is to provide different priorities and service qualities for different network traffic to meet the needs of different users.

[0108] This embodiment provides a routing scheduling method based on the service priority of the virtual power plant. This method is based on the Q-learning reinforcement learning algorithm, and the specific description is as follows:

[0109] The Q-learning reinforcement learning algorithm is a reinforcement learning algorithm based on value iteration. By learning the value of the state-action (Q value), it selects the optimal action. The present invention schedules various services in the power communication network based on the service priority and service flow to optimize the allocation and use of network resources; defines as the network state selects the behavior under and follows the estimated value of the optimal policy, and updates it by using immediate rewards and discounted rewards.

[0110] The update of the Q value is as follows:

[0111]

[0112] Among them, : the current network state; : the current action; : the learning rate; : the current reward; : the discount factor; : the next state; is the optional action in the next network state; : the maximum of the next state value.

[0113] Updated network status Indicates: . : Current network topology; : Business queue lengths of each node; : Priorities of each service, the smaller the value, the higher the priority (Level 1 is the highest, Level 4 is the lowest).

[0114] Current reward Indicates: . Indicates different weight coefficients, used to adjust the weights of priority and delay in the current reward in it.

[0115] Such as Figure 2 shown, this embodiment provides a routing and scheduling method based on the business priority of a virtual power plant, and the specific steps are as follows:

[0116] Step 1: Initialize value table , all values are set to 0. Define the state space and action space. Set the learning rate , discount factor and exploration rate .

[0117] Step 2: When a new service request accesses, according to its service priority and the parameters of the service flow model , add the new service request to the corresponding business queue .

[0118] Step 3: Observe the updated network status .

[0119] Step 4: Use greedy policy to select an action . Set a small value, use probability to greedily select the optimal action, and use probability to randomly select one from all possible optional actions.

[0120] Step 5: Perform service scheduling and path selection according to the selected action to meet the requirements of the virtual power plant service flow model.

[0121] Step 6: After performing the action, observe the next network status and the obtained current reward .

[0122] Step 7: Update the value according to the value update formula.

[0123] Step 8: Transfer the updated network state s to the next network state , and repeat Steps 4 to 7 until the termination condition is met. The termination condition here can be the number of iterations or the value converges; after reaching the termination condition, finally output the optimal routing and scheduling policy combination.

[0124] It should be noted that the order of Step 1 and Step 2 can be swapped or they can be carried out synchronously.

[0125] It can be seen that this method comprehensively considers the virtual power plant service traffic and service priorities, solves the problem of multi-concurrent flows in the virtual power plant network, and thus the problem of tight routing and network resource allocation caused by this phenomenon; at the same time, this method proposes a Q-learning routing and scheduling scheme based on the sum of flows, obtains the basic network information, uses an improved Q-learning algorithm to find paths, and improves the performance of the entire virtual power plant network.

[0126] The present invention provides a routing and scheduling method based on the service priorities of virtual power plants. Compared with traditional routing and scheduling methods, it has the following advantages:

[0127] This method provides a highly adaptive and intelligent routing and scheduling method for virtual power plant service representation. This method responds to new service requests by updating the network state in real time, and uses the Q-learning reinforcement learning algorithm to iteratively optimize the routing strategy, so as to achieve efficient and reliable routing and scheduling under the conditions of limited and dynamically changing network resources. The pre-constructed virtual power plant service flow model contains detailed information such as service priorities, types, paths, bandwidth requirements, and delay requirements, providing accurate guidance for routing and scheduling. This method effectively improves the utilization rate of network resources by intelligently allocating network resources to higher-priority services and meeting their specific bandwidth and delay requirements, while reducing the risk of service transmission. In addition, this method also ensures the continuous optimization and convergence of the routing and scheduling strategy through the finely represented current reward and Q-value update formula to adapt to the changing network environment and service requirements. In summary, this method not only improves the transmission efficiency and quality of virtual power plant services, but also enhances the stability and security of the entire network system.

[0128] The above embodiments are only one of the implementation manners that can implement the technical solution of the present invention. The scope of protection required by the present invention is not limited only by this embodiment, but also includes any changes, substitutions, and other implementation manners that are easily conceivable by those skilled in the art within the technical scope disclosed by the present invention.

Claims

1. A routing scheduling method based on virtual power plant service priority, characterized in that: include: Based on the service priority of the new service request received, the current network status of the pre-built virtual power plant service flow model is updated; wherein the service priority is included in the virtual power plant service flow model, including: When a new service request is received, the service priority P of the new service request and the parameters of the virtual power plant service flow model are calculated. , add the new service request to the corresponding service queue, update the current network status of the pre-built virtual power plant service flow model, and obtain the updated network status; Based on the Q-learning reinforcement learning method, the updated network status is iterated until the termination condition is met, and the optimal routing scheduling strategy is output, including: initialization Value Table ,all Set the value to 0; define the state space and action space; set the learning rate , Discount Factor and exploration rate ; Observe the updated network status ; Where T is the current network topology; is the service queue length of each node; Action selection process: Use Greedy strategy to choose actions ; Scheduling process: Based on the selected action Carry out business scheduling and path selection to meet the requirements of the virtual power plant business flow model; State Observation Process: Action After the execution is completed, observe the next network status and the current rewards obtained ; Value update process: According to the current Value Update Formula Update value; Transfer the updated network state s to the next network state , repeat the action selection process, scheduling process, state observation process and The value update process continues until the termination condition is met and the optimal routing scheduling strategy is output; Among them, the The value update formula is as follows: in, is the updated network status; For the current action; is the learning rate; For current rewards; is the discount factor; For the next network state; For optional actions in the next network state; The maximum value for the next state value; The virtual power plant service flow model is constructed based on service type, service priority, service flow path, service bandwidth requirement and service delay requirement, and is specifically expressed as follows: in, is the type of business; For business priorities; To meet business bandwidth requirements; Service latency requirements; is the business flow path, represented as a sequence of nodes and connections: ;in, represents a node in the network, 1≤i≤k; and k≥2.

2. The routing scheduling method based on virtual power plant service priority according to claim 1 is characterized in that: Current Rewards It is expressed as: ;in, : Represents different weight coefficients, used to adjust the priority and delay in the current reward The weight in .

3. The routing scheduling method based on virtual power plant service priority according to claim 1 is characterized in that: The virtual power plant business flow model uses a virtual power plant data flow for interaction; wherein the virtual power plant data flow includes an acquisition flow and a control flow.

4. The routing scheduling method based on virtual power plant service priority according to claim 1 is characterized in that: Virtual power plant services are managed at four levels, including real-time services, quasi-real-time services, non-real-time services and default services; the service priorities of service requests corresponding to the four-level virtual power plant services are 1 to 4.

5. A routing scheduling system based on virtual power plant service priority, used to implement the steps of the routing scheduling method based on virtual power plant service priority according to any one of claims 1 to 4, characterized in that: include: A service request access module, used to update the current network status of a pre-built virtual power plant service flow model based on the service priority of the new service request being accessed; wherein the virtual power plant service flow model includes the service priority; The scheduling strategy output module is used to iterate the updated network status based on the Q-learning reinforcement learning method until the termination condition is met and output the optimal routing scheduling strategy; The virtual power plant service flow model is constructed based on service type, service priority, service flow path, service bandwidth requirement and service delay requirement, and is specifically expressed as follows: in, is the type of business; For business priorities; To meet business bandwidth requirements; Service latency requirements; is the business flow path, represented as a sequence of nodes and connections: ;in, represents a node in the network, 1≤i≤k; and k≥2.

6. A routing scheduling device, characterized in that: include: Memory for storing computer programs; A processor is used to implement the steps of the routing scheduling method based on virtual power plant business priority as described in any one of claims 1-4 when executing the computer program.

7. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it is used to implement the steps of the routing scheduling method based on virtual power plant business priority as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Optimal routing scheduling method based on reinforcement learning for power wireless network

    CN116828548A