Wireless energy transmission equipment combination selection method and device based on deep reinforcement learning and ant colony optimization
Through deep reinforcement learning and ant colony optimization methods, an efficient wireless energy transmission combination solution is generated, which solves the problem of insufficient energy supply in outdoor high-intensity tasks, and realizes efficient energy transmission and equipment energy supply.
Patent Information
- Application Number
- CN202510304188.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-08-08
AI Technical Summary
In outdoor high-intensity tasks, the energy consumption of smart wearable devices has increased dramatically, resulting in frequent charging demands, while wireless charging efficiency is low and affected by charging distance and number of devices, it is difficult to meet energy supply needs.
Using a method based on deep reinforcement learning and ant colony optimization, the heuristic information selected by the device is generated through the graph neural network, pheromone concentration is initialized, and the candidate energy transmission combination scheme is generated using the ant colony optimization algorithm, and the maximum power scheme is selected through iterative optimization, triggering the transmission device to perform wireless energy transmission.
It realizes a combination solution of efficient calculation and high power wireless energy transmission, meets the equipment energy supply needs, and improves the efficiency of the combination solution generation and the system's real-time response capabilities.
Smart Images

Figure CN120447713A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wireless charging technology, and in particular to a method and device for selecting a combination of wireless energy transmission equipment based on deep reinforcement learning and ant colony optimization. Background Art
[0002] In recent years, smart wearable devices, as sensory information devices that can be worn on the human body, have penetrated into fields such as healthcare, smart education, and gaming and entertainment. These devices have not only changed people's daily lives but also significantly improved their quality of life, gradually becoming an indispensable part of human life. However, when users wear these devices for high-intensity outdoor tasks (such as real-time location tracking or health monitoring), the device's energy consumption increases dramatically, requiring users to charge frequently. However, in most cases, users often find it difficult to find available charging equipment in outdoor environments far away from traditional wired charging facilities.
[0003] Researchers have proposed a radio frequency-based wireless charging technology that offers a convenient, efficient, safe, and low-cost energy supply solution. However, its charging efficiency is still lower than wired charging and is affected by charging distance and the number of transmitting devices. To meet the energy supply needs of devices, multiple wireless charging points need to be combined.
[0004] Therefore, how to design an efficient algorithm to calculate a high-power wireless energy transmission combination solution becomes the key to solving this problem. Summary of the Invention
[0005] In view of this, an embodiment of the present application provides a method and device for selecting a combination of wireless energy transmission equipment based on deep reinforcement learning and ant colony optimization, which can efficiently calculate a high-power wireless energy transmission combination solution to meet the energy supply needs of the equipment.
[0006] A first aspect of an embodiment of the present application provides a method for selecting a combination of wireless energy transmission devices based on deep reinforcement learning and ant colony optimization, comprising:
[0007] Input the transmission device parameters and the receiving device parameters, and generate heuristic information for device selection through the pre-trained graph neural network model;
[0008] Initialize the selection probability matrix of all ants and the pheromone concentration between transmission devices;
[0009] Using the heuristic information and pheromone concentration, generating candidate energy transmission combination schemes through an ant colony optimization algorithm;
[0010] Selecting the optimal solution with the maximum energy transmission power from the candidate combination solutions through iterative optimization;
[0011] The transmitting device in the optimal solution is triggered to perform wireless energy transmission to the receiving device.
[0012] A second aspect of an embodiment of the present application provides a device for selecting a combination of wireless energy transmission devices based on deep reinforcement learning and ant colony optimization, comprising:
[0013] A graph neural network model module, which is used to input transmission device parameters and receiving device parameters and generate heuristic information for device selection through a pre-trained graph neural network model;
[0014] Initialization module, used to initialize the selection probability matrix of all ants and the pheromone concentration between transmission devices;
[0015] An ant colony optimization calculation module, configured to utilize the heuristic information and pheromone concentration to generate candidate energy transmission combination schemes through an ant colony optimization algorithm;
[0016] An iterative calculation module, configured to select an optimal solution with the maximum energy transmission power from the candidate combination solutions through iterative optimization;
[0017] The output module is used to trigger the transmitting device in the optimal solution to perform wireless energy transmission to the receiving device.
[0018] A third aspect of an embodiment of the present application provides an electronic device, comprising a processor, a memory, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, the electronic device implements the wireless energy transmission device combination selection method based on deep reinforcement learning and ant colony optimization as provided in the first aspect of an embodiment of the present application.
[0019] A fourth aspect of the embodiments of the present application provides a computer program product, including a computer program. When the computer program is executed, the method according to the first aspect of the embodiments of the present application is executed.
[0020] The first aspect of the embodiments of the present application provides a method for selecting a combination of wireless energy transmission devices based on deep reinforcement learning and ant colony optimization. By inputting transmission device parameters and receiving device parameters, heuristic information for device selection is generated through a pre-trained graph neural network model; the selection probability matrix of all ants and the pheromone concentration between transmission devices are initialized; the heuristic information and pheromone concentration are used to generate candidate energy transmission combination schemes through an ant colony optimization algorithm; the optimal scheme with the maximum energy transmission power is selected from the candidate combination schemes through iterative optimization; and the transmission device in the optimal scheme is triggered to perform wireless energy transmission to the receiving device.
[0021] It can be understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0023] Figure 1 This is a flow chart of a method for selecting a combination of wireless energy transmission devices based on deep reinforcement learning and ant colony optimization provided in one embodiment of the present application;
[0024] Figure 2 This is a combined optimization framework diagram of a wireless energy transmission device combination selection method based on deep reinforcement learning and ant colony optimization provided by another embodiment of the present application;
[0025] Figure 3 This is a flowchart of a method for selecting a combination of wireless energy transmission devices based on deep reinforcement learning and ant colony optimization, provided in another embodiment of the present application;
[0026] Figure 4 This is a flowchart of a method for selecting a combination of wireless energy transmission devices based on deep reinforcement learning and ant colony optimization, provided in another embodiment of the present application;
[0027] Figure 5 This is an outdoor scene diagram provided by an embodiment of the present application;
[0028] Figure 6 This is a flowchart of a method for selecting a combination of wireless energy transmission devices based on deep reinforcement learning and ant colony optimization, provided in another embodiment of the present application;
[0029] Figure 7 is a schematic diagram of the DRL-A model of this application;
[0030] Figures 8-10 The experimental results of each algorithm on datasets of size 50, 100 and 200 are respectively;
[0031] Figures 11-13 These are schematic diagrams comparing the performance of each algorithm under different problem scales;
[0032] Figure 14 This is a structural diagram of a wireless energy transmission device combination selection device based on deep reinforcement learning and ant colony optimization provided in an embodiment of the present application;
[0033] Figure 15 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0034] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0035] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0036] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0037] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0038] like Figure 1 As shown, the wireless energy transmission device combination selection method based on deep reinforcement learning and ant colony optimization provided in the embodiment of the present application includes the following steps S101 to S105:
[0039] Step S101: Input transmission device parameters and receiving device parameters, and generate heuristic information for device selection through a pre-trained graph neural network model.
[0040] Step S102: Initialize the selection probability matrix of all ants and the pheromone concentration between transmission devices.
[0041] Step S103: using the heuristic information and the pheromone concentration, generating candidate energy transmission combination schemes through an ant colony optimization algorithm.
[0042] Step S104: selecting the optimal solution with the maximum energy transmission power from the candidate combination solutions through iterative optimization.
[0043] Step S105: triggering the transmitting device in the optimal solution to perform wireless energy transmission to the receiving device.
[0044] In applications, such as Figure 2 As shown, the overall framework of the wireless energy transmission device combination selection method based on deep reinforcement learning and ant colony optimization provided by the embodiment of the present application is mainly composed of four modules: outdoor environment module, RF-SWD module, DRL-A algorithm module and solution module. The outdoor environment module describes that in the OETC-SWD problem environment, the energy provider can provide the following data: Number of Transmission Devices, Energy Capacity of Transmission Device and Transmission Device Emission Power. The energy demander provides Capacity Requirements of Receiving Devices data and Device Distance data. These data will be used to calculate the transmission power and provide data support for the subsequent energy combination optimization steps.
[0045] The RF-SWD module uses the Device Distance data in the problem environment and the Transmission Device Emission Power data for each device to calculate the energy transmission power received by the demand device from each transmission device. This energy transmission power data is then integrated with other data from the problem environment to form a complete dataset for the OETC-SWD problem. This dataset serves as a training source for the algorithm and also tests its effectiveness in solving the problem.
[0046] The DRL-A algorithm module consists of two submodules: the DRL submodule and the ACO submodule. The DRL submodule extracts data features from energy transmission devices and maps them into a neural network to generate heuristic information for device selection. The ant colony algorithm uses the selection probabilities in this heuristic information and default pheromones to construct a solution. The constructed solution is then evaluated for quality and compared with the average solution. The comparison result is converted into a loss and fed back into the graph neural network for optimization. This process is repeated until the set number of iterations is reached. The ACO submodule inputs the problem instance data into the graph neural network trained by the DRL submodule to generate heuristic information for that instance and initialize relevant parameters such as the ant colony and pheromones. This heuristic information, combined with the initial pheromones, guides the ant colony algorithm in constructing a solution. After each iteration, the pheromone values are updated based on the device information selected in the solution. Through repeated iterations, the module continuously optimizes the solution, ultimately selecting the optimal energy combination and its corresponding maximum energy transmission power.
[0047] The solution module obtains the optimal energy combination solution calculated by the algorithm module and issues energy transmission instructions to the corresponding energy providers based on the solution. Ultimately, energy users receive the maximum energy transmission power service while fully meeting their energy needs.
[0048] This embodiment of the present application integrates deep reinforcement learning with an ant colony optimization algorithm to automate the selection of wireless energy transmission device combinations. By using a pre-trained graph neural network to generate heuristic information, this approach avoids the limitations of traditional ant colony algorithms, which rely on manually designed heuristic rules, and significantly improves the efficiency of generating combination solutions. Iterative optimization directly outputs a maximum transmission power solution and triggers the device to execute the transmission, forming a closed-loop process from decision-making to execution, enhancing the system's real-time responsiveness in dynamic environments.
[0049] In one embodiment, Figure 3 As shown, step S104 includes the following steps:
[0050] Step S1041: The ant selects a transmission device based on the current pheromone concentration and heuristic information.
[0051] In the application, all ants will select a device from the dataset based on pheromone and heuristic information. The probability P of the ant selecting the next device i is (i|j) The calculation formula is as follows:
[0052]
[0053] Among them, τ ij represents the pheromone concentration between device i and device j, η ijrepresents the heuristic information between device i and device j, U J represents the set of all unselected devices that can be selected after selecting device i. α and β are control parameters that adjust the weighting relationship between heuristic information and pheromone concentration. This probability calculation formula balances the influence of pheromone concentration and heuristic information through exponential weighting, dynamically adjusting the algorithm's exploration and exploitation balance. The introduction of α and β parameters makes the algorithm adaptable to different problem characteristics, retaining the positive feedback mechanism of the ant colony algorithm while incorporating learned device selection rules. This dual guidance mechanism effectively avoids local optimality traps and significantly increases the probability of discovering the global optimal solution.
[0054] In the application, when traversing the combination, the concentration of pheromones released by ants is related to the power of the energy transmission combination of the device. The pheromones of all combinations will gradually evaporate over time. The pheromone update formula is as follows:
[0055] τ′ ij =(1-ρ)·τ ij +Δτ ij , 0<ρ<1;
[0056]
[0057] Among them, τ′ ij represents the latest pheromone concentration between device i and device j, ρ represents the volatility of pheromone, τ ij represents the old pheromone concentration between device i and device j, Δτ ij represents the newly added pheromone concentration of all ants between device i and device j, represents the newly added pheromone concentration between device i and device j by the ωth ant, Q represents the proportional factor controlling the pheromone increment, and f ω (s) represents the total energy transmission power of the devices in the energy combination solution selected by the ωth ant. This pheromone update mechanism uses the volatility coefficient ρ to gradually forget outdated paths while dynamically enhancing the mark of high-quality paths based on solution quality. The incremental design, in which the Q factor is positively correlated with the transmission power, ensures that high-quality solutions receive stronger pheromone reinforcement. This adaptive update strategy maintains population diversity while accelerating the propagation of high-quality solutions, enabling the algorithm to quickly focus on high-performance device combinations.
[0058] Step S1042: Dynamically calculate the total energy capacity of the selected devices:
[0059] If the energy demand of the receiving device is not met, the next transmitting device will be selected;
[0060] If the demand is reached or exceeded, the current ant selection process is terminated.
[0061] In applications, such as Figure 4 As shown, if the energy capacity C in the current combination scheme s is full, the device selection step is ended. If there is still remaining space, the pre-selected devices are added to the combination scheme and the current energy capacity C is updated.
[0062] Step S1043: Record the candidate combination solutions generated by the current ants.
[0063] Step S1044: update the pheromone concentration and iteration status, and return to the step of executing the ants selecting the transmission device according to the current pheromone concentration and heuristic information until the iteration termination condition is met.
[0064] In application, the candidate energy transmission combination scheme must meet the following conditions:
[0065] Select a subset from the set of transmitting devices so that the total energy transmission of the subset meets the needs of the receiving device and maximizes the total energy transmission power;
[0066] The combined relationships between transmission devices are represented by a graph structure, with nodes in the graph representing transmission devices. By modeling the combined relationships between devices using a graph structure, the combinatorial optimization problem is transformed into a graph node selection problem, effectively capturing the synergistic effects between devices. While meeting energy demand constraints, the objective function is to maximize transmission power, ensuring that the selected solution is both functional and economical. This structured modeling approach improves the solvability of complex constrained problems and provides an optimized mathematical framework for coordinated energy supply across multiple devices.
[0067] This embodiment of the application achieves precise control under resource constraints by dynamically calculating the total energy capacity of selected devices and setting termination conditions. This method determines the energy demand fulfillment status in real time during the ant selection process, preventing resource waste caused by overselection while ensuring that the combined solution strictly meets the receiver's requirements. The combination of step-by-step selection and capacity verification improves the algorithm's feasibility and convergence speed in complex constraint scenarios.
[0068] In one embodiment, the method further includes training a graph neural network model, including:
[0069] Step S201: Acquire historical transmission device parameters, including the number of transmission devices, energy capacity, and transmission power.
[0070] Step S202: Obtain corresponding historical receiving device demand parameters, including the energy demand of the receiving device and the device distance.
[0071] In the application, an OETC-SWD problem environment is constructed, and multiple energy providers are deployed in an outdoor environment, each of which is equipped with wireless transmission equipment. Figure 5This figure shows an outdoor wireless environment where multiple energy providers (#1, #2, #3, and #4) provide energy to a single energy user via wireless transmission technology. The figure uses different icons to represent the selected energy transmission device, unselected devices, wireless transmission radio frequency technology, and energy capacity of the energy transmission device.
[0072] Step S203: Calculate the energy transmission power of each transmission device based on the acquired data.
[0073] Step S204: Integrate historical transmission equipment parameters, historical receiving equipment requirement parameters, and the calculated energy transmission power into a structured data set.
[0074] Step S205: Obtain a graph neural network model based on training of the structured data set.
[0075] In the application, the energy transmission power characteristics, energy capacity characteristics, and inter-device selection relationship characteristics are extracted from the energy transmission equipment through feature extraction. These three features are processed through linear transformation, feature aggregation, activation function, and normalization to generate an embedded representation of the device. This embedded information is input into the input layer of the multi-layer perceptron.
[0076] In the hidden layer, the embedded information is further processed through linear transformation and activation function to generate heuristic information prediction values;
[0077] The output layer compresses the predicted value into the range of 0 to 1 through the sigmoid function to obtain the selection probability of the device, that is, the heuristic information of device selection.
[0078] This embodiment of the present application provides multidimensional feature input for graph neural networks by constructing a structured dataset and integrating historical transmission parameters with real-time computation results. This design enables the model to learn the underlying associations between devices, enhancing the generalization capabilities of heuristic information generation. The training process based on historical data effectively captures the relationship between device performance and spatial distribution, significantly improving the algorithm's adaptability to new scenarios and its decision-making reliability.
[0079] In one embodiment, Figure 6 As shown, it also includes:
[0080] After each iteration, the loss function is calculated
[0081]
[0082]
[0083] Update the parameters of the graph neural network model through backpropagation to tune the graph neural network model;
[0084] Among them, θ represents the training parameters of the network, N represents the number of trajectories generated in a batch, and f(s i ) represents the total energy transfer value of the solution found by the i-th ant, b is the average value of the total energy transfer value of all solutions, a i,t represents the device selected by the ant i at step t in its trajectory, z i,t represents the state of the i-th ant at step t, including the state of the selectable equipment and the state of the remaining energy demand, π θ (a i,t |z i,t ) means in state z i,t The probability of each device being selected is , T represents the total number of steps of selecting devices in each ant’s trajectory, logπ θ (a i,t |z i,t ) indicates that the i-th ant selects device a at step t i,t The logarithmic probability of .
[0085] In the application, after calculating the loss function, the gradient of the loss with respect to the network parameters is transferred layer by layer to the weights and biases of each layer through the backpropagation algorithm, including the part of extracting features and generating heuristic information in the graph neural network.
[0086] In the application, network parameters are updated using the AdamW optimization algorithm, which employs an adaptive weight adjustment strategy and incorporates momentum and regularization mechanisms to ensure rapid model convergence and prevent overfitting. The updated parameters are used to generate new heuristic information. The graph neural network extracts relationships between devices, generating a more accurate probability distribution for device selection. These probability distributions are then combined with pheromones and used by the ant algorithm to construct a solution. Consequently, as the algorithm continues to iterate and optimize, the neural network's ability to generate heuristic information gradually improves.
[0087] The embodiment of this application uses a loss function design based on the REINFORCE algorithm, directly optimizing the expected return of the combined solution through policy gradient updates. Using the mean of the batch trajectory as a baseline effectively reduces the variance of the policy update and accelerates model convergence. The combination of backpropagation and the AdamW optimizer enables refined adjustment of neural network parameters, improving training stability while preventing overfitting and ensuring continuous improvement in the quality of heuristic information generation.
[0088] In applications, such as Figure 7 As shown in Figure 3, the DRL-A model consists of two parts: the training model and the inference model.
[0089] 1. In the training model, we define a real-world energy transmission device combination sequence as a graph structure to represent the problem instance. Each node in the graph represents an energy transmission device. Feature extraction is used to extract the energy transmission power characteristics, energy capacity characteristics, and inter-device selection relationship characteristics from the energy transmission devices. These three features are processed through linear transformation, feature aggregation, activation function, and normalization to generate an embedded representation of the device. This embedded information is input into the input layer of the multilayer perceptron.
[0090] In the hidden layer, the embedded information is further processed through linear transformation and activation function to generate heuristic information prediction values. Finally, the output layer uses the sigmoid function to compress the prediction value into the range of 0 to 1, resulting in the device selection probability, which is the heuristic information for device selection.
[0091] In the experiment, the number of layers of the graph neural network is set to 12, and the node features of the kth layer are set to Edge Features<i,j> Set as The feature representation formula of the k+1th layer is as follows:
[0092]
[0093] in, are learnable parameters, SiLU represents SiLU activation function, BN represents batch normalization, NA represents the neighborhood aggregation function, and in this study, NA represents mean pooling. represents the sigmoid activation function, ⊙ is the Hadamard product.
[0094] The extracted edge features are mapped to embedded node features using a 3-layer multi-layer perceptron (MLP) with skip connections The output layer uses the sigmoid function to normalize the output, and other layers use SiLU as the activation function.
[0095] 2. In the inference phase of the model, the input is a set of energy transmission devices, each device corresponds to a node in the graph, and the goal is to find the optimal energy transmission combination solution. The inference process consists of the following key steps:
[0096] a. The trained graph neural network acts as a heuristic "expert" and uses the trained model to generate a heuristic information matrix, which is combined with the pheromone matrix to guide the construction of the solution.
[0097] b. During the solution construction process, the ant algorithm gradually selects actions based on the weighted calculation rules of heuristic information and pheromones to build a complete solution. The device selection at each step follows the ant algorithm's probabilistic model to balance exploration and exploitation.
[0098] c. After the construction is completed, the pheromone matrix will be dynamically updated. The evaporation mechanism gradually weakens the pheromone strength of the old path to prevent the model from falling into the local optimum, while the enhancement mechanism increases the pheromone strength of the path with high-quality solutions to further improve its selection probability.
[0099] d. Through multiple iterations, the model generates multiple possible combinations and evaluates the energy transfer efficiency of each. Finally, the optimal solution with the highest energy transfer efficiency is selected from all generated combinations. This optimal solution is used for actual energy reception and transmission operations, achieving efficient resource allocation.
[0100] The entire inference process leverages dynamic pheromone updates and heuristic information generated by the trained model to ensure the quality of the solution is continuously optimized. Ultimately, this process provides an accurate and efficient optimal solution for the energy transmission device, improving the efficiency of energy transmission.
[0101] To validate the effectiveness of our method, we conducted comparative experiments using three classic metaheuristic methods: AntSystem (AS), Elitist Ant System (EAS), and MAX-MIN Ant System (MMAS). Compared to the original ant colony system, EAS additionally enhances the pheromones along the optimal path at each iteration, accelerating convergence and improving the algorithm's optimization capabilities. MMAS, on the other hand, builds on the original ant colony system by imposing upper and lower limits on pheromone concentration. These restrictions help the algorithm avoid premature convergence to local optima while maintaining its exploration capabilities. We implemented the DRL-A algorithm based on these three methods and evaluated it on test cases. Figures 8-10 The experimental results of each algorithm on datasets of 50, 100, and 200 are shown. The horizontal axis represents the number of algorithm iterations, and the vertical axis represents the optimal energy combination transmission value found by the algorithm in the test dataset. eli represents the algorithm based on EAS, and min_max represents the algorithm based on MMAS. The results show that the DRL-A algorithm consistently outperforms other ant colony algorithms on datasets of varying problem sizes.
[0102] In addition, the performance of the POMO algorithm and the greedy algorithm were evaluated on the same test set, and their results were compared with the DRL-A algorithm. Figures 11-13The performance of each algorithm at different problem scales is shown, with the horizontal axis representing the algorithm and the vertical axis representing the maximum transmission power of the optimal energy transmission solution calculated by the algorithm. The title corresponds to the different problem scales. Experimental results show that the DRL-A algorithm can find the optimal energy combination solution at all tested scales.
[0103] The software and hardware environments tested are as follows:
[0104] (1)CPU: Intel(R)Core(TM)i9-9900K 3.60GHz
[0105] (2) Graphics card: NVIDIA GeForce RTX 2080Ti
[0106] Experimental Data: Based on the problem size, data for each test case is generated as the energy transfer capacity and energy transfer power of the transmission equipment. The test dataset designed for the OETC-SWD problem contains 100 test cases.
[0107] Experimental parameter settings: The number of ant colonies and the number of iterations are set according to the problem scale and the convergence performance of the algorithm. Table 1 provides other specific parameter information:
[0108] Table 1 Experimental parameters
[0109]
[0110]
[0111] We compared other popular algorithms, including the POMO algorithm, Greedy algorithm, and the primitive ant colony algorithm (ACO) [16, 53]. We also organized the benchmark test results into a table and introduced the optimal value to facilitate a clearer comparison of the advantages and disadvantages of each algorithm. The benchmark test results are shown in Table 2:
[0112] Table 2 Baseline test results
[0113]
[0114] In Table 2, Value represents energy transmission power. A larger value indicates a higher energy transmission power for the combined solution under the same energy transmission requirements. Gap represents the gap between the maximum energy transmission power calculated by the algorithm and the actual optimal solution. A smaller value indicates a smaller gap from the optimal energy transmission power value.
[0115] These benchmark results further validate the efficiency of the DRL-A algorithm. In summary, the DRL-A algorithm can quickly converge to the optimal solution within a relatively small number of iterations and outperforms similar algorithms on problems of varying scales without changing the network model or heuristic parameters.
[0116] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0117] The present application also provides an apparatus for selecting a wireless energy transmission device combination based on deep reinforcement learning and ant colony optimization, which is configured to execute the steps described in the aforementioned method for selecting a wireless energy transmission device combination based on deep reinforcement learning and ant colony optimization. The apparatus for selecting a wireless energy transmission device combination based on deep reinforcement learning and ant colony optimization can be a virtual appliance within an electronic device, executed by a processor within the electronic device, or can be the electronic device itself.
[0118] like Figure 14 As shown, the wireless energy transmission device combination selection device 100 based on deep reinforcement learning and ant colony optimization provided in an embodiment of the present application includes:
[0119] A graph neural network model module 101 is configured to input transmission device parameters and receiving device parameters and generate heuristic information for device selection through a pre-trained graph neural network model;
[0120] Initialization module 102, used to initialize the selection probability matrix of all ants and the pheromone concentration between transmission devices;
[0121] The ant colony optimization calculation module 103 is used to generate candidate energy transmission combination schemes through an ant colony optimization algorithm using heuristic information and pheromone concentration;
[0122] Iterative calculation module 104, configured to select an optimal solution with the maximum energy transmission power from candidate combination solutions through iterative optimization;
[0123] The output module 105 is used to trigger the transmitting device in the optimal solution to perform wireless energy transmission to the receiving device.
[0124] In application, each module in the wireless energy transmission equipment combination selection device based on deep reinforcement learning and ant colony optimization can be a software program module, or can be implemented through different logic circuits integrated in the processor, or can be implemented through multiple distributed processors.
[0125] like Figure 15As shown, the embodiment of the present application further provides an electronic device 200, including: at least one processor 201 ( Figure 15 Only one processor is shown in the figure), a memory 202, and a computer program 203 stored in the memory 202 and executable on at least one processor 201. When the processor 201 executes the computer program 203, the steps in the above-mentioned various method embodiments are implemented.
[0126] In applications, electronic devices may include, but are not limited to, processors and memories. Those skilled in the art will appreciate that Figure 15 The electronic device is merely an example and does not limit the electronic device. The electronic device may include more or fewer components than shown in the figure, or may include a combination of certain components or different components.
[0127] In applications, the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0128] In applications, in some embodiments, the memory can be an internal storage unit of an electronic device, such as a hard disk or memory of the electronic device. In other embodiments, the memory can also be an external storage device of the electronic device, such as a plug-in hard disk equipped on the electronic device, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card, etc. Furthermore, the memory can also include both an internal storage unit of the electronic device and an external storage device. The memory is used to store an operating system, application programs, a boot loader (BootLoader), data, and other programs, such as the program code of a computer program. The memory can also be used to temporarily store data that has been output or is about to be output.
[0129] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.
[0130] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0131] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.
[0132] An embodiment of the present application provides a computer program product, including a computer program. When the computer program product runs on an electronic device, the electronic device can implement the steps in the above-mentioned various method embodiments when executing the computer program product.
[0133] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. Among them, the computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device that can carry the computer program code to the device / electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electric carrier signal, a telecommunication signal and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, a computer-readable medium cannot be an electric carrier signal or a telecommunication signal.
[0134] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0135] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0136] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0137] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0138] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A method for selecting a combination of wireless energy transmission devices based on deep reinforcement learning and ant colony optimization, characterized in that: include: Input the transmission device parameters and the receiving device parameters, and generate heuristic information for device selection through the pre-trained graph neural network model; Initialize the selection probability matrix of all ants and the pheromone concentration between transmission devices; Using the heuristic information and pheromone concentration, generating candidate energy transmission combination schemes through an ant colony optimization algorithm; Selecting the optimal solution with the maximum energy transmission power from the candidate combination solutions through iterative optimization; The transmitting device in the optimal solution is triggered to perform wireless energy transmission to the receiving device.
2. The method for selecting a combination of wireless energy transmission devices based on deep reinforcement learning and ant colony optimization according to claim 1, wherein: The selecting the optimal solution with the maximum energy transmission power from the candidate combination solutions through iterative optimization includes: The ants select a transmission device based on the current pheromone concentration and the heuristic information; Dynamically calculate the total energy capacity of the selected equipment: If the energy demand of the receiving device is not met, the next transmitting device will be selected; If the demand is reached or exceeded, the current ant selection process is terminated; Record the candidate combination solutions generated by the current ant; The pheromone concentration and iteration status are updated, and the step of selecting a transmission device according to the current pheromone concentration and the heuristic information is returned to execution until the iteration termination condition is met.
3. The method for selecting a combination of wireless energy transmission devices based on deep reinforcement learning and ant colony optimization according to claim 1, wherein: Also includes: Obtain historical transmission equipment parameters, including the number of transmission equipment, energy capacity, and transmission power; Obtain the corresponding historical receiving device demand parameters, including the energy demand of the receiving device and the device distance; Calculating the energy transmission power of each transmission device based on the acquired data; Integrate historical transmission equipment parameters, historical receiving equipment demand parameters, and calculated energy transmission power into a structured data set; A graph neural network model is obtained by training based on the structured dataset.
4. The method for selecting a combination of wireless energy transmission devices based on deep reinforcement learning and ant colony optimization according to claim 1, wherein: Also includes: After each iteration, the loss function is calculated Updating the parameters of the graph neural network model through back propagation to tune the graph neural network model; Among them, θ represents the training parameters of the network, N represents the number of trajectories generated in a batch, and f(s i ) represents the total energy transfer value of the solution found by the i-th ant, b is the average value of the total energy transfer value of all solutions, a i,t represents the device selected by the ant i at step t in its trajectory, z i,t represents the state of the i-th ant at step t, including the state of the selectable equipment and the state of the remaining energy demand, π θ (a i,t |z i,t ) means in state z i,t The probability of each device being selected is , T represents the total number of steps of selecting devices in each ant’s trajectory, logπ θ (a i,t |z i,t ) indicates that the i-th ant selects device a at step t i,t The logarithmic probability of .
5. The method for selecting a combination of wireless energy transmission devices based on deep reinforcement learning and ant colony optimization according to claim 2, wherein: When the ant selects a transmission device based on the current pheromone concentration and the heuristic information, the probability P of the ant selecting the next device i is (i|j) The calculation formula is as follows: Among them, τ ij represents the pheromone concentration between device i and device j, η ij represents the heuristic information between device i and device j, U J represents the set of all unselected devices that can be selected after device i is selected. α and β are control parameters used to adjust the weight relationship between heuristic information and pheromone concentration.
6. The method for selecting a combination of wireless energy transmission devices based on deep reinforcement learning and ant colony optimization according to claim 2, wherein: When the ants select a transmission device based on the current pheromone concentration and the heuristic information, the pheromone update formula is as follows: the ij =(1-ρ)·τ ij +Δt ij ,0<ρ<1; Among them, τ′ ij represents the latest pheromone concentration between device i and device j, ρ represents the volatility of pheromone, τ ij represents the old pheromone concentration between device i and device j, Δτ ij represents the newly added pheromone concentration of all ants between device i and device j, represents the newly added pheromone concentration between device i and device j by the ωth ant, Q represents the proportional factor controlling the pheromone increment, and f ω (s) represents the total energy transmission power of the devices in the energy combination scheme selected by the ωth ant.
7. The method for selecting a combination of wireless energy transmission devices based on deep reinforcement learning and ant colony optimization according to claim 1, wherein: The candidate energy transmission combination scheme must meet the following conditions: Select a subset from the set of transmitting devices so that the total energy transmission of the subset meets the needs of the receiving device and maximizes the total energy transmission power; The combination relationship between transmission devices is represented by a graph structure, and the nodes in the graph represent transmission devices.
8. A wireless energy transmission equipment combination selection device based on deep reinforcement learning and ant colony optimization, characterized in that: include: A graph neural network model module, which is used to input transmission device parameters and receiving device parameters and generate heuristic information for device selection through a pre-trained graph neural network model; Initialization module, used to initialize the selection probability matrix of all ants and the pheromone concentration between transmission devices; An ant colony optimization calculation module, configured to utilize the heuristic information and pheromone concentration to generate candidate energy transmission combination schemes through an ant colony optimization algorithm; An iterative calculation module, configured to select an optimal solution with the maximum energy transmission power from the candidate combination solutions through iterative optimization; The output module is used to trigger the transmitting device in the optimal solution to perform wireless energy transmission to the receiving device.
9. An electronic device, characterized in that: The electronic device comprises a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the electronic device implements the method according to any one of claims 1 to 7.
10. A computer program product, characterized in that The invention comprises a computer program, which, when being executed, enables the method according to any one of claims 1 to 7 to be performed.