A reinforcement learning enhanced adaptive large neighborhood search method
Patent Information
- Application Number
- CN202511311325.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2026-09-22
AI Technical Summary
权重更新效率低:传统ALNS依赖手工设计的算子选择器,缺乏对长期收益的预测,导致权重更新盲目
1、收益提升:实验表明,RL-ALNS在200-600任务规模下,平均利润率较ALNS/TPF提升2.31%,较传统ALNS提升36.52%。
Smart Images

Figure CN122797682A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of satellite communication resource scheduling, specifically involving an adaptive large neighborhood search method enhanced by reinforcement learning. Background Technology
[0002] In existing technologies, the Adaptive Large Neighborhood Search (ALNS) algorithm is the mainstream method for solving AEOSSP, but it has the following drawbacks: Low efficiency of weight updates: Traditional ALNS relies on manually designed operator selectors, which lack prediction of long-term returns, resulting in blind weight updates.
[0003] Limited task insertion location: Traditional methods only use fixed insertion rules (such as minimum insertion cost), which lacks diversity and limits the adaptability of the solution.
[0004] Local Optimality Trap: Existing methods are prone to getting trapped in local optima in complex scenarios, resulting in low search efficiency.
[0005] The existing solution proposes a tabu-based adaptive large neighborhood search algorithm (ALNS / TPF), which improves performance by introducing a tabu strategy and fast insertion (FI), but still does not solve the above problems.
[0006] Therefore, there is an urgent need for a more intelligent weight update mechanism and a diversified insertion position selection strategy to improve the efficiency of operator weight update, avoid blind search and local optima, provide diversified task insertion position selection strategies, enhance the adaptability of the algorithm to different scenarios, and improve the overall benefit of the scheduling scheme by dynamically optimizing the search process through reinforcement learning. Summary of the Invention
[0007] To address the problems of existing technologies, this invention provides an adaptive large neighborhood search method with reinforcement learning enhancement (RL-ALNS). The core components of this algorithm are a reinforcement learning controller, neighborhood operators (including destruction and repair operators), and a position operator.
[0008] The reinforcement learning-enhanced controller employs a temporal difference (TD) learning mechanism to evaluate the long-term potential of position operators and neighborhood operators rather than their immediate performance. This enables it to accurately determine the applicability of operators in different scenarios and identify their performance differences at each stage of the search process, quickly adjusting weights to prioritize the most suitable operator.
[0009] The diverse operators construct an "operator library" that covers the needs of multiple scenarios, providing rich policy options for reinforcement learning booster controllers and ensuring that they can still make efficient decisions in complex environments.
[0010] Specifically, the algorithm framework is as follows:
[0011] Specifically SA initialization phase: The algorithm first constructs an initial feasible solution, denoted as the solution vector. And create a taboo list Used to record prohibited operations to avoid duplicate searches. Sets the temperature decay function. Control the temperature drop curve of the simulated annealing process to determine the maximum number of iterations. As one of the termination conditions, the reward set is also initialized. ,in, This represents the fourth level of reward intensity and is a set of destruction operators. Repair operator set and position operator set Assign initial weight vectors respectively .
[0012] SB Iterative Optimization Loop: When the termination condition is not met, the following steps are executed repeatedly: SB1 Operator Dynamic Selection: Based on the roulette wheel (abbreviated as RS in pseudocode), the destructive operator for the current iteration is selected from three sets of operators. Repair operator and position operator Its selection probability is determined by the weight vector. Decide.
[0013] SB2 Deconstruction Process: First, the current solution... Execute the destruction operator Remove some elements, then repair the operator. Reconstructing to generate intermediate solutions This process should refer to the contraindication list. Avoid invalid operations. For intermediate solutions... Applying position operators Select the insertion position to generate a new candidate solution. .
[0014] SB3 Solution Acceptance Criterion: If the objective function value of the new solution is... Better than the current solution, or meets the simulated annealing acceptance criterion (determined by temperature). If the probability of accepting a suboptimal solution is controlled, then a new solution is accepted: the current solution is updated to... Reward levels are assigned based on the degree of improvement in solution quality. Otherwise, reject the new solution and record zero reward.
[0015] SC Adaptive Adjustment: Based on reward, the reinforcement learning controller (RLC) adjusts the input. Real-time updates of the weight vectors of the three types of operators This enables dynamic optimization of the operator selection strategy.
[0016] SD Taboo Management: Reward-Based and intermediate solutions Update the taboo list Operations related to high-reward solutions are added to the taboo list to avoid repeated access in the near future, while some taboo operations are released in low-reward areas.
[0017] SE temperature decay: Perform temperature update By gradually reducing the simulated annealing temperature, the algorithm transitions from a highly exploratory approach (high temperature) in the early stages to a more exploitative approach (low temperature) in the later stages.
[0018] SF Termination Detection: When a preset termination condition is met, exit the loop and output the historical best solution. .
[0019] Furthermore, in the process of scheduling a single agile satellite, the algorithm flow is introduced using 7 tasks as an example. Tasks 1-4 are set as scheduled tasks, and tasks 5-7 are tasks to be scheduled. The algorithm flow is as follows: S1. Destroy task 2 and 4 to free up scheduling space for new task; S2. Repair operations filter tasks to be scheduled based on a priority sorting mechanism; S21. The task priority sequence is: Task 6 > Task 7 > Task 5 > Task 4 > Task 2; S3. Insertion operations attempt to schedule high-priority tasks sequentially; Tasks 6, 7, and 5 were successfully inserted in sequence. Tasks 4 and 2 failed due to a conflict. S4. Generate a new scheduling sequence; The new scheduling sequence is as follows: Task 1, Task 3, Task 6, Task 7, and Task 5 are arranged in sequence.
[0020] Throughout the process, the tabu policy ensures the diversity of search paths by prohibiting immediate actions on recently adjusted tasks. The reinforcement learning-based augmentation controller, based on the temporal difference (TD) algorithm, dynamically adjusts the weight parameters of each operator to achieve continuous optimization and iterative updates of the policy.
[0021] The effects of the invention are: 1. Increased profitability: Experiments show that, with a task scale of 200-600, RL-ALNS has an average profit margin that is 2.31% higher than ALNS / TPF and 36.52% higher than traditional ALNS.
[0022] 2. Dynamic adaptability: The reinforcement learning controller adjusts the operator weights in real time to quickly escape local optima.
[0023] 3. Efficient insertion: POs improves operator diversity through a multi-index hybrid strategy, thereby enhancing the algorithm's adaptability in different scenarios and increasing the task insertion success rate.
[0024] 4. Wide applicability: The algorithm framework can be extended to other combinatorial optimization problems, such as the traveling salesman problem and the vehicle routing problem. Attached Figure Description
[0025] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0026] Figure 1 This is a comparison chart of the reinforcement learning-enhanced adaptive large neighborhood search algorithm (RL-ALNS) and the ALNS framework; Figure 2 Flowchart of the RL-ALNS algorithm, an adaptive large neighborhood search algorithm for reinforcement learning enhancement; Figure 3 This is a diagram illustrating the quick insertion process of FI. Detailed Implementation
[0027] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of this application, including various details to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0028] This invention proposes an adaptive large neighborhood search (ALNS) algorithm enhanced by reinforcement learning, which includes the following core improvements: Reinforcement learning controller: Based on temporal-difference (TD) learning, dynamically adjust the weights of operators to predict future returns and optimize the selection strategy.
[0029] Specifically: the reinforcement learning controller is used to define the state set, action set, and reward function; The state set: This includes: the current solution quality index and the operator weight vector, which describe complete information about the system during the iteration cycle; The set of actions: Includes: selectable neighborhood search operators, generated by policy function decisions; The reward function: Instant rewards The improvement in the quality of new solutions brought about by the selected operator is quantified, and the cumulative reward is calculated. Discount factor Balances short-term and long-term returns. Used to enhance the returns of new solutions.
[0030] The reinforcement learning controller uses the TD(0) update value function to dynamically adjust the operator weights: in For learning rate, For timing difference error, Calculated using historical weights for smoothing.
[0031] The reinforcement learning controller selects operators using roulette wheel selection based on operator weights. The probability of each operator being selected is:
[0032] Operator The probability of being selected and its weight Proportional, of which Indicates the total number of operators. , Representing the first , Each operator at time step The state.
[0033] in
[0034] Hybrid neighborhood operation: Combining destruction and repair operators to dynamically generate candidate solutions.
[0035] The destruction operator includes: Random removal: Randomly removing tasks to break the current solution, explore new search areas, and avoid the algorithm getting stuck in local optima is a common problem in complex scheduling.
[0036] Lowest Profit Removal: Prioritize the removal of tasks with lower profits, allowing resources to be allocated to high-value tasks, which aligns with the goal of maximizing overall profit in AEOSSP.
[0037] Lowest unit profit removal: Remove tasks with the lowest unit profit (profit divided by duration), emphasizing efficiency—tasks with low returns per unit of time have lower priority when resources are limited.
[0038] Longest transfer time removal: Prioritize removing tasks that require long transfer times, as long transfer times will compress the available time of subsequent tasks. Removing them helps improve the overall scheduling feasibility of AEOSSP.
[0039] Historical transfer time removal: Remove tasks with transfer times far exceeding their historical minimums, and prioritize retaining tasks with typically shorter transfer times using historical data to optimize satellite path planning.
[0040] Most Opportunity Window Removal: Prioritize removing tasks with the longest Visible Time Window (VTW), as tasks with fewer time windows are difficult to reschedule. Protecting them ensures that critical tasks with limited flexibility are prioritized.
[0041] The repair operator includes: Highest Profit Restoration: Prioritize inserting the highest-profit tasks and restoring high-value tasks, in line with the goal of maximizing overall profit.
[0042] Highest unit profit repair: Prioritize inserting tasks with the highest unit profit (profit / duration), focusing on efficiency, and prioritizing tasks with the highest return per unit time when resources are limited.
[0043] Shortest transfer time fix: Prioritize inserting tasks with the shortest required transfer time. Shortening the transfer time frees up more time for subsequent tasks, improving the feasibility of AEOSSP scheduling.
[0044] Historical transfer time restoration: Prioritize inserting tasks with transfer times close to the historical minimum, and utilize historical data to prioritize habitual short-term transfer tasks to optimize satellite movement efficiency.
[0045] Least chance window fix: Prioritize inserting tasks with the shortest time window. Tasks with low flexibility should be inserted as early as possible to avoid conflicts, ensuring that critical tasks are completed first.
[0046] Position Operators (POs): Two heuristic rules are designed to provide diverse insertion position choices; these two heuristic rules are as follows: 1. Minimum Insertion Cost (MIC) operator Insertion cost is calculated by iterating through all feasible insertion positions within the visible time window (VTW).
[0047] in For the task of updating after insertion Earliest start time; For task transfer time, select the lowest cost location as the insertion position.
[0048] 2. Minimum Idle Time (MIT) Position Operator Idle time definition: The gap between adjacent tasks = the start time of the subsequent task. -Previous task completion time : Choose the location with the shortest idle time (reduce redundant gaps and improve resource utilization).
[0049] The position operator POs is weighted by coefficients. The optimal insertion position is selected by dynamically balancing the MIC (Minimum Cost) and MIT (Minimum Idle Time) operators using the following formula. : The positional operator Pos supports five weight configurations, enhancing flexibility: →Full Priority Idle Minimization (MIT) →Fully Prioritize Cost Minimization (MIC) → Hybrid Strategy Core design logic: By quantifying insertion cost and idle time, combined with dynamic weight adjustment, adaptive optimization of task insertion position is achieved, maximizing scheduling flexibility and benefits under resource constraints.
[0050] The position operator POs can combine insertion cost and idle time weights to generate a composite index to select the optimal insertion position. By configuring different weights, an "operator library" covering various scenario requirements is constructed, providing rich policy choices for the reinforcement learning augmentation controller, enabling it to achieve efficient decision-making under different environmental conditions.
[0051] The reinforcement learning-enhanced adaptive large neighborhood search algorithm (RL-ALNS) also applies tabu strategies, including tabu removal, tabu insertion, and immediate tabu, to prevent repeated searches.
[0052] like Figure 1 The figure shown is a comparison between the reinforcement learning-enhanced adaptive large neighborhood search algorithm (RL-ALNS) and the ALNS framework.
[0053] Compared to the traditional ALNS algorithm, RL-ALNS introduces POs in addition to the destruction and repair operators used for neighborhood operations, and uses a reinforcement learning-enhanced controller to uniformly update and select all operators.
[0054] Figure 2 The flowchart of the Reinforcement Learning-Enhanced Adaptive Large Neighborhood Search (RL-ALNS) algorithm is shown below. like Figure 2 As shown, the RL-ALNS framework proposed in this embodiment includes four core components: tabu policy, neighborhood operation, insertion operation, and reinforcement learning boosting controller.
[0055] Among them, the neighborhood operation focuses on exploring the solution space by replacing the current task, while the insertion operation aims to embed the new task reasonably into the scheduling sequence. In the insertion operation, the composite operator of POs integrates multiple heuristic strategies to help locate the optimal insertion position.
[0056] The weight parameters of all operators are uniformly managed and dynamically selected by the reinforcement learning augmentation controller. The operators include destruction operators, repair operators, and POs.
[0057] As a key control mechanism, the tabu policy effectively avoids algorithm loops and reduces invalid searches by prohibiting repeated operations on recently modified tasks.
[0058] Taking a scheduling instance containing 7 tasks as an example, Among them, tasks 1-4 are already assigned tasks, and tasks 5-7 are tasks to be scheduled. The algorithm flow is as follows: S1. Destroy task 2 and 4 to free up scheduling space for new task; S2. Repair operations filter tasks to be scheduled based on a priority sorting mechanism; S21. The task priority sequence is Task 6 > Task 7 > Task 5 > Task 4 > Task 2; S3. Insertion operations attempt to schedule high-priority tasks sequentially; Tasks 6, 7, and 5 were successfully inserted in sequence. Tasks 4 and 2 failed due to a conflict. S4. Generate a new scheduling sequence; The new scheduling sequence is as follows: Task 1, Task 3, Task 6, Task 7, and Task 5 are arranged in sequence.
[0059] Throughout the process, the tabu policy ensures the diversity of search paths by prohibiting immediate actions on recently adjusted tasks. The reinforcement learning-based augmentation controller, based on the temporal difference (TD) algorithm, dynamically adjusts the weight parameters of each operator to achieve continuous optimization and iterative updates of the policy.
[0060] Figure 3 To illustrate the rapid insertion process of FI, specifically, task i is inserted at position j+1. Adjacent tasks are slid to the left or right to create idle time for task i. If there is sufficient idle time, the task insertion is successful, and the transition times of all tasks in the sequence are updated.
[0061] If scheduling is performed for a single agile satellite, the following steps are executed: SA. Initialization: The ground station acquires the target satellite identifier and generates a scheduling file based on the target satellite. The file to be scheduled includes: a set of candidate tasks. The ground station sets satellite orbit parameters, mission priorities, and time windows; SB. Neighborhood operations: Remove tasks using a random destruction operator, and sort tasks to be inserted by priority using a repair operator; SC. Insertion operation: Select the task insertion position using POs, calculate idle time and insertion cost, and generate a new solution; SD. Weight Update: The reinforcement learning controller updates the operator weights based on the reward of the new solution; SE. Taboo Management: Re-insertion or removal of recently modified tasks and updates to the taboo list are prohibited.
[0062] The effects of the invention are: 1. Increased profitability: Experiments show that, with a task scale of 200-600, RL-ALNS has an average profit margin that is 2.31% higher than ALNS / TPF and 36.52% higher than traditional ALNS.
[0063] 2. Dynamic adaptability: The reinforcement learning controller adjusts the operator weights in real time to quickly escape local optima.
[0064] 3. Efficient insertion: POs improves operator diversity through a multi-index hybrid strategy, thereby enhancing the algorithm's adaptability in different scenarios and increasing the task insertion success rate.
[0065] 4. Wide applicability: The algorithm framework can be extended to other combinatorial optimization problems (such as the traveling salesman problem, vehicle routing problem, etc.).
[0066] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A reinforcement learning-enhanced adaptive large neighborhood search method, characterized in that, To schedule a single agile satellite, the following steps are performed: SA. Initialization: The ground station server constructs an initial feasible solution, denoted as the solution vector. And create a taboo list Used to record prohibited operations to avoid duplicate searches; Set temperature decay function Control the temperature drop curve of the simulated annealing process to determine the maximum number of iterations. As one of the termination conditions; Initialize the reward set ,in Indicates the intensity of the fourth-level reward; To destroy operator sets Repair operator set and position operator set Assign initial weight vectors respectively ; SB. Iterative optimization loop: When the termination condition is not met, the following steps are executed repeatedly: SB1. Dynamic Operator Selection: Based on a roulette wheel selection mechanism, the destructive operator for the current iteration is selected from three sets of operators. Repair operator and position operator Its selection probability is determined by the weight vector. Decide; SB2. Deconstruction process: For the current solution Execute the destruction operator Remove some elements; By repairing the operator Reconstructing to generate intermediate solutions ; for intermediate solutions Applying position operators Select the insertion position to generate a new candidate solution. ; SB3. Solution Acceptance Criterion: If the objective function value of the new solution is... If the solution is better than the current solution, or satisfies the simulated annealing acceptance criterion, then accept the new solution: update the current solution to... Reward levels are assigned based on the degree of improvement in solution quality. Otherwise, reject the new solution and record zero reward. SC. Adaptive Adjustment: The reinforcement learning controller (RLC) adjusts the response based on reward. Real-time updates of the weight vectors of the three types of operators This enables dynamic optimization of the operator selection strategy; SD. Taboo Management: Based on Rewards and intermediate solutions Update the taboo list Operations related to high-reward solutions are added to the taboo list to prevent recent repeated visits, while some taboo operations are released in low-reward areas. SE. Temperature decay: Perform temperature update. Gradually reduce the simulated annealing temperature to allow the algorithm to transition from a highly exploratory approach in the early stages to a more exploitative approach in the later stages. SF. Termination Detection: When a preset termination condition is met, exit the loop and output the historical best solution. .
2. The adaptive large neighborhood search method with reinforcement learning enhancement according to claim 1, characterized in that, In the scheduling process for a single agile satellite, seven tasks are set up. Set tasks 1-4 as scheduled tasks and tasks 5-7 as tasks to be scheduled. The algorithm flow is as follows: S1. Destroy task 2 and 4 to free up scheduling space for new task; S2. Repair operations filter tasks to be scheduled based on a priority sorting mechanism; S21. The task priority sequence is Task 6 > Task 7 > Task 5 > Task 4 > Task 2; S3. Insertion operations attempt to schedule high-priority tasks sequentially; Tasks 6, 7, and 5 were successfully inserted in sequence. Tasks 4 and 2 failed due to a conflict. S4. Generate a new scheduling sequence; The new scheduling sequence is as follows: Task 1, Task 3, Task 6, Task 7, and Task 5 are arranged in that order. Throughout the process, the tabu policy ensures the diversity of search paths by prohibiting immediate actions on recently adjusted tasks, while the reinforcement learning boosting controller dynamically adjusts the weight parameters of each operator based on the temporal difference TD algorithm to achieve continuous optimization and iterative updates of the policy.