Automatic algorithm design method for solving vehicle path planning problem by using large language model

Through the AutoDH framework and two-level MDP intelligent selection heuristic method, the subjectivity and limitations of manual design algorithms in the existing technology are solved, and the cost-effectiveness of LLMs design is used to achieve more flexible, adaptable and efficient algorithm design.

CN120106324APending Publication Date: 2025-06-06NORTHWEST UNIV
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510046296.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art has subjectivity and limitations when designing algorithms manually, and when using large language models (LLMs) for automatic algorithm design, no research has been conducted on the most cost-effective heuristics of using deep reinforcement learning (DRL) to intelligently select the design of LLMs.

Method used

By building the AutoDH framework, defining the decision-making form of agents, obtaining status information of vehicle path planning problems (CVRPs), using two-level MDP intelligent selection heuristics, designing reward mechanisms that integrate the quality improvement of solutions, heuristic time costs and calling LLMs API costs, and training agents to optimize the solutions of CVRPs.

Benefits of technology

It improves the flexibility and adaptability of algorithm design, realizes the intelligent selection of the most cost-effective heuristic algorithm, enhances the universality and pertinence of algorithms, and improves the efficiency and effectiveness of algorithm design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106324A_ABST
    Figure CN120106324A_ABST
Patent Text Reader

Abstract

The invention relates to the field of automatic algorithm design, in particular to an automatic algorithm design method for solving a vehicle path planning problem by utilizing a large language model, which comprises the following steps of: S1, constructing an AutoDH framework, and defining a decision form of an intelligent agent; s2, defining heuristics in an LLM pool, an improvement pool and a disturbance pool; s3, acquiring state information of the CVRPs, and intelligently selecting a heuristic mode through two-stage MDP; s4, designing a reward mechanism fusing solution quality improvement, heuristic time cost and LLMs API calling cost; and S5, training the intelligent agent to optimize the solution of the CVRPs. The method can intelligently select the most cost-effective heuristic algorithm according to the state information of the current CVRPs, and improves the flexibility and adaptability of algorithm design.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of automatic algorithm design, and in particular to an automatic algorithm design method for solving vehicle path planning problems using a large language model. Background Art

[0002] The field of automatic algorithm design has developed rapidly in recent years, especially in solving combinatorial optimization problems (COPs). Combinatorial optimization problems involve finding optimal solutions in discrete sets and are widely used in many fields such as logistics, bioinformatics, and energy management. However, manually designing algorithms not only requires a lot of domain-specific knowledge, but is also very time-consuming.

[0003] Since there are a lot of algorithms for solving COPs, and manually designing these algorithms requires a lot of expertise and time, the reliance on the designer's experience and knowledge introduces subjectivity and limitations in algorithm design. In addition, with the development of artificial intelligence, neural networks have been widely used in combinatorial optimization, and automatic algorithm design has become a rapidly developing and highly promising research direction. In particular, the increasing influence of large language models (LLMs) in practical applications has made automatic algorithm design using LLMs an interesting research area. However, building solutions using LLMs faces challenges related to learning and inference. In terms of algorithm design, no research has yet used deep reinforcement learning (DRL) to intelligently select the most cost-effective heuristics for LLMs design. Summary of the invention

[0004] The present invention aims to provide an automatic algorithm design method for solving vehicle path planning problems using a large language model, which can intelligently select the most cost-effective heuristic algorithm based on the current state information of CVRPs, thereby improving the flexibility and adaptability of algorithm design.

[0005] The present invention is achieved through the following technical solutions:

[0006] An automatic algorithm design method for solving a vehicle path planning problem using a large language model comprises the following steps:

[0007] S1: Build the AutoDH framework and define the decision-making form of the intelligent agent;

[0008] S2: Define the heuristics in LLM pool, improvement pool and perturbation pool;

[0009] S3: Obtain the status information of CVRPs and intelligently select heuristics through two-level MDP;

[0010] S4: Design a reward mechanism to integrate solution quality improvement, heuristic time cost, and LLMs API call cost;

[0011] S5: Train the agent to optimize the solution of CVRPs.

[0012] Preferably, the decision form of the agent in S1 is defined as a multi-dimensional decision space, the first dimension stores the heuristic type to be executed, denoted as Heur-type, and the second dimension stores the parameter configuration of the corresponding heuristic, denoted as Heur-config.

[0013] Preferably, the LLM pool in S2 contains three LLM heuristics, namely subpath reconstruction LLM, intelligent design LLM and intelligent perturbation LLM, the improvement pool contains improved heuristics composed of 27 local search heuristic methods designed by experts, and the perturbation pool contains perturbation heuristics that can significantly perturb the current solution.

[0014] Furthermore, the sub-path reconstruction LLM focuses on the optimization at the solution level, and reconstructs the sub-path with the largest distance in the current solution using information such as the distance matrix, node requirements, and vehicle capacity;

[0015] The intelligent design LLM is used to create new heuristics;

[0016] The smart perturbation LLM adjusts parameters by analyzing the code of previously designed heuristics and their rewards to generate more cost-effective heuristics.

[0017] Preferably, the two-level MDP in S3 includes a high-level MDP (M1) and a low-level MDP (M2), so that the agent intelligently selects and generates heuristics at the global level (M1) and the local level (M2), forming a closed-loop learning system, wherein:

[0018] High-level MDP (M1): responsible for the agent to select the best heuristic method from the LLM pool, improvement pool and perturbation pool;

[0019] Low-level MDP (M2): LLMs are responsible for generating new heuristics based on hints, contextual information, and the performance scores of previously designed heuristics.

[0020] Preferably, the reward mechanism in S4 integrates three key elements: quality improvement of solution, running time of heuristics, and cost of calling LLMs API;

[0021] The improvement in solution quality is calculated by calculating the distance difference ΔQ = L between the solution obtained after the current heuristic application and the previous solution. t-1 -L t To measure, where L t-1 is the solution distance before applying the heuristic, L t is the solution distance after application;

[0022] The running time (T) of the heuristic algorithm is based on the time of the whole process from receiving the hint to generating the heuristic code and the execution time of the traditional heuristic, and the more efficient algorithm is selected;

[0023] The cost (C) of calling the LLMs API includes the number of input and output tokens and their prices, calculated as C = I t I p +O t ·O p , where I t is the number of input tokens, I p is the price of the input token, O t is the number of output tokens, O p is the price of the output token to reflect the economic cost of using LLMs;

[0024] These three factors are combined together through weighted coefficients α, β, and γ to form a comprehensive reward function R, that is, Used to balance the relationship between solution quality improvement, computing efficiency and economic cost.

[0025] Furthermore, the S4 also includes a feedback mechanism, that is, the agent updates its policy network according to the received rewards, optimizes the network parameters, and evaluates the performance of AutoDH in terms of solution quality, running time and cost-effectiveness by comparing the solutions generated by AutoDH with the solutions of other baseline methods, and feeds back the performance scores to LLMs to optimize subsequent heuristic choices.

[0026] Preferably, the method of training the agent in S5 is specifically to repeat S3 to S4 until the strategy network of the agent converges, or reaches a preset number of iterations or time limit.

[0027] The present invention has the following beneficial effects:

[0028] (1) This application provides a new automatic algorithm design framework AutoDH by combining a large language model (LLM) and traditional heuristic methods. The framework allows the agent to intelligently select the most cost-effective heuristic algorithm from the LLM pool, improvement pool, and perturbation pool according to the current state, instead of using a single pre-selected heuristic for all problem instances and the entire optimization process, thereby enhancing the universality and pertinence of the algorithm and improving the flexibility and adaptability of algorithm design.

[0029] (2) The LLM pool in this application can automatically design algorithms to enhance the current solution based on given hints. This real-time optimization capability enables AutoDH to perform real-time optimization at the solution and heuristic levels, improving the efficiency and effectiveness of the algorithm;

[0030] (3) This application defines the automatic algorithm design problem as a reinforcement learning task and models it as a two-level MDP. This design establishes a closed-loop learning system with a feedback mechanism, which improves the quality and generalization of the heuristic design.

[0031] (4) This application proposes a new reward mechanism that comprehensively considers the improvement of solution quality, the time cost of the heuristic, and the cost of calling the LLM API, and then trains the intelligent agent to optimize the solution of CVRPs, making the effectiveness of the selected heuristic more comprehensive and objective. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 A diagram showing the design steps of AutoDH according to the present invention;

[0033] Figure 2 This is the overall architecture diagram of AutoDH of the present invention;

[0034] Figure 3 is an example of the AutoDH framework of the present invention;

[0035] Figure 4 It is a frequency diagram of heuristic usage in the heuristic pool of the present invention. DETAILED DESCRIPTION

[0036] The present invention is further described in detail below in conjunction with specific embodiments, which are intended to explain the present invention rather than to limit it.

[0037] refer to Figure 1 and Figure 2 As shown, by intelligently selecting heuristic algorithms and integrating the advantages of LLM, the present application provides a new solution for the automatic algorithm design of COPs, especially in the capacity vehicle routing problem (CVRPs), which improves the efficiency and effect of automatic algorithm design. Specifically, an automatic algorithm design method for solving vehicle routing problems using a large language model is provided, including the following steps:

[0038] S1: Define the problem and context

[0039] First, the problem instance of CVRPs is defined, including parameters such as node location, customer demand, and vehicle capacity. These parameters constitute the state space of CVRPs.

[0040] S2: Initialize the AutoDH framework

[0041] Initialize the AutoDH framework, including the agent’s policy network, LLM pool, improvement pool, and perturbation pool. The policy network consists of a convolutional neural network, an attention layer, and a multi-layer perceptron (MLP) to output the probability of selecting each heuristic.

[0042] LLM pool: includes three LLMs, each of which targets a specific task and is responsible for subpath reconstruction, intelligent design and intelligent perturbation respectively. LLM heuristics: consists of three LLMs, including subpath reconstruction LLM, intelligent design LLM and intelligent perturbation LLM, which can automatically design algorithms to enhance the current solution based on given hints.

[0043] Among them, subpath reconstruction LLM: focuses on solution-level optimization, and uses information such as distance matrix, node requirements, and vehicle capacity to reconstruct the subpath with the largest distance in the current solution; intelligent design LLM: aims to create new heuristics; intelligent perturbation LLM: adjusts parameters by analyzing the codes and rewards of previously designed heuristics to generate more cost-effective heuristics. Their codes are h28 to h30 respectively.

[0044] Improvement Pool: The Improvement Pool consists of 27 expert-designed local search heuristics that are used to fine-tune and improve the current solution. Improvement Heuristics: It consists of local search heuristics that are used to fine-tune and improve the current solution, including 2Opt (h1), swap (h2), symmetric swap (h5 to h7, h24), and relocation (h3, h8 to h10, h25 to h27). These heuristics focus on making subtle adjustments to the current solution to improve the quality of the solution.

[0045] Perturbation pool: contains heuristics that can significantly perturb the current solution to explore new solution spaces. Perturbation heuristics: contains three large neighborhood search heuristics designed by experts that can significantly perturb the current solution to explore new solution spaces and avoid falling into local optimality. These heuristics include random-permute, random-exchange, and cyclic-exchange. Their codes are h31 to h33 respectively.

[0046] S3: Agent-Environment Interaction

[0047] The agent starts to interact with the environment and first obtains the current state information of CVRPs, including static state and dynamic state. Static state information such as node coordinates and vehicle capacity, and dynamic state information such as the current solution and historical usage of heuristics.

[0048] S4: Selection Heuristics

[0049] The agent uses the policy network to select a heuristic based on the current state information, refer to Figure 4, this selection is based on the ε-greedy strategy, which selects the current optimal heuristic with a certain probability, while retaining a certain probability to explore other heuristics. The specific decision-making process includes the following:

[0050] High-level MDP (M1): Responsible for the agent to select the best heuristic from the LLM pool, improvement pool, and perturbation pool. This process involves analyzing static and dynamic states, including problem-specific characteristics and historical usage of the current solution, and then using an ε-greedy strategy to select the heuristic that is most likely to improve the quality of the solution based on the probability output by the policy network. The reward for the selected action is calculated based on a weighted sum of the improvement in solution quality, the runtime cost of the heuristic, and the cost of calling the LLMs API.

[0051] Low-level MDP (M2): responsible for LLMs to generate new heuristics based on prompts, contextual information, and performance scores of previously designed heuristics. The state space of LLMs consists of natural language prompts and dialogue context information used to guide heuristic generation, which helps LLMs better understand the task and generate more accurate heuristics. The actions of M2 are the heuristics generated by LLMs, while the reward mechanism is the same as M1, which aims to encourage LLMs to design more effective heuristics and feed back the performance scores to LLMs through a feedback mechanism to optimize subsequent heuristic designs.

[0052] The two-level MDP design of the AutoDH framework allows the agent to intelligently select and generate heuristics at both the global level (M1) and the local level (M2), forming a closed-loop learning system. This two-level structure not only maintains the quality of the solution, but also effectively balances the use of computing resources and adaptively adjusts the heuristic method to suit different instances and requirements of CVRPs. In this way, AutoDH is able to achieve efficient optimization of CVRPs while demonstrating cost-effectiveness compared to expert-designed heuristics.

[0053] S5: Apply heuristics and receive feedback

[0054] The selected heuristics are applied to the current solution of CVRPs and the agent receives feedback based on the application results. This feedback is provided through a reward mechanism that takes into account three key factors: the improvement in the quality of the understanding, the running time of the heuristics, and the cost of calling the LLMs API.

[0055] Specifically, the quality improvement of the solution is achieved by calculating the distance difference (ΔQ = L t-1 -L t ) is used to measure, where L t-1 is the solution distance before applying the heuristic, Lt is the solution distance after application; the running time of the heuristic algorithm (T) takes into account the time of the entire process from receiving the hint to generating the heuristic code, as well as the execution time of the traditional heuristic to encourage the selection of more efficient algorithms; the cost of calling the LLMs API (C) includes the number of input and output tokens and their prices, calculated as C = I t I p +O t ·O p , where I t is the number of input tokens, I p is the price of the input token, O t is the number of output tokens, O p is the price of the output token to reflect the economic cost of using LLMs; these factors are combined together through weighting coefficients α, β, and γ to form a comprehensive reward function R, that is, In this way, the relationship between solution quality improvement, computational efficiency and economic cost is balanced. This multi-factor fusion reward mechanism is the core of the AutoDH framework's intelligent selection heuristic method. It not only pursues the improvement of solution quality, but also takes into account the economy and efficiency of computing resources, which helps to achieve more efficient and economical algorithm design in practical applications.

[0056] S6: Update policy network

[0057] The agent updates its policy network based on the rewards it receives, optimizing the network parameters using a gradient ascent method so that future decisions are more likely to choose heuristics that lead to higher rewards.

[0058] S7: Validation and Evaluation

[0059] The effectiveness of AutoDH is verified on CVRP instances of different sizes, which includes generating CVRP instances of different problem sizes (e.g., N = 20, 50, 100) and evaluating the performance of the agent. The performance of AutoDH is evaluated in terms of solution quality, running time, and cost-effectiveness by comparing the solutions generated by AutoDH with those of other baseline methods.

[0060] S8: Iterative Optimization

[0061] Repeat S3 to S7 to train the agent to optimize the solution of CVRPs until the agent's policy network converges or a preset number of iterations or time limit is reached.

[0062] During the training process, the setting of model parameters and training parameters is crucial, including optimizer, learning rate, discount factor, exploration rate (ε), etc. For example, the parameters of AutoSAF in Table 1 below, the setting of these parameters ensures that the agent can effectively learn and converge to a good strategy. Through such a training process, the agent can learn how to intelligently select the most appropriate heuristic method from the LLM pool, improvement pool and perturbation pool to optimize the solution when facing different CVRP instances. This training not only improves the quality of the solution, but also improves the efficiency of computing resources, making AutoDH an effective automatic algorithm design framework for solving CVRPs. The experimental results are shown in Table 2.

[0063] Table 1 Parameters of AutoSAF

[0064]

[0065]

[0066] Table 2 Experimental results of AutoSAF for random CVRP

[0067]

[0068] Specific operation process: Figure 3 In the CVRP instance shown, there are 20 customers and 2 vehicles, each with a capacity of 10 units. Customers' demands are randomly distributed between 1 and 9 units. The agent first obtains the state information of this instance, including problem-specific features and historical usage of the current solution, and then selects a heuristic such as 2-opt through the policy network to optimize the current solution. After applying 2-opt, the quality of the solution is improved. The agent calculates the reward based on the improvement, the running time of the heuristic, and the LLMs API call cost, and then updates its policy network based on this reward, and then continues to interact with the environment, repeating this process until the optimal or near-optimal solution is found.

[0069] Example 1

[0070] Task: Optimize a given sub-solution in the Capacity Constrained Vehicle Routing Problem (CVRP) to minimize the travel distance. Please output the optimized sub-solution directly. The optimization process must ensure that each customer's demand is met without exceeding the vehicle's capacity, and no node in the sub-solution is visited multiple times or left unvisited. The input includes a distance matrix representing the distances between locations in the sub-solution, a list representing customer demands, an integer representing the vehicle's capacity, and the current node sequence. An optimized node sequence should be returned that minimizes the travel distance as much as possible.

[0071] Input parameters:

[0072] Distance Matrix: A 2D list representing the distances between locations in a sub-solution. Distance Matrix = {Distance Matrix} · Demand: A list representing the demand of each customer;

[0073] Demand = {demand}·vehiclecapacity: an integer representing the total vehicle capacity;

[0074] VehicleCapacity = {VehicleCapacity} · SubSolution: A list representing the current node sequence in the subsolution;

[0075] subsolution = {subsolution} · subsolution total distance: a floating point number representing the total travel distance of the subsolution before optimization;

[0076] Sub-solution total distance = {sub-solution total distance}.

[0077] Output:

[0078] Optimized sub-solutions: a list representing the optimized node sequence;

[0079] difference: A floating point number representing the total distance difference between the perturbed solution and the original solution.

[0080] Example 2

[0081] Task: Design a heuristic function to optimize the Capacitated Vehicle Routing Problem (CVRP). The input includes the distance matrix between all locations, customer demands, vehicle capacities, and the current solution. The function should minimize the total distance traveled while ensuring that each customer's demand is met and the vehicle capacity is not exceeded.

[0082] Function name: Intelligent design Input parameters:

[0083] Problem: Class (the class containing the distance matrix and customer requirements has been designed and does not need to be redefined);

[0084] Solution: List (a list of routes, each route is a list of customer indexes, representing the order in which customers are visited).

[0085] Output:

[0086] Perturbation solution: List (an optimized route list, each route is a customer index list);

[0087] Difference: Float (difference in total distance between the perturbed solution and the original solution);

[0088] In function design, if you need to calculate the distance between nodes, the need to visit nodes, or the capacity of vehicles, use the following functions:

[0089] Calculate the distance between indicators (question, indicator 1, indicator 2): used to calculate the distance between two nodes;

[0090] problem.capacity[index]: The demand to access the nodes in the problem, where index represents the index of the node and problem.capacity[0] represents the total capacity of the vehicles.

[0091] Example 3

[0092] Task: Please redesign a heuristic function to optimize the Capacity Constrained Vehicle Routing Problem (CVRP) based on the heuristic usage of the past t steps. The input includes the distance matrix between all locations, customer demands, vehicle capacities, and the current solution. The function aims to minimize the total travel distance while ensuring that each customer's demand is satisfied without exceeding the vehicle capacity.

[0093] Based on past heuristic usage and their rewards, redesign a heuristic to improve the current solution. The heuristics designed by the previous LLMs are encoded as:

[0094] Code 1:............................................

[0095] Reward 1: ..................................................

[0096] Code: ................................................

[0097] Reward: ................................

[0098] Function name: Intelligent modification of input parameters:

[0099] Problem: Classes containing attributes such as distance matrix, customer demand, vehicle capacity, etc.

[0100] Solutions: A list representing the current solution (a set of paths), where each path is a list of customer indices representing the order in which the customers visited them.

[0101] Output:

[0102] LLMs solution: optimize the list of paths, where each path is a list of customer indicators;

[0103] difference: a floating point number indicating the total distance difference between the perturbed solution and the original solution;

[0104] When designing functions, use the following functions to calculate distances between nodes, node visit requirements, or vehicle capacity:

[0105] CalculateDistanceBetweenIndicators(problem, indicator1, indicator2): Calculates the distance between two nodes;

[0106] problem.capacity[index]: the demand to access the nodes in the problem, where 'index' represents the node index;

[0107] problem.capacity[0]: represents the total capacity of the vehicle.

[0108] The above contents are further detailed descriptions of the present invention in combination with specific embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these inventions. For ordinary technicians in the technical field of the present invention, without departing from the concept of the present invention, they can also make several simple deductions or substitutions, which should be regarded as belonging to the protection scope of the present invention.

Claims

1. An automatic algorithm design method for solving vehicle path planning problems using a large language model, characterized in that: The following steps are involved: S1: Build the AutoDH framework and define the decision-making form of the intelligent agent; S2: Define the heuristics in LLM pool, improvement pool and perturbation pool; S3: Obtain the status information of CVRPs and intelligently select heuristics through two-level MDP; S4: Design a reward mechanism to integrate solution quality improvement, heuristic time cost, and LLMs API call cost; S5: Train the agent to optimize the solution of CVRPs.

2. The automatic algorithm design method for solving vehicle path planning problems using a large language model according to claim 1, characterized in that: The decision form of the agent in S1 is defined as a multi-dimensional decision space, where the first dimension stores the heuristic type to be executed, denoted as Heur-type, and the second dimension stores the parameter configuration of the corresponding heuristic, denoted as Heur-config.

3. The automatic algorithm design method for solving vehicle path planning problems using a large language model according to claim 1, characterized in that: The LLM pool in S2 contains three LLM heuristics, namely subpath reconstruction LLM, intelligent design LLM and intelligent perturbation LLM. The improvement pool contains improved heuristics composed of 27 local search heuristic methods designed by experts. The perturbation pool contains perturbation heuristics that can significantly perturb the current solution.

4. The automatic algorithm design method for solving vehicle path planning problems using a large language model according to claim 3 is characterized in that: The sub-path reconstruction LLM focuses on solution-level optimization and reconstructs the sub-path with the largest distance in the current solution using information such as distance matrix, node requirements, and vehicle capacity; The intelligent design LLM is used to create new heuristics; The smart perturbation LLM adjusts parameters by analyzing the code of previously designed heuristics and their rewards to generate more cost-effective heuristics.

5. The automatic algorithm design method for solving vehicle path planning problems using a large language model according to claim 1, characterized in that: The two-level MDP in S3 includes a high-level MDP (M1) and a low-level MDP (M2), which enables the agent to intelligently select and generate heuristics at the global level (M1) and the local level (M2), forming a closed-loop learning system, in which: High-level MDP (M1): responsible for the agent to select the best heuristic method from the LLM pool, improvement pool and perturbation pool; Low-level MDP (M2): LLMs are responsible for generating new heuristics based on hints, contextual information, and the performance scores of previously designed heuristics.

6. The automatic algorithm design method for solving vehicle path planning problems using a large language model according to claim 1, characterized in that: The reward mechanism in S4 integrates three key elements: the improvement of understanding quality, the running time of the heuristic, and the cost of calling the LLMs API; The improvement in solution quality is calculated by calculating the distance difference ΔQ = L between the solution obtained after the current heuristic application and the previous solution. t-1 -L t To measure, where L t-1 is the solution distance before applying the heuristic, L t is the solution distance after application; The running time (T) of the heuristic algorithm is based on the time of the whole process from receiving the hint to generating the heuristic code and the execution time of the traditional heuristic, and the more efficient algorithm is selected; The cost (C) of calling the LLMs API includes the number of input and output tokens and their prices, calculated as C = I t I p +O t ·O p , where I t is the number of input tokens, I p is the price of the input token, O t is the number of output tokens, O p is the price of the output token to reflect the economic cost of using LLMs; These three factors are combined together through weighted coefficients α, β, and γ to form a comprehensive reward function R, that is, Used to balance the relationship between solution quality improvement, computing efficiency and economic cost.

7. The automatic algorithm design method for solving vehicle path planning problems using a large language model according to claim 6, characterized in that: The S4 also includes a feedback mechanism, namely, the agent updates its policy network according to the received rewards, optimizes the network parameters, and evaluates the performance of AutoDH in terms of solution quality, running time, and cost-effectiveness by comparing the solutions generated by AutoDH with those of other baseline methods, and feeds the performance scores back to LLMs to optimize subsequent heuristic choices.

8. The automatic algorithm design method for solving vehicle path planning problems using a large language model according to claim 1, characterized in that: The method of training the agent in S5 is specifically to repeat S3 to S4 until the strategy network of the agent converges or reaches a preset number of iterations or time limit.

Citation Information

Cited By

  • Power grid metering material verification and distribution collaborative optimization method and device and storage medium

    CN121052760A

  • Power grid metering material verification and distribution collaborative optimization method and device and storage medium

    CN121052760B

  • Logistics distribution path acquisition method, apparatus and device, and storage medium

    CN122022659A

  • A logistics distribution path acquisition method, device, equipment and storage medium

    CN122022659B

  • Capacity constraint vehicle path modeling and solving method based on large language model

    CN122288543A