CTSP solving method and system based on convolution enhanced proxy attention mechanism

By constructing a neural policy network and reinforcement learning algorithm based on convolution enhancement agent attention mechanism, combined with traditional optimizers, the problem of inefficiency of existing TSP solvers is solved, and the effect of efficient generation of high-quality path solutions is achieved.

CN120354086AActive Publication Date: 2025-07-22SHAOXING UNIVERSITY

Patent Information

Application Number
CN202510833906.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-07-22
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

The existing mainstream neural TSP solvers are less efficient when processing large-scale graph examples, which is difficult to meet the needs of efficiency and scalability in practical applications. They lack effective coordination between neural networks and traditional optimization paradigms, making it difficult to achieve a good balance between solution efficiency and the quality of solutions.

Method used

A neural strategy network based on convolution enhancement agent attention mechanism is built, combined with reinforcement learning algorithm training and generate a preliminary optimal path scheme, and then finely optimized through traditional optimizers to achieve both the efficient and quality of the path generation process.

Benefits of technology

It significantly improves the inference efficiency and scalability of the model in large-scale graph examples, generates high-quality path solutions, and achieves a good balance between computing resources and the quality of solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354086A_ABST
    Figure CN120354086A_ABST
Patent Text Reader

Abstract

The invention provides a CTSP solving method and system based on a convolution enhanced agency attention mechanism, and relates to the technical field of data processing, and the method comprises the steps: obtaining city data and salesperson data; based on the city data and the salesman data, constructing a complete graph used for describing CTSP problem instances; based on the complete graph, constructing a constrained Markov decision process model, and modeling a path generation process as a step-by-step decision mechanism; a neural strategy network based on convolution enhancement and an agency attention mechanism is constructed, and the neural strategy network is used for solving a constrained Markov decision process model; training the neural strategy network based on the convolution enhancement and agency attention mechanism by using a reinforcement learning algorithm; generating a preliminary optimal path scheme through the trained neural strategy network based on convolution enhancement and an agency attention mechanism; and inputting the optimal path scheme into a traditional optimizer for fine optimization to obtain a final optimal path scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a CTSP solving method and system based on a convolutional enhanced proxy attention mechanism. Background Art

[0002] The Travelling Salesman Problem (TSP), as a representative difficult problem in combinatorial optimization problems, has long played an important role in practical scenarios such as transportation scheduling, logistics path optimization, and power system layout. Its extended form, the Multi-Travelling Salesman Problem (MTSP), further introduces a collaborative path planning mechanism among multiple travelling salesmen to meet the requirements of complex task allocation and multi-agent systems. With the development of intelligent manufacturing, urban intelligent transportation, and intelligent logistics systems, the requirement for heterogeneous path planning capabilities is increasing day by day, giving rise to more complex model forms and algorithm mechanisms.

[0003] In recent years, Neural Combinatorial Optimization, as an emerging strategy, has gradually become an effective tool for solving such problems. Represented by pointer networks and attention mechanism-based models (such as the Attention Model, AM), the autoregressive (AR) paradigm has shown strong modeling capabilities in tasks such as TSP and VRP (Vehicle Routing Problem). Further research such as the POMO strategy and non-autoregressive (NAR) heatmap methods has achieved remarkable results in solution space expression, diversity enhancement, and inference efficiency.

[0004] However, existing mainstream neural TSP solvers are inefficient in processing large-scale graph instances and are difficult to meet the requirements of high efficiency and scalability in practical applications. At the same time, existing methods usually rely on a single neural network or traditional optimization paradigm, lacking effective coordination between the two, and it is difficult to achieve a good balance between solution efficiency and solution quality. Summary of the Invention

[0005] In view of the above deficiencies of the prior art, the purpose of the embodiments of the present invention is to provide a CTSP solving method based on a convolutional enhanced proxy attention mechanism, which can solve the technical problems that existing mainstream neural TSP solvers are inefficient in processing large-scale graph instances and are difficult to meet the requirements of high efficiency and scalability in practical applications. At the same time, existing methods usually rely on a single neural network or traditional optimization paradigm, lacking effective coordination between the two, and it is difficult to achieve a good balance between solution efficiency and solution quality.

[0006] In the first aspect of the embodiments of the present invention, a CTSP solving method based on a convolutional enhanced proxy attention mechanism is proposed, including: S1: Obtain city data and salesman data; S2: Based on the city data and the salesman data, construct a complete graph for describing the CTSP problem instance; S3: Based on the complete graph, construct a constrained Markov decision process model. By defining the state space, action space, transition dynamics, and reward function of the constrained Markov decision process model, model the path generation process of the CTSP problem as a step-by-step decision-making mechanism; S4: Construct a neural policy network based on convolutional enhancement and proxy attention mechanism, where the neural policy network is used to solve the constrained Markov decision process model; S5: Use a reinforcement learning algorithm to train the neural policy network based on convolutional enhancement and proxy attention mechanism; S6: Generate a preliminary optimal path plan through the trained neural policy network based on convolutional enhancement and proxy attention mechanism; S7: Input the optimal path plan into a traditional optimizer for fine optimization to obtain the final optimal path plan.

[0007] In the second aspect of the embodiments of the present invention, a CTSP solving system based on a convolutional enhanced proxy attention mechanism is proposed, including: a processor and a memory; The memory stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor, it implements the steps of the CTSP solving method based on the convolutional enhanced proxy attention mechanism as described in the first aspect.

[0008] In the third aspect of the embodiments of the present invention, a readable storage medium is proposed. A program or instruction is stored on the readable storage medium. When the program or instruction is executed by a processor, it implements the steps of the CTSP solving method based on the convolutional enhanced proxy attention mechanism as described in the first aspect.

[0009] The beneficial effects brought by the technical solutions provided by the embodiments of the present invention at least include: In the embodiments of the present invention, by constructing a neural policy network based on convolutional enhancement and proxy attention mechanism, combining the attention structure and the local convolutional feature extraction module, the inference efficiency and scalability of the model in large-scale graph instances are significantly improved. At the same time, a reinforcement learning algorithm is introduced to train the network, enabling the model to have an efficient policy exploration ability and quickly generate preliminary high-quality solutions. On this basis, a traditional optimizer is further introduced for fine optimization, realizing the effective integration of the fast modeling advantage of the neural network and the accurate solution ability of the traditional algorithm, taking into account both the quality of the solution and the computational efficiency, and achieving a good balance between the solution accuracy and the computational resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The drawings are only for the purpose of showing specific embodiments and are not considered as limitations of the present invention. Throughout the drawings, the same reference numerals represent the same components. Obviously, the drawings in the following description are only some embodiments described in the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0011] Figure 1 is a schematic flowchart of a CTSP solving method based on a convolutional enhancement proxy attention mechanism provided by an embodiment of the present invention; Figure 2 is a schematic structural diagram of a CTSP solving system based on a convolutional enhancement proxy attention mechanism provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0012] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all of them. It should be understood that these descriptions are only exemplary and not intended to limit the scope of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0013] The CTSP solving method based on the convolutional enhancement proxy attention mechanism provided by the embodiments of the present invention will be described in detail below with reference to the drawings through specific embodiments and their application scenarios.

[0014] Refer to the accompanying drawings of the specification Figure 1 , which shows a schematic flowchart of a CTSP solving method based on a convolutional enhancement proxy attention mechanism provided by an embodiment of the present invention.

[0015] The embodiments of the present invention provide a CTSP solving method based on a convolutional enhancement proxy attention mechanism, which may include the following steps: S1: Obtain city data and salesperson data.

[0016] In a possible implementation, the city data includes: the location coordinates of each city and the color markers assigned to each city, where each city includes one or more color attributes.

[0017] The salesperson data includes: each traveling salesman is assigned a color marker.

[0018] In the embodiments of the present invention, the practice of assigning color markers to each city and salesperson effectively realizes the management of constraints, the flexibility and sharing of path planning, the optimization of computational efficiency, and enhances the scalability of the model.

[0019] S2: Based on the city data and salesperson data, construct a complete graph for describing the CTSP problem instance.

[0020] Among them, TSP is a multi-salesperson constrained path planning problem, which requires collaborative access to all nodes of the entire graph while satisfying color access constraints and optimizing the path cost.

[0021] In a possible implementation, S2 specifically includes: S201: Define a set of city nodes.

[0022] S202: Assign coordinates and color attributes to each node in the node set.

[0023] S203: Define an edge set.

[0024] S204: Calculate the Euclidean distance between each pair of nodes as the edge weight.

[0025] S205: Add access color constraints to each salesperson, where the access color constraint is that each salesperson can only visit city nodes with matching colors.

[0026] S206: Based on the node set, edge set, edge set, and access color constraints, construct a complete graph for describing the CTSP problem instance.

[0027] Specifically, a colored traveling salesman problem (CTSP) is defined on the complete graph where the vertex set represents the cities to be visited, and the edge set contains all undirected connections between cities (i.e., any two cities are directly reachable). Each salesperson is assigned a unique color in the set . The problem requires for each color . Design a Hamiltonian cycle for the salesperson that follows specific color constraints, ultimately achieving full coverage (visiting all cities exactly once), while optimizing the global path cost, such as minimizing the total distance or the longest single-path distance. The complete graph structure means the number of edges Each node is associated with a planar coordinate vector The color mapping function assigns a non-empty subset of colors to each node where . Let represent the set of cities that the salesperson (associated with color ) can visit, which satisfies the color compatibility condition: .

[0028] It should be noted that any multi-colored node (where ) essentially belongs to multiple salesperson-specific sets , thus allowing cross-agent task sharing while maintaining color constraints.

[0029] In the embodiments of the present invention, by uniformly modeling city nodes, coordinate information, color attributes, salesperson access rules, and edge weights as a structured complete graph, not only is an explicit expression of color access constraints achieved, but also a clear graph structure basis is provided for subsequent path generation and optimization algorithms, further improving the efficiency, flexibility, and controllability of path planning.

[0030] S3: Based on the complete graph, construct a constrained Markov decision process model. By defining the state space, action space, transition dynamics, and reward function of the constrained Markov decision process model, model the path generation process of the CTSP problem as a step-by-step decision-making mechanism.

[0031] Among them, the constrained Markov decision process model (Constrained Markov Decision Process, CMDP) is an extended modeling framework that introduces constraint conditions on the basis of the standard Markov decision process (MDP). Its goal is not only to maximize the cumulative reward, but also to meet certain cost or resource constraints.

[0032] In a possible implementation, S3 specifically includes: S301: Based on the complete graph, define the state space of the constrained Markov decision process model, where the states in the state space include the identifier of the current salesperson, the set of currently visited nodes, the current node position, and the color attributes carried by the current salesperson.

[0033] In the embodiment of the present invention, this step clearly expresses the decision context for constructing the current path, including: the identity of the salesperson, the current location, the historical access status, and the color attribute, providing comprehensive state input information for the policy function and facilitating the learning of effective decisions.

[0034] S302: Define the action space of the constrained Markov decision process model based on the complete graph.

[0035] It should be noted that the action space is defined as the node selection operation that satisfies the color constraint. For example, the action that meets the color requirements of the target node is added to the current path.

[0036] In the embodiment of the present invention, the node selection process is explicitly transformed into a set of legal actions constrained by color, ensuring that the path construction is only carried out in accessible cities, realizing effective action screening in the policy generation stage, and avoiding invalid calculations and illegal paths.

[0037] S303: Design the transition dynamics of the constrained Markov decision process model.

[0038] It should be noted that the transition dynamics achieve constraint compliance through deterministic state updates and invalid action masking.

[0039] In the embodiment of the present invention, through deterministic state updates, the path generation process is made controllable. At the same time, combined with the invalid action masking mechanism, illegal states such as repeated access or unauthorized access can be automatically excluded, strengthening the path legality.

[0040] S304: Set the reward function of the constrained Markov decision process model: ; where represents the immediate reward at the t -th time step, represents the node to be visited at the t +1-th time step, represents the node visited at the t -th time step, represents the Euclidean distance norm.

[0041] S305: Define the policy function of the constrained Markov decision process model.

[0042] S306: Based on the state space, action space, transition dynamics, reward function, and policy function, construct a decision-making mechanism for gradually generating a path solution that meets the constraint conditions: ; where Represents the policy function, Represents a valid path solution, Represents a valid path solution, Represents the path of the th salesperson, Represents the node selected by the th salesperson at the th time step, Represents the t node selected by the

[0043] th salesperson at the

[0044] -1 time step. Simply put, we input the CTSP (Constrained Traveling Salesman Problem) instance into the encoder, converting the node coordinates and their attributes into a unified representation space. Then, the decoder generates the solution autoregressively. At each time step, based on the node embeddings selected at the previous time point and the color attributes carried by the salesperson, the decoder calculates the probability of each node being selected in the next state. Through the masking operation, the algorithm masks the nodes that have already been visited and any options that violate the color constraints, thereby determining the node to be selected in the next state. This process continues until all the nodes of all cities have been visited.

[0045] S4: Construct a neural policy network based on convolutional enhancement and proxy attention mechanism, where the neural policy network is used to solve the constrained Markov decision process model.

[0046] Among them, the neural policy network is an intelligent path generation policy generator that integrates graph structure modeling, context awareness, and attention mechanism. It can generate an optimized path through step-by-step probability selection while satisfying the access constraints, and is the core decision-making module for solving complex path planning problems such as CTSP.

[0047] In a possible implementation manner, S4 specifically includes: S401: Process the complete graph through the problem initialization network to generate the initial input vector of the neural policy network.

[0048] In a possible implementation manner, S401 specifically includes: S4011: Extract the spatial coordinates and color attributes of each node in the complete graph, and construct a node feature aggregation tensor according to the spatial coordinates and color attributes of each node.

[0049] S4012: Perform a linear projection transformation on each coordinate in the node feature aggregation tensor to obtain a coordinate embedding: ; Among them, represents the position embedding, represents the weight matrix of the warehouse node coordinates, represents element-wise multiplication, represents the i coordinates of the th node city, represents the bias matrix of the warehouse node coordinates, represents the weight matrix of the city node coordinates,

[0050] S4013: Perform a linear projection transformation on the color attribute vector in the node feature aggregation tensor to obtain a color embedding: ; Among them, represents the color embedding, represents the weight matrix of the warehouse node color attributes, represents the i th node city's color attribute, represents the bias matrix of the warehouse node color attributes, represents the weight matrix of the city node color attributes, represents the bias matrix of the city node color attributes.

[0051] S4014: Fuse the coordinate embedding and the color embedding to generate an initial input vector for the neural policy network: ; Among them, represents the initial input vector.

[0052] In the embodiment of the present invention, independent linear projections are respectively performed on the coordinate and color information to achieve dual representations of space and semantics. At the same time, the initial input vector has both geometric position and access constraint context information, providing an expression basis for deep network modeling.

[0053] S402: Extract the deep features of the initial input vector through the convolutional enhancement and proxy attention mechanism in the encoder to obtain a deep feature embedding vector.

[0054] In a possible implementation manner, S402 specifically includes: S4021: Extract the global semantic features in the initial input vector through the proxy attention mechanism: ; Among them, represents the proxy attention mechanism, represents the attention function, represents the query matrix, represents the proxy matrix, represents the value matrix.

[0055] S4022: Extract local geometric features in the initial input vector through convolution enhancement: ; Among them, represents the convolution enhancement embedding of the i th layer, represents the activation function, represents the convolution operation, represents the value matrix.

[0056] S4023: Fuse the global semantic features and local geometric features, and perform non-linear transformation and residual connection through a feed-forward neural network to obtain a deep feature embedding vector: ; Among them, represents the node embedding vector of the i th layer, represents the root mean square normalization layer, represents the deep feature embedding vector, represents the feed-forward layer.

[0057] In the embodiment of the present invention, S403: Construct the context variable of the current time step based on the deep feature embedding vector and the historical state information of path construction, where the context variable includes the average embedding vector of the graph, the node embedding vector of the node selected in the previous step, and the node color attribute.

[0058] In the embodiment of the present invention, by integrating the proxy attention mechanism and the convolution enhancement module in the encoder, jointly extracting the global semantic features and local geometric features of the initial input vector, and fusing and constructing a high-dimensional deep embedding representation, the representation ability and discrimination ability of the neural policy network for node features are significantly improved, and the perception effect of the model on complex graph structures and color constraint semantics is enhanced.

[0059] In a possible implementation manner, S403 specifically includes: S4031: Perform an average calculation on the deep feature embedding vector to obtain an average embedding vector.

[0060] S4032: Extract the node embedding vector and node color attribute of the node selected in the previous step in the historical path.

[0061] S4033: Concatenate the average embedding vector, the node embedding vector, and the node color attribute to construct the context variable at the current time step: ; Wherein, represents the context variable, represents the concatenation function, represents the average embedding vector, represents the embedding vector of the warehouse node, represents the color carried by the salesperson starting from the warehouse node, represents the t- embedding vector of the node selected at the 1st time step, represents the t- color attribute carried by the salesperson at the 1st time step.

[0062] In the embodiment of the present invention, by fusing the average semantic representation of all nodes in the graph, the embedding vector of the node selected in the previous step in the historical path, and its color attribute, the context variable at the current time step is constructed, enabling the policy network to perceive the historical state and graph structure information in the path construction process, thereby realizing the context-aware node selection mechanism and improving the rationality and coherence of path generation.

[0063] S404: Based on the context variable and the node embedding vector, determine the node selection probability distribution through the proxy attention mechanism and the convolutional enhancement module in the decoder.

[0064] Among them, the proxy attention (Agent Attention) is an improved self-attention mechanism, aiming to reduce the computational complexity while retaining the global semantic modeling ability.

[0065] Among them, the convolutional enhancement module is an improved method for compensating for the weakness of the attention mechanism in perceiving local structures. It captures local geometric patterns and neighborhood structure information by adding a lightweight convolutional network before or after the attention module, retains the spatial position information, and enhances the feature diversity.

[0066] In a possible implementation manner, S404 specifically includes: S4041: Use the context variable as the query vector, the node embedding vector as the key vector and the value vector, and calculate the attention score of each candidate node through the attention layer jointly composed of the proxy attention mechanism and the convolutional enhancement: ; Wherein, represents the attention score of the i th candidate node, C represents the hyperparameter, represents the hyperbolic tangent operator, d represents the dimension of the embedding vector.

[0067] Optionally, C = 10.

[0068] S4042: Perform validity judgment on each candidate node, and filter out the target candidate nodes that meet the gating conditions: ; where, represents the color of the node selected at the t -th time step, represents the color of the node selected at the t- 1-th time step, represents the empty set, represents the nodes visited at the -th time step.

[0069] S4043: Normalize the attention scores of each target candidate node to generate a node selection probability distribution: ; where, represents the probability that the edge is connected to the i-th candidate node at the t -1-th time step, represents the feature representation of the edge between the connected nodes at the t- 1-th time step for the i -th candidate node, e represents the exponential function, represents the i -th attention score of the candidate node, represents the value of the attention score of the i-th candidate node after exponential transformation, n represents the total number of target candidate nodes, j represents the index variable for summation, represents the sum of the exponential attention scores of all candidate nodes.

[0070] In the embodiments of the present invention, by using the context variable as the query vector and combining the node embedding vector as the key-value input into the attention calculation layer that fuses the agent attention and the convolutional enhancement mechanism, the importance of each candidate node at the current time step is accurately modeled; at the same time, the color compatibility and access uniqueness rules are combined for legal screening, and a differentiable node selection probability distribution is output through normalization, so as to realize a constraint-aware and context-driven path selection strategy, significantly improving the compliance and strategy quality of path generation.

[0071] S405: Select the target node accessed at the current time step by means of greedy selection according to the node selection probability distribution, and append the target node to the current path sequence.

[0072] S406: Update the current path construction state and the candidate node set according to the target node, and repeat the process of S403 to S405 until all city nodes have been visited.

[0073] In the embodiment of the present invention, by making greedy node decisions based on the node selection probability distribution at each time step, and dynamically updating the candidate node set and the path state during the path construction process, the step-by-step construction of the city path sequence that satisfies the color access constraint is realized. This path generation method has a clear control flow, good computational efficiency and trainability, and can efficiently complete the complete coverage of the city nodes in the entire graph while maintaining legality, thereby generating an optimal initial path plan with reasonable structure and compliant constraints.

[0074] S5: Use the reinforcement learning algorithm to train the neural policy network based on convolutional enhancement and proxy attention mechanism.

[0075] Among them, reinforcement learning (RL) is a machine learning method in which a learning agent interacts with the environment, obtains a reward signal through trial and error, and thus learns an optimal policy.

[0076] Optionally, the training steps of the model specifically include: In each training round, resample the input graph instance according to the path generation policy to generate a node permutation sequence.

[0077] Based on each graph instance and the corresponding starting node, generate a path trajectory through the policy function.

[0078] Calculate the cumulative reward corresponding to each sampled path: ; Among them, represents the cumulative reward corresponding to the sampled path , represents a sampled path, represents the moment when the path ends, represents the discount factor, represents the immediate reward obtained at the t-th time step.

[0079] It should be noted that the discount factor is used to reduce the weight of future rewards to reflect the view that immediate rewards are more certain and valuable than long-term rewards.

[0080] Using the REINFORCE method, the expected return of the policy function is estimated by the cumulative reward on the sampled trajectory: ; where, represents the gradient estimate with respect to the parameter which is used to update the parameters of the policy function to maximize the expected value of the cumulative reward, represents the output of the line function, which depends on the current state or context G , represents the sampled trajectory corresponding cumulative reward, represents that this represents the policy function with respect to the parameter log-likelihood gradient.

[0081] Calculate the shared baseline based on the POMO policy.

[0082] Specifically, using the POMO structure, in each training round, from N different starting nodes to generate N greedy trajectories, and using the average reward of these trajectories as the shared baseline to reduce the variance of the gradient estimate.

[0083] Repeat the processes of path sampling, reward evaluation, baseline calculation and gradient update, and use the optimizer to update and iteratively optimize the policy network parameters until the model reaches the expected convergence condition or the training round upper limit on the training set.

[0084] In the embodiment of the present invention, by introducing a reinforcement learning algorithm to train the neural policy network constructed based on convolutional enhancement and proxy attention mechanism, the model can continuously learn to generate high-quality path solutions through path sampling, reward feedback and policy optimization without unsupervised annotation. The REINFORCE method is used for unbiased policy gradient estimation, and combined with the POMO multi-start greedy solution as the shared baseline, effectively reducing the training variance and improving the policy stability and convergence speed. This method significantly enhances the path optimization ability of the model under the premise of meeting the color constraints.

[0085] S6: Generate a preliminary optimal path plan through the trained neural policy network based on convolutional enhancement and proxy attention mechanism.

[0086] Specifically, the model uses the trained neural policy network, receives the graph instance input, and constructs paths for each salesman in turn according to the color constraints through encoding embedding, context construction and greedy policy inference, and finally generates a set of preliminary feasible path plans that meet the constraints.

[0087] In the embodiments of the present invention, by using the trained neural policy network, in the inference stage, the color traveling salesman problem instance is embedded and context-aware modeled, and combined with the greedy selection strategy, a preliminary feasible path solution that meets the color constraints is quickly generated. This path generation process does not rely on traditional heuristic or enumeration algorithms, and has advantages such as explicit constraint control, full-graph structure perception, and decoding efficiency. It can significantly improve the path construction speed and scalability while ensuring the legality of the solution, and provide a high-quality initial solution with a reasonable structure for subsequent fine optimization.

[0088] S7: Input the optimal path solution into a traditional optimizer for fine optimization to obtain the final optimal path solution.

[0089] Optionally, the traditional optimizer includes Gurobi, COPT, and OR-Tools.

[0090] In the embodiments of the present invention, based on the initial solution with a reasonable structure and compliant colors provided by the neural model, the traditional optimizer can focus on local path adjustment or global reconstruction, thereby significantly shortening the optimization convergence time and obtaining a better path length or cost metric.

[0091] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include: In the embodiments of the present invention, by constructing a neural policy network based on convolutional enhancement and proxy attention mechanism, combining the attention structure and the local convolutional feature extraction module, the inference efficiency and scalability of the model in large-scale graph instances are significantly improved. At the same time, a reinforcement learning algorithm is introduced to train the network, enabling the model to have an efficient policy exploration ability and quickly generate a preliminary high-quality solution. On this basis, a traditional optimizer is further introduced for fine optimization, realizing the effective integration of the fast modeling advantage of the neural network and the accurate solving ability of the traditional algorithm, taking into account both the quality of the solution and the computational efficiency, and achieving a good balance between the solution accuracy and computational resources.

[0092] Refer to the attached Figure 2 description, which shows a schematic structural diagram of a CTSP solving system based on a convolutional enhancement proxy attention mechanism provided in the embodiments of the present invention.

[0093] The embodiments of the present invention provide a CTSP solving system 20 based on a convolutional enhancement proxy attention mechanism, including: a processor 201 and a memory 202; The memory 202 stores programs or instructions that can run on the processor 201. When the programs or instructions are executed by the processor 201, they implement the steps of the above-mentioned CTSP solving method based on a convolutional enhancement proxy attention mechanism and can achieve the same technical effects. To avoid repetition, the present invention will not elaborate further.

[0094] It should be understood that the processor 201 in the embodiments of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0095] It should also be understood that the memory 202 in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0096] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that contains one or more sets of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0097] It should be understood that in various embodiments of the present invention, the order numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0098] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0099] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.

[0100] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.

[0101] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0102] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0103] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0104] The embodiment of the present invention provides a readable storage medium including: programs or instructions are stored on the readable storage medium. When the programs or instructions are executed by a processor, the steps of the above-mentioned CTSP solving method based on the convolutional enhanced proxy attention mechanism are implemented, and the same technical effects can be achieved. To avoid repetition, the present invention will not be described in detail again.

[0105] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present invention, rather than to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. Any changes or replacements that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.

Claims

1. A CTSP solution method based on a convolutional enhanced proxy attention mechanism, characterized in that Including: S1: Obtain city data and salesman data; S2: Based on the city data and the salesman data, construct a complete graph for describing the CTSP problem instance; S3: Based on the complete graph, construct a constrained Markov decision process model. By defining the state space, action space, transition dynamics, and reward function of the constrained Markov decision process model, model the path generation process of the CTSP problem as a step-by-step decision-making mechanism; S4: Construct a neural policy network based on convolutional enhancement and proxy attention mechanism, where the neural policy network is used to solve the constrained Markov decision process model; S5: Use the reinforcement learning algorithm to train the neural policy network based on convolutional enhancement and proxy attention mechanism; S6: Generate a preliminary optimal path plan through the trained neural policy network based on convolutional enhancement and proxy attention mechanism; S7: Input the optimal path plan into a traditional optimizer for fine optimization to obtain the final optimal path plan.

2. The CTSP solving method based on the convolutional enhanced proxy attention mechanism according to claim 1, wherein The city data includes: the location coordinates of each city and the color markers assigned to each city, where each city includes one or more color attributes; The salesman data includes: assigning a color marker to each traveling salesman.

3. The CTSP solving method based on the convolutional enhanced proxy attention mechanism according to claim 1, characterized in that The specific content of S2 includes: S201: Define the city node set; S202: Assign coordinates and color attributes to each node in the node set; S203: Define the edge set; S204: Calculate the Euclidean distance between each node as the edge weight; S205: Add access color constraints to each salesman, where the access color constraint is that each salesman can only access city nodes with matching colors; S206: Based on the node set, the edge set, the edge set, and the access color constraint, construct a complete graph for describing the CTSP problem instance.

4. The CTSP solving method based on a convolutional enhanced proxy attention mechanism according to claim 3, characterized in that The specific content of S3 includes: S301: Based on the complete graph, define the state space of the constrained Markov decision process model, where the states in the state space include the identifier of the current salesman, the set of currently visited nodes, the current node position, and the color attributes carried by the current salesman; S302: Based on the complete graph, define the action space of the constrained Markov decision process model; S303: Design the transition dynamics of the constrained Markov decision process model; S304: Set the reward function of the constrained Markov decision process model: ; Among them, represents the immediate reward at the t -th time step, represents the node to be visited at the t +1-th time step, represents the node visited at the t -th time step, represents the Euclidean distance norm; S305: Define the policy function of the constrained Markov decision process model; S306: Based on the state space, the action space, the transition dynamics, the reward function, and the policy function, construct a decision-making mechanism for gradually generating a path solution that meets the constraint conditions: ; Among them, represents the policy function, represents a valid path solution, represents the number of salespersons, represents the th path of the salesperson, represents the th salesperson's selected node at the th time step, represents the complete graph, represents the th salesperson's selected node at the t -1th time step.

5. The CTSP solving method based on the convolutional enhanced proxy attention mechanism according to claim 1, wherein, The neural policy network includes a problem initialization network, an encoder, and a decoder; the specific content of S4 includes: S401: Process the complete graph through the problem initialization network to generate the initial input vector of the neural policy network; S402: Extract the depth features of the initial input vector through the convolution enhancement and proxy attention mechanism in the encoder to obtain a depth feature embedding vector; S403: Construct the context variable at the current time step based on the depth feature embedding vector and the historical state information of path construction, where the context variable includes the average embedding vector of the graph, the node embedding vector of the previously selected node, and the node color attribute; S404: Based on the context variable and the node embedding vector, determine the node selection probability distribution through the proxy attention mechanism and the convolution enhancement module in the decoder; S405: According to the node selection probability distribution, select the target node accessed at the current time step by means of greedy selection, and append the target node to the current path sequence; S406: Update the current path construction state and the candidate node set according to the target node, and repeat the process of S403 to S405 until all city nodes are visited.

6. The CTSP solving method based on the convolutional enhanced proxy attention mechanism according to claim 5, characterized in that The S401 specifically includes: S4011: Extract the spatial coordinates and color attributes of each node in the complete graph, and construct a node feature aggregation tensor according to the spatial coordinates and color attributes of each node; S4012: Perform a linear projection transformation on each coordinate in the node feature aggregation tensor to obtain a coordinate embedding; ; Among them, represents the positional embedding, represents the weight matrix of the warehouse node coordinates, represents element-wise multiplication, represents the i coordinates of the represents the bias matrix of the warehouse node coordinates, represents the weight matrix of the city node coordinates, represents the bias matrix of the city node coordinates; S4013: Perform a linear projection transformation on the color attribute vector in the node feature aggregation tensor to obtain a color embedding; ; Among them, represents color embedding, represents the weight matrix of the color attributes of the warehouse nodes, represents the i th color attribute of the node city, represents the bias matrix of the color attributes of the warehouse nodes, represents the weight matrix of the color attributes of the city nodes, represents the bias matrix of the color attributes of the city nodes; S4014: Fuse the coordinate embedding and the color embedding to generate the initial input vector of the neural policy network; ; Among them, represents the initial input vector.

7. The CTSP solving method based on the convolutional enhanced proxy attention mechanism according to claim 5, characterized in that The S402 specifically includes: S4021: Extract the global semantic features in the initial input vector through the proxy attention mechanism; ; Among them, represents the proxy attention mechanism, represents the attention function, represents the query matrix, represents the proxy matrix, represents the value matrix; S4022: Extract the local geometric features in the initial input vector through the convolution enhancement; ; Among them, represents the convolutional enhanced embedding of the i th layer, represents the activation function, represents the convolution operation, represents the value matrix; S4023: Fuse the global semantic features and the local geometric features, and perform a non-linear transformation and residual connection through a feed-forward neural network to obtain the depth feature embedding vector; ; in, Indicates i The node embedding vector of the layer, represents the RMS normalization layer, represents the deep feature embedding vector, represents a feed-forward layer.

8. The CTSP solving method based on the convolutional enhanced proxy attention mechanism according to claim 5, characterized in that The S403 specifically includes: S4031: Perform an average calculation on the depth feature embedding vector to obtain an average embedding vector; S4032: Extract the node embedding vector and the node color attribute of the previously selected node in the historical path; S4033: Connect the average embedding vector, the node embedding vector, and the node color attribute to construct the context variable at the current time step; ; Among them, represents a context variable, represents a concatenation function, represents an average embedding vector, represents the embedding vector of a warehouse node, represents the color carried by the salesperson starting from the warehouse node, represents the t- embedding vector of the selected node at the 1st time step, represents the t- color attribute carried by the salesperson at the 1st time step.

9. The CTSP solving method based on the convolutional enhanced proxy attention mechanism according to claim 5, wherein, The S404 specifically includes: S4041: Take the context variable as the query vector, the node embedding vector as the key vector and the value vector, and calculate the attention score of each candidate node through the attention layer jointly composed of the proxy attention mechanism and the convolution enhancement; ; Among them, represents i the attention score of the C th candidate node, represents a hyperparameter, d represents the dimension of the embedding vector; S4042: Perform a validity judgment on each candidate node to screen out the target candidate nodes that meet the gating conditions; ; Among them, represents the color of the node selected at the t th time step, represents the color of the node selected at the t- 1st time step, represents the empty set, represents the nodes visited at the th time step; S4043: Normalize the attention scores of each target candidate node to generate the node selection probability distribution; ; Among them, represents the probability that the edge at t -1 time steps is connected to the i-th candidate node, represents the edge between the connected nodes at t- 1 time step, the i -th candidate node's feature representation, e represents the exponential function, represents the i -th candidate node's attention score, represents the value after the exponential transformation of the i-th candidate node's attention score, n represents the total number of target candidate nodes, j represents the summation index variable, represents the sum of the exponential attention scores of all candidate nodes.

10. A CTSP solving system based on a convolutional enhanced proxy attention mechanism, characterized in that, Includes: A processor and a memory; The memory stores a program or instructions that can run on the processor. When the program or instructions are executed by the processor, the steps of the CTSP solving method based on the convolutional enhanced proxy attention mechanism according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Multicolor spraying path planning method and system based on hybrid heuristic algorithm

    CN114648162A

  • Method for solving coloring traveling salesman problem with time window constraint

    CN116523159A

  • Large-scale traveling salesman problem solving method based on deep reinforcement learning

    CN119047671A

  • Optical remote sensing image salient target detection method based on double-branch progressive interaction and multi-direction feature enhancement

    CN119992068A

  • Real-time map route construction method and system

    CN120146346A

Cited By

  • Multi-starting-point sequence decision reinforcement learning method based on dynamic mask attention

    CN121119029A