A Multi-Decoder Dynamic Attention Model for Container Rehandling

The multi-decoder dynamic attention model enhances CRP solution diversity and quality by using a feature enhancer and multiple decoders trained with reinforcement learning, addressing information loss and solution diversity issues in existing models.

CN118965953BActive Publication Date: 2025-07-15CHONGQING NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410935787.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-12
Publication Date
2025-07-15
Estimated Expiration
2044-07-12

AI Technical Summary

Technical Problem

The existing dynamic attention model has problems of information loss and insufficient diversity in the container rushing problem, which leads to the lack of diversity in the generated solutions and it is difficult to find the optimal solution.

Method used

The multi-decoder dynamic attention model designed based on the Transformer architecture is adopted, the container stack semantics are extracted through feature enhancers and sampled using multiple decoders. The model is trained in combination with a reinforcement learning algorithm of greedy baselines to increase the diversity of outputs and the diversity of decoders.

Benefits of technology

It improves the diversity and efficiency of container rumbling operations, reduces the number of rumbling times, and improves the quality of solutions, especially in large-scale problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118965953B_ABST
    Figure CN118965953B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of container stowing and unstowing, and specifically relates to a multi-decoder dynamic attention model for container unstowing, including a multi-decoder dynamic attention model designed based on the Transformer architecture and including feature enhancement; training the multi-decoder dynamic attention model using a reinforcement learning algorithm based on a greedy baseline: inputting a bay configuration, extracting the container stack semantics through a feature enhancer to form a bay configuration containing the inter-stack semantics; using the trained multi-decoder dynamic attention model to predict the unstowing operation sequence of the input bay. This method designs a stack feature enhancer to enhance the input bay configuration information, helps the encoder capture the relationship between stacks, and the decoder dynamic attention model trains multiple construction strategies to increase the diversity of the output. During the training process, each decoder learns different solution patterns and is regularized through a divergence loss to force the decoder to output different probability distributions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of container turning and re-turning, and particularly to a multi-decoder dynamic attention model for container turning. Background Art

[0002] In a container terminal yard, container turning is a common and important operation. Reducing the number of container turnings can not only improve the operation efficiency of the terminal, but also reduce the moving distance of handling equipment, thereby reducing the non-productive costs of the terminal. Therefore, it is particularly important to study the container turning problem and reduce the number of container turnings in the yard. The container turning problem (Container Relocation Problem, CRP), also known as the block relocation problem (Block Relocation Problem, BRP), is an important combinatorial optimization problem that appears in yard management. The goal of CRP is to minimize the total number of turnings (or re-handlings) required to retrieve all containers from the bays.

[0003] CRP has multiple variants and can be classified from the following two perspectives. First, consider the uniqueness of the priority values of the containers. In CRP with different priorities, each container is assigned a unique priority. Therefore, the retrieval order of the containers is determined. On the contrary, in CRP with duplicate priorities (or group priorities), the priorities within a group of containers can be the same, and the retrieval order among them is arbitrary. Another classification method is to consider the restrictions on turning. In restricted CRP (RCRP), only the containers in the stack where the highest-priority container is located are allowed to be relocated. On the contrary, in unrestricted CRP (UCRP), there is no such restriction, and any container on the top layer of any stack can be relocated.

[0004] Deep learning has shown good capabilities in solving NP-hard decision problems and achieved remarkable results in various fields. For example, deep reinforcement learning (DRL) is used to solve the Go game. However, for the CRP problem, the research on using deep learning methods is still limited. Different from the traditional attention model (AM), the dynamic attention model (DAM) considers dynamic state information. After each operation, the new bay construction is re-input into the DAM to accurately determine the next operation. In the DAM, the bay construction is represented by a two-dimensional matrix and used as the input of the DAM. The multi-head attention layer extracts the features of the stacks in the bay construction. Based on these stack features, the decoder gradually constructs the solution in an autoregressive manner. Autoregressive means that the decoder selects a target stack according to the output probability distribution each time and adds it to the partial solution until the complete solution is constructed. To obtain a better solution, a sampling strategy is adopted to generate a set of solutions from the trained neural network model and then select the optimal one from this set of solutions.

[0005] However, DAM has the following defects. First, directly inputting the two-dimensional matrix into the neural network model may lose some important information of the stacks, and the multi-head attention may not be able to extract some key features. Second, in the decoding stage of DAM, a single decoder is used to output the solution, and the generated solution will lack diversity. Intuitively, increasing the diversity of the solutions will potentially bring better solutions. Because for CRP, there are multiple optimal solutions, and by generating diverse solutions, the possibility of finding the optimal solution can be increased. In addition, in a set of solutions generated by sampling, due to insufficient generation diversity, some solutions are the same, which will reduce the search range of the solution space. The existing method DAM only trains one construction strategy and samples solutions from this strategy. The diversity only comes from sampling, but since the probability distribution is determined, this results in insufficient diversity. Summary of the Invention

[0006] The purpose of the present invention is to provide a multi-decoder dynamic attention model for container reshuffling, aiming to solve the problem of insufficient diversity in the existing dynamic attention model that only trains one construction strategy.

[0007] To achieve the above purpose, the present invention provides a multi-decoder dynamic attention model for container reshuffling, including the following steps:

[0008] Design a multi-decoder dynamic attention model with feature enhancement based on the Transformer architecture;

[0009] Train the multi-decoder dynamic attention model based on the reinforcement learning algorithm with a greedy baseline;

[0010] Input a bay structure, extract the container stack semantics through a feature enhancer, and form a bay structure containing the inter-stack semantics;

[0011] Use the trained multi-decoder dynamic attention model to predict the sequence of container repositioning operations for the input bay.

[0012] Among them, the multi-decoder dynamic attention model designed based on the Transformer architecture and including feature enhancement includes a designed feature enhancer, an encoder, and a decoder.

[0013] Among them, the specific method of the designed feature enhancer:

[0014] Define five original features representing the inter-stack semantics based on heuristic rules and manual experience;

[0015] Based on the input bay structure, calculate the five original features and represent them in matrix form to obtain the inter-stack semantic feature matrix;

[0016] Connect the bay structure matrix and the inter-stack semantic feature matrix to form a bay structure matrix containing the inter-stack semantics.

[0017] Among them, the original features include: the height percentage of the current stack, the vacancy quantity percentage of the stack, the stack priority, the priority of the target container, and the number of disorderly stacked containers in the stack.

[0018] Among them, the specific method of the designed feature encoder:

[0019] Input the output of the feature enhancer into a long short-term memory network to form the initial embedding required by the Transformer;

[0020] Input the output of the long short-term memory network into the encoding end of the Transformer, and extract the global semantic features of the bay structure through the self-attention mechanism and residual connection.

[0021] A multi - decoder dynamic attention model for container re - stowing in the present invention designs a multi - decoder dynamic attention model with feature enhancement based on the Transformer architecture; trains the multi - decoder dynamic attention model using a reinforcement learning algorithm based on a greedy baseline: inputs a bay structure, extracts the container stack semantics through a feature enhancer to form a bay structure containing inter - stack semantics; uses the trained multi - decoder dynamic attention model to predict the re - stowing operation sequence of the input bay. This method designs a stack feature enhancer to enhance the input bay structure information, effectively helping the encoder capture the mutual relationship between stacks. The decoder dynamic attention model trains multiple construction strategies to increase the diversity of the output. It encodes the stack information using an encoder similar to Transformer and samples different solutions using multiple different attention decoders with non - shared parameters. During training, each decoder learns different solution patterns and is regularized through the Kullback - Leibler (KL) divergence loss to force the decoder to output different probability distributions, solving the problem of insufficient diversity in existing dynamic attention models that only train one construction strategy. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following - described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0023] Figure 1 is the overall flowchart of a multi - decoder dynamic attention model for container re - stowing provided by the present invention.

[0024] Figure 2 is a schematic diagram of training the multi - decoder dynamic attention model using a reinforcement learning algorithm based on a greedy baseline.

[0025] Figure 3 is a schematic diagram of predicting the re - stowing operation sequence of the input bay using the trained multi - decoder dynamic attention model. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] The following details the embodiments of the present invention. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are intended to explain the present invention, and should not be construed as a limitation of the present invention.

[0027] Please refer toFigures 1 to 3 , the present invention provides a multi - decoder dynamic attention model for container re - stowing, including the following steps:

[0028] S1 Design a multi - decoder dynamic attention model with feature enhancement based on the Transformer architecture;

[0029] Specifically, designing a multi - decoder dynamic attention model with feature enhancement based on the Transformer architecture includes designing a feature enhancer, an encoder, and a decoder. Based on the Transformer network architecture, the multi - decoder dynamic attention model (MDDAM) is used to train multiple construction strategies to increase the diversity of outputs. The encoding end of MDDAM integrates a stack feature enhancer (SFE) to improve the encoder's ability to capture the semantics between container stacks; the decoding end integrates multiple decoders, each decoder based on multi - head attention, and the diversity between decoders is maintained by KL divergence.

[0030] The specific way to design the feature enhancer (SFE):

[0031] S101 Define five original features representing the semantics between stacks based on heuristic rules and manual experience;

[0032] Specifically, based on heuristic rules and manual experience, define five original features representing the semantics between stacks. These five features are: the height percentage of the current stack (1 / T, 2 / T, …, 1); the percentage of the number of empty spaces in the stack; the stack priority; the priority of the target container; the number of unordered stacks (US) in the stack.

[0033] S102 Calculate the five original features based on the input bay construction and represent them in matrix form to obtain the semantic feature matrix between stacks;

[0034] Specifically, according to the input bay construction, calculate five semantic features between stacks and represent them in matrix form. The rows of the matrix have 5 components, and each component corresponds to a semantic feature value between stacks.

[0035] S103 Connect the bay construction matrix and the semantic feature matrix between stacks to form a bay construction matrix containing the semantics between stacks.

[0036] The specific way to design the encoder:

[0037] S111 Input the output of the feature enhancer into a long - short - term memory network to form the initial embedding required by Transformer;

[0038] Specifically, the output of SFE enters a long - short - term memory network (LSTM) to form the initial embedding required by Transformer.

[0039] S112 inputs the output of the long short-term memory network into the encoding end of the Transformer, and extracts the global semantic features of the bay structure through the self-attention mechanism and residual connection.

[0040] Specifically, the output of the LSTM enters the encoding end of the Transformer, and the global semantic features of the bay structure are extracted through the self-attention mechanism and residual connection.

[0041] Design a decoder.

[0042] Design a number of decoders with the same structure, and each decoder consists of a multi-head attention network;

[0043] Define the KL divergence between decoders and add it as a regular term to the loss function in training to maintain the diversity of the output solutions;

[0044] The ablation experiment verifies the effectiveness of the SFE and multi-decoder (MD) designs.

[0045] Specifically, using the average number of box flips as a metric, both SFE and MD effectively improve the quality of the CRP solution, especially when the problem scale is large, the effect is more obvious.

[0046] S2 trains the multi-decoder dynamic attention model based on the reinforcement learning algorithm of the greedy baseline;

[0047] Specifically, each decoder in multiple decoders outputs a different probability distribution, and independently samples a solution π d , to obtain a separate reinforcement learning loss:

[0048] where L(π d ) is the sequence length of the solution π d , that is, the total number of box flips. Since the stack selection action is sampled from the probability distribution and the total number of box flips is used as the cost calculation, due to the uncertainty of sampling, the costs recorded in different episodes have a large variance, thus affecting the convergence speed (Li et al., 2022). To reduce the variance and accelerate the learning speed, a baseline b(X) is usually introduced. The baseline we adopt is similar to the baseline in (Kool et al., 2019). b(X) is the baseline policy p θ*The total number of rehandling times for deterministic greedy rollout. The baseline policy is the currently best-trained model. At the end of each training cycle, once the currently trained model is better than the baseline policy, we replace the baseline policy with the current model. The parameters θ of the model can be updated using the REINFORCE algorithm by gradient descent:

[0049]

[0050] where represents the REINFORCE loss gradient, is the gradient of the KL divergence between the decoder output probability distributions, and k KL is the coefficient of the KL loss. The negative sign in front of k KL is because we want to maximize the diversity between decoders. Generally speaking, for each state encountered by each decoder, the KL loss should be calculated during training to encourage the decoder to generate different solutions. But we only apply the KL loss at the first step for the following reasons: on the one hand, it is to avoid expensive calculations, and on the other hand, for the same instance, after choosing different steps, the state of the instance is different, and it is meaningless to pursue different outputs for different instances at this time.

[0051] For the sampled solution of each decoder obtained by the current model p θ (line 6), the greedy solution of each decoder is obtained on the best model p θ* (line 7). Lines 8 - 9 define the loss function of DRL. The parameters are optimized by the Adam optimizer (line 10). If the solution of the model p θ is better than the solution of the best model p θ* , then the parameters θ * are updated (lines 12 - 13).

[0052]

[0053]

[0054] Randomly generate 64,000 instances, whose input is an S×T bay structure, where S represents the number of stacks in the bay and T represents the height of the bay; the output is the rehandling operation sequence π = (π1, π2, …, π K ), where π i (i ∈ {1, 2, …, K}) represents the i-th rehandling operation.

[0055] Set the number of iterations and the number of iteration steps, and adopt Figure 2The reinforcement learning algorithm shown is used to train MDDAM. In the figure, epoch represents the current training round of the model, and step represents the number of iteration steps in a certain training round. Step++ and epoch++ respectively mean: step = step + 1 and epoch = epoch + 1.

[0056] S3 Input a bay structure, extract the container stack semantics through the feature enhancer to form a bay structure containing the inter-stack semantics.

[0057] Specifically, the feature enhancer (SFE) extracts the inter-stack semantics of the bay structure, which is expressed by a matrix of S×5. Each row of the matrix represents the semantic features of a certain stack: the percentage of stack height, the percentage of stack vacancy quantity, the stack priority, the priority of the target container, and the number of unordered stackings (US) in the stack. Where S and T respectively represent the number of stacks and the layer height limit in the bay structure.

[0058] The bay structure of S×T is concatenated with the inter-stack semantic features of S×5, and a bay structure of S×(T + 5) is output.

[0059] S4 Use the trained multi-decoder dynamic attention model to predict the container relocation operation sequence of the input bay.

[0060] Specifically, use the trained MDDAM to predict the container relocation operation sequence. For a newly input bay structure, MDDAM processes each container one by one, and predicts its container relocation operation π i for the containers that need to be relocated until all containers in the bay structure are processed, and MDDAM outputs the container relocation operation sequence π of this bay.

[0061] Specific method: Step 1: Input the bay structure for the container relocation task into the multi-decoder dynamic attention model (MDDAM).

[0062] Step 2: Determine the target container to be processed according to the priority order of the containers.

[0063] Specifically, the higher the priority of the container, the more it needs to be processed first.

[0064] Step 3: Judge whether there is an obstructive container above the target container. If not, jump to Step 4, otherwise jump to Step 5.

[0065] Specifically, the container without an obstructive container can be directly retrieved, and the meaning of retrieval is to delete the container from the bay structure.

[0066] Step 4: Delete the target container from the bay structure and jump to Step 2.

[0067] Specifically, after deletion, the number of containers in the bay structure will be reduced by one.

[0068] Step Five: Use MDDAM to predict and perform the unboxing operations on all the obstructive containers above the target container.

[0069] Specifically, after execution, the number of containers in the bay structure will not be reduced, but the distribution of the containers in the bay structure has changed.

[0070] Step Six: Determine whether the containers in the bay have been processed. If so, jump to Step Seven; otherwise, jump to Step Four.

[0071] Specifically, the completion of processing the containers in the bay structure means that the bay is empty, and the unboxing task for this bay is completed.

[0072] Step Seven: Output the unboxing operation sequence corresponding to this bay and end.

[0073] The above-disclosed is only a preferred embodiment of a multi-decoder dynamic attention model for container unboxing of the present invention. Of course, the scope of the rights of the present invention cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.

Claims

1. A multi-decoder dynamic attention model for container repositioning, characterized in that, It includes the following steps: Design a multi-decoder dynamic attention model with feature enhancement based on the Transformer architecture; Train the multi-decoder dynamic attention model based on the reinforcement learning algorithm with a greedy baseline; Input a bay structure, extract the container stack semantics through a feature enhancer to form a bay structure containing inter-stack semantics; Use the trained multi-decoder dynamic attention model to predict the sequence of container repositioning operations for the input bay; The design of the multi-decoder dynamic attention model with feature enhancement based on the Transformer architecture includes designing a feature enhancer, an encoder, and a decoder; The specific method for designing the feature enhancer: Define five original features representing inter-stack semantics based on the Transformer architecture, heuristic rules, and manual experience; Based on the input bay structure, calculate the five original features and represent them in matrix form to obtain an inter-stack semantic feature matrix; Concatenate the bay structure matrix and the inter-stack semantic feature matrix to form a bay structure matrix containing inter-stack semantics; Among them, the specific method for defining five original features representing inter-stack semantics based on the Transformer architecture, heuristic rules, and manual experience is: Based on heuristic rules and manual experience, define five original features representing inter-stack semantics; these five features are: the height percentage of the current stack; the percentage of empty spaces in the stack; the stack priority; the priority of the target container; the number of disorderly stacks in the stack; Among them, the height percentage of the current stack is 1 / T, 2 / T, …, 1; Among them, the specific method for calculating the five original features based on the input bay structure and representing them in matrix form to obtain an inter-stack semantic feature matrix is: According to the input bay structure, calculate five inter-stack semantic features and represent them in matrix form; the rows of the matrix have 5 components, and each component corresponds to an inter-stack semantic feature value; The specific method for designing the encoder: Input the output of the feature enhancer into a long short-term memory network to form the initial embedding required by the Transformer; Input the output of the long short-term memory network into the encoding end of the Transformer, and extract the global semantic features of the bay structure through the self-attention mechanism and residual connection; Among them, the specific content of training the multi-decoder dynamic attention model based on the reinforcement learning algorithm with a greedy baseline includes: Each of the multiple decoders outputs a different probability distribution and independently samples a solution to obtain separate reinforcement learning losses: where is the sequence length of the solution, i.e., the total number of container reshuffles; b(X) is the total number of container reshuffles for the deterministic greedy rollout of the baseline policy ; the parameters θ of the model are updated using the REINFORCE algorithm by gradient descent: ​ Among them represents the REINFORCE loss gradient is the gradient of the KL divergence between the decoder output probability distributions, k KL is the coefficient of the KL loss, k KL The negative sign in front is because we hope to maximize the diversity among the decoders; Among them, the specific content of inputting a bay structure, extracting the container stack semantics through a feature enhancer to form a bay structure containing inter-stack semantics includes: The feature enhancer extracts the inter-stack semantics of the bay structure, which is expressed by an S×5 matrix. Each row of the matrix represents the semantic features of a certain stack: the height percentage of the stack, the percentage of empty spaces in the stack, the stack priority, the priority of the target container, and the number of disorderly stacks in the stack, where S and T respectively represent the number of stacks and the height limit in the bay structure; Perform a concatenation operation on the S×T bay structure and the S×5 inter-stack semantic features, and output an S×(T + 5) bay structure; Among them, specifically including using the trained multi-decoder dynamic attention model to predict the sequence of operations for turning over boxes at the input bay location: Use the trained MDDAM to predict the sequence of container flipping operations; for a newly input bay structure, MDDAM processes each container one by one, and for the containers that need to be flipped, predicts their flipping operations π i , until all containers in the bay structure are processed, and MDDAM outputs the sequence of container flipping operations π for this bay.