Production line layout optimization method based on constraint-aware transformer and soft permutation matrix
Patent Information
- Application Number
- CN202511879559.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-12
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2045-12-12
AI Technical Summary
传统方法主要依赖数学规划与元启发式算法:数学规划方法虽能保证小规模最优性,但面临“组合爆炸”问题,难以应对大规模实际场景;元启发式算法(如遗传算法、模拟退火)通过智能搜索获得满意解,但优化过程依赖大量领域知识与参数调优,且作为“黑箱”运行,难以嵌入先验知识并提供可解释性洞察
本发明在输入处理阶段,将离散设备集合映射为高维连续嵌入特征,打破传统数学规划面临的离散空间组合爆炸困境,使后续梯度优化成为可能;约束感知Transformer编码器通过自注意力机制深度挖掘设备间的全局依赖关系,其双预测头架构同步完成布局质量评估与细粒度位置分布预测,有效克服全局优化目标与局部分配决策难以协同的难题;针对硬约束满足机制缺失的缺陷,可微Sinkhorn层将位置分布迭代转换为满足双随机条件的软置换矩阵,在概率意义上内生地建模设备与位置间的“一一对应”分配约束,避免传统方法“生成-验证”范式的低效试错;随后匈牙利算法将软分配转化为严格的离散硬布局方案,确保最终输出在物理空间上完全可行;在训练过程中,多任务损失函数通过约束损失的梯度反向传播,驱动模型主动学习并内化工艺约束模式,而非依赖人工调优或事后惩罚,从而显著提升在复杂多约束场景下的收敛稳定性与解的质量一致性。
Smart Images

Figure CN122047152B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of intelligent manufacturing and industrial engineering technology, and in particular to a production line layout optimization method based on constraint-aware Transformer and soft permutation matrix. Background Technology
[0002] Production line layout optimization is a core aspect of manufacturing system design. Its solution space grows factorially with the number of equipment, classifying it as an NP-hard problem. Traditional methods primarily rely on mathematical programming and metaheuristic algorithms. While mathematical programming can guarantee small-scale optimality, it faces the "combinatorial explosion" problem, making it difficult to handle large-scale real-world scenarios. Metaheuristic algorithms (such as genetic algorithms and simulated annealing) obtain satisfactory solutions through intelligent search, but the optimization process depends heavily on domain knowledge and parameter tuning, and operates as a "black box," making it difficult to embed prior knowledge and provide interpretable insights. More importantly, traditional methods generally employ a "generate-verify" iterative paradigm, penalizing constraint violations through cost functions rather than intrinsically and explicitly modeling and adhering to hard process constraints during solution generation. This results in low search efficiency and difficulty converging to the feasible region in complex constraint scenarios. In recent years, deep learning has shown its potential in combinatorial optimization, but its direct application to layout problems still faces three major challenges: the lack of differentiable connections between discrete permutations and continuous models, which makes end-to-end optimization impossible; the lack of hard-satisfaction mechanisms for complex multi-constraints (such as fixed equipment positions, prohibition of adjacency, etc.); and the difficulty in co-optimizing global layout quality and fine-grained position distribution. These defects seriously restrict the practical application effectiveness of production line layout optimization in smart manufacturing. Summary of the Invention
[0003] To address the aforementioned shortcomings, the present invention aims to propose a production line layout optimization method based on constraint-aware Transformer and soft permutation matrix. This method constructs an end-to-end differentiable learning framework, synchronously learning the global quality and fine-grained positional distribution of the layout through a constraint-aware Transformer model. It utilizes a differentiable Sinkhorn layer and the Hungarian algorithm to achieve a soft-hard mapping from continuous probabilities to discrete permutations. Furthermore, it combines a multi-task curriculum learning strategy to dynamically coordinate quality optimization and constraint satisfaction, thereby achieving efficient output of high-quality layout schemes while satisfying complex process hard constraints.
[0004] To achieve this objective, the present invention adopts the following technical solution: The production line layout optimization method based on constraint-aware Transformer and soft permutation matrix includes the following steps: S1: Obtain the equipment set information of the production line and the preset process constraint set, and convert the equipment set information into equipment embedding features; S2: Input the device embedding features into the constraint-aware Transformer model, extract global dependency features through the encoder of the constraint-aware Transformer model, and generate the position distribution information of each device in all layout positions based on the global dependency features; S3: Based on the location distribution information, a soft permutation matrix is generated through differentiable Sinkhorn iteration. The soft permutation matrix satisfies the double random condition and represents the allocation relationship between the device and the layout location in probabilistic form. S4: Based on the soft permutation matrix, the discrete allocation relationship between devices and layout positions is determined by the Hungarian algorithm, and a hard layout scheme is generated based on the discrete allocation relationship; S5: During the training of the constraint-aware Transformer model, the degree of violation of each process constraint is calculated based on the soft permutation matrix and a constraint loss is generated. The constraint loss and the layout quality prediction loss are combined to form a multi-task loss function to drive the constraint-aware Transformer model to simultaneously optimize layout quality and constraint satisfaction.
[0005] Preferably, in step S2, extracting global dependency features through the encoder of the constraint-aware Transformer model includes: A combined embedding representation is generated for each device and its corresponding layout location. Satisfying the relation: ; in, This represents the device embedding vector obtained by mapping device semantic information. This represents the position embedding vector that characterizes the absolute position of the device in the layout sequence. The combined embedded representation is input to a Transformer encoder, where a multi-head self-attention module calculates attention weights by aggregating interaction information across all devices. Satisfying the relation: ; in, , , The representations are respectively represented by combined embeddings. The query matrix, key matrix, and value matrix obtained through linear transformation This represents the dimension of the query and key vectors in the single-head attention mechanism; The Transformer encoder performs a nonlinear transformation on the output of the multi-head self-attention module through a feedforward neural network, and combines layer normalization and Dropout operations to output encoded global dependency features.
[0006] Preferably, in step S2, generating the location distribution information of each device at all layout locations based on global dependency features includes: The global dependency features are processed in parallel using a dual-predictor-head architecture of a constraint-aware Transformer model, which includes a quality prediction head and a location prediction head. The quality prediction head includes performing a global average pooling operation on the global dependent features to aggregate and obtain a feature vector representing the overall layout information. The feature vector representing the overall layout information is then mapped to a comprehensive quality score within a preset numerical range through a two-layer MLP. The location prediction head maps the global dependency features into a three-dimensional tensor through a linear transformation layer, which serves as a matrix representation of the location distribution information. The dimensions of the three-dimensional tensor correspond to the batch size, layout sequence length, and number of devices, respectively, where each element represents the original distribution score of the corresponding device assigned to a specific layout position.
[0007] Preferably, step S3 includes: According to the preset temperature parameters The original distribution scores in the location distribution information are scaled to obtain a scaled logarithmic probability matrix S, which satisfies the following relationship: ; in, This represents the original distribution score in the location distribution information. A temperature parameter used to control the sharpness of a probability distribution; For the logarithmic probability matrix Perform alternating row and column normalization to make... The soft permutation matrix is obtained by satisfying the double random condition and through exponential operation. Satisfying the relation: ; in, This represents the final generated soft permutation matrix. This represents the logarithmic probability matrix that satisfies the double random condition after alternating row and column normalization. Indicates exponentiation; The soft permutation matrix Each element in the matrix represents the probability value of the corresponding device being assigned to a specific layout position, and the soft permutation matrix... The sum of the elements in each row and the sum of the elements in each column are both 1.
[0008] Preferably, step S4 includes: For each sample, the soft permutation matrix Construct the cost matrix The cost matrix Satisfying the relation: ; in, Representing the cost matrix The Middle Line number Column elements, Represents the first in the soft permutation matrix The device was assigned to the first The probability value of each layout position; For the cost matrix The Hungarian algorithm is applied to find the perfect match with the minimum total cost, resulting in the device index array. and position index array ,in Indicates the first Device index in each matching pair. Indicates the first The location index to which the device is assigned in each matching pair. Indicates the index of the matching pair; According to the device index array and position index array Generate hard permutation matrix The hard permutation matrix The generation satisfies the following relation: ; In this case, each element of the hard permutation matrix is either 0 or 1, and each row and each column has exactly one 1; According to the hard permutation matrix The hard layout scheme is generated, which is a determined arrangement sequence of devices in the layout position.
[0009] Preferably, in step S1, converting the device set information into device embedded features includes: The device set information is formalized as a finite set including N devices. ; The preset set of process constraints is formalized into a set of binary constraints. Each constraint Indicates equipment Must be assigned to the device Before; Among them, feasible layout schemes include the device set. permutation sequence Permutation sequence Satisfy: Total number of layout positions Equal to the total number of devices and sequence Each device in It appears exactly once.
[0010] Preferably, step S5 includes: For the set of process constraints Each constraint in computing devices and equipment Desired location, equipment Expected position and equipment Expected position Satisfying the relation: , ; in, Indicates the length of the layout sequence. Represents the soft permutation matrix medium equipment Assigned to position The probability value, Represents the soft permutation matrix medium equipment Assigned to position The probability value, Indicates the position index; Based on desired location and Calculate the set of process constraints Each constraint in The degree of violation, and generate the constraint loss. The constraint loss Satisfying the relation: ; in, Represents the set of process constraints The total number of constraints in the middle, Representing constraints Preset weights, This represents the preset tolerance range parameter; The constraint loss With layout quality prediction loss Perform weighted summation to construct the multi-task loss function. Among them, the layout quality prediction loss Satisfying the relation: ; in, Indicates the size of the current training batch. The constraint-aware Transformer is the first... The overall quality score predicted for each sample. Indicates the first Each sample corresponds to a preset target quality score.
[0011] One of the above technical solutions has the following advantages or beneficial effects: In the input processing stage, this invention maps a set of discrete devices into high-dimensional continuous embedded features, breaking the dilemma of discrete space combinatorial explosion faced by traditional mathematical programming and making subsequent gradient optimization possible. The constraint-aware Transformer encoder deeply mines the global dependencies between devices through a self-attention mechanism. Its dual prediction head architecture simultaneously completes layout quality assessment and fine-grained position distribution prediction, effectively overcoming the difficulty of coordinating global optimization objectives and local allocation decisions. To address the lack of hard constraint satisfaction mechanisms, the differentiable Sinkhorn layer iteratively transforms the position distribution into a soft permutation matrix that satisfies double random conditions, endogenously modeling the "one-to-one" allocation constraints between devices and positions in a probabilistic sense, avoiding the inefficient trial and error of the traditional "generate-verify" paradigm. Subsequently, the Hungarian algorithm transforms the soft allocation into a strict discrete hard layout scheme, ensuring that the final output is completely feasible in physical space. During training, the multi-task loss function drives the model to actively learn and internalize process constraint patterns through backpropagation of the gradient of constraint loss, rather than relying on manual tuning or ex-post penalties, thereby significantly improving the convergence stability and solution quality consistency in complex multi-constraint scenarios. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0013] Figure 1 This is a flowchart of a production line layout optimization method based on constraint-aware Transformer and soft permutation matrix provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the architecture for production line layout optimization based on constraint-aware Transformer and soft permutation matrix provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the Transformer encoder architecture provided in an embodiment of the present invention; Figure 4This is a flowchart of the Sinkhorn layer provided in an embodiment of the present invention; Figure 5 This is a schematic diagram comparing the layout sequences of different models provided in the embodiments of the present invention. Detailed Implementation
[0014] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0015] In this invention, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0016] Production line layout optimization methods based on constraint-aware Transformer and soft permutation matrix, such as Figure 1 and Figure 2 As shown, a preferred embodiment of the present invention includes the following steps: S1: Obtain the equipment set information of the production line and the preset process constraint set, and convert the equipment set information into equipment embedding features; It should be noted that the equipment set information refers to the structured data of all equipment to be laid out in the production line, including at least equipment identifiers and equipment type attributes. In this embodiment, the equipment set information specifically refers to a set of 12 types of equipment, including 1 SP loading machine, 4 welding machines, 6 small material loading machines, and 1 unloading station. The process constraint set refers to the sequence constraints between equipment predefined by process engineers according to the production process logic. Each constraint is represented by a tuple (a, b), meaning that equipment a must be placed before equipment b. Equipment embedding features are a technique for converting discrete equipment identifiers into high-dimensional continuous vector representations. A learnable embedding layer (nn.Embedding) is used to achieve the mapping. During training, the embedding layer automatically adjusts its parameters through backpropagation, making semantically similar equipment closer in the vector space. The role of embedding technology is to overcome the sparsity limitations of traditional one-hot encoding, capturing the functional semantics and process attributes of equipment in a compact continuous space, providing a dense and expressive input foundation for subsequent neural network processing.
[0017] Understandably, traditional methods that directly manipulate device indexes or manually design features struggle to automatically learn deep relationships between devices. By introducing a learnable device embedding mechanism, the system can autonomously discover the role differences and collaborative relationships of different devices in the process flow, such as automatically identifying the functional distinction between welding and feeding equipment. This transformation process not only provides the subsequent Transformer encoder with continuous vectors that meet its input requirements, but more importantly, it pre-injects optimizable structural information into the embedding space, allowing the same device to have different semantic expressions in different positions, laying the foundation for position-sensitive layout optimization. The embedding operation smoothly maps the entire optimization problem from a discrete combinatorial space to a continuous parameter space, making gradient-based end-to-end optimization possible and fundamentally solving the problem of low search efficiency in discrete spaces in traditional methods.
[0018] S2: Input the device embedding features into the constraint-aware Transformer model, extract global dependency features through the encoder of the constraint-aware Transformer model, and generate the position distribution information of each device in all layout positions based on the global dependency features; It should be noted that the constraint-aware Transformer model refers to a neural network model based on a pure encoder architecture. Its core consists of a multi-head self-attention mechanism and a feedforward neural network, specifically designed to handle complex process constraints and spatial relationships between devices. The encoder is the fundamental component of the Transformer model, composed of multiple identical stacked layers. Each layer contains a multi-head self-attention module and a feedforward network, responsible for feature extraction and context encoding of the input sequence. Global dependency features refer to the feature representations in the encoder output that contain long-distance interaction relationships between devices, capable of capturing the strength of process constraints and spatial synergies between any two devices. Location distribution information refers to the non-normalized score generated by the model for each device at all candidate locations (…). Its dimensions are This represents the degree of preference each device has for each layout location; a higher value indicates that the model considers the device to be more suitable for the corresponding location.
[0019] Understandably, step S2 leverages the powerful sequence modeling capabilities of the Transformer encoder to automatically mine implicit process associations and explicit constraint relationships between devices from their embedded features. Traditional methods rely on manually designed adjacency matrices or distance metrics, making it difficult to capture high-order nonlinear dependencies between devices. The Transformer, through its self-attention mechanism, can dynamically calculate the correlation weights between any pair of devices, for example, automatically recognizing the process logic that welding equipment must be located after the loading equipment. The multi-head attention mechanism further extends this process to multiple subspaces, capturing diverse dependency patterns in parallel across different representation spaces, such as spatial proximity, functional similarity, or process sequence. This feature extraction process transforms discrete device relationship modeling into continuous feature similarity calculation, allowing constraint information to be incorporated into model parameters in a differentiable manner, providing a feature foundation rich in constraint semantics for subsequently generating positional distributions that conform to process logic. Through end-to-end training, the encoder learns to push device pairs that violate constraints further away in the feature space and bring device pairs that satisfy constraints closer together, thereby internalizing constraints at the feature level.
[0020] S3: Based on the location distribution information, a soft permutation matrix is generated through differentiable Sinkhorn iteration. The soft permutation matrix satisfies the double random condition and represents the allocation relationship between the device and the layout location in probabilistic form. It should be noted that differentiable Sinkhorn iteration refers to a matrix normalization algorithm based on optimal transport theory. It performs alternating normalization operations on the rows and columns of the matrix to satisfy the double randomness condition, and all operations are differentiable, supporting gradient backpropagation. The double randomness condition means that the sum of the elements in each row and the sum of the elements in each column of the matrix are 1. In layout optimization, this corresponds to the hard constraint that each device must be assigned to a location and each location must accommodate one device. The flowchart of the Sinkhorn layer is as follows... Figure 4 As shown. A soft permutation matrix is a probability matrix generated through Sinkhorn iteration, whose elements take values between 0 and 1, representing the probability of allocation between a device and a location rather than a deterministic allocation.
[0021] Understandably, step S3 involves converting the non-normalized score output by the position prediction head into a probability matrix that conforms to the allocation constraints, while maintaining the differentiability of the entire computation process, allowing the constraint information to be backpropagated via gradients. Traditional methods directly perform combinatorial search in the discrete space, resulting in a non-differentiable optimization process. The Sinkhorn iteration relaxes discrete constraints into continuous probability constraints, transforming the non-differentiable combinatorial problem into a differentiable matrix optimization problem while preserving the constraint semantics. This relaxation strategy allows the model to gradually adjust the allocation probabilities through gradient descent during training, causing high-probability allocation schemes to gradually converge towards satisfying the process constraints. The double random condition strictly guarantees the integrity of the allocation in a probabilistic sense, avoiding the equipment conflict or position vacancy problems common in traditional methods. The introduction of the temperature parameter implements the annealing strategy; the initial smooth distribution supports extensive exploration, while the later sharp distribution promotes precise convergence. This progressive optimization mechanism significantly improves the training stability and final solution quality under complex constraint scenarios.
[0022] S4: Based on the soft permutation matrix, the discrete allocation relationship between devices and layout positions is determined by the Hungarian algorithm, and a hard layout scheme is generated based on the discrete allocation relationship; It should be noted that the Hungarian algorithm is a combinatorial optimization algorithm for solving the assignment problem. It can find the perfect match that minimizes the total cost in polynomial time and is suitable for transforming soft assignment probabilities into hard assignment decisions. A hard placement scheme refers to a discrete permutation sequence in which each device is deterministically assigned to a unique position, where there is no probability or ambiguity, satisfying the physical requirements of actual deployment. The cost matrix is the input to the Hungarian algorithm, and its elements represent the cost of assigning a specific device to a specific position. In this embodiment, negative probability is used as the cost to maximize the total assignment probability.
[0023] Understandably, step S4 involves converting the probabilistic soft assignments used in the training phase into deterministic hard assignments usable in the inference phase, ensuring that the final output layout scheme can be directly used for actual production line deployment. Although the soft permutation matrix satisfies the assignment constraints in a probabilistic sense, practical engineering applications require explicit correspondences between equipment positions. The Hungarian algorithm, as a classic combinatorial optimization tool, can quickly solve for the optimal hard assignment based on the soft assignment matrix, with a computational complexity of O(n log n). For production lines with 12 to 20 devices, the solution can be completed in milliseconds. By constructing a negative probability cost matrix, the algorithm's objective shifts from minimizing the cost to maximizing the total allocation probability, ensuring that the hard allocation scheme aligns with the model's soft allocation tendency. This hard transformation mechanism is executed only during the inference phase and does not participate in training backpropagation, thus separating differentiable optimization from discrete decision-making. This maintains the numerical stability of the training process while guaranteeing the physical feasibility of the final solution. The generation of the hard layout scheme marks the final mapping from the continuous optimization space to the discrete solution space, serving as a crucial bridge to resolving the contradiction between differentiable training and discrete inference.
[0024] S5: During the training of the constraint-aware Transformer model, the degree of violation of each process constraint is calculated based on the soft permutation matrix and a constraint loss is generated. The constraint loss and the layout quality prediction loss are combined to form a multi-task loss function to drive the constraint-aware Transformer model to simultaneously optimize layout quality and constraint satisfaction.
[0025] It should be noted that constraint loss is a differentiable function that measures the degree to which the layout scheme predicted by the model violates the preset process constraints. Its calculation is based on a soft permutation matrix rather than hard assignment, allowing gradients to be effectively backpropagated. Layout quality prediction loss is a function that evaluates the difference between the model's predicted quality score and the true quality label, typically measured using mean squared error. The multi-task loss function is a total optimization objective formed by linearly combining the losses of multiple sub-tasks according to their weights. This allows the model to balance multiple objectives during a single training process. The curriculum learning strategy refers to the technique of dynamically adjusting the weights of different task losses during training, initially focusing on quality learning and gradually increasing constraint guidance in later stages to achieve progressive optimization from simple to complex. The desired position is the weighted average of the device's position in the layout, with the weights being the assignment probabilities in the soft permutation matrix, transforming discrete-order constraints into continuously optimizable objectives.
[0026] Understandably, the core purpose of step S5 is to establish a unified optimization framework, enabling the model to simultaneously learn to maximize layout quality and strictly adhere to process constraints, thus resolving the dilemma of separating quality and constraint optimization in traditional methods. Traditional methods typically optimize quality first and then verify constraints, leading to a large amount of ineffective searching. By constructing a multi-task loss function, quality optimization and constraint satisfaction share the same set of model parameters, and gradient signals simultaneously guide the model towards both objectives. The innovation of the constraint loss lies in using a soft permutation matrix to calculate the desired position, thus integrating discrete order relation constraints (such as equipment...) Must be in the device Previously, the constraint was transformed into a difference constraint of the desired position in a continuous space, thus making the originally non-differentiable discrete constraint differentiable. The constraint loss weights are dynamically adjusted through a course learning strategy. Smaller in the early stages of training The model should prioritize learning the basic patterns of high-quality layouts to avoid being trapped in local optima due to premature strict constraints; the value should be gradually increased in later stages. The value strengthens constraint satisfaction, ensuring that the final model output is both high-quality and compliant. The above joint optimization mechanism realizes the endogenous modeling of constraints, enabling the model to internalize the constraint logic during the representation learning stage, rather than punishing it afterward, which significantly improves search efficiency and constraint satisfaction rate.
[0027] like Figure 5This demonstrates the significant value of the constraint-aware Transformer model in engineering applications. In terms of quality assurance, the output solutions from the constraint-aware Transformer model consistently achieve a high quality score of 99.9+, guaranteeing 100% compliance with process constraints and effectively reducing engineering risks. Regarding efficiency improvement, the model can generate ready-to-use layout solutions, shortening the traditional design cycle from weeks to less than a day. The stable output characteristics of the constraint-aware Transformer model also provide a reliable technical foundation for automated layout optimization in intelligent manufacturing systems.
[0028] Preferably, in step S2, extracting global dependency features through the encoder of the constraint-aware Transformer model includes: A combined embedding representation is generated for each device and its corresponding layout location. Satisfying the relation: ; in, This represents the device embedding vector obtained by mapping device semantic information. This represents the position embedding vector that characterizes the absolute position of the device in the layout sequence. The combined embedded representation is input to a Transformer encoder, where a multi-head self-attention module calculates attention weights by aggregating interaction information across all devices. Satisfying the relation: ; in, , , The representations are respectively represented by combined embeddings. The query matrix, key matrix, and value matrix obtained through linear transformation This represents the dimension of the query and key vectors in the single-head attention mechanism; The Transformer encoder performs a nonlinear transformation on the output of the multi-head self-attention module through a feedforward neural network, and combines layer normalization and Dropout operations to output encoded global dependency features.
[0029] It's important to note that the combined embedding representation refers to the joint feature representation formed by fusing the device embedding vector and the location embedding vector element-wise. Its dimensionality remains consistent with that of a single embedding vector. This fusion mechanism ensures the model can distinguish semantic differences of the same device in different locations. The multi-head self-attention module is a core component of the Transformer encoder. It captures different aspects of the relationships between devices by computing multiple attention heads in parallel. Each attention head independently computes the query, key, and value matrices and generates attention weights. The query matrix Q, key matrix K, and value matrix V are three feature matrices obtained by performing three independent linear transformations on the combined embedding representation. Q is used to initiate the query, K is used for matching, and V is used for information aggregation. All three participate in the calculation of the attention weights. Single-head attention dimension. This represents the feature dimension of the query vector and key vector in each attention head, and its value is the total embedding dimension divided by the number of attention heads. A feedforward neural network is a multilayer perceptron structure consisting of two linear layers and activation functions. It performs a non-linear transformation on the attention output to enhance the model's expressive power. Layer normalization is an operation that standardizes the mean and variance of the neural network layer outputs, which can stabilize the training process and accelerate convergence. Dropout is a regularization technique that randomly masks some neurons during training to prevent overfitting.
[0030] Understandably, traditional methods, employing fixed handcrafted features or shallow networks, struggle to capture high-order interactions and long-range dependencies between devices. Combinatorial embedding mechanisms, by explicitly fusing positional information, enable the model to perceive the different semantics of devices at different positions, such as the difference in process role a SP feeder plays depending on its position at the beginning or middle of a sequence. Multi-head self-attention mechanisms, by decomposing the feature space into multiple subspaces, allow the model to capture diverse dependency patterns in parallel across different representation spaces, such as focusing on process sequence constraints in one head and on device functional similarity in another. Scaling dot product attention... The correlation scores between devices are calculated, and after softmax normalization, attention weights are obtained. This weight matrix reflects the strength of the influence between devices; device pairs that violate constraints receive lower weights. A feedforward neural network performs a non-linear mapping on the attention-weighted features to uncover deeper combined features. Layer normalization and the introduction of Dropout alleviate the gradient vanishing and overfitting problems in deep network training, ensuring that the encoder can stably learn complex constraint patterns. This architecture transforms discrete constraint relationships into continuous feature similarity calculations, allowing constraint information to be embedded into model parameters in a differentiable manner, providing a feature foundation rich in constraint semantics for subsequent location prediction.
[0031] Preferably, in step S2, generating the location distribution information of each device at all layout locations based on global dependency features includes: The global dependency features are processed in parallel using a dual-predictor-head architecture of a constraint-aware Transformer model, which includes a quality prediction head and a location prediction head. The quality prediction head includes performing a global average pooling operation on the global dependent features to aggregate and obtain a feature vector representing the overall layout information. The feature vector representing the overall layout information is then mapped to a comprehensive quality score within a preset numerical range through a two-layer MLP. The location prediction head maps the global dependency features into a three-dimensional tensor through a linear transformation layer, which serves as a matrix representation of the location distribution information. The dimensions of the three-dimensional tensor correspond to the batch size, layout sequence length, and number of devices, respectively, where each element represents the original distribution score of the corresponding device assigned to a specific layout position.
[0032] It should be noted that the dual-prediction-head architecture refers to a neural network structure design that starts from the globally dependent features extracted from the same encoder and processes them in parallel through two functionally independent network branches, performing quality assessment and position prediction tasks respectively. The quality prediction head is a sub-network composed of a global average pooling layer and two multilayer perceptron (MLP) layers cascaded together, responsible for compressing sequence features into a scalar quality score. The global average pooling operation involves calculating the arithmetic mean of all position feature vectors output by the Transformer encoder along the sequence dimension, converting the variable-length sequence representation into a fixed-length global feature vector. The Transformer encoder architecture is as follows... Figure 3 As shown. A two-layer MLP refers to a feedforward network structure containing two linear transformation layers with a non-linear activation function inserted in between. The first layer maps the pooled feature vectors to a higher-dimensional space (e.g., 128→512), and the second layer compresses them to the target output dimension (e.g., 512→1). The overall quality score refers to the scalar value output after the MLP mapping, usually normalized to the [0,100] interval, used to quantitatively evaluate the overall quality of the layout scheme. The position prediction head refers to a lightweight network branch consisting of a single linear transformation layer, whose function is to independently map each device feature vector output by the encoder to the raw score at all positions. A linear transformation layer refers to a neural network layer that performs matrix multiplication between the weight matrix and the input features and adds a bias term, without containing a non-linear activation function. The raw distribution score refers to the value directly output by the linear transformation layer without softmax normalization ( The value of reflects the strength of the device's preference for a specific location; a positive value indicates preference, a negative value indicates rejection, and the absolute value reflects the confidence level.
[0033] Understandably, traditional methods often treat quality assessment and location allocation as two independent stages, leading to information gaps and inconsistencies in optimization objectives. The dual-predictor architecture generates two types of outputs simultaneously from the same encoder features, enabling quality assessment to feed back into location prediction, and location distribution information to provide detailed support for quality judgment, achieving synergistic optimization of the two tasks. The quality prediction head aggregates interaction information from all devices through global average pooling to gain a comprehensive understanding of the overall layout, such as automatically determining whether material flow is smooth and whether equipment utilization is balanced. Pooling eliminates differences in location dimensions, generating a global representation independent of sequence length, making the quality score scale-invariant. The two-layer MLP structure mines the complex mapping relationship between layout quality and equipment configuration through nonlinear transformation, compressing high-dimensional features into a single score for rapid solution selection. The location prediction head retains the complete sequence dimension of the encoder output, independently generating a location preference distribution for each device, enabling the model to model the competition and constraint satisfaction between devices at a fine-grained level. The linear transformation layer directly outputs... Instead of probabilistic approaches, this avoids premature gradient vanishing caused by normalization, preserving the numerical optimization space. This dual-head design allows the model to simultaneously obtain macroscopic quality judgments and microscopic allocation details in a single forward propagation, supporting downstream multi-objective joint optimization.
[0034] Preferably, step S3 includes: According to the preset temperature parameters The original distribution scores in the location distribution information are scaled to obtain a scaled logarithmic probability matrix S, which satisfies the following relationship: ; in, This represents the original distribution score in the location distribution information. A temperature parameter used to control the sharpness of a probability distribution; For the logarithmic probability matrix Perform alternating row and column normalization to make... The soft permutation matrix is obtained by satisfying the double random condition and through exponential operation. Satisfying the relation: ; in, This represents the final generated soft permutation matrix. This represents the logarithmic probability matrix that satisfies the double random condition after alternating row and column normalization. Indicates exponentiation; The soft permutation matrix Each element in the matrix represents the probability value of the corresponding device being assigned to a specific layout position, and the soft permutation matrix... The sum of the elements in each row and the sum of the elements in each column are both 1.
[0035] It should be noted that the temperature parameter It is a positive real-valued hyperparameter used to control the smoothness or sharpness of the probability distribution. Its value directly affects the numerical range of the logarithmic probability matrix. A smaller value indicates a higher probability distribution. The value makes the distribution approach the one-hot form, and a larger value... The value maintains a smooth distribution; this parameter is dynamically adjusted during training using an annealing strategy to balance exploration and exploitation. The logarithmic probability matrix is the intermediate matrix after temperature scaling, its elements being unnormalized logarithmic probability values, serving as input for the Sinkhorn iteration. Row normalization involves normalizing each row of the matrix so that the sum of the exponents of all elements in that row equals 1, corresponding to the constraint that the sum of the probabilities of each device across all positions is 1. Column normalization involves normalizing each column of the matrix so that the sum of the exponents of all elements in that column equals 1, corresponding to the constraint that the sum of the probabilities of each position being occupied by all devices is 1. The double randomness condition is the mathematical property that the sum of the elements in each row and the sum of the elements in each column of the matrix are both 1, corresponding to the one-to-one correspondence constraint between devices and positions in layout optimization. Exponentiation refers to raising the elements of the logarithmic probability matrix to the power of the natural constant e, converting the normalized logarithmic probabilities into actual probability values.
[0036] Understandably, traditional combinatorial optimization searches directly in the discrete space, resulting in a non-differentiable objective function and making it impossible to apply efficient optimization algorithms such as gradient descent. Sinkhorn iteration, by introducing temperature scaling and alternating normalization mechanisms, relaxes discrete permutation constraints into continuous double-random constraints, achieving differentiability while preserving constraint semantics. The temperature scaling operation is achieved by dividing by the parameter... To control the sharpness of the distribution, a larger distribution is needed in the early stages of training. The value smooths the distribution, allowing the model to explore different allocation possibilities extensively and avoid getting trapped in local optima too early; smaller values in the later stages... The sharp distribution of values forces the model to focus on high-confidence assignments, improving the quality of the solution. Alternating row and column normalization constitutes a fixed-point iterative process. By repeatedly strengthening row and column constraints, the matrix gradually converges to a doubly stochastic state. This process consists entirely of differentiable basic operations, and the gradient can be smoothly backpropagated to the location prediction head. Exponential operations transform the logarithmic domain calculation results back to the probability domain, obtaining directly interpretable device-location assignment probabilities. This differentiable constraint modeling mechanism enables hard process constraints to be intrinsically satisfied through numerical optimization, rather than through ex-post correction using traditional penalty functions, significantly improving training efficiency and convergence stability in complex constraint scenarios.
[0037] Preferably, step S4 includes: For each sample, the soft permutation matrix Construct the cost matrix The cost matrix Satisfying the relation: ; in, Representing the cost matrix The Middle Line number Column elements, Represents the first in the soft permutation matrix The device was assigned to the first The probability value of each layout position; For the cost matrix The Hungarian algorithm is applied to find the perfect match with the minimum total cost, resulting in the device index array. and position index array ,in Indicates the first Device index in each matching pair. Indicates the first The location index to which the device is assigned in each matching pair. Indicates the index of the matching pair; According to the device index array and position index array Generate hard permutation matrix The hard permutation matrix The generation satisfies the following relation: ; In this case, each element of the hard permutation matrix is either 0 or 1, and each row and each column has exactly one 1; According to the hard permutation matrix The hard layout scheme is generated, which is a determined arrangement sequence of devices in the layout position.
[0038] It should be noted that the cost matrix is the input matrix of the Hungarian algorithm, where each element represents the cost of assigning a specific device to a specific location. This embodiment uses negative probability as the cost, transforming the original problem's objective of maximizing the assignment probability into minimizing the total cost. The Hungarian algorithm is a combinatorial optimization algorithm that solves the assignment problem in polynomial time. It achieves optimal matching by systematically finding augmenting paths. This embodiment calls... Function implementation, which is based on efficient Algorithm. A perfect match is a set of edges in a bipartite graph where every vertex is connected to exactly one matching edge. In the layout problem, this corresponds to assigning exactly one device to exactly one location, ensuring the completeness and validity of the allocation scheme. Device index array. This is one of the matching results returned by the Hungarian algorithm. It stores the device identifiers that participated in the matching in order, and the array length is the total number of devices. Position index array It is another result array returned by the algorithm, storing and The optimal location identifier corresponding to the device in the array. and Constituting the first A hard permutation matrix is a binary matrix with only 0 or 1 elements, where each row and column has exactly one 1. It strictly represents the one-to-one correspondence between devices and locations and is the final mathematical representation of the layout scheme. The determined permutation sequence is an ordered arrangement of devices directly mapped from the hard permutation matrix. The device allocation at each location is unique and unambiguous, and can be directly used for physical production line deployment.
[0039] Understandably, while soft permutation matrices satisfy allocation constraints probabilistically, their elements are continuous probability values, making them unsuitable for directly guiding physical equipment installation. The Hungarian algorithm, by solving for optimal matching in a bipartite graph, quickly calculates a unique optimal hard allocation based on the soft allocation matrix. Its polynomial time complexity ensures efficient conversion even in large-scale production lines. A negative probability strategy is employed when constructing the cost matrix, cleverly transforming the probability maximization problem into a cost minimization problem, aligning the algorithm's objective with the model's learning goal and ensuring maximum consistency between the hard allocation scheme and the soft allocation tendency. The binary nature of the hard permutation matrix completely eliminates allocation uncertainty; each row's unique 1 value explicitly specifies the unique location of the equipment, and each column's unique 1 value explicitly specifies the unique equipment at that location, avoiding illegal schemes such as equipment conflicts or empty locations. This hard conversion mechanism is executed only during the inference phase and does not participate in training backpropagation, thus separating the differentiable optimization process from the discrete decision-making process. This maintains the numerical stability of the training process while ensuring the physical feasibility of the final scheme. The generated deterministic permutation sequence, as the final output, can be directly used to guide production line construction without additional verification of allocation legality, achieving seamless integration from model prediction to engineering application.
[0040] Preferably, in step S1, converting the device set information into device embedded features includes: The device set information is formalized as a finite set including N devices. ; The preset set of process constraints is formalized into a set of binary constraints. Each constraint Indicates equipment Must be assigned to the device Before; Among them, feasible layout schemes include the device set. permutation sequence Permutation sequence Satisfy: Total number of layout positions Equal to the total number of devices and sequence Each device in It appears exactly once.
[0041] It should be noted that finite sets It refers to the mathematical set consisting of all the equipment to be laid out in the production line, whose elements... Representing specific equipment instances, the finiteness of the set is determined by the actual number of equipment on the production line. This formalization determines that discrete device objects are transformed into computable mathematical symbols. (Binary constraint set) It is an abstract representation of process constraints, each constraint To orderly The formal encoding of the order of devices, which mandates that devices in the generated arrangement... The index position must be smaller than the device. The index position is used to transform the technological logic into verifiable mathematical conditions. Permutation sequence refers to a collection of equipment A permutation of all elements is mathematically a bijective function. Each position in the sequence This corresponds to a specific slot in the layout space. The constraint that a device appears exactly once is a fundamental property of permutations, ensuring that each device is assigned exactly one slot and each slot accommodates exactly one device, avoiding illegal schemes such as duplicate device deployment or empty slots. This property is mathematically equivalent to the row and column constraints of a double random matrix.
[0042] Understandably, traditional methods typically describe constraints using natural language or heuristic rules, making them difficult for computers to automatically parse and verify. By abstracting the device set into a finite set... The system can clearly define the boundaries and scale of optimization objects, providing a basis for vocabulary size and index mapping for the embedding layer. Binary constraint set. Formalizing the process transforms complex process sequence relationships (such as the SP feeder must precede the welding equipment) into concise ordered pairs, converting constraint satisfaction verification into a comparison of position indexes, significantly reducing computational complexity. (Permutation sequence) The definition strictly limits the feasible solution space to the set of all permutations of N devices, although its size is enormous ( (There are multiple possibilities), but the mathematical structure is clear, providing a theoretical basis for the double random constraints of the subsequent Sinkhorn layer. The property that each occurrence occurs exactly once ensures a one-to-one correspondence between the solution space and the legality of the physical layout, avoiding the waste of time in generating invalid solutions. This formalization process transforms engineering experience into a computable mathematical structure, enabling constraint information to be explicitly injected into the model in symbolic form, providing a formal guarantee for constraint-aware training, while ensuring that the final hard layout solution strictly satisfies physical feasibility.
[0043] Preferably, step S5 includes: For the set of process constraints Each constraint in computing devices and equipment Desired location, equipment Expected position and equipment Expected position Satisfying the relation: , ; in, Indicates the length of the layout sequence. Represents the soft permutation matrix medium equipment Assigned to position The probability value, Represents the soft permutation matrix medium equipment Assigned to position The probability value, Indicates the position index; Based on desired location and Calculate the set of process constraints Each constraint in The degree of violation, and generate the constraint loss. The constraint loss Satisfying the relation: ; in, Represents the set of process constraints The total number of constraints in the middle, Representing constraints Preset weights, This represents the preset tolerance range parameter; The constraint loss With layout quality prediction loss Perform weighted summation to construct the multi-task loss function. Among them, the layout quality prediction loss Satisfying the relation: ; in, Indicates the size of the current training batch. The constraint-aware Transformer is the first... The overall quality score predicted for each sample. Indicates the first Each sample corresponds to a preset target quality score.
[0044] It should be noted that the desired location refers to equipment The weighted average of positions in the layout sequence, with weights determined by the soft permutation matrix. medium equipment The probability of assigning each position determines the computation, transforming the discrete order relation constraint into a differentiable objective in a continuous space. (Constraint loss) It is a differentiable function that measures the degree to which a layout scheme violates process constraints, expressed as the expected position difference and the tolerance range. In comparison, the ReLU style The activation function transforms constraint violations into non-negative loss values. Multi-task loss function. The overall optimization objective is a linear combination of constraint loss and quality loss with dynamic weights, which drives the model to learn both quality and constraint patterns simultaneously through gradient backpropagation. (Layout quality prediction loss) The mean squared error (MSE) calculation model is used to predict the difference between the quality score and the expert-annotated target score. The squared term penalizes large errors, guiding the model to accurately predict layout quality. Weights This is a differentiated importance coefficient set for different constraints. Key constraints can be assigned higher weights, while secondary constraints have lower weights, thus distinguishing the priority of constraints. Tolerance interval parameter. It is a small positive number used to establish a buffer band around the expected position difference, avoiding excessive penalty for small position deviations and improving the model's robustness to noise.
[0045] Understandably, step S5 involves constructing a unified joint optimization framework that enables the model to simultaneously learn to maximize layout quality and strictly adhere to process constraints, thus addressing the inefficiency caused by the separation of quality and constraint optimization in traditional methods. By using probabilistic weighting with soft permutation matrices, discrete device sequence constraints are transformed into continuous comparisons of desired positions, making the originally non-differentiable order constraints differentiable, thereby supporting gradient backpropagation. Regarding constraints... If the equipment The expected position is greater than the device If the desired position is not found, the constraint is violated. The loss function generates a gradient signal that propagates back to the position prediction head, guiding the model to adjust the probability allocation. The expected position is shifted forward. The introduction of weights allows the model to distinguish the strictness of different constraints. For example, the critical path constraint can be set to a weight of 5, and the secondary constraint can be set to a weight of 1, so that the optimization process prioritizes satisfying the critical links. The parameter settings reflect the fault-tolerance principle of engineering practice. No loss occurs when the expected position difference is within a threshold, preventing the model from overfitting and reducing generalization ability due to small deviations. The multi-task loss function uses a weighted summation method to fuse quality loss and constraint loss, ensuring that a single gradient signal simultaneously contains both quality and constraint information, enabling end-to-end joint training. The course learning strategy dynamically adjusts the constraint loss weights. Early training Smaller models prioritize learning basic layout patterns, and later... Increase the intensity of constraint satisfaction to avoid the narrow optimization space caused by strict constraints in the early stages of training, which could lead to suboptimal solutions.
[0046] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0047] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.
Claims
1. A production line layout optimization method based on constraint-aware Transformer and soft permutation matrix, characterized in that, Includes the following steps: S1: Obtain the equipment set information of the production line and the preset process constraint set, and convert the equipment set information into equipment embedding features; S2: Input the device embedding features into the constraint-aware Transformer model, extract global dependency features through the encoder of the constraint-aware Transformer model, and generate the position distribution information of each device in all layout positions based on the global dependency features; S3: Based on the location distribution information, a soft permutation matrix is generated through differentiable Sinkhorn iteration. The soft permutation matrix satisfies the double random condition and represents the allocation relationship between the device and the layout location in probabilistic form. S4: Based on the soft permutation matrix, the discrete allocation relationship between devices and layout positions is determined by the Hungarian algorithm, and a hard layout scheme is generated based on the discrete allocation relationship; S5: During the training of the constraint-aware Transformer model, the degree of violation of each process constraint is calculated based on the soft permutation matrix and a constraint loss is generated. The constraint loss and the layout quality prediction loss are combined to form a multi-task loss function to drive the constraint-aware Transformer model to simultaneously optimize layout quality and constraint satisfaction.
2. The production line layout optimization method based on constraint-aware Transformer and soft permutation matrix according to claim 1, characterized in that, In step S2, extracting global dependency features through the encoder of the constraint-aware Transformer model includes: A combined embedding representation is generated for each device and its corresponding layout location. Satisfying the relation: ; in, This represents the device embedding vector obtained by mapping device semantic information. This represents the position embedding vector that characterizes the absolute position of the device in the layout sequence. The combined embedded representation is input to the Transformer encoder, where a multi-head self-attention module calculates attention weights by aggregating interaction information between all devices. Satisfying the relation: ; in, , , The representations are respectively represented by combined embeddings. The query matrix, key matrix, and value matrix obtained through linear transformation This represents the dimension of the query and key vectors in the single-head attention mechanism; The Transformer encoder performs a nonlinear transformation on the output of the multi-head self-attention module through a feedforward neural network, and combines layer normalization and Dropout operations to output encoded global dependency features.
3. The production line layout optimization method based on constraint-aware Transformer and soft permutation matrix according to claim 1, characterized in that, In step S2, generating the location distribution information of each device at all layout locations based on global dependency features includes: The global dependency features are processed in parallel using a dual-predictor-head architecture of a constraint-aware Transformer model, which includes a quality prediction head and a location prediction head. The quality prediction head includes performing a global average pooling operation on the global dependent features to aggregate and obtain a feature vector representing the overall layout information. The feature vector representing the overall layout information is then mapped to a comprehensive quality score within a preset numerical range through a two-layer MLP. The location prediction head maps the global dependency features into a three-dimensional tensor through a linear transformation layer, which serves as a matrix representation of the location distribution information. The dimensions of the three-dimensional tensor correspond to the batch size, layout sequence length, and number of devices, respectively, where each element represents the original distribution score of the corresponding device assigned to a specific layout position.
4. The production line layout optimization method based on constraint-aware Transformer and soft permutation matrix according to claim 3, characterized in that, Step S3 includes: According to the preset temperature parameters The original distribution scores in the location distribution information are scaled to obtain a scaled logarithmic probability matrix S, which satisfies the following relationship: ; in, This represents the original distribution score in the location distribution information. A temperature parameter used to control the sharpness of a probability distribution; For the logarithmic probability matrix Perform alternating row and column normalization to make... The soft permutation matrix is obtained by satisfying the double random condition and through exponential operation. Satisfying the relation: ; in, This represents the final generated soft permutation matrix. This represents the logarithmic probability matrix that satisfies the double random condition after alternating row and column normalization. Indicates exponentiation; The soft permutation matrix Each element in the matrix represents the probability value of the corresponding device being assigned to a specific layout position, and the soft permutation matrix... The sum of the elements in each row and the sum of the elements in each column are both 1.
5. The production line layout optimization method based on constraint-aware Transformer and soft permutation matrix according to claim 4, characterized in that, Step S4 includes: For each sample, the soft permutation matrix Construct the cost matrix The cost matrix Satisfying the relation: ; in, Representing the cost matrix The Middle Line 1 Column elements, Represents the first in the soft permutation matrix The device was assigned to the first The probability value of each layout position; For the cost matrix The Hungarian algorithm is applied to find the perfect match with the minimum total cost, resulting in the device index array. and position index array ,in Indicates the first Device index in each matching pair. Indicates the first The location index to which the device is assigned in each matching pair. Indicates the index of the matching pair; According to the device index array and position index array Generate hard permutation matrix The hard permutation matrix The generation satisfies the following relation: ; In this case, each element of the hard permutation matrix is either 0 or 1, and each row and each column has exactly one 1; According to the hard permutation matrix The hard layout scheme is generated, which is a determined arrangement sequence of devices in the layout position.
6. The production line layout optimization method based on constraint-aware Transformer and soft permutation matrix according to claim 1, characterized in that, In step S1, converting the device set information into device embedded features includes: The device set information is formalized as a finite set including N devices. ; The preset set of process constraints is formalized into a set of binary constraints. Each constraint Indicates device Must be assigned to the device Before; Among them, feasible layout schemes include the device set. permutation sequence Permutation sequence Satisfy: Total number of layout positions Equal to the total number of devices and sequence Each device in It appears exactly once.
7. The production line layout optimization method based on constraint-aware Transformer and soft permutation matrix according to claim 6, characterized in that, Step S5 includes: For the set of process constraints Each constraint in computing devices and equipment Desired location, equipment Expected position and equipment Expected position Satisfying the relation: , ; in, Indicates the length of the layout sequence. Represents the soft permutation matrix medium equipment Assigned to position The probability value, Represents the soft permutation matrix medium equipment Assigned to position The probability value, Indicates the position index; Based on desired location and Calculate the set of process constraints Each constraint in The degree of violation, and generate the constraint loss. The constraint loss Satisfying the relation: ; in, Represents the set of process constraints The total number of constraints in the middle, Representing constraints Preset weights, This represents the preset tolerance range parameter; The constraint loss With layout quality prediction loss Perform weighted summation to construct the multi-task loss function. Among them, the layout quality prediction loss Satisfying the relation: ; in, Indicates the size of the current training batch. The constraint-aware Transformer is the first... The overall quality score predicted for each sample. Indicates the first Each sample corresponds to a preset target quality score.
Citation Information
Patent Citations
Transform and decision fusion-based modulation identification method
CN115238748A
Street building group automatic layout and iterative optimization method based on TFM architecture
CN120296840A