Assembly line variant path planning system and method based on graph deep reinforcement learning
The assembly line variant path planning system based on graph deep reinforcement learning solves the problems of local optimization and cold start with small data in assembly line variant path planning, and achieves efficient and accurate assembly line variant path planning, thereby improving the flexibility and adaptability of the assembly line.
Patent Information
- Application Number
- CN202511004286.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-11-21
AI Technical Summary
Existing assembly line variation path planning methods struggle to avoid local optimization when dealing with high-dimensional and complex information, and suffer from cold start problems with small datasets, leading to unreliable and inefficient planning results and an inability to achieve the global optimal solution.
An assembly line variant path planning system based on graph deep reinforcement learning is adopted. A neural network model is constructed by graph neural network and reinforcement learning strategy. Combined with domain experience and mechanism model, a reward function is designed for training to identify the optimal variant path.
It significantly improves the efficiency and quality of assembly line variation planning, can handle complex change transfer directed graphs, shortens reverse modeling time, improves model adaptability and accuracy, avoids local optimization, and achieves the global optimal solution.
Smart Images

Figure CN120996308A_ABST
Abstract
Description
Technical Field
[0001] This disclosure pertains to the field of modern assembly line transformation, and particularly relates to an assembly line variant path planning system and method based on graph deep reinforcement learning. Background Technology
[0002] In the face of increasingly complex and rapidly changing manufacturing demands, modern assembly lines require high flexibility and adaptability to cope with product variations, process adjustments, and market demand fluctuations. However, existing assembly lines struggle to effectively handle complex environmental information with high dimensions and multiple constraints. Traditional metaheuristic algorithms or experience-based methods are significantly insufficient in information processing capabilities and decision-making efficiency when dealing with the massive network relationships of parts, equipment states, process parameters, and multiple constraints in assembly lines, making it difficult to perform efficient optimization in complex high-dimensional spaces. Moreover, they are prone to getting trapped in local optima and failing to obtain the global optimum. Furthermore, existing assembly line planning methods are often based on the designer's intuitive experience or fixed optimization strategies, easily converging to local optima in complex solution spaces, failing to identify truly global optimal or better assembly line variation paths, thus limiting further improvements in production line performance.
[0003] Equally noteworthy is the "cold start" problem in existing assembly line modifications with limited data, resulting in poor model adaptability. Purely data-driven planning models suffer significant performance degradation due to insufficient training data when lacking sufficient historical data or facing new assembly line variations. This makes it difficult for them to quickly adapt to new environments and operating conditions, leading to unreliable planning results.
[0004] Therefore, a system capable of constructing a high-precision and highly adaptable assembly line variation path planning system, thereby significantly improving the efficiency and quality of assembly line variation planning, is in high demand. Summary of the Invention
[0005] This disclosure aims to address the aforementioned technical problems of existing assembly line variant path planning methods in handling high-dimensional complex information, avoiding local optimization, dealing with cold starts with small data, and achieving data and knowledge hybrid driving. It provides an assembly line variant path planning system and method based on graph deep reinforcement learning.
[0006] The technical solution disclosed herein is:
[0007] An assembly line variant path planning system based on graph deep reinforcement learning, comprising:
[0008] The input layer is used to receive the initial information of the variation requirements of the assembly line to be planned, as the data source;
[0009] The knowledge and mechanism layer is used to provide constraints.
[0010] The core processing module interacts with the input layer to convert the initial information into a graph structure, performs graph deep reinforcement learning using graph neural networks and reinforcement learning strategies, and constructs a neural network model. It sets a reward function for matching the preferred target of change propagation, and trains the neural network model by using the reward function of the preferred target of change propagation as the training basis and the constraints as the standard, to obtain the optimal variant path.
[0011] The output layer interacts with the core processing module to output the optimal variant path.
[0012] The knowledge and mechanism layer includes: a domain experience base and an incentive model base;
[0013] Both the domain experience library and the incentive model library interact with the core processing module to provide constraints.
[0014] The core processing module includes:
[0015] The graph modeling unit interacts with the input layer to convert the initial information into a graph structure.
[0016] The graph deep reinforcement learning unit interacts with the graph modeling unit to perform graph deep reinforcement learning based on the graph structure using graph neural networks and reinforcement learning strategies, thereby obtaining graph information data after deep learning.
[0017] The reward function and optimization objective evaluation unit interacts with the graph deep reinforcement learning unit to calculate the reward based on the graph information data after deep learning, according to the change propagation optimization mathematical model, to obtain the final reward result, and to obtain the optimal variant path based on the final reward result;
[0018] The reward function and the optimization objective evaluation unit interact with the output layer to output the optimal variant path.
[0019] The core processing module also includes: a knowledge fusion unit;
[0020] The knowledge fusion unit interacts with the reward function and optimization target evaluation unit to provide constraints when constructing the optimal mathematical model for change propagation.
[0021] The step of converting the initial information into a graph structure includes:
[0022] The initial information is converted into text using a multimodal model of graphics and text;
[0023] Multiplying the visual_embedding in the image-text multimodal model by the text_emedding yields a [N,N] matrix to be used.
[0024] The values on the diagonal of the matrix to be used are multiplied sequentially to obtain paired feature inner products;
[0025] The pairwise feature inner products are compared, and the largest value is selected as the choice.
[0026] The image architecture is obtained through an image editor, and positional information is added to the image architecture. The feature map is then expanded into a sequence to obtain the priority of any patch.
[0027] The obtained selections are matched with the positions of the highest priority tiles until the graph structure is obtained.
[0028] The assembly line variant path planning system based on graph deep reinforcement learning further includes: an application and feedback layer;
[0029] The application and feedback layer interacts with the core processing module to collect actual data from the assembly line using the optimal variant path; and sends the actual data to the reward function and optimization target evaluation unit in the core processing module to optimize the optimal variant path.
[0030] A method for planning variant paths in an assembly line includes:
[0031] Receive the initial information on the proposed assembly line variation requirements as the data source;
[0032] The initial information is converted into a graph structure, and graph deep reinforcement learning is performed using graph neural networks and reinforcement learning strategies to construct a neural network model and obtain graph information data after deep learning.
[0033] Set a reward function for matching the preferred target of change propagation, and train the neural network model by using the reward function of the preferred target of change propagation as the training basis and the constraints as the standard, so as to obtain the optimal variant path;
[0034] Collect actual data from the assembly line using the optimal variant path; and optimize the optimal variant path using the actual data.
[0035] The beneficial effects of this disclosure include at least the following:
[0036] The assembly line variation path planning system disclosed herein is capable of handling non-Euclidean space data and possesses powerful perception and intelligent sequential decision-making capabilities. It effectively handles the complex change transfer directed graph during assembly line variation, overcoming the shortcomings of traditional algorithms in processing high-dimensional information and avoiding local optimization. Simultaneously, it establishes a graph sample modeling approach under the difference between forward and reverse optimization objectives, including indicator prediction classes and self-organizing pattern improvement classes, to meet different planning needs. It also provides a dimensional linkage skeleton for instantiation and reconstruction, ultimately achieving rapid modeling of the virtual-physical mapping of the entire assembly process, significantly shortening the reverse modeling time of the physical assembly line. This disclosure has the advantages of high accuracy and high adaptability, and can significantly improve the efficiency and quality of assembly line variation planning. Attached Figure Description
[0037] Figure 1 This is the overall architecture diagram of the system described in this disclosure;
[0038] Figure 2 This is a detailed architecture diagram of a deep reinforcement learning model;
[0039] Figure 3 This is a flowchart of the method described in this disclosure. Detailed Implementation
[0040] The present application will now be further described with reference to the accompanying drawings.
[0041] Terminology Explanation:
[0042] Graph deep reinforcement learning is an extension of deep reinforcement learning to graph-structured data. It utilizes graph neural networks to process non-Euclidean space data, effectively capturing complex topological relationships and feature dependencies between graph nodes, thus demonstrating excellent performance in areas such as network optimization, recommender systems, and path planning.
[0043] Knowledge / mechanism model fusion: This refers to the method of combining prior knowledge such as domain expert experience, physical laws, mathematical formulas, and engineering principles, or mechanism-based simulation models, with data-driven machine learning models, such as deep learning models. This fusion aims to overcome the "cold start" problem of purely data-driven models with small amounts of data, and improve the model's generalization ability, interpretability, and robustness.
[0044] Assembly line transformation path planning: In the context of intelligent manufacturing, when an assembly line faces demands for product upgrades, process adjustments, capacity expansion, or flexible production, it involves dynamically adjusting and optimizing its structural layout, equipment configuration, and process flow, and planning a series of change steps and paths from the current state to the target state. Its goal is to achieve efficient, low-cost, and highly flexible rapid reconfiguration and optimization of the production line.
[0045] Metaheuristic algorithms are a class of iterative search methods used to solve complex optimization problems. They find approximate optimal solutions by simulating natural or physical phenomena, such as genetic algorithms, particle swarm optimization, and simulated annealing. However, when dealing with high-dimensional, multi-constraint, and nonlinear problems, their performance may be limited by the search space, making them prone to getting trapped in local optima.
[0046] Current assembly line variation path planning methods, when applied to complex manufacturing scenarios, often fall short in handling the increasingly complex, high-dimensional information present in assembly lines, such as equipment interconnections, process flows, product characteristics, and logistics relationships. They struggle to comprehensively perceive and analyze this information, leading to biased planning decisions. Furthermore, due to their inherent search mechanisms, traditional methods suffer from low computational efficiency when dealing with large-scale, non-convex assembly line variation path planning problems. They are prone to getting trapped in local optima, failing to converge to the globally optimal or suboptimal variation path, thus impacting the overall performance of the production line. Simultaneously, existing data-driven planning models heavily rely on large amounts of high-quality training data. In practical applications, especially during new product introductions or significant production line modifications, historical data is often insufficient, leading to a "cold start" problem, poor generalization ability, and weak adaptability to unknown or abnormal operating conditions. Moreover, existing methods fail to fully capture the domino effect that a single change can trigger during assembly line modifications, causing chain reactions across different equipment, workstations, and processes. If the complexity of this change transfer is not effectively modeled and predicted, the planning results will not match the actual situation. At the same time, the existing planning methods are disconnected from the virtual-physical mapping and reverse modeling process of the physical assembly line, resulting in a long verification cycle for the planning scheme and making it difficult to achieve rapid iterative optimization, thereby prolonging the production line reconfiguration time.
[0047] This disclosure provides a method for assembly line variation path planning based on graph deep reinforcement learning. The core of this method lies in constructing an intelligent planning model that can effectively handle the complexity of assembly line variations, overcome the cold start problem with small data, and identify the optimal variation path. This disclosure treats the essence of assembly line variation path planning decision-making as a complex directed graph problem of change propagation, aiming to identify the optimal variation path through intelligent decision-making. Furthermore, this method models the assembly line variation process as a graph structure and designs a reward function that accurately matches the change propagation optimization objective. By using the input change propagation optimization mathematical model as the training basis, the neural network model constructed based on graph deep reinforcement learning is repeatedly trained until the model output results stabilize and converge, thereby intelligently identifying the optimal initial variation path. This optimal path will provide a key dimensional linkage skeleton for subsequent instantiation and reconstruction, ultimately achieving rapid modeling of the virtual-physical mapping of the entire assembly process and significantly shortening the reverse modeling time of the physical assembly line. Meanwhile, to address the "cold start" problem of the model under small data conditions, this disclosure also innovatively integrates knowledge / mechanism models into the learning model to reduce the required training data and computational load, and enhance the model's adaptability to different environments and working conditions. Specific Implementation Example 1:
[0049] like Figure 1 A variant path planning system for assembly lines based on graph deep reinforcement learning includes: an input layer, a knowledge and mechanism layer, a core processing module, an output layer, and an application and feedback layer.
[0050] Specifically, the input layer is responsible for receiving the initial information regarding assembly line variation requirements and, as the data source in this embodiment, aggregating simulation data and real-time operational data. Its function is to provide comprehensive basic information about the current state of the assembly line and potential change directions for the subsequent planning process.
[0051] Specifically, the knowledge and mechanism layer, as an independent knowledge system, includes a domain experience base and various mechanism model libraries. Its role is to emphasize the importance of prior knowledge and physical mechanisms in this disclosure, and through interaction with the planning engine, to provide principled guidance and necessary constraints for intelligent decision-making.
[0052] Specifically, the core processing module, namely the intelligent planning and decision-making module, includes a graph modeling unit, a graph deep reinforcement learning model, a reward function and optimization objective evaluation unit, and a knowledge fusion module. Its main functions are data transformation, policy learning, result evaluation, and knowledge integration. The graph deep reinforcement learning model contains graph neural networks and reinforcement learning strategies; the reward function and optimization objective evaluation unit is used to calculate rewards based on the change propagation optimization mathematical model; the knowledge fusion module includes embedding knowledge into the learning model; and the graph modeling unit can convert assembly line information into a graph structure.
[0053] Specifically, the output layer's function is to explicitly output the "optimal initial variant path," representing the intelligent decision result obtained after processing by the planning engine. This path is a sequence of optimal variant solutions generated based on the comprehensive optimization objective, and is directly provided to downstream applications.
[0054] Specifically, the application and feedback layer demonstrates the downstream application value of the planning results and possible feedback loops. Its role is to promote the practical application of the planning results in areas such as "instantiated reconstruction," "rapid modeling of virtual-physical mapping," and "reducing the reverse modeling time of physical assembly lines," reflecting the ultimate goals and benefits of this disclosure, and supporting continuous system optimization.
[0055] Through the above technical solution, this embodiment achieves a comprehensive, data- and knowledge-driven digital description of the assembly line variation path planning process, constructing a high-fidelity intelligent planning system and laying a solid foundation for subsequent detailed algorithm implementation and application. This multi-dimensional and multi-level information definition and integration facilitates efficient hierarchical and refined management of planning data and models, laying a solid foundation for building a comprehensive intelligent variation digital twin of the assembly line.
[0056] like Figure 2 As shown, this embodiment abstracts the complex decision-making process of assembly line variation path planning into a sequential decision problem on a change transfer directed graph, and utilizes the powerful capabilities of graph deep reinforcement learning to solve it. This process mainly includes constructing the change transfer directed graph, designing and optimizing the reward function, and training and converging the deep reinforcement learning model.
[0057] Specifically, the graph data input module transforms assembly line variation-related information into a "directed graph of change propagation" as model input. This graph dynamically and comprehensively reflects assembly line elements and their interrelationships. Various equipment, product manufacturing characteristics, key workstations, process steps, and other relevant attributes within the assembly line are defined as nodes in the graph. Each node carries its own attribute vector, providing rich input information for the graph neural network. Nodes are connected by directed edges, representing the propagation of changes or mutual influences. The weights quantify the strength or frequency of these influences, such as logistical intensity or the number of times equipment and manufacturing characteristics are matched, collectively constructing a topology that captures the complex relationships and change propagation paths of the assembly line.
[0058] Specifically, the graph neural network processing module serves as the core processing unit. This module receives input graph data and its role is to effectively process the graph structure data to learn complex variation strategies. This module employs advanced GNN architectures such as GCN, GAT, or Graph SAGE as the main body of the policy network and / or value network. GNNs effectively integrate node features and graph topology information through iterative information aggregation mechanisms, extracting node embeddings or graph representations containing rich semantics, thereby understanding the complex topology and intrinsic relationships of the assembly line variation graph. The graph neural network processing module can utilize a GNN network module.
[0059] Specifically, the reinforcement learning decision module is responsible for transforming the features processed by the GNN into executable decisions. It contains key components of reinforcement learning: the policy network receives the features extracted by the GNN and outputs the probability distribution of different variant actions in the current assembly line variant state, such as adding equipment, adjusting the workstation layout, and modifying process parameters, to guide the agent's decision-making; the value network evaluates the expected long-term reward of the current assembly line variant state or state-action pair; and the action space defines all possible variant actions that the model can execute.
[0060] Specifically, the reward calculation and model optimization loop provides a crucial feedback mechanism to drive the training and optimization of the deep reinforcement learning model. It receives the actions performed by the reinforcement learning agent and the new state from environmental feedback, and calculates and generates reward signals based on a pre-defined "change propagation optimization mathematical model." The reward function is the core driver of the model's learning of the optimal strategy. By quantifying multiple optimization indicators such as mutation time, cost, production line flexibility, product quality, production efficiency, resource utilization, and potential risk avoidance, it precisely matches the "change propagation optimization objective." After each decision, the simulated environment calculates the reward value in real time and feeds it back to the agent; positive rewards encourage beneficial actions, while negative rewards punish unfavorable actions.
[0061] Specifically, the output module is used to output the final decision result of the trained graph deep reinforcement learning model, namely the "optimal variant strategy" or "optimal variant path sequence", to provide accurate and efficient decision support for the variants of the actual assembly line.
[0062] Through the above technical solution, this disclosure accurately transforms the assembly line variation path planning problem into a sequential decision-making process within a graph deep reinforcement learning framework. By constructing a refined change propagation directed graph, designing a reasonable reward function, and training an efficient graph neural network model, this disclosure can intelligently identify the optimal variation path, thereby overcoming the shortcomings of traditional methods in handling high-dimensional complex information and avoiding local optimization, and providing core algorithmic support for intelligent variation in assembly lines. Specific Implementation Example 2:
[0064] like Figure 3A method for planning variant paths in an assembly line, based on the graph deep reinforcement learning-based assembly line variant path planning system as described in Specific Embodiment 1, includes: receiving initial information of variant requirements of the assembly line to be planned as a data source; converting the initial information into a graph structure, performing graph deep reinforcement learning using a graph neural network and reinforcement learning strategies to construct a neural network model and obtain graph information data after deep learning; setting a reward function for matching the preferred target of variant propagation, training the neural network model by using the reward function of the preferred target of variant propagation as the training basis and the constraints as the standard, and obtaining the optimal variant path; collecting actual data of the assembly line applying the optimal variant path; and optimizing the optimal variant path using the actual data.
[0065] This disclosure also provides an embodiment:
[0066] An electronic device includes: a storage medium and a processing unit; wherein the storage medium is used to store a computer program, and the processing unit exchanges data with the storage medium for executing the computer program through the processing unit when performing urban local-scale carbon emission calculations and carbon neutrality predictions, performing the steps of the assembly line variant path planning method as described in Specific Embodiment 2.
[0067] A computer-readable storage medium storing a computer program; when the computer program is run, it executes the steps of the assembly line variation path planning method as described in Specific Embodiment 2.
[0068] In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wireline, optical fiber, RF, etc., or any suitable combination thereof.
[0069] The above disclosures only cover a few specific implementation scenarios. However, this disclosure is not limited to these, and any variations that can be conceived by those skilled in the art should fall within the protection scope of this disclosure. The serial numbers in this disclosure are for descriptive purposes only and do not represent the superiority or inferiority of the implementation scenarios.
Claims
1. A variant path planning system for assembly lines based on graph deep reinforcement learning, characterized in that, include: The input layer is used to receive the initial information of the variation requirements of the assembly line to be planned, as the data source; The knowledge and mechanism layer is used to provide constraints. The core processing module interacts with the input layer to convert the initial information into a graph structure, and uses graph neural networks and reinforcement learning strategies to perform graph deep reinforcement learning and build a neural network model. Set a reward function for matching the preferred target of change propagation, and train the neural network model by using the reward function of the preferred target of change propagation as the training basis and the constraints as the standard, so as to obtain the optimal variant path; The output layer interacts with the core processing module to output the optimal variant path.
2. The assembly line variant path planning system based on graph deep reinforcement learning according to claim 1, characterized in that: The knowledge and mechanism layer includes: a domain experience base and an incentive model base; Both the domain experience library and the incentive model library interact with the core processing module to provide constraints.
3. The assembly line variant path planning system based on graph deep reinforcement learning according to claim 1, characterized in that, The core processing module includes: The graph modeling unit interacts with the input layer to convert the initial information into a graph structure. The graph deep reinforcement learning unit interacts with the graph modeling unit to perform graph deep reinforcement learning based on the graph structure using graph neural networks and reinforcement learning strategies, thereby obtaining graph information data after deep learning. The reward function and optimization objective evaluation unit interacts with the graph deep reinforcement learning unit to calculate the reward based on the graph information data after deep learning, according to the change propagation optimization mathematical model, to obtain the final reward result, and to obtain the optimal variant path based on the final reward result; The reward function and the optimization objective evaluation unit interact with the output layer to output the optimal variant path.
4. The assembly line variant path planning system based on graph deep reinforcement learning according to claim 3, characterized in that, The core processing module also includes: a knowledge fusion unit; The knowledge fusion unit interacts with the reward function and optimization target evaluation unit to provide constraints when constructing the optimal mathematical model for change propagation.
5. The assembly line variant path planning system based on graph deep reinforcement learning according to claim 3, characterized in that, The step of converting the initial information into a graph structure includes: The initial information is converted into text using a multimodal model of graphics and text; Multiplying the visual_embedding in the image-text multimodal model by the text_emedding yields a [N,N] matrix to be used. The values on the diagonal of the matrix to be used are multiplied sequentially to obtain paired feature inner products; The pairwise feature inner products are compared, and the largest value is selected as the choice. The image architecture is obtained through an image editor, and positional information is added to the image architecture. The feature map is then expanded into a sequence to obtain the priority of any patch. The obtained selections are matched with the positions of the highest priority tiles until the graph structure is obtained.
6. The assembly line variant path planning system based on graph deep reinforcement learning according to claim 1, characterized in that, Also includes: Application and feedback layers; The application and feedback layer interacts with the core processing module to collect actual data from the assembly line using the optimal variant path; and sends the actual data to the reward function and optimization target evaluation unit in the core processing module to optimize the optimal variant path.
7. A method for planning variant paths in an assembly line, based on the assembly line variant path planning system based on graph deep reinforcement learning as described in any one of claims 1-5, characterized in that, include: Receive the initial information on the proposed assembly line variation requirements as the data source; The initial information is converted into a graph structure, and graph deep reinforcement learning is performed using graph neural networks and reinforcement learning strategies to construct a neural network model and obtain graph information data after deep learning. Set a reward function for matching the preferred target of change propagation, and train the neural network model by using the reward function of the preferred target of change propagation as the training basis and the constraints as the standard, so as to obtain the optimal variant path; Collect actual data from the assembly line using the optimal variant path; and optimize the optimal variant path using the actual data.