Intelligent dynamic scheduling method and system for spacecraft final assembly workshop

By constructing a multi-evaluation network collaborative security near-end strategy optimization model, the problems of static and passive scheduling, conflicting optimization objectives, and insufficient security in the spacecraft assembly workshop scheduling were solved, achieving efficient and safe dynamic scheduling that meets the requirements of real-time performance and flexibility.

CN122155312APending Publication Date: 2026-06-05SHANGHAI SPACE PRECISION MACHINERY RES INST

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI SPACE PRECISION MACHINERY RES INST
Filing Date
2026-04-28
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Traditional spacecraft assembly workshop scheduling methods suffer from static and passive nature, single or conflicting optimization objectives, insufficient safety, and high computational complexity. They are difficult to achieve dynamic scheduling that integrates multiple objectives and ensures safety, and cannot meet the requirements for real-time performance and flexibility.

Method used

A multi-evaluation network collaborative security near-end policy optimization model is constructed using real-time multi-source data. Through the multi-evaluation network collaborative security near-end policy optimization algorithm, security constraints are embedded to construct a decision optimization problem with security constraints, realizing multi-objective fusion and dynamic scheduling. The neural network is used for online update and rescheduling decisions.

Benefits of technology

It significantly improves scheduling response speed and production efficiency, reduces labor costs, enhances human-machine safety and the satisfaction rate of safety constraints for pyrotechnic products, and supports flexible production scheduling for complex production modes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122155312A_ABST
    Figure CN122155312A_ABST
Patent Text Reader

Abstract

The application provides a kind of aerospace vehicle assembly workshop intelligent dynamic scheduling method and system, method includes steps S1: by real-time connection assembly workshop central control system to obtain real-time multi-source data;Step S2: training multi-evaluation network coordination safety near end strategy optimization model, optimal control sequence in limited time domain is solved at each sampling time, only the first control action is adopted, the next sampling time is optimized based on last control action and state, after completing all control action optimization, reward model is based on neural network back update;Repeat the above process until convergence;Step S3: the multi-evaluation network coordination safety near end strategy optimization model trained is deployed to carry out workshop state update and rescheduling decision, generate and issue new scheduling plan.The application realizes the real-time response to the abnormal disturbance such as response emergency order, equipment failure, realizes the safe, efficient, flexible dynamic self-adapting scheduling of aerospace vehicle assembly workshop.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of intelligent manufacturing and artificial intelligence technology, specifically to an intelligent dynamic scheduling method and system for aerospace vehicle assembly workshops. Background Technology

[0002] Spacecraft final assembly is a core link in aerospace product manufacturing. Its production process is characterized by multi-model mixed-line production, complex process paths, close human-machine collaboration, frequent dynamic disturbances, and extremely high safety requirements. The spacecraft final assembly workshop adopts a dual-line layout of Assembly and Testing Line 1 and Assembly and Testing Line 2, supporting the co-production of multiple spacecraft models. Due to the long assembly cycle and unpredictable pace of spacecraft products, and the involvement of flammable and explosive materials such as pyrotechnics, extremely high demands are placed on the safety, real-time performance, and flexibility of the scheduling system.

[0003] Traditional spacecraft assembly shop scheduling methods primarily rely on advanced planning and scheduling modules within manufacturing execution systems (MES), which are mostly based on mixed-integer programming or heuristic rules for scheduling. However, existing technologies suffer from the following significant drawbacks: Static vs. Passive: Traditional scheduling methods typically generate a static, globally optimal plan before production begins. However, the actual production process in a spacecraft assembly workshop is fraught with uncertainties, such as emergency orders, sudden equipment failures, material supply delays, product quality anomalies, and temporary staff shortages. Faced with these dynamic disturbances, static plans often quickly become ineffective, requiring manual intervention for adjustments. This intervention suffers from high response delays, and the adjustment results are difficult to guarantee as optimal, severely impacting production efficiency and order delivery rates.

[0004] Single or conflicting optimization objectives: The optimization objectives of a spacecraft assembly workshop are multi-dimensional, requiring the pursuit of production efficiency (minimizing total completion time), consideration of resource utilization (maximizing production line balance), and attention to the working conditions of operators (minimizing personnel movement). These objectives often conflict with each other. Traditional methods struggle to effectively balance multiple conflicting objectives, typically focusing on only one or two primary objectives, resulting in poor overall efficiency in scheduling.

[0005] Insufficient safety assurance: Spacecraft assembly workshops involve flammable and explosive materials such as pyrotechnics, requiring extremely high safety standards. Existing scheduling methods fail to embed key safety factors such as pyrotechnic safety constraints, human-machine safety distance constraints, and exclusive use of explosion-proof workstations into the scheduling algorithm, posing potential safety hazards.

[0006] High computational complexity: For large-scale, multi-constraint spacecraft assembly workshop scheduling problems, precise mathematical programming methods often face combinatorial explosion problems, resulting in excessively long solution times and failing to meet the requirements of real-time dynamic scheduling. While heuristic rules are fast, the quality of solutions is usually poor, falling far short of the global optimum.

[0007] Therefore, there is an urgent need for an intelligent dynamic scheduling method suitable for spacecraft assembly workshops, which can achieve multi-objective fusion, ensure safe dynamic scheduling, and have high real-time performance and anti-interference characteristics, so as to ensure the safe, efficient and stable operation of human-machine collaborative production in spacecraft assembly workshops. Summary of the Invention

[0008] To address the shortcomings of existing technologies, the purpose of this invention is to provide an intelligent dynamic scheduling method and system for aerospace vehicle assembly workshops.

[0009] A method for intelligent dynamic scheduling in a spacecraft assembly workshop according to the present invention includes: Step S1: Obtain real-time multi-source data by connecting to the central control system of the final assembly workshop; The real-time multi-source data includes static basic data and dynamic data; Step S2: Construct a workshop state training set based on the real-time multi-source data and train a multi-evaluation network collaborative safety near-end strategy optimization model accordingly; solve for the optimal control sequence in the finite time domain at each sampling time, adopt only the first control action, and re-optimize based on the previous control action and state at the next sampling time. After completing the optimization of all control actions, perform reverse update of the neural network based on the reward model; repeat the above process until convergence. Step S3: Deploy the trained multi-evaluation network collaborative security near-end strategy optimization model. When abnormal disturbances are detected, input real-time information into the central control system of the final assembly workshop and trigger the rescheduling mechanism to complete the workshop status update and rescheduling decision, and generate and issue a new scheduling plan.

[0010] Preferably, step S2 includes the following sub-steps: Step S2.1: Construct a set of production scheduling objective functions containing three optimization objectives; The goal is to minimize the maximum completion time of all pending work orders; The objective is to minimize the total distance traveled by personnel and the variance of workload among personnel. The objective is to minimize the variance of the planned working hours load for each workstation; The formulas are expressed as follows:

[0011]

[0012]

[0013] in , This is a collection of all work orders. For work orders Completion time; For the gathering of all operators, For scheduling periods, For personnel During the period Within the mission's movement distance, For personnel Total effective work hours within the scheduling cycle Represents variance. and These are the weighting coefficients; For the set of all workstations, To schedule workstations within the field of view The sum of standard working hours for all assigned processes; Step S2.2: Define process sequence constraints, resource quantity constraints, equipment and workstation exclusivity constraints, human-machine safety distance constraints, explosive buffer capacity constraints, and personnel qualification constraints. Model the scheduling problem as a constrained Markov decision process, and transform various safety constraints into safety cost terms of the safety near-end strategy optimization algorithm, embedding them into the algorithm's immediate cost function to construct a decision optimization problem with safety constraints. Step S2.3: Transform the acquired real-time multi-source data into a standardized state space representation; the state vector includes workstation state sub-vectors, logistics equipment state sub-vectors, personnel state sub-vectors, material inventory state sub-vectors, and work order queue state sub-vectors; Step S2.4: Define the action space, and the decision output is a composite action vector. The plan is parsed into four types of structured plans: material delivery instruction set, work order execution instruction set, personnel assignment instruction set, and equipment control instruction set; at any decision-making moment... The model's decision output is a structured action vector, formally represented as:

[0014] in, For the action space, For material distribution instruction set, For the work order execution instruction set, Assign instruction sets to personnel. For equipment control instruction set; Step S2.5: Construct a multi-evaluation network collaborative optimization mechanism. Construct independent Critic networks for the three optimization objectives respectively. Merge the independent advantage functions into a unified comprehensive advantage signal through a linear weighted fusion function with learnable weights. Update the policy network using the pruning objective function optimized by the near-end strategy. Step S2.6: Construct and execute the training and online update process of the safety near-end policy optimization algorithm. Complete the offline training of the algorithm based on the historical operation data of the workshop and the sample data generated by the simulation environment. Combine the real-time operation status of the workshop and the scheduling decision feedback to complete the online dynamic update of the algorithm. Quantify the safety constraints into algorithm penalty terms and embed them into the cost function and update mechanism of the algorithm to realize the continuous optimization of the algorithm under safety constraints.

[0015] Preferably, the process sequence constraint requires that each product process strictly follow the preset process route. The subsequent process cannot be started if the previous process is not completed, and a certain period of waiting is set between the critical assembly process and the inspection process. That is, for any work order , It is a set of work orders, and the processes contained therein must be executed according to a preset partial order relationship; set up For process The direct successor process must satisfy the following:

[0016] in, and Each is a process The start and end times, Allow time for necessary inter-process preparation or material transfer. For work orders The set of partial order relations of the process steps; actions that violate this constraint will be considered infeasible. Resource quantity constraints limit the number of AGVs configured on each production line to no more than The number of units and operators shall not be less than Each AGV has a rated load, and the number of tooling fixtures at each workstation matches the requirements of the corresponding process, meaning that at any decision-making moment... The total amount of any type of resource allocated to a task must not exceed its current available amount, as expressed by the formula:

[0017] in, for The set of all selected scheduling actions at any given time. ( ) is an action Resources The demand, For resources exist Total available time; Equipment and workstation exclusivity constraints require that high-precision testing equipment cannot be occupied by other work orders while performing testing tasks, and that the same operator can only be assigned to one workstation at a time. and Each represents a process. In the work order Start time and processing time For 0-1 decision variables, when the work order Assigned to work station The value is 1 if the condition is met, and 0 otherwise; for any two different work orders If they might use the same workstation Then the following mutual exclusion conditions must be met:

[0018]

[0019]

[0020] in, For a sufficiently large positive number, This is a sequence indicator variable used to ensure that two work orders are at the workstation. The processing sequences on each part do not overlap; Safety constraints include human-machine safety distance constraints, explosive buffer capacity constraints, exclusive use of explosion-proof workstations constraints, electrostatic discharge detection constraints, and personnel qualification constraints. Human-machine safety distance constraint definition real-time safety cost function:

[0021] in For a moment The safety cost function; For a moment The condition of the workshop; For a moment The action; To preset a safe distance threshold, To perform the action The predicted minimum Euclidean distance between all AGVs and all personnel is calculated using the following formula: ,when The action was refused; In the formula For AGV Position coordinates; For personnel The location coordinates.

[0022] The buffer capacity constraint for pyrotechnic materials limits the maximum number of line-side buffers for flammable and explosive materials. at any time Total inventory of pyrotechnic products at the production line satisfy ,when When an alert is triggered, New delivery orders for this material are prohibited at this time; The exclusivity constraint of the explosion-proof workstation requires that the assembly process involving pyrotechnics must be performed at the explosion-proof workstation. During the pyrotechnics process, no other tasks may be assigned to this workstation. The workstation status is marked as locked until the process is completed and a safety inspection is passed. Electrostatic discharge (ESD) detection requirements stipulate that ESD detection must be completed before the start of the pyrotechnics process, and the detection signal... Only when Only when the time is right can the process be started; Personnel qualification constraints encode personnel qualifications as vectors The qualification levels are represented by five dimensions, and the process requirements are coded as vectors. Only when For all Only upon establishment can the personnel be assigned to perform the procedure.

[0023] Preferably, step S2.5 includes: Construct three Critic networks with identical structures but independent parameters. The input layer dimension of the network structure is... The hidden layers consist of two fully connected layers, each with 256 neurons. The output layer dimension is... The activation function used is ReLU, and the three networks output... , , Each evaluation network independently computes its advantage function.

[0024] in , For the first An evaluation network computation advantage function; For the first Instant rewards for each optimization goal; For the first The state value function of a Critic network; For the first Parameters of a Critic network; This is the discount factor.

[0025] Each reward function is defined as follows: This represents the change in completion time (negative values ​​indicate penalties). This represents the change in maximum completion time. For changes in personnel efficiency, Defined as the change in the total distance traveled by people. This represents the personnel load vector. This represents the change in production line balance. This represents the workstation load vector. Employing learnable weights The linear weighted fusion method yields the comprehensive advantage function. The constraints are and Weight parameters Update the policy network synchronously via gradient descent; Policy network updates use PPO pruning objective function

[0026] In the formula Tailor the objective function for PPO. For policy network parameters; Importance sampling ratio, The action probability output by the policy network. The action probabilities output by the old policy network. To trim hyperparameters, This is the pruning function. Parameters of each Critic network. By minimizing the time-series difference error update, the th Loss function of a Critic network for:

[0027] Preferably, the training and online update process for constructing and executing a secure near-end policy optimization algorithm includes: The security near-end policy optimization is based on the standard PPO algorithm, which quantifies security constraints as penalty terms and embeds them into the objective function. The core optimization objective is:

[0028] in, For interactive trajectories; Weighted rewards for multiple objectives; The cost function is violated due to safety constraints; This is a safety penalty coefficient; For parameterized random policy networks, For strategy parameters; The safety cost function is composed of a weighted average of various constraint penalty terms: .

[0029] in Pyrotechnics / Explosion-proof / Electrostatic Confinement Weights The remaining constraints are equally distributed among the remaining weights. For constraint categories The penalty function.

[0030] The definitions of various penalty items are as follows: The penalty for process sequence constraints is:

[0031] in Indicate process for The preceding process.

[0032] The resource quantity constraint penalty is:

[0033] The penalty for human-machine safe distance constraints is:

[0034] The minimum Euclidean distance between the AGV and the personnel; The core update formula of the Safe-PPO algorithm: The enhanced advantage function is ; , which is the original multi-objective advantage function; The policy network update uses a pruned objective function to ensure update stability.

[0035] in Importance sampling ratio, To trim hyperparameters, The action probability output by the policy network. The probability of the action output by the old policy network; Policy networks maximize Updated Adam optimizer parameters: , ; In the formula For the policy network learning rate, , These are the first and second momentum coefficients of the Adam optimizer.

[0036] The evaluation network update module considers three types of Critic networks and updates by minimizing the TD error:

[0037] in, , Adam optimizer learning rate ; Perform multi-objective weight updates The updated normalized weights ensure that the constraints are satisfied; the offline training process trains the dataset. Number of iterations Batch size Update round number .

[0038] Update policy network: ,in Let the loss function be the policy network. The learning rate of the policy network. Represents the loss function Regarding policy network parameters The gradient; Updated evaluation network: ,in To evaluate the network's loss function, To evaluate the learning rate of the network, Represents the loss function Regarding the first Evaluation network parameters The gradient; Update weights: ,in For the first An adaptive weight for evaluating the network, The learning rate for weight updates. Represents the expectation operator. For the first The target return of an evaluation network These are Lagrange multipliers for safety constraints, used to balance the relationship between reward maximization and safety constraints. For safety cost function; Normalized weights Ensure that the sum of the weights of all evaluation networks is 1.

[0039] Calculate the average reward on the validation set Constraint violation rate , ;like , Save the optimal strategy parameters .

[0040] Preferably, the convergence criteria include: Average reward of validation set Rate of change over 200 consecutive rounds ; constraint violation rate ; Strategy parameter update magnitude .

[0041] Preferably, the online update process in step S2.6 includes: Collect real-time data, in order to Data collection status Cache the most recent Step trajectory data, batch size Updates are triggered periodically; the online strategy fine-tuning sets the learning rate to 1 / 10 of the offline stage. Update round number to To avoid overfitting; the evaluation network is updated incrementally only; real-time safety constraint monitoring and calculation. Record the type and severity of the violation; five consecutive violations of the same constraint: increase and If there are no violations for 100 consecutive steps, restore the baseline value; back up before each update. Rollback occurs when efficiency declines or violation rates increase.

[0042] Preferably, step S2 further includes introducing a model predictive control framework for rolling optimization to establish a workshop state transition model: ; in This is an approximation of the state transition function in a deep neural network. For process noise, based on the current state and control action sequence Predicting the future Step state sequence Subsequently at each sampling time Solving finite-time domain problems The optimal control sequence within the range, with the optimization objective being to minimize the cumulative cost function.

[0043] in For immediate cost function, For the terminal cost function, As a discount factor; only the first action of the optimal control sequence is adopted. At the next sampling time Reacquire the actual state Calculate the prediction error The solution is then optimized again based on the actual conditions; a fixed sampling period is also set. Introduce an event triggering mechanism and define an event set: When the event When it occurs, rescheduling is triggered immediately and is not limited by a fixed sampling period.

[0044] Preferably, step S3 includes: The workshop status is polled periodically over a preset time period, and a threshold for status change is defined. When any state variable changes The system determines an event has occurred upon receiving an external event notification; after the event occurs, it completes a state space update within seconds; the scheduling model completes the neural network forward computation and outputs a new action sequence within a preset time period after receiving the updated state, and completes plan parsing and distribution; simultaneously, it defines a set of switching decision points. ,for The new process is executed according to the new plan, and the original execution plan is maintained for processes that have already started until completion; the event detection data, state update data and execution feedback data of the rescheduling decision mentioned above are all used as online update incremental samples, and the algorithm completes small batch incremental updates after rescheduling; a special sample pool is established for high-frequency disturbance events for focused training and updating.

[0045] According to the present invention, an intelligent dynamic scheduling system for a spacecraft assembly workshop includes: The data interface module is used to connect to the central control system of the final assembly workshop in real time to obtain real-time multi-source data; The state awareness and construction module is used to construct the real-time multi-source data into a normalized state space vector; The intelligent scheduling decision module has a built-in pre-trained multi-evaluation network collaborative security near-end policy optimization model, which is used to receive the state space vector and output the optimal action vector; The security constraint management module is used to embed various security constraints into the intelligent scheduling decision module and perform security verification during decision-making. The dynamic response module is used to trigger rescheduling when abnormal disturbances are detected, enabling the central control system in the final assembly workshop to complete the workshop status update and rescheduling decision. The plan parsing and distribution module is used to parse the optimal action vector output by the intelligent scheduling decision module into a scheduling plan, and distribute it to the central control system of the final assembly workshop for execution through the data interface module; The algorithm training and update module is used to perform offline training and online updating of the multi-evaluation network collaborative security near-end policy optimization model.

[0046] Compared with the prior art, the present invention has the following beneficial effects: 1. The response speed of this invention is significantly improved, with a scheduling decision response time of less than 3 seconds and a rescheduling completion time of less than 30 seconds. This is more than 10 times faster than traditional mathematical programming methods (which typically take more than 300 seconds to solve), thus meeting the needs of real-time dynamic scheduling.

[0047] 2. The production efficiency of this invention is significantly improved. Through multi-objective optimization and MPC rolling optimization, the total completion time is reduced, the production line balance rate is improved, and the waiting time at bottleneck workstations is reduced, thereby reducing the waiting waste caused by bottleneck workstations or production line imbalances and maximizing the overall output of the workshop.

[0048] 3. By incorporating personnel movement efficiency into the optimization objective, this invention reduces personnel movement distance, increases AGV utilization, reduces energy consumption per unit product, reduces labor intensity, and improves the efficiency of manual operations per unit time, thereby indirectly reducing labor costs.

[0049] 4. By introducing a safety reinforcement learning algorithm and a safety shield mechanism, this invention reduces human-machine safety distance violations to zero, transforming safety from a passive, sensor-based emergency obstacle avoidance to an active, pre-planned safety scheduling, which greatly improves the safety of human-machine hybrid working environments.

[0050] 5. This invention embeds safety constraints for pyrotechnic products into the scheduling algorithm, including upper limit constraints for line-side buffers, exclusive constraints for explosion-proof workstations, and electrostatic discharge detection constraints. The safety constraint satisfaction rate for pyrotechnic products reaches 100%, fundamentally eliminating potential safety hazards of pyrotechnic products.

[0051] 6. This invention supports complex production modes. The method and system of this invention can effectively cope with complex scenarios such as long-cycle and short-cycle dual production lines, multiple workstations, and multiple models co-produced in aerospace vehicle assembly workshops. Its data-driven and self-learning characteristics enable it to adapt to the process flow and production cycle of different products, and achieve highly flexible production scheduling. Attached Figure Description

[0052] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0053] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.

[0054] like Figure 1 As shown, a method for intelligent dynamic scheduling in a spacecraft assembly workshop includes: Step S1: Obtain real-time multi-source data by connecting to the central control system of the final assembly workshop; The real-time multi-source data includes static basic data and dynamic data; Step S2: Construct a workshop state training set based on the real-time multi-source data and train a multi-evaluation network collaborative safety near-end strategy optimization model accordingly; solve for the optimal control sequence in the finite time domain at each sampling time, adopt only the first control action, and re-optimize based on the previous control action and state at the next sampling time. After completing the optimization of all control actions, perform reverse update of the neural network based on the reward model; repeat the above process until convergence. Step S3: Deploy the trained multi-evaluation network collaborative security near-end strategy optimization model. When abnormal disturbances are detected, input real-time information into the central control system of the final assembly workshop and trigger the rescheduling mechanism to complete the workshop status update and rescheduling decision, and generate and issue a new scheduling plan.

[0055] In one embodiment, the above steps specifically include: Step S1: Connect to the central control system of the final assembly workshop in real time to obtain static basic data including process layout, process flow of each product model, bill of materials, standard work tasks, production shifts and rest time, as well as dynamic data including orders to be produced, line-side inventory, real-time equipment status, personnel location and status; in: Workstation status data was collected for two production lines, Assembly Line 1 and Assembly Line 2. Each production line contains... Each workstation contains [number] workstations and each workstation contains [number] workstations. Each station collects specific information including the currently occupied work order ID, the executed process number, the remaining processing time, and the ready / fault / maintenance status.

[0056] The status data of logistics equipment is collected from AGV devices within the workshop, specifically including AGV number and location coordinates. Load material code, task destination, remaining distance / time, and idle / busy / charging / fault status.

[0057] Personnel status data, including personnel ID and location coordinates. Work status, skill level vector, and list of operable process types.

[0058] Material inventory status data includes material code, storage location, available quantity, and batch information, with pyrotechnic materials being separately identified and monitored; The work order queue status data includes work order number, product model code, planned quantity, completed quantity, remaining process list, delivery date, and priority weight.

[0059] The aforementioned static basic data provides basic prior information such as process, resources, and constraints for the offline training of the safety near-end strategy optimization algorithm, while the dynamic data provides real-time state samples for the online update of the algorithm. Among them, safety-related dynamic data such as pyrotechnics and human-machine positions are the core basis for adjusting the safety penalty parameters when the algorithm is updated.

[0060] Step S2: Based on the data obtained in step S1, construct a set of production scheduling objective functions containing three optimization objectives: minimizing the maximum completion time of all pending work orders, minimizing the total personnel movement distance and the variance of workload among personnel, and minimizing the variance of planned working hours at each workstation. The formula is expressed as follows:

[0061]

[0062] ; in , This is a collection of all work orders. For work orders Completion time; For the gathering of all operators, For scheduling periods, For personnel During the period Within the mission's movement distance, For personnel Total effective work hours within the scheduling cycle Represents variance. and These are the weighting coefficients; For the set of all workstations, To schedule workstations within the field of view The sum of standard working hours for all assigned processes.

[0063] Step S3: Define process sequence constraints, resource quantity constraints, equipment and workstation exclusivity constraints, human-machine safety distance constraints, explosive buffer capacity constraints, and personnel qualification constraints. Model the scheduling problem as a constrained Markov decision process, and transform various safety constraints into safety cost terms of the safety near-end strategy optimization algorithm, embedding them into the algorithm's immediate cost function to construct a decision optimization problem with safety constraints. The constraints defined above include process constraints, resource constraints, and safety constraints. Among them, the process sequence constraint requires that each product process strictly follow the preset process route. The subsequent process cannot start until the previous process is completed. In addition, a certain period of waiting is set between the critical assembly process and the inspection process. That is, for any work order , Let be a set of work orders, whose contained processes must be executed according to a preset partial order relationship. For process The direct successor process must satisfy the following: ; in, and Each is a process The start and end times, Allow time for necessary inter-process preparation or material transfer. For work orders The set of partial order relations of the processes; actions that violate this constraint will be considered infeasible.

[0064] Resource quantity constraints limit the number of AGVs configured on each production line to no more than The number of units and operators shall not be less than Each AGV has a rated load, and the number of tooling fixtures at each workstation matches the requirements of the corresponding process, meaning that at any decision-making moment... The total amount of any type of resource (materials, equipment, personnel) allocated to a task must not exceed its current availability. The formula is expressed as: ; in, for The set of all selected scheduling actions at any given time. ( ) is an action Resources The demand, For resources exist The total amount of time available.

[0065] Equipment and workstation exclusivity constraints require that high-precision testing equipment cannot be occupied by other work orders while performing testing tasks, and that the same operator can only be assigned to one workstation at a time. and Each represents a process. In the work order Start time and processing time For 0-1 decision variables, when the work order Assigned to work station The value is 1 if the condition is met, and 0 otherwise. For any two distinct work orders... If they might use the same workstation Then the following mutual exclusion conditions must be met:

[0066]

[0067]

[0068] in, For a sufficiently large positive number, This is a sequence indicator variable used to ensure that two work orders are at the workstation. The processing order on each part does not overlap.

[0069] Safety constraints include human-machine safety distance constraints, explosive buffer capacity constraints, explosion-proof workstation exclusivity constraints, electrostatic discharge detection constraints, and personnel qualification constraints. Among these, the human-machine safety distance constraint defines the real-time safety cost function. , For a moment The safety cost function, For a moment The condition of the workshop For a moment The action, To preset a safe distance threshold, To perform the action The predicted minimum Euclidean distance between all AGVs and all personnel is calculated using the following formula: ,when When the action was refused, For AGV Position coordinates; For personnel Location coordinates. The limit on the number of line-side buffers for flammable and explosive materials is set by the pyrotechnics buffer capacity constraint. at any time Total inventory of pyrotechnic products at the production line satisfy ,when When an alert is triggered, New delivery tasks for this material are prohibited at this time; the exclusivity constraint of the explosion-proof workstation requires that assembly processes involving pyrotechnics must be performed at the explosion-proof workstation, and no other tasks may be assigned to this workstation during the pyrotechnic process. The workstation status is marked as locked until the process is completed and a safety test is passed; the electrostatic discharge test constraint requires that an electrostatic discharge test must be completed before the pyrotechnic process is started, and the test signal... Only when The process can only be started when the time is right; personnel qualification constraints encode personnel qualifications as vectors. The qualification levels are represented by five dimensions, and the process requirements are coded as vectors. Only when For all Only upon establishment can the personnel be assigned to perform the procedure.

[0070] All the aforementioned constraints' quantitative indicators and penalty rules are embedded into the safe near-end policy optimization algorithm as hard constraint penalty terms. During offline training, the algorithm optimizes the penalty term coefficients based on the execution samples of various constraints. During online updates, it dynamically adjusts the penalty term weights based on the real-time execution of workshop constraints. When the constraint threshold is adjusted, the algorithm will simultaneously update the penalty term parameters and retrain the network.

[0071] Step S4: Transform the real-time multi-source data obtained in step S1 into a standardized state space representation. The state vector consists of workstation state sub-vectors, logistics equipment state sub-vectors, personnel state sub-vectors, material inventory state sub-vectors, and work order queue state sub-vectors. The state space is encoded as a multidimensional vector, and the workstation state sub-vector is represented as a two-dimensional matrix. Where 2 corresponds to two production lines, For each production line workstation, For each workstation, 4 is the feature dimension, and the matrix element values ​​consist of occupancy status, work order ID, remaining working hours, and personnel ID; the logistics equipment status sub-vector is represented as a matrix. , The number of AGVs is represented by 6, and the feature dimensions include AGV number and location. Coordinates, position Coordinates, load material ID, task status, and power consumption; personnel status sub-vectors are represented as matrices. , The number of personnel is 6, and the feature dimensions include personnel ID and location. Coordinates, position Coordinates, work status code, skill level, current task ID; material inventory status sub-vector is represented as a vector. , The vector represents the number of material types, with each element representing the current quantity of each material in the online edge buffer. Explosive materials are located at the end of the vector. Each location is stored separately; the work order queue status sub-vector is represented as a matrix. , 7 represents the length of the work order queue, and 7 represents the feature dimension, including work order ID, product model, planned quantity, completed quantity, remaining number of processes, delivery date, and priority. The aforementioned standardized multidimensional vector provides standardized input layer features for the policy network and evaluation network. The independent encoding structure of safety-related features such as pyrotechnics and human-machine positions can provide structured support for safety feature gradient calculation and parameter optimization during algorithm updates. Furthermore, during the online update process, the algorithm will adjust the feature weights of the network input layer according to the real-time feature distribution of the state space.

[0072] The data processing steps are as follows: First, perform data cleaning operations, using... The criteria remove outliers, fill in missing values ​​using linear interpolation, and standardize the timestamp format to ISO 8601. Subsequently, continuous numerical features are subjected to Min-Max normalization, with the normalization mapping formula being: Mapping eigenvalues ​​to The time interval is then used to align the multidimensional information from the central control system in the final assembly workshop with time, using the current system time as the reference, and taking the data from each data source within the specified time interval. arrive The latest data within the time window is used to construct a unified state vector; finally, the workstation load rate is extracted. AGV utilization rate Staff work efficiency As auxiliary features, the dataset after the aforementioned processing is divided into training, validation, and test sets in a 7:2:1 ratio to provide high-quality samples for offline training. The real-time data processing results provide noise-free, time-aligned real-time samples for online algorithm updates. During the online update process, the algorithm will also dynamically adjust the extreme value parameters of data normalization and the extraction weights of feature engineering based on the feature distribution of the real-time data.

[0073] Step S5: Define the action space, and the decision output is a composite action vector. The data is parsed into four types of structured planning tables: material delivery instruction set, work order execution instruction set, personnel assignment instruction set, and equipment control instruction set; at any decision-making moment... The model's decision output is a structured action vector, formally represented as: .in, For the action space, For material distribution instruction set, For the work order execution instruction set, Assign instruction sets to personnel. This is the set of equipment control instructions.

[0074] The output is parsed in real time into the following four types of structured plan tables, which are then sent to the corresponding systems in the workshop for execution via standardized interface protocols: Material distribution plan based on Generate. Each record contains fields such as task ID, AGV number, material code, quantity, pickup point, drop-off point, and latest arrival time, which are used to drive the warehousing and logistics system to perform precise delivery.

[0075] Work process execution plan based on Each record clearly identifies the work order number, process number, assigned production line or workstation, and planned start and end times, serving as the basis for the central control system in the final assembly workshop to dispatch production and monitor progress.

[0076] Personnel positioning plan based on Generate. Each record specifies the employee's employee number, task process, target workstation, and task duration, and is sent to the central control system in the final assembly workshop for display on personnel terminals or workshop dashboards.

[0077] Equipment control plan based on Generate. Each record contains the equipment number and control command code, which is sent to the central control system in the final assembly workshop to generate an equipment control plan and then control the equipment to execute.

[0078] Step S6: Construct a multi-evaluation network collaborative optimization mechanism. Construct independent Critic networks for the three optimization objectives respectively. Merge the independent advantage functions into a unified comprehensive advantage signal through a linear weighted fusion function with learnable weights. Update the policy network using the pruning objective function optimized by the proximal strategy. The implementation method of the multi-evaluation network collaborative optimization mechanism is as follows: First, three Critic networks with identical structures but independent parameters are constructed. The input layer dimension of the network structure is... The hidden layers consist of two fully connected layers, each with 256 neurons. The output layer dimension is... The activation function used is ReLU, and the three networks output... , , Each evaluation network independently calculates its advantage function.

[0079] in , For the first An evaluation network computation advantage function; For the first Instant rewards for each optimization goal; For the first The state value function of a Critic network; For the first Parameters of a Critic network; This is the discount factor. Each reward function is defined as follows: This represents the change in completion time (negative values ​​indicate penalties). This represents the change in maximum completion time. For changes in personnel efficiency, Defined as the change in the total distance traveled by people. For personnel load vector, and These are parameter coefficients; This represents the change in production line balance. The workstation load vector is then used; learnable weights are then applied. The linear weighted fusion method yields:

[0080] The constraints are and Weight parameters Update the policy network synchronously via gradient descent; Policy network updates use the PPO pruning objective function:

[0081] in Tailor the objective function for PPO. For policy network parameters; Importance sampling ratio, The action probability output by the policy network. The action probabilities output by the old policy network. To trim hyperparameters, The pruning function; parameters of each Critic network. By minimizing the time-series difference error update, the th The loss function of each Critic network is .

[0082] The aforementioned multi-evaluation network and policy network constitute the core network structure of the secure near-end policy optimization algorithm. The training and updating of the algorithm involves the collaborative updating of this policy network and three independent Critic networks. During offline training, the Adam optimizer is used to perform global optimization of the network parameters, and the learning rate is set. Batch size When updating online, a small-batch incremental update strategy is adopted, adjusting network parameters based on real-time scheduling feedback samples, while maintaining gradient-priority updates for safety constraint penalty terms.

[0083] Step S7: Construct and execute the training and online update process of the Safe Proximity Policy Optimization (Safe-PPO) algorithm. Complete offline training of the algorithm based on historical workshop operation data and sample data generated by the simulation environment. Complete online dynamic update of the algorithm by combining the real-time operation status of the workshop and scheduling decision feedback. Quantify safety constraints into algorithm penalty terms and embed them into the cost function and update mechanism of the algorithm to achieve continuous optimization of the algorithm under safety constraints. The training and online update process of the Safe Proximity Policy Optimization (Safe-PPO) algorithm is implemented as follows: Based on the standard PPO algorithm, the security near-end policy optimization quantifies security constraints as penalty terms and embeds them into the objective function. The core optimization objective is: .

[0084] in:

[0085] For interactive trajectories; This is a weighted reward system for multiple objectives; The cost function is violated due to safety constraints; Security penalty coefficient (offline) (Online adaptive adjustment) For parameterized random policy networks, These are strategy parameters. The safety cost function is composed of a weighted average of various constraint penalty terms: .

[0086] in Pyrotechnics / Explosion-proof / Electrostatic Confinement Weights The remaining constraints are equally distributed among the remaining weights. For constraint categories The penalty function; definitions of various penalty terms: the penalty for process sequence constraints is...

[0087] in Indicate process for Preceding processes; resource quantity constraints and penalties Human-machine safety distance constraint penalty is , For safe distance threshold, The minimum Euclidean distance between the AGV and the personnel.

[0088] The core update formula of the Safe-PPO algorithm is: the enhanced advantage function is... , This is the original multi-objective advantage function. The policy network update uses a pruned objective function to ensure update stability: .in Importance sampling ratio, To trim hyperparameters, The action probability output by the policy network. The action probabilities output by the old policy network; the policy network maximizes... Updated Adam optimizer parameters: , In the formula For the policy network learning rate, , These are the first and second momentum coefficients of the Adam optimizer; The evaluation network update module considers three types of Critic networks and updates by minimizing the TD error: , , Adam optimizer learning rate . Perform multi-objective weight updates The updated normalized weights ensure that the constraints are satisfied. Offline training process: Training dataset. Number of iterations Batch size Update round number Update the policy network: ,in Let the loss function be the policy network. The learning rate of the policy network. Represents the loss function Regarding policy network parameters Gradient of the evaluation network; update the evaluation network: ,in To evaluate the network's loss function, To evaluate the learning rate of the network, Represents the loss function Regarding the first Evaluation network parameters Gradient; Update weights: ,in For the first An adaptive weight for evaluating the network, The learning rate for weight updates. Represents the expectation operator. For the first The target return of an evaluation network These are Lagrange multipliers for safety constraints, used to balance the relationship between reward maximization and safety constraints. For safety cost function; normalized weights Ensure that the sum of the weights of all evaluation networks is 1, and calculate the average reward of the validation set. Constraint violation rate , ;like , Save the optimal strategy parameters .

[0089] Convergence criteria: 1. Average reward of the validation set Rate of change over 200 consecutive rounds ; 2. Constraint violation rate ; 3. Strategy parameter update magnitude .

[0090] The online dynamic update process first collects real-time data, in order to... Data collection status Cache the most recent Step trajectory data, batch size Updates are triggered periodically. The online strategy fine-tuning sets the learning rate to 1 / 10 of the offline phase. Update round number to To avoid overfitting, the evaluation network is updated incrementally only. Real-time safety constraint monitoring and calculation are performed. Record the type and severity of the violation; five consecutive violations of the same constraint: increase and If there are 100 consecutive steps without violations, restore the baseline value. Backup before each update. Rollback occurs when efficiency declines or violation rates increase.

[0091] Algorithm deployment: 1. Encapsulate the model as a RESTful API and deploy it to edge computing nodes; 2. Central control system Call the API once, input Output optimal action ; 3. The action analysis translates into a scheduling plan that is then sent to the central control system in the final assembly workshop; 4. Record key indicators in real time for parameter optimization.

[0092] Step S8: Introduce Model Predictive Control (MPC) framework for rolling optimization, establish a workshop state transition prediction model, solve for the optimal control sequence in the finite time domain at each sampling time, adopt only the first control action, and re-optimize based on the previous control action and state at the next sampling time; The implementation process of the model predictive control framework is as follows: First, establish a workshop state transition model. ; in This is an approximation of the state transition function in a deep neural network. For process noise, based on the current state and control action sequence Predicting the future Step state sequence Subsequently at each sampling time Solving finite-time domain problems The optimal control sequence within the given range, with the optimization objective being to minimize the cumulative cost function:

[0093] in For immediate cost function, For the terminal cost function, As a discount factor; only the first action of the optimal control sequence is issued. At the next sampling time Reacquire the actual state Calculate the prediction error The solution is then optimized again based on the actual conditions; a fixed sampling period is also set. Introduce an event triggering mechanism and define event sets. When the event When it occurs, rescheduling is triggered immediately and is not limited by a fixed sampling period.

[0094] The rolling optimization results and prediction error data of the model predictive control framework are used as core feedback samples for online updates. The algorithm adjusts the weight parameters of the state transition model according to the magnitude of the prediction error, optimizes the policy network output distribution of the algorithm based on the execution effect of the optimal control sequence of rolling optimization, and the rescheduling decision results triggered by events are also stored as incremental samples in the algorithm's experience replay pool for online incremental updates of the algorithm.

[0095] Step S9: Deploy the trained multi-evaluation network collaborative security near-end policy optimization model, receive the current workshop status in real time, output the optimal action sequence and parse it into a collaborative scheduling plan, and send it to the central control system of the final assembly workshop for execution via the MQTT protocol. When abnormal disturbances are detected, the system completes the status update and rescheduling decision within seconds, generates and sends a new scheduling plan, and realizes closed-loop dynamic adaptive scheduling.

[0096] The dynamic response mechanism is implemented in the following ways: First, the workshop status is polled every 30 seconds, and a status change threshold is defined. When any state variable changes An event is determined to have occurred upon receiving an external event notification; after an event occurs, the state space is updated within seconds, specifically including marking the faulty device as unavailable and recording the estimated repair time. Add new orders to the work order queue and set their priority. Update material inventory quantities; within 30 seconds of receiving the updated status, the scheduling model completes the neural network forward calculation and outputs a new action sequence, and completes plan parsing and distribution; simultaneously, it defines a set of switching decision points. ,for The new process is executed according to the new plan. For processes that have already started, the original execution plan is maintained until completion to avoid production fluctuations caused by frequent switching. The event detection data, status update data and execution feedback data of the aforementioned abnormal disturbances are all used as online update incremental samples. The algorithm completes small-batch incremental updates after rescheduling. For high-frequency disturbance events such as emergency order insertion and equipment failure, the algorithm will establish a special sample pool for key training and updates to ensure that the algorithm's decision response capability to similar disturbances is continuously optimized.

[0097] A smart dynamic scheduling system for a spacecraft assembly workshop includes: The data interface module is used to connect to the central control system of the final assembly workshop in real time to obtain real-time multi-source data; The state awareness and construction module is used to construct the real-time multi-source data into a normalized state space vector; The intelligent scheduling decision module has a built-in pre-trained multi-evaluation network collaborative security near-end policy optimization model, which is used to receive the state space vector and output the optimal action vector; The security constraint management module is used to embed various security constraints into the intelligent scheduling decision module and perform security verification during decision-making. The dynamic response module is used to trigger rescheduling when abnormal disturbances are detected, enabling the central control system in the final assembly workshop to complete the workshop status update and rescheduling decision. The plan parsing and distribution module is used to parse the optimal action vector output by the intelligent scheduling decision module into a scheduling plan, and distribute it to the central control system of the final assembly workshop for execution through the data interface module; The algorithm training and update module is used to perform offline training and online updating of the multi-evaluation network collaborative security near-end policy optimization model.

[0098] In one embodiment, the state awareness and construction module receives real-time data and constructs a state space vector, wherein the state vector dimension is... ; The intelligent scheduling decision module incorporates a pre-trained multi-evaluation network collaborative security near-end policy optimization model. The input layer of the policy network structure is... The layer has two hidden layers with 512 neurons each, and the output layer is... Dimensionality, activation functions used are ReLU and Softmax; The safety constraint management module includes a rule-based hard constraint checker, a neural network-based trajectory predictor, a safety shield mechanism, and a safety shield response time. millisecond; The security shield mechanism is implemented as follows: the trajectory predictor uses a lightweight neural network, and the current state is input as input. and candidate actions Output the future Predicted position sequence of AGVs and personnel within seconds: ; The collision detector calculates the minimum distance between any AGV and any person in the predicted trajectory: ; Action arbiter in The action was approved and executed in a timely manner. The action is then rejected and replaced with a predefined safety action. This means that the AGV waits in place until the safe distance is restored.

[0099] The trajectory prediction data, collision detection results, and action adjudication results of the safety shield mechanism are transmitted to the algorithm training and update module in real time. As safety samples for the safety near-end policy optimization algorithm, the algorithm optimizes the decision boundary of safety actions based on these samples during offline training. During online updates, the weight parameters of trajectory prediction and the penalty term coefficient of safety distance are adjusted according to the real-time feedback of these samples. When the safety shield mechanism detects a safety violation risk, it will trigger the algorithm's emergency update mechanism to quickly optimize the safety decision output of the policy network.

[0100] The dynamic response module polls for status changes every 30 seconds, makes rescheduling decisions, and issues plans. The plan parsing and distribution module parses the action vectors into four types of plan tables and distributes them to the corresponding execution units of the central control system through the data interface module.

[0101] The algorithm training and update module includes an offline training unit, an online update unit, a sample management unit, and a parameter optimization unit. The offline training unit completes global training of the algorithm based on historical workshop data and simulation samples. The online update unit receives real-time feedback data from each module to complete incremental updates of the algorithm. The sample management unit classifies, stores, and dynamically filters training samples, validation samples, and real-time feedback samples. The parameter optimization unit dynamically adjusts core parameters of the algorithm, such as the penalty term coefficient and network learning rate, based on the execution of safety constraints and scheduling effects.

[0102] Example 1 I. System Architecture and Workflow The system in this embodiment is primarily deployed on an edge computing server and tightly integrated with the central control system of the aerospace vehicle assembly workshop. The data interface module communicates bidirectionally with the workshop's central control system via the MQTT protocol. Uplink data is obtained by subscribing to the MQTT topic ` / workshop / status / + / +` to retrieve status updates for all devices. A typical device status JSON message format is as follows: {"deviceId": "AGV-01", "timestamp": "2026-01-11T10:30:05Z", "status":"idle", "location": [10.5, 25.2], "battery": 85}. The downlink command publishes a JSON command to the MQTT topic / workshop / logistics / agv01 / task: {"taskId": "T20260111-101", "type":"transport", "source": "CentralStorage-A3", "destination": "Line1-Station2-Buffer", "material": "SKU-X01", "quantity": 10}.

[0103] The state awareness and construction module receives raw data from the data interface module, performs data cleaning, alignment, and fusion, and maps the fused data into standardized state vectors. .

[0104] State vector dimension ; in For each production line workstation, For each workstation, For the number of AGVs, For the number of personnel, For the number of material types, This represents the length of the work order queue.

[0105] The intelligent scheduling decision-making module incorporates a pre-trained multi-evaluation network collaborative security near-end policy optimization model. The policy network structure is the input layer. The innermost layer has two hidden layers, each with 512 neurons, and the output layer has... The three Critic networks have the same structure, using ReLU and Softmax activation functions. The input layer... The system has two layers: a hidden layer with 256 neurons each, and a 1-dimensional output layer. The safety constraint management module includes a rule-based hard constraint checker, a neural network-based trajectory predictor, and a safety shield mechanism. The trajectory predictor uses a lightweight neural network, with the input being... The two layers, the main layer and the hidden layer, each have 128 neurons, and the output is... Dimension (10-step prediction, each step) AGV and (x, y coordinates of each person). The security shield response time is less than 50 milliseconds.

[0106] The dynamic response module polls for status changes every 30 seconds to complete rescheduling decisions and plan issuance.

[0107] II. Model Training The model was trained offline. First, a high-fidelity digital twin simulation environment was built, which included a complete process model, equipment operation model, AGV motion model, and personnel behavior model of the aerospace vehicle assembly workshop, which could accurately simulate the dynamic operation process of the workshop.

[0108] In the simulation environment, a FIFO scheduling strategy is run to collect initial trajectory data, forming an experience playback pool with a capacity of [missing information]. Transfer samples . The neural network model was trained using the collected sample data. The policy network and each Critic network employed the Adam optimizer with a learning rate of [missing information]. Batch size PPO clipping hyperparameters Discount factor GAE parameters Training iterations The network parameters are updated every 2048 steps. The network parameter update process is as follows: Policy network parameters renew:

[0109] in For learning rate, This is to prune the gradient of the objective function with respect to the policy parameters.

[0110] Critic network parameters renew:

[0111] Learnable weights renew:

[0112] The constraints are and The constraints are satisfied by using the projected gradient descent method.

[0113] Once the model converges on the training set and performs well in independent test scenarios, it is deployed to the actual scheduling system. Initial deployment uses a shadow mode, where the model only makes decisions and records data but does not actually execute them; human decision-making is used for comparison and verification. After the model's performance stabilizes, it is switched to a fully autonomous decision-making mode.

[0114] III. Safety Constraint Handling The core of the security constraint management module is the security shield mechanism, the implementation details of which are as follows: The trajectory predictor uses a lightweight neural network, taking the current state as input. and candidate actions Output the future Predicted position sequences of AGVs and personnel within seconds. The network structure is an input layer. Two layers of 128 neurons each: a hidden layer and an output layer. dimension.

[0115] The collision detector calculates the minimum distance between any AGV and any person in the predicted trajectory: .

[0116] Action arbiter in The action was approved and executed in a timely manner. The action is rejected and replaced with a predefined safety action, i.e., the AGV waits in place until the safe distance is restored.

[0117] Safety restraints for pyrotechnic devices: at any time Total inventory of pyrotechnic products at the production line satisfy Item. When When an alert is triggered, New delivery tasks for this material are prohibited during this period. Assembly processes involving pyrotechnics must be performed at explosion-proof stations. No other tasks may be assigned to these stations while pyrotechnic processes are being performed. The station status must be marked as locked until the process is completed and a safety inspection is passed. Static electricity elimination testing is required before starting any pyrotechnic process. The detection signal... Only when The process can only begin when the time is right.

[0118] IV. Model Predictive Control Rolling Optimization The specific implementation of MPC rolling optimization is as follows: at each sampling time... Based on the current state Solve the following optimization problem:

[0119] Constraints: , .

[0120] in For state The set of possible actions is determined by both safety and process constraints. A rolling time-domain optimization strategy is adopted. 1. At any time Solving the above optimization problem yields the optimal control sequence. ; 2. Only the first control action is adopted. ; 3. At that moment Obtain the actual state ; 4. Calculate the prediction error ; 5. Based on actual conditions Resolve the optimization problem and set a fixed sampling period. Seconds, and an event triggering mechanism is introduced.

[0121] Define event collection When the event When this occurs, a rescheduling is triggered immediately, without being limited by a fixed sampling period.

[0122] V. Dynamic Response and Rescheduling The specific implementation of the dynamic response mechanism is as follows: The event detection module polls the workshop status every 30 seconds and defines a threshold for status changes. (Normalized state values), when any state variable changes Alternatively, it can determine that an event has occurred upon receiving an external event notification.

[0123] The fast state update module completes the state space update within 3 seconds of the event occurring. 1. Mark the faulty device as unavailable and record the estimated repair time. ; 2. Add the new order to the work order queue and set its priority. ,in Prioritize orders based on their fundamental priority. This represents the change in the urgency of the order. Within 3 seconds of receiving the updated status, the rescheduling decision module completes the forward computation of the neural network, outputs a new action sequence, and completes the plan parsing and distribution.

[0124] The smooth handover strategy module defines the set of handover decision points. ,for The new plan will be followed for the new process, while the original plan will be maintained for processes that have already started until completion, in order to avoid production fluctuations caused by frequent switching.

[0125] When an abnormal disturbance is detected, the system completes the state update and rescheduling decision within 3 seconds, generates and issues a new scheduling plan, and realizes closed-loop dynamic adaptive scheduling.

[0126] VI. Online Learning and Continuous Optimization The algorithm training and update module includes an offline training unit, an online update unit, a sample management unit, and a parameter optimization unit.

[0127] The offline training unit completes global training of the algorithm based on historical workshop data and simulation samples. The training dataset is divided into training set, validation set, and test set in a 7:2:1 ratio.

[0128] The online update unit receives real-time feedback data from each module to complete the incremental update of the algorithm. The update strategy is mini-batch incremental update, with a batch size of 32 and a learning rate of 1 / 10 of that used in offline training.

[0129] The sample management unit implements the classification, storage, and dynamic filtering of training samples, validation samples, and real-time feedback samples. The experience replay pool has a capacity of [missing information]. For each sample, a priority experience playback mechanism is used, prioritizing the sampling of samples with larger TD errors.

[0130] The parameter optimization unit dynamically adjusts core parameters of the algorithm, such as the penalty term coefficient and the network learning rate, based on the execution of safety constraints and scheduling effects. The adjustment rule for the penalty term coefficient is as follows: ; in To adjust the step size, To control the violation rate, The target violation rate (usually set to 0).

[0131] Through the detailed implementation methods described above, this invention constructs a closed-loop, self-learning, adaptive, and inherently safe intelligent scheduling system, which can lead the spacecraft assembly workshop towards a higher level of intelligent and flexible production.

[0132] Those skilled in the art will understand that, besides implementing the system and its various devices, modules, and units provided by this invention in the form of purely computer-readable program code, the same functions can be achieved entirely through logical programming of the method steps, making the system and its various devices, modules, and units of this invention function in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, the system and its various devices, modules, and units provided by this invention can be considered as a hardware component, and the devices, modules, and units included therein for implementing various functions can also be considered as structures within the hardware component; alternatively, the devices, modules, and units for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0133] Specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Unless otherwise specified, the embodiments and features described in this application can be arbitrarily combined with each other.

Claims

1. A method for intelligent dynamic scheduling in a spacecraft assembly workshop, characterized in that, include: Step S1: Obtain real-time multi-source data by connecting to the central control system of the final assembly workshop; The real-time multi-source data includes static basic data and dynamic data; Step S2: Construct a workshop state training set based on the real-time multi-source data and train a multi-evaluation network collaborative safety near-end strategy optimization model accordingly; the training process is as follows: solve for the optimal control sequence in the finite time domain at each sampling time, adopt only the first control action, and re-optimize based on the previous control action and state at the next sampling time. After completing the optimization of all control actions, perform reverse update of the neural network based on the reward model; repeat the training process until convergence; Step S3: Deploy the trained multi-evaluation network collaborative security near-end strategy optimization model. When abnormal disturbances are detected, input real-time information into the central control system of the final assembly workshop and trigger the rescheduling mechanism to complete the workshop status update and rescheduling decision, and generate and issue a new scheduling plan.

2. The intelligent dynamic scheduling method for aerospace vehicle assembly workshop according to claim 1, characterized in that, Step S2 includes the following sub-steps: Step S2.1: Construct a set of production scheduling objective functions containing three optimization objectives; The goal is to minimize the maximum completion time of all pending work orders; The objective is to minimize the total distance traveled by personnel and the variance of workload among personnel. The objective is to minimize the variance of the planned working hours load for each workstation; The formulas are expressed as follows: in , This is a collection of all work orders. For work orders Completion time; For the gathering of all operators, For scheduling periods, For personnel During the period Within the mission's movement distance, For personnel Total effective work hours within the scheduling cycle Represents variance. and These are the weighting coefficients; For the set of all workstations, To schedule workstations within the field of view The sum of standard working hours for all assigned processes; Step S2.2: Define process sequence constraints, resource quantity constraints, equipment and workstation exclusivity constraints, human-machine safety distance constraints, explosive buffer capacity constraints, and personnel qualification constraints. Model the scheduling problem as a constrained Markov decision process, and transform various safety constraints into safety cost terms of the safety near-end strategy optimization algorithm, embedding them into the algorithm's immediate cost function to construct a decision optimization problem with safety constraints. Step S2.3: Transform the acquired real-time multi-source data into a standardized state space representation; the state vector includes workstation state sub-vectors, logistics equipment state sub-vectors, personnel state sub-vectors, material inventory state sub-vectors, and work order queue state sub-vectors; Step S2.4: Define the action space, and the decision output is a composite action vector. The plan is parsed into four types of structured plans: material delivery instruction set, work order execution instruction set, personnel assignment instruction set, and equipment control instruction set; at any decision-making moment... The model's decision output is a structured action vector, formally represented as: in, For the action space, For material distribution instruction set, For the work order execution instruction set, Assign instruction sets to personnel. For equipment control instruction set; Step S2.5: Construct a multi-evaluation network collaborative optimization mechanism. Construct independent Critic networks for the three optimization objectives respectively. Merge the independent advantage functions into a unified comprehensive advantage signal through a linear weighted fusion function with learnable weights. Update the policy network using the pruning objective function optimized by the near-end strategy. Step S2.6: Construct and execute the training and online update process of the safety near-end policy optimization algorithm. Complete the offline training of the algorithm based on the historical operation data of the workshop and the sample data generated by the simulation environment. Combine the real-time operation status of the workshop and the scheduling decision feedback to complete the online dynamic update of the algorithm. Quantify the safety constraints into algorithm penalty terms and embed them into the cost function and update mechanism of the algorithm to realize the continuous optimization of the algorithm under safety constraints.

3. The intelligent dynamic scheduling method for aerospace vehicle assembly workshop according to claim 2, characterized in that, The process sequence constraint requires that each product process strictly follow the preset process route. The subsequent process cannot start until the preceding process is completed. Furthermore, a certain period of waiting is set between critical assembly processes and inspection processes; that is, for any work order… , It is a set of work orders, and the processes contained therein must be executed according to a preset partial order relationship; set up For process The direct successor process must satisfy the following: ; in, and Each is a process The start and end times, Allow time for necessary inter-process preparation or material transfer. For work orders The set of partial order relations of the process steps; actions that violate this constraint will be considered infeasible. Resource quantity constraints limit the number of AGVs configured on each production line to no more than The number of units and operators shall not be less than Each AGV has a rated load, and the number of tooling fixtures at each workstation matches the requirements of the corresponding process, meaning that at any decision-making moment... The total amount of any type of resource allocated to a task must not exceed its current available amount, as expressed by the formula: in, for The set of all selected scheduling actions at any given time. ( ) is an action Resources The demand, For resources exist Total available time; Equipment and workstation exclusivity constraints require that high-precision testing equipment cannot be occupied by other work orders while performing testing tasks, and that the same operator can only be assigned to one workstation at a time. and Each represents a process. In the work order Start time and processing time For 0-1 decision variables, when the work order Assigned to work station The value is 1 if the condition is met, and 0 otherwise; for any two different work orders If they might use the same workstation Then the following mutual exclusion conditions must be met: in, For a sufficiently large positive number, This is a sequence indicator variable used to ensure that two work orders are at the workstation. The processing sequences on each part do not overlap; Safety constraints include human-machine safety distance constraints, explosive buffer capacity constraints, exclusive use of explosion-proof workstations constraints, electrostatic discharge detection constraints, and personnel qualification constraints. Human-machine safety distance constraint definition real-time safety cost function: in For a moment The safety cost function; For a moment The condition of the workshop; For a moment The action; To preset a safe distance threshold, To perform the action The predicted minimum Euclidean distance between all AGVs and all personnel is calculated using the following formula: ,when The action was refused; In the formula For AGV Position coordinates; For personnel Position coordinates; The buffer capacity constraint for pyrotechnic materials limits the maximum number of line-side buffers for flammable and explosive materials. At any time Total inventory of pyrotechnic products at the production line satisfy ,when When an alert is triggered, New delivery orders for this material are prohibited at this time; The exclusivity constraint of the explosion-proof workstation requires that the assembly process involving pyrotechnics must be performed at the explosion-proof workstation. During the pyrotechnics process, no other tasks may be assigned to this workstation. The workstation status is marked as locked until the process is completed and a safety inspection is passed. Electrostatic discharge (ESD) detection requirements stipulate that ESD detection must be completed before the start of the pyrotechnics process, and the detection signal... Only when Only when the time is right can the process be started; Personnel qualification constraints encode personnel qualifications as vectors The qualification levels are represented by five dimensions, and the process requirements are coded as vectors. Only when For all Only upon establishment can the personnel be assigned to perform the procedure.

4. The intelligent dynamic scheduling method for aerospace vehicle assembly workshop according to claim 2, characterized in that, Step S2.5 includes: Construct three Critic networks with identical structures but independent parameters. The input layer dimension of the network structure is... The hidden layers consist of two fully connected layers, each with 256 neurons. The output layer dimension is... The activation function used is ReLU, and the three networks output... , , Each evaluation network independently computes its advantage function. in , For the first An evaluation network computation advantage function; For the first Instant rewards for each optimization goal; For the first The state value function of a Critic network; For the first Parameters of a Critic network; Discount factor; Each reward function is defined as follows: For the change in completion time, This represents the change in maximum completion time. For changes in personnel efficiency, Defined as the change in the total distance traveled by people. For personnel load vector, and These are parameter coefficients; This represents the change in production line balance. This represents the workstation load vector. Employing learnable weights The linear weighted fusion method is obtained The constraints are and Weight parameters Update the policy network synchronously via gradient descent; Policy network updates use PPO pruning objective function in Tailor the objective function for PPO. For policy network parameters; The importance sampling ratio, The action probability output by the policy network. The action probabilities output by the old policy network. To trim hyperparameters, The pruning function; parameters of each Critic network. By minimizing the time-series difference error update, the th Loss function of a Critic network for: 。 5. The intelligent dynamic scheduling method for aerospace vehicle assembly workshop according to claim 2, characterized in that, The process of building and executing the training and online update of the safe near-end policy optimization algorithm includes: The security near-end policy optimization is based on the standard PPO algorithm, which quantifies security constraints as penalty terms and embeds them into the objective function. The core optimization objective is: in, For interactive trajectories; This is a weighted reward system for multiple objectives; Safety constraints violate the cost function; This is a safety penalty coefficient; For parameterized random policy networks, For strategy parameters; The safety cost function is composed of a weighted average of various constraint penalty terms: ; in Pyrotechnics / Explosion-proof / Electrostatic Confinement Weights The remaining constraints are equally distributed among the remaining weights. For constraint categories The penalty function; The definitions of various penalty items are as follows: The penalty for process sequence constraints is: in Indicate process for The preceding process; The resource quantity constraint penalty is: The penalty for human-machine safe distance constraints is: The minimum Euclidean distance between the AGV and the personnel; The core update formula of the Safe-PPO algorithm: The enhanced advantage function is ; , which is the original multi-objective advantage function; The policy network update uses a pruned objective function to ensure update stability. in The importance sampling ratio, To trim hyperparameters, The action probability output by the policy network. The probability of the action output by the old policy network; Policy networks maximize Updated Adam optimizer parameters: , ; In the formula For the policy network learning rate, , These are the first and second momentum coefficients of the Adam optimizer; The evaluation network update module considers three types of Critic networks and updates by minimizing the TD error: in, , Adam optimizer learning rate ; Perform multi-objective weight updates The updated normalized weights ensure that the constraints are satisfied; the offline training process trains the dataset. Number of iterations Batch size Update round number ; Update policy network: ,in Let the loss function be the policy network. The learning rate of the policy network. Represents the loss function Regarding policy network parameters The gradient; Updated evaluation network: ,in To evaluate the network's loss function, To evaluate the learning rate of the network, Represents the loss function Regarding the first Evaluation network parameters The gradient; Update weights: ,in For the first An adaptive weight for evaluating the network, The learning rate for weight updates. Represents the expectation operator. For the first The target return of an evaluation network These are Lagrange multipliers for safety constraints, used to balance the relationship between reward maximization and safety constraints. For safety cost function; Normalized weights Ensure that the sum of the weights of all evaluation networks is 1; Calculate the average reward on the validation set Constraint violation rate , ;like , Save the optimal strategy parameters .

6. The intelligent dynamic scheduling method for aerospace vehicle assembly workshop according to claim 5, characterized in that, Convergence criteria include: Average reward of validation set Rate of change over 200 consecutive rounds ; constraint violation rate ; Strategy parameter update magnitude .

7. The intelligent dynamic scheduling method for aerospace vehicle assembly workshop according to claim 5, characterized in that, The online update process in step S2.6 includes: Collect real-time data, in order to Data collection status Cache the most recent Step trajectory data, batch size Updates are triggered periodically; the online strategy fine-tuning sets the learning rate to 1 / 10 of the offline stage. Update round number to To avoid overfitting; the evaluation network is updated incrementally only; real-time safety constraint monitoring and calculation. Record the type and severity of the violation; five consecutive violations of the same constraint: increase and If there are no violations for 100 consecutive steps, restore the baseline value; back up before each update. Rollback occurs when efficiency declines or violation rates increase.

8. The intelligent dynamic scheduling method for aerospace vehicle assembly workshop according to claim 1, characterized in that, Step S2 further includes introducing a model predictive control framework for rolling optimization and establishing a workshop state transition model. in This is an approximation of the state transition function in a deep neural network. For process noise, based on the current state and control action sequence Predicting the future Step state sequence Subsequently at each sampling time Solving finite-time domain problems The optimal control sequence within the range, with the optimization objective being to minimize the cumulative cost function. in For immediate cost function, For the terminal cost function, As a discount factor; only the first action of the optimal control sequence is adopted. At the next sampling time Reacquire the actual state Calculate the prediction error The solution is then optimized again based on the actual conditions; a fixed sampling period is also set. Introduce an event triggering mechanism and define an event set: When the event When it occurs, rescheduling is triggered immediately and is not limited by a fixed sampling period.

9. The intelligent dynamic scheduling method for aerospace vehicle assembly workshop according to claim 1, characterized in that, Step S3 includes: The workshop status is polled periodically over a preset time period, and a threshold for status change is defined. When any state variable changes The system determines an event has occurred upon receiving an external event notification; after the event occurs, it completes a state space update within seconds; the scheduling model completes the neural network forward computation and outputs a new action sequence within a preset time period after receiving the updated state, and completes plan parsing and distribution; simultaneously, it defines a set of switching decision points. ,for The new process is executed according to the new plan, and the original execution plan is maintained for processes that have already started until completion; the event detection data, state update data and execution feedback data of the rescheduling decision mentioned above are all used as online update incremental samples, and the algorithm completes small batch incremental updates after rescheduling; a special sample pool is established for high-frequency disturbance events for focused training and updating.

10. An intelligent dynamic scheduling system for a spacecraft assembly workshop, characterized in that, include: The data interface module is used to connect to the central control system of the final assembly workshop in real time to obtain real-time multi-source data; The state awareness and construction module is used to construct the real-time multi-source data into a normalized state space vector; The intelligent scheduling decision module has a built-in pre-trained multi-evaluation network collaborative security near-end policy optimization model, which is used to receive the state space vector and output the optimal action vector; The security constraint management module is used to embed various security constraints into the intelligent scheduling decision module and perform security verification during decision-making. The dynamic response module is used to trigger rescheduling when abnormal disturbances are detected, enabling the central control system in the final assembly workshop to complete the workshop status update and rescheduling decision. The plan parsing and distribution module is used to parse the optimal action vector output by the intelligent scheduling decision module into a scheduling plan, and distribute it to the central control system of the final assembly workshop for execution through the data interface module; The algorithm training and update module is used to perform offline training and online updating of the multi-evaluation network collaborative security near-end policy optimization model.