A multi-device linkage intelligent scheduling method and system for an automobile disassembly production line

By acquiring multi-source data from the automobile dismantling production line and using particle swarm optimization and deep reinforcement learning models to optimize equipment scheduling, the problem of equipment load imbalance was solved, production efficiency and resource utilization were improved, and physical trial and error and resource waste were avoided.

CN121787871BActive Publication Date: 2026-05-08JIANGSU SUBEI FEIJIU CAR HOME APPLIANCES DISMANTLING REGENERAT
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JIANGSU SUBEI FEIJIU CAR HOME APPLIANCES DISMANTLING REGENERAT
Filing Date
2026-03-09
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

In dynamic operation scenarios, the load imbalance of multiple devices on the automobile dismantling production line leads to both local congestion and resource idleness.

Method used

By acquiring multi-source operational data, a task scheduling priority matrix is ​​constructed using the particle swarm optimization algorithm. Combined with a deep reinforcement learning model, policy interaction is performed in a virtual simulation environment to generate load adjustment parameters, thereby achieving dynamic scheduling and resource optimization among devices.

Benefits of technology

It effectively solved the problem of imbalance between equipment, improved the average output efficiency of the production line, avoided downtime accidents caused by physical trial and error, made full use of the redundant capacity of the entire line, and ensured the continuity and smoothness of material flow.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121787871B_ABST
    Figure CN121787871B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of intelligent manufacturing, and discloses a kind of automobile disassembly production line multi-device linkage intelligent scheduling method and system.The method comprises the following steps: obtaining multi-source operation data to obtain equipment load distribution matrix and material flow vector; according to the equipment load distribution matrix and the material flow vector, a particle swarm algorithm is used to obtain a task scheduling priority matrix and determine a linkage bottleneck area; a deep reinforcement learning model is used to output a task allocation scheme in a virtual simulation environment; according to the task allocation scheme, determine the congestion area and generate the path correction vector; schedule available alternative equipment to perform load transfer to obtain load adjustment parameters; calculate the progress deviation vector according to the load adjustment parameters and update the operation parameters of multi-device cooperation. This method can solve the cascading imbalance problem caused by the difference in the capabilities of heterogeneous devices in the disassembly line.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent manufacturing technology, and in particular to a method and system for intelligent scheduling of multiple devices in an automobile dismantling production line. Background Technology

[0002] Currently, mainstream automotive dismantling production lines typically rely on traditional industrial control systems for operation management. In one existing technology, the production line employs a static scheduling mode based on preset cycle times. This means that the central controller sets fixed operating speeds and start / stop logic for heterogeneous equipment such as crushers, magnetic separators, and non-ferrous metal sorters based on theoretical average working hours. Each piece of equipment operates relatively independently, connected only by a simple serial connection via physical conveyor belts. However, when faced with non-standard input scenarios such as significant differences in the condition of scrapped vehicles and varying degrees of corrosion in parts, this static scheduling mode struggles to cope with random fluctuations in process times. When the processing time of a particular process increases or decreases, the fixed transmission cycle time cannot dynamically adapt, causing a disconnect between the logistics status within the production line and the equipment's operational capacity.

[0003] Existing technologies present a technical problem where the load of multiple devices in a car dismantling production line becomes unbalanced under dynamic operating scenarios, leading to both localized congestion and resource idleness. Summary of the Invention

[0004] This invention provides a method and system for intelligent scheduling of multiple devices in an automotive dismantling production line, in order to solve the technical problem in the prior art where the load linkage of multiple devices in an automotive dismantling production line is unbalanced under dynamic operation scenarios, resulting in both local congestion and resource idleness.

[0005] Firstly, in order to solve the above-mentioned technical problems, the present invention provides a multi-equipment linkage intelligent scheduling method for an automobile dismantling production line, comprising:

[0006] Acquire multi-source operational data from the automobile dismantling production line and preprocess it to obtain the equipment load distribution matrix and material flow vector;

[0007] Based on the equipment load distribution matrix and the material flow vector, the particle swarm optimization algorithm is used to iteratively optimize and obtain the task scheduling priority matrix; based on the task scheduling priority matrix and the preset processing capacity threshold, the blocking nodes are determined, and based on the blocking nodes and the preset theoretical throughput of the equipment, the linkage bottleneck area and the severity of the linkage bottleneck area are determined.

[0008] If the severity exceeds a preset safety threshold, a virtual simulation environment is constructed based on the linkage bottleneck area; a pre-trained deep reinforcement learning model is used to perform policy interaction and search in the virtual simulation environment, and a task allocation scheme is output.

[0009] According to the task allocation scheme, the material occupancy rate within the preset workshop space topology is calculated. If the material occupancy rate exceeds the preset congestion threshold, the congested area is determined based on the spatial distribution of the material occupancy rate, and a path correction vector for the congested area is generated.

[0010] Available alternative devices located outside the congested area are dispatched to take over the pending tasks within the congested area, perform load transfer, and obtain load adjustment parameters;

[0011] Based on the load adjustment parameters, a progress deviation vector is calculated using a preset production progress simulation model, and the operating parameters of multi-device collaboration are updated based on the progress deviation vector.

[0012] Secondly, the present invention provides a multi-equipment linkage intelligent scheduling system for an automobile dismantling production line, comprising:

[0013] The data processing module is used to acquire multi-source operating data of the automobile dismantling production line and perform preprocessing to obtain the equipment load distribution matrix and material flow vector.

[0014] The bottleneck identification module is used to perform iterative optimization using the particle swarm optimization algorithm based on the equipment load distribution matrix and the material flow vector to obtain a task scheduling priority matrix; to determine the blocking nodes based on the task scheduling priority matrix and a preset processing capacity threshold; and to determine the linkage bottleneck area and the severity of the linkage bottleneck area based on the blocking nodes and a preset theoretical equipment throughput.

[0015] The strategy optimization module is used to construct a virtual simulation environment based on the linkage bottleneck area if the severity exceeds a preset safety threshold; and to perform strategy interaction and search in the virtual simulation environment using a pre-trained deep reinforcement learning model to output a task allocation scheme.

[0016] The path monitoring module is used to calculate the material occupancy rate within the preset workshop space topology according to the task allocation scheme. If the material occupancy rate exceeds the preset congestion threshold, the congested area is determined according to the spatial distribution of the material occupancy rate, and a path correction vector for the congested area is generated.

[0017] The load transfer module is used to schedule available alternative devices located outside the congested area to take over the pending tasks within the congested area, perform load transfer, and obtain load adjustment parameters;

[0018] The feedback control module is used to calculate the progress deviation vector using a preset production progress simulation model based on the load adjustment parameters, and update the multi-device collaborative operation parameters based on the progress deviation vector.

[0019] Compared with the prior art, the present invention has the following beneficial effects:

[0020] (1) This invention effectively solves the problem of linkage discontinuity caused by inconsistent data dimensions and static cycle control among heterogeneous equipment by performing variance-based weighted fusion of multi-source operating data and constructing a task scheduling priority matrix using particle swarm optimization. Traditional systems often rely solely on a single current or counting signal for control, which is insufficient to reflect the actual load of the equipment. This solution firstly uses variance-weighted fusion of multi-dimensional features such as current and vibration to accurately quantify the real-time health and load status of the equipment; then, it uses particle swarm optimization to iteratively search for the optimal task scheduling priority at the global level, rather than adhering to a preset process sequence. This enables the production line to dynamically adjust the task allocation weights of upstream and downstream based on the real-time processing capacity of the equipment, realizing the transformation from fixed-cycle production to capacity-driven production and reducing the risk of cascading blockage caused by sudden high loads upstream.

[0021] (2) This invention overcomes the technical bottleneck of traditional rule-based scheduling in dealing with random disturbances of non-standard materials by constructing a virtual simulation environment containing virtual models and using deep reinforcement learning models for policy interaction and search. The automobile dismantling process has high uncertainty (such as rust jamming and irregular parts), and the cost of trial and error on the physical production line is extremely high. This solution introduces a virtual simulation environment as a sandbox and uses the powerful high-dimensional state perception and long-term reward prediction capabilities of deep reinforcement learning models to pre-rehearse and evaluate actions such as adjusting task priorities and modifying transmission rhythm in the virtual space. Only the solutions that generate positive rewards (increased output and reduced delay) are issued for execution. This ensures the safety and global optimality of the scheduling strategy, avoids downtime accidents that may be caused by physical trial and error, and realizes adaptive intelligent decision-making for complex dynamic scenarios, which greatly improves the average output efficiency of the production line.

[0022] (3) This invention generates a path correction vector by calculating the material occupancy rate within the workshop space topology, and schedules equipment in non-congested areas to perform load transfer accordingly, thus achieving deep decoupling and utilization of resources based on physical space constraints. Existing scheduling methods are mostly limited to adjusting single equipment parameters (such as simple deceleration), which often cannot solve the deadlock in physical space. This solution maps abstract load data back to the physical coordinate system of the workshop, intuitively delineates congested hot zones, and generates specific path correction vectors; furthermore, it actively awakens idle resources (available alternative equipment) in non-congested areas to share the flow in congested areas. This space-for-time strategy not only quickly clears local bottlenecks, but also makes full use of the redundant capacity of the entire production line, effectively eliminating the waste of resources caused by local overload and global idleness, and ensuring the continuity and smoothness of material flow at the physical space level. Attached Figure Description

[0023] Figure 1 This is a schematic diagram of the intelligent scheduling method for multi-equipment linkage in an automobile dismantling production line provided in the first embodiment of the present invention;

[0024] Figure 2 This is a schematic diagram of the structure of the intelligent scheduling system for multi-equipment linkage in an automobile dismantling production line provided in the second embodiment of the present invention. Detailed Implementation

[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0026] Reference Figure 1 The first embodiment of the present invention provides a method for intelligent scheduling of multiple devices in an automobile dismantling production line, comprising the following steps:

[0027] S11: Obtain multi-source operating data of the automobile dismantling production line and perform preprocessing to obtain the equipment load distribution matrix and material flow vector;

[0028] S12, Based on the equipment load distribution matrix and the material flow vector, the particle swarm optimization algorithm is used to iteratively optimize and obtain the task scheduling priority matrix; Based on the task scheduling priority matrix and the preset processing capacity threshold, the blocking nodes are determined, and the linkage bottleneck area and the severity of the linkage bottleneck area are determined based on the blocking nodes and the preset theoretical throughput of the equipment.

[0029] S13, if the severity exceeds a preset safety threshold, then a virtual simulation environment is constructed based on the linkage bottleneck area; using a pre-trained deep reinforcement learning model, strategy interaction and search are performed in the virtual simulation environment, and a task allocation scheme is output.

[0030] S14. According to the task allocation scheme, calculate the material occupancy rate in the preset workshop space topology. If the material occupancy rate exceeds the preset congestion threshold, determine the congested area according to the spatial distribution of the material occupancy rate, and generate a path correction vector for the congested area.

[0031] S15, dispatch available alternative equipment located outside the congested area to take over the pending tasks within the congested area, perform load transfer, and obtain load adjustment parameters;

[0032] S16, based on the load adjustment parameters, calculate the progress deviation vector using a preset production progress simulation model, and update the multi-device collaborative operation parameters based on the progress deviation vector.

[0033] In step S11, multi-source operating data of the automobile dismantling production line is acquired and preprocessed to obtain the equipment load distribution matrix and material flow vector, including:

[0034] Extract the timestamps of the multi-source running data, and use a linear interpolation algorithm to align data streams of different frequencies to a preset unified time axis;

[0035] Calculate the numerical variance of each aligned data stream within a preset sliding window;

[0036] Calculate the reciprocal of the numerical variance and normalize the reciprocal to obtain the fusion weight;

[0037] The multi-source operating data of each device are weighted and summed according to the index dimensions using the fusion weights to generate the device load distribution matrix that represents the real-time load status of each device.

[0038] In one implementation, the multi-source operational data originates from a heterogeneous sensor array deployed on various key pieces of equipment (such as crushers, magnetic separators, and sorters) of the dismantling production line. For any given piece of equipment, the collected data includes, but is not limited to, the real-time current value of the drive motor, the vibration acceleration amplitude of the equipment body, and the temperature value of key components. Since the sampling frequencies of different sensors differ significantly, direct fusion would cause timing misalignment. Therefore, this embodiment first establishes a preset unified timeline.

[0039] It should be noted that the sampling interval of the preset unified time axis is not subjectively set, but determined based on the sampling theorem. Specifically, the system traverses the raw data streams of all sensor data sources (such as temperature data), and this embodiment uses a linear interpolation algorithm for alignment processing. For any target time on the unified time axis, the nearest sampling point before and after that time is located in the low-frequency data stream; the time distance between the target time and these two sampling points is calculated, and the reciprocal of the time distance is used as a weight to perform a weighted average of the values ​​of the two sampling points, thereby obtaining the interpolation result for the target time.

[0040] In one implementation, to evaluate the stability of sensor data in characterizing device load, this embodiment calculates the numerical variance of each data stream within a preset sliding window. It is worth noting that the length of the preset sliding window is determined through autocorrelation analysis of historical data. A historical data sequence from stable device operation is selected, and the autocorrelation function of the sequence is calculated. The lag time corresponding to the autocorrelation coefficient decreasing to the reciprocal of the natural constant (approximately 0.368) is identified, and this lag time is determined as the length of the sliding window. This method ensures that the window covers the typical fluctuation cycle of the device's operating state. After determining the window, the system calculates the sum of squares of the differences between all data points within the window and the average value, and then divides this sum by the number of data points to obtain the numerical variance.

[0041] In one implementation, this embodiment uses a reciprocal method to calculate the fusion weights, automatically reducing the impact weight of highly volatile data (which typically contains high noise). The reciprocal of the variance calculated above is taken as the original weight. To avoid calculation errors caused by a zero variance, a very small positive number (e.g., one millionth) is added to the variance before calculating the reciprocal. Subsequently, normalization processing is performed, dividing the original weight of a certain indicator of a certain device by the sum of the original weights of all indicators of that device to obtain the final fusion weight of that indicator.

[0042] For example, suppose the crusher has two indicators: current and vibration. If, within the current window, the variance of the current data is smaller, its reciprocal, calculated as its original weight, is 0.8; and the variance of the vibration data is larger, its reciprocal, calculated as its original weight, is 0.2. The sum of the two is 1.0. In this case, the normalized fusion weight of the current indicator is 0.8, and the weight of the vibration indicator is 0.2. This indicates that in subsequent calculations, the system will rely more on the more stable current indicator.

[0043] In one implementation, the equipment load distribution matrix is ​​generated using the fusion weights. First, the data for each indicator is standardized (i.e., the mean is subtracted and then divided by the standard deviation) to eliminate differences in different physical units (such as amperes and hertz). Next, for each piece of equipment, the standardized values ​​of each indicator are multiplied by their corresponding fusion weights, and all products are summed to obtain the single fusion load value for that equipment at the current moment. This process is repeated for all equipment on the production line, and the real-time fusion load values ​​of all equipment are arranged in order of equipment number, forming a column vector. This column vector is the equipment load distribution matrix. This matrix visually reflects the real-time pressure distribution of each physical node on the entire production line.

[0044] It should be noted that the material flow vector is used to quantitatively describe the rate and direction of material flow between each process, and serves as the spatial basis for subsequent bottleneck identification and path planning. In this embodiment, the material flow vector is collected in real time by photoelectric encoders and RFID readers deployed on each conveyor belt, diverter, and automated guided vehicle to collect the quantity of material passing through each monitoring point and the corresponding task identifier within a unit of time. Based on the workshop equipment topology diagram, a directed graph model is established, where nodes represent equipment or buffer zones, and directed edges represent material transmission paths. Finally, the quantity of material passing through each edge in the directed graph within a preset sampling period is arranged into a column vector according to the equipment number, and this vector is aligned with the equipment load distribution matrix on the time axis to obtain the material flow vector. Positive values ​​in this vector indicate forward material flow, while negative values ​​(which can be determined by the direction sensor) indicate backflow or reverse flow abnormalities.

[0045] In step S12, based on the equipment load distribution matrix and the material flow vector, the particle swarm optimization algorithm is used for iterative optimization to obtain the task scheduling priority matrix, including:

[0046] Construct a particle swarm, defining the position of each particle as a potential task scheduling priority matrix;

[0047] Define a fitness function that is inversely proportional to the average turnaround time of the entire production line and inversely proportional to the load variance of each equipment node;

[0048] The velocity and position of the particle are updated according to the fitness function until a preset number of iterations is reached or a preset convergence condition is met, and the position of the optimal particle is determined as the task scheduling priority matrix.

[0049] In one implementation, this embodiment first constructs a particle swarm, where each particle represents a possible scheduling strategy. Specifically, the position of each particle is encoded as a... A two-dimensional matrix (i.e., the potential task scheduling priority matrix), where This indicates the total number of tasks pending processing in the current production line buffer. This represents the number of available devices. Elements in the matrix. Let be a continuous floating-point number within the closed interval [0,1], representing the recommended priority for the j-th device to handle the i-th task. During the initialization phase, the system uses a uniformly distributed random function to generate a preset number (e.g., 50) of particles to form the initial population. It should be noted that choosing continuous floating-point numbers instead of discrete integers as the encoding method is to avoid information loss due to rounding operations during speed updates, ensuring the smoothness of the search space and effective gradient propagation.

[0050] In one implementation, to quantitatively evaluate the merits of each particle (i.e., each priority strategy), this embodiment defines a fitness function with dual optimization objectives. The specific calculation process is as follows: Fast discrete event simulation is performed based on the priority matrix of the current particle. For each device, the execution queue of tasks to be processed is simulated according to the priority values ​​defined in the matrix from largest to smallest; the average turnaround time for the entire process is calculated. The total time taken for all tasks from entering the buffer to leaving the terminal device during the simulation is recorded, and its average value is calculated. The smaller this value, the higher the production efficiency; the load variance of the device nodes is calculated. The total running time of each device at the end of the simulation is statistically analyzed, and the variance of these running time values ​​is calculated. The smaller this value, the more balanced the load distribution among the devices, with no obvious bottlenecks or idleness; the fitness value is calculated using a weighted summation formula. Specifically, the reciprocal of the average turnaround time is multiplied by a first weighting coefficient, and the reciprocal of the load variance is multiplied by a second weighting coefficient.

[0051] It is worth noting that the first and second weighting coefficients are determined based on the normalized standard deviation of historical data. Specifically, the fluctuation ranges of two indicators, turnover time and load variance, are statistically analyzed within a preset period (such as the most recent 30 days). The larger the fluctuation range of an indicator, the smaller its corresponding weighting coefficient, in order to balance the difference in magnitude between the two objectives.

[0052] In another implementation, the particle swarm update process strictly follows the velocity-position update formula. In each iteration, the system traverses all particles in the swarm and updates their velocity vectors based on three parts: first, their own inertial velocity, maintaining the current search momentum; second, the cognitive part, i.e., the vector pointing from the particle's current position to its own historical best position, guiding the particle to recall individual experience; and third, the social part, i.e., the vector pointing from the particle's current position to the global historical best position of the entire swarm, guiding the particle towards collective intelligence. After updating the velocity, the old position is superimposed with the new velocity to obtain the new position. If the element value in the new position exceeds the range [0,1], a boundary clamping operation is performed to forcibly correct it to the boundary value (0 or 1).

[0053] For example, suppose that, after deduction, the strategy corresponding to particle A results in an average turnaround time of 20 minutes and a load variance of 5.0; while the strategy corresponding to particle B results in an average turnaround time of 18 minutes, but a load variance as high as 20.0. If the system sets the first weight (efficiency weight) to 0.6 and the second weight (balance weight) to 0.4, and the data normalization process has been completed, calculations show that although particle A is slightly slower, its overall fitness may be higher than that of particle B due to its extremely high load balance (small variance, large reciprocal). This reflects the scheduling philosophy of this invention: to slightly reduce the speed of a single machine in order to maintain overall line balance.

[0054] It should be noted that the preset convergence condition is determined using a stagnant window mechanism. Specifically, a sliding window is set to record the rate of change of the globally optimal fitness value over several consecutive iterations. It is worth noting that the length K of this sliding window (i.e., the number of observation generations) is dynamically determined based on the dimension of the solution space (i.e., the number of tasks multiplied by the number of devices). Specifically, this embodiment sets the K value to be proportional to the square root of the solution space dimension, or based on the average number of generations required for the algorithm to escape local optima in offline testing. For example, through backtesting analysis of historical data, it is found that the algorithm requires an average of 15 generations to escape local optima on a certain scale of problem; therefore, the K value is set to 1.2 times this average (e.g., 18 generations). This setting ensures that the window length is sufficient to cover the algorithm's oscillation period, effectively preventing it from getting trapped in local optima due to premature convergence determination in flat areas of the solution space, while also avoiding ineffective computation and wasted computing power caused by an excessively large window.

[0055] For example, if it is detected that the rate of change of the global optimal fitness value is consistently lower than a preset minimum threshold within a sliding window of length K, then the algorithm is considered to have converged to a steady state. It should be noted that this minimum threshold is determined based on the precision error limit of computer floating-point arithmetic, and is typically set to 10 times the precision of floating-point numbers to ignore minor fluctuations caused by computational noise. The iteration is then immediately terminated, and the current global optimal position is output as the final task scheduling priority matrix.

[0056] It should be noted that determining the blocked node based on the task scheduling priority matrix and the preset processing capacity threshold requires first parsing the task scheduling priority matrix. The values ​​stored in this matrix are used to characterize the recommended priority of different devices in handling different tasks. The higher the value, the higher the priority of the device in handling the task. Combining the real-time material flow direction information recorded in the material flow vector, the system counts the list of tasks waiting to be processed in the inlet buffer of each device at the current moment, and predicts the list of tasks that will arrive at the device within the preset time window based on the real-time processing progress of the upstream devices.

[0057] For each task to be processed, the system assigns a weight coefficient to the task based on its priority value in the priority matrix corresponding to the target device. Then, for each device, all its tasks to be processed (including tasks currently in the queue and tasks predicted to arrive soon) are summed according to their respective weight coefficients to obtain the overall expected load of the device. This overall expected load reflects both the number of tasks to be processed and the importance of the tasks themselves, and can more accurately characterize the real-time processing pressure that the device will face.

[0058] After completing the above calculations, the system compares the overall expected load of each device with the preset processing capacity threshold. If the overall expected load of a device is greater than its preset processing capacity threshold, the device is determined to be in an overload state and marked as a blocked node; otherwise, it is considered to be in a normal state.

[0059] The preset processing capacity threshold is set by collecting instantaneous load data of each device during its historical normal operation cycle, removing abnormal values ​​caused by faults or planned maintenance, and obtaining a load sample set that reflects the normal operating conditions of the equipment. Then, the statistical distribution of this sample set is calculated, and its high quantile is selected as the typical maximum processing capacity reference value of the equipment. To ensure that the equipment operation has sufficient safety margin and avoid the risk of equipment fatigue and failure caused by long-term full-load operation, the above-mentioned typical maximum processing capacity reference value is multiplied by a preset safety factor, and finally the preset processing capacity threshold of the equipment is obtained. This safety factor can be set differently according to different equipment types, process requirements and production targets, and is usually between 0.8 and 0.95.

[0060] In step S12, the bottleneck region and its severity are determined based on the blocked node and a preset theoretical throughput of the equipment, including:

[0061] Obtain the real-time task backlog of the blocked node, and calculate the difference between the real-time task backlog and the preset theoretical throughput of the device;

[0062] Divide the difference by the theoretical throughput of the device to obtain the quantified severity.

[0063] The node with the highest degree of congestion and its adjacent upstream node are identified as the linkage bottleneck area.

[0064] In one implementation, the system first obtains the real-time task backlog by using a visual sensor or photoelectric counter deployed at the entrance of the blocked node (i.e., the high-risk device identified in the previous steps). It should be noted that, to address the issue of consistency in physical dimensions, the real-time task backlog does not refer solely to the static inventory quantity in the current buffer, but rather to the total task load that the node needs to process within a preset evaluation time window. Specifically, the calculation logic involves reading the actual amount of material currently stuck in the buffer, adding it to the predicted amount of material expected to flow in from upstream equipment within the time window, and the sum of these two amounts is the real-time task backlog.

[0065] It is worth noting that the length of the preset evaluation time window is determined based on the minimum operating cycle time of the equipment. Specifically, the system obtains the standard single operation cycle time of this type of equipment and sets the evaluation time window length as an integer multiple of this cycle time (e.g., 10 times) to ensure that the evaluation results are statistically significant and to avoid misjudgments caused by instantaneous flow fluctuations.

[0066] In one implementation, this embodiment defines the preset theoretical throughput of the equipment as the maximum amount of material the equipment can process within the aforementioned evaluation time window (i.e., rated processing rate multiplied by the time window length). Subsequently, the difference between the real-time task backlog and the theoretical throughput of the equipment is calculated. This difference directly reflects the capacity gap within the current time window. A positive difference indicates a risk of backlog; a negative difference indicates excess capacity.

[0067] For example, suppose the crusher's evaluation time window is set to 10 minutes. Currently, the buffer has 5 cars awaiting crushing, and it is expected that 3 more cars will be delivered upstream within the next 10 minutes, resulting in a real-time task backlog of 8 cars. Referring to the equipment parameters, the crusher's rated throughput is 0.5 cars per minute, so its theoretical throughput within 10 minutes is 5 cars. Therefore, the difference is 3 cars.

[0068] In one implementation, to normalize the comparison of congestion levels for devices of different specifications, this embodiment divides the difference by the theoretical throughput of the device to obtain a dimensionless percentage value, i.e., the quantified severity. Continuing the example above, the severity is 0.6 (i.e., 60%). This value indicates that the current node's overload has reached 60% of its rated processing capacity, belonging to a highly congested state. The system iterates through all devices marked as congested nodes, calculates their respective severity values, and sorts them in descending order.

[0069] It should be noted that, considering the significant reverse propagation characteristics of production line congestion (i.e., downstream congestion can force upstream shutdowns), simply locking down the current node cannot completely solve the linkage problem. Therefore, this embodiment defines the most severe congestion node as the primary bottleneck and identifies its direct predecessor node (i.e., the adjacent upstream node) by querying a preset workshop equipment topology diagram. The system delineates these two nodes (the primary bottleneck plus the direct upstream node) as the linkage bottleneck area. This area definition method ensures that subsequent virtual simulation and load transfer strategies can simultaneously cover both the congestion point and the material supply source, thereby achieving root-cause unblocking.

[0070] In step S13, a virtual simulation environment is constructed based on the linkage bottleneck area, including:

[0071] Obtain the physical parameters and current operating status of the equipment within the aforementioned bottleneck area;

[0072] A virtual model consistent with the geometric parameters of the physical entity is constructed based on the physical parameters of the device.

[0073] The current running state is mapped to the virtual model, and the interaction running rules between the virtual models are configured using the task scheduling priority matrix to generate the virtual simulation environment.

[0074] It should be noted that the preset safety threshold is determined based on sensitivity analysis from offline simulation experiments. During the system debugging phase, a simulation model encompassing the entire production line is constructed to simulate bottleneck areas of varying severity, for example, the severity gradually increases from 10% to 100%, and the changing patterns of key performance indicators of the production line, such as overall production rate, congestion propagation speed, and equipment idle rate, are observed. When the severity exceeds a certain critical value, the performance indicators deteriorate significantly, for example, the production rate drops by more than 5%, or the congestion spreads to multiple upstream nodes within a preset time. This critical value is then recorded as the preset safety threshold. This threshold can be calibrated according to the process characteristics and operational goals of different production lines and is fixed in the system to trigger subsequent deep reinforcement learning optimization processes.

[0075] In one implementation, the system first defines the simulation boundary and performs refined modeling only on the bottleneck area and its adjacent buffer nodes to reduce computational overhead. Then, it retrieves static equipment physical parameters from a pre-configured equipment lifecycle management database.

[0076] It should be noted that the equipment lifecycle management database is pre-built by the system during the initial deployment phase. Specifically, the database is built by parsing the electronic technical specifications (datasheets) provided by the equipment manufacturers, extracting the geometric and kinematic parameters of each device, and combining this with historical maintenance logs to input theoretical failure rate data, thus establishing a structured dataset indexed by the device's unique identifier.

[0077] In one implementation, the system synchronizes its current operating status from the field data acquisition and monitoring control system. It's worth noting that this embodiment establishes a connection with the underlying PLC controller via the OPC UA industrial communication protocol, reading register values ​​in real time to obtain the operating status. The physical parameters of the equipment include, but are not limited to, the three-dimensional geometric dimensions of the equipment used for collision detection, the kinematic constraints of moving parts (such as the maximum stroke and rotation angle of the robotic arm), the rated processing cycle time, and the theoretical mean time between failures. The current operating status includes the currently processed task ID, the real-time occupancy of the input / output buffers, the wear index of core components, and instantaneous energy consumption readings.

[0078] In one implementation, a virtual model is constructed based on the physical parameters of the device. This embodiment employs a hybrid modeling technique. At the geometric level, a simplified bounding box model is generated based on the device's CAD data to ensure that the footprint in the virtual space strictly matches the physical entity, thereby verifying the feasibility of path planning. At the logical level, a behavioral model based on discrete events is constructed. This behavioral model defines the device's state machine (e.g., idle, running, fault, blocked) and incorporates a random disturbance generator based on probability distribution (e.g., using Weibull distribution to simulate random shutdown faults), thus enabling the virtual model to possess realistic dynamic response characteristics.

[0079] For example, for a hydraulic shearing machine, its virtual model includes not only length, width, and height... The geometric entity also includes a logic module. This logic module is configured to have a normal distribution with a mean of 45 seconds and a standard deviation of 3 seconds for a single shearing operation, and is configured to have a 0.5% probability of triggering a jamming state every 100 runs.

[0080] In one implementation, to ensure the simulation's starting point is synchronized with physical reality, the system performs a state injection operation. The real-time operating status of the physical devices (e.g., the current buffer remaining space equals 2) is directly assigned to the corresponding virtual model variables. The task scheduling priority matrix generated in the preceding steps is used to configure the interaction rules between virtual models. Specifically, this matrix is ​​loaded into the virtual simulation engine's scheduler as the basis for resource arbitration decisions. When multiple virtual devices simultaneously compete for the same batch of materials or vie for downstream conveying channels, the simulation engine queries the corresponding priority value in the matrix. Devices with higher priority gain priority processing or passage rights, while devices with lower priority are forced into a waiting state.

[0081] It's worth noting that, to simulate realistic physical interaction constraints, the system also incorporates blocking propagation rules. If the buffer of a downstream virtual device is full (i.e., the physical capacity limit has been reached), the output port of the upstream virtual device is forcibly locked, prohibiting material flow until the downstream releases space. Through this dual mechanism of priority-driven decision-making and physical constraint restrictions, the generated virtual simulation environment can faithfully reproduce the evolution trend of bottleneck areas under different scheduling strategies.

[0082] In step S13, a pre-trained deep reinforcement learning model is used to perform policy interaction and search in the virtual simulation environment, outputting a task allocation scheme, including:

[0083] Collect the device queue length, remaining processing capacity, and buffer occupancy rate of the virtual simulation environment to construct a state space vector;

[0084] Using the policy network of the deep reinforcement learning model, action instructions are output according to the state space vector. The action instructions include adjusting task priority weights or modifying transport band beat parameters.

[0085] Execute the action instructions in the virtual simulation environment, and calculate the output increase rate per unit time and the delay reduction rate of key processes after execution.

[0086] The comprehensive reward value is calculated based on the production increase rate and the delay reduction rate, and the policy network is updated using the comprehensive reward value until the comprehensive reward value meets the preset convergence condition, and the task allocation scheme under the corresponding action sequence is output.

[0087] In one implementation, the system first constructs a state space vector to drive decision-making. This embodiment uses a preset simulation time step as the sampling period to extract key feature data from the virtual model in real time, including the number of tasks to be processed at the input ports of each bottleneck device as the device queue length; obtaining the difference between the theoretical maximum output of each device under the current operating conditions and the current actual load, and defining the ratio of this difference to the theoretical maximum output as the remaining processing capacity; the system also reads the ratio of the occupied space of the transmission belt or buffer area between each device to the total physical capacity as the buffer occupancy rate. To eliminate the influence of different physical dimensions on neural network calculations, this embodiment standardizes the above data by subtracting the historical mean from each value and dividing by the standard deviation. The processed data is then concatenated into a one-dimensional tensor to form the state space vector.

[0088] It should be noted that the deep reinforcement learning model used in this embodiment is specifically built based on the actor-critic architecture. The policy network (i.e., the actor network) exemplarily adopts a fully connected feedforward neural network structure. This network includes an input layer whose dimension is consistent with the dimension of the state space vector; several hidden layers (e.g., three hidden layers with 128, 64, and 32 neurons respectively), with Rectified Linear Units (ReLU) used as activation functions between layers to introduce non-linear features; and an output layer whose dimension corresponds to the dimension of the action instruction. The output layer uses a hyperbolic tangent function (Tanh) to constrain the output value within a preset action range. To address the problem of excessive time consumption when searching from scratch, this embodiment employs a hybrid training strategy of offline warm-start and online fine-tuning.

[0089] Specifically, the system pre-collects operation logs from several historical periods of the production line, parses the sensor data and control commands in the logs, and formats them into a sequence of tuples consisting of states, actions, and rewards to construct an offline dataset. The behavior cloning algorithm is then used to perform supervised learning pre-training on the aforementioned policy network, enabling its initial weight parameters to mimic historical best-practice scheduling experiences.

[0090] In one implementation, the policy network outputs action instructions based on the input state space vector and sends them to the virtual environment for simulation. After the simulation, the system quantitatively evaluates the execution effect of the instructions. First, the total output within the current simulation cycle is calculated and compared with the baseline output before the action is executed. The ratio of the difference between the two to the baseline output is calculated to obtain the output improvement rate per unit time. Simultaneously, the longest dwell time on the critical path of the production line is identified, and the difference between the baseline delay before the action and the simulation delay after the action is calculated. The ratio of this difference to the baseline delay is defined as the critical process delay reduction rate. These two indicators objectively reflect the effectiveness of the scheduling strategy from the dimensions of throughput and timeliness, respectively.

[0091] It is worth noting that, in order to guide the policy network to evolve in a direction that aligns with the current production goals, this embodiment constructs a calculation logic for the comprehensive reward value. Specifically, the output increase rate per unit time is multiplied by a preset first reward weight to obtain the output reward component; the delay reduction rate of the key process is multiplied by a preset second reward weight to obtain the delay reward component; and the output reward component and the delay reward component are added together to obtain the comprehensive reward value.

[0092] It should be noted that the first and second reward weights are not fixed, but are determined based on the current production strategy. For example, when the system is in capacity-first mode, the first weight is set to a larger value (e.g., 0.7) and the second weight to a smaller value (e.g., 0.3) to focus on improving throughput; when in on-time delivery mode, the weight ratios are reversed.

[0093] For example, this online fine-tuning process is a rapidly iterative closed loop. The system uses a proximal policy optimization algorithm to update the policy network parameters. It's worth noting that the key hyperparameters in the algorithm are selected based on the following criteria: the learning rate is set to 0.03% to ensure a balance between convergence speed and stability; the discount factor is set to a value close to one (e.g., 0.99) to focus on long-term cumulative rewards; and the pruning parameter is set to a small value (e.g., 0.2) to limit the policy update magnitude and prevent performance collapse. The system sets a preset convergence condition where the variance of the average comprehensive reward value over several consecutive training rounds is less than a preset minimum threshold (e.g., 0.5%). When this condition is met, the model is considered to have converged to a steady state. At this point, the system stops updating parameters, uses the converged policy network to infer the current state, and outputs a deterministic action sequence, i.e., a combination of specific priority adjustment values ​​and beat parameters, as the final task allocation scheme.

[0094] In step S14, a path correction vector for the congested area is generated, including:

[0095] The preset workshop space topology is divided into a grid map, and the material occupancy rate is mapped to the access cost weight of the grid.

[0096] Analyze the task allocation scheme to determine the starting position and target processing position of the material to be processed;

[0097] Using the A* path search algorithm, search for the minimum cost path from the starting position to the target processing position in the grid map;

[0098] Calculate the difference between the coordinates of the key nodes of the minimum cost path and the coordinates of the corresponding nodes of the original path, and generate the path correction vector.

[0099] In one implementation, the system first discretizes the preset workshop space topology to construct a two-dimensional grid map. It's important to note that to ensure the physical feasibility of the generated paths, the grid size is not arbitrarily set, but rather determined based on the minimum projected area of ​​the material transport vehicle. Specifically, the system obtains the physical width and length of the automated guided vehicle (AGV) or overhead conveyor, selects the larger of the two as the baseline size, and sets the grid side length to an integer multiple of this baseline size, thus ensuring that each grid can accommodate a complete transport unit.

[0100] In one implementation, the system maps the material occupancy rate calculated in the preceding steps to the passage cost weight for each grid cell. It's worth noting that this mapping employs piecewise nonlinear penalty logic. Specifically, the system pre-sets safety and congestion thresholds.

[0101] It should be noted that the congestion threshold is not subjectively set, but determined based on deadlock probability analysis of historical workshop operation data. For example, the inflection point value of a sudden increase in deadlock occurrence rate (such as 85%) is selected as the threshold. When the material occupancy rate of the grid is lower than the safety threshold, the passage cost weight is equal to the basic passage constant; when it is between the two, the passage cost weight increases exponentially with the occupancy rate, forming a soft isolation barrier; when it exceeds the congestion threshold, it is marked as an impassable area, and the cost weight is set to infinity.

[0102] In one implementation, the A* path search algorithm is used to search the grid map for the minimum-cost path from the starting position to the target processing position. It should be noted that, considering the possibility that no effective path may be found under extreme congestion conditions (i.e., the cost of all feasible paths tends to infinity or the search times out), this embodiment specifically includes a fallback scheduling logic. If the A* path search algorithm returns an empty path, the system will automatically retrieve the nearest preset temporary buffer or emergency avoidance zone to the current starting position, replan the temporary path to that avoidance zone, and generate a path correction vector pointing to that avoidance zone to prevent scheduling deadlock.

[0103] For example, if the search is successful, to address the issue of inconsistent node counts between the new path and the original path, this embodiment introduces key point alignment logic. First, all geometric inflection points of the new path are extracted; then, based on a dynamic time warping algorithm or the minimum Euclidean distance principle, a mapping is established between the key nodes of the new path and their corresponding nodes in the original path; finally, the coordinate difference is calculated to generate the path correction vector. This vector precisely quantifies the spatial detour and direction that materials need to take to avoid congestion.

[0104] In step S15, available alternative devices located outside the congested area are scheduled to take over the pending tasks within the congested area, perform load transfer, and obtain load adjustment parameters, including:

[0105] Obtain the set of all candidate devices located outside the congested area;

[0106] Traverse the set of candidate devices and check the current operating status and process attributes of each candidate device;

[0107] Select the equipment whose current operating status is idle and whose process attributes match the task to be processed as the available alternative equipment;

[0108] Generate scheduling instructions for the available alternative equipment to obtain the load adjustment parameters.

[0109] In one implementation, the system first performs a spatial reverse search based on the congested areas identified in the previous steps. Specifically, the system loads a preset workshop equipment layout coordinate system, identifies all equipment whose physical coordinates do not fall within the congested areas, and defines them as an initial set of candidate equipment. It should be noted that, to ensure the feasibility of load transfer at the logistics level, the system further calls a preset logistics network topology map when constructing this set, performs reachability analysis, and eliminates isolated equipment that lacks a direct physical transmission link (such as a conveyor belt connection or automated guided vehicle lane) to the congested areas, ensuring that the candidate equipment is physically logistics reachable.

[0110] In one implementation, the system performs a deep capability verification on each device in the candidate device set. First, it reads the device's controller register data via the industrial control bus to obtain its current operating status. It's worth noting that this embodiment strictly categorizes status codes; a device is only considered idle if its status code indicates "power-on ready" and its task queue is empty. Devices in fault-related shutdown, planned maintenance, or energy-saving hibernation states are explicitly excluded to prevent ineffective scheduling. Second, the system retrieves process attributes from the device master data management system. These attributes represent a standardized set of capability tags, including key indicators such as maximum shear thickness and processing accuracy level.

[0111] In one implementation, the system performs feature matching filtering to determine available alternative equipment. The process requirement tags of the current task to be processed are obtained and compared with the process attributes of candidate equipment. It should be noted that this embodiment uses capability coverage logic to determine whether process attributes match. Specifically, a match is determined if and only if the process capability parameter value of the candidate equipment is greater than or equal to the process requirement parameter value of the task, and the candidate equipment possesses all the specific functional tags required by the task. That is, the capability set of the candidate equipment must be a superset of the task requirement set.

[0112] For example, if the task to be processed requires cutting a thickness of 3 mm, and candidate device A has the capability to cut a thickness of 5 mm (i.e., backward compatible), then it is determined to be a match; if candidate device B only has the capability to cut a thickness of 2 mm, then it is determined to be a mismatch. The system filters out all devices that simultaneously meet the idle state and matching conditions as available alternative devices.

[0113] In one implementation, the system generates scheduling instructions for the available alternative equipment. These instructions include logistics route switching instructions (such as controlling the operation of a diverter) and job initiation instructions. Simultaneously, the system calculates the load adjustment parameters. It should be noted that these parameters are a quantified gain vector used to guide the intensity of equipment operation. Specifically, the calculation logic is as follows: based on the ratio of the estimated working hours of the transfer task to the standard capacity of the alternative equipment, the expected load rate increment of the alternative equipment is calculated; simultaneously, based on the increase in logistics transportation distance, the speed compensation coefficient of the transmission system is calculated. This vector, containing the load rate increment and the speed compensation coefficient, constitutes the load adjustment parameters, providing data support for subsequent global collaborative correction.

[0114] In step S16, based on the load adjustment parameters, a schedule deviation vector is calculated using a preset production schedule simulation model, and the operating parameters of the multi-device collaboration are updated based on the schedule deviation vector, including:

[0115] If the magnitude of the schedule deviation vector exceeds the preset tolerance range, a global cycle time compensation coefficient for multi-device collaboration is generated based on the direction of the schedule deviation vector.

[0116] The global beat compensation coefficient is used to correct the operating power or transmission speed of each device.

[0117] In one implementation, the system first invokes a pre-defined production schedule simulation model to perform feedforward extrapolation. It's important to note that this model is not a graphical 3D physical simulation, but rather a mathematical recursive model based on discrete time steps, internally storing the standard operating hours and process dependency paths for each standard process on the production line. The system inputs the load adjustment parameters generated in the preceding steps (i.e., the expected change in operating hours due to load transfer) into this model, and, starting from the current moment, extrapolates the projected output within a future preset period (e.g., one hour). Subsequently, the system obtains a snapshot of the target schedule set in the master production schedule. The difference between the extrapolated projected completion time and the target planned time is defined as the time dimension deviation, and the difference between the projected output quantity and the target planned quantity is defined as the output dimension deviation. The system combines the deviation values ​​of these two dimensions to construct a two-dimensional schedule deviation vector.

[0118] In one implementation, the system calculates the magnitude of the schedule deviation vector to quantify the overall severity of the current production status deviating from the plan, and compares this magnitude with a preset tolerance range. It's worth noting that this tolerance range is determined through statistical process control analysis of historical delivery data, typically set to three standard deviations of the historical average deviation. If the magnitude falls within the tolerance range, it indicates that the current minor fluctuations are normal system noise and require no intervention; if it exceeds the tolerance range, the system activates a global compensation mechanism.

[0119] In one implementation, the system analyzes the direction of the schedule deviation vector to determine the adjustment strategy. If the vector points to the lagging or gap quadrant, it indicates that the overall capacity of the production line is insufficient and acceleration is needed; if it points to the leading or backlog quadrant, it indicates that there is excess capacity and speed can be appropriately reduced to save energy. Based on this direction, the system generates a global cycle time compensation coefficient.

[0120] It should be noted that the calculation of this coefficient employs proportional-integral control logic. Specifically, the system first calculates the product of the current deviation magnitude and the preset proportional gain coefficient as the proportional term, used to quickly respond to the current instantaneous error. Simultaneously, the system accumulates historical deviation values ​​over a period of time, calculates the product of the accumulated error and the preset integral gain coefficient as the integral term, used to eliminate steady-state errors. The final global clock compensation coefficient equals the reference value (usually a value of one) plus the sum of the proportional and integral terms.

[0121] For example, the proportional gain coefficient and integral gain coefficient are not arbitrarily set, but determined through offline simulation calibration. That is, before the system goes online, a typical step disturbance signal (such as simulating a shutdown accident) is input, the system response curve is observed, and the two coefficients are adjusted using the Ziegler-Nichols tuning method until the system's settling time is minimized and the overshoot is reduced.

[0122] In one implementation, the global cycle time compensation coefficient is used to correct the operating power or transmission speed of each device. For conveyor belt devices, the set linear speed is directly multiplied by the coefficient to achieve overall acceleration or deceleration of material transmission; for processing equipment (such as crushers), the rated operating power or feed speed is multiplied by the coefficient.

[0123] It should be noted that, to prevent equipment damage due to overclocking, the system verifies the physical limit parameters (such as maximum rated speed) of each device before performing corrections. If the calculated target speed exceeds the physical limit, it is forcibly clamped to a safe limit value, and the actual executable compensation coefficient is recalculated to ensure the engineering safety of the scheduling instructions. Through this global fine-tuning of the cycle time, the system can bring the overall production schedule back to the predetermined track without disrupting the local load balance.

[0124] In summary, this invention constructs a dynamic task scheduling priority matrix that reflects equipment capabilities in real time by fusing multi-source operational data from an automotive dismantling production line and iteratively optimizing based on a particle swarm optimization algorithm, breaking the limitations of traditional static cycle time control. It utilizes a deep reinforcement learning model in a virtual simulation environment for offline hot-start and online fine-tuning strategy search, achieving adaptive optimal task allocation for scenarios with random disturbances involving non-standard materials. Furthermore, by combining spatial topology-based congestion area identification and path correction vector generation, as well as load transfer scheduling for idle equipment in non-congested areas, it effectively alleviates physical space-based logistical deadlocks and eliminates local bottlenecks. Finally, through global cycle time compensation closed-loop control based on a production progress simulation model, it achieves dynamic correction of multi-equipment collaborative operating parameters, thereby significantly improving the overall load balance, system throughput, and resource utilization efficiency of the automotive dismantling production line under complex operating conditions.

[0125] Reference Figure 2 The second embodiment of the present invention provides a multi-equipment linkage intelligent scheduling system for an automobile dismantling production line, comprising:

[0126] The data processing module is used to acquire multi-source operating data of the automobile dismantling production line and perform preprocessing to obtain the equipment load distribution matrix and material flow vector.

[0127] The bottleneck identification module is used to perform iterative optimization using the particle swarm optimization algorithm based on the equipment load distribution matrix and the material flow vector to obtain a task scheduling priority matrix; to determine the blocking nodes based on the task scheduling priority matrix and a preset processing capacity threshold; and to determine the linkage bottleneck area and the severity of the linkage bottleneck area based on the blocking nodes and a preset theoretical equipment throughput.

[0128] The strategy optimization module is used to construct a virtual simulation environment based on the linkage bottleneck area if the severity exceeds a preset safety threshold; and to perform strategy interaction and search in the virtual simulation environment using a pre-trained deep reinforcement learning model to output a task allocation scheme.

[0129] The path monitoring module is used to calculate the material occupancy rate within the preset workshop space topology according to the task allocation scheme. If the material occupancy rate exceeds the preset congestion threshold, the congested area is determined according to the spatial distribution of the material occupancy rate, and a path correction vector for the congested area is generated.

[0130] The load transfer module is used to schedule available alternative devices located outside the congested area to take over the pending tasks within the congested area, perform load transfer, and obtain load adjustment parameters;

[0131] The feedback control module is used to calculate the progress deviation vector using a preset production progress simulation model based on the load adjustment parameters, and to update the multi-device collaborative operating parameters based on the progress deviation vector. It should be noted that the intelligent scheduling system for multi-device linkage of an automobile dismantling production line provided in this embodiment of the invention executes all the process steps of the intelligent scheduling method for multi-device linkage of an automobile dismantling production line described in the above embodiment. The working principles and beneficial effects of both correspond one-to-one, and therefore will not be elaborated further.

[0132] It should be noted that the system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the system embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.

[0133] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.

Claims

1. A method for intelligent scheduling of multiple devices in an automobile dismantling production line, characterized in that, include: Acquire multi-source operational data from the automobile dismantling production line and preprocess it to obtain the equipment load distribution matrix and material flow vector; Based on the equipment load distribution matrix and the material flow vector, the particle swarm optimization algorithm is used to iteratively optimize and obtain the task scheduling priority matrix; based on the task scheduling priority matrix and the preset processing capacity threshold, the blocking nodes are determined, and based on the blocking nodes and the preset theoretical throughput of the equipment, the linkage bottleneck area and the severity of the linkage bottleneck area are determined. If the severity exceeds a preset safety threshold, a virtual simulation environment is constructed based on the linkage bottleneck area; a pre-trained deep reinforcement learning model is used to perform policy interaction and search in the virtual simulation environment, and a task allocation scheme is output. According to the task allocation scheme, the material occupancy rate within the preset workshop space topology is calculated. If the material occupancy rate exceeds the preset congestion threshold, the congested area is determined based on the spatial distribution of the material occupancy rate, and a path correction vector for the congested area is generated. Available alternative devices located outside the congested area are dispatched to take over the pending tasks within the congested area, perform load transfer, and obtain load adjustment parameters; Based on the load adjustment parameters, the progress deviation vector is calculated using a preset production progress simulation model, and the operating parameters of multi-device collaboration are updated based on the progress deviation vector. The step of determining the bottleneck region and its severity based on the blocked node and a preset theoretical throughput of the equipment includes: Obtain the real-time task backlog of the blocked node, and calculate the difference between the real-time task backlog and the preset theoretical throughput of the device; Divide the difference by the theoretical throughput of the device to obtain the quantified severity. The node with the highest degree of congestion and its adjacent upstream node are identified as the linkage bottleneck area. The step of utilizing a pre-trained deep reinforcement learning model to perform policy interaction and search in the virtual simulation environment and output a task allocation scheme includes: Collect the device queue length, remaining processing capacity, and buffer occupancy rate of the virtual simulation environment to construct a state space vector; Using the policy network of the deep reinforcement learning model, action instructions are output according to the state space vector. The action instructions include adjusting task priority weights or modifying transport band beat parameters. Execute the action instructions in the virtual simulation environment, and calculate the output increase rate per unit time and the delay reduction rate of key processes after execution. The comprehensive reward value is calculated based on the production increase rate and the delay reduction rate, and the policy network is updated using the comprehensive reward value until the comprehensive reward value meets the preset convergence condition, and the task allocation scheme under the corresponding action sequence is output.

2. The intelligent scheduling method for multi-equipment linkage in an automobile dismantling production line according to claim 1, characterized in that, The process of acquiring multi-source operational data from the automotive dismantling production line and preprocessing it to obtain the equipment load distribution matrix and material flow vector includes: Extract the timestamps of the multi-source running data, and use a linear interpolation algorithm to align data streams of different frequencies to a preset unified time axis; Calculate the numerical variance of each aligned data stream within a preset sliding window; Calculate the reciprocal of the numerical variance and normalize the reciprocal to obtain the fusion weight; The multi-source operating data of each device are weighted and summed according to the index dimensions using the fusion weights to generate the device load distribution matrix that represents the real-time load status of each device.

3. The intelligent scheduling method for multi-equipment linkage in an automobile dismantling production line according to claim 1, characterized in that, The step of obtaining the task scheduling priority matrix by iterative optimization using the particle swarm optimization algorithm based on the equipment load distribution matrix and the material flow vector includes: Construct a particle swarm, defining the position of each particle as a potential task scheduling priority matrix; Define a fitness function that is inversely proportional to the average turnaround time of the entire production line and inversely proportional to the load variance of each equipment node; The velocity and position of the particle are updated according to the fitness function until a preset number of iterations is reached or a preset convergence condition is met, and the position of the optimal particle is determined as the task scheduling priority matrix.

4. The intelligent scheduling method for multi-equipment linkage in an automobile dismantling production line according to claim 1, characterized in that, The construction of the virtual simulation environment based on the aforementioned bottleneck area includes: Obtain the physical parameters and current operating status of the equipment within the aforementioned bottleneck area; A virtual model consistent with the geometric parameters of the physical entity is constructed based on the physical parameters of the device. The current running state is mapped to the virtual model, and the interaction running rules between the virtual models are configured using the task scheduling priority matrix to generate the virtual simulation environment.

5. The intelligent scheduling method for multi-equipment linkage in an automobile dismantling production line according to claim 1, characterized in that, The generation of path correction vectors for the congested area includes: The preset workshop space topology is divided into a grid map, and the material occupancy rate is mapped to the access cost weight of the grid. Analyze the task allocation scheme to determine the starting position and target processing position of the material to be processed; Using the A* path search algorithm, search for the minimum cost path from the starting position to the target processing position in the grid map; Calculate the difference between the coordinates of the key nodes of the minimum cost path and the coordinates of the corresponding nodes of the original path, and generate the path correction vector.

6. The intelligent scheduling method for multi-equipment linkage in an automobile dismantling production line according to claim 1, characterized in that, The scheduling of available alternative equipment located outside the congested area to take over the pending tasks within the congested area, perform load transfer, and obtain load adjustment parameters, including: Obtain the set of all candidate devices located outside the congested area; Traverse the set of candidate devices and check the current operating status and process attributes of each candidate device; Select the equipment whose current operating status is idle and whose process attributes match the task to be processed as the available alternative equipment; Generate scheduling instructions for the available alternative equipment to obtain the load adjustment parameters.

7. The intelligent scheduling method for multi-equipment linkage in an automobile dismantling production line according to claim 1, characterized in that, The step of calculating the progress deviation vector using a preset production progress simulation model based on the load adjustment parameters, and updating the multi-device collaborative operation parameters based on the progress deviation vector, includes: If the magnitude of the schedule deviation vector exceeds the preset tolerance range, a global cycle time compensation coefficient for multi-device collaboration is generated based on the direction of the schedule deviation vector. The global beat compensation coefficient is used to correct the operating power or transmission speed of each device.

8. A multi-equipment linkage intelligent scheduling system for an automobile dismantling production line, characterized in that, include: The data processing module is used to acquire multi-source operating data of the automobile dismantling production line and perform preprocessing to obtain the equipment load distribution matrix and material flow vector. The bottleneck identification module is used to perform iterative optimization using the particle swarm optimization algorithm based on the equipment load distribution matrix and the material flow vector to obtain a task scheduling priority matrix; to determine the blocking nodes based on the task scheduling priority matrix and a preset processing capacity threshold; and to determine the linkage bottleneck area and the severity of the linkage bottleneck area based on the blocking nodes and a preset theoretical equipment throughput. The strategy optimization module is used to construct a virtual simulation environment based on the linkage bottleneck area if the severity exceeds a preset safety threshold; and to perform strategy interaction and search in the virtual simulation environment using a pre-trained deep reinforcement learning model to output a task allocation scheme. The path monitoring module is used to calculate the material occupancy rate within the preset workshop space topology according to the task allocation scheme. If the material occupancy rate exceeds the preset congestion threshold, the congested area is determined according to the spatial distribution of the material occupancy rate, and a path correction vector for the congested area is generated. The load transfer module is used to schedule available alternative devices located outside the congested area to take over the pending tasks within the congested area, perform load transfer, and obtain load adjustment parameters; The feedback control module is used to calculate the progress deviation vector using a preset production progress simulation model based on the load adjustment parameters, and update the multi-device collaborative operation parameters based on the progress deviation vector. The step of determining the bottleneck region and its severity based on the blocked node and a preset theoretical throughput of the equipment includes: Obtain the real-time task backlog of the blocked node, and calculate the difference between the real-time task backlog and the preset theoretical throughput of the device; Divide the difference by the theoretical throughput of the device to obtain the quantified severity. The node with the highest degree of congestion and its adjacent upstream node are identified as the linkage bottleneck area. The step of utilizing a pre-trained deep reinforcement learning model to perform policy interaction and search in the virtual simulation environment and output a task allocation scheme includes: Collect the device queue length, remaining processing capacity, and buffer occupancy rate of the virtual simulation environment to construct a state space vector; Using the policy network of the deep reinforcement learning model, action instructions are output according to the state space vector. The action instructions include adjusting task priority weights or modifying transport band beat parameters. Execute the action instructions in the virtual simulation environment, and calculate the output increase rate per unit time and the delay reduction rate of key processes after execution. The comprehensive reward value is calculated based on the production increase rate and the delay reduction rate, and the policy network is updated using the comprehensive reward value until the comprehensive reward value meets the preset convergence condition, and the task allocation scheme under the corresponding action sequence is output.

Citation Information

Patent Citations

  • Multi-AGV global planning method based on network congestion model

    CN113516429A

  • Optimal path planning medical waste recovery scheduling system based on A* algorithm

    CN116222602A