Production line unit task allocation method and system based on reinforcement learning

By using a reinforcement learning-based production line unit task allocation method, real-time monitoring of orders and equipment status, and combining genetic algorithms and particle swarm optimization, the problem of multi-objective balance in production line task allocation is solved, achieving efficient and flexible task allocation, improving production efficiency and equipment utilization, and reducing energy consumption and costs.

CN121504043APending Publication Date: 2026-02-10DALIAN UNIV OF TECH +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511677464.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-17
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing production line task allocation methods struggle to balance task completion time, equipment utilization, energy consumption, and changeover costs in dynamic environments, leading to resource waste and inefficiency, and an inability to quickly respond to multidimensional constraints and changes in order priorities.

Method used

A reinforcement learning-based task allocation method for production line units is adopted. Through data acquisition, multi-objective function calculation, particle swarm optimization, and adaptive adjustment mechanism, order demand and equipment status are monitored in real time to generate a highly adaptive task allocation scheme. Combined with genetic algorithm and particle swarm optimization algorithm, multi-objective dynamic balance is achieved.

Benefits of technology

It improves production efficiency and equipment utilization, reduces energy consumption and production costs, enhances system flexibility and scalability, and can quickly respond to order priority fluctuations and equipment failures in highly dynamic environments, ensuring the real-time nature and effectiveness of task allocation schemes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121504043A_ABST
    Figure CN121504043A_ABST
Patent Text Reader

Abstract

The invention discloses a production line unit task allocation method based on reinforcement learning, which comprises the following steps: extracting a specific task instruction of each production line unit from a final task allocation matrix, verifying the feasibility of the instruction under the constraint of completion time through a simulation execution module, determining an instruction set passing verification, and performing task allocation on the instruction set; wherein the instruction set is obtained by fusing the solution of the multi-target conflict; according to the verified instruction set, production line feedback data such as actual execution time and energy consumption records are obtained, deviation is analyzed from the feedback data, if it is judged that the deviation is larger than a preset threshold value, a self-adaptive adjustment mechanism is triggered, and a corrected distribution strategy is obtained; and the optimized resource allocation indexes are extracted from the corrected allocation strategy, the indexes are pushed to the production line equipment through the real-time distribution system, the execution monitoring cycle after pushing is determined, and the monitoring cycle is obtained by continuously tracking the low-efficiency risk.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of production line scheduling technology, and in particular to a method and system for assigning production line unit tasks based on reinforcement learning. Background Technology

[0002] Task allocation within production line units is a crucial aspect of intelligent manufacturing, directly impacting production efficiency, cost control, and product quality. In modern industry, production lines need to dynamically allocate tasks based on order demand, equipment status, and resource constraints to achieve high-efficiency, low-cost, and highly flexible production goals. However, traditional task allocation methods rely on static rules or human experience, making it difficult to cope with dynamically changing production environments, leading to resource waste and inefficiency. Existing methods have limitations in handling dynamic production scenarios, failing to comprehensively consider the balance between task completion time, equipment utilization, energy consumption levels, and changeover costs. Especially in highly dynamic environments, task allocation decisions require rapid responses to multi-dimensional constraints, while existing solutions lack flexibility and a global perspective, making it difficult to achieve holistic optimization allocation strategies.

[0003] The core technical challenge lies in constructing an evaluation system that integrates multiple objectives and enables real-time optimization of task allocation in a dynamic environment. Production line task allocation needs to consider multiple dimensions simultaneously, including task completion time, equipment utilization, energy consumption, and changeover costs, but these objectives often conflict. For example, pursuing the shortest completion time may lead to frequent task changes, increasing additional changeover costs and energy consumption. Furthermore, real-time optimization in a dynamic environment further exacerbates the technical challenges. Since order priorities may change at any time during production, or equipment may malfunction, allocation decisions need to integrate multiple pieces of information and quickly generate the optimal solution within a very short period. Therefore, in actual production, designing an evaluation system that can dynamically balance task completion time, equipment utilization, and changeover costs under conflicting objectives, and rapidly generating highly adaptable allocation decisions in a real-time changing production environment, becomes a key issue in production line unit task allocation. Summary of the Invention

[0004] In order to solve the above-mentioned technical problems, the present invention provides a method and system for assigning tasks to production line units based on reinforcement learning.

[0005] The technical solution of this invention is implemented as follows: A reinforcement learning-based method for task allocation in production line units includes: Extract all current production line order demand data and equipment status information from sensors and databases, and determine a preliminary task sequence list for potential order priority fluctuations and equipment failure signals in a highly dynamic environment; Based on the preliminary task sequence list, calculate the multi-objective function value, filter out the combination of completion time, equipment utilization and switching cost from the task sequence list, and determine if the function value exceeds the preset threshold. If so, adjust the sequence to balance these objectives and obtain an optimized task allocation candidate set. Real-time environmental change data is obtained from the optimized task allocation candidate set. By comparing the matching degree between the change data and the candidate set, the allocation parameters that need to be updated are determined. For the updated allocation parameters, a new task allocation scheme is generated using the particle swarm optimization algorithm. The trade-off between equipment utilization and switching cost is iteratively calculated from the parameters. If the trade-off is lower than a preset threshold, the process is repeated until the condition is met, and the final task allocation matrix is ​​obtained. From the final task allocation matrix, extract the specific task instructions for each production line unit, verify the feasibility of the instructions under the completion time constraint through the simulation execution module, and determine the instruction set that has passed the verification.

[0006] A reinforcement learning-based production line unit task allocation system includes: The data acquisition module is used to extract current production line order demand data and equipment status information from sensors and databases, and generate a preliminary task sequence list to cope with order priority fluctuations and equipment failure signals in a highly dynamic environment. The optimization filtering module is used to calculate multi-objective function values, filter out combinations that include completion time, equipment utilization and switching costs, adjust the task sequence to balance these objectives, and generate an optimized task allocation candidate set. The parameter update module is used to acquire real-time environmental change data, compare the matching degree between the change data and the candidate set, and determine the allocation parameters that need to be updated to adapt to the highly dynamic environment. The particle swarm optimization module is used to generate new task allocation schemes through the particle swarm optimization algorithm, iteratively calculate the trade-off between equipment utilization and switching costs, and generate the final task allocation matrix. The instruction verification module is used to extract specific task instructions for production line units from the final task allocation matrix, verify the feasibility of the instructions through the simulation execution module, and determine the instruction set that passes the verification. The strategy correction module is used to analyze deviations based on production line feedback data, trigger an adaptive adjustment mechanism, correct the allocation strategy, and push the optimized resource allocation indicators to the production line equipment to form an execution monitoring loop.

[0007] Compared with the prior art, the present invention has the following advantages: 1. This invention obtains order demand data and equipment status information through a real-time monitoring system, and combines it with reinforcement learning algorithms to quickly respond to order priority fluctuations and equipment failure signals in highly dynamic environments, generating highly adaptable task allocation schemes, effectively solving the problem of insufficient adaptability of existing technologies in dynamic scenarios; 2. This invention comprehensively considers multiple dimensions such as task completion time, equipment utilization, energy consumption level and switching cost. Through the synergistic effect of genetic algorithm and particle swarm optimization algorithm, it achieves dynamic balance between multiple objectives, avoiding the resource waste and inefficiency caused by single objective optimization in the prior art. 3. This invention can quickly optimize and adjust the task allocation scheme in a real-time environment, verify the feasibility of instructions through the simulation execution module, and trigger an adaptive adjustment mechanism based on production line feedback data to ensure the real-time performance and effectiveness of the task allocation scheme, thereby improving production efficiency and product quality. 4. By optimizing task allocation, we reduced equipment idle time and task switching costs, improved equipment utilization and production efficiency, reduced energy consumption and production costs, and enhanced the company's market competitiveness. 5. This invention employs a reinforcement learning algorithm, which can optimize task allocation from a global perspective, avoiding the problem of poor overall performance caused by local optimization in existing technologies. At the same time, it enhances the flexibility and scalability of the system, enabling it to better meet diverse and personalized market demands. Attached Figure Description

[0008] Figure 1 This is a flowchart of a production line unit task allocation method based on reinforcement learning, as described in Example 1. Figure 2 This is a framework diagram of a production line unit task allocation system based on reinforcement learning, as shown in Example 2. Detailed Implementation

[0009] To make the objectives, features, and advantages of this invention more apparent and understandable, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described below are only some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0010] Example 1 like Figure 1 As shown, this embodiment provides a production line unit task allocation method based on reinforcement learning, including: Extract all current production line order demand data and equipment status information from sensors and databases, and determine a preliminary task sequence list for potential order priority fluctuations and equipment failure signals in a highly dynamic environment; Based on the preliminary task sequence list, calculate the multi-objective function value, filter out the combination of completion time, equipment utilization and switching cost from the task sequence list, and determine if the function value exceeds the preset threshold. If so, adjust the sequence to balance these objectives and obtain an optimized task allocation candidate set. Real-time environmental change data is obtained from the optimized task allocation candidate set. By comparing the matching degree between the change data and the candidate set, the allocation parameters that need to be updated are determined. For the updated allocation parameters, a new task allocation scheme is generated using the particle swarm optimization algorithm. The trade-off between equipment utilization and switching cost is iteratively calculated from the parameters. If the trade-off is lower than a preset threshold, the process is repeated until the condition is met, and the final task allocation matrix is ​​obtained. From the final task allocation matrix, extract the specific task instructions for each production line unit, verify the feasibility of the instructions under the completion time constraint through the simulation execution module, and determine the instruction set that has passed the verification.

[0011] Furthermore, the preliminary task sequence list is obtained by sorting the extracted data, specifically as follows: By acquiring order demand data and equipment status information from sensors and databases through a real-time monitoring system, the fluctuation amplitude of order priority fluctuations in a highly dynamic environment is determined, and fluctuation adjustment parameters are obtained. The fluctuation adjustment parameters are used to process equipment fault signals. If the fault signal exceeds a preset threshold, the equipment status information is synchronously updated from the database to determine the fault recovery sequence. Based on the fault recovery sequence fusion sensor fusion processing results, the task sequence generation is optimized by data sorting to obtain a preliminary task sequence list.

[0012] Specifically, the real-time monitoring system includes a variety of sensors deployed on the production line, such as temperature sensors, vibration sensors, and pressure sensors. These sensors are used to collect equipment operating parameters in real time, such as equipment speed, temperature changes, and load conditions. At the same time, order demand data, including order quantity, delivery deadline and material requirements, is extracted from the database; This approach ensures the timeliness and accuracy of data collection, providing a foundation for subsequent processing. Furthermore, in response to potential fluctuations in order priority in a highly dynamic environment, the system will monitor changes in order data. For example, when a sudden order is inserted or the priority of an existing order is adjusted, the system detects fluctuations by comparing the current priority value with historical values; The specific process involves calculating the priority difference; if the difference exceeds a preset threshold, it is marked as a fluctuation signal. This detection mechanism helps to respond quickly to changes in market demand in the production line environment and avoid production delays; Preferably, for the identification of equipment fault signals, the system analyzes the status information extracted from the sensors; In one possible implementation, if the vibration sensor detects an abnormal increase in amplitude, it combines the temperature sensor data to determine whether the fault is caused by overheating. It should be noted that the determination of fault signals is based on multi-sensor data fusion, such as integrating readings from different sensors through a weighted averaging method to generate a comprehensive fault index. This method is applicable to various equipment in manufacturing production lines, such as assembly line robots or conveyor belt systems, to ensure that faults are detected early. Based on the extracted data, the system determines a preliminary list of task sequences; Specifically, the task sequence list is generated through a sorting algorithm, such as a priority-based sorting method that combines order requirements with equipment status. First, list all tasks to be performed, and then sort them according to order priority and device availability; High-priority orders are placed at the beginning of the sequence, while tasks related to faulty equipment are delayed or reassigned; This sorting process demonstrates flexibility in highly dynamic environments and can adapt to sudden changes; For example, on an electronic product assembly line, when the equipment status shows that a welding machine has malfunctioned, the system will retrieve the availability information of backup equipment from the database and adjust the task sequence to transfer the welding task to the backup machine. Meanwhile, considering order priority fluctuations, if a high-end order suddenly gains priority, the system will reorder the sequence to ensure that the task for that order is executed first. Through this implementation, the production line maintains continuous operation; In another embodiment, for a food processing production line, the real-time monitoring system extracts data from humidity and speed sensors to detect equipment status such as abnormal conveyor belt speed, which may indicate a fault signal. In response to order demand, such as the priority fluctuation of bulk bread orders, the system calculates the degree of fluctuation, for example, by quantifying it through the rate of change of priority scores; Then, generate a task sequence list, prioritizing the baking tasks of high-priority orders; This scenario demonstrates the versatility of technological solutions within the same manufacturing field; Furthermore, the data extraction process must ensure real-time performance; In one implementation, sensor data is transmitted wirelessly to a central database and updated every second, while order data is retrieved from an enterprise resource planning system. This integration method supports rapid response in highly dynamic environments. For example, in an automotive parts production line, when a stamping machine malfunction signal occurs, the system immediately extracts data and sorts tasks to avoid production interruption. Understandably, the sorting of the task sequence list is based on multiple factors. The specific process includes first filtering out valid tasks, and then applying sorting rules, such as priority descending order combined with device load balancing. For example, if two orders have the same priority, the device with the lower load will be selected to execute the order based on the device status. This detailed sorting mechanism improves the overall efficiency of the production line and maintains stable output amidst dynamic changes; For example, in a textile production line, the system monitors the temperature sensor data of the dyeing equipment. If a fault signal caused by fluctuation is detected, the task sequence will be adjusted to move the dyeing task to a backup device. Meanwhile, fluctuations in order priority, such as urgent export orders, will trigger sequence rearrangement to ensure timely delivery; This method achieves resource optimization through data-driven sorting; Overall, through the above steps, the system can effectively handle dynamic factors in the manufacturing line, generate a reliable list of task sequences, and support efficient production.

[0013] Furthermore, the optimized task assignment candidate set includes: Using the task sequence list, a genetic algorithm is used to calculate the multi-objective function value. The genetic algorithm processes the sequence through selection, crossover, and mutation operations to obtain the multi-objective function value. From this, sequences containing the combination of completion time, the equipment utilization rate value, and the switching cost combination are selected to obtain the preliminary selected sequences. For the preliminary screening sequence, its function value is obtained, wherein the function value is obtained by weighted summation of the combination of completion time, the combination of equipment utilization rate and the combination of switching cost; If the function value exceeds a preset threshold, the order of tasks in the sequence is adjusted to balance the combination of completion times and the device utilization value. The adjustment is achieved by swapping the positions of adjacent tasks, and the adjusted sequence is determined. Based on the adjusted sequence, its switching cost combination is obtained, wherein the switching cost combination is obtained by accumulating the task switching overhead between sequences; If the switching cost combination exceeds a preset threshold, the task positions are further swapped to reduce the switching cost combination, resulting in an optimized sequence set. From the optimized sequence set, the task allocation candidates are fused using a dynamic resource allocation method, wherein the dynamic resource allocation is achieved by allocating device resources according to the completion time combination to obtain an optimized task allocation candidate set.

[0014] Specifically, in one implementation, the initial task sequence list is generated by collecting information on various tasks within the production workshop.

[0015] For example, these tasks include parts processing, equipment maintenance, and material transportation on the assembly line, aiming to provide basic data for subsequent optimization.

[0016] Specifically, the list records the estimated execution time, required equipment type, and possible sequence arrangements for each task to ensure efficient scheduling within the manufacturing field. Furthermore, the process of using a genetic algorithm to calculate multi-objective function values ​​first requires understanding the basic principles of genetic algorithms. It is a search optimization method that simulates natural evolution, finding the optimal solution through population iteration. In this embodiment, the genetic algorithm is applied to task sequence optimization. The initial population consists of multiple randomly generated sequence combinations, each representing a task allocation scheme.

[0017] For example, the population size can be set to 50 to 100 individuals to cover sufficient diversity.

[0018] It should be noted that the calculation of the multi-objective function value involves a comprehensive evaluation of three indicators: completion time, equipment utilization, and switching cost. Completion time refers to the total duration of all tasks from start to finish, obtained by summing the execution time of each task and considering parallel or serial relationships; equipment utilization reflects the proportion of idle time of the equipment.

[0019] For example, it can be quantified by calculating the ratio of total equipment uptime to available time; switchover costs include the preparation time and resource consumption required for task transitions, such as the cost of equipment reset or worker adjustments. These metrics are weighted and summed to form a function value, with weights set according to actual production needs to balance different objectives.

[0020] In one possible implementation, the specific steps for filtering out combinations from the list of task sequences are as follows: First, the initial list is encoded, with each sequence represented as a chromosome.

[0021] For example, a sequence might be encoded as a string of numbers, where each number corresponds to a task ID. Then, the crossover operation of a genetic algorithm is applied, swapping portions of two parent sequences to generate offspring, in order to explore new combinations.

[0022] For example, in a manufacturing workshop scenario, if task A is part grinding and task B is assembly, the intersection may generate an optimized path from A to B, reducing switching costs.

[0023] Preferably, the mechanism for determining whether a function value exceeds a preset threshold is achieved by comparing the calculated multi-objective value with the threshold, which can be set based on historical production data.

[0024] For example, if the function value exceeds a threshold, it indicates that the current sequence is unbalanced in terms of completion time or cost, triggering an adjustment. The process of adjusting the sequence to balance these objectives involves mutation operations, where mutation introduces diversity by randomly changing the order of tasks in the sequence in genetic algorithms.

[0025] For example, moving a task from the middle of the sequence to the end can reduce the total completion time.

[0026] Specifically, within an iteration cycle, the algorithm evaluates the fitness of all individuals, and the fitness function integrates three objectives.

[0027] For example, fitness = w1 * (1 / completion time) + w2 * equipment utilization - w3 * switching cost, where w1, w2, and w3 are weighting coefficients to ensure balanced optimization.

[0028] Understandably, the iterative process of genetic algorithms typically runs for multiple generations, such as 10 to 50, until it converges to the Pareto optimal front, a non-dominated set of solutions where no single solution is superior to all others on all objectives. In production scheduling, this front helps in selecting a compromise sequence of tasks.

[0029] For example, one solution might prioritize equipment utilization, while another focuses on minimizing switching costs, thus generating multiple candidate solutions.

[0030] For example, in a specific manufacturing scenario, suppose there are 5 tasks to be assigned to 3 machines. The initial sequence list might include a sequence where task 1 precedes task 2. After calculation by the genetic algorithm, if the function value indicates that the completion time exceeds a threshold, the process is adjusted to execute tasks 1 and 3 in parallel to improve machine utilization. This adjustment not only balances the objectives but also reduces the economic losses caused by production delays. Furthermore, the optimized task assignment candidate set is formed by selecting elite individuals after the algorithm converges; these individuals represent the best sequence combinations under multiple objectives.

[0031] In one embodiment, the size of the candidate set can be controlled to within 10, for decision-makers to choose from.

[0032] For example, the sequence with the lowest switching cost can be selected based on the actual workshop load.

[0033] Specifically, throughout the process, the implementation of the genetic algorithm emphasizes its versatility in task scheduling.

[0034] For example, population parameters can be adjusted across different production batches to adapt to changes in task volume, demonstrating the flexibility of the technical solution. This method enables efficient allocation of manufacturing resources, avoiding waste caused by equipment idleness or excessive switching.

[0035] In one embodiment, the results show that the optimization method reduced the average completion time by 15% and increased equipment utilization to over 85% in a simulated production environment, while keeping the switching costs within budget. These results were obtained by objectively comparing the initial and optimized sequences, demonstrating the practical value of the solution in the manufacturing field.

[0036] Furthermore, the determination of the allocation parameters that need to be updated specifically includes: By obtaining real-time environmental change data from the optimized task allocation candidate set, and comparing the similarity between the sudden order insertion and equipment maintenance notification and the candidate set, the matching degree is determined. Based on the matching degree comparison results, the deviation between the allocated parameters and the changed data is compared to determine the allocated parameters that need to be updated, and a parameter set reflecting the highly dynamic environment is obtained. Also includes: The adaptive indicators in the parameter set are obtained, and the expanded task redistribution scheme is determined by fusing them with the real-time environmental change data.

[0037] Specifically, in one implementation, the optimized task allocation candidate set refers to a set of multiple task allocation schemes generated by an initial optimization algorithm in the logistics scheduling system. These schemes, for example, include different combinations of vehicle allocations for orders, taking into account static factors such as distance and load. When acquiring real-time environmental change data, the system first monitors external input interfaces to capture sudden order insertions, i.e., temporarily added new delivery demands, or equipment maintenance notifications, such as unavailability caused by sudden vehicle malfunctions.

[0038] Specifically, the system can collect this data in real time through integrated sensors or external APIs to ensure data timeliness. Furthermore, the impact is assessed by comparing the degree of matching between the changing data and the candidate set.

[0039] For example, the matching degree can be calculated based on key indicators, such as comparing the location and time of a sudden order with the existing allocated paths in the candidate set, calculating the deviation value, and considering it as a low match if the path overlap rate is lower than a threshold.

[0040] It should be noted that this comparison process involves quantitative analysis. First, features of the changing data are extracted, such as order urgency and location coordinates. Then, these features are compared item by item with parameters in the candidate set, such as vehicle availability and route planning, to identify any mismatches. Based on this comparison, the allocation parameters that need to be updated are further determined.

[0041] In one possible implementation, the system analyzes parameters with high mismatches. For example, if a vehicle becomes unavailable due to a device maintenance notification, the relevant allocation parameters, such as vehicle ID and load allocation, are marked as items that need to be updated.

[0042] Understandably, this step is implemented through a threshold judgment mechanism. For example, a matching threshold of 80% is set, and an update is triggered if the value is lower than this, ensuring that adjustments are made only to parameters that have a significant impact.

[0043] Preferably, the updated parameters reflect the adaptability to highly dynamic environments.

[0044] Specifically, in logistics and delivery scenarios, the update process can use heuristic algorithms to adjust parameters, such as reallocating sudden orders to backup vehicles and optimizing routes to minimize delays. This adaptability is reflected in the system's ability to respond quickly to changes. For example, when inserting orders during peak periods, the overall efficiency remains above 95% of the original level after updating parameters, thus supporting continuous operation. In another embodiment, task allocation in the manufacturing field is considered, such as process allocation on an assembly line. When acquiring real-time change data, the system monitors equipment maintenance notifications, such as machine downtime alarms, and compares them with the candidate set to calculate the impact of process delays. If the matching degree is low, update the allocation parameters such as process sequence and personnel configuration to adapt to the dynamic environment and ensure production continuity; Further expanding, in e-commerce warehousing task allocation, the insertion of sudden orders can trigger data acquisition, and the matching degree comparison focuses on the fit between inventory location and allocation plan; After determining the update parameters, the system adjusts parameters such as robot path parameters to reflect adaptability and achieve efficient response by reducing waiting time.

[0045] For example, in the scenario described above, the matching degree calculation process includes the following steps: first, standardizing the changing data, such as converting order insertions into vector representations; then, calculating the cosine similarity with the candidate set vectors to ensure objective evaluation of bias. This method enhances the robustness of the system.

[0046] In one embodiment, after the parameters are updated, the adaptability can be verified, for example by simulating a high-dynamic environment test to confirm that the updated parameters still maintain a balanced distribution as the frequency of change increases.

[0047] It should be noted that these implementation methods are all limited to the field of task scheduling, such as logistics or manufacturing, to ensure versatility without exceeding the original context.

[0048] Furthermore, obtaining the final task allocation matrix specifically includes: The allocation parameters are updated, and a particle swarm optimization algorithm is used, where particles represent task allocation variables. The swarm searches for the optimal position through velocity updates to generate the task allocation scheme and obtain the initial values ​​for equipment utilization calculation and switching cost assessment. The equipment utilization calculation is based on the ratio of task load to equipment capacity, and the switching cost assessment is based on the product of the number of task migrations and migration costs. Based on the initial value, the trade-off value between the equipment utilization calculation and the switching cost assessment is iteratively calculated from the task allocation scheme, wherein the trade-off value is determined by summing the equipment utilization calculation multiplied by a first weight and the switching cost assessment multiplied by a second weight; If the tradeoff value is lower than a preset threshold, the particle swarm optimization algorithm is iterated again for the tradeoff value, wherein the particle positions are updated to minimize the tradeoff value, and the adjusted task allocation scheme is obtained. Based on the adjusted task allocation scheme, it is determined that the trade-off value meets the condition, wherein meeting the condition means that the trade-off value is not lower than a preset threshold, and the final matrix is ​​obtained to generate the corresponding task allocation matrix.

[0049] Specifically, in one implementation, for the updated allocation parameters, the parameters of the particle swarm optimization algorithm are first initialized, including the number of particles, initial positions, and velocities. These positions represent possible task allocation schemes, with each scheme corresponding to a task-to-device mapping. Particle swarm optimization is a swarm intelligence-based optimization method that searches for optimal solutions by simulating the foraging behavior of bird flocks. Each particle updates its velocity and position based on its historical best position and the global best position, thereby generating a new task allocation scheme.

[0050] Specifically, information such as task requirements and equipment capacity are extracted from the updated allocation parameters and used as input for the initial position of the particles.

[0051] For example, in a cloud computing task scheduling scenario, allocation parameters might include the task's computational load and the device's processing capacity, while particle positions represent the task-to-virtual machine allocation matrix. During algorithm iteration, a fitness function is calculated for each particle, which integrates a trade-off between device utilization and switching costs. Device utilization refers to the proportion of actual device usage time to total available time, calculated as the ratio of task execution time to device capacity; switching costs involve the overhead incurred when migrating a task from one device to another, such as data transmission latency and restart time, which can be quantified using a predefined cost matrix. Furthermore, the trade-off is calculated using a weighted summation method.

[0052] For example, a comprehensive index is obtained by multiplying the equipment utilization rate by a weighting factor and then subtracting the weighted value of switching costs. If this tradeoff value is lower than a preset threshold, such as 0.8, it indicates that the current solution has not achieved the expected balance. The algorithm then iterates again, updates the particle velocity and position, and continues to generate new solutions. This iterative process continues until the tradeoff value meets the condition, outputting the final task allocation matrix. This matrix uses rows to represent tasks and columns to represent equipment, with an element value of 1 indicating allocation and 0 otherwise. Preferably...

[0053] In one possible implementation, multi-objective optimization is considered, and an inertial weight parameter is introduced to adjust the particle velocity update formula, thereby improving the algorithm's convergence speed.

[0054] For example, in edge computing task allocation, when a task involves real-time data processing, the solution with low switching costs is prioritized to reduce latency. In this way, the technical solution can adapt to different load scenarios within the same domain, such as high-concurrency tasks or low-load maintenance, achieving allocation flexibility.

[0055] It should be noted that the effectiveness of this method lies in maximizing equipment utilization and minimizing switching costs through iterative optimization. In actual task scheduling, it can improve the overall system efficiency without introducing additional complexity.

[0056] For example, in another embodiment, for task allocation on mobile devices, updated parameters include remaining battery power and network bandwidth. The particle swarm optimization algorithm also calculates trade-off values, iterates if the values ​​are below a threshold, and finally guides task migration with a matrix to balance energy consumption and performance.

[0057] Understandably, this iterative judgment mechanism enhances the robustness of the solution and makes it suitable for various computing resource allocation scenarios.

[0058] Furthermore, the instruction to confirm successful verification specifically includes: Extract production line unit instructions from the task allocation matrix, and simulate the running trajectory of the production line unit instructions under time constraint checks through the execution module to obtain the feasibility determination result; For the portion of the feasibility determination results that has been verified, multi-objective conflict resolution is incorporated to determine the instruction set after the conflict resolution is incorporated. Obtain simulation data from the execution module, perform simulation based on the instruction set after conflict resolution and integration, and determine the verification set. If the verified set satisfies the time constraint check, then the final instruction set is extracted from the verified set to obtain the instruction set extraction result.

[0059] Specifically, in one implementation, extracting the specific task instructions for each production line unit from the final task allocation matrix first requires understanding the structure of the task allocation matrix. This matrix is ​​typically represented in rows and columns, where rows correspond to different production line units, such as robotic arm units or welding units on an assembly line, and columns correspond to task sequences, such as part assembly or quality inspection. The extraction process involves traversing the matrix elements to identify the task instructions assigned to each unit.

[0060] For example, for a robotic arm unit, the instructions might include grasping a specific part and moving it to a designated location. This extraction ensures a one-to-one correspondence between instructions and production line units, providing a foundation for subsequent verification. Furthermore, the extracted task instructions are verified using a simulation execution module to check their feasibility under completion time constraints. The simulation execution module is a virtual environment simulator that builds a model based on actual production line parameters, such as unit processing speed, task dependencies, and time windows.

[0061] For example, in the simulation, instructions are input into the module, which executes the virtual tasks one by one, calculating the time taken for each step. For instance, it might take 5 seconds for the robotic arm to grasp a part, combined with a total time constraint, such as the entire production line cycle not exceeding 30 minutes. Feasibility is determined by comparing the total simulation time with the constraint value. If the simulation time exceeds the constraint, it is marked as infeasible.

[0062] It should be noted that the resolution of multi-objective conflicts has been incorporated into the instruction set.

[0063] Specifically, multi-objective conflicts may include contradictions between time optimization and resource allocation. For example, a task instruction may require high-speed execution but increase energy consumption. The resolution process is achieved through prioritization or trade-off mechanisms, adjusting allocations during matrix generation. For instance, the speed of some non-critical tasks may be reduced to meet overall time constraints, resulting in an optimized instruction set. This integration ensures that conflicts are considered during instruction set extraction, avoiding later adjustments.

[0064] In one possible implementation, once the verified instruction set is determined, it can be applied to actual production line scenarios.

[0065] For example, on an electronics assembly line, instructions are extracted and simulated for verification. If all units complete their tasks within a specified time, that instruction set is selected. The simulation module records potential bottlenecks, such as task backlog in a particular unit, and outputs analysis results through logs to aid in optimization.

[0066] Preferably, the simulation execution module is constructed with production line dynamics in mind.

[0067] Specifically, this module uses a state machine model to represent unit state transitions, such as the change from idle to task execution. During verification, the module simulates concurrent tasks, such as multiple units operating simultaneously, and calculates the cumulative time. In this way, it ensures that the feasibility of the instruction set is not only for single tasks but also for overall coordination.

[0068] For example, when validating a welding unit's task instructions, the simulation module inputs instruction parameters, such as welding duration and cooling time, combined with time constraints such as not exceeding 10 minutes, and runs the simulation. If the simulation shows a total duration of 8 minutes, the validation is successful. This detailed simulation helps identify hidden problems, such as inter-task interference, thereby improving production line efficiency.

[0069] Understandably, the final determination of the instruction set depends on multiple iterative simulations.

[0070] In one embodiment, if the initial verification fails, the system backtracks to the matrix adjustment conflict, re-extracts the instructions, and verifies them until the conditions are met. This iterative process ensures the robustness of the instruction set and reduces downtime in practical applications.

[0071] Specifically, details of multi-objective conflict resolution include using weighted functions to evaluate objectives.

[0072] For example, a time objective is assigned a weight of 1, and a resource objective is assigned a weight of 0.8. The optimal allocation scheme is selected by calculating the total score. This method is reflected in the instruction set as adjusted task parameters, ensuring that conflicts are resolved during verification. Furthermore, this method is also applicable in production line expansion scenarios, such as adding new units. When extracting the matrix, a new row is included, and simulation verification checks the overall time constraints, thereby demonstrating the versatility of the solution.

[0073] In one embodiment, the verified instruction set, once output, can be directly imported into the production line control system for automated execution. This technical feature, corresponding to the claims, ensures a complete process from extraction to verification, providing efficient task management in the production line field.

[0074] Furthermore, the revised allocation strategy includes: By obtaining execution time records and energy consumption data monitoring through production line feedback, the difference between the actual value and the planned value is calculated from the execution time records and energy consumption data monitoring as the deviation value analysis to obtain the deviation amount; For the deviation amount, if it is determined that the deviation amount is greater than a preset threshold, an adjustment mechanism is triggered to determine the updated parameters. Based on the updated parameters, an adaptive parameter update correction allocation strategy is adopted to obtain the corrected allocation strategy.

[0075] Specifically, in one implementation, the system first obtains production line feedback data based on the verified instruction set; A verified instruction set refers to a set of task instructions that have passed security and logic verification. These instructions are used to guide the operation of production line equipment. For example, in a manufacturing production line, the instruction set may include the movement path of a robotic arm or an assembly sequence. The feedback data is obtained through a sensor network, such as real-time collection of actual execution time, i.e., the specific duration from task start to completion, and energy consumption records, such as the power consumption value of the device during operation. These data are extracted from the production line database to ensure accuracy and timeliness. In this way, the system can collect information reflecting the actual operating status of the production line, providing a basis for subsequent analysis; Furthermore, analyzing deviations from feedback data involves comparing the differences between actual and expected values; The process of deviation analysis begins by defining expected values. For example, the expected execution time is based on historical averages or simulation calculations, while the expected energy consumption is estimated based on equipment specifications and task load.

[0076] Specifically, deviation calculations can employ difference formulas, such as subtracting the expected time from the actual execution time to obtain the time deviation, and similarly calculating energy consumption deviation. This analysis helps identify anomalies in the production line; for example, a significant extension of actual time may indicate equipment failure or material delays. Statistical methods, such as calculating the mean and standard deviation of the deviation, further quantify the degree of deviation, ensuring the reliability of the analysis results.

[0077] It's important to note that determining whether the deviation exceeds a preset threshold is based on a threshold comparison using the analysis results. The preset threshold can be set according to the production line type. For example, in an automotive assembly line, the time deviation threshold might be set to 5 minutes, and the energy consumption threshold to 10% exceeding expectations. This step is implemented through conditional logic; if the deviation exceeds the threshold, it is marked as an abnormal state. This judgment process is simple and efficient, enabling rapid response to production line changes and preventing small deviations from accumulating into major problems.

[0078] In one possible implementation, if the deviation exceeds a preset threshold, an adaptive adjustment mechanism is triggered. This mechanism is the core component, designed to dynamically optimize production line resource allocation. The specific process includes first assessing the cause of the deviation, such as analyzing log data to determine if the time deviation stems from excessive equipment load. Then, the mechanism activates an optimization algorithm, such as using a genetic algorithm to simulate multiple adjustment schemes and iteratively select the optimal path.

[0079] For example, in an electronics assembly line, if energy consumption deviations are too large, the mechanism may adjust machine speeds or reallocate tasks to backup equipment. Furthermore, adaptive adjustment involves multiple feedback loops: after the initial adjustment, data is collected again to compare new deviations; if the deviation still exceeds a threshold, the iteration continues until it converges. The principle behind this mechanism is to use real-time data to drive decision-making, enabling the production line to respond flexibly. In this way, not only can current problems be corrected, but overall efficiency can also be improved, such as reducing energy waste by up to 15%, thereby supporting continuous production needs.

[0080] Preferably, the corrected allocation strategy is obtained through the output of an adaptive adjustment mechanism. The corrected strategy specifically includes updating the task allocation table, for example, transferring high-energy-consuming tasks to more efficient devices, or adjusting the execution order to shorten the total time.

[0081] In one embodiment, for a textile production line, if feedback indicates a large deviation in the execution time of a certain machine, the strategy may offload some tasks to parallel machines to ensure load balancing. This strategy is generated based on the calculation results of the adjustment mechanism and stored in the system for future use.

[0082] For example, in a semiconductor manufacturing production line scenario, after the system obtains feedback data, it analyzes the deviation. If the energy consumption record shows that it exceeds the threshold, the mechanism is triggered to adjust the cooling system parameters and obtain correction strategies such as optimizing the chip test sequence, thereby maintaining the stability of the production line.

[0083] Understandably, the above process ensures the versatility of the technical solution in the manufacturing sector; for example, in food processing lines, time deviations can be analyzed to adjust packaging speed. Furthermore, this adaptive mechanism can handle multivariate deviations, such as considering the combined deviations of time and energy consumption, improving the accuracy of adjustments through weighted calculation of a comprehensive index. In another implementation, deviation analysis can be incorporated into machine learning models to predict potential deviations, but still limited to feedback data.

[0084] For example, the entire process, from data acquisition to strategy correction, forms a closed loop, supporting long-term production line optimization.

[0085] Furthermore, the execution monitoring loop after determining the push includes: The optimized resource allocation indicators are obtained from the source of the strategy correction, and the indicators are pushed to the production line equipment through the real-time distribution system to determine the execution monitoring loop after the push. During the execution of the monitoring cycle, dynamic adjustment data of the production line is acquired, and if the allocation efficiency assessment is lower than a preset threshold, the cycle risk tracking is performed to address the risk of low efficiency. Through the cyclical risk tracking, risk reduction is confirmed, and a backup allocation plan is generated for the production line through dynamic adjustments, determining the optimal extraction of resources for the backup plan.

[0086] Furthermore, it also includes: Based on the verified instruction set, obtain production line feedback data such as actual execution time and energy consumption records, analyze the deviation from the feedback data, and determine if the deviation is greater than the preset threshold. If so, trigger the adaptive adjustment mechanism to obtain the corrected allocation strategy. From the revised allocation strategy, the optimized resource allocation indicators are extracted and pushed to the production line equipment through the real-time distribution system to determine the execution monitoring loop after the push.

[0087] Furthermore, the revised allocation strategy specifically includes: By obtaining execution time records and energy consumption data monitoring through production line feedback, the difference between the actual value and the planned value is calculated from the execution time records and energy consumption data monitoring as the deviation value analysis to obtain the deviation amount; For the deviation amount, if it is determined that the deviation amount is greater than a preset threshold, an adjustment mechanism is triggered to determine the updated parameters. Based on the updated parameters, an adaptive parameter update and correction allocation strategy is adopted to obtain the corrected allocation strategy.

[0088] Specifically, in one implementation, the system first acquires production line feedback data based on a validated instruction set. The validated instruction set refers to a set of task instructions that have undergone safety and logical verification. These instructions guide the operation of production line equipment; for example, in a manufacturing production line, the instruction set might include the movement path of a robotic arm or an assembly sequence. The acquisition of feedback data is specifically achieved through a sensor network, such as real-time collection of actual execution time (the specific duration from task initiation to completion) and energy consumption records, such as the power consumption of the equipment during operation. This data is extracted from the production line database to ensure accuracy and real-time performance. In this way, the system can collect information reflecting the actual operating status of the production line, providing a basis for subsequent analysis. Further, analyzing deviations from the feedback data involves comparing the differences between actual and expected values. The deviation analysis process first defines expected values; for example, expected execution time is derived based on historical averages or simulation calculations, while expected energy consumption is estimated based on equipment specifications and task load.

[0089] Specifically, deviation calculations can employ difference formulas, such as subtracting the expected time from the actual execution time to obtain the time deviation, and similarly calculating energy consumption deviation. This analysis helps identify anomalies in the production line; for example, a significant extension of actual time may indicate equipment failure or material delays. Statistical methods, such as calculating the mean and standard deviation of the deviation, further quantify the degree of deviation, ensuring the reliability of the analysis results.

[0090] It's important to note that determining whether the deviation exceeds a preset threshold is based on a threshold comparison using the analysis results. The preset threshold can be set according to the production line type. For example, in an automotive assembly line, the time deviation threshold might be set to 5 minutes, and the energy consumption threshold to 10% exceeding expectations. This step is implemented through conditional logic; if the deviation exceeds the threshold, it is marked as an abnormal state. This judgment process is simple and efficient, enabling rapid response to production line changes and preventing small deviations from accumulating into major problems.

[0091] In one possible implementation, if the deviation exceeds a preset threshold, an adaptive adjustment mechanism is triggered. This mechanism is the core component, designed to dynamically optimize production line resource allocation. The specific process includes first assessing the cause of the deviation, such as analyzing log data to determine if the time deviation stems from excessive equipment load. Then, the mechanism activates an optimization algorithm, such as using a genetic algorithm to simulate multiple adjustment schemes and iteratively select the optimal path.

[0092] For example, in an electronics assembly line, if energy consumption deviations are too large, the mechanism may adjust machine speeds or reallocate tasks to backup equipment. Furthermore, adaptive adjustment involves multiple feedback loops: after the initial adjustment, data is collected again to compare new deviations; if the deviation still exceeds a threshold, the iteration continues until it converges. The principle behind this mechanism is to use real-time data to drive decision-making, enabling the production line to respond flexibly. In this way, not only can current problems be corrected, but overall efficiency can also be improved, such as reducing energy waste by up to 15%, thereby supporting continuous production needs.

[0093] Preferably, the corrected allocation strategy is obtained through the output of an adaptive adjustment mechanism. The corrected strategy specifically includes updating the task allocation table, for example, transferring high-energy-consuming tasks to more efficient devices, or adjusting the execution order to shorten the total time.

[0094] In one embodiment, for a textile production line, if feedback indicates a large deviation in the execution time of a certain machine, the strategy may offload some tasks to parallel machines to ensure load balancing. This strategy is generated based on the calculation results of the adjustment mechanism and stored in the system for future use.

[0095] For example, in a semiconductor manufacturing production line scenario, after the system obtains feedback data, it analyzes the deviation. If the energy consumption record shows that it exceeds the threshold, the mechanism is triggered to adjust the cooling system parameters and obtain correction strategies such as optimizing the chip test sequence, thereby maintaining the stability of the production line.

[0096] Understandably, the above process ensures the versatility of the technical solution in the manufacturing sector; for example, in food processing lines, time deviations can be analyzed to adjust packaging speed. Furthermore, this adaptive mechanism can handle multivariate deviations, such as considering the combined deviations of time and energy consumption, improving the accuracy of adjustments through weighted calculation of a comprehensive index. In another implementation, deviation analysis can be incorporated into machine learning models to predict potential deviations, but still limited to feedback data.

[0097] Furthermore, the execution monitoring loop after determining the push specifically includes: The optimized resource allocation indicators are obtained from the source of the strategy correction, and the indicators are pushed to the production line equipment through the real-time distribution system to determine the execution monitoring loop after the push. During the execution of the monitoring cycle, dynamic adjustment data of the production line is acquired, and if the allocation efficiency assessment is lower than a preset threshold, the cycle risk tracking is performed to address the risk of low efficiency. Through the cyclical risk tracking, risk reduction is confirmed, and a backup allocation plan is generated for the production line through dynamic adjustments, determining the optimal extraction of resources for the backup plan.

[0098] Specifically, in one implementation, the process of extracting optimized resource allocation indicators from the revised allocation strategy first involves processing the revision mechanism of the allocation strategy. The revised allocation strategy is typically adjusted based on historical data and real-time feedback from the production line.

[0099] For example, in a manufacturing production line, strategies might include optimizing equipment utilization and material allocation ratios. A data extraction module reads key parameters, such as resource allocation ratios and priority indicators, from the strategy file. These indicators are extracted into a structured data format for subsequent processing.

[0100] It should be noted that the extraction process ensures the accuracy of the indicators and verifies data integrity through validation algorithms, thus laying the foundation for optimized resource allocation. Furthermore, these indicators are pushed to production line equipment through a real-time distribution system. This real-time distribution system can adopt an IoT-based framework.

[0101] For example, the metric data can be packaged into a message queue using a wireless network and pushed to the device terminal.

[0102] In one possible implementation, the system includes a central server and edge nodes. After receiving metrics from the extraction module, the server immediately distributes them to automated equipment on the production line, such as assembly robots or conveyor belt controllers. This push mechanism supports low-latency transmission, ensuring that metrics reach the equipment within seconds, thereby enabling real-time application of optimized allocation during production.

[0103] For example, determining the execution monitoring loop after a push involves setting monitoring parameters. After the push is completed, the system starts a loop process that checks the device execution status at fixed time intervals.

[0104] For example, in an automotive parts production line, a monitoring cycle collects equipment operating data, such as production speed and resource consumption rate, and compares it with pushed metrics. If the deviation exceeds a threshold, an alarm is triggered. This determination process is automated through scripts, ensuring the continuity of monitoring.

[0105] Specifically, the process of continuously tracking inefficiency risks in the monitoring loop needs to be explained in detail. Inefficiency risks refer to potential problems where production efficiency falls below preset standards, such as equipment idleness or resource waste. The monitoring loop uses an iterative algorithm to continuously collect real-time data.

[0106] For example, sensors monitor the load and output of production line equipment, and then a risk index is calculated. The calculation process involves comparing the current efficiency value with the optimization target; when the accumulated difference exceeds a certain level, it is identified as a risk. Specifically...

[0107] In one embodiment, the loop runs once per minute, tracking variables including resource utilization and failure rate, and assessing risk trends by accumulating deviation values. This tracking mechanism helps identify problems early and achieve stable control of the production process.

[0108] Preferably, in another implementation, the extraction of optimized resource allocation metrics can be aided by a machine learning model. The model is trained based on historical allocation data to predict the optimal metric value, which is then extracted from the correction strategy.

[0109] For example, on an electronics assembly line, the model analyzes past strategy correction records to extract optimization metrics such as manpower allocation and machine time. This approach enhances the intelligence of the extraction and supports more complex production line scenarios.

[0110] Understandably, push notifications from a real-time distribution system can be implemented through a cloud platform. After a push notification, the determination of the monitoring loop includes a logging step, recording the execution result of each push event for loop initialization. The monitoring loop then continuously tracks risks through a data stream engine.

[0111] For example, the engine processes sensor data streams, calculates efficiency metrics in real time, and adjusts allocations when risks increase.

[0112] For example, in a food processing production line, indicators are extracted from the corrective strategy and pushed to the packaging equipment, with monitoring cycles tracking risks such as inefficiencies caused by material waste. Through this cycle, the system can achieve dynamic risk management. Furthermore, the continuous tracking of the monitoring cycle can be extended to multi-device coordination.

[0113] For example, the loop not only tracks individual devices but also analyzes the efficiency chain of the entire production line, making overall adjustments if a risk in one area affects downstream processes. This extended demonstration of the technology's versatility allows it to be adapted to production lines of different sizes within the same field.

[0114] In one embodiment, the technical effect of the entire process is reflected in improved production efficiency. Through real-time push and monitoring, the probability of inefficiency is reduced, and the optimized execution of resource allocation is ensured.

[0115] Example 2 like Figure 2 As shown, this embodiment provides a reinforcement learning-based production line unit task allocation system to implement the reinforcement learning-based production line unit task allocation method, including: The data acquisition module is used to extract all order demand data and equipment status information of the current production line from sensors and databases, and to determine a preliminary task sequence list for possible order priority fluctuations and equipment failure signals in a highly dynamic environment. Specifically, this involves acquiring order demand data and equipment status information through a real-time monitoring system, determining the fluctuation range of order priority, processing equipment fault signals, updating equipment status information, determining fault recovery sequences, and integrating sensor processing results to optimize the data sorting of task sequence generation, thereby obtaining a preliminary task sequence list.

[0116] The optimization filtering module is used to calculate multi-objective function values ​​based on the initial task sequence list, filter out combinations that include completion time, equipment utilization, and switching costs from the task sequence list, and determine if the function value exceeds a preset threshold. If so, the sequence is adjusted to balance these objectives, resulting in an optimized task allocation candidate set. Specifically, this involves using a genetic algorithm to calculate multi-objective function values, processing sequences through selection, crossover, and mutation operations to obtain multi-objective function values, screening out preliminary sequences, adjusting the task order based on function values ​​to balance completion time and equipment utilization, further exchanging task positions to reduce switching costs, obtaining an optimized sequence set, and using a dynamic resource allocation method to merge task allocation candidates to obtain an optimized task allocation candidate set.

[0117] The parameter update module is used to obtain real-time environmental change data from the optimized task allocation candidate set, and determine the allocation parameters that need to be updated by comparing the matching degree between the change data and the candidate set. Specifically, this involves comparing the similarity between sudden order insertions and equipment maintenance notifications and the candidate set to determine the matching degree; comparing the deviation between the allocation parameters and the changing data based on the matching degree comparison results to determine the allocation parameters that need to be updated; obtaining a parameter set that reflects the highly dynamic environment; acquiring the adaptability index in the parameter set; and determining the expanded task redistribution scheme by fusing it with real-time environmental change data.

[0118] The particle swarm optimization module is used to generate a new task allocation scheme based on the updated allocation parameters using the particle swarm optimization algorithm. It iteratively calculates the trade-off between equipment utilization and switching costs from the parameters. If the trade-off is lower than a preset threshold, it iterates again until the conditions are met to obtain the final task allocation matrix. Specifically, this involves employing a particle swarm optimization algorithm, where particles represent task allocation variables. The swarm searches for the optimal position by updating its velocity, generating a task allocation scheme and obtaining initial values ​​for equipment utilization calculation and switching cost assessment. The algorithm iteratively calculates the trade-off between equipment utilization calculation and switching cost assessment from the task allocation scheme. If the trade-off is lower than a preset threshold, the particle swarm optimization algorithm is iterated again, updating particle positions to minimize the trade-off, obtaining an adjusted task allocation scheme, determining whether the trade-off meets the conditions, and obtaining the final task allocation matrix.

[0119] The instruction verification module is used to extract the specific task instructions for each production line unit from the final task allocation matrix, verify the feasibility of the instructions under the completion time constraint through the simulation execution module, and determine the instruction set that passes the verification. Specifically, this involves extracting production line unit instructions from the task allocation matrix, simulating the running trajectory of the production line unit instructions under time constraint checks through the execution module to obtain feasibility determination results, incorporating multi-objective conflict resolution into the verified parts of the feasibility determination results, determining the instruction set after conflict resolution integration, obtaining simulation data from the execution module, performing simulations based on the instruction set after conflict resolution integration, determining the verified set, and if the verified set satisfies the time constraint check, extracting the final instruction set from the verified set to obtain the instruction set extraction result.

[0120] The strategy correction module is used to obtain production line feedback data such as actual execution time and energy consumption records based on the verified instruction set, analyze the deviation from the feedback data, and determine if the deviation is greater than the preset threshold. If so, an adaptive adjustment mechanism is triggered to obtain the corrected allocation strategy. From the corrected allocation strategy, the optimized resource allocation indicators are extracted and pushed to the production line equipment through the real-time distribution system to determine the execution monitoring loop after the push. Specifically, this involves obtaining execution time records and energy consumption data monitoring through production line feedback; calculating the difference between actual and planned values ​​from these records as a deviation value analysis to obtain the deviation amount; determining if the deviation amount exceeds a preset threshold, triggering an adjustment mechanism to update parameters; using the updated parameters, employing adaptive parameter updates to correct the allocation strategy to obtain the corrected allocation strategy; obtaining optimized resource allocation indicators from the strategy correction source; pushing these indicators to production line equipment through a real-time distribution system; determining the execution monitoring loop after the push; acquiring dynamic adjustment data from the production line during the execution monitoring loop; determining if the allocation efficiency assessment is lower than a preset threshold, then performing cyclical risk tracking to address the low efficiency risk; confirming risk reduction through cyclical risk tracking; generating a backup allocation plan based on the dynamic adjustments of the production line; and determining the resource optimization extraction for the backup plan.

Claims

1. A production line unit task allocation method based on reinforcement learning, characterized in that, include: Extract all current order demand data and equipment status information from sensors and databases, and determine a preliminary task sequence list for potential order priority fluctuations and equipment failure signals in a highly dynamic environment; Based on the preliminary task sequence list, calculate the multi-objective function value, filter out the combination of completion time, equipment utilization and switching cost from the task sequence list, and determine if the function value exceeds the preset threshold. If so, adjust the sequence to balance these objectives and obtain an optimized task allocation candidate set. Real-time environmental change data is obtained from the optimized task allocation candidate set. By comparing the matching degree between the change data and the candidate set, the allocation parameters that need to be updated are determined. For the updated allocation parameters, a new task allocation scheme is generated using the particle swarm optimization algorithm. The trade-off between equipment utilization and switching cost is iteratively calculated from the parameters. If the trade-off is lower than a preset threshold, the process is repeated until the condition is met, and the final task allocation matrix is ​​obtained. From the final task allocation matrix, extract the specific task instructions for each production line unit, verify the feasibility of the instructions under the completion time constraint through the simulation execution module, and determine the instruction set that has passed the verification.

2. The production line unit task allocation method based on reinforcement learning according to claim 1, characterized in that, The preliminary task sequence list is obtained by sorting the extracted data, specifically as follows: By acquiring order demand data and equipment status information from sensors and databases through a real-time monitoring system, the fluctuation amplitude of order priority fluctuations in a highly dynamic environment is determined, and fluctuation adjustment parameters are obtained. The fluctuation adjustment parameters are used to process equipment fault signals. If the fault signal exceeds a preset threshold, the equipment status information is synchronously updated from the database to determine the fault recovery sequence. Based on the fault recovery sequence fusion sensor fusion processing results, the task sequence generation is optimized by data sorting to obtain a preliminary task sequence list.

3. The production line unit task allocation method based on reinforcement learning according to claim 1, characterized in that, The optimized task assignment candidate set includes: Using the task sequence list, a genetic algorithm is used to calculate the multi-objective function value. The genetic algorithm processes the sequence through selection, crossover, and mutation operations to obtain the multi-objective function value. From this, sequences containing the combination of completion time, the equipment utilization rate value, and the switching cost combination are selected to obtain the preliminary selected sequences. For the preliminary screening sequence, its function value is obtained, wherein the function value is obtained by weighted summation of the combination of completion time, the combination of equipment utilization rate and the combination of switching cost; If the function value exceeds a preset threshold, the task order in the sequence is adjusted to balance the completion time combination and the device utilization value. The adjustment is achieved by swapping the positions of adjacent tasks, and the adjusted sequence is determined. Based on the adjusted sequence, its switching cost combination is obtained, wherein the switching cost combination is obtained by accumulating the task switching overhead between sequences; If the switching cost combination exceeds a preset threshold, the task positions are further swapped to reduce the switching cost combination, resulting in an optimized sequence set. From the optimized sequence set, the task allocation candidates are fused using a dynamic resource allocation method, wherein the dynamic resource allocation is achieved by allocating device resources according to the completion time combination to obtain an optimized task allocation candidate set.

4. The production line unit task allocation method based on reinforcement learning according to claim 1, characterized in that, The determination of the allocation parameters that need to be updated specifically includes: By obtaining real-time environmental change data from the optimized task allocation candidate set, and comparing the similarity between sudden order insertions and equipment maintenance notifications and the candidate set, the matching degree is determined. Based on the matching degree comparison results, the deviation between the allocated parameters and the changed data is compared to determine the allocated parameters that need to be updated, and a parameter set reflecting the highly dynamic environment is obtained. Also includes: The adaptive indicators in the parameter set are obtained, and the expanded task redistribution scheme is determined by fusing them with the real-time environmental change data.

5. The production line unit task allocation method based on reinforcement learning according to claim 1, characterized in that, The process of obtaining the final task allocation matrix specifically includes: The allocation parameters are updated, and a particle swarm optimization algorithm is used, where particles represent task allocation variables. The swarm searches for the optimal position through velocity updates to generate the task allocation scheme and obtain the initial values ​​for equipment utilization calculation and switching cost assessment. The equipment utilization calculation is based on the ratio of task load to equipment capacity, and the switching cost assessment is based on the product of the number of task migrations and migration costs. Based on the initial value, the trade-off value between the equipment utilization calculation and the switching cost assessment is iteratively calculated from the task allocation scheme, wherein the trade-off value is determined by summing the equipment utilization calculation multiplied by a first weight and the switching cost assessment multiplied by a second weight; If the tradeoff value is lower than a preset threshold, the particle swarm optimization algorithm is iterated again for the tradeoff value, wherein the particle positions are updated to minimize the tradeoff value, and the adjusted task allocation scheme is obtained. Based on the adjusted task allocation scheme, it is determined that the trade-off value meets the condition, wherein meeting the condition means that the trade-off value is not lower than a preset threshold, and the final matrix is ​​obtained to generate the corresponding task allocation matrix.

6. The production line unit task allocation method based on reinforcement learning according to claim 1, characterized in that, The specific instructions for confirming successful verification include: Extract production line unit instructions from the task allocation matrix, and simulate the running trajectory of the production line unit instructions under time constraint checks through the execution module to obtain the feasibility determination result; For the portion of the feasibility determination results that has been verified, multi-objective conflict resolution is incorporated to determine the instruction set after the conflict resolution is incorporated. Obtain simulation data from the execution module, perform simulation based on the instruction set after conflict resolution and integration, and determine the verification set. If the verified set satisfies the time constraint check, then the final instruction set is extracted from the verified set to obtain the instruction set extraction result.

7. The production line unit task allocation method based on reinforcement learning according to claim 1, characterized in that, The revised allocation strategy includes: By obtaining execution time records and energy consumption data monitoring through production line feedback, the difference between the actual value and the planned value is calculated from the execution time records and energy consumption data monitoring as the deviation value analysis to obtain the deviation amount; For the deviation amount, if it is determined that the deviation amount is greater than a preset threshold, an adjustment mechanism is triggered to determine the updated parameters. Based on the updated parameters, an adaptive parameter update correction allocation strategy is adopted to obtain the corrected allocation strategy.

8. The production line unit task allocation method based on reinforcement learning according to claim 1, characterized in that, The execution monitoring loop after determining the push includes: The optimized resource allocation indicators are obtained from the source of the strategy correction, and the indicators are pushed to the production line equipment through the real-time distribution system to determine the execution monitoring loop after the push. During the execution of the monitoring cycle, dynamic adjustment data of the production line is acquired, and if the allocation efficiency assessment is lower than a preset threshold, the cycle risk tracking is performed to address the risk of low efficiency. Through the cyclical risk tracking, risk reduction is confirmed, and a backup allocation plan is generated for the production line through dynamic adjustments, determining the optimal extraction of resources for the backup plan.

9. The production line unit task allocation method based on reinforcement learning according to claim 1, characterized in that, Also includes: Based on the verified instruction set, obtain production line feedback data such as actual execution time and energy consumption records, analyze the deviation from the feedback data, and determine if the deviation is greater than the preset threshold. If so, trigger the adaptive adjustment mechanism to obtain the corrected allocation strategy. From the revised allocation strategy, the optimized resource allocation indicators are extracted and pushed to the production line equipment through the real-time distribution system to determine the execution monitoring loop after the push.

10. A reinforcement learning-based production line unit task allocation system, used to implement the reinforcement learning-based production line unit task allocation method according to any one of claims 1-9, characterized in that, include: The data acquisition module is used to extract current production line order demand data and equipment status information from sensors and databases, and generate a preliminary task sequence list to cope with order priority fluctuations and equipment failure signals in a highly dynamic environment. The optimization filtering module is used to calculate multi-objective function values, filter out combinations that include completion time, equipment utilization and switching costs, adjust the task sequence to balance these objectives, and generate an optimized task allocation candidate set. The parameter update module is used to acquire real-time environmental change data, compare the matching degree between the change data and the candidate set, and determine the allocation parameters that need to be updated to adapt to the highly dynamic environment. The particle swarm optimization module is used to generate new task allocation schemes through the particle swarm optimization algorithm, iteratively calculate the trade-off between equipment utilization and switching costs, and generate the final task allocation matrix. The instruction verification module is used to extract specific task instructions for production line units from the final task allocation matrix, verify the feasibility of the instructions through the simulation execution module, and determine the instruction set that passes the verification. The strategy correction module is used to analyze deviations based on production line feedback data, trigger an adaptive adjustment mechanism, correct the allocation strategy, and push the optimized resource allocation indicators to the production line equipment to form an execution monitoring loop.

Citation Information

Cited By

  • Cross-warehouse scheduling migration method and device, equipment and medium

    CN122048248A

  • Migration method and device for cross-warehouse scheduling, equipment and medium

    CN122048248B