Distributed processing method and system for automobile manufacturing industry data based on digital twin
Through real-time collection and dynamic sorting, the problems of rigid priority of simulation tasks and unbalanced resource allocation in the digital twin system are solved, rapid response to sudden production events and intelligent allocation of resources are achieved, and the production efficiency and adaptability of the automobile manufacturing system are improved.
Patent Information
- Application Number
- CN202511113471.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2045-08-11
AI Technical Summary
When faced with emergencies, the existing digital twin automobile manufacturing system is unable to dynamically adjust the execution priority and resource allocation of simulation tasks, resulting in delayed execution of key decision support tasks, serious resource contention and load imbalance problems, and difficulty adapting to rapid changes in business needs and dynamic adjustments to resource constraints.
Through business status monitoring, production abnormal events, decision urgency parameters and simulation task types are collected in real time, a business priority evaluation data set is constructed, simulation tasks are dynamically sorted and load balancing and preemptive scheduling are performed, and performance feedback and optimization are performed in combination with an adaptive scheduling mechanism.
It achieves accurate perception and intelligent analysis of multi-dimensional business information in the automotive manufacturing environment, ensures timely response to key tasks, avoids resource contention, and improves computing resource utilization efficiency and system responsiveness.
Smart Images

Figure CN120596237B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a method and system for distributed processing of automotive manufacturing industry data based on digital twins. Background Art
[0002] With the rapid development of digital manufacturing technology, digital twin technology has been widely used in the automotive manufacturing field. By constructing a virtual mirror of the physical factory, it enables real-time monitoring, simulation analysis, and optimization decision-making of the production process. Existing digital twin automotive manufacturing systems typically adopt a distributed computing architecture, assigning complex simulation computing tasks to multiple computing nodes for parallel execution. These tasks include workshop production simulation, logistics distribution simulation, and equipment maintenance simulation. Traditional task scheduling methods are mainly based on static priority allocation and polling scheduling algorithms, which manage the execution order and resource usage of simulation tasks through preset task priority rules and fixed resource allocation strategies. These methods can meet basic scheduling requirements in scenarios with relatively stable system loads and relatively simple task types. They also have the advantages of simple implementation and low computational overhead.
[0003] However, existing technologies have obvious shortcomings: when emergencies such as equipment failures, production anomalies, or urgent order changes occur during the automobile manufacturing process, traditional static scheduling methods are unable to dynamically adjust the execution priority of simulation tasks based on the business urgency and task importance, resulting in delayed execution of key decision support tasks and affecting production efficiency. Existing resource allocation strategies usually adopt an average allocation or fixed allocation mode, lacking in-depth analysis of the resource demand characteristics of different simulation tasks. When multiple resource-intensive tasks are executed simultaneously, resource contention and load imbalance problems are prone to occur, causing some computing nodes to be overloaded while other nodes have idle resources. In addition, traditional scheduling methods lack an effective performance feedback mechanism and are unable to dynamically optimize scheduling strategies based on task execution and resource utilization efficiency. They are difficult to adapt to the actual requirements of rapid changes in business needs and dynamic adjustment of resource constraints in the automobile manufacturing environment. Summary of the Invention
[0004] This application provides a distributed processing method and system for automobile manufacturing industry data based on digital twins, which is used to solve the problem of lack of intelligent scheduling mechanism when multiple simulation tasks are executed concurrently in the existing technology, and improve the response speed of key simulation tasks and the utilization efficiency of computing resources.
[0005] In a first aspect, the present application provides a method for distributed processing of automotive manufacturing industry data based on digital twins, the method comprising:
[0006] Through business status monitoring, production abnormal events, decision urgency parameters, and simulation task types in the digital twin system are collected and processed in real time to obtain a business priority evaluation data set;
[0007] Prioritizing multiple simulation tasks according to the business priority evaluation data set to obtain a dynamic priority queue and a resource demand matrix;
[0008] The dynamic priority queue is input into a distributed simulation engine, and load balancing is performed on the CPU occupancy, memory usage, and GPU utilization of the computing nodes to obtain a node resource allocation plan;
[0009] Perform preemptive scheduling on high-priority simulation tasks according to the node resource allocation scheme, migrate low-priority tasks to idle nodes, and obtain a task execution mapping table;
[0010] The simulation progress and computing resource consumption in the task execution mapping table are dynamically monitored and processed through an adaptive scheduling mechanism to obtain scheduling performance feedback data and update the dynamic priority queue.
[0011] In a second aspect, the present application provides a digital twin-based automobile manufacturing industry data distributed processing system, the digital twin-based automobile manufacturing industry data distributed processing system comprising:
[0012] The acquisition module is used to collect and process production abnormal events, decision urgency parameters, and simulation task types in the digital twin system in real time through business status monitoring to obtain a business priority evaluation data set;
[0013] A sorting module is used to perform priority sorting on multiple simulation tasks according to the business priority evaluation data set to obtain a dynamic priority queue and a resource demand matrix;
[0014] An input module is used to input the dynamic priority queue into a distributed simulation engine, perform load balancing on the CPU occupancy, memory usage, and GPU utilization of the computing nodes, and obtain a node resource allocation plan;
[0015] A scheduling module is used to perform preemptive scheduling on high-priority simulation tasks according to the node resource allocation scheme, migrate low-priority tasks to idle nodes, and obtain a task execution mapping table;
[0016] The monitoring module is used to dynamically monitor the simulation progress and computing resource consumption in the task execution mapping table through an adaptive scheduling mechanism, obtain scheduling performance feedback data and update the dynamic priority queue.
[0017] In a third aspect, a distributed processing device for automobile manufacturing industry data based on digital twins is provided, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the distributed processing device for automobile manufacturing industry data based on digital twins executes the above-mentioned distributed processing method for automobile manufacturing industry data based on digital twins.
[0018] In a fourth aspect, a computer-readable storage medium is provided, in which instructions are stored. When the computer-readable storage medium is run on a computer, the computer executes the above-mentioned distributed processing method of automobile manufacturing industry data based on digital twins.
[0019] In the technical solution provided by this application, production abnormal events, decision urgency parameters and simulation task types in the digital twin system are collected and processed in real time through business status monitoring, realizing accurate perception and intelligent analysis of multi-dimensional business information in the automobile manufacturing environment. Compared with the traditional static task allocation method, it can dynamically adjust the importance assessment of simulation tasks according to the actual production situation to ensure that key business needs are responded to first. The technical feature of prioritizing multiple simulation tasks according to the business priority evaluation data set solves the problems of task priority fixation and inaccurate resource demand analysis in the existing technology by constructing a dynamic priority queue and a resource demand matrix, so that different types of tasks such as workshop production simulation, logistics distribution simulation and equipment maintenance simulation can be reasonably sorted according to business value and resource characteristics, avoiding the situation where important tasks are delayed due to resource contention. The technical feature of inputting the dynamic priority queue into the distributed simulation engine and performing load balancing on the computing nodes realizes the intelligent allocation and dynamic adjustment of computing resources by real-time monitoring of CPU occupancy, memory usage and GPU utilization, effectively solving the system performance bottleneck problem caused by unbalanced resource allocation in traditional methods.
[0020] The technical feature of preemptively scheduling high-priority simulation tasks according to the node resource allocation scheme, by migrating low-priority tasks to idle nodes and generating a task execution mapping table, realizes the refined management and dynamic optimization of computing resources, ensures that urgent simulation tasks can obtain the necessary computing resources and complete execution in a timely manner, and significantly improves the digital twin system's response to sudden production events. The key technical feature of dynamically monitoring and processing the simulation progress and computing resource consumption in the task execution mapping table through the adaptive scheduling mechanism establishes a complete performance feedback and optimization closed loop, which can dynamically adjust the scheduling strategy and update the dynamic priority queue according to the actual execution effect, solving the technical defects of the traditional scheduling method that lacks adaptive capabilities. In the application field of digital twins in automobile manufacturing, the priority sorting algorithm and preemptive scheduling algorithm of the present invention fully consider the business characteristics and resource constraints of the manufacturing environment. Through algorithmic features such as business importance score calculation and resource matching evaluation, the scheduling decision is closer to the actual production needs. Compared with the general task scheduling method, it has higher adaptability and execution efficiency when processing complex simulation tasks in automobile manufacturing scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0022] Figure 1 This is a schematic diagram of an embodiment of a method for distributed processing of automobile manufacturing industry data based on digital twins in an embodiment of the present application;
[0023] Figure 2 This is a schematic diagram of an embodiment of a distributed processing system for automobile manufacturing industry data based on digital twins in an embodiment of the present application;
[0024] Figure 3 It is a schematic block diagram of the structure of the distributed processing equipment for automobile manufacturing industry data based on digital twins in an embodiment of the present invention. DETAILED DESCRIPTION
[0025] The embodiments of the present application provide a method and system for distributed processing of automotive manufacturing industry data based on digital twins. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or that are inherent to these processes, methods, products or devices.
[0026] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 In the embodiments of the present application, an embodiment of a distributed processing method for automobile manufacturing industry data based on digital twins includes:
[0027] Step S101: Through business status monitoring, production abnormal events, decision urgency parameters, and simulation task types in the digital twin system are collected and processed in real time to obtain a business priority evaluation data set;
[0028] Step S102: Prioritize multiple simulation tasks according to the business priority evaluation data set to obtain a dynamic priority queue and resource demand matrix;
[0029] Step S103: Input the dynamic priority queue into the distributed simulation engine, perform load balancing on the CPU occupancy, memory usage, and GPU utilization of the computing nodes, and obtain a node resource allocation plan;
[0030] Step S104: preemptively schedule high-priority simulation tasks according to the node resource allocation plan, migrate low-priority tasks to idle nodes, and obtain a task execution mapping table;
[0031] Step S105 : Dynamically monitor the simulation progress and computing resource consumption in the task execution mapping table through the adaptive scheduling mechanism to obtain scheduling performance feedback data and update the dynamic priority queue.
[0032] It is understandable that the execution subject of this application can be a distributed processing system for automobile manufacturing industry data based on digital twins, or a terminal or server, which is not limited here. The embodiment of this application is explained by taking the server as the execution subject as an example.
[0033] Specifically, business status monitoring collects production anomalies, decision urgency parameters, and simulation task types from the digital twin system. Production anomalies include equipment failure signals, production line downtime events, and quality anomaly alarms. The system identifies these anomaly signals and generates production anomaly event identification codes by analyzing the fault code and alarm level. It also determines anomaly severity parameters based on the anomaly's impact scope and severity. Decision urgency parameters are obtained by analyzing the current production plan execution status, order delivery time, and inventory buffer level. The system quantifies these parameters to generate numerical indicators reflecting decision timeliness. Simulation task types include workshop production simulation tasks, logistics and distribution simulation tasks, and equipment maintenance simulation tasks. The system categorizes and labels different task types based on anomaly severity parameters, generating simulation task type labels and computational complexity weights. The decision urgency parameters, time window constraints, and simulation task type labels are then weighted and integrated. A weighted algorithm is used to calculate the business importance score of each simulation task, which is then ranked by production impact to form a business priority assessment dataset. When performing priority sorting based on the business priority assessment dataset, the system numerically compares the business importance scores of each simulation task to generate a simulation task priority sequence and grouping tasks of the same priority. The resource requirement matrix construction process includes resource statistics processing of the CPU requirements, memory usage, and GPU computing power of workshop production simulation tasks, logistics distribution simulation tasks, and equipment maintenance simulation tasks. The resource requirements of each task are derived through historical execution data and task complexity analysis to form a single-task resource consumption vector. The system performs concurrent analysis on simulation tasks within the same priority task group, calculates execution time estimates and resource conflict probabilities, and identifies dependencies between tasks and concurrent execution constraints. The single-task resource consumption vectors are arranged in a matrix according to the simulation task priority sequence to generate a resource requirement matrix and resource pre-allocation strategy. A dynamic priority queue is also constructed for subsequent scheduling.
[0034] After the distributed simulation engine receives the dynamic priority queue, the task scheduler monitors the current CPU occupancy, memory usage, and GPU utilization of each computing node in real time, and collects node resource state vectors and available resource capacity data through performance counters and resource monitoring agents. During the load balancing process, the system evaluates the load distribution of the master node, computing nodes, and storage nodes in the distributed simulation engine, calculates the load imbalance coefficient, and identifies resource bottlenecks. When a high-load computing node is detected, the system selects migration candidates for its simulation tasks, matches and analyzes resource-intensive tasks with lightweight tasks, and generates a list of task migration candidates and target node matching pairs. Constraint checking ensures that the migration task does not exceed the node capacity limit by verifying the remaining number of CPU cores, available memory size, and GPU memory space of each target node, forming a feasibility allocation matrix and generating a node resource allocation plan based on the principle of maximizing resource utilization.
[0035] Preemptive scheduling determines the priority status of currently running simulation tasks based on the node resource allocation plan. The system scans the task priorities in the run queues of each compute node, extracts the priority values of workshop production simulation tasks, logistics distribution simulation tasks, and equipment maintenance simulation tasks, and forms a current task priority list. By comparing the priorities of newly arrived high-priority simulation tasks with those of currently running tasks, it sets a preemption threshold and identifies low-priority tasks that need to be preempted. This generates a priority difference matrix and a set of preemption candidate tasks. The system calculates the degree of match between the resource utilization of each compute node and the resource requirements of the new task, determines whether the resource release conditions for preemptive scheduling are met, and generates preemptive execution instructions for tasks that meet the conditions. Preempted low-priority simulation tasks are required to save their state, including execution context and intermediate computation results, to generate task checkpoint data and task recovery information. An idle node searcher is then used to locate idle compute nodes that meet the resource requirements, complete task redeployment and state recovery, and establish a task execution mapping table to record the execution location and resource utilization of each simulation task on the distributed compute nodes.
[0036] The adaptive scheduling mechanism monitors simulation progress and computing resource consumption in the task execution mapping table, collecting the execution progress percentage and remaining execution time of each simulation task in real time. It also compiles statistics on the completion status of workshop production simulation tasks, logistics distribution simulation tasks, and equipment maintenance simulation tasks, generating simulation progress monitoring data and task completion time predictions. The system calculates resource consumption statistics for each compute node, including CPU usage fluctuations, memory usage fluctuations, and GPU computing load. It also calculates the resource consumption rate and peak load per unit time, generating computing resource consumption data and node performance metrics. Performance evaluation analyzes the execution efficiency and resource utilization efficiency of the current scheduling strategy, compiling statistics on task waiting time, response latency, and throughput, and identifying scheduling efficiency evaluation indicators and performance bottlenecks. By correlating computing resource consumption data with scheduling efficiency evaluation indicators, it identifies scheduling performance optimization directions and resource allocation adjustment needs. Scheduling performance feedback data is generated, and corrective actions are taken based on task priority deviations and execution efficiency anomalies. The system adjusts the business importance scores and priority sorting rules of simulation tasks, and updates the dynamic priority queue to complete closed-loop optimization.
[0037] In a specific embodiment, the process of executing step S101 may specifically include the following steps:
[0038] Perform status recognition and processing on equipment failure signals, production line downtime events, and quality abnormality alarms in automobile manufacturing workshops to obtain production abnormality event identification codes and abnormality severity level parameters;
[0039] According to the production abnormal event identification code, the current production plan execution status, order delivery time, and inventory buffer level are quantified to obtain the decision urgency parameters and time window constraints.
[0040] Based on the abnormal severity level parameters, the workshop production simulation tasks, logistics distribution simulation tasks and equipment maintenance simulation tasks to be executed are classified and labeled to obtain the simulation task type label and computational complexity weight;
[0041] The decision urgency parameter, time window constraint, and simulation task type label are weighted and integrated to obtain the business importance score of each simulation task.
[0042] The business importance scores are sorted according to the degree of production impact to obtain a business priority assessment data set.
[0043] Specifically, the status identification and processing of equipment failure signals, production line shutdown events and quality abnormality alarms in automobile manufacturing workshops are carried out by parsing the original signal data from the ANDON system, MES system and IOT sensors. The equipment failure signal contains the faulty equipment number, fault type code and fault occurrence timestamp, the production line shutdown event contains the shutdown line body identification, shutdown reason classification and shutdown duration, and the quality abnormality alarm contains the batch number of unqualified products, defect type and detection location information. The state recognition processing formats and standardizes these raw data through the signal analysis module, maps the fault type code in the equipment fault signal to a standardized fault classification, converts the shutdown cause classification of the production line shutdown event into a unified event type code, and converts the defect type of the quality abnormality alarm into a quality problem level code. Finally, a production abnormality event identification code containing the event type, occurrence location, time information and impact range is generated. At the same time, the abnormality severity level parameter is calculated based on the number of equipment affected by the fault, the downtime duration and the severity of the quality problem. This parameter is obtained by weighted calculation of the fault impact factor, time impact factor and range impact factor. The fault impact factor reflects the degree of loss of production capacity due to equipment failure, the time impact factor reflects the degree of delay in production progress due to the downtime duration, and the range impact factor reflects the scope of the abnormal event on related workstations and production lines.
[0044] When performing urgency quantification based on production exception event identification codes, the ERP system retrieves the current production plan execution status, including the planned completion progress, the deviation between actual and target production output, and the execution status of key processes. Order delivery time information, including order deadlines, remaining delivery time, and customer priority, is then extracted from the order management system. Inventory buffer level data, including raw material inventory, semi-finished product buffer inventory, and finished product inventory levels, is also obtained from the WMS system. This urgency quantification process analyzes the exception type and impact scope in the production exception event identification code, combined with the progress deviation in the current production plan execution status, to calculate the threat level of the exception event to production plan completion. The urgency value increases when the exception event affects key processes and production schedule is already delayed. It further increases when the order delivery date is approaching and the customer priority is high. The urgency value peaks when the inventory buffer level is insufficient to cope with the production disruption. The decision urgency parameter is derived by comprehensively calculating the schedule threat factor, delivery risk factor, and inventory risk factor. The time window constraint is calculated by subtracting the current time from the order deadline and the estimated recovery time. When the time window constraint is less than the safety threshold, the decision must be completed within the specified time.
[0045] When classifying and marking simulation tasks based on the anomaly severity parameters, workshop production simulation tasks primarily simulate the production line's operating status, workstation workflow, and product assembly process. Logistics and distribution simulation tasks primarily simulate material distribution routes, AGV scheduling strategies, and buffer management. Equipment maintenance simulation tasks primarily simulate equipment maintenance plans, spare parts consumption forecasts, and maintenance resource allocation. Classification and marking processes classify tasks based on the numerical range of the anomaly severity parameters. When the anomaly severity parameter exceeds the production line shutdown threshold, the workshop production simulation task is marked as a high-priority task. When the anomaly affects the material distribution route or AGV operation, the logistics and distribution simulation task is marked as a medium-priority task. When the anomaly involves equipment failure but does not affect overall line production, the equipment maintenance simulation task is marked as a low-priority task. The simulation task type label records the business type, priority level and execution constraints of the task in an encoded manner. The computational complexity weight is quantified based on the scale of the simulation model, the number of computing nodes and the simulation duration. The workshop production simulation task has the highest complexity weight because it involves complex production logic and a large number of equipment status calculations. The logistics distribution simulation task has a medium complexity weight because it involves path optimization and scheduling algorithms. The equipment maintenance simulation task has a lower complexity weight because it mainly performs status prediction and resource calculations.
[0046] The weighted fusion process numerically integrates the decision urgency parameter, time window constraint, and simulation task type label. The decision urgency parameter reflects the timeliness requirement for task execution, the time window constraint reflects the time limit for task completion, and the simulation task type label reflects the business importance of the task. The weighted fusion algorithm first normalizes the decision urgency parameter, converting urgency values of different dimensions into standardized values between zero and one. It then performs an inverse transformation on the time window constraint, assigning a larger weight to the tighter time window. The priority level in the simulation task type label is directly mapped to the weight coefficient. The business importance score is calculated through a weighted summation. The weight coefficient for the decision urgency parameter, the weight coefficient for the time window constraint, and the weight coefficient for the simulation task type label are set to 0.2. This weighted summation is then normalized to form the final business importance score, which reflects the overall importance and execution priority of the simulation task in the current business scenario.
[0047] The sorting process arranges tasks in descending order based on their business importance scores, with production impact serving as the primary basis for sorting. When multiple simulation tasks have the same business importance score, they are secondary sorted by comparing their production impact. The production impact is calculated by analyzing the scope and intensity of the simulation task's impact on overall production efficiency, product quality, and delivery schedule. The business priority assessment dataset contains the sorted simulation task list, its corresponding business importance scores, its production impact assessment, and execution constraints. This dataset serves as the basic input for subsequent priority queue construction and resource allocation.
[0048] In a specific embodiment, the process of executing step S102 may specifically include the following steps:
[0049] Perform numerical comparison processing on the business importance scores of each simulation task in the business priority evaluation data set to obtain the simulation task priority sequence and the grouping of tasks with the same priority;
[0050] According to the simulation task priority sequence, the CPU demand, memory usage, and GPU computing power of the workshop production simulation task, logistics distribution simulation task, and equipment maintenance simulation task are statistically processed to obtain the single-task resource consumption vector.
[0051] Based on the grouping of tasks with the same priority, the execution time estimation and resource conflict probability of the simulation tasks in each group are analyzed concurrently to obtain the dependency graph between tasks and the concurrent execution constraints.
[0052] The single-task resource consumption vector is arranged in a matrix according to the simulation task priority sequence to obtain the resource demand matrix and resource pre-allocation strategy;
[0053] The high-priority simulation tasks in the task dependency graph are queued according to the execution order to obtain a dynamic priority queue.
[0054] Specifically, when the business importance scores of each simulation task in the business priority assessment dataset are numerically compared, the comparison algorithm traverses the business importance scores of all simulation tasks in the dataset and uses a quick sort algorithm to sort the scores in descending order. The business importance score serves as the primary key for sorting, and the task identification code serves as the secondary key to ensure the uniqueness and stability of the sorting result. The numerical comparison process first extracts the business importance score corresponding to each simulation task, and then determines the relative priority relationship of the tasks through pairwise comparison. When the business importance scores of two tasks are equal, they are sorted by comparing the inherent priorities of the task types. The workshop production simulation task has the highest inherent priority, followed by the equipment maintenance simulation task, and the logistics distribution simulation task has the lowest priority. The simulation task priority sequence is formed by the sorted task list. Each position in the sequence corresponds to a simulation task and its priority number. The priority number increases from one, and the smaller the value, the higher the priority. Tasks of the same priority level are grouped by identifying simulation tasks with the same business importance score. The grouping algorithm traverses adjacent tasks in the priority sequence and compares whether their business importance scores are equal. Equal tasks are classified into the same group and assigned the same group identifier. The grouping result records the number of tasks contained in each group, the distribution of task types, and the average resource requirements.
[0055] Resource statistics processing calculates the resource requirements for each simulation task based on the task type and computational complexity weight in the simulation task priority sequence. The CPU requirement is determined by analyzing the computational intensity and parallel processing requirements of the simulation model. The workshop production simulation task has the highest CPU requirement because it needs to simulate complex production processes and equipment status changes. The logistics and distribution simulation task has a medium CPU requirement because it involves path optimization and scheduling algorithms. The equipment maintenance simulation task has a lower CPU requirement because it mainly performs status prediction and simple calculations. Memory usage is calculated by evaluating the data size and cache requirements of the simulation model. The workshop production simulation task requires loading a large amount of equipment status data, process parameters, and product models, so it has the largest memory usage. The logistics and distribution simulation task requires storing material information, path data, and scheduling status, and has a medium memory usage. The equipment maintenance simulation task mainly stores equipment historical data and maintenance records, and has a relatively small memory usage. GPU computing power is primarily used for 3D model rendering and graphical display. Workshop production simulation tasks require rendering complex 3D production line scenes and equipment animations, necessitating the highest GPU computing power. Logistics and distribution simulation tasks require rendering AGV paths and material flow animations, requiring moderate GPU computing power. Equipment maintenance simulation tasks primarily display equipment status charts and maintenance progress, requiring relatively low GPU computing power. The single-task resource consumption vector is formed by combining the CPU demand, memory usage, and GPU computing power of each simulation task into a three-dimensional vector. The first element of the vector represents the normalized value of the CPU demand, the second element represents the normalized value of the memory usage, and the third element represents the normalized value of the GPU computing power.
[0056] Concurrency analysis and processing estimate execution time and calculate resource conflict probabilities based on the characteristics of simulation tasks within the same priority group. Execution time estimates are determined by analyzing the computational complexity, model size, and historical execution data of the simulation tasks. The estimation algorithm considers both the serial and parallel computational components of the task. The execution time of the serial computation component is directly related to CPU performance, while the execution time of the parallel computation component depends on the number of available CPU cores and parallel efficiency. Resource conflict probabilities are calculated by analyzing the degree of overlap in resource requirements among different simulation tasks within the same group. The probability of resource conflict increases when two tasks simultaneously require the same type of computing resources and the total amount of resources is limited. The conflict probability calculation considers the peak time, duration, and total capacity of the resource pool. The inter-task dependency graph is constructed by analyzing the data and logical dependencies between simulation tasks. Data dependencies refer to the output of one task being the input of another, while logical dependencies refer to constraints on the order in which tasks are executed. The dependency graph is represented as a directed graph, with nodes representing simulation tasks, edges representing dependency relationships, and edge weights indicating the strength of the dependency. Concurrent execution constraints are determined based on the probability of resource conflict and the dependencies between tasks. When the probability of resource conflict exceeds the threshold, related tasks cannot be executed concurrently. When there is a strong dependency, the dependent task must be executed before the dependent task.
[0057] Matrix permutation processing arranges the single-task resource consumption vectors in the order of the simulation task priority sequence to form a resource demand matrix. The rows of the matrix represent different simulation tasks, the columns represent different types of computing resources, and the values of the matrix elements represent the demand for specific resources by the corresponding tasks. The resource demand matrix analyzes the distribution characteristics and peak load of the overall resource demand through matrix operations. The sum of the rows gives the total resource demand of each task, and the sum of the columns gives the total demand of each resource. The maximum eigenvalue of the matrix reflects the concentration of resource demand. The resource pre-allocation strategy is formulated based on the resource demand matrix and the current available resource status. The strategy includes priority rules for resource allocation, resource reservation mechanism, and dynamic adjustment strategy. High-priority tasks are given priority in resource allocation, and a certain proportion of capacity is reserved for critical resources to cope with sudden demands. The dynamic adjustment strategy adjusts the allocation plan according to real-time resource usage.
[0058] The queue construction process constructs a dynamic priority queue based on the topological sorting results and priority information in the inter-task dependency graph. Topological sorting ensures that dependent tasks are executed in the correct order, and priority information ensures that important tasks are given priority execution opportunities. The queue construction algorithm first topologically sorts the dependency graph, identifying tasks without predecessor nodes as candidate tasks for execution. Then, the highest-priority task is selected from the candidate tasks and added to the queue. Completed tasks are removed from the dependency graph, and the in-degree of their successor tasks is reduced by one. When the in-degree of a successor task reaches zero, it becomes a new candidate task. The dynamic priority queue uses a priority queue data structure that supports dynamic insertion and deletion of tasks. The head of the queue is always the task with the highest priority that meets the execution conditions. The queue length reflects the number of tasks to be executed, and the queue update frequency is dynamically adjusted based on the progress of task execution and the arrival of new tasks.
[0059] In a specific embodiment, the process of executing step S103 may specifically include the following steps:
[0060] The dynamic priority queue is input into the task scheduler of the distributed simulation engine to monitor the current CPU occupancy, memory usage, and GPU utilization of each computing node in real time to obtain the node resource state vector and available resource capacity data;
[0061] Based on the node resource state vector, the load distribution of the master node, computing node and storage node in the distributed simulation engine is balanced and evaluated to obtain the load imbalance coefficient and resource bottleneck identification mark;
[0062] Based on the load imbalance coefficient, the simulation tasks of the high-load computing nodes are selected for migration candidates, and resource-intensive tasks are matched and analyzed with lightweight tasks to obtain a list of task migration candidates and target node matching pairs.
[0063] The task migration candidate list is checked against the available resource capacity data, and the remaining number of CPU cores, available memory size, and GPU memory space of each target node are verified to obtain a feasibility allocation matrix.
[0064] The feasibility allocation matrix is optimized and allocated according to the principle of maximizing resource utilization to obtain the node resource allocation plan and the load balancing configuration parameters of the distributed simulation engine.
[0065] Specifically, after the dynamic priority queue is input into the distributed simulation engine's task scheduler, the task scheduler sends a resource monitoring request to each compute node via a network communication protocol. Upon receiving the request, the resource monitoring agent of each compute node immediately obtains real-time data on the current CPU occupancy, memory usage, and GPU utilization. CPU occupancy is obtained by reading the operating system's processor performance counters, which reflects the ratio of the number of CPU cores currently executing tasks to the total number of cores. Memory usage is obtained by querying the system memory manager to obtain the allocated memory size and total memory capacity. GPU utilization is obtained by accessing the graphics card driver interface to obtain the GPU core's workload status and video memory usage. Real-time monitoring utilizes a timed polling mechanism, with each compute node actively reporting resource status data at a set interval. The task scheduler then formats and normalizes this data, unifying resource data from different nodes into the same data format and value range. The node resource status vector is a three-dimensional vector formed by combining the CPU occupancy, memory usage, and GPU utilization of each compute node. Each element in the vector corresponds to the usage status of a resource type, and the modulus of the vector reflects the overall load level of the node. Available resource capacity data is calculated by subtracting the current usage from the total resource capacity of each node. This includes the number of available CPU cores, remaining memory size, and free GPU memory capacity. This data represents the maximum amount of resources that the node can allocate to new tasks.
[0066] The balance assessment process analyzes the load distribution characteristics of different types of nodes in the distributed simulation engine based on the node resource state vector. The master node is responsible for task scheduling and system coordination, typically with high CPU utilization but relatively low memory and GPU utilization. Compute nodes handle specific simulation computations, with high CPU, memory, and GPU utilization. Storage nodes primarily handle data read and write operations, with high memory utilization but low CPU and GPU utilization. Load distribution is quantitatively assessed by calculating the standard deviation and coefficient of variation of each node's resource utilization. The standard deviation reflects the degree of load dispersion between nodes, while the coefficient of variation reflects the relative differences in load distribution. The load imbalance coefficient is calculated by calculating the difference in resource utilization between the highest-loaded and lowest-loaded nodes. When the imbalance coefficient exceeds a preset threshold, it indicates that the load distribution is uneven and needs to be adjusted. The calculation of the imbalance coefficient takes into account the combined load of the three resources: CPU, memory, and GPU, and a comprehensive evaluation result is obtained through weighted averaging. The resource bottleneck identification flag is determined by analyzing whether the resource utilization rate of each node is close to full load. When the utilization rate of any resource of a node exceeds the warning threshold, the bottleneck flag is set. The bottleneck flag contains information such as the bottleneck node identifier, bottleneck resource type, and bottleneck severity.
[0067] The migration candidate selection process identifies high-load compute nodes requiring task migration based on the load imbalance coefficient. The selection algorithm first selects nodes with loads exceeding the average as migration source nodes. It then analyzes the characteristics of the simulation tasks running on these nodes, classifying them into resource-intensive and lightweight tasks based on their resource consumption. Resource-intensive tasks are simulation tasks with high CPU, memory, or GPU requirements, typically including complex workshop production simulations and large-scale logistics and distribution simulations. Lightweight tasks are simulation tasks with relatively low resource requirements, primarily simple equipment maintenance simulations and condition monitoring tasks. Matching analysis is performed by calculating the resource complementarity between resource-intensive and lightweight tasks. When one resource-intensive task primarily consumes CPU resources and another lightweight task primarily consumes memory resources, the two tasks have good complementarity and can be executed in parallel on the same node. The task migration candidate list contains information such as the identity of the task to be migrated, its current node, resource demand characteristics, and migration priority. The migration priority is determined based on a combination of the task's business importance and migration cost. The target node matching is performed by analyzing the available resources of low-load nodes and matching them with the resource requirements of the tasks to be migrated. The matching algorithm takes into account factors such as resource capacity constraints, network latency, and data transmission costs.
[0068] Constraint checking verifies the resource constraints of each task to be migrated from the candidate task list against the target node. This verification process includes capacity checks on the target node's remaining CPU cores, available memory, and GPU memory. The remaining CPU cores are calculated by subtracting the number of currently active cores from the target node's total CPU cores. Capacity verification ensures that the CPU requirements of the task to be migrated do not exceed the remaining cores. Available memory is obtained by querying the target node's memory manager to obtain the current free memory capacity, verifying whether the task's memory requirements are within the available range. GPU memory capacity verification verifies the available memory capacity by accessing the target node's graphics card status information to ensure that the task's GPU memory requirements are met. Constraint checking also includes verification of network bandwidth, storage space, and software environment to ensure that the task can execute normally after migration. A feasibility allocation matrix records the feasibility of each task to be migrated and each target node. The rows of the matrix represent the task to be migrated, and the columns represent the target nodes. The values of the matrix elements represent the feasibility score of the match, which comprehensively considers factors such as resource satisfaction, migration cost, and execution efficiency.
[0069] The optimal allocation process applies the resource utilization maximization principle to the feasibility allocation matrix to optimize task allocation. The goal of the optimization algorithm is to maximize overall resource utilization and minimize load imbalance while satisfying resource constraints. Resource utilization maximization is achieved by calculating the average resource utilization of all nodes and finding the allocation solution that maximizes this metric. The optimization process uses a heuristic search algorithm to select the highest-scoring task-node matching pair from the feasibility allocation matrix. The matching pair is then checked to see if it results in resource conflicts or constraint violations. If the matching is feasible, the task is assigned to the corresponding node and the matrix is updated. Otherwise, the suboptimal matching pair is selected and further attempts are made. The node resource allocation solution records the finalized task allocation results, including the target node assigned to each simulation task, the amount of allocated resources, and the expected execution time. The load balancing configuration parameters of the distributed simulation engine include load monitoring frequency, migration trigger threshold, resource reservation ratio, and dynamic adjustment strategy. These parameters are set and adjusted based on the current load distribution and system performance requirements.
[0070] In a specific embodiment, the process of executing step S104 may specifically include the following steps:
[0071] According to the node resource allocation plan, the priority status of the currently running workshop production simulation task, logistics distribution simulation task and equipment maintenance simulation task is preempted and processed to obtain the high-priority task preemption flag and the low-priority task migration flag;
[0072] Based on the high-priority task preemption flag, the target computing node's running state is suspended, the execution context and intermediate calculation results of the low-priority simulation task being executed are saved, and the task checkpoint data and task recovery information are obtained;
[0073] Input the task checkpoint data into the idle node searcher, perform matching and searching on the idle computing nodes that meet the resource requirements in the distributed simulation engine, and obtain the available node list and node resource margin parameters;
[0074] Redeploy low-priority simulation tasks according to the available node list, transmit task recovery information to the target idle node, and restore the task execution status to obtain the task migration completion status and node load update data;
[0075] A mapping relationship is established between the task migration completion status and the node load update data, and the execution location and resource occupancy of each simulation task on the distributed computing node are recorded to obtain a task execution mapping table.
[0076] Specifically, the preemption determination process compares the priority status of currently running workshop production simulation tasks, logistics distribution simulation tasks, and equipment maintenance simulation tasks based on the task priority information in the node resource allocation plan. The determination algorithm first extracts information about newly arrived high-priority tasks from the node resource allocation plan, including task identification, priority value, and resource requirement parameters. It then queries the list of simulation tasks currently executing on each compute node to obtain the current priority status of each running task. The priority status includes information such as the task's business importance score, execution progress, and remaining time. Preemption determination is performed by comparing the priority value of the new task with the priority value of the currently running task. When the new task's priority value is significantly higher than that of a running task and both require the same type of computing resources, the preemption condition is triggered. Preemption determination also considers the benefits of resource release and the cost of preemption. If the preempted task is nearing completion or preemption would result in the loss of a large number of intermediate results, the preemption determination algorithm lowers the preemption priority. The preemption flag for high-priority tasks is generated by setting a Boolean flag and preemption target information. A true flag indicates that a preemption operation is required. The preemption target information includes the identification of the preempted task, its node, and the expected amount of resources released. The low-priority task migration flag is set for the preempted task. The flag information includes the task ID, current execution status and migration urgency. The migration urgency is calculated based on the remaining execution time of the task and the business importance.
[0077] Task suspension processing controls the target compute node's running state based on the preemption flag of the high-priority task. The suspension algorithm first sends a task suspension instruction to the target node. Upon receiving the instruction, the node immediately stops executing the preempted task and prohibits the submission of new computation instructions. The execution context saves key information, including the task's program counter, register state, memory allocation table, and file handles. This information records the complete execution state of the task at the moment of suspension, ensuring accurate resumption of task execution from the point of suspension. Intermediate computation results are saved, including the simulation model's current state data, completed computation steps, and temporary variable values. For workshop production simulation tasks, this requires saving workstation status, product progress, and equipment operating parameters. For logistics and distribution simulation tasks, this requires saving AGV location, cargo status, and routing information. For equipment maintenance simulation tasks, this requires saving equipment health, maintenance progress, and resource consumption records. State preservation utilizes an incremental snapshot mechanism, saving only the data that has changed since the initial state, reducing storage space usage and data transmission time. Task checkpoint data is serialized to convert the execution context and intermediate computation results into a storable and transferable data format. Checkpoint data contains metadata such as the data version number, creation timestamp, and data integrity checksum. Task recovery information includes the environment configuration, dependencies, and resource allocation requirements required to restart the task. The environment configuration specifies the software version and system parameters required for the task to run. The dependencies describe the relationship between the task and other tasks or data. The resource allocation requirements clearly define the minimum amount of resources required to resume the task.
[0078] After receiving task checkpoint data, the idle node searcher first parses the resource requirements contained in the data, including the number of CPU cores, memory capacity, and GPU memory requirements. It then traverses all compute nodes in the distributed simulation engine, querying each node's current resource usage and available resource capacity. The matching search process utilizes a multi-dimensional evaluation algorithm, encompassing resource matching, network connection quality, node load history, and task execution performance. Resource matching is determined by measuring the degree of fit between a node's available resources and the task's required resources. A node is considered a match when its available CPU core count is equal to or greater than the task's required number and its available memory capacity meets the minimum requirements. Network connection quality is assessed by measuring inter-node network latency and bandwidth capacity. Nodes with good network connectivity are selected to minimize data transmission time and task migration costs. Node load history is analyzed to assess node stability and reliability by analyzing resource usage patterns and task execution success rates over time, avoiding the selection of nodes with frequent failures or significant performance fluctuations. The available node list contains information on all candidate nodes that meet resource and performance requirements. The list is sorted by comprehensive evaluation score, with the highest-scoring node being selected as the target node for task migration. The node resource margin parameters record the remaining resource capacity of each available node after being allocated to migration tasks, including the number of remaining CPU cores, remaining memory size, and remaining GPU memory capacity. These parameters are used for resource allocation decisions and load balancing adjustments.
[0079] The redeployment process selects the optimal target node for task migration based on a list of available nodes. The deployment algorithm first creates the task execution environment on the target node, including allocating the required computing resources, establishing network connections, and installing necessary software dependencies. Task recovery information is transmitted over the network, sending task checkpoint data and recovery information to the target node. Compression and encryption techniques are used during transmission to ensure data integrity and security. Upon receipt, the target node decompresses and decrypts the data and then verifies its integrity and validity. Task execution status recovery involves rebuilding the task's memory space, restoring the program execution context, and reconnecting to external resources. Memory space reconstruction is achieved by allocating a memory area of the same size as the original node and loading the saved data. Program execution context restoration is achieved by setting the program counter, register state, and stack pointer. External resource reconnection involves reopening files, reestablishing network connections, and reinitializing hardware devices. The task migration completion status is determined by monitoring the task's launch and initial execution results on the target node. Migration is considered complete when the task successfully launches and produces the expected computational output. The migration completion status includes information such as the migration success flag, migration duration, and resource usage. The node load update data includes the amount of resources released by the original node and the amount of new resources occupied by the target node. These data are used to update the global resource status and load distribution information of the distributed simulation engine.
[0080] The mapping process correlates and analyzes task migration completion status and node load updates to establish a correspondence between each simulation task and compute node. This mapping records the compute node where each simulation task is currently located, the specific resources it occupies, and the expected completion time. Execution location information includes the node's physical identifier, network address, and resource configuration parameters. The physical identifier uniquely identifies the node's location within the distributed system, the network address is used to establish a communication connection with the node, and the resource configuration parameters describe the node's hardware specifications and performance characteristics. Resource usage records the CPU core number, memory address range, and GPU memory block identifier used by the task on the node. This information is used for resource conflict detection and resource release management, enabling accurate release of computing resources based on resource usage upon task completion. The task execution mapping table uses a relational data structure to store mappings. Each row in the table represents a simulation task, and columns include fields such as task identifier, node identifier, resource usage details, execution status, and update timestamp. The mapping table supports fast query and dynamic update operations, promptly updating corresponding mapping records when task status changes or node configuration adjustments are made.
[0081] In a specific embodiment, the process of executing the step of preempting the priority status of the currently running workshop production simulation task, logistics distribution simulation task, and equipment maintenance simulation task according to the node resource allocation plan may specifically include the following steps:
[0082] Perform task priority scanning on the running queues of each computing node in the node resource allocation plan, extract the priority values of the currently executed workshop production simulation tasks, logistics distribution simulation tasks, and equipment maintenance simulation tasks, and obtain the current task priority list and node task distribution status;
[0083] According to the current task priority list, the newly arrived high-priority simulation task is compared with the priority value of the running task, and the preemption threshold is set and the low-priority task that needs to be preempted is identified to obtain the priority difference matrix and the preemption candidate task set;
[0084] Based on the priority difference matrix, the resource occupancy of each computing node is matched with the resource requirements of the new task to determine whether the resource release conditions for preemptive scheduling are met, and the resource matching score and preemption feasibility flag are obtained.
[0085] The preemption candidate task set and resource matching score are comprehensively processed for decision making, and a preemption execution instruction is generated for the high-priority simulation task that meets the preemption conditions, and the preemption flag of the high-priority task and the preemption target node identifier are obtained;
[0086] The necessity of migration is determined for the preempted low-priority simulation tasks in the preemption candidate task set, tasks that need to be migrated to other nodes for continued execution are marked, and a low-priority task migration flag and a migration task identifier list are obtained.
[0087] Specifically, the task priority scanning process traverses and analyzes the run queues of each compute node in the node resource allocation scheme. The run queue is a task execution sequence maintained by each compute node, recording information about all currently executing and pending simulation tasks. The scanning algorithm first accesses the task manager of each compute node, queries the list of tasks in the current run queue, and then extracts basic information about each task, including task ID, task type, and business priority value. The priority value of a workshop production simulation task is calculated based on its impact on production efficiency and urgency parameters. The priority value of a logistics distribution simulation task is based on the timeliness requirements of material delivery and inventory risk assessment. The priority value of an equipment maintenance simulation task is determined based on the severity of the equipment failure and the urgency of the maintenance. Priority values are standardized on a scale of zero to one hundred, with higher values indicating higher task priority. The scanning process standardizes the priority values of tasks of the same type on different nodes to ensure consistency across node priority comparisons. The current task priority list is formed by aggregating the scan results from all nodes. The list contains fields such as task ID, node, task type, priority value, and current execution status. The list is sorted in descending order by priority value to facilitate subsequent comparative analysis. The node task distribution status records the quantity distribution and load of different types of simulation tasks on each computing node, including the number of workshop production simulation tasks, logistics distribution simulation tasks, and equipment maintenance simulation tasks carried by each node, as well as the resource ratio occupied by each type of task and execution progress information.
[0088] The comparison and judgment process performs a priority comparison analysis on the newly arrived high-priority simulation task based on the current task priority list. The comparison algorithm first obtains the priority value of the new task, and then compares the value with each running task in the current task priority list. The priority value comparison adopts a numerical difference calculation method. The priority difference is obtained by subtracting the running task priority value from the new task priority value. A positive value indicates that the new task has a higher priority, a negative value indicates that the running task has a higher priority, and a zero value indicates that the two tasks have equal priority. The preemption threshold is a preset critical value. When the priority difference exceeds the preemption threshold, the new task is considered to have sufficient priority advantage to be preempted. The setting of the preemption threshold takes into account the balance between task switching cost and preemption benefit. A too low threshold will lead to frequent task preemption and affect the overall execution efficiency. A too high threshold will reduce the response speed of high-priority tasks. Low-priority task identification traverses the current task priority list and filters out running tasks whose priority values are lower than the new task and whose priority difference exceeds the preemption threshold. These tasks are marked as preemption candidates. The priority difference matrix uses a two-dimensional structure to record the priority difference between new tasks and running tasks. The rows of the matrix represent newly arrived high-priority tasks, and the columns represent currently running tasks. The values of the matrix elements represent the priority difference between the corresponding tasks. A larger difference indicates a stronger need for preemption. The preemption candidate task set contains information about all running tasks that meet the preemption criteria. Each element in the set records key information such as the candidate task's identity, node location, current resource usage, and expected resource release amount.
[0089] The matching calculation process analyzes the compatibility of each compute node's resource usage with the new task's resource requirements based on a priority difference matrix. The calculation process first obtains the new task's resource requirements, including the number of CPU cores, memory capacity, and GPU memory size. It then queries the preemption candidate's current resource usage, including allocated CPU cores, occupied memory, and GPU memory usage. Resource release conditions are determined by comparing the candidate task's resource type and quantity with the new task's resource requirements. Basic resource release conditions are considered met when the number of CPU cores released by the candidate task is greater than or equal to the new task's CPU requirement and the released memory capacity meets the new task's memory requirements. Preemptive scheduling also considers resource continuity and compatibility requirements for resource release. The new task must require a consecutively numbered CPU core combination, a continuous address range for memory, and independent GPU memory blocks for GPU memory. If the candidate task's released resources do not meet these continuity requirements, the preemption condition cannot be met, even if sufficient resources are available. The resource matching score is quantitatively calculated by comprehensively evaluating resource quantity matching, resource quality matching, and resource release efficiency. Resource quantity matching reflects the degree of consistency between the quantity of released resources and the required resources. Resource quality matching considers the performance characteristics of the released resources and the performance requirements of the new task. Resource release efficiency evaluates the time cost of preemption and the speed of resource recovery. The preemption feasibility flag is marked by setting a Boolean value and a feasibility level. A Boolean value of true indicates that preemption is technically feasible. The feasibility level is divided into high, medium, and low levels, reflecting the recommendation and execution priority of the preemption operation.
[0090] The comprehensive decision-making process correlates and analyzes the set of preemption candidate tasks with their resource matching scores. The decision-making algorithm comprehensively considers multiple factors, including priority difference, resource matching score, preemption cost, and execution benefit. The priority difference reflects the necessity and urgency of preemption; a larger difference indicates a higher business value. The resource matching score reflects the technical feasibility of preemption; a higher score indicates a higher probability of successful preemption. Preemption cost includes the time overhead of task switching, network costs for data transmission, and system resource management overhead. Execution benefit includes the business value of early completion of high-priority tasks and the improvement in overall system efficiency. Comprehensive decision-making is performed using a weighted scoring model, assigning a weight coefficient to each evaluation factor. A weighted summation is used to calculate the comprehensive decision score, and the preemption solution with the highest score is selected for execution. High-priority simulation tasks that meet the preemption criteria are screened through comprehensive decision-making. Screening criteria include a comprehensive decision score exceeding the execution threshold, a minimum resource matching score meeting the minimum requirement, and a true preemption feasibility flag. Preemption execution instructions contain specific preemption steps and parameter settings. These instructions include the identifier of the preempted task, the preemption execution time, the resource recovery method, and the startup parameters for the new task. High-priority tasks are marked for preemption by setting a flag bit and a preemption level. The flag bit indicates whether preemption is performed, and the preemption level reflects the urgency and execution priority of the preemption. The preemption target node identifier records information about the specific compute node executing the preemption operation, including the node's physical identifier, network address, and current load status.
[0091] The migration necessity determination process analyzes preempted low-priority simulation tasks in the preemption candidate set. The determination algorithm evaluates whether the preempted task should be migrated to another node for continued execution. This migration necessity assessment considers factors such as the task's remaining execution time, business importance, and migration cost. When the preempted task has a long remaining execution time and possesses a certain business value, migrating to another node for continued execution is more economical than terminating it directly. The business importance of a task is assessed by analyzing its impact on the overall production process and the loss of execution delays. Workshop production simulation tasks typically have high business importance because their results directly impact production decisions. Logistics and distribution simulation tasks have medium business importance because they affect material supply and inventory management. Equipment maintenance simulation tasks have relatively low business importance because they are primarily used for preventive maintenance planning. Migration cost assessment includes the time cost of preserving task state, the network cost of data transmission, and the management cost of resource allocation at the target node. When the migration cost exceeds the benefit of continuing the task, the task is terminated rather than migrated. Migration necessity determination is performed by comparing the value of continuing the task with the migration cost. When the execution value significantly exceeds the migration cost, the task is marked as requiring migration; otherwise, it is terminated. Low-priority tasks are marked for migration by setting a migration flag and migration urgency. The migration flag indicates whether the task requires migration, and the migration urgency reflects the time requirement and priority of the migration operation. The migration task identifier list contains information about all preempted tasks that need to be migrated. Each entry in the list records key information such as the task identifier, current node, target node candidate, and migration time window.
[0092] In a specific embodiment, the process of executing step S105 may specifically include the following steps:
[0093] The execution progress percentage and remaining execution time of each simulation task in the task execution mapping table are collected and processed in real time. The completion status of workshop production simulation tasks, logistics distribution simulation tasks, and equipment maintenance simulation tasks are counted to obtain simulation progress monitoring data and task completion time prediction values.
[0094] Based on the simulation progress monitoring data, the resource consumption statistics of each computing node's CPU usage changes, memory usage fluctuations, and GPU computing load are processed. The resource consumption rate and peak load per unit time are calculated to obtain computing resource consumption data and node performance indicators.
[0095] Based on the predicted task completion time, the execution efficiency and resource utilization efficiency of the current scheduling strategy are evaluated. Task waiting time, response delay, and throughput indicators are analyzed to obtain scheduling efficiency evaluation indicators and performance bottleneck identification results.
[0096] Correlate and analyze computing resource consumption data with scheduling efficiency evaluation indicators to identify scheduling performance optimization directions and resource allocation adjustment needs, and obtain scheduling performance feedback data and scheduling strategy optimization suggestions;
[0097] The task priority deviation and execution efficiency anomaly in the scheduling performance feedback data are corrected, and the business importance score and priority sorting rules of the simulation tasks are adjusted to obtain an updated dynamic priority queue.
[0098] Specifically, real-time data collection and processing continuously monitors the execution status of each simulation task in the task execution mapping table. This data collection process accesses the task monitoring interface of each compute node at regular intervals through a timed polling mechanism to obtain the current execution status of the task. The execution progress percentage is calculated by calculating the ratio of completed workload to total workload. The workload of workshop production simulation tasks is measured by simulation time step and computational complexity; the workload of logistics and distribution simulation tasks is measured by the number of path calculations and scheduling decisions; and the workload of equipment maintenance simulation tasks is measured by the number of state prediction steps and maintenance plan generation. The remaining execution time is predicted by analyzing the task's historical execution rate and remaining workload. The prediction algorithm considers rate variations and resource availability fluctuations during task execution. Tasks execute more slowly during computationally intensive phases and more quickly during data transfer phases. Completion status statistics categorize and summarize the execution status of different types of simulation tasks. The statistical processing identifies the number of completed tasks, the progress distribution of ongoing tasks, and the queue length of tasks waiting to be executed. The average execution time and completion rate trends for each type of task are also analyzed. Simulation progress monitoring data includes real-time progress information, status change records, and abnormal event logs for each simulation task. The data uses a time series structure to record the historical trajectory of task status changes. Task completion time predictions are calculated using time series prediction methods such as linear regression and exponential smoothing. The prediction process takes into account the seasonality and cyclical nature of task execution. When the workshop production load is high, the execution time of simulation tasks is relatively prolonged. When computing resources are sufficient, the task execution time is significantly shortened.
[0099] Resource consumption statistics analyze resource usage patterns and trends for each compute node based on simulation progress monitoring data. A statistical algorithm correlates the relationship between simulation task execution progress and node resource consumption. CPU utilization fluctuations are measured by monitoring the workload and frequency regulation of the processor cores. When a workshop production simulation task enters a complex computational phase, CPU utilization rises sharply, while when the task is waiting for data, CPU utilization returns to a baseline level. Memory usage fluctuations are analyzed by tracking the frequency and scale of memory allocation and release operations. Logistics and distribution simulation tasks frequently load route data and scheduling information, leading to periodic memory fluctuations. Equipment maintenance simulation tasks maintain relatively stable memory usage due to relatively stable datasets. GPU computational load is assessed by monitoring graphics processing unit (GPU) core utilization and video memory access frequency. 3D visualization rendering and model computation are the primary sources of GPU load, and GPU computational load increases significantly as simulation scene complexity increases. The resource consumption rate per unit time is calculated by taking the time derivative of resource usage. This rate reflects the intensity and trend of resource demand for task execution. A high rate indicates a computationally intensive task, while a low rate indicates a waiting or data transfer phase. Peak load is determined by identifying the maximum value and duration of resource utilization. The timing of peak load is closely related to the execution phase of the simulation task. For workshop production simulation tasks, peak load typically occurs during the production line optimization calculation phase, while for logistics and distribution simulation tasks, peak load occurs during the path planning and conflict detection phase. Computing resource consumption data is stored in a multidimensional time series format, recording the consumption history and statistical characteristics of various resource types. Node performance metrics include key performance parameters such as average response time, resource utilization efficiency, and task processing throughput.
[0100] Performance evaluation quantitatively analyzes the overall performance of a scheduling strategy based on predicted task completion times. The evaluation algorithm measures scheduling effectiveness by comparing the deviation between actual execution results and expected targets. Execution efficiency is assessed by calculating the ratio of task completion time to the theoretical optimal time. The theoretical optimal time assumes an ideal execution environment with unlimited resources and no scheduling overhead. Actual execution time includes scheduling overhead such as task waiting, resource allocation, and context switching. Resource utilization efficiency is assessed by analyzing the ratio of resource idle time to total available time. An efficient scheduling strategy maximizes resource utilization and minimizes idle time. Task waiting time is calculated by recording the interval between task submission and execution start. The length of this waiting time reflects the responsiveness of the scheduling algorithm and the efficiency of resource allocation. Long waiting times for high-priority tasks indicate room for optimization in the scheduling strategy. Response latency is assessed by measuring the interval between the start of task execution and the generation of the first valid output. Response latency is related to task initialization overhead and data loading time. Optimized scheduling strategies can reduce response latency through preloading and caching mechanisms. Throughput is measured by counting the number of tasks completed per unit time and the amount of data processed. Throughput reflects the overall processing capacity of the distributed simulation engine and the effectiveness of the scheduling strategy. The scheduling efficiency evaluation index is derived by comprehensively weighting various performance indicators. The weight allocation is determined based on business priorities and performance goals. The performance bottleneck identification results are derived by analyzing the limiting factors and improvement potential of each indicator. When CPU resources become a limiting factor, it is recommended to add computing nodes. When network bandwidth becomes a bottleneck, it is recommended to optimize the data transmission strategy.
[0101] The correlation analysis process statistically correlates and causally analyzes the resource consumption data with the scheduling efficiency evaluation indicators. The correlation algorithm identifies the quantitative relationship between resource consumption and scheduling efficiency through correlation coefficient calculation and regression analysis. The correlation analysis of resource consumption and task completion time reveals the degree of influence of different resource types on execution efficiency. When CPU resources are sufficient, task completion time is significantly shortened. When memory capacity is insufficient, task execution time is significantly prolonged. GPU performance limitations mainly affect visualization-related simulation tasks. The scheduling performance optimization direction is determined by identifying performance bottlenecks and improper resource allocation. When it is found that a certain type of task often encounters resource contention, it is recommended to adjust the resource allocation strategy. When it is found that the load distribution is uneven, it is recommended to improve the load balancing mechanism. The resource configuration adjustment requirement is determined by analyzing the resource utilization rate distribution and task execution mode. The demand analysis considers the peak resource demand and average utilization level. When the peak resource demand is much higher than the average level, it is recommended to increase the resource buffer. When the average utilization rate is low, it is recommended to optimize the resource allocation algorithm. The scheduling performance feedback data includes historical trends of performance indicators, abnormal event records, and optimization suggestions. The feedback data uses a structured format for automated analysis and decision support. The scheduling strategy optimization suggestions are generated through expert rules and machine learning methods. The suggestions include parameter adjustment schemes, algorithm improvement directions, and resource configuration optimization strategies.
[0102] The correction process makes targeted adjustments to the problems and abnormalities identified in the scheduling performance feedback data. The correction algorithm analyzes the root causes of task priority deviation and the influencing factors of execution efficiency abnormalities. Task priority deviation is identified by comparing the difference between expected priority and actual execution priority. The causes of deviation include inaccurate business importance evaluation, resource demand prediction deviation, and improper scheduling algorithm parameter setting. Execution efficiency abnormalities are identified by detecting task execution time exceeding the expected range or abnormal fluctuations in resource utilization. Abnormal situations include task execution time far exceeding expectations, resource consumption exceeding the normal range, and abnormal increase in task execution failure rate. Business importance score adjustment is made by re-evaluating the impact and urgency of tasks on production processes. The adjustment process considers the latest changes in business demand and production plan adjustments. When the business value of a certain simulation task changes, the importance score is adjusted accordingly. Priority sorting rule adjustment is made by modifying the priority calculation formula and weight allocation. The adjustment goal is to make the priority sorting better reflect the actual business demand and resource constraints. When it is found that the current sorting rule causes important tasks to be delayed, the rule is optimized in a timely manner. Dynamic priority queue updating is done by recalculating the priority values of tasks and adjusting the queue sorting. The updating process ensures that the priority order of tasks in the queue is consistent with the latest business demand and resource status.
[0103] The above describes the automobile manufacturing industry data distributed processing method based on digital twin in the embodiment of the present application. The following describes the automobile manufacturing industry data distributed processing system based on digital twin in the embodiment of the present application. Figure 2 In the embodiments of the present application, an embodiment of a distributed processing system for automobile manufacturing industry data based on digital twins includes:
[0104] The acquisition module is used to collect and process production abnormal events, decision urgency parameters, and simulation task types in the digital twin system in real time through business status monitoring to obtain a business priority evaluation data set;
[0105] A sorting module is used to perform priority sorting on multiple simulation tasks according to the business priority evaluation data set to obtain a dynamic priority queue and a resource demand matrix;
[0106] An input module is used to input the dynamic priority queue into a distributed simulation engine, perform load balancing on the CPU occupancy, memory usage, and GPU utilization of the computing nodes, and obtain a node resource allocation plan;
[0107] A scheduling module is used to perform preemptive scheduling on high-priority simulation tasks according to the node resource allocation scheme, migrate low-priority tasks to idle nodes, and obtain a task execution mapping table;
[0108] The monitoring module is used to dynamically monitor the simulation progress and computing resource consumption in the task execution mapping table through an adaptive scheduling mechanism, obtain scheduling performance feedback data and update the dynamic priority queue.
[0109] above Figure 2 The automobile manufacturing industry data distributed processing system based on digital twins in the embodiment of the present invention is described in detail from the perspective of modular functional entities. The automobile manufacturing industry data distributed processing equipment based on digital twins in the embodiment of the present invention is described in detail from the perspective of hardware processing.
[0110] Reference Figure 3 In the embodiment of the present invention, a distributed processing device for automobile manufacturing industry data based on digital twin is also provided. The distributed processing device for automobile manufacturing industry data based on digital twin can be a server, and its internal structure can be as follows: Figure 3As shown. The automobile manufacturing industry data distributed processing device based on digital twins includes a processor, memory, display screen, input device, network interface and database connected through a system bus. Among them, the computer-designed processor is used to provide computing and control capabilities. The memory of the automobile manufacturing industry data distributed processing device based on digital twins includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the automobile manufacturing industry data distributed processing device based on digital twins is used to store the corresponding data in this embodiment. The network interface of the automobile manufacturing industry data distributed processing device based on digital twins is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the above method is implemented.
[0111] Those skilled in the art will understand that Figure 3 The structure shown is only a block diagram of part of the structure related to the solution of the present invention, and does not constitute a limitation on the digital twin-based automobile manufacturing industry data distributed processing equipment to which the solution of the present invention is applied.
[0112] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions. When the instructions are run on a computer, the computer executes the steps of the distributed processing method for automobile manufacturing industry data based on digital twins.
[0113] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0114] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a digital twin-based automotive manufacturing industry data distributed processing device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.
[0115] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A distributed processing method for automobile manufacturing industry data based on digital twins, characterized by: The method comprises: Through business status monitoring, production abnormal events, decision urgency parameters, and simulation task types in the digital twin system are collected and processed in real time to obtain a business priority evaluation data set; Prioritizing multiple simulation tasks according to the business priority evaluation data set to obtain a dynamic priority queue and a resource demand matrix; The dynamic priority queue is input into a distributed simulation engine, and load balancing is performed on the CPU occupancy, memory usage, and GPU utilization of the computing nodes to obtain a node resource allocation plan; Preemptively schedule high-priority simulation tasks according to the node resource allocation scheme, migrate low-priority tasks to idle nodes, and obtain a task execution mapping table; The simulation progress and computing resource consumption in the task execution mapping table are dynamically monitored and processed through an adaptive scheduling mechanism to obtain scheduling performance feedback data and update the dynamic priority queue.
2. The method for distributed processing of automobile manufacturing industry data based on digital twins according to claim 1 is characterized in that: The business status monitoring is used to collect and process production abnormal events, decision urgency parameters, and simulation task types in the digital twin system in real time to obtain a business priority evaluation data set, including: Perform status recognition and processing on equipment failure signals, production line downtime events, and quality abnormality alarms in automobile manufacturing workshops to obtain production abnormality event identification codes and abnormality severity level parameters; Performing emergency quantification on the current production plan execution status, order delivery time, and inventory buffer level according to the production abnormal event identification code to obtain decision urgency parameters and time window constraints; Based on the abnormal severity level parameters, the workshop production simulation tasks, logistics distribution simulation tasks and equipment maintenance simulation tasks to be executed are classified and labeled to obtain simulation task type labels and computational complexity weights; Performing weighted fusion processing on the decision urgency parameter, the time window constraint condition, and the simulation task type label to obtain a business importance score for each simulation task; The business importance scores are sorted according to the degree of production impact to obtain the business priority assessment data set.
3. The method for distributed processing of automobile manufacturing industry data based on digital twins according to claim 1 is characterized in that: The prioritizing of the plurality of simulation tasks according to the business priority evaluation data set to obtain a dynamic priority queue and a resource requirement matrix includes: Performing numerical comparison processing on the business importance scores of each simulation task in the business priority evaluation data set to obtain a simulation task priority sequence and a grouping of tasks with the same priority; According to the simulation task priority sequence, resource statistics processing is performed on the CPU demand, memory usage, and GPU computing power of the workshop production simulation task, the logistics distribution simulation task, and the equipment maintenance simulation task to obtain a single task resource consumption vector; Based on the grouping of tasks with the same priority, concurrent analysis and processing are performed on the execution time estimation and resource conflict probability of the simulation tasks in each group to obtain a dependency graph between tasks and concurrent execution constraints; Performing matrix arrangement processing on the single-task resource consumption vector according to the simulation task priority sequence to obtain the resource demand matrix and resource pre-allocation strategy; The high-priority simulation tasks in the inter-task dependency graph are queued according to the execution order to obtain the dynamic priority queue.
4. The method for distributed processing of automobile manufacturing industry data based on digital twins according to claim 1 is characterized in that: The dynamic priority queue is input into a distributed simulation engine, and load balancing is performed on the CPU occupancy, memory usage, and GPU utilization of the computing nodes to obtain a node resource allocation plan, including: Input the dynamic priority queue into the task scheduler of the distributed simulation engine, perform real-time monitoring and processing on the current CPU occupancy, memory usage and GPU utilization of each computing node, and obtain the node resource state vector and available resource capacity data; Performing a load balance evaluation process on the master node, computing node, and storage node in the distributed simulation engine according to the node resource state vector to obtain a load imbalance coefficient and a resource bottleneck identification mark; Performing migration candidate selection processing on the simulation tasks of the high-load computing nodes based on the load imbalance coefficient, performing matching analysis on the resource-intensive tasks and the lightweight tasks, and obtaining a task migration candidate list and a target node matching pair; Perform constraint checking on the task migration candidate list and the available resource capacity data, perform capacity verification on the remaining number of CPU cores, available memory size, and GPU memory space of each target node, and obtain a feasibility allocation matrix; The feasibility allocation matrix is optimized and allocated according to the principle of maximizing resource utilization to obtain the node resource allocation scheme and the load balancing configuration parameters of the distributed simulation engine.
5. The method for distributed processing of automobile manufacturing industry data based on digital twins according to claim 1 is characterized in that: The high-priority simulation tasks are preemptively scheduled according to the node resource allocation scheme, and low-priority tasks are migrated to idle nodes to obtain a task execution mapping table, including: According to the node resource allocation plan, a preemption determination process is performed on the priority status of the currently running workshop production simulation task, logistics distribution simulation task, and equipment maintenance simulation task to obtain a high-priority task preemption flag and a low-priority task migration flag; Performing task suspension processing on the target computing node's running state based on the high-priority task preemption flag, saving the execution context and intermediate calculation results of the low-priority simulation task being executed, and obtaining task checkpoint data and task recovery information; Input the task checkpoint data into the idle node searcher, perform matching and searching on the idle computing nodes that meet the resource requirements in the distributed simulation engine, and obtain the available node list and node resource margin parameters; Re-deploy low-priority simulation tasks according to the available node list, transmit the task recovery information to the target idle node and restore the task execution status, and obtain the task migration completion status and node load update data; A mapping relationship is established between the task migration completion status and the node load update data, and the execution position and resource occupancy of each simulation task on the distributed computing node are recorded to obtain the task execution mapping table.
6. The method for distributed processing of automobile manufacturing industry data based on digital twins according to claim 5 is characterized in that: The priority status of the currently running workshop production simulation task, logistics distribution simulation task and equipment maintenance simulation task is preempted and determined according to the node resource allocation scheme to obtain a high-priority task preemption flag and a low-priority task migration flag, including: Performing a task priority scan on the running queues of each computing node in the node resource allocation scheme, extracting the priority values of the currently executed workshop production simulation tasks, logistics distribution simulation tasks, and equipment maintenance simulation tasks, and obtaining the current task priority list and node task distribution status; Compare and determine the priority values of the newly arrived high-priority simulation task and the currently running task according to the current task priority list, set a preemption threshold and identify the low-priority tasks that need to be preempted, and obtain a priority difference matrix and a set of preemption candidate tasks; Based on the priority difference matrix, the resource occupancy of each computing node and the resource requirements of the new task are matched and calculated to determine whether the resource release conditions of preemptive scheduling are met, thereby obtaining a resource matching score and a preemptive feasibility flag; Performing a comprehensive decision-making process on the set of preemption candidate tasks and the resource matching score, generating a preemption execution instruction for the high-priority simulation task that meets the preemption conditions, and obtaining the preemption flag of the high-priority task and the preemption target node identifier; The preempted low-priority simulation tasks in the preemption candidate task set are subjected to migration necessity determination processing, tasks that need to be migrated to other nodes for continued execution are marked, and the low-priority task migration flags and migration task identifier list are obtained.
7. The method for distributed processing of automobile manufacturing industry data based on digital twins according to claim 1 is characterized in that: The method of dynamically monitoring the simulation progress and computing resource consumption in the task execution mapping table through the adaptive scheduling mechanism to obtain scheduling performance feedback data and update the dynamic priority queue includes: The execution progress percentage and remaining execution time of each simulation task in the task execution mapping table are collected and processed in real time, and the completion status of the workshop production simulation task, the logistics distribution simulation task, and the equipment maintenance simulation task are counted to obtain simulation progress monitoring data and task completion time prediction values; Perform resource consumption statistics on the CPU usage changes, memory usage fluctuations, and GPU computing load of each computing node based on the simulation progress monitoring data, calculate the resource consumption rate and peak load per unit time, and obtain computing resource consumption data and node performance indicators; Based on the task completion time prediction value, the execution efficiency and resource utilization efficiency of the current scheduling strategy are evaluated, and the task waiting time, response delay and throughput indicators are analyzed to obtain scheduling efficiency evaluation indicators and performance bottleneck identification results; Correlation analysis is performed on the computing resource consumption data and the scheduling efficiency evaluation index to identify scheduling performance optimization directions and resource configuration adjustment requirements, thereby obtaining scheduling performance feedback data and scheduling strategy optimization suggestions; Correction processing is performed on task priority deviations and execution efficiency anomalies in the scheduling performance feedback data, and the business importance scores and priority sorting rules of the simulation tasks are adjusted to obtain the updated dynamic priority queue.
8. A distributed processing system for automobile manufacturing industry data based on digital twins, characterized by: For implementing the automobile manufacturing industry data distributed processing method based on digital twin according to any one of claims 1 to 7, the automobile manufacturing industry data distributed processing system based on digital twin comprises: The acquisition module is used to collect and process production abnormal events, decision urgency parameters, and simulation task types in the digital twin system in real time through business status monitoring to obtain a business priority evaluation data set; A sorting module is used to perform priority sorting on multiple simulation tasks according to the business priority evaluation data set to obtain a dynamic priority queue and a resource demand matrix; An input module is used to input the dynamic priority queue into a distributed simulation engine, perform load balancing on the CPU occupancy, memory usage, and GPU utilization of the computing nodes, and obtain a node resource allocation plan; A scheduling module is used to perform preemptive scheduling on high-priority simulation tasks according to the node resource allocation scheme, migrate low-priority tasks to idle nodes, and obtain a task execution mapping table; The monitoring module is used to dynamically monitor the simulation progress and computing resource consumption in the task execution mapping table through an adaptive scheduling mechanism, obtain scheduling performance feedback data and update the dynamic priority queue.
9. A distributed processing device for automobile manufacturing industry data based on digital twins, characterized in that: It includes a memory and a processor, the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, it implements the distributed processing method of automobile manufacturing industry data based on digital twins as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the processor executes the distributed processing method for automobile manufacturing industry data based on digital twins as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Multitask scheduling method and device based on heterogeneous distributed cluster
CN118227291A
Industrial digital factory cooperative work method and cloud platform
CN120295258A