Task scheduling method and device in circuit simulation, medium and product
By sub-matrix division and dynamic task scheduling of sparse matrix, the load imbalance problem in sparse matrix solving is solved, and the computational efficiency of large-scale simulation circuit simulation is improved.
Patent Information
- Application Number
- CN202510788368.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-13
AI Technical Summary
In the prior art, sparse matrix solution has load imbalance problems in large-scale simulation circuit simulation, resulting in low computational efficiency.
By obtaining the target sparse matrix, dividing it into multiple submatrices, and assigning the calculation tasks to the computing node cluster, the monitor periodically monitors the load information and task execution status, calculates the node weight, judges whether the scheduling conditions are met based on the weight, and migrates the lighter tasks to the heavier node to achieve load balancing.
It effectively avoids load imbalance in sparse matrix solution and improves computing efficiency.
Smart Images

Figure CN120295740A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of circuit simulation, and particularly relates to a task scheduling method, device, medium and product in circuit simulation. Background Art
[0002] In large-scale analog circuit simulation, the solution of sparse matrices is a core link. The circuit netlist generates a linear equation system Ax = b through the modified nodal analysis method; where A is a sparse matrix representing the connection relationship of components (resistors, capacitors, transistors, etc.) in the circuit, and b is an excitation vector, such as the input values of voltage sources and current sources. Since current integrated circuits may contain billions of components, the matrix scale can reach more than one million orders, and since each component only affects a few nodes connected to it, the proportion of non-zero elements in matrix A is extremely low.
[0003] The solution efficiency of sparse matrices directly affects the simulation performance. Therefore, accelerating the solution of sparse matrices is a key technical challenge in large-scale analog circuit simulation. The current solution of sparse matrices mainly adopts a static partitioning strategy, and its typical process includes matrix reordering, uniform task block partitioning, and multi-node parallel computing. However, this static partitioning strategy does not consider the non-uniform distribution characteristics of non-zero elements in sparse matrices. For example, some task blocks may contain dense non-zero elements, while other blocks are sparse or even all zero. This partitioning method will cause some nodes to be computationally intensive and overloaded, while other nodes are idle due to too light tasks or ineffective calculations, resulting in load imbalance in the solution of the entire sparse matrix and low computational efficiency.
[0004] In summary, how to solve the load imbalance problem in the solution of sparse matrices during circuit simulation and significantly improve the computational efficiency is an issue to be solved currently. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a task scheduling method, device, medium and product in circuit simulation, which can solve the load imbalance problem in the solution of sparse matrices during circuit simulation and significantly improve the computational efficiency. The specific solutions are as follows: In the first aspect, the present application discloses a task scheduling method in circuit simulation, including: Obtain a target sparse matrix generated based on the connection relationship between circuit components in a target circuit netlist, and divide the target sparse matrix into multiple sub-matrices; wherein, the target circuit netlist is the circuit netlist used during circuit simulation, and each sub-matrix corresponds to a computing task; Allocate each computing task to a preset computing node cluster to utilize each computing node in the computing node cluster to execute the corresponding computing task; Periodically monitor the load information and task execution status of each computing node using a preset monitor, and calculate the node weights of each computing node based on the load information and task execution status; Judge whether the current situation meets the preset task scheduling conditions based on the node weights of each node. If it meets, schedule the unexecuted target computing tasks in the first computing node to the second computing node; the node weight of the first computing node is less than that of the second computing node.
[0006] Optionally, before dividing the target sparse matrix into multiple sub-matrices, it further includes: Determine the target sorting method according to the matrix characteristics of the target sparse matrix; where the target sorting method is any one of the minimum degree sorting method, spectral sorting method, and graph partitioning method; Use the target sorting method to rearrange the rows and columns in the target sparse matrix to obtain the re-ordered target sparse matrix.
[0007] Optionally, dividing the target sparse matrix into multiple sub-matrices includes: Count the number of non-zero elements in each row and each column of the target sparse matrix to obtain non-zero element distribution information; Divide the target sparse matrix into multiple sub-matrices based on the non-zero element distribution information; where the difference in the number of non-zero elements between different sub-matrices is within a preset range.
[0008] Optionally, allocating each computing task to a preset computing node cluster includes: Determine the preset computing node cluster and determine the number of computing nodes in the computing node cluster; Based on the number of each computing task and the number of computing nodes, evenly allocate each computing task to the task queues corresponding to each computing node in the computing node cluster.
[0009] Optionally, the load information includes CPU utilization and current remaining memory, and the task execution status is the current length of the task queue; Correspondingly, calculating the node weights of each computing node based on the load information and task execution status includes: Use the preset weight coefficients to perform weighted calculations on the CPU utilization, current remaining memory, and current length corresponding to each computing node to obtain the node weights of the computing nodes.
[0010] Optionally, the node weights are calculated based on a preset weight formula; where the preset weight formula is: ; Where , , are all weight coefficients.
[0011] Optionally, after periodically monitoring the load information and task execution status of each computing node using a preset monitor, it further includes: If the CPU utilization rate of any computing node is greater than the preset utilization rate threshold for a continuous preset number of cycles, or if the current remaining memory of any computing node is lower than the preset memory threshold, then mark any computing node as an abnormal computing node, and schedule the unexecuted computing tasks in any computing node to the remaining computing nodes.
[0012] Optionally, determining whether the current meets the preset task scheduling condition based on each node weight includes: Calculating the weight average value based on each node weight, and calculating the corresponding weight standard deviation using each node weight and the weight average value; Determining whether the weight standard deviation is greater than the preset standard deviation threshold; If it is greater, then determine that the current meets the preset task scheduling condition, otherwise determine that the current does not meet the preset task scheduling condition.
[0013] Optionally, scheduling the unexecuted target computing task in the first computing node to the second computing node includes: Determine the first computing node with a node weight lower than the first preset threshold, and determine the second computing node with a node weight higher than the second preset threshold; where the first preset threshold is less than the second preset threshold; Determine the unexecuted target computing task according to the task execution status corresponding to the first computing node, and schedule the target computing task from the first computing node to the second computing node.
[0014] Optionally, scheduling the unexecuted target computing task in the first computing node to the second computing node includes: Sorting each computing node in descending order of node weight to obtain a sorting result; Determine several computing nodes at the preset tail position in the sorting result as the first computing node, and determine several computing nodes at the preset head position in the sorting result as the second computing node; Determine the unexecuted target computing task according to the task execution status corresponding to the first computing node, and schedule the target computing task from the first computing node to the second computing node.
[0015] Optionally, the task scheduling method in the circuit simulation of the present application further includes: Monitoring whether each computing task has been executed and completed; If all have been executed and completed, then obtain the task execution results in each computing node, and summarize each task execution result to obtain the solution result of the target sparse matrix.
[0016] Optionally, in the process of scheduling the unexecuted target computing tasks in the first computing node to the second computing node, the following steps are further included: If there are multiple second computing nodes, for the unexecuted target computing tasks in the first computing node, it is determined whether there are computing tasks adjacent to the target computing tasks among the computing tasks corresponding to each second computing node based on the target position relationship; wherein, the target position relationship is the position relationship of the sub-matrices corresponding to each computing task in the target sparse matrix; If there are, the target computing task is scheduled to the second computing node where the computing task adjacent to the target computing task is located.
[0017] In a second aspect, the present application discloses an electronic device, including: A memory for storing a computer program; A processor for executing the computer program to implement the steps of the task scheduling method in the circuit simulation disclosed above.
[0018] In a third aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the steps of the task scheduling method in the circuit simulation disclosed above are implemented.
[0019] In a fourth aspect, the present application discloses a computer program product including computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the task scheduling method in the circuit simulation disclosed above are implemented.
[0020] It can be seen that the present application obtains a target sparse matrix generated based on the connection relationship between circuit elements in a target circuit netlist, and divides the target sparse matrix into multiple sub-matrices; wherein, the target circuit netlist is the circuit netlist used in circuit simulation, and each sub-matrix corresponds to a computing task; allocates each computing task to a preset computing node cluster to use each computing node in the computing node cluster to execute the corresponding computing task; uses a preset monitor to periodically monitor the load information and task execution status of each computing node, and calculates the node weight of each computing node based on the load information and task execution status; determines whether the preset task scheduling condition is currently satisfied based on each node weight, and if so, schedules the unexecuted target computing tasks in the first computing node to the second computing node; the node weight of the first computing node is less than the node weight of the second computing node.
[0021] Beneficial effects: In this application, first, a target sparse matrix is obtained based on the connection relationships between circuit elements in the target circuit netlist, where the target circuit netlist is the circuit netlist used in circuit simulation; then the target sparse matrix is partitioned to obtain multiple sub-matrices, and each sub-matrix corresponds to a computing task. Further, each computing task is first assigned to a preset computing node cluster so that each computing node in the computing node cluster starts to execute the corresponding computing task. After that, this application uses a preset monitor to periodically monitor the load information and task execution status of each computing node to dynamically monitor the load and task execution of each computing node, and calculate the node weights of each computing node accordingly. In this application, it is stipulated that the higher the node weight, the more suitable it is to process computing tasks. Finally, it is determined whether the current situation meets the preset task scheduling condition based on each node weight. If it meets, the unexecuted target computing task in the first computing node is scheduled to the second computing node, where the node weight of the first computing node is less than that of the second computing node. That is, when the current situation meets the preset task scheduling condition, the unexecuted target computing task in the computing node with a lower node weight is scheduled to the second computing node with a higher node weight, avoiding the situation where some computing nodes have high time consumption and load due to undertaking the computing tasks corresponding to dense sub-matrices. By scheduling, the unexecuted tasks of this node are migrated to relatively idle computing nodes, which can avoid the load imbalance problem in sparse matrix solution and improve the computing efficiency at the same time. Description of the Drawings
[0022] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0023] Figure 1 It is a flowchart of a task scheduling method in circuit simulation disclosed in the present application; Figure 2 It is a schematic diagram of a specific processing architecture disclosed in the present application; Figure 3 It is a flowchart of a specific task scheduling method in circuit simulation disclosed in the present application; Figure 4 It is a schematic diagram of the structure of a monitor disclosed in the present application; Figure 5 It is a schematic diagram of the structure of a scheduler disclosed in the present application; Figure 6 It is a schematic diagram of the structure of a task scheduling device in circuit simulation disclosed in the present application; Figure 7 A structural diagram of an electronic device disclosed in this application. Specific implementation manners
[0024] The following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0025] In large-scale analog circuit simulation, the solution of sparse matrices is the core link. The solution efficiency of sparse matrices directly affects the simulation performance. Therefore, accelerating the solution of sparse matrices is a key technical challenge in large-scale analog circuit simulation. Currently, the solution of sparse matrices mainly adopts a static partitioning strategy. However, this static partitioning strategy does not consider the non-uniform distribution characteristics of the non-zero elements of the sparse matrix. For example, some task blocks may contain dense non-zero elements, while other blocks are sparse or even all zero. This partitioning method will cause some nodes to be computationally intensive and overloaded, while other nodes are idle due to too light tasks or ineffective calculations, resulting in load imbalance in the solution of the entire sparse matrix and low computational efficiency.
[0026] Therefore, the embodiments of this application disclose a task scheduling method, device, medium, and product in circuit simulation, which can solve the load imbalance problem in the solution of sparse matrices during circuit simulation and significantly improve the computational efficiency.
[0027] See Figure 1 As shown, the embodiments of this application disclose a task scheduling method in circuit simulation. The method includes: Step S11: Obtain a target sparse matrix generated based on the connection relationship between circuit elements in the target circuit netlist, and divide the target sparse matrix into multiple sub-matrices; wherein, the target circuit netlist is the circuit netlist used during circuit simulation, and each sub-matrix corresponds to a computational task.
[0028] In this embodiment, during the actual analog circuit simulation process, it is first necessary to generate a linear equation system Ax = b for the circuit netlist used in circuit simulation through the modified nodal analysis method; where A is the sparse matrix, that is, the target sparse matrix to be solved in this embodiment, and the target sparse matrix specifically represents the connection relationship between circuit elements.
[0029] Further, considering that the scale of the target sparse matrix can reach over a million orders, and the serial solution speed of a sparse matrix of a million orders is extremely slow, in order to optimize the computing efficiency and resource utilization rate, the present application divides the target sparse matrix to obtain multiple sub-matrices. Among them, each sub-matrix corresponds to a computing task, so that multiple sub-matrices can be assigned to different computing nodes for parallel processing. It should be noted that each sub-matrix is a square matrix with the same number of rows and columns.
[0030] In addition, it should be pointed out that before dividing the target sparse matrix into multiple sub-matrices, it also includes: determining a target sorting method according to the matrix characteristics of the target sparse matrix; where the target sorting method is any one of the minimum degree sorting method, the spectral sorting method, and the graph partitioning method; using the target sorting method to re-arrange the rows and columns in the target sparse matrix to obtain a re-ordered target sparse matrix. That is, after obtaining the target sparse matrix and before dividing the target sparse matrix into multiple sub-matrices, the sparse matrix can also be preprocessed, such as re-ordering the target sparse matrix. The purpose is to reduce the number of filled zeros and reduce the matrix bandwidth. By re-arranging the rows and columns of the matrix, the non-zero elements are gathered together as much as possible and close to the diagonal, thereby improving the efficiency of locally solving the sub-matrices on each computing node subsequently.
[0031] In the specific implementation, the corresponding target sorting method can be determined according to the matrix characteristics of the target sparse matrix, so as to use the target sorting method to re-arrange the rows and columns in the target sparse matrix to obtain a re-ordered target sparse matrix; the target sorting method includes but is not limited to the minimum degree sorting method, the spectral sorting method, and the graph partitioning method. For example, if the target sparse matrix is a symmetric matrix and the minimization of filled zeros is pursued, the minimum degree sorting method can be preferentially used; if the target sparse matrix is an asymmetric matrix or has a relatively complex matrix structure, the spectral sorting method or the graph partitioning method is preferentially selected; for ultra-large-scale matrices (such as those with more than a million orders), the graph partitioning method is adopted.
[0032] Step S12: Assign each computing task to a preset computing node cluster to use each computing node in the computing node cluster to execute the corresponding computing task.
[0033] In this embodiment, after dividing the target sparse matrix into multiple sub-matrices, the computing tasks corresponding to each sub-matrix are assigned to a preset computing node cluster to use each computing node in the computing node cluster to start executing the corresponding computing task. It can be understood that for each divided sub-matrix, local solution can be performed on its own independent computing node. The computing node can specifically be a multi-core processor, a distributed computing node, etc. Among them, the local solution method can adopt a direct method (such as LU decomposition) or an iterative method (such as the conjugate gradient method).
[0034] Step S13: Periodically monitor the load information and task execution status of each computing node by using a preset monitor, and calculate the node weight of each computing node based on the load information and task execution status.
[0035] In this embodiment, the preset monitor is also used to periodically monitor the load information and task execution status of each computing node, so as to dynamically monitor the load and task execution of each computing node, and calculate the node weight of each computing node accordingly. In this embodiment, it is specified that the higher the node weight, the more suitable it is to process computing tasks.
[0036] Step S14: Determine whether the current preset task scheduling condition is satisfied based on each node weight. If it is satisfied, schedule the unexecuted target computing task in the first computing node to the second computing node; the node weight of the first computing node is less than the node weight of the second computing node.
[0037] In this embodiment, it is determined whether the current preset task scheduling condition is satisfied based on each node weight. If it is satisfied, the unexecuted target computing task in the first computing node is scheduled to the second computing node, where the node weight of the first computing node is less than the node weight of the second computing node. That is, when the current preset task scheduling condition is satisfied, the unexecuted target computing task in the computing node with a lower node weight is scheduled to the second computing node with a higher node weight, avoiding the situation where some computing nodes have high time consumption and load due to undertaking the corresponding computing tasks of dense sub-matrices. By scheduling, the unexecuted tasks of this node are migrated to a relatively idle computing node, which can avoid the load imbalance problem in sparse matrix solving and improve the computing efficiency at the same time.
[0038] In a specific implementation manner, scheduling the unexecuted target computing task in the first computing node to the second computing node includes: determining the first computing node whose node weight is lower than the first preset threshold, and determining the second computing node whose node weight is higher than the second preset threshold; where the first preset threshold is less than the second preset threshold; determining the unexecuted target computing task according to the task execution status corresponding to the first computing node, and scheduling the target computing task from the first computing node to the second computing node.
[0039] It can be understood that when the preset task scheduling conditions are currently met, the purpose is to schedule the unexecuted computing tasks in some computing nodes with heavy loads to the computing nodes with lighter loads for execution to achieve load balancing. In this application, the node weights of each computing node are calculated, and it is stipulated that the higher the node weight, the more suitable it is to process computing tasks, that is, it means that its load is lower. Therefore, this application sets a first preset threshold and a second preset threshold, and the first preset threshold is less than the second preset threshold. Then, the computing nodes with node weights lower than the first preset threshold are denoted as first computing nodes, and the computing nodes with node weights higher than the second preset threshold are denoted as second computing nodes. Then, the unexecuted target computing tasks are determined according to the task execution status corresponding to the first computing nodes, and thus the target computing tasks are scheduled from the first computing nodes to the second computing nodes for execution. In this way, the unexecuted target computing tasks in the computing nodes with lower node weights are scheduled to the computing nodes with higher node weights. Among them, the specific values of the first preset threshold and the second preset threshold can be custom-set according to the actual situation, and the specific values are not limited in this embodiment.
[0040] In another specific implementation manner, scheduling the unexecuted target computing tasks in the first computing nodes to the second computing nodes includes: sorting each computing node in descending order of node weight to obtain a sorting result; determining several computing nodes at a preset tail position in the sorting result as the first computing nodes, and determining several computing nodes at a preset head position in the sorting result as the second computing nodes; determining the unexecuted target computing tasks according to the task execution status corresponding to the first computing nodes, and scheduling the target computing tasks from the first computing nodes to the second computing nodes.
[0041] That is, in addition to setting two different thresholds to determine nodes with lower weights and computing nodes with higher weights, the present application can first sort all computing nodes in descending order of node weights to obtain a sorting result, and then determine several computing nodes located at a preset tail position in the sorting result as the first computing nodes, and determine several computing nodes located at a preset head position in the sorting result as the second computing nodes; wherein, the preset tail position corresponds to the position of the minimum node weight in the sorting result, and the preset head position corresponds to the position of the maximum node weight in the sorting result. For example, 3 computing nodes can be selected as the first computing nodes starting from the preset tail position, that is, the 3 computing nodes with the smallest node weights are selected as the first computing nodes, and 3 computing nodes can be selected as the second computing nodes starting from the preset head position in the same way, that is, the 3 computing nodes with the largest node weights are selected as the second computing nodes; it should be noted that the number of the first computing nodes can be the same as or different from the number of the second computing nodes, and the present application does not limit this. Finally, the unexecuted target computing tasks are determined according to the task execution status corresponding to the first computing nodes, and the target computing tasks are scheduled from the first computing nodes to the second computing nodes, so that the unexecuted target computing tasks in the computing nodes with lower node weights can also be scheduled to the computing nodes with higher node weights.
[0042] It should also be noted that in the process of scheduling the unexecuted target computing tasks in the first computing nodes to the second computing nodes, it further includes: if the number of the second computing nodes is multiple, for the unexecuted target computing tasks in the first computing nodes, based on the target position relationship, it is determined whether there are computing tasks adjacent to the target computing tasks among the computing tasks corresponding to each second computing node; wherein, the target position relationship is the position relationship of the sub-matrices corresponding to each computing task in the target sparse matrix; if there is, the target computing task is scheduled to the second computing node where the computing task adjacent to the target computing task is located.
[0043] That is, if multiple second computing nodes are selected, when scheduling the unexecuted target computing tasks in the first computing node to the second computing nodes, the positional relationship of the sub-matrices corresponding to each computing task in the target sparse matrix can also be referred to, and adjacent computing tasks are preferably scheduled to the same computing node for execution. The core purpose is to reduce the communication overhead between nodes and the complexity of cross-node synchronization. For example, in a 10×10 sparse matrix, the elements located in the first 5 rows and the first 5 columns are constructed into a sub-matrix, and the elements located in the first 5 rows and the last 5 columns are also constructed into a sub-matrix. Then these two sub-matrices are adjacent in the sparse matrix. Therefore, for the unexecuted target computing tasks in each first computing node in this embodiment, it is necessary to determine whether there are computing tasks adjacent to the target computing task among the computing tasks corresponding to each second computing node based on the positional relationship of the sub-matrices corresponding to each computing task in the target sparse matrix. If so, the target computing task is scheduled to the second computing node where the computing task adjacent to the target computing task is located for execution.
[0044] On the contrary, if there are no computing tasks adjacent to the target computing task among the computing tasks corresponding to each second computing node, the unexecuted target computing tasks in the first computing node can be randomly scheduled to multiple second computing nodes. However, it should be noted that during the random scheduling process, the balance of task computing amounts should also be maintained among multiple second computing nodes, rather than scheduling all the unexecuted computing tasks in multiple first computing nodes to the same second computing node. In the specific implementation, the total number of unexecuted target computing tasks in each first computing node can also be first counted, and the number of second computing nodes is determined, and then these target computing tasks are evenly distributed to each second computing node.
[0045] In addition, the method of the present application further includes: monitoring whether all computing tasks have been completed; if all have been completed, obtaining the task execution results in each computing node, and summarizing each task execution result to obtain the solution result of the target sparse matrix. It can be understood that in the present application, a large target sparse matrix is split into multiple sub-matrices and calculated in parallel in multiple computing nodes. Therefore, the results calculated in each computing node are local solution results. Therefore, after monitoring that all computing tasks have been completed, it is also necessary to obtain the task execution results in each computing node and summarize each task execution result to obtain the solution result of the target sparse matrix. During the summarization process, global iterative correction can also be performed to ensure that the summarized result meets the accuracy requirements of the overall system. Finally, post-processing steps such as residual calculation and error analysis are performed to verify the accuracy of the solution.
[0046] In addition, the processing architecture applicable to the present application can be as Figure 2As shown in the figure, it specifically includes a preprocessing module, a monitor, a scheduler, and computing nodes. Among them, the preprocessing module is mainly used to reorder and partition sub-matrices of the input large-scale sparse matrix; the monitor is used to monitor the load information and task execution status of each computing node in real time and calculate the node weights; the scheduler is used to schedule computing tasks in real time according to the node weights; the computing nodes are mainly used to complete the local calculation of computing tasks and the synchronization of necessary data.
[0047] It can be seen that in this application, first, a target sparse matrix generated based on the connection relationship between circuit elements in the target circuit netlist is obtained, where the target circuit netlist is the circuit netlist used in circuit simulation; then the target sparse matrix is partitioned to obtain multiple sub-matrices, and each sub-matrix corresponds to a computing task. Further, each computing task is first assigned to a preset computing node cluster, so that each computing node in the computing node cluster starts to execute the corresponding computing task. After that, this application will use a preset monitor to periodically monitor the load information and task execution status of each computing node to dynamically monitor the load and task execution conditions of each computing node, and calculate the node weights of each computing node based on this. In this application, it is stipulated that the higher the node weight, the more suitable it is to process computing tasks. Finally, it is determined whether the current preset task scheduling condition is met based on each node weight. If it is met, the unexecuted target computing task in the first computing node is scheduled to the second computing node, where the node weight of the first computing node is less than that of the second computing node. That is, when the current preset task scheduling condition is met, the unexecuted target computing task in the computing node with a lower node weight is scheduled to the second computing node with a higher node weight, avoiding the situation where some computing nodes have high time consumption and load due to undertaking computing tasks corresponding to dense sub-matrices. By scheduling, the unexecuted tasks of this node are migrated to a relatively idle computing node, which can avoid the load imbalance problem in sparse matrix solving and improve the computing efficiency at the same time.
[0048] See Figure 3 As shown in the figure, the embodiment of this application discloses a specific task scheduling method in circuit simulation. Compared with the previous embodiment, this embodiment further explains and optimizes the technical solution, which specifically includes: Step S21: Obtain a target sparse matrix generated based on the connection relationship between circuit elements in the target circuit netlist; where the target circuit netlist is the circuit netlist used in circuit simulation.
[0049] Step S22: Count the number of non-zero elements in each row and each column of the target sparse matrix to obtain non-zero element distribution information, and divide the target sparse matrix into multiple sub-matrices based on the non-zero element distribution information; where each sub-matrix corresponds to a computing task, and the difference in the number of non-zero elements between different sub-matrices is within a preset range.
[0050] In this embodiment, when dividing the target sparse matrix into multiple sub-matrices, the non-zero element distribution information of the target sparse matrix can be specifically referred to. Among them, the non-zero element distribution information can be obtained by counting the number of non-zero elements in each row and each column of the target sparse matrix. By dividing the target sparse matrix based on the non-zero element distribution information, the purpose is to make the difference in the number of non-zero elements between different sub-matrices within a preset range. That is, when initially dividing the sub-matrices, try to ensure that the number of non-zero elements in each sub-matrix is approximately the same, so that the computational workloads of the corresponding computing tasks do not differ too much. Therefore, when initially allocating each computing task to different computing nodes, the load balance can be maintained as much as possible.
[0051] Step S23: Determine a preset computing node cluster, and determine the number of computing nodes in the computing node cluster. Then, based on the number of each computing task and the number of computing nodes, evenly allocate each computing task to the task queues corresponding to each computing node in the computing node cluster, so as to use each computing node in the computing node cluster to execute the corresponding computing task.
[0052] In this embodiment, it is necessary to determine a preset computing node cluster to determine the number of computing nodes in the cluster, that is, how many computing nodes in the current cluster are used to execute computing tasks. Further, based on the number of each computing task and the number of computing nodes, this application evenly allocates each computing task to the task queues corresponding to each computing node in the computing node cluster, so as to use each computing node in the computing node cluster to execute the corresponding computing task. For example, assume there are 100 computing tasks and 20 computing nodes, then these 100 computing tasks are evenly allocated to 20 computing nodes, and each computing node needs to execute 5 computing tasks.
[0053] Step S24: Use a preset monitor to periodically monitor the load information and task execution status of each computing node; the load information includes CPU utilization rate and current remaining memory, and the task execution status is the current length of the task queue.
[0054] In this embodiment, considering that as the calculation progresses, the load level of the nodes will also change accordingly. For example, even if two sub-matrices have the same number of non-zero elements, due to the different positions of the non-zero elements in the sub-matrices, the required calculation time and the number of computing resources invoked are also different. Therefore, this application uses a monitor to real-time monitor the status of the computing nodes, providing data support for dynamic scheduling to ensure that the system can quickly respond to load changes. In the specific implementation, the monitor will periodically monitor the load information and task execution status of each computing node. For example, it will perform a monitoring every 5 seconds. Among them, the load information includes but is not limited to CPU utilization rate and current remaining memory. For example, it can also include network bandwidth, and the task execution status can specifically be the current length of the task queue. In addition, it can also include task completion rate, calculation error rate, etc.
[0055] It should be noted that after periodically monitoring the load information and task execution status of each computing node by using a preset monitor, it further includes: if the CPU utilization rate of any computing node is greater than the preset utilization rate threshold in continuously preset number of cycles, or if the current remaining memory of any computing node is lower than the preset memory threshold, then mark any computing node as an abnormal computing node, and schedule the unexecuted computing tasks in any computing node to the remaining computing nodes.
[0056] That is, if it is monitored that the CPU utilization rate of any computing node is greater than the preset utilization rate threshold in continuously preset number of cycles. For example, the CPU utilization rate > 90% lasts for 1 minute, indicating that this computing node is already in an overloaded state, or if it is monitored that the current remaining memory of any computing node is lower than the preset memory threshold, it means that this computing node is currently not suitable for executing computing tasks. Therefore, if the above two situations occur, directly mark this computing node as an abnormal computing node, and schedule the unexecuted computing tasks in this computing node to the remaining computing nodes for execution. Among them, the specific values of the preset utilization rate threshold and the preset memory threshold can be custom-set according to the actual situation, and the embodiments of this application do not limit this.
[0057] Step S25: Use the preset weight coefficient to perform weighted calculation on the CPU utilization rate, current remaining memory, and current length corresponding to each computing node to obtain the node weight of the computing node.
[0058] In this embodiment, after the monitor obtains the CPU utilization rate, current remaining memory, and current length of the task queue corresponding to each computing node, it uses the preset weight coefficient to perform weighted calculation on the CPU utilization rate, current remaining memory, and current length corresponding to each computing node to obtain the node weight of the corresponding computing node.
[0059] The node weight is calculated based on a preset weight formula; among them, the preset weight formula is: ; wherein, 、 、 are all weight coefficients, representing the priorities of different indicators; the higher the node weight, the more suitable the node is for receiving computing tasks.
[0060] Figure 4 is a structural schematic diagram of a monitor disclosed in the present application, which specifically includes a data acquisition module, a data processing module, an anomaly detection module, and a communication interface module. Among them, the data acquisition module is used to deploy lightweight agents on each computing node to periodically collect local load information (such as CPU utilization rate, current remaining memory, network bandwidth, etc.) and task status (such as the current length of the task queue, task completion rate, etc.); the data processing module is used to aggregate the raw data reported by multiple nodes according to a time window (such as 10 seconds) to generate a global load view, and then calculate the node weight according to a preset weight formula; the anomaly detection module is used to detect whether there is an anomaly in the node, such as whether it is overloaded or faulty, and mark it as an abnormal computing node, and then the scheduler stops allocating computing tasks to this abnormal computing node; the communication interface module interacts with the scheduler to push the node weight and anomaly events to the scheduler.
[0061] Step S26: Calculate the average weight based on the weights of each node, and calculate the corresponding weight standard deviation using the weights of each node and the average weight, and then determine whether the weight standard deviation is greater than a preset standard deviation threshold.
[0062] In this embodiment, after obtaining the node weights of each computing node, it is necessary to calculate the average weight, and calculate the corresponding weight standard deviation using the weights of each node and the average weight. Then, it is determined whether the current preset task scheduling condition is satisfied by judging whether the weight standard deviation is greater than the preset standard deviation threshold.
[0063] Step S27: If it is greater, it is determined that the current preset task scheduling condition is satisfied, and the unexecuted target computing tasks in the first computing node are scheduled to the second computing node; otherwise, it is determined that the current preset task scheduling condition is not satisfied; the node weight of the first computing node is less than the node weight of the second computing node.
[0064] In a specific embodiment, if the weight standard deviation is greater than the preset standard deviation threshold, it is determined that the current preset task scheduling condition is satisfied, and the unexecuted target computing tasks in the first computing node are scheduled to the second computing node; the node weight of the first computing node is less than the node weight of the second computing node.
[0065] In another specific embodiment, if the weight standard deviation is not greater than the preset standard deviation threshold, it is determined that the current preset task scheduling condition is not satisfied, so there is no need to perform task scheduling.
[0066] Among them, in the actual scenario, the specific value of the preset standard deviation threshold can be customized, and this embodiment does not limit it. For example, it can be set to 25%, 30%, etc.
[0067] In addition, when it is monitored that the idle time of the task queue corresponding to a certain computing node exceeds a certain threshold, such as more than 10 seconds, task scheduling can also be performed, that is, the idle node can actively obtain the unexecuted computing tasks from the overloaded nodes.
[0068] Figure 5 It is a schematic structural diagram of a scheduler disclosed in this application, which specifically includes a task queue management module, a load balancing strategy module, a task migration module, and a communication interface module. Among them, the task queue management module is mainly used to maintain the status of all tasks to be assigned, in execution, and completed, such as task ID, computing volume, dependency relationship, etc., and can also dynamically adjust the queue order according to the task urgency (such as deadline, computing volume); the load balancing strategy module is used to allocate tasks based on the computing volume of the computing tasks and the node weights; the task scheduling module triggers scheduling according to the alarms of the monitor (such as the standard deviation of the weights exceeds the preset standard deviation threshold). It should be noted that the scheduling strategy is to only schedule the calculation of tasks that have not started, to avoid interrupting the running computing tasks, and to migrate the calculation of adjacent tasks to the same computing node to reduce communication overhead; the communication interface module interacts with the monitor on the one hand to receive node weights and alarm information, and on the other hand interacts with the computing nodes to issue task instructions, receive completion notifications, and synchronize scheduling requests.
[0069] Among them, for the more specific processing process of the above step S21, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be described herein again.
[0070] It can be seen that the present application first divides the target sparse matrix based on the non-zero element distribution information, aiming to make the difference in the number of non-zero elements between different sub-matrices within a preset range. That is, when initially dividing the sub-matrices, try to ensure that the number of non-zero elements in each sub-matrix is approximately the same, so that the computational workloads of the corresponding computational tasks do not differ too much. Thus, when initially allocating each computational task to different computing nodes, load balancing is maintained as much as possible. Further, as the calculation progresses, the load level of the nodes will also change accordingly. The present application uses a monitor to continuously monitor the status of the computing nodes to provide data support for dynamic scheduling, ensuring that the system can quickly respond to load changes. The scheduler first determines whether the preset task scheduling condition is met by obtaining the node weights of each computing node calculated by the monitor. If it is met, the unexecuted target computational tasks in the computing node with a lower node weight are scheduled to the computing node with a higher node weight, so as to avoid the load imbalance problem in sparse matrix solution and improve the sparse matrix solution speed and the performance of circuit simulation at the same time.
[0071] See Figure 6 As shown, an embodiment of the present application discloses a task scheduling device in circuit simulation. The device includes: A matrix division module 11, configured to obtain a target sparse matrix generated based on the connection relationship between circuit elements in a target circuit netlist, and divide the target sparse matrix into multiple sub-matrices; wherein, the target circuit netlist is the circuit netlist used in circuit simulation, and each sub-matrix corresponds to a computational task.
[0072] A task allocation module 12, configured to allocate each computational task to a preset computing node cluster, so as to use each computing node in the computing node cluster to execute the corresponding computational task.
[0073] A weight calculation module 13, configured to periodically monitor the load information and task execution status of each computing node by using a preset monitor, and calculate the node weight of each computing node based on the load information and task execution status.
[0074] A task scheduling module 14, configured to determine whether the preset task scheduling condition is met based on each node weight. If it is met, the unexecuted target computational task in the first computing node is scheduled to the second computing node; the node weight of the first computing node is less than the node weight of the second computing node.
[0075] It can be seen that in this application, first, a target sparse matrix generated based on the connection relationships between circuit elements in the target circuit netlist is obtained, where the target circuit netlist is the circuit netlist used in circuit simulation; then the target sparse matrix is partitioned to obtain a plurality of sub-matrices, and each sub-matrix corresponds to a computing task. Further, each computing task is first assigned to a preset computing node cluster, so that each computing node in the computing node cluster starts to execute the corresponding computing task. After that, this application uses a preset monitor to periodically monitor the load information and task execution status of each computing node, so as to dynamically monitor the load and task execution conditions of each computing node, and calculate the node weights of each computing node based on this. In this application, it is stipulated that the higher the node weight, the more suitable it is to process computing tasks. Finally, it is determined whether the preset task scheduling condition is met based on each node weight. If it is met, the unexecuted target computing task in the first computing node is scheduled to the second computing node, where the node weight of the first computing node is less than the node weight of the second computing node. That is, when the preset task scheduling condition is met currently, the unexecuted target computing task in the computing node with a lower node weight is scheduled to the second computing node with a higher node weight, avoiding the situation where some computing nodes have high time consumption and load due to undertaking the computing tasks corresponding to dense sub-matrices. By scheduling and migrating the unexecuted tasks of this node to a relatively idle computing node, the load imbalance problem in sparse matrix solution can be avoided, and at the same time, the computing efficiency is improved.
[0076] Since the embodiments of the device part correspond to the above embodiments, the embodiments of the device part are described with reference to the embodiments of the above method part and will not be elaborated here.
[0077] Figure 7 It is a schematic structural diagram of an electronic device provided by an embodiment of this application. Specifically, it may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the task scheduling method in circuit simulation executed by the electronic device disclosed in any of the foregoing embodiments.
[0078] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of this application, and no specific limitation is made here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.
[0079] Among them, the processor 21 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 21 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for the rendering and drawing of the content to be displayed on the display screen. In some embodiments, the processor 21 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0080] In addition, the memory 22, as a carrier for resource storage, may be a read-only memory, a random access memory, a disk, or an optical disc, etc. The resources stored thereon include an operating system 221, a computer program 222, data 223, etc., and the storage method may be temporary storage or permanent storage.
[0081] Among them, the operating system 221 is used to manage and control each hardware device and the computer program 222 on the electronic device 20 to enable the processor 21 to perform operations and processing on the massive data 223 in the memory 22. It may be Windows, Unix, Linux, etc. In addition to the computer program that can be used to complete the task scheduling method in the circuit simulation executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include a computer program that can be used to complete other specific tasks. The data 223 may include not only the data transmitted by external devices received by the electronic device, but also the data collected by its own input / output interface 25, etc.
[0082] Furthermore, the embodiments of the present application also disclose a computer-readable storage medium. When the computer program stored in the storage medium is loaded and executed by a processor, the steps of the task scheduling method in the circuit simulation disclosed in any of the foregoing embodiments are implemented.
[0083] An embodiment of the present invention also discloses a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the task scheduling method in the circuit simulation disclosed in any of the foregoing embodiments are implemented.
[0084] In this specification, the various embodiments are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method part.
[0085] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this document can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0086] The steps of the method or algorithm described in combination with the embodiments disclosed in this document can be directly implemented by hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, compact disc read-only memory (CD-ROM), or any other form of storage medium well-known in the technical field.
[0087] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0088] The above has introduced in detail a task scheduling method, device, medium and product in circuit simulation provided by the present invention. Specific examples are used in this text to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A task scheduling method in circuit simulation, characterized in that Including: Obtain a target sparse matrix generated based on the connection relationship between circuit elements in a target circuit netlist, and divide the target sparse matrix into multiple sub-matrices; wherein, the target circuit netlist is the circuit netlist used in circuit simulation, and each sub-matrix corresponds to a calculation task; Allocate each of the calculation tasks to a preset calculation node cluster, so as to use each calculation node in the calculation node cluster to execute the corresponding calculation task; Periodically monitor the load information and task execution status of each calculation node by using a preset monitor, and calculate the node weight of each calculation node based on the load information and the task execution status; Judge whether the preset task scheduling condition is satisfied currently based on each node weight. If it is satisfied, schedule the unexecuted target calculation task in the first calculation node to the second calculation node; the node weight of the first calculation node is less than the node weight of the second calculation node.
2. The task scheduling method in circuit simulation according to claim 1, wherein Before dividing the target sparse matrix into multiple sub-matrices, it further includes: Determine a target sorting method according to the matrix characteristics of the target sparse matrix; wherein, the target sorting method is any one of the minimum degree sorting method, the spectral sorting method, and the graph partitioning method; Use the target sorting method to rearrange the rows and columns in the target sparse matrix to obtain the rearranged target sparse matrix.
3. The task scheduling method in circuit simulation according to claim 1, wherein The dividing the target sparse matrix into multiple sub-matrices includes: Count the number of non-zero elements in each row and each column of the target sparse matrix to obtain non-zero element distribution information; Divide the target sparse matrix into multiple sub-matrices based on the non-zero element distribution information; wherein, the difference in the number of non-zero elements between different sub-matrices is within a preset range.
4. The task scheduling method in circuit simulation according to claim 1, characterized in that, The allocating each of the calculation tasks to a preset calculation node cluster includes: Determine a preset calculation node cluster and determine the number of calculation nodes in the calculation node cluster; Based on the number of each calculation task and the number of calculation nodes, evenly allocate each calculation task to the task queues corresponding to each calculation node in the calculation node cluster.
5. The task scheduling method in circuit simulation according to claim 4, characterized in that, The load information includes CPU utilization rate and current remaining memory, and the task execution status is the current length of the task queue; Correspondingly, the calculating the node weight of each calculation node based on the load information and the task execution status includes: Use a preset weight coefficient to perform weighted calculation on the CPU utilization rate, current remaining memory, and current length corresponding to each calculation node to obtain the node weight of the calculation node.
6. The task scheduling method in circuit simulation according to claim 5, wherein The node weight is calculated based on a preset weight formula; wherein, the preset weight formula is: ; Among them, , , are all weighting coefficients.
7. The task scheduling method in circuit simulation according to claim 5, wherein After periodically monitoring the load information and task execution status of each calculation node by using a preset monitor, it further includes: If the CPU utilization rate of any calculation node is greater than a preset utilization threshold in a continuous preset number of cycles, or if the current remaining memory of any calculation node is lower than a preset memory threshold, mark the any calculation node as an abnormal calculation node, and schedule the unexecuted calculation tasks in the any calculation node to the remaining calculation nodes.
8. The task scheduling method in circuit simulation according to claim 1, wherein Judging whether the current situation meets the preset task scheduling conditions based on the weights of the nodes includes: Calculating the average weight based on the weights of the nodes, and calculating the corresponding standard deviation of weights by using the weights of the nodes and the average weight; Judging whether the standard deviation of weights is greater than a preset standard deviation threshold; If it is greater, it is determined that the current situation meets the preset task scheduling conditions, otherwise it is determined that the current situation does not meet the preset task scheduling conditions.
9. The task scheduling method in circuit simulation according to claim 1, wherein, Scheduling the unexecuted target calculation tasks in the first calculation node to the second calculation node includes: Determining a first calculation node with a node weight lower than a first preset threshold, and determining a second calculation node with a node weight higher than a second preset threshold; wherein, the first preset threshold is less than the second preset threshold; Determining the unexecuted target calculation tasks according to the task execution status corresponding to the first calculation node, and scheduling the target calculation tasks from the first calculation node to the second calculation node.
10. The task scheduling method in circuit simulation according to claim 1, wherein Scheduling the unexecuted target calculation tasks in the first calculation node to the second calculation node includes: Sorting the calculation nodes in descending order of node weight to obtain a sorting result; Determining several calculation nodes at a preset tail position in the sorting result as the first calculation node, and determining several calculation nodes at a preset head position in the sorting result as the second calculation node; Determining the unexecuted target calculation tasks according to the task execution status corresponding to the first calculation node, and scheduling the target calculation tasks from the first calculation node to the second calculation node.
11. The task scheduling method in circuit simulation according to claim 1, wherein It further includes: Monitoring whether all the calculation tasks have been executed; If all have been executed, obtaining the task execution results in each calculation node, and summarizing the task execution results to obtain the solution result of the target sparse matrix.
12. The task scheduling method in circuit simulation according to any one of claims 1 to 11, characterized in that, During the process of scheduling the unexecuted target calculation tasks in the first calculation node to the second calculation node, it further includes: If the number of the second calculation nodes is multiple, for the unexecuted target calculation tasks in the first calculation node, judging whether there are calculation tasks adjacent to the target calculation tasks among the calculation tasks corresponding to each second calculation node based on the target position relationship; wherein, the target position relationship is the position relationship of the sub-matrices corresponding to the calculation tasks in the target sparse matrix; If there are, scheduling the target calculation tasks to the second calculation node where the calculation task adjacent to the target calculation task is located.
13. An electronic device, characterized in that, It includes: A memory for storing a computer program; A processor for executing the computer program to implement the steps of the task scheduling method in the circuit simulation according to any one of claims 1 to 12.
14. A computer-readable storage medium, characterized in that, For storing a computer program; wherein, when the computer program is executed by the processor, the steps of the task scheduling method in the circuit simulation according to any one of claims 1 to 12 are implemented.
15. A computer program product comprising computer programs / instructions, characterized in that, When the computer program / instructions are executed by the processor, the steps of the task scheduling method in the circuit simulation according to any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
Deleted graph-based parallel decomposition method for circuit sparse matrix in circuit simulation
CN102156777A
Adaptive parallel LU decomposition method aiming at circuit simulation
CN102426619A
Circuit simulation method and device, electronic equipment and computer readable storage medium
CN116070584A
Calculation task scheduling method, electronic equipment, storage medium and product
CN119829898A
Iterative Solution Using Compressed Inductive Matrix for Efficient Simulation of Very-Large Scale Circuits
US20160328508A1