Memory pressure regulation method and device for chip design simulation task

CN122221779BActive Publication Date: 2026-09-29ZHEJIANG QUSU TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610697969.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-20
Publication Date
2026-09-29
Estimated Expiration
2046-05-20

AI Technical Summary

Technical Problem

该方案可在云系统环境下实现任务的跨节点迁移恢复,但其存在以下不足:必须依赖集群中存在一台内存更大的空闲备用节点作为第二计算资源;跨节点传输会话信息产生网络输入输出开销,延长任务恢复时间;在回归测试高峰期,集群中往往不存在完全空闲的备用节点,导致迁移策略失效

Benefits of technology

[0020]本申请提供的芯片设计仿真任务的内存压力调控方法的有益效果包括:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122221779B_ABST
    Figure CN122221779B_ABST
Patent Text Reader

Abstract

The application provides a memory pressure regulation method and device for chip design simulation tasks, and relates to the field of electric digital data processing.The method comprises the following steps: monitoring memory usage state data in a single server node; selecting a target simulation task when the growth slope indicates that the node is about to trigger memory exchange; if the task is in the simulation stage, calling a simulation tool interface to generate a live image file and store it locally, and actively terminating the process to recover the memory; after the remaining resources recover, reading the image file and resubmitting the task in the original node to make it resume execution from the corresponding simulation progress point. The application is used to solve the problems that the existing scheme cannot completely release the memory, still has the risk of memory exchange, depends on the idle standby node for cross-node migration, and has network transmission overhead. The application can completely release the memory in a single node and directly avoid memory exchange, significantly improves the simulation resource utilization efficiency, and does not need to depend on the standby node to ensure the smooth running of high-priority tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of electronic digital data processing technology, and in particular to a method and device for controlling memory pressure in chip design simulation tasks. Background Technology

[0002] Large-scale logic simulations during the chip design verification phase, especially memory-intensive tasks such as system-on-a-chip (SoC) full-chip simulation, require the simulation cluster to run hundreds of simulation tasks simultaneously during peak regression testing periods. The memory footprint of a single task can reach tens of gigabytes (GB). Assuming sufficient server resources, a single server can execute multiple simulation tasks concurrently. However, when the total memory demand of concurrent tasks approaches the server's physical memory limit, the operating system triggers a memory swapping mechanism, writing some simulation task memory pages to disk. Since disk I / O speeds are much slower than memory access speeds, any swapped-out memory page accessed in subsequent clock cycles will cause a sharp drop in the running speed of the corresponding simulation task, with performance degradation reaching over 90%. Simultaneously, the operating system's Out-of-Memory (OOM) killer may forcibly terminate some simulation tasks to ensure the continued operation of the remaining tasks, resulting in the waste of simulation computations that took several days and the valuable Electronic Design Automation (EDA) license resources.

[0003] In existing technologies, a static resource pre-allocation strategy is typically employed: sufficient resources are pre-locked for a task before it is submitted to the simulation cluster server, and the submission of a new task is rejected when the remaining resources of the server are insufficient to meet the requirements of the new task. This strategy can ensure the resource security of submitted tasks when resources are plentiful. However, the resource requirements of chip design simulation tasks fluctuate dynamically during the analysis, elaboration, and simulation phases. If resources are reserved based on the peak demand, it will lead to a high server resource vacancy rate. If the reserved resources are insufficient to cover the peak resource demand at a specific moment, it may still induce the aforementioned problems of memory swapping slowdown or forced task termination.

[0004] To address memory pressure, some job scheduling systems (such as IBM's Load Sharing Facilities (LSF)) provide a load-stop mechanism for tasks on the same node based on a load threshold. When the available memory on a compute node falls below a preset threshold, the system can suspend some tasks at the operating system level, allowing them to resume execution once resources are restored. However, this mechanism has the following limitations: the suspension operation only pauses the process; the memory pages of the suspended tasks remain resident in physical memory, failing to truly release memory resources. Under sustained memory pressure, these pages may still be swapped out to the hard drive by the operating system, failing to fundamentally prevent simulation performance degradation. Furthermore, the suspension strategy is based only on a single dimension such as job priority and submission time, without optimizing for the characteristics of chip design simulation tasks (such as differences in support for saving data at different simulation stages, and the correlation between runtime and natural termination).

[0005] To address memory pressure issues, another cross-node migration solution, such as a method for executing computing tasks on cloud systems, involves saving the session information of the computing task and pausing it when the memory usage of the first computing resource reaches a threshold. Then, a second computing resource with more memory than the first resource is identified from among multiple computing resources, and the task is resumed on the second computing resource based on the session information. This solution can achieve cross-node migration and recovery of tasks in a cloud system environment, but it has the following drawbacks: it must rely on the existence of a spare node with more memory in the cluster as the second computing resource; cross-node transmission of session information incurs network input / output overhead, prolonging task recovery time; and during peak regression testing periods, there are often no completely idle spare nodes in the cluster, causing the migration strategy to fail.

[0006] In other words, existing same-node operating system (OS) level suspension solutions cannot completely release memory resources and still carry the risk of memory swapping; cross-node migration solutions rely on idle backup nodes and incur network transmission overhead. Therefore, there is an urgent need for a memory pressure control solution that can completely reclaim memory resources within a single server node, fundamentally avoid memory swapping, and does not require backup nodes. Summary of the Invention

[0007] In view of this, embodiments of this application provide a method and apparatus for regulating memory pressure in chip design simulation tasks, so as to eliminate or improve one or more defects existing in the prior art.

[0008] One aspect of this application provides a method for controlling memory pressure in chip design simulation tasks, including: Within a server node of a chip design simulation cluster that has not triggered memory swapping, continuously monitor the memory usage status data of that server node; If the growth rate of the current memory usage data exceeds a preset slope threshold indicating that the server node is about to trigger memory swapping, then at least one of the chip design simulation tasks running on the server node is selected as the current target simulation task. If the target simulation task is currently in the simulation phase, the function interface of the simulation tool corresponding to the target simulation task is called to generate a field image file for resuming execution from the current simulation progress point, and the field image file is stored in the local or directly connected storage path of the server node. After confirming that the field image file has been generated, the process of the target simulation task is actively terminated to reclaim the memory resources occupied by the target simulation task. If the current memory usage data indicates that the remaining available resources have recovered to the preset recovery threshold, then the field image file stored locally or in the direct-connected storage path is read, and the target simulation task corresponding to the field image file is resubmitted on the server node, so that the target simulation task resumes execution of chip design simulation from the simulation progress point corresponding to the field image file.

[0009] In some embodiments of this application, the memory usage status data includes: the memory occupancy rate of the server node monitored within a preset monitoring time window.

[0010] In some embodiments of this application, selecting at least one of the chip design simulation tasks running on the server node as the current target simulation task includes: Obtain the priority, runtime, and current memory usage of each chip design simulation task running on the server node; Based on the priority, runtime, and current memory usage, calculate the suspension score for each chip design simulation task. According to the order of the suspension scores from high to low, select at least one chip design simulation task as the current target simulation task. The expression for the suspended score is shown in Formula (I): (one) In formula (1), The suspended score is given. The value represents the priority, and a larger value indicates a higher priority. The duration of operation. The current memory usage, It is an exponentially decaying function that monotonically decreases with the duration R of operation. , and These are the weighting coefficients; The exponential decay function The expression is shown in formula (II): (two) in, This is a preset upper limit for the runtime reference. The expected total simulation time for the chip design simulation task is estimated, where k is the attenuation steepness coefficient.

[0011] In some embodiments of this application, selecting at least one chip design simulation task from those running on the server node as the current target simulation task further includes: Before obtaining the priority, runtime, and current memory usage of each chip design simulation task running on the server node, first count the number of chip design simulation tasks currently running on the server node for each user. The users are sorted in descending order of the number of chip design simulation tasks they are running, and the top N users are selected as candidate users, where N is a positive integer. Correspondingly, the priority, runtime, and current memory usage of each chip design simulation task running on the server node are obtained, including: Obtain the priority, runtime, and current memory usage of the chip design simulation task currently running on the server node for each of the candidate users.

[0012] In some embodiments of this application, before reading the site image file stored locally or in a direct-attached storage path, the method further includes: If the current memory usage data indicates that the remaining available resources have recovered to the preset recovery threshold, then the resource requirements declared by the target simulation task corresponding to the field image file at the time of the original submission are obtained from the preset suspended task record. Compare the resource requirements with the currently available resources of the server node; If the current available resources of the server node meet the resource requirements, then it is determined to read the site image file.

[0013] In some embodiments of this application, reading the site image file stored locally or in a direct-attached storage path includes: Following a last-in-first-out (LIFO) order, the field image files stored locally or in a directly connected storage path are read one by one; wherein, LIFO means that the chip design simulation task that was most recently terminated is resubmitted and resumed execution first.

[0014] In some embodiments of this application, the active termination of the target simulation task process includes: Invoke the job termination command of the job scheduling system running on the server node to forcibly terminate the process of the target simulation task.

[0015] In some embodiments of this application, after the process of actively terminating the target simulation task is completed, the method further includes: Send a notification message to the user to whom the terminated target simulation task belongs; wherein the notification message is used to inform the user that the target simulation task has been terminated, the reason for the termination, and the storage path of the field image file; If a migration instruction for a terminated target simulation task is received from the user, the record corresponding to the target simulation task is deleted from the preset suspended task record to avoid resuming the execution of the target simulation task on the server node.

[0016] Another aspect of this application provides a memory pressure control device for chip design simulation tasks, which can be deployed as an independent middleware program on each simulation cluster server node. This device achieves closed-loop control of memory pressure for chip design simulation tasks within each server node by calling the command-line interface of the job scheduling system and the functional interfaces of the simulation tools. The memory pressure control device for chip design simulation tasks includes: The memory pressure monitoring module is used to continuously monitor the memory usage status data of a server node in a chip design simulation cluster that has not triggered memory swapping. The execution suspension module is configured to select at least one chip design simulation task running on the server node as the current target simulation task if the growth slope of the current memory usage status data exceeds a preset slope threshold indicating that the server node is about to trigger memory swapping; if the target simulation task is currently in the simulation stage, the module calls the function interface of the simulation tool corresponding to the target simulation task to generate a field image file for resuming execution from the current simulation progress point, and stores the field image file in the local or directly connected storage path of the server node; after confirming that the field image file has been generated, the module actively terminates the process of the target simulation task to reclaim the memory resources occupied by the target simulation task; The task recovery module is used to read the field image file stored locally or in a direct-connected storage path if the current memory usage status data indicates that the remaining available resources have recovered to a preset recovery threshold, and resubmit the target simulation task corresponding to the field image file on the server node so that the target simulation task resumes execution of chip design simulation from the simulation progress point corresponding to the field image file.

[0017] A third aspect of this application provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a memory pressure control method for the chip design simulation task.

[0018] A fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the memory pressure control method for the chip design simulation task.

[0019] The fifth aspect of this application provides a computer program product comprising a computer program that, when executed by a processor, implements the memory pressure control method for the chip design simulation task.

[0020] The beneficial effects of the memory pressure control method for chip design simulation tasks provided in this application include: (1) By continuously monitoring the growth slope of memory usage status data within a single server node of the chip design simulation cluster, the rising trend of memory pressure can be predicted before memory swapping is triggered by the server node, thereby proactively intervening before memory swapping actually occurs.

[0021] (2) For the selected target simulation task, if it is in the simulation stage, the simulation tool's function interface is called to generate a field image file and store it locally or in a directly connected storage path. Then, the process of the target simulation task is actively terminated to completely reclaim the memory resources it occupies. If it is in the analysis or refinement stage, the process is terminated directly to reclaim the memory. The above method can redistribute the reclaimed memory resources to other simulation tasks that continue to run on the server node, fundamentally avoiding the simulation performance degradation caused by the operating system swapping memory pages to the hard disk due to continuous memory resource shortage. At the same time, it can prevent the abnormal termination of simulation tasks by the operating system's memory overflow termination mechanism.

[0022] (3) After the remaining available resources of the server node recover to the preset recovery threshold, the field image file stored locally or in the direct-connected storage path is read and the corresponding target simulation task is resubmitted on the same server node, so that it continues to execute from the simulation progress point when it was saved. This same-node in-situ recovery mechanism does not rely on the existence of idle spare nodes in the cluster and has no network overhead for transmitting field image files across nodes, which can effectively shorten the task recovery time.

[0023] By combining the above technical means, this method can achieve complete reclamation and closed-loop control of the memory resources of the controlled simulation task within a single server node, improve the resource utilization efficiency of the simulation cluster server, ensure the smooth operation of high-priority simulation tasks under memory pressure, and accelerate the release of electronic design automation licenses to improve the circulation efficiency of license resources.

[0024] Additional advantages, objectives, and features of this application will be set forth in part in the description which follows, and will in part become apparent to those skilled in the art upon review of the following description, or may be learned by practice of the application. The objectives and other advantages of this application can be realized and obtained by means of the structures specifically pointed out in the specification and drawings.

[0025] Those skilled in the art will understand that the purposes and advantages that can be achieved with this application are not limited to those specifically described above, and that the above and other purposes that this application can achieve will be more clearly understood from the following detailed description. Attached Figure Description

[0026] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, do not constitute a limitation thereof. The components in the drawings are not drawn to scale but are merely for illustrating the principles of this application. For ease of illustration and description of certain parts of this application, corresponding portions in the drawings may be enlarged, i.e., may appear larger relative to other components in an exemplary device actually manufactured according to this application. In the drawings: Figure 1 This is a schematic diagram of the first process of a memory pressure control method for chip design simulation tasks in one embodiment of this application.

[0027] Figure 2 This is a second flowchart illustrating a memory pressure control method for chip design simulation tasks in one embodiment of this application.

[0028] Figure 3 This is a schematic diagram of the third method for controlling memory pressure in a chip design simulation task according to an embodiment of this application.

[0029] Figure 4 This is a schematic diagram of the memory pressure control device for chip design simulation tasks in one embodiment of this application. Detailed Implementation

[0030] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the embodiments and accompanying drawings. Here, the illustrative embodiments and their descriptions are used to explain this application, but are not intended to limit it.

[0031] It should also be noted that, in order to avoid obscuring this application with unnecessary details, only the structures and / or processing steps closely related to the solution according to this application are shown in the accompanying drawings, while other details that are not closely related to this application are omitted.

[0032] It should be emphasized that the term "including / comprises" as used herein refers to the presence of a feature, element, step, or component, but does not exclude the presence or addition of one or more other features, elements, steps, or components.

[0033] It should also be noted that, unless otherwise specified, the term "connection" in this article can refer not only to a direct connection, but also to an indirect connection involving an intermediary.

[0034] In the following description, embodiments of the present application will be illustrated with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar parts, or the same or similar steps.

[0035] To address the issues of existing same-node operating system-level suspension schemes failing to completely release memory and still posing memory swapping risks, and cross-node migration schemes relying on idle backup nodes and incurring network transmission overhead, this application provides a memory pressure control method for chip design simulation tasks, a memory pressure control device for chip design simulation tasks, an electronic device, a computer-readable storage medium, and a computer program product for executing the memory pressure control method for chip design simulation tasks. By predicting pressure based on the memory growth slope before memory swapping is triggered on a node, actively selecting tasks in the simulation stage to save the current image and terminate the process to completely reclaim memory, and resuming execution in the same location on the same node after resources are restored, complete memory release can be achieved within a single node, directly avoiding memory swapping. This significantly improves the efficiency of simulation resource utilization and ensures the smooth operation of high-priority tasks without relying on backup nodes.

[0036] The following examples will provide a detailed description.

[0037] Based on this, embodiments of this application provide a method for regulating the memory pressure of a chip design simulation task, which can be implemented by a memory pressure regulation device for the chip design simulation task. This memory pressure regulation device can be deployed as an independent middleware program on each simulation cluster server node. This device achieves closed-loop regulation of the memory pressure of the chip design simulation task within each server node by calling the command-line interface of the job scheduling system and the functional interfaces of the simulation tools. See also... Figure 1 The memory pressure control method for the chip design simulation task specifically includes the following: Step 100: Continuously monitor the memory usage status data of a server node in the chip design simulation cluster that has not triggered memory swapping.

[0038] In this system, the control device continuously monitors the memory usage status data of a server node within a chip design simulation cluster that is not experiencing memory swapping. This memory usage status data refers to quantitative indicators that reflect the current memory resource occupancy of the server node, such as memory utilization rate, used memory amount, and remaining available memory amount. Taking IBM's load sharing facility job scheduling system as an example, the control device can periodically obtain dynamic load information such as the server node's memory utilization rate and CPU utilization rate through the load viewing command (lsload), serving as a source of memory usage status data.

[0039] Step 200: If the growth slope of the current memory usage status data exceeds a preset slope threshold used to indicate that the server node is about to trigger memory swapping, then select at least one of the chip design simulation tasks running on the server node as the current target simulation task; if the target simulation task is currently in the simulation stage, then call the function interface of the simulation tool corresponding to the target simulation task to generate a field image file for resuming execution from the current simulation progress point, and store the field image file in the local or directly connected storage path of the server node; after confirming that the field image file has been generated, actively terminate the process of the target simulation task to reclaim the memory resources occupied by the target simulation task.

[0040] It should be noted that if only one chip design simulation task is selected, it will be used as the target simulation task; if multiple chip design simulation tasks are selected, each selected chip design simulation task will be used as a target simulation task, and subsequent operations such as generating and storing field image files and active termination will be performed for each target simulation task.

[0041] Specifically, the control device analyzes the continuously collected memory usage data and calculates its growth slope. The growth slope refers to the magnitude of change in memory usage data per unit time, used to characterize the rate of increase in memory pressure. When the growth slope of the current memory usage data exceeds a preset slope threshold, it indicates that the server node is about to trigger memory swapping, meaning the memory pressure is rapidly approaching the critical point that triggers operating system memory swapping. At this time, the control device selects at least one chip design simulation task running on the server node as the current target simulation task. The chip design simulation task refers to a simulation process used to verify chip logic design, such as a full-chip simulation task for a system-on-a-chip (SoC).

[0042] For the selected target simulation task, the control device determines its current simulation stage. Chip design simulation tasks typically go through the analysis stage, refinement stage, and simulation stage sequentially. If the target simulation task is currently in the simulation stage, it indicates that the simulation engine has complete internal state serialization capabilities, and a state saving operation can be performed. The control device calls the function interface of the simulation tool corresponding to the target simulation task, such as the save function of the Synopsys VCS simulation tool, to generate a state image file for resuming execution from the current simulation progress point. The state image file contains all current register states, memory states, combinational logic signal values, and the internal state of the simulation engine for the target simulation task. The control device stores the state image file in the local storage of the server node or in a network storage path directly connected to the server node to ensure fast retrieval during subsequent recovery.

[0043] After confirming the successful generation of the site image file, the control device actively terminates the process of the target simulation task. Taking the IBM LSF job scheduling system as an example, the process of the target simulation task can be forcibly terminated via a job termination command. After the process terminates, the memory resources occupied by the target simulation task are immediately reclaimed by the operating system, thereby freeing up memory space that can be allocated to other running simulation tasks.

[0044] Understandably, if the target simulation task is currently in the analysis or refinement stage, it indicates that the simulation engine does not yet possess complete internal state serialization capabilities and cannot perform on-site saving operations. In this case, the control device skips the step of calling the simulation tool's function interface to generate an on-site image file and directly and proactively terminates the process of the target simulation task to reclaim the memory resources it occupies.

[0045] Step 300: If the current memory usage status data indicates that the remaining available resources have recovered to the preset recovery threshold, then read the field image file stored locally or in the direct-connected storage path, and resubmit the target simulation task corresponding to the field image file on the server node, so that the target simulation task resumes execution of chip design simulation from the simulation progress point corresponding to the field image file.

[0046] Specifically, the control device continuously monitors the memory usage status data of the server node. When the monitored memory usage status data indicates that the remaining available resources have recovered to the preset recovery threshold, it indicates that the memory pressure on the server node has been effectively alleviated, and the conditions for resuming the execution of the terminated task are met. At this time, the control device reads the field image file previously stored locally or in a direct-attached storage path, and resubmits the target simulation task corresponding to the field image file on the same server node. Taking the IBM LSF job scheduling system as an example, the target simulation task can be resubmitted through the job submission command (bsub). After resubmission, the target simulation task continues to execute the chip design simulation from the simulation progress point corresponding to the field image file, thereby achieving closed-loop recovery of the controlled task.

[0047] Through the above steps, this embodiment can predict memory pressure and actively intervene before memory swapping actually occurs within a single server node, completely reclaim the memory resources occupied by the controlled task, and resume task execution in place after the resources are restored, effectively avoiding simulation performance degradation and abnormal task termination caused by memory pressure.

[0048] Understandably, for a target simulation task that is terminated while in the analysis or refinement stage, since the on-site image file has not been saved, after the control device resubmits the target simulation task on the same server node, the target simulation task will start executing the chip design simulation from scratch, rather than resuming execution from a certain simulation progress point.

[0049] In addition, the following lifecycle management can be performed on the on-site image file: When a target simulation task is terminated, the control device establishes and stores a mapping relationship between the original job ID of the terminated chip design simulation task, the new job ID obtained after subsequent resubmission, and the storage path of the field image file. This mapping relationship can be recorded in a preset suspended task record or stored in a separate image management table.

[0050] When a resubmitted chip design simulation task completes normally, or is terminated again by the control device due to memory pressure, the control device performs corresponding retention or deletion operations on the field image file according to a preset cleanup strategy. For example, for a task that completes normally, the control device can retain the field image file for a certain period of time (e.g., 7 days) after its completion for debugging and traceability, and delete it after the expiration period; for a task that is terminated again, the control device can overwrite the old file with a newly generated field image file, or retain multiple historical versions for backtracking.

[0051] By establishing the above mapping relationship and implementing the cleanup strategy, the control device can perform closed-loop management of the field image files generated during the control process, avoiding expired or invalid field image files from occupying storage space for a long time.

[0052] As described above, the memory pressure control method for chip design simulation tasks provided in this application continuously monitors the growth slope of memory usage data within a single server node of the chip design simulation cluster. This allows for the prediction of rising memory pressure trends before memory swapping is triggered on the server node, enabling proactive intervention before actual memory swapping occurs. For selected target simulation tasks, if they are in the simulation phase, the simulation tool's functional interface is called to generate a live image file and store it locally or in a directly connected storage path. Subsequently, the process of the target simulation task is actively terminated to completely reclaim its occupied memory resources. If it is in the analysis or refinement phase, the process is directly terminated to reclaim memory. This method can redistribute the reclaimed memory resources to other simulation tasks continuing to run on the server node, fundamentally avoiding simulation performance degradation caused by the operating system swapping memory pages to the hard disk due to continuous memory resource shortages. It also prevents abnormal termination of simulation tasks by the operating system's memory overflow termination mechanism. Once the remaining available resources of the server node recover to the preset recovery threshold, the on-site image file stored locally or in a directly connected storage path is read, and the corresponding target simulation task is resubmitted on the same server node, allowing it to continue execution from the simulation progress point at the time of saving. This on-site recovery mechanism does not rely on idle spare nodes in the cluster and avoids the network overhead of transmitting on-site image files across nodes, effectively shortening task recovery time. Combining the above technical means, this method can achieve complete reclamation and closed-loop control of the memory resources of the controlled simulation task within a single server node, improving the resource utilization efficiency of the simulation cluster server, ensuring the smooth operation of high-priority simulation tasks under memory pressure, and accelerating the release of electronic design automation licenses to improve the circulation efficiency of high-value license resources.

[0053] To further and more reliably intervene before memory swapping actually occurs, in a memory pressure control method for chip design simulation tasks provided in this application embodiment, the memory usage status data of the memory pressure control method for chip design simulation tasks includes: the memory occupancy rate of the server node monitored within a preset monitoring time window.

[0054] The preset monitoring time window refers to the observation period set by the control device for calculating the growth slope of memory usage data, such as a continuous 5 minutes or 10 minutes. Within this preset monitoring time window, the control device periodically collects the memory occupancy rate of the server node and calculates the growth slope within that period based on the difference in memory occupancy rate between the first and last sampling times. Memory occupancy rate refers to the percentage of currently used memory in the total physical memory of the server node, which can be obtained through operating system commands or job scheduling system commands.

[0055] Taking a server node in a chip design simulation cluster as an example, the control device sets a preset monitoring time window of 5 minutes and a preset slope threshold of memory usage exceeding 20% ​​within 5 minutes. The control device samples the server node's memory usage at 40% in minute 0 and 65% in minute 5. Calculations show that the memory usage increase within this 5-minute monitoring time window is 25 percentage points (65% - 40% = 25%), exceeding the preset slope threshold (20%). At this point, the control device determines that the growth slope of the current memory usage data meets the preset condition, indicating that the server node is ready to trigger memory swapping. It then selects the target simulation task and executes subsequent processing.

[0056] As can be seen from the above description, the memory pressure control method for chip design simulation tasks provided in this application uses quantifiable memory occupancy and its change within a fixed time window as the triggering basis, making the control device more accurate and stable in identifying the trend of memory pressure increase, avoiding false triggering due to instantaneous fluctuations or noise data, and thus more reliably intervening before memory swapping actually occurs.

[0057] To further address the issue of how to scientifically select target simulation tasks in multi-task concurrent scenarios to avoid blind adjustment, this application provides a memory pressure control method for chip design simulation tasks, see [link to relevant documentation]. Figure 2 The step 200 of the memory pressure control method for the chip design simulation task, which involves selecting at least one of the chip design simulation tasks running on the server node as the first implementation of the current target simulation task, specifically includes the following: Step 210: If the growth slope of the current memory usage status data is detected to exceed the preset slope threshold used to indicate that the server node is about to trigger memory swapping, then obtain the priority, runtime and current memory usage of each chip design simulation task running on the server node.

[0058] Step 220: Calculate the suspension score for each chip design simulation task based on the priority, runtime, and current memory usage.

[0059] Step 230: Select at least one chip design simulation task as the current target simulation task according to the order of the suspension scores from high to low.

[0060] The expression for the suspended score is shown in Formula (I): (one) In formula (1), The suspended score is given. The value represents the priority, and a larger value indicates a higher priority. The duration of operation. The current memory usage, It is an exponentially decaying function that monotonically decreases with the duration R of operation. , and These are the weighting coefficients, and their sum can be equal to 1.

[0061] The exponential decay function The expression is shown in formula (II): (two) in, This is a preset upper limit for the runtime reference. The expected total simulation time for the chip design simulation task is estimated, where k is the attenuation steepness coefficient.

[0062] It should be noted that in this embodiment, priority is the scheduling priority assigned to the chip design simulation task when it is submitted to the job scheduling system. A higher value indicates a higher priority, and the system prioritizes resource allocation and execution for high-priority tasks. Runtime is the cumulative running time of the chip design simulation task from the start of execution to the current moment, in minutes. Current memory usage is the actual amount of physical memory currently occupied by the chip design simulation task, in gigabytes. The suspension score is a weighted score calculated by combining priority, runtime, and current memory usage. It is used to quantitatively evaluate the suitability of each chip design simulation task as a target simulation task; higher scores result in higher priority selection. Weighting coefficients correspond to the weights of priority, runtime, and current memory usage in the suspension score calculation and can be configured according to cluster load characteristics and control strategies. The exponential decay function is a function that monotonically decreases with runtime. Its purpose is to significantly suppress the suspension score of tasks with longer run times, thereby reducing the probability of selection. The preset upper limit for runtime is a reference threshold used to define the runtime of long-cycle tasks and short-cycle tasks, such as 10,000 minutes. The estimated total simulation time is based on historical data or user-predicted total runtime required to complete the chip design simulation task. The decay steepness coefficient is a parameter that controls the decay rate of the exponential decay function near the estimated total simulation time; the larger the value of k, the steeper the score decreases as it approaches expected completion.

[0063] As can be seen from the above description, the memory pressure control method for chip design simulation tasks provided in this application calculates the suspension score by weighting three dimensions: comprehensive priority, running time, and current memory usage, and introduces an exponential decay function to suppress the probability of tasks nearing completion being selected. This method can prioritize the release of tasks with high memory usage and minimal impact on the overall progress, thereby maximizing memory release benefits while ensuring the smooth operation of high-priority tasks.

[0064] To further address the issue of how to fairly and efficiently select target simulation tasks in multi-user shared server node scenarios, this application provides a memory pressure control method for chip design simulation tasks, see [link to relevant documentation]. Figure 3 The second implementation method of selecting at least one of the chip design simulation tasks running on the server node as the current target simulation task in step 200 of the memory pressure control method for the chip design simulation task specifically includes the following: Step 201: If the growth slope of the current memory usage status data is detected to exceed the preset slope threshold used to indicate that the server node is about to trigger memory swapping, then before obtaining the priority, runtime and current memory usage of each chip design simulation task running on the server node, first count the number of chip design simulation tasks currently running on the server node for each user. Step 202: Sort the users in descending order of the number of chip design simulation tasks they have run, and select the top N users from the sorted users as candidate users, where N is a positive integer. Step 203: Obtain the priority, runtime, and current memory usage of the chip design simulation task currently running on the server node for each candidate user.

[0065] Step 220: Calculate the suspension score for each chip design simulation task based on the priority, runtime, and current memory usage.

[0066] Step 230: Select at least one chip design simulation task as the current target simulation task according to the order of the suspension scores from high to low.

[0067] The candidate users are the top N users selected after counting the number of tasks currently running for each user and sorting them from most to least. Only the chip design simulation tasks under the candidate user's name will proceed to the subsequent suspension score calculation stage; tasks from other users will not be included in this selection. N is a preset positive integer used to limit the number of candidate users. The value of N can be configured according to the scale of concurrent users on the server node and the control strategy. When N is 1, only the user with the highest score (i.e., the user running the most tasks) is selected as a candidate user; when N is larger, the candidate range can be expanded to take into account the tasks of more users.

[0068] As can be seen from the above description, the memory pressure control method for chip design simulation tasks provided in this application can avoid affecting the normal simulation progress of other users due to a single user occupying too many resources by prioritizing the selection of candidate ranges from users with a large number of running tasks. At the same time, it can prioritize releasing the memory resources of batch task users to alleviate the overall memory pressure more quickly and achieve load balancing control in a multi-user environment.

[0069] To further address the issue of unclear implementation methods for terminating the target simulation task process, a memory pressure control method for chip design simulation tasks is provided in an embodiment of this application, see [link to relevant documentation]. Figure 2 or Figure 3 The memory pressure control method for the chip design simulation task, after step 230 in step 200, further includes the following: Step 240: If the target simulation task is currently in the simulation stage, call the function interface of the simulation tool corresponding to the target simulation task to generate a field image file for resuming execution from the current simulation progress point, and store the field image file in the local or directly connected storage path of the server node.

[0070] Step 250: After confirming that the field image file has been generated, call the job termination command of the job scheduling system running on the server node to forcibly terminate the process of the target simulation task, so as to reclaim the memory resources occupied by the target simulation task.

[0071] The job scheduling system is system software deployed on server nodes to manage the submission, scheduling, execution, and resource allocation of computing tasks, such as the IBM Load Sharing Facility job scheduling system. The job termination command is a command-line interface provided by the job scheduling system to forcibly terminate a specified job process. Taking the IBM LSF job scheduling system as an example, the job termination command is "bkill". After executing this command, the target job enters the exit state (EXIT), and the computing resources it occupies are immediately reclaimed by the operating system. After confirming that the field image file of the target simulation task has been successfully generated, the control device calls the job termination command of the LSF job scheduling system running on the server node, specifying the job identifier of the target simulation task as a parameter. Upon receiving the job termination command, LSF forcibly terminates the process of the target simulation task, and the target simulation task immediately enters the exit state (EXIT). After the process terminates, the memory resources occupied by the target simulation task are immediately reclaimed by the operating system and made available for other running tasks.

[0072] It should be noted that even if the target simulation task was enabled with the automatic restart function via the automatic restart option (-r) of the job submission command when it was originally submitted, it will not be automatically restarted by the IBM LSF job scheduling system after being forcibly terminated by the job termination command, thus ensuring that the task will not be repeatedly scheduled and executed before the resources have been restored.

[0073] As can be seen from the above description, the memory pressure control method for chip design simulation tasks provided in the embodiments of this application... Forced process termination is achieved by calling the job termination command natively of the job scheduling system. Without intruding into the job scheduling system kernel or modifying the simulation tool source code, memory resources can be completely reclaimed within a single server node. It is easy to deploy and highly versatile.

[0074] To further address the issue that users of terminated tasks are unaware of the control process and unable to make independent decisions, a memory pressure control method for chip design simulation tasks is provided in this application embodiment. (See also...) Figure 2 or Figure 3The memory pressure control method for the chip design simulation task, after step 250 in step 200, further includes the following: Step 260: Send a notification message to the user to whom the terminated target simulation task belongs; wherein the notification message is used to inform the user that the target simulation task has been terminated, the reason for the termination, and the storage path of the field image file.

[0075] Step 270: If a migration instruction for the terminated target simulation task is received from the user, the record corresponding to the target simulation task is deleted from the preset suspended task record to avoid resuming the execution of the target simulation task on the server node.

[0076] The notification message is an electronic message, such as an email or system message, automatically sent by the control device to the user of the target simulation task after it has been terminated. The notification message includes the fact that the target simulation task has been terminated, the reason for termination (e.g., server node memory pressure triggering control), and the storage path of the generated field image file. The migration instruction is a feedback command sent by the user to the control device after receiving the notification message, indicating their decision to migrate the target simulation task to another server node for resumption of execution. This instruction instructs the control device to delete the corresponding record of the target simulation task from the preset suspended task record. The preset suspended task record is a data structure used by the control device to persistently store key information about terminated target simulation tasks.

[0077] Taking a target simulation task running on a server node as an example, after the control device actively terminates the process of the target simulation task, it automatically sends a notification email to the registered email address of the user to whom the task belongs.

[0078] If the user decides to migrate the target simulation task to another server node for resumption of execution, the user sends a migration command to the control device. Upon receiving the migration command, the control device deletes the record corresponding to the target simulation task from the preset suspended task record. Thereafter, even if the remaining available resources of the original server node recover to the preset recovery threshold, the control device will not repeat the execution of the target simulation task on the original server node. The user can manually resume the execution of the target simulation task on another server node using the aforementioned field image file.

[0079] If the user does not send a migration command, the record of the target simulation task will remain in the preset suspended task record. Once the remaining available resources of the original server node recover to the preset recovery threshold, the control device will automatically resume execution according to the methods described in steps 100 to 300.

[0080] As can be seen from the above description, the memory pressure control method for chip design simulation tasks provided in this application embodiment gives users flexible scheduling rights over their own simulation tasks by actively notifying users and providing migration options. At the same time, through the linkage deletion mechanism of migration instructions and suspended task records, it avoids the control device from repeatedly restoring tasks that have been migrated by users on the original server node, thus ensuring the consistency between the control process and the user's autonomous decision-making.

[0081] To further address the issue of resource stress or memory swapping jitter immediately triggering again after task resumption due to a lack of prior assessment of resource sufficiency, a memory pressure control method for chip design simulation tasks provided in this application embodiment is described below. Figure 2 or Figure 3 Step 300 of the memory pressure control method for the chip design simulation task specifically includes the following: Step 310: If the current memory usage status data indicates that the remaining available resources have recovered to the preset recovery threshold, then obtain the resource requirements declared by the target simulation task corresponding to the field image file at the time of the original submission from the preset suspended task record.

[0082] Step 320: Compare the resource requirement with the currently available resources of the server node.

[0083] Step 330: If the current available resources of the server node meet the resource requirements, then determine to read the site image file.

[0084] The pre-set suspended task record is a persistent set of information recorded by the control device when a target simulation task is terminated, used to save key information about the terminated task. This record includes at least the job identifier of the target simulation task at the time of termination, the original submission command, the storage path of the field image file, the simulation stage at the time of termination, and the resource requirements declared at the time of original submission. The resource requirements declared at the time of original submission are the resource allocation quotas specified by the user or system for the target simulation task when it is first submitted through the job scheduling system, including but not limited to memory reservations and the number of CPU cores. Taking the IBM LSF job scheduling system as an example, the resource requirement declaration at the time of original submission can be obtained through the job history viewing command.

[0085] Taking a server node as an example, the control device continuously monitors the memory usage status data of that server node. At a certain moment, the control device detects that the current memory usage status data indicates that the remaining available resources have recovered to the preset recovery threshold (for example, the remaining available resources have recovered to more than 30% of the total resources), meeting the conditions for starting the recovery process.

[0086] The control device reads a pending task record from the preset suspended task log. The target simulation task corresponding to this record declared the following resource requirements upon initial submission: 16 gigabytes of memory reservation and 4 CPU cores. The control device obtains the above resource requirement information through the job history viewing command and compares it with the currently available resources of the server node. The query reveals that the server node currently has 24 gigabytes of available memory and 8 available CPU cores.

[0087] Upon comparison, the currently available memory (24GB) exceeds the memory reservation declared in the task (16GB), and the currently available number of CPU cores (8 cores) exceeds the number of cores declared in the task (4 cores). Therefore, the currently available resources meet the resource requirements of the target simulation task. Consequently, the control device determines to read the field image file and resubmit the target simulation task on the same server node according to the methods described in steps 100 to 300, thus resuming execution from the corresponding simulation progress point.

[0088] If the currently available resources do not meet the resource requirements of the target simulation task, for example, if the currently available memory is only 8 gigabytes, which is less than the 16 gigabytes declared in the task, the control device will not read the field image file for the time being and will keep the record in the preset suspended task record, waiting for the next round of resource checks.

[0089] As can be seen from the above description, the memory pressure control method for chip design simulation tasks provided in this application actively compares the original resource requirements of the task with the currently available resources before reading the field image file, ensuring that the recovery operation is only performed when resources are sufficient, thereby avoiding invalid submissions and repeated adjustments, and ensuring the stable operation of the task after recovery.

[0090] To further address the issue of determining the recovery order when multiple terminated tasks are restored to their resources, a memory pressure control method for chip design simulation tasks is provided in an embodiment of this application. (See also...) Figure 2 or Figure 3 The memory pressure control method for the chip design simulation task, after step 330 in step 300, further includes the following: Step 340: Read the field image files stored locally or in the direct-connected storage path one by one in the Last-In-First-Out (LIFO) order; wherein, LIFO means that the chip design simulation task that was most recently terminated is resubmitted and resumed execution first.

[0091] Step 350: The target simulation task corresponding to the field image file is resubmitted on the server node so that the target simulation task resumes execution of chip design simulation from the simulation progress point corresponding to the field image file.

[0092] Last-In-First-Out (LIFO) is a data reading order rule that means the last data stored is read out first. In this scheme, LIFO means that the field image file corresponding to the most recently terminated chip design simulation task is read and resumed first, while tasks that were terminated earlier are placed at the end of the recovery queue.

[0093] As can be seen from the above description, the memory pressure control method for chip design simulation tasks provided in this application embodiment restores terminated tasks one by one in a last-in-first-out order, which can prioritize the resumption of execution of tasks with shorter waiting times, reduce task backlog caused by control, and avoid new resource competition caused by improper restoration order.

[0094] From a software perspective, this application also provides a memory pressure control device for all or part of the chip design simulation task in the memory pressure control method for executing the aforementioned chip design simulation task, see [link to relevant documentation]. Figure 4 The memory pressure control device for the chip design simulation task specifically includes the following components: The memory pressure monitoring module 10 is used to continuously monitor the memory usage status data of a server node in a chip design simulation cluster that has not triggered memory swapping. The execution suspension module 20 is configured to, if the growth slope of the current memory usage status data exceeds a preset slope threshold indicating that the server node is about to trigger memory swapping, select at least one chip design simulation task running on the server node as the current target simulation task; if the target simulation task is currently in the simulation stage, call the function interface of the simulation tool corresponding to the target simulation task to generate a field image file for resuming execution from the current simulation progress point, and store the field image file in the local or directly connected storage path of the server node; after confirming that the field image file has been generated, actively terminate the process of the target simulation task to reclaim the memory resources occupied by the target simulation task; The task recovery module 30 is used to read the field image file stored locally or in a direct-connected storage path if the current memory usage status data indicates that the remaining available resources have recovered to a preset recovery threshold, and resubmit the target simulation task corresponding to the field image file on the server node so that the target simulation task resumes the execution of chip design simulation from the simulation progress point corresponding to the field image file.

[0095] The embodiment of the memory pressure control device for chip design simulation tasks provided in this application can be used to execute the processing flow of the embodiment of the memory pressure control method for chip design simulation tasks in the above embodiment. Its function will not be repeated here, but can be referred to the detailed description of the embodiment of the memory pressure control method for chip design simulation tasks in the above embodiment.

[0096] The memory pressure regulation device for the chip design simulation task can be deployed in either a server or a client device to regulate memory pressure during the simulation. The specific deployment can be chosen based on the processing power of the client device and the limitations of the user's usage scenario. This application does not impose any limitations on this. If all operations are performed in the client device, the client device may further include a processor for the specific processing of memory pressure regulation during the chip design simulation task.

[0097] The aforementioned client device may have a communication module (i.e., a communication unit) that can communicate with a remote server to achieve data transmission with the server. The server may include a server on the task scheduling center side; in other implementation scenarios, it may also include a server on an intermediate platform, such as a server on a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, a server cluster consisting of multiple servers, or a distributed server structure.

[0098] The server and the client device can communicate using any suitable network protocol, including those not yet developed as of the date of this application. Such network protocols may include, for example, TCP / IP, UDP / IP, HTTP, HTTPS, etc. Furthermore, such network protocols may also include RPC (Remote Procedure Call Protocol) and REST (Representational State Transfer Protocol) protocols used on top of the aforementioned protocols.

[0099] As can be seen from the above description, the memory pressure control device for chip design simulation tasks provided in this application embodiment can completely release memory within a single node and directly avoid memory swapping, which can significantly improve the efficiency of simulation resource utilization, and can ensure the smooth operation of high-priority tasks without relying on backup nodes.

[0100] To further illustrate the above embodiments, this application also provides a specific application example of a memory pressure control method for chip design simulation tasks executed by a memory pressure control device for chip design simulation tasks. In this application example, the control device deployed on each server node in the chip design simulation cluster continuously monitors the memory usage status data of the server node. This is particularly suitable for dynamic simulation steps in the chip design verification stage where the memory resource requirements of a single server are relatively large. Taking the IBM Load Sharing Facility (LSF) job scheduling system as an example, the total available resources of a specified server node can be queried using the lshosts command (used to display static resource information of each server in the LSF cluster, including configuration parameters such as hostname, host type, host model, maximum number of CPUs, maximum memory, and maximum swap space). The resource usage of a specified server node can be queried using the load viewing command (used to display dynamic load information of each server in the LSF cluster, including real-time status data such as current CPU utilization, memory utilization, swap space utilization, and average load over the last 1 minute / 5 minutes / 15 minutes). Before the server node triggers memory swapping, the control device predicts the increasing trend of memory pressure based on the growth slope of memory usage data. It proactively selects a subset of chip design simulation tasks as target simulation tasks, calls the simulation tool's function interface to save a live image file, and then proactively terminates the process of the target simulation tasks to completely reclaim the memory resources they occupy, thereby ensuring the smooth operation of high-priority tasks. Once the remaining available resources of the server node recover, the execution of the target simulation tasks is resumed from the saved live image file. This method avoids the controlled target simulation tasks from entering a memory swapping state due to memory pressure, prevents them from being abnormally terminated by the operating system, avoids damage to simulation waveforms, and ensures simulation throughput.

[0101] Logic simulation of large-scale chips typically requires a significant amount of computing and memory resources, and the demand for computing and memory resources in chip design simulation tasks fluctuates dynamically during the analysis, refinement, and simulation phases.

[0102] In practical implementation, the control device deployed on each server node of the chip design simulation cluster receives chip design simulation tasks in real time while continuously monitoring the memory usage status data (including CPU utilization and memory occupancy) of the server node. When the remaining available resources of the server node are detected to be occupied to a certain extent (e.g., the remaining available resources are less than 10% of the total, this threshold can be configured empirically), the control device selects at least one target simulation task from the chip design simulation tasks running on that server node according to specified rules. In addition, when the growth slope of the memory usage status data of the server node exceeds a preset slope threshold indicating that the server node is about to trigger memory swapping (e.g., within a preset monitoring time window, the increase in the memory occupancy rate of the server node exceeds a preset percentage threshold, this threshold can be configured empirically, such as 10% or 20%, but the memory pressure rise rate corresponding to the preset slope threshold must be lower than the memory pressure rise rate required to trigger memory swapping), the control device also selects a target simulation task according to specified rules.

[0103] Taking the IBM Load Sharing Facility (LSF) job scheduling system as an example, this job scheduling system is an independent external program. The control device achieves functional interaction by calling LSF basic commands, and there are no version requirements. Specifically, for the selected target simulation task, its current simulation stage is first determined: if the target simulation task is currently in the simulation stage, the functional interface (such as the save function of the VCS simulation tool) of the simulation tool corresponding to the target simulation task (such as Synopsys VCS simulation tool, Cadence Xcelium simulation tool, or Mentor Questa simulation tool, etc.) is called to save the field image file, generate a field image file for resuming execution from the current simulation progress point, and store the field image file in the local or directly connected storage path of the server node; after confirming that the field image file has been generated, the job termination command is called to forcibly terminate the process of the target simulation task, and the target simulation task enters the exit state. It should be noted that even if the target simulation task had automatic restart enabled via the automatic restart option (-r) in the job submission command during the original submission, it will not be automatically restarted by the job scheduling system after being forcibly terminated by the job termination command. In this way, the memory resources occupied by the target simulation task are released in advance. If the target simulation task is currently in the analysis or refinement phase, since saving the live image file cannot be performed at this stage, the job termination command is directly invoked to forcibly terminate the process of the target simulation task, thereby releasing the server resources it occupies in advance.

[0104] Information about terminated target simulation tasks is recorded in a pre-defined suspended task record. The record includes the job identifier of the target simulation task at the time of termination, the original submission command (which can be obtained through the job history viewing command (bhist -ljobid)), the storage path of the on-site image file, and the simulation stage at the time of termination.

[0105] The control device continuously monitors the memory usage status data of the server node. When the current memory usage status data indicates that the remaining available resources of the server node have recovered to a preset recovery threshold (e.g., the remaining available resources have recovered to more than 30% of the total resources), it reads the information of the tasks to be recovered from the suspended task record one by one according to the specified rules. Generally, the order of restarting tasks is the reverse of the suspension operation, that is, last-in-first-out, with the most recently terminated target simulation task being resubmitted for recovery first. For target simulation tasks in the simulation stage that have saved the field image file, the corresponding field image file is read, and the target simulation task is resubmitted on the same server node through the job submission command, with recovery parameters added to the submission command (such as the simulation execution recovery command (simv-r) of the VCS simulation tool), so that it continues to execute the chip design simulation from the corresponding simulation progress point; for target simulation tasks in the analysis or refinement stage, since the field image file has not been saved, the target simulation task is resubmitted on the same server node through the job submission command, so that it starts executing the chip design simulation from the beginning.

[0106] The above method enables the memory resources occupied by the controlled target simulation task to be completely reclaimed and redistributed to other simulation tasks that continue to run. This ensures that the computing and memory resources of the running tasks are sufficient, avoids the simulation performance degradation caused by the partial memory occupied by the task being swapped out to the hard disk, and prevents the target simulation task from being abnormally terminated by the operating system and losing the running time, thus keeping the server resource utilization at a high level.

[0107] The specific implementation process is as follows: 1. Resource monitoring and trigger preparation The control device deployed on each server node continuously monitors the resource usage of the server node. When the remaining available resources of the server node are detected to be lower than a preset threshold (e.g., the remaining available resources are lower than 10% of the total), the control device initiates a suspension operation process.

[0108] 2. Screening of target simulation tasks to be suspended The control device selects at least one target simulation task from the chip design simulation tasks running on the current server node according to specified rules. The selection rules comprehensively consider the following dimensions: (a) Task submission priority: Tasks with lower priority are selected first; (b) Task runtime: Tasks with shorter runtimes are selected first. Prioritizing tasks with shorter runtimes is because tasks with longer runtimes may be nearing their natural end, and once they end naturally, the computing and memory resources they occupy can be completely released.

[0109] Furthermore, when the control device is run with server administrator privileges, it also needs to comprehensively evaluate the distribution of the number of tasks currently running for each user. If a user has significantly more tasks running than other users, it is assumed that the user may be executing non-urgent batch tasks such as regression testing, and their tasks will be prioritized as target simulation tasks. Generally, the more tasks a user has running, the higher the priority of their tasks being selected.

[0110] The following is an example of the target simulation task selection algorithm: Assuming the control device is operated by a regular user or the range of users to be screened has been determined, the following algorithm is used to calculate the suspension score of each candidate task, and the target simulation task is selected based on the score ranking: Input: The set of all running chip design simulation tasks on the current server node. .

[0111] Each task's attributes include: (1) Priority P, with a value range of 1 to 10. The larger the value, the higher the priority. (2) Runtime R, in minutes; (3) Current memory usage in M, in gigabytes.

[0112] In addition to the aforementioned formula (I), the formula for calculating the suspended score can also be simplified to the following form: That is, in addition to the aforementioned formula (II), It can also be simplified to: =(10000-R); where, , and The weighting coefficients are configurable; for example, they can be configured as follows: =5、 =0.1、 =2. The formula uses 10,000 minutes as the upper limit of the running time. When the running time R of a task exceeds 10,000 minutes, the second term (10,000-R) becomes negative, which significantly reduces the suspension score of the task, so as to avoid selecting tasks that have been running for a very long time and are close to natural termination.

[0113] The control device calculates the suspension score for each task, sorts them in descending order of score, and selects the task with the highest score as the target simulation task. If more memory resources need to be released, subsequent tasks are selected in descending order of score until the estimated amount of memory to be released meets the resource recovery requirements.

[0114] 3. Processing flow of target simulation task After selecting the target simulation tasks, the control device sequentially performs processing operations on each target simulation task. The following explanation uses the IBM LSF job scheduling system as the task submission system and VCS as the simulation tool. It is understood that other simulation tools, such as the parallel logic simulation platform (Cadence Xcelium) and the advanced verification simulation platform (Mentor Questa), can also be used. Specific details are as follows: Step 3.1: Simulation Phase Determination: Chip design simulation tasks sequentially go through the analysis phase, refinement phase, and simulation phase within their lifecycle. The execution of the simulation phase depends on the executable file generated in the previous phase (e.g., the simv executable file generated by VCS). If this executable file exists in the task's runtime directory, it indicates that the task has entered the simulation phase. Since the simulation engine only possesses complete internal state serialization capabilities and can perform state saving operations during the simulation phase, it is necessary to determine the current simulation phase of the target simulation task before executing subsequent steps.

[0115] If the target simulation task is currently in the analysis or refinement stage, since saving the current state cannot be performed at this stage, the process will proceed directly to step 3.5 for further processing.

[0116] It should be noted that the task-stage detection is performed periodically at a preset frequency (e.g., every few minutes). If the task is in the analysis stage at the moment of detection but immediately switches to the simulation stage, the impact of this boundary case on the overall control effect is negligible due to the high detection frequency, and no special handling is required. In addition, for target simulation tasks that are in the simulation stage but have been running for too short a time (e.g., only a few minutes), the cost-effectiveness of saving the execution state is not high, and you can directly jump to step 3.5 for processing.

[0117] Step 3.2: Generate save instruction marker: For a target simulation task that is in the simulation stage and meets the runtime requirements, the control device creates a marker file named "save.txt" (save document) in the running directory of the task and writes the content "marker to be saved" into the file.

[0118] Step 3.3: Perform on-site saving and anomaly monitoring: The target simulation task has a resident monitoring process (pre-integrated into the simulation task startup script). This monitoring process detects in real time whether a "save.txt" file exists in the running directory and checks whether the file contains a "prepared to save" flag. Once this flag is detected, the monitoring process calls the built-in save function of the simulation tool. This save function packages all the current states of the target simulation task—including the current values ​​of all registers, memory, combinational logic, and the internal state of the simulation engine—into a on-site image file and stores this on-site image file in a specified storage path (e.g., the local storage of the server node or directly connected network storage). After successful saving, the monitoring process appends a "save successful" flag to the "save.txt" file.

[0119] The control device synchronously monitors whether the on-site image file has been saved successfully. The scenarios and handling methods for saving failure include: (1) The target simulation task finishes running just before the save action is executed. In this case, the save function will not be triggered, and therefore no on-site image file will be generated. The control device can further confirm the running status of the target simulation task through the job viewing command (bjobs jobid); if it is confirmed that the task has ended, the current processing flow will be terminated.

[0120] (2) Insufficient disk space. Server administrators typically deploy independent storage monitoring scripts for early warning. If saving fails due to insufficient disk space, the simulation tool will output the corresponding error message. After capturing this error message, the control device can record the failure event and trigger an alarm, while skipping the current processing of the target simulation task.

[0121] Step 3.4: Task Termination and Information Recording: When the control device detects a "save successful" flag in the "save.txt" file, it calls the job termination command to forcibly terminate the process of the target simulation task, and the target simulation task enters the exit state. It should be noted that even if the target simulation task had automatic restart enabled via the automatic restart option (-r) in the original job submission command, it will not be automatically restarted by LSF after being forcibly terminated by the job termination command. Through the above method, the memory resources occupied by the target simulation task are completely reclaimed.

[0122] Subsequently, the control device records the relevant information of the target simulation task in a preset suspended task list file. The recorded content includes, but is not limited to: the job identifier of the target simulation task at the time of termination, the original submission command (which can be obtained by viewing the command in the job history), the storage path of the on-site image file, the simulation stage at the time of termination, and the termination timestamp.

[0123] Step 3.5: User Notification and Autonomous Migration: After the target simulation task is terminated, the control device automatically sends a notification email to the user to whom the task belongs. The email content includes: the target simulation task has been terminated, the reason for termination, the storage path of the field image file, and optional operation instructions for the user. The user can decide whether to migrate the target simulation task to another server node to resume operation according to their own needs.

[0124] If a user decides to migrate the target simulation task to another server node for resumption, they must first notify the control device so that the control device can delete the corresponding record of the target simulation task from the suspended task list file. This prevents the control device from repeatedly submitting the resumption of the target simulation task after the original server node's resources are restored. After the user completes the notification, they can manually resume the execution of the target simulation task on other server nodes using the field image file.

[0125] 4. Resource recovery and resumption of target simulation tasks The control device continuously monitors the memory usage status data of the server node. When the monitored memory usage status data indicates that the remaining available resources of the server node have recovered to a preset recovery threshold (e.g., the remaining available resources have recovered to more than 30% of the total resources), the task recovery process is initiated. Specific details are as follows: Step 4.1: Resource Pre-Check Before Recovery: Before resubmitting the terminated target simulation task, the control device first obtains the resource requirements (including memory reservation, number of CPU cores, etc.) declared by the target simulation task at the time of the original submission using the job history viewing command (bhist -l jobid), and compares these resource requirements with the currently available resources of the server node. The resubmission operation is only performed if the currently available resources of the server node meet the resource requirements. This resource pre-check mechanism effectively avoids repeated processing (i.e., jitter) caused by resource shortages being triggered again immediately after the target simulation task is restored.

[0126] Step 4.2: Task recovery order: The control device reads the information of the tasks to be recovered from the suspended task list file in a last-in-first-out order, that is, the target simulation task that was most recently terminated is given priority to be resubmitted and resumed.

[0127] Step 4.3: Differentiate recovery methods based on simulation stage: The control device employs different recovery methods based on the simulation stage at which the target simulation task was terminated, as recorded in the suspended task list file. (1) If the target simulation task is terminated in the analysis or refinement stage, and the field image file is not saved, the target simulation task is directly resubmitted through the job submission command so that the chip design simulation is executed from scratch.

[0128] (2) If the target simulation task is terminated while it is in the simulation stage and the field image file has been saved, then read the corresponding field image file, resubmit the target simulation task through the job submission command, and add the recovery parameter in the submission command. Taking the VCS simulation tool as an example, use the simulation execution recovery command (simv-r<field image file path>) to make the target simulation task load the field image file and continue to execute the chip design simulation from the simulation progress point when it was saved.

[0129] After restarting, the target simulation task reads the field image file, reconstructs all internal states, and continues execution from the save point. For the user, this means the target simulation task resumes after a pause, rather than starting from scratch, thus avoiding the repeated consumption of previous simulation time.

[0130] Step 4.4: Explanation of Consistency Restoration and Scheduling Conflicts: Regarding the consistency of simulation results after restoration from the field image file, this consistency is guaranteed by the internal mechanisms of the simulation tools (such as VCS, Xcelium, Questa, etc.), and the control device does not need to set up an additional verification mechanism. Regarding the compatibility issue between resubmitted jobs and the LSF job scheduling system's own scheduling policies (such as job priority, resource reservation, etc.), when resubmitting, the job submission command submits the task to LSF. LSF will determine whether to allow the task to be submitted normally and allocate resources according to its own scheduling policy, and the control device does not need to perform any special processing in this regard.

[0131] 5. Supplementary Explanation on Handling Abnormal Scenarios The following is a summary of the exception handling logic involved in the above process: Step 5.1: Insufficient Disk Space: As described in Step 3.3, disk space shortages are typically alerted by the server administrator through a separate monitoring script. If saving the live image file fails due to insufficient disk space, the simulation tool will output the corresponding error message. The control device will capture and record the failure event and trigger an alarm, while simultaneously skipping the current processing of that target simulation task.

[0132] Step 5.2: Early Task Termination: As described in Step 3.3, if the target simulation task finishes running just before the save action is executed, the save function will not be triggered, and the field image file will not be generated. The control device will terminate the current processing flow after confirming the task status through the job viewing command.

[0133] Step 5.3: Server Downtime: If a server node crashes, all running, unfinished tasks will typically fail. For such catastrophic anomalies, it is recommended that a separate system disaster recovery program handle the situation. Within the scope of this solution, for target simulation tasks whose state snapshots were successfully saved before the crash, users can resubmit the task using these snapshots after the server recovers, thus saving some initial simulation time.

[0134] Step 5.4: Multi-server instance coordination: In a multi-server cluster environment with shared network storage, the control device running on each server node only manages chip design simulation tasks submitted to its own node. The control device obtains a task list limited to the current server node through a job viewing command. The management boundaries between instances are clear, and there is no situation where multiple instances manage the same task simultaneously.

[0135] Step 5.5: Manual Resumption by User: As described in Step 3.5 above, after the target simulation task is terminated, the control device has sent a notification email to the user. If the user chooses to manually resume the task on another server node, they must first notify the control device to delete the corresponding record from the suspended task list file. Typically, terminated tasks will not be automatically resumed by the control device in a very short time, giving the user ample time to perform the notification operation.

[0136] The above-mentioned monitoring, screening, processing, and recovery steps are repeatedly executed during the operation of the control device to achieve continuous closed-loop control of the memory pressure of each server node in the chip design simulation cluster.

[0137] In other words, this application example can prevent the controlled target simulation task from entering the memory swapping state due to memory pressure when a large number of simulation tasks are submitted, prevent the controlled target simulation task from being abnormally terminated by the operating system, reduce the idle rate of simulation cluster server resources, and improve the resource utilization efficiency of simulation cluster server.

[0138] Based on this, the memory pressure control method and apparatus for chip design simulation tasks provided in this application example have the following advantages compared to existing same-node operating system-level suspension schemes and cross-node migration schemes: (1) Regarding memory resource reclamation, this application actively terminates the process of the target simulation task to reclaim the memory resources it occupies. This allows the physical memory occupied by the controlled task to be completely released and redistributed to other simulation tasks that continue to run on the server node. Unlike the existing same-node operating system-level suspension scheme where processes are only suspended while memory pages remain in physical memory, this application can fundamentally avoid the situation where memory pages are swapped out to memory swapping due to the memory resources not being truly released, thereby eliminating the simulation performance degradation caused by memory swapping.

[0139] (2) Regarding the control triggering mechanism, this application monitors the growth slope of the server node's memory usage status data, enabling it to predict the rising trend of memory pressure before the server node actually triggers memory swapping, and proactively intervene before memory swapping occurs. Unlike the passive response mechanism based on static thresholds in existing solutions, this application can precisely lock the intervention timing at the critical window where memory pressure rises rapidly but memory swapping has not yet been triggered, achieving predictive proactive protection for the controlled task.

[0140] (3) Regarding the control location and recovery method, this application completes the entire closed-loop process of saving the on-site image file of the target simulation task, terminating the process, recycling resources, and resuming execution within the same server node. Unlike existing cross-node migration schemes that rely on the existence of idle spare nodes with larger memory in the cluster and require cross-node transmission of session information, generating network input / output overhead, this application does not rely on spare nodes and has no cross-node network transmission overhead, which can significantly shorten the task recovery time and is suitable for actual deployment environments where there are no idle spare nodes in the cluster.

[0141] (4) Regarding adaptation during the simulation phase, this application adopts a differentiated handling strategy to address the differences in the support for field save operations in the analysis, refinement, and simulation phases of chip design simulation tasks. For target simulation tasks in the simulation phase, the simulation tool's functional interface can be called to generate a field image file and resume execution from the save point; for target simulation tasks in the analysis or refinement phases, the field save step can be skipped to directly terminate the process to quickly release resources, and the process can be resubmitted for execution after resource recovery. This hierarchical handling method can avoid simulation tool anomalies caused by blindly initiating field save operations, and improve the high compatibility and robustness of the control system when managing various simulation tasks.

[0142] (5) Regarding the pre-recovery resource check, before reading the site image file and resubmitting the target simulation task, this application obtains the resource requirements declared at the time of the original task submission and compares them with the currently available resources. The recovery operation is only performed when the resources meet the requirements. This mechanism can effectively avoid repeated adjustments and task jitter caused by resource shortages being triggered again immediately after the target simulation task is recovered, and ensure the stable operation of the simulation task after recovery.

[0143] (6) Regarding the selection of target simulation tasks, this application calculates a suspended score by weighting three dimensions: task priority, runtime, and current memory usage, and selects target simulation tasks based on the score ranking. This multi-dimensional selection method can prioritize the release of tasks with high memory usage and minimal impact on the overall simulation progress, maximizing memory release benefits while ensuring the smooth operation of high-priority tasks. Furthermore, in a multi-user environment, this application can prioritize the selection of target simulation tasks from users with a large number of running tasks, achieving load balancing control in multi-user scenarios and preventing a single user from occupying too many resources and affecting the normal simulation progress of other users.

[0144] (7) Regarding the task recovery order, this application adopts the last-in-first-out order to recover the terminated target simulation tasks one by one, which can enable tasks with shorter waiting times to be recovered first, reduce task backlog caused by regulation, and make the overall waiting time distribution of each regulated task more balanced.

[0145] (8) Regarding user interaction, this application can proactively send notification messages to users after the target simulation task is terminated, informing them of the task status and the storage path of the on-site image file, and supporting users to decide whether to migrate the task to other server nodes for resumption of execution. This mechanism can give users flexible scheduling rights over their own simulation tasks, and through the linkage between user feedback and the deletion of suspended task records, it can avoid the control device from repeatedly resuming tasks that have been migrated by users on the original server node, ensuring the consistency between the control process and the user's independent decision-making.

[0146] (9) In terms of storage management, this application establishes a mapping relationship between the original job identifier of the terminated task, the new job identifier obtained after resubmission, and the storage path of the field image file, and performs retention or deletion operations on the field image file according to the preset cleanup strategy. This enables closed-loop lifecycle management of the field image file generated during the control process, and avoids expired or invalid field image files occupying storage space for a long time.

[0147] (10) Regarding electronic design automation license resources, this application adopts the method of actively terminating the target simulation task process. This process termination method belongs to the normal exit process of the simulation tool, which can trigger the return verification of electronic design automation licenses in an instant. Compared with the state of license not being released in time (false occupation) caused by the forced termination of the operating system, it can accelerate the release and circulation of EDA license resources.

[0148] (11) In terms of system deployment and versatility, this application operates as a middleware device independent of the job scheduling system and simulation tools. It achieves closed-loop control by calling the standard command line interface of the job scheduling system and the standard function interface of the simulation tools. It does not require customized modification of the existing electronic design automation toolchain and cluster management system. It can seamlessly adapt to different versions of the job scheduling system and a variety of mainstream simulation tools. It is easy to deploy and has strong versatility.

[0149] This application also provides an electronic device, which may include a processor, a memory, a receiver, and a transmitter. The processor is used to execute the memory pressure control method for chip design simulation tasks mentioned in the above embodiments. The processor and the memory can be connected via a bus or other means, taking a bus connection as an example. The receiver can be connected to the processor and the memory via wired or wireless means.

[0150] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the aforementioned memory pressure control method for chip design simulation tasks. The computer-readable storage medium can be a tangible storage medium, such as random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, floppy disks, hard disks, removable storage disks, CD-ROMs, or any other form of storage medium known in the art.

[0151] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the aforementioned memory pressure control method for chip design simulation tasks.

[0152] Those skilled in the art will understand that the exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Whether implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application. When implemented in hardware, it can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. The programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave.

[0153] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.

[0154] In this application, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or in place of features of other embodiments.

[0155] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to the embodiments of this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for controlling memory pressure in chip design simulation tasks, characterized in that, include: Within a server node of a chip design simulation cluster that has not triggered memory swapping, continuously monitor the memory usage status data of that server node; If the growth slope of the current memory usage status data exceeds a preset slope threshold used to indicate that the server node is about to trigger memory swapping, then the priority, runtime, and current memory usage of each chip design simulation task running on the server node are obtained. Based on the priority, runtime, and current memory usage of each chip design simulation task, a suspension score is calculated for each chip design simulation task. At least one chip design simulation task is selected as the current target simulation task in descending order of the suspension score. If the target simulation task is currently in the simulation stage, the function interface of the simulation tool corresponding to the target simulation task is called to generate a live image file for resuming execution from the current simulation progress point, and the live image file is stored in the local or directly connected storage path of the server node. After confirming that the live image file has been generated, the process of the target simulation task is actively terminated to reclaim the memory resources occupied by the target simulation task. If the current memory usage data indicates that the remaining available resources have recovered to the preset recovery threshold, then the field image file stored locally or in the direct-connected storage path is read, and the target simulation task corresponding to the field image file is resubmitted on the server node, so that the target simulation task resumes execution of chip design simulation from the simulation progress point corresponding to the field image file.

2. The memory pressure control method for chip design simulation tasks according to claim 1, characterized in that, The memory usage status data includes the memory occupancy rate of the server node monitored within a preset monitoring time window.

3. The memory pressure control method for chip design simulation tasks according to claim 1, characterized in that, The expression for the suspended score is shown in Formula (I): (one) In formula (1), The suspended score is given. The value represents the priority, and a larger value indicates a higher priority. The duration of operation. The current memory usage, It is an exponentially decaying function that monotonically decreases with the duration R of operation. , and These are the weighting coefficients; The exponential decay function The expression is shown in formula (II): (two) in, This is a preset upper limit for the runtime reference. The expected total simulation time for the chip design simulation task is estimated, where k is the attenuation steepness coefficient.

4. The memory pressure control method for chip design simulation tasks according to claim 3, characterized in that, The step of selecting at least one chip design simulation task as the current target simulation task also includes: Before obtaining the priority, runtime, and current memory usage of each chip design simulation task running on the server node, first count the number of chip design simulation tasks currently running on the server node for each user. The users are sorted in descending order of the number of chip design simulation tasks they are running, and the top N users are selected as candidate users, where N is a positive integer. Correspondingly, the priority, runtime, and current memory usage of each chip design simulation task running on the server node are obtained, including: Obtain the priority, runtime, and current memory usage of the chip design simulation task currently running on the server node for each of the candidate users.

5. The memory pressure control method for chip design simulation tasks according to claim 1, characterized in that, Before reading the site image file stored locally or in a directly connected storage path, the method further includes: If the current memory usage data indicates that the remaining available resources have recovered to the preset recovery threshold, then the resource requirements declared by the target simulation task corresponding to the field image file at the time of the original submission are obtained from the preset suspended task record. Compare the resource requirements with the currently available resources of the server node; If the current available resources of the server node meet the resource requirements, then it is determined to read the site image file.

6. The memory pressure control method for chip design simulation tasks according to claim 1, characterized in that, The reading of the site image file stored locally or in a directly connected storage path includes: Following a last-in-first-out (LIFO) order, the field image files stored locally or in a directly connected storage path are read one by one; wherein, LIFO means that the chip design simulation task that was most recently terminated is resubmitted and resumed execution first.

7. The memory pressure control method for chip design simulation tasks according to claim 1, characterized in that, The process of actively terminating the target simulation task includes: Invoke the job termination command of the job scheduling system running on the server node to forcibly terminate the process of the target simulation task.

8. The memory pressure control method for chip design simulation tasks according to claim 1, characterized in that, After the process of actively terminating the target simulation task is completed, the following steps are also included: Send a notification message to the user to whom the terminated target simulation task belongs; wherein the notification message is used to inform the user that the target simulation task has been terminated, the reason for the termination, and the storage path of the field image file; If a migration instruction for a terminated target simulation task is received from the user, the record corresponding to the target simulation task is deleted from the preset suspended task record to avoid resuming the execution of the target simulation task on the server node.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the memory pressure control method for chip design simulation tasks as described in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the memory pressure control method for the chip design simulation task as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Memory-tracking resource manager for elastic distributed graph-processing system

    US20250094224A1