A method and apparatus for distributing statistics of application computing resources of a supercomputer
Patent Information
- Application Number
- CN202211582336.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-09
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2042-12-09
AI Technical Summary
[0006]本发明要解决的技术问题在于,针对现有技术的上述缺陷,提供一种超级计算机的应用计算资源分布统计方法及装置,旨在解决现有技术中超级计算机上各个应用的计算资源记录数据量很大,无法直观的反映超级计算机上的应用情况的问题
[0055]本发明提供一种超级计算机的应用计算资源分布统计方法及装置,所述方法包括:获取作业计算资源信息,识别所述作业计算资源信息中的作业号以及与所述作业号对应的计算资源使用量;获取作业日志信息,根据所述作业号查找所述作业日志信息,得到与所述作业号相匹配的应用名称;对匹配到同一应用名称的作业号对应的计算资源使用量进行统计,得到每个应用名称对应的计算资源累计使用量;将所述应用名称按照预设分类规则进行分类处理,按照分类处理后的应用名称所对应的计算资源累计使用量生成应用计算资源分布图。本发明通过对各个应用名称对应的计算资源使用量进行统计,并按照分类结果生成应用计算资源分布图,从应用计算资源分布图上可以直观地查看超级计算机的应用情况。
Smart Images

Figure CN116089065B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of supercomputer technology, and in particular to a method and apparatus for statistical analysis of application computing resource distribution in supercomputers. Background Technology
[0002] Supercomputing has become an indispensable tool for exploring and transforming the world. A supercomputer is a computer capable of processing massive amounts of data and performing high-speed calculations that are impossible for a regular personal computer. Its main characteristics include two aspects: enormous data storage capacity and extremely fast data processing speed. Therefore, it can perform tasks in various fields that are incapable of being done by humans or ordinary computers. Supercomputers can be used to unravel the mysteries of life and thought, explore the mechanisms of drug action, discover and design new materials, simulate new energy devices, predict disasters, simulate fluid dynamics, explore the unknown realms of matter and the universe, conduct large-scale data mining, and process large-scale complex social networks. According to incomplete statistics, there are hundreds of applications in high-performance computing, and when AI and big data applications are included, the number of applications for supercomputing can reach thousands.
[0003] Therefore, understanding the application usage on supercomputers is very meaningful, such as tracking the number of application calls and the number of users. The application usage on a supercomputer can also be reflected by statistically analyzing the computing resource usage of each application. For example, if the Fluent application used one CPU core for 10 hours in a week, then the Fluent application's weekly core-hour usage is 10 core-hours. This information can serve as a key focus for research support work in computing centers, provide insights into the development trends of the supercomputing industry, and serve as an important basis for research on next-generation supercomputers. For instance, if the VASP application for materials calculation used the most computing resources in the past week, then investment in the VASP application will increase (e.g., increase manpower, material resources, and financial resources); similarly, if the WRF application has seen the largest increase in computing resources in the past few weeks, then the importance attached to this application will also increase accordingly.
[0004] Each job already has a corresponding log file, which includes the application name for that job. The supercomputer also records the computing resources of each job over a specific time period in another file. However, the system's setting to scan job information every few minutes (e.g., every 5 minutes) results in nearly 20,000 records of application information per job per day, along with a massive number of corresponding computing resource records. Therefore, directly searching for each application's computing resource records to understand the application status on the supercomputer is extremely inconvenient.
[0005] Therefore, existing technologies have shortcomings and need to be improved and developed. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a method and apparatus for statistical analysis of application computing resource distribution on a supercomputer, in order to address the above-mentioned deficiencies of the prior art. This method aims to solve the problem that the amount of computing resource recording data for various applications on a supercomputer is too large to intuitively reflect the application status on the supercomputer.
[0007] The technical solution adopted by this invention to solve the technical problem is as follows:
[0008] A statistical method for the distribution of application computing resources in a supercomputer, comprising:
[0009] Obtain job computing resource information, and identify the job number and the computing resource usage corresponding to the job number in the job computing resource information;
[0010] Obtain job log information, search for the job log information based on the job number, and obtain the application name that matches the job number;
[0011] The computing resource usage corresponding to job numbers that match the same application name is statistically analyzed to obtain the cumulative computing resource usage for each application name;
[0012] The application names are categorized according to preset classification rules, and an application computing resource distribution map is generated based on the cumulative computing resource usage corresponding to the categorized application names.
[0013] In one implementation, obtaining job computing resource information and identifying the job number and the computing resource usage corresponding to the job number in the job computing resource information includes:
[0014] Upload the task calculation resource information for the first preset time period to the directory where the task is located;
[0015] Identify the job number in the directory where the task is located and the computing resource usage corresponding to the job number.
[0016] In one implementation, the step of obtaining job log information and searching the job log information according to the job number to obtain the application name matching the job number includes:
[0017] Obtain the operation log information within the second preset time period;
[0018] Iterate through the job numbers in the directory where the task is located, and find the application name that matches each job number in the job log information;
[0019] Save the job number, the computing resource usage corresponding to the job number, and the application name matching the job number to the first record file;
[0020] The first preset time period falls within the second preset time period.
[0021] In one implementation, after saving the job number, the computing resource usage corresponding to the job number, and the application name matching the job number to the first record file, the method further includes:
[0022] If the first record file contains an application name that matches the job number but is empty, then the application name of the job number will be assigned the first preset name.
[0023] In one implementation, the step of statistically analyzing the computing resource usage corresponding to job numbers matching the same application name to obtain the cumulative computing resource usage for each application name includes:
[0024] Save each application name and its corresponding computing resource usage to the second record file;
[0025] The same application names in the second record file are deduplicated, and the computing resource usage corresponding to each application name after deduplication is accumulated to obtain the cumulative computing resource usage corresponding to each application name.
[0026] In one implementation, after calculating the computing resource usage corresponding to job numbers matching the same application name and obtaining the cumulative computing resource usage for each application name, the method further includes:
[0027] The application names are sorted according to the cumulative usage of computing resources to obtain the original application computing resource data.
[0028] In one implementation, the application names are categorized according to preset classification rules, and an application computing resource distribution map is generated based on the cumulative computing resource usage corresponding to the categorized application names, including:
[0029] Save the application name and corresponding cumulative computing resource usage belonging to the computing node, the first preset name, or the self-developed application to the first result file, and save the application name and corresponding cumulative computing resource usage not belonging to the computing node, the first preset name, or the self-developed application to the second result file.
[0030] Obtain the preset application dictionary and iterate through the application names in the second result file;
[0031] Save the application names and corresponding cumulative computing resource usage that belong to the preset application dictionary but do not belong to the platform application in the second result file to the first result subfile;
[0032] Save the application names and corresponding cumulative computing resource usage that belong to the preset application dictionary and platform application in the second result file to the second result sub-file, and replace the application names in the second result sub-file with the first preset name;
[0033] Save the application names in the second result file that do not belong to the preset application dictionary to the third result sub-file;
[0034] Create an application computing resource distribution map based on the first result file, the first result sub-file, the second result sub-file, and the third result sub-file.
[0035] In one implementation, an application computing resource distribution map is created based on the first result file, the first result sub-file, the second result sub-file, and the third result sub-file, including:
[0036] Obtain the first result file, the second result sub-file, and the third result sub-file, and classify the application names in the first result file, the second result sub-file, and the third result sub-file into a second preset name;
[0037] The cumulative usage of computing resources corresponding to each of the second preset names is added together to obtain the total usage of the second computing resources corresponding to the second preset name;
[0038] Generate an application computing resource distribution map based on the second total computing resource usage and the first result sub-file;
[0039] The application computing resource distribution map includes: the computing resource percentage corresponding to each application name in the first result sub-file, and the computing resource percentage corresponding to the second preset name.
[0040] In one implementation, an application computing resource distribution map is created based on the first result file, the first result sub-file, the second result sub-file, and the third result sub-file, including:
[0041] Obtain the first result file, the first result sub-file, the second result sub-file, and the third result sub-file;
[0042] Sort the application names in the first result subfile according to the cumulative computing resource usage from largest to smallest to obtain the computing resource ranking result;
[0043] Based on the sorting results of the computing resources, a preset number of application names are obtained as target application names, and the application names other than the target application names in the first result sub-file are classified as a third preset name.
[0044] The application names in the first result file, the second result sub-file, and the third result sub-file are all classified as the third preset name;
[0045] The cumulative usage of computing resources corresponding to each of the third preset names is added together to obtain the total usage of the third computing resources corresponding to the third preset name.
[0046] Based on the total usage of the third computing resources and the target application name and its corresponding cumulative computing resource usage, an application computing resource distribution map is generated.
[0047] The application computing resource distribution map includes: the proportion of computing resources corresponding to each of the target application names, and the proportion of computing resources corresponding to the third preset name.
[0048] The present invention also provides a supercomputer application computing resource distribution statistics device, comprising:
[0049] The identification module is used to obtain job computing resource information and identify the job number and the computing resource usage corresponding to the job number in the job computing resource information.
[0050] The search module is used to obtain job log information, search the job log information according to the job number, and obtain the application name that matches the job number.
[0051] The statistics module is used to count the computing resource usage corresponding to job numbers that match the same application name, and to obtain the cumulative computing resource usage for each application name.
[0052] The generation module is used to classify the application names according to preset classification rules and generate an application computing resource distribution map based on the cumulative computing resource usage corresponding to the classified application names.
[0053] The present invention also provides a terminal, comprising: a memory, a processor, and a supercomputer application computing resource distribution statistics program stored in the memory and executable on the processor, wherein the supercomputer application computing resource distribution statistics program, when executed by the processor, implements the steps of the supercomputer application computing resource distribution statistics method as described above.
[0054] The present invention also provides a computer-readable storage medium storing a computer program that can be executed to implement the steps of the supercomputer application computing resource distribution statistics method as described above.
[0055] This invention provides a method and apparatus for statistically analyzing the distribution of application computing resources in a supercomputer. The method includes: acquiring job computing resource information, identifying the job number and the computing resource usage corresponding to the job number in the job computing resource information; acquiring job log information, searching the job log information according to the job number to obtain the application name matching the job number; statistically analyzing the computing resource usage corresponding to job numbers matching the same application name to obtain the cumulative computing resource usage for each application name; classifying the application names according to a preset classification rule, and generating an application computing resource distribution map based on the cumulative computing resource usage corresponding to the classified application names. This invention, by statistically analyzing the computing resource usage corresponding to each application name and generating an application computing resource distribution map according to the classification results, allows for a direct and intuitive view of the supercomputer's application status. Attached Figure Description
[0056] Figure 1 This is a flowchart of a preferred embodiment of the statistical method for the distribution of application computing resources of supercomputers in this invention.
[0057] Figure 2 This is a flowchart of a specific embodiment of the application computing resource distribution statistics method for supercomputers in this invention.
[0058] Figure 3 This is a distribution map of application computing resources in this invention.
[0059] Figure 4 This is a functional principle block diagram of a preferred embodiment of the supercomputer application computing resource distribution statistics device of the present invention.
[0060] Figure 5 This is a functional principle block diagram of a preferred embodiment of the terminal in this invention. Detailed Implementation
[0061] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0062] On a supercomputer, if application computing resource distribution statistics are performed periodically (e.g., weekly), up to 140,000 records need to be processed. Current technology cannot quickly and accurately determine the distribution of application computing resources over a given period.
[0063] Please see Figure 1 , Figure 1 This is a flowchart illustrating the distribution statistics of application computing resources in the supercomputer of this invention. For example... Figure 1As shown, the application computing resource distribution statistics of the supercomputer described in this embodiment of the invention include:
[0064] Step S100: Obtain job computing resource information and identify the job number and the computing resource usage corresponding to the job number in the job computing resource information.
[0065] Specifically, the supercomputer calculates the computing resource usage for each job and saves it as job computing resource information. This information includes: job number, username, user group name, number of job cores, submission time, start time, end time, computing resource usage, and cost.
[0066] In one implementation, step S100 specifically includes:
[0067] Step S110: Upload the job calculation resource information within the first preset time period to the directory where the task is located;
[0068] Step S120: Identify the job number in the directory where the task is located and the computing resource usage corresponding to the job number.
[0069] Specifically, in this embodiment, the task's computational resource information for the first preset time period is uploaded to the task's directory, allowing the task to be executed. This embodiment can pre-set the first preset time period, for example, setting it to one week. This way, the application computational resource distribution for the past seven days is calculated every week, allowing users to intuitively view the supercomputer's weekly application status.
[0070] like Figure 1 As shown, the application computing resource distribution statistics of the supercomputer described in this embodiment also include:
[0071] Step S200: Obtain job log information, search the job log information according to the job number, and obtain the application name that matches the job number.
[0072] Specifically, the supercomputer stores job log information corresponding to each job, which includes: job number, username, node name, and application name. Therefore, job log information can be retrieved based on the job number to obtain the application name matching the job number.
[0073] In one implementation, step S200 specifically includes:
[0074] Step S210: Obtain the job log information within the second preset time period;
[0075] Step S220: Traverse the job numbers in the directory where the task is located, and find the application name that matches each job number in the job log information;
[0076] Step S230: Save the job number, the computing resource usage corresponding to the job number, and the application name matching the job number to the first record file.
[0077] Specifically, if the first preset time period is set to one week, then the second preset time period can be set to ten days. This means acquiring the job log information from the past ten days and matching the application name among all acquired job log information, thus avoiding information omissions. In this embodiment, the first preset time period is set to fall within the second preset time period, preventing information omissions and ensuring the accuracy of information acquisition. Furthermore, this embodiment stores the original data—job number, the computing resource usage corresponding to the job number, and the application name matching the job number—in a first record file, facilitating subsequent data traceability and improving the accuracy of application computing resource processing.
[0078] In one embodiment, the step S230 is followed by:
[0079] Step S240: If there is an empty application name matching the job number in the first record file, then the application name of the job number is assigned the first preset name.
[0080] Specifically, the log file may contain jobs that do not display application names. In this case, the application name corresponding to the job number is assigned a first preset name, for example, "other," so that each job has an application name for subsequent classification and statistics. This embodiment assigns an application name to jobs with empty application names instead of discarding them directly. This is because discarding them directly would change the computing resource allocation of other applications, reducing accuracy. Therefore, this embodiment ensures the accuracy of application computing resource distribution statistics.
[0081] like Figure 1 As shown, the application computing resource distribution statistics of the supercomputer described in this embodiment also include:
[0082] Step S300: Calculate the computing resource usage corresponding to the job numbers that match the same application name, and obtain the cumulative computing resource usage corresponding to each application name.
[0083] Specifically, in order to clearly reflect the distribution of computing resources for various applications of the supercomputer, the computing resource usage corresponding to the same application name is statistically analyzed to obtain the cumulative computing resource usage for each application within the first preset time period.
[0084] In one implementation, step S300 specifically includes:
[0085] Step S310: Save each application name and its corresponding computing resource usage to the second record file;
[0086] Step S320: Deduplicate the same application names in the second record file, and sum up the computing resource usage corresponding to each application name after deduplication to obtain the cumulative computing resource usage corresponding to each application name.
[0087] Specifically, the computing resource usage of applications with the same name needs to be summed to obtain the total computing resource usage of that application within the first preset time period. Therefore, duplicate application names in the second record file are deduplicated to obtain the cumulative computing resource usage corresponding to each application name. This embodiment saves the deduplicated data in the second record file, allowing each step in the computing resource distribution statistics process to be saved in its corresponding folder, facilitating subsequent retrieval of processing data at each stage and improving the accuracy of application computing resource processing.
[0088] In one embodiment, after step S300, the method further includes: sorting each application name according to the cumulative usage of computing resources to obtain the original application computing resource data.
[0089] Specifically, the cumulative usage of computing resources can be sorted by the original application name, providing a clear view of the usage of each application within the supercomputer. In this embodiment, the original application computing resource data can be stored in a specific folder for easy access to the original application computing resource ranking data later.
[0090] like Figure 1 As shown, the application computing resource distribution statistics of the supercomputer described in this embodiment also include:
[0091] Step S400: Classify the application names according to preset classification rules, and generate an application computing resource distribution map based on the cumulative computing resource usage corresponding to the classified application names.
[0092] Specifically, this embodiment also categorizes the application names on the supercomputer. This is because the applications on supercomputers are very diverse. Some are highly targeted applications for various disciplines, such as the VASP application for materials calculation; some are platform applications, such as MATLAB, Python, and Anaconda; some are user-developed applications; and some are named after computing nodes (a small number of abnormal jobs prevent the acquisition of application information). Categorizing the application names in this embodiment can more intuitively reflect the application status of the supercomputer.
[0093] In one implementation, step S400 specifically includes:
[0094] Step S410: Save the application name and corresponding cumulative computing resource usage of the computing node, the first preset name or the self-compiled application to the first result file, and save the application name and corresponding cumulative computing resource usage of the application that does not belong to the computing node, the first preset name or the self-compiled application to the second result file.
[0095] Step S420: Obtain the preset application dictionary and traverse the application names in the second result file;
[0096] Step S430: Save the application names and corresponding cumulative computing resource usage that belong to the preset application dictionary but do not belong to platform applications in the second result file to the first result sub-file;
[0097] Step S440: Save the application names and corresponding cumulative computing resource usage of the application belonging to the preset application dictionary and platform application in the second result file to the second result sub-file, and replace the application names in the second result sub-file with the first preset name;
[0098] Step S450: Save the application names in the second result file that do not belong to the preset application dictionary to the third result sub-file;
[0099] Step S460: Create an application computing resource distribution map based on the first result file, the first result sub-file, the second result sub-file, and the third result sub-file.
[0100] This embodiment stores a preset application dictionary. Specifically, the preset application dictionary may include two columns: the first column is the application keyword, and the second column is the application name. When determining the application used by the running job, keyword matching (case-insensitive) is performed on the command information of the running job to obtain the corresponding application name by matching the keywords.
[0101] This embodiment categorizes applications on the supercomputer into four types. The first type consists of applications stored in a preset application dictionary, excluding platform applications such as MATLAB, Python, and Anaconda. This is because there is a distinction between platform applications and non-platform applications. Platform applications are generally widely used applications. By statistically analyzing the computing resource usage of platform applications, we can understand their usage patterns, but we cannot determine the discipline or field to which the application belongs.
[0102] The second category consists of platform applications stored in the pre-defined application dictionary, such as MATLAB, Python, and Anaconda. Statistics on this category can clearly identify which specific discipline or field consumes the most computing resources, allowing for increased investment in those applications based on their respective domains.
[0103] The third category is applications that are not in the preset application dictionary. In other words, the preset application dictionary generally stores common application names, but some running jobs do not match the application name. In this case, the last two columns of the application command line are selected as the application name, which does not exist in the preset application dictionary.
[0104] The fourth category includes node names, the first preset name, or custom applications. This means that in the job log information, a job might only be a node, in which case the application name will be displayed as the computation node; there might also be cases where there is no application name (i.e., the application name is empty), in which case the first preset name (such as "other") will be assigned; there might also be user-compiled application software, i.e., custom applications. These three types of application names cannot be found in the preset application dictionary. Some running jobs may not obtain application command-line information, for example, if the job is submitted and then quickly deleted by the user.
[0105] Finally, the information on these four types of application computing resources is output to a CSV file. Using WPS software on Windows, the final application computing resource distribution map is obtained.
[0106] This embodiment categorizes all application names in detail, making it convenient for users to intuitively query the status of various applications on the supercomputer within a certain time period.
[0107] In the first embodiment, step S460 specifically includes:
[0108] Step S461a: Obtain the first result file, the second result sub-file, and the third result sub-file, and classify the application names in the first result file, the second result sub-file, and the third result sub-file into the second preset name;
[0109] Step S462a: Add up the cumulative usage of computing resources corresponding to each of the second preset names to obtain the total usage of the second computing resources corresponding to the second preset name;
[0110] Step S463a: Generate an application computing resource distribution map based on the second total computing resource usage and the first result sub-file;
[0111] Specifically, the application computing resource distribution map includes: the computing resource percentage corresponding to each application name in the first result sub-file, and the computing resource percentage corresponding to the second preset name. It can be understood that the sum of all application names in the first result file, first result sub-file, second result sub-file, and third result sub-file represents the total computing resource usage; the computing resource percentage corresponding to an application name refers to the proportion of the cumulative computing resource usage of a certain application name to the total computing resource usage; the computing resource percentage corresponding to the second preset name refers to the proportion of the total second computing resource usage to the total computing resource usage. That is, when creating the computing resource distribution map in Windows, the application names in the first result file, second result sub-file, and third result sub-file are all replaced with the second preset name (e.g., "other"), directly displaying the total computing resource percentage corresponding to the second preset name, and also displaying the computing resource percentage of each application name in the first result sub-file. This embodiment displays the computing resource percentage of each application name in the first result sub-file in detail, while other application names are categorized into the second preset name, highlighting the computing resource percentage of applications of various disciplines or fields without changing the actual percentage, making the application computing resource distribution map more intuitive.
[0112] In the second embodiment, step S460 specifically includes:
[0113] Step S461b: Obtain the first result file, the first result sub-file, the second result sub-file, and the third result sub-file;
[0114] Step S462b: Sort the application names in the first result subfile according to the cumulative computing resource usage from largest to smallest to obtain the computing resource sorting result;
[0115] Step S463b: Obtain a preset number of application names based on the computing resource sorting results, use them as target application names, and classify the application names other than the target application names in the first result sub-file as a third preset name;
[0116] Step S464b: Classify the application names in the first result file, the second result sub-file, and the third result sub-file into the third preset name;
[0117] Step S465b: Add up the cumulative usage of computing resources corresponding to each of the third preset names to obtain the total usage of the third computing resources corresponding to the third preset name;
[0118] Step S466b: Generate an application computing resource distribution map based on the total usage of the third computing resources, the target application name, and the corresponding cumulative usage of computing resources.
[0119] Specifically, the application computing resource distribution map includes: the computing resource percentage corresponding to each of the target application names, and the computing resource percentage corresponding to the third preset name. When creating the computing resource distribution map in Windows, the application names in the first result file, the second result sub-file, and the third result sub-file are all replaced with the third preset name (or can be set to "other"). Furthermore, the application names with the highest cumulative computing resource usage in the first result sub-file are displayed as target application names, and those with lower cumulative computing resource usage are also categorized as the third preset name. For example, the first 8 application names in the first result sub-file are used as target application names and displayed in the application computing resource distribution map, while the others are categorized as the third preset name; that is, the application computing resource distribution map only displays 9 application names.
[0120] This embodiment displays the proportion of computing resources for each application name with high cumulative computing resource usage in the first result sub-file, while other application names are categorized into the third preset name. This highlights the applications with high computing resource usage without changing the actual proportion, making the application computing resource distribution map more intuitive.
[0121] The following specific embodiments are provided for illustration. Please refer to them. Figure 2 .
[0122] Step A1: Upload the task time information for the past seven days to the directory where the task program is located. The task will be executed every Monday.
[0123] Step A2: Traverse each job in the job core time information, obtain the job log information for the past ten days, find the application name, and save the job number, core time usage, and application name to the first record file;
[0124] Step A3: Traverse the first record file and check if the application name exists; if not, proceed to step A4; if yes, proceed to step A5.
[0125] Step A4: Assign the application name the value "other";
[0126] Step A5: Save the application name and core usage to the second record file;
[0127] Step A6: Deduplicate the application names in the second log file to obtain all application names selected from the log for this task;
[0128] Step A7: Iterate through all application names, sum up the core usage under the same application name, and obtain the cumulative core usage for each application name;
[0129] Step A8: Sort the data according to the cumulative core time usage to obtain the original application core time data sorted by core time usage;
[0130] Step A9: Determine if the application name is Node, Other, or Custom; if yes, proceed to Step A10; otherwise, proceed to Step A11.
[0131] Step A10: Save the application name and core usage to file 1; and then proceed to step A17.
[0132] Step A11: Save the application name and core usage to file 2;
[0133] Step A12: Traverse file 2 and determine if the application name is in the application dictionary; if not, proceed to step A13; if yes, proceed to step A14.
[0134] Step A13: Save the application name and core usage to the date-3.csv file; and then proceed to step A17;
[0135] Step A14: Determine if the application is a platform application; if not, proceed to step A15; if yes, proceed to step A16.
[0136] Step A15: Save the application name and core usage to the date-1.csv file; and then proceed to step A17;
[0137] Step A16: Replace the application name with "other", save the application name and core usage to the date-2.csv file; and then execute step A17.
[0138] Step A17: Output to the final result file.
[0139] Specifically, in this embodiment, core time is used as the unit of computing resources. Before 12:00 noon every Monday, the core time information of the past week is uploaded to the program running directory taiyi.csv. The sample content of the core time file is as follows: job number, username, user group name, number of cores for the job, submission time, start time, end time, core time for the job, and cost.
[0140] Execute the statistics script appData.sh:
[0141] [xx@login04~]#cd / xx / xx / jobidAndApp /
[0142] [xx@login04 jobidAndApp]#sh appData.sh
[0143] The contents of appData.sh include:
[0144] Iterate through each job in taiyi.csv and find the application name in the job logs collected over the past ten days (to avoid missing information). Save the job number, core usage, and application name to the file jodidAppname. For example, in the job log information recorded on November 4, 2022, the corresponding columns would be: job number, username, node name, and application.
[0145] Iterate through `jodidAppname`, checking if the application name exists; if not, assign the application name the value "other". Save the application name and core usage records to a file `finalLin`. Remove duplicate application names from the `finalLin` file to obtain all application names selected from the logs for this task. Iterate through all application names from the previous step, summing the corresponding core usage records in the `finalLin` file to obtain the application name and the cumulative core usage.
[0146] Sort the data according to the core time size to obtain the original application core time data file sorted by core time.
[0147] Application names are categorized. The application dictionary `keyWordApp` includes two columns: the first column is the application keyword, and the second column is the application name. The first category includes applications listed in `keyWordApp`, but not platform-based applications like MATLAB, Python, or Anaconda. The second category includes platform-based applications listed in the dictionary for core time statistics, such as MATLAB, Python, and Anaconda; these applications are replaced with "other". The third category includes applications not listed in the dictionary; the last two columns of the application command line are used as the application name for core time statistics. The fourth category includes node names, "other", or custom applications.
[0148] The date-1.csv file is generated by executing modifyData.sh, which is executed within the script appData.sh. The contents of this file are the applications and computing resources (excluding MATLAB, Zibian, Python, etc.) that exist in the keyWordApp file.
[0149] The date-2.csv file is generated by executing modifyData.sh, which is executed within the script appData.sh. The content of this file consists of the computing resource files for platform applications such as MATLAB, R, Python, Anaconda, Miniconda, and Orca that exist in the keyWordApp file. These applications are replaced with "other" to generate the file.
[0150] The date-3.csv file is generated by executing modifyData.sh, which is executed within the script appData.sh. This file contains files saved by applications that do not exist in the keyWordApp file. For example, a user-compiled application might have an executable named Allrun, which does not exist in the application dictionary.
[0151] File 1 is generated by executing modifyData.sh, which is executed within the script appData.sh. A small number of applications whose application names are the same as the names of compute nodes are stored in file 1.
[0152] Finally, the information on these four types of application computing resources is output to a CSV file. Using WPS Office on Windows, the final application computing resource distribution map is obtained, as shown below. Figure 3 As shown in the figure, in the past week, other, VASP, and Gaussian applications accounted for more than 10% of computing resources, while LAMMPS and TURBT applications accounted for 6%.
[0153] In one embodiment, such as Figure 4 As shown, based on the above-mentioned method for statistically analyzing the distribution of application computing resources in a supercomputer, the present invention also provides a corresponding device for statistically analyzing the distribution of application computing resources in a supercomputer, comprising:
[0154] The identification module 100 is used to acquire job computing resource information and identify the job number and the computing resource usage corresponding to the job number in the job computing resource information.
[0155] The lookup module 200 is used to obtain job log information, search the job log information according to the job number, and obtain the application name that matches the job number.
[0156] The statistics module 300 is used to count the computing resource usage corresponding to job numbers that match the same application name, and to obtain the cumulative computing resource usage for each application name.
[0157] The generation module 400 is used to classify the application names according to preset classification rules and generate an application computing resource distribution map according to the cumulative usage of computing resources corresponding to the classified application names.
[0158] In one embodiment, the present invention also provides a terminal, such as... Figure 5 As shown, it includes: a memory 20, a processor 10, and a supercomputer application computing resource distribution statistics program 30 stored on the memory 20 and executable on the processor 10. When the supercomputer application computing resource distribution statistics program 30 is executed by the processor 10, it implements the steps of the supercomputer application computing resource distribution statistics method as described above.
[0159] The present invention also provides a computer-readable storage medium storing a computer program that can be executed to implement the steps of the supercomputer application computing resource distribution statistics method as described above.
[0160] In summary, this invention discloses a method and apparatus for statistically analyzing the distribution of application computing resources in a supercomputer. The method includes: acquiring job computing resource information, identifying the job number and the computing resource usage corresponding to the job number in the job computing resource information; acquiring job log information, searching the job log information according to the job number to obtain the application name matching the job number; statistically analyzing the computing resource usage corresponding to job numbers matching the same application name to obtain the cumulative computing resource usage for each application name; classifying the application names according to a preset classification rule, and generating an application computing resource distribution map based on the cumulative computing resource usage corresponding to the classified application names. This invention, by statistically analyzing the computing resource usage corresponding to each application name and generating an application computing resource distribution map according to the classification results, allows for a direct and intuitive view of the supercomputer's application status.
[0161] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.
Claims
1. A method for statistically analyzing the distribution of application computing resources in a supercomputer, characterized in that, The method includes: Obtain job computing resource information, and identify the job number and the computing resource usage corresponding to the job number in the job computing resource information; Obtain job log information, search for the job log information based on the job number, and obtain the application name that matches the job number; The computing resource usage corresponding to job numbers that match the same application name is statistically analyzed to obtain the cumulative computing resource usage for each application name; The application names are classified according to preset classification rules, and an application computing resource distribution map is generated based on the cumulative computing resource usage corresponding to the classified application names. The application names are categorized according to preset classification rules, and an application computing resource distribution map is generated based on the cumulative computing resource usage corresponding to the categorized application names, including: Save the application name and corresponding cumulative computing resource usage belonging to the computing node, the first preset name, or the self-developed application to the first result file, and save the application name and corresponding cumulative computing resource usage not belonging to the computing node, the first preset name, or the self-developed application to the second result file. Obtain the preset application dictionary and iterate through the application names in the second result file; Save the application names and corresponding cumulative computing resource usage that belong to the preset application dictionary but do not belong to the platform application in the second result file to the first result subfile; Save the application names and corresponding cumulative computing resource usage that belong to the preset application dictionary and platform application in the second result file to the second result sub-file, and replace the application names in the second result sub-file with the first preset name; Save the application names in the second result file that do not belong to the preset application dictionary to the third result sub-file; Create an application computing resource distribution map based on the first result file, the first result sub-file, the second result sub-file, and the third result sub-file.
2. The method for statistically analyzing the distribution of application computing resources in a supercomputer according to claim 1, characterized in that, The step of obtaining job computing resource information and identifying the job number and the computing resource usage corresponding to the job number includes: Upload the task calculation resource information for the first preset time period to the directory where the task is located; Identify the job number in the directory where the task is located and the computing resource usage corresponding to the job number.
3. The method for statistically analyzing the distribution of application computing resources in a supercomputer according to claim 2, characterized in that, The step of obtaining job log information, and searching for the job log information based on the job number to obtain the application name matching the job number, includes: Obtain the operation log information within the second preset time period; Iterate through the job numbers in the directory where the task is located, and find the application name that matches each job number in the job log information; Save the job number, the computing resource usage corresponding to the job number, and the application name matching the job number to the first record file; The first preset time period falls within the second preset time period.
4. The method for statistically analyzing the distribution of application computing resources in a supercomputer according to claim 3, characterized in that, After saving the job number, the computing resource usage corresponding to the job number, and the application name matching the job number to the first record file, the method further includes: If the first record file contains an application name that matches the job number but is empty, then the application name of the job number will be assigned the first preset name.
5. The method for statistically analyzing the distribution of application computing resources in a supercomputer according to claim 1, characterized in that, The step of statistically analyzing the computing resource usage corresponding to job numbers matching the same application name to obtain the cumulative computing resource usage for each application name includes: Save each application name and its corresponding computing resource usage to the second record file; The same application names in the second record file are deduplicated, and the computing resource usage corresponding to each application name after deduplication is accumulated to obtain the cumulative computing resource usage corresponding to each application name.
6. The method for statistically analyzing the distribution of application computing resources in a supercomputer according to claim 1, characterized in that, After calculating the cumulative computing resource usage for each application name by tracking the computing resource usage of job numbers matched with the same application name, the process further includes: The application names are sorted according to the cumulative usage of computing resources to obtain the original application computing resource data.
7. The method for statistically analyzing the distribution of application computing resources in a supercomputer according to claim 1, characterized in that, Based on the first result file, the first result sub-file, the second result sub-file, and the third result sub-file, an application computing resource distribution map is created, including: Obtain the first result file, the second result sub-file, and the third result sub-file, and classify the application names in the first result file, the second result sub-file, and the third result sub-file into a second preset name; The cumulative usage of computing resources corresponding to each of the second preset names is added together to obtain the total usage of the second computing resources corresponding to the second preset name; Generate an application computing resource distribution map based on the second total computing resource usage and the first result sub-file; The application computing resource distribution map includes: the computing resource percentage corresponding to each application name in the first result sub-file, and the computing resource percentage corresponding to the second preset name.
8. The method for statistically analyzing the distribution of application computing resources in a supercomputer according to claim 1, characterized in that, Based on the first result file, the first result sub-file, the second result sub-file, and the third result sub-file, an application computing resource distribution map is created, including: Obtain the first result file, the first result sub-file, the second result sub-file, and the third result sub-file; Sort the application names in the first result subfile according to the cumulative computing resource usage from largest to smallest to obtain the computing resource ranking result; Based on the sorting results of the computing resources, a preset number of application names are obtained as target application names, and the application names other than the target application names in the first result sub-file are classified as a third preset name. The application names in the first result file, the second result sub-file, and the third result sub-file are all classified as the third preset name; The cumulative usage of computing resources corresponding to each of the third preset names is added together to obtain the total usage of the third computing resources corresponding to the third preset name. Based on the total usage of the third computing resources and the target application name and its corresponding cumulative computing resource usage, an application computing resource distribution map is generated. The application computing resource distribution map includes: the proportion of computing resources corresponding to each of the target application names, and the proportion of computing resources corresponding to the third preset name.
9. A device for statistically analyzing the distribution of application computing resources in a supercomputer, characterized in that, The device includes: The identification module is used to acquire job computing resource information and identify the job number and the computing resource usage corresponding to the job number in the job computing resource information. The search module is used to obtain job log information, search the job log information according to the job number, and obtain the application name that matches the job number. The statistics module is used to count the computing resource usage corresponding to job numbers that match the same application name, and to obtain the cumulative computing resource usage for each application name. The generation module is used to classify the application names according to preset classification rules and generate an application computing resource distribution map based on the cumulative computing resource usage corresponding to the classified application names. The application names are categorized according to preset classification rules, and an application computing resource distribution map is generated based on the cumulative computing resource usage corresponding to the categorized application names, including: Save the application name and corresponding cumulative computing resource usage belonging to the computing node, the first preset name, or the self-developed application to the first result file, and save the application name and corresponding cumulative computing resource usage not belonging to the computing node, the first preset name, or the self-developed application to the second result file. Obtain the preset application dictionary and iterate through the application names in the second result file; Save the application names and corresponding cumulative computing resource usage that belong to the preset application dictionary but do not belong to the platform application in the second result file to the first result subfile; Save the application names and corresponding cumulative computing resource usage that belong to the preset application dictionary and platform application in the second result file to the second result sub-file, and replace the application names in the second result sub-file with the first preset name; Save the application names in the second result file that do not belong to the preset application dictionary to the third result sub-file; Create an application computing resource distribution map based on the first result file, the first result sub-file, the second result sub-file, and the third result sub-file.
10. A terminal, characterized in that, include: The supercomputer includes a memory, a processor, and a supercomputer application computing resource distribution statistics program stored in the memory and executable on the processor. When the supercomputer application computing resource distribution statistics program is executed by the processor, it implements the steps of the supercomputer application computing resource distribution statistics method as described in any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that can be executed to implement the steps of the application computing resource distribution statistics method for supercomputers as described in any one of claims 1 to 8.