Accelerator card resource monitoring management method and device, chip and storage medium
Patent Information
- Application Number
- CN202611142006.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-30
- Publication Date
- 2026-09-04
AI Technical Summary
[0004]本申请实施例提供一种加速卡资源监控管理方法、装置、芯片及存储介质,可以解决相关技术中存在的不同系统接口下设备标识难以统一、运行状态分散以及故障定位困难的技术问题
[0010] Fifthly, embodiments of this application provide a chip including a processor coupled to a transceiver for executing the technical solution provided in the first aspect of this application. In one possible design, the chip can also be a dedicated hardware structure for implementing the technical solution provided in the first aspect.
Smart Images

Figure CN122691908A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence hardware management, and more specifically to an accelerator card resource monitoring and management method, device, chip, and storage medium. Background Technology
[0002] As the scale of deep learning models and the amount of training data continue to grow, the requirements for computing resources, storage bandwidth, and inter-device communication efficiency in artificial intelligence computing tasks are further increasing. Deep learning accelerator cards (DLCs), as an important computing component in artificial intelligence servers, have been widely used in scenarios such as model training, inference, and multi-card collaborative computing.
[0003] In multi-GPU deployment environments, a single host typically mounts multiple deep learning accelerator cards. Each deep learning accelerator card connects to the host system via the PCIe bus, a peripheral interconnect standard, and can form a specific multi-GPU communication topology through the Inter-Chip Connect (ICC) interface to support high-speed data exchange and aggregated communication during parallel training. To ensure stable system operation, maintenance personnel typically need to monitor various aspects such as device identification, link status, resource usage, power consumption and temperature, firmware version, process usage, and fault status. However, in multi-GPU scenarios, existing accelerator card management methods may present inconsistent device identifiers, device numbers, and sorting methods across different system interfaces or management tools. Maintenance personnel need to manually verify the specific hardware devices, which is complex and prone to location errors. Therefore, existing accelerator card management methods suffer from technical problems such as difficulty in unifying device identifiers across different system interfaces, scattered operating statuses, and difficulties in fault location, necessitating a new solution. Summary of the Invention
[0004] This application provides an accelerator card resource monitoring and management method, device, chip, and storage medium, which can solve the technical problems in related technologies such as difficulty in unifying device identification under different system interfaces, scattered operating status, and difficulty in fault location.
[0005] In a first aspect, embodiments of this application provide a method for monitoring and managing accelerator card resources, the method comprising: Obtain first device identification information from the target computing device to represent the physical connection relationship of the accelerator card, and obtain second device identification information to represent the logical registration relationship of the accelerator card; Based on the format-normalized second device identification information and the first device identification information, a mapping relationship between the physical device identification and logical device index of the accelerator card is established to obtain the accelerator card device list, and the accelerator cards in the accelerator card device list are associated with device identity information. Monitor the operating status data of each accelerator card in the accelerator card device list; based on the data source of the operating status data, perform resource status analysis, hot status estimation, interconnection status conversion and abnormal status identification respectively, use the mapping relationship as the correlation benchmark between the operating status data corresponding to different data sources, collect the resource status analysis results, hot status estimation results and abnormal status identification results into the device status record associated with the corresponding logical device index, and associate the interconnection status conversion results with the logical device indexes at both ends of the corresponding interconnection link to obtain the comprehensive status information corresponding to each accelerator card; The comprehensive status information is bound to the logical device index of the corresponding accelerator card, and the accelerator card device list is updated according to the abnormal status identification results in the comprehensive status information to generate accelerator card resource monitoring and management results.
[0006] Secondly, embodiments of this application provide an accelerator card resource monitoring and management device, which has functions corresponding to the accelerator card resource monitoring and management method provided in the first aspect above. These functions can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions, and these modules can be software and / or hardware.
[0007] In one embodiment, the accelerator card resource monitoring and management device includes: The acquisition module is configured to acquire first device identification information representing the physical connection relationship of the accelerator card from the target computing device, and to acquire second device identification information representing the logical registration relationship of the accelerator card. The mapping module is configured to establish a mapping relationship between the physical device identifier and the logical device index of the accelerator card based on the format-normalized second device identifier information and the first device identifier information, to obtain an accelerator card device list, and associate the accelerator cards in the accelerator card device list with device identity information. The monitoring module is configured to monitor the operating status data of each accelerator card in the accelerator card device list; based on the data source of the operating status data, it performs resource status analysis, hot status estimation, interconnection status conversion and abnormal status identification respectively; using the mapping relationship as the correlation benchmark between the operating status data corresponding to different data sources, it collects the resource status analysis results, hot status estimation results and abnormal status identification results into the device status record associated with the corresponding logical device index, and associates the interconnection status conversion results with the logical device indexes at both ends of the corresponding interconnection link to obtain the comprehensive status information corresponding to each accelerator card; The management module is configured to bind the comprehensive status information with the logical device index of the corresponding accelerator card, and update the accelerator card device list based on the abnormal status identification results in the comprehensive status information, thereby generating accelerator card resource monitoring and management results.
[0008] Thirdly, embodiments of this application provide a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the accelerator card resource monitoring and management method as described in the first aspect.
[0009] Fourthly, embodiments of this application provide a computing device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the accelerator card resource monitoring and management method described in the first aspect.
[0010] Fifthly, embodiments of this application provide a chip including a processor coupled to a transceiver for executing the technical solution provided in the first aspect of this application. In one possible design, the chip can also be a dedicated hardware structure for implementing the technical solution provided in the first aspect.
[0011] Compared to existing technologies, this application's embodiments, by acquiring the physical connection relationships and logical registration relationships of accelerator cards and normalizing the format of device identification information from different sources, establish a mapping relationship between physical device identifiers and logical device indexes. This allows the accelerator card device list to simultaneously reflect hardware connection locations and software logical numbers, thereby reducing the operational costs and positioning errors associated with manual device verification in multi-card environments. Furthermore, by performing resource status parsing, hot status estimation, interconnection status conversion, and abnormal status identification according to the data sources of the operating status data, resource usage, hot status, interconnection link status, and abnormal status can be associated and managed under a unified logical device index, thereby improving the completeness and consistency of accelerator card operating status monitoring. Finally, by binding comprehensive status information with the logical device index of the corresponding accelerator card and updating the accelerator card device list based on the abnormal status identification results, abnormal devices can be identified in a timely manner and matched with specific logical devices, thereby improving the fault location efficiency, operation and maintenance efficiency, and system reliability of multi-card accelerator card clusters. Attached Figure Description
[0012] The objectives, features, and advantages of the embodiments of this application will become readily understood by referring to the accompanying drawings and the detailed description of the embodiments. Wherein: Figure 1 This is a flowchart illustrating the accelerator card resource monitoring and management method in this application embodiment; Figure 2This is a schematic diagram of the format alignment and bidirectional mapping process in an embodiment of this application; Figure 3 This is a schematic diagram of the overall system architecture of the accelerator card resource monitoring and management method according to an embodiment of this application; Figure 4 This is a schematic diagram of the interconnection topology of the accelerator card resource monitoring and management method according to an embodiment of this application; Figure 5 This is a schematic diagram of the comprehensive status information panel output of the accelerator card resource monitoring and management method according to an embodiment of this application; Figure 6 This is a schematic diagram of the structure of the accelerator card resource monitoring and management device according to an embodiment of this application; Figure 7 This is a schematic diagram of the structure of a computing device according to an embodiment of this application; Figure 8 This is a schematic diagram of a server structure in one embodiment of this application. Detailed Implementation
[0013] This application provides an accelerator card resource monitoring and management method and apparatus, which can be applied to accelerator card computing systems in scenarios such as artificial intelligence model training, model inference, multi-card collaborative computing, and accelerator card operation and maintenance. The accelerator card computing system may include a target computing device, a resource monitoring and management device, and at least one accelerator card installed in the target computing device. The target computing device can connect to each accelerator card via a Peripheral Component Interconnect Express (PCIe) bus. Multiple accelerator cards can also form an interconnect topology through an Inter-Chip Connect (ICC) interface for data exchange or aggregated communication.
[0014] The accelerator card can be a hardware acceleration device for performing artificial intelligence computing tasks, such as a Deep Learning Controller (DLC). The accelerator card can include one or more processor units, which can be a Tensor Processing Unit (TPU) or other processors used to perform matrix operations, tensor operations, and neural network calculations. The accelerator card may also include high-bandwidth memory (HBM), a power management chip, device firmware, a PCIe interface, and high-speed interconnect ports between devices.
[0015] The accelerator card resource monitoring and management device is the executing entity of the accelerator card resource monitoring and management method provided in this application embodiment. It can be deployed on the target computing device as a resource monitoring and management program, command-line tool, or system service, or it can be deployed on a management device that is communicatively connected to the target computing device. The resource monitoring and management device can obtain the device identification information and operating status data of each accelerator card by calling the virtual file system, device enumeration interface, sensor acquisition program, firmware information program, and link detection program of the target computing device through the System Management Interface (SMI).
[0016] The device identification information involved in this application embodiment includes, but is not limited to: first device identification information for indicating the physical connection relationship of accelerator cards, and second device identification information for indicating the logical registration relationship of accelerator cards. Since the device identifications provided by different system interfaces may use different address formats, the resource monitoring and management device can normalize the format of the corresponding device identifications, establish a mapping relationship between physical device identifications and logical device indexes, generate an accelerator card device list, and associate each accelerator card with its corresponding device serial number and other device identity information.
[0017] The operational status data involved in this application embodiment may include storage resource data, power consumption data, power domain operating parameters, clock data, firmware version data, interconnect detection data, and abnormal device enumeration data. The accelerator card resource monitoring and management device can perform resource status parsing, thermal status calculation, interconnect status transition, and abnormal status identification according to the source of the operational status data, and use the mapping relationship as the association benchmark between different data sources to aggregate the corresponding processing results into the device status record associated with the corresponding logical device index. The interconnect status transition result can be associated with the logical device indexes at both ends of the corresponding interconnect link.
[0018] In the current accelerator card management process, different system interfaces or management tools may use inconsistent device identifiers, logical numbers, and data output formats. Maintenance personnel typically need to manually verify the device information and operating status of the same accelerator card under different interfaces. When multiple accelerator cards are configured in a target computing device, device mismatches are prone to occur, and the process of locating abnormal devices and abnormal interconnect links is quite complex.
[0019] This application embodiment normalizes the format of device identifiers provided by different system interfaces, establishes a mapping relationship between physical device identifiers and logical device indexes, and aggregates operational status data from different sources based on this mapping relationship. After identifying an abnormal accelerator card, the accelerator card device list can be updated according to the logical device index of the abnormal accelerator card, generating accelerator card resource monitoring and management results including device identity, resource status, hot status, interconnection status, and abnormal status, thereby reducing manual verification of different device identifiers and multi-source operational status data.
[0020] Reference Figure 1 , Figure 1 This is a flowchart illustrating a method for monitoring and managing accelerator card resources, provided in an embodiment of this application. The method can be executed by an accelerator card resource monitoring and management device, and includes steps 101-105: Step 101: Obtain first device identification information from the target computing device to represent the physical connection relationship of the accelerator card, and obtain second device identification information to represent the logical registration relationship of the accelerator card.
[0021] In this embodiment, the target computing device can be a computing device equipped with one or more accelerator cards and capable of running accelerator card drivers and resource monitoring and management programs. For example, it can be a computing node in an artificial intelligence server, training server, inference server, or server cluster. The accelerator card can be a deep learning accelerator card, which can be equipped with a tensor processor, high-bandwidth memory, power management devices, peripheral component interconnection standard interfaces, and high-speed interconnect ports between devices.
[0022] Here, the physical connection relationships of the accelerator cards are used to reflect their connection or enumeration location in the target computing device's bus system. The first device identification information can be device address information provided by the target computing device through the relevant system interface of the Peripheral Component Interconnect Express (PCIe), such as the Bus Device Function (BDF) address. Different accelerator cards have different physical device identifiers, thus allowing differentiation between different physical accelerator cards in the target computing device based on the first device identification information.
[0023] The logical registration relationship of the accelerator card reflects the logical device number used by the accelerator card driver when registering the accelerator card in the operating system. The second device identification information may include device address information obtained from logical device nodes in the virtual file system. Logical device nodes can use sequentially numbered node names, with each node corresponding to a logical device index. For example, logical device indices can be sequentially numbered 0, 1, 2, etc., used to reference the corresponding accelerator card in the driver, device file, and resource monitoring management program.
[0024] The first device identification information and the second device identification information can be provided by different system interfaces. Since different system interfaces may use different address fields, character formats, or domain prefixes, the device identification information presented by the same accelerator card under the two interfaces may not be directly matched with strings. Therefore, step 101 obtains the two types of device identification information separately for subsequent establishment of the correspondence between physical accelerator cards and logical device indexes.
[0025] As an optional embodiment, step 101, obtaining first device identification information representing the physical connection relationship of accelerator cards and second device identification information representing the logical registration relationship of accelerator cards from the target computing device, includes: traversing the logical device nodes corresponding to each accelerator card in the virtual file system of the target computing device; reading the bus address information of the corresponding accelerator card from the logical device node and using the bus address information as the first device identification information; calling a bus enumeration tool to obtain the enumeration output information of each accelerator card in the target computing device; extracting the enumeration address information of the corresponding accelerator card from the enumeration output information and using the enumeration address information as the second device identification information; wherein, the first device identification information and the second device identification information respectively correspond to the device identification presented by the same accelerator card under different system interfaces.
[0026] Specifically, in step 101, the logical device nodes corresponding to each accelerator card in the virtual file system of the target computing device can be traversed, and the bus address information of the corresponding accelerator card can be read from each logical device node. The read bus address information is used as the first device identification information. The logical device nodes are registered in the operating system by the accelerator card driver, and different logical device nodes correspond to different logical device indices. For example, logical device nodes dlc0, dlc1 to dlcN can be traversed sequentially, and the bus address file under each logical device node can be read. For logical device node dlc0, the read bus address information can be 0000:4f:00.0, which is a complete bus device function address containing the domain identifier, bus identifier, device identifier, and function identifier.
[0027] Simultaneously, the bus enumeration tool in the target computing device can be invoked to enumerate each accelerator card connected to the target computing device, obtaining enumeration output information. The enumeration output information may include enumeration address, device type, and device description. The resource monitoring and management program can scan the enumeration output information line by line, filter the device records corresponding to the accelerator cards, and extract the enumeration address information from the corresponding device records, using this enumeration address information as the second device identification information. For example, the enumeration address information output by the bus enumeration tool for the same accelerator card mentioned above could be 4f:00.0, which includes the bus identifier, device identifier, and function identifier, but does not include the domain identifier.
[0028] The first and second device identification information come from logical device nodes in the virtual file system and the bus enumeration tool, respectively, representing the device identification presented by the same accelerator card through different system interfaces. In the example above, the first device identification information 0000:4f:00.0 and the second device identification information 4f:00.0 correspond to the same accelerator card; the main difference between them is whether or not a domain identifier is included. After obtaining the two types of device identification information, the corresponding device identifiers can be formatted and matched in subsequent steps to determine the correspondence between physical devices and logical device nodes in the bus enumeration results.
[0029] By obtaining the device identifier of the same accelerator card from the logical device node and the bus enumeration information respectively, the identification basis of the device under the driver registration interface and the bus enumeration interface can be retained, providing a data foundation for establishing the mapping relationship between physical device identifiers and logical device indexes in the future, and reducing device correspondence errors caused by the two system interfaces using different address formats or device arrangement orders.
[0030] Step 102: Based on the format-normalized second device identification information and the first device identification information, establish a mapping relationship between the physical device identification and logical device index of the accelerator card to obtain the accelerator card device list, and associate the accelerator cards in the accelerator card device list with device identity information.
[0031] Format normalization refers to adjusting the format of device identifiers provided by different system interfaces so that the device identifiers to be matched have the same address field composition and representation. For example, a BDF address provided by one system interface may include a domain identifier, while a BDF address provided by another system interface may not. In this case, by adding a domain prefix, standardizing the character format, or adjusting the address field format, the second device identifier information can be matched with the first device identifier information.
[0032] The physical device identifier is used to determine the physical enumeration location of the accelerator card in the target computing device bus system, and the logical device index is used to determine the logical number of the accelerator card in the driver and resource monitoring management program. The mapping relationship is used to record the correspondence between the two. For example, the mapping relationship can indicate that one BDF address corresponds to logical device index 0, and another BDF address corresponds to logical device index 1. The resource monitoring management program can thus find the logical device index based on the physical device identifier, and can also determine the physical device identifier of the corresponding accelerator card based on the logical device index.
[0033] The accelerator card device list is a set of device records based on the above mapping relationship. Each device record can correspond to one accelerator card and includes at least the physical device identifier and logical device index of that accelerator card. In some implementations, the device records can also be sorted according to the logical device index, and bus connection status such as PCIe link speed and PCIe link width can be written into the device records.
[0034] Device identification information is a relatively stable hardware identifier used to distinguish different physical accelerator cards, such as a device serial number or product number. The logical device index may be affected by changes in device enumeration order, driver loading process, or system configuration, while the device serial number usually remains constant with the physical accelerator card. Associating accelerator cards in the device list with their device identification information allows maintenance personnel to further identify specific physical accelerator cards based on the logical device index, and also facilitates the location and replacement of abnormal accelerator cards.
[0035] As an optional embodiment, in step 102, based on the format-normalized second device identifier information and the first device identifier information, a mapping relationship between the physical device identifier and the logical device index of the accelerator card is established to obtain the accelerator card device list. This includes: extracting the index of the first device identifier information to obtain the logical device index corresponding to each first device identifier information, and establishing a first mapping table, which is used to represent the correspondence between the first device identifier information and the logical device index; performing format normalization processing on the second device identifier information so that the second device identifier information and the first device identifier information have the same address format; matching the format-normalized second device identifier information with the first device identifier information in the first mapping table to determine the logical device index corresponding to the second device identifier information; and sorting the successfully matched accelerator cards according to the order of the logical device indexes to generate the accelerator card device list.
[0036] In the above embodiments, a logical device index can be extracted based on the node identifier of the logical device node where the first device identifier information is located, and the first device identifier information can be associated with the corresponding logical device index to establish a first mapping table. For example, when the first device identifier information 0000:4f:00.0 is read from the logical device node dlc0, the logical device index 0 can be extracted from the node identifier dlc0, and the correspondence between 0000:4f:00.0 and the logical device index 0 can be recorded in the first mapping table; when the first device identifier information 0000:52:00.0 is read from the logical device node dlc1, the correspondence between the first device identifier information and the logical device index 1 can be recorded accordingly. Thus, the first mapping table uses the first device identifier information as the query key and the logical device index as the corresponding mapping value.
[0037] The second device identifier information extracted from the bus enumeration information can be normalized to have the same address format as the first device identifier information. For example, if the second device identifier information is a bus device function address 4f:00.0 without a field prefix, a preset field prefix can be added to it to obtain the normalized second device identifier information 0000:4f:00.0. It should be noted that the format normalization process is not limited to adding a field prefix; it can also include unifying the address field composition or character representation. The specific processing method can be determined based on the format differences between the two types of device identifier information.
[0038] Subsequently, using the normalized second device identifier information as the query key, the system searches for identical first device identifier information in the first mapping table. Taking 0000:4f:00.0 as an example, when identical first device identifier information exists in the first mapping table, its corresponding logical device index 0 is determined as the logical device index corresponding to the original second device identifier information 4f:00.0. By processing each bus enumeration information in the same way, a mapping relationship between the physical device identifiers and logical device indices of each accelerator card can be established. Second device identifier information that fails to match a record in the first mapping table can be excluded from the set of successfully matched accelerator cards, and the corresponding matching error information can be retained.
[0039] After matching is complete, the successfully matched accelerator cards can be sorted according to their logical device indices to generate an accelerator card device list. For example, if the output order of the bus enumeration tool is 52:00.0 and 4f:00.0, which correspond to logical device indices 1 and 0 respectively, the generated accelerator card device list will still be arranged according to logical device indices 0 and 1. Each device record in the accelerator card device list can include a physical device identifier, a logical device index, and corresponding bus address information for use in subsequent status acquisition and device management processes.
[0040] The above processing can eliminate the address format differences between different system interfaces and avoid the impact of changes in the bus enumeration order on the logical arrangement of the accelerator cards. This ensures that the physical devices obtained by bus enumeration can be stably mapped to the logical device index registered by the driver, providing a unified device association basis for the subsequent collection of the operating status data of each accelerator card according to the logical device index.
[0041] Further optionally, in the above embodiments, matching the format-normalized second device identifier information with the first device identifier information in the first mapping table to determine the logical device index corresponding to the second device identifier information includes: extracting the bus device function address without a field prefix from the enumeration output information; determining whether the bus device function address and the first device identifier information have the same address field; if the bus device function address does not contain the field identifier of the first device identifier information, then adding a preset field prefix to the bus device function address to obtain the format-normalized second device identifier information; using the format-normalized second device identifier information as the query key, performing a search operation in the first mapping table; if the first device identifier information that is consistent with the format-normalized second device identifier information is found in the first mapping table, then determining the logical device index corresponding to the first device identifier information as the logical device index corresponding to the second device identifier information.
[0042] In the above embodiments, the enumeration output information of the bus enumeration tool can be parsed line by line to filter the device records corresponding to the accelerator card, and the bus device function address without the field prefix can be extracted from the device records. The bus device function address can include a bus identifier, a device identifier, and a function identifier. For example, from the enumeration output information 4f:00.0 AIaccelerator controller, 4f:00.0 can be extracted as the second device identifier information.
[0043] Next, the extracted bus device function address is compared with the address field of the first device identification information to determine whether they have the same address field. The first device identification information is, for example, the complete address 0000:4f:00.0 read from the logical device node, including the domain identifier, bus identifier, device identifier, and function identifier. When the second device identification information 4f:00.0 does not contain the corresponding domain identifier, a preset domain prefix 0000: can be added to it to obtain the format-normalized second device identification information 0000:4f:00.0. When the second device identification information already contains a domain identifier and has the same address field as the first device identification information, it can be directly used for subsequent matching.
[0044] For example, it can be determined whether the extracted bus device function address and the first device identification information have the same address field by parsing the address field and comparing the number of fields. Specifically, the two types of address information can be separated by colons and periods to identify the domain identifier, bus identifier, device identifier, and function identifier respectively, and the field types and number of fields contained in the two can be compared.
[0045] For example, if the first device identifier is 0000:4f:00.0, after parsing according to the address separator, we can obtain the domain identifier 0000, bus identifier 4f, device identifier 00, and function identifier 0. The extracted bus device function address is 4f:00.0, which, after parsing, only yields the bus identifier 4f, device identifier 00, and function identifier 0. By comparison, it can be determined that this bus device function address lacks the domain identifier from the first device identifier, but the bus identifier, device identifier, and function identifier are consistent. In this case, a domain prefix corresponding to the first device identifier can be added to the bus device function address to obtain the complete address 0000:4f:00.0.
[0046] In another implementation, a preset matching rule corresponding to the complete bus device function address format can be used to perform format matching on the extracted address. The preset matching rule can be used to detect whether the address sequentially contains a domain field, a bus field, a device field, and a function field. If the matching result shows that the extracted address only contains the bus field, device field, and function field, it is determined that it lacks a domain field; if it contains all of the above fields, it is determined that it has the same address field composition as the first device identification information.
[0047] Furthermore, the composition of the address field can be determined based on the number or position of the delimiters. For example, the full address 0000:4f:00.0 contains two colons separating the field field, bus field, and device field, while the short address 4f:00.0 contains only one colon. Therefore, the number of colons can be used to determine whether the extracted bus device function address contains a field identifier. In contrast, using field parsing or preset matching rules can simultaneously verify the format of each field, making it suitable for scenarios where the enumerated output may contain other text content.
[0048] Subsequently, using the format-normalized second device identifier information as the lookup key, the first device identifier information with the same address is searched in the first mapping table. For example, if the first mapping table records the correspondence between 0000:4f:00.0 and logical device index 0, then after finding this address, logical device index 0 is determined as the logical device index corresponding to the second device identifier information 4f:00.0. For other enumerated addresses, the same format normalization and mapping lookup method can be used to determine the logical device index corresponding to each enumerated address.
[0049] For example, the first mapping table also records the correspondence between 0000:52:00.0 and logical device index 1, and 0000:57:00.0 and logical device index 2. For the bus device function address 52:00.0 extracted from the enumeration output information, a field prefix of 0000: can be added to obtain the format-normalized second device identifier information 0000:52:00.0, and this address is used as the lookup key in the first mapping table. After finding a matching first device identifier, its corresponding logical device index 1 is determined as the logical device index corresponding to the enumeration address 52:00.0. Similarly, after format normalization and mapping lookup of the enumeration address 57:00.0, its corresponding logical device index can be determined to be 2.
[0050] When the bus enumeration tool outputs the devices in the order of 57:00.0, 4f:00.0, and 52:00.0, after the above format normalization and mapping lookup, the corresponding logical device indices can be determined to be 2, 0, and 1, respectively. Subsequently, the corresponding device records can be organized according to the order of logical device indices 0, 1, and 2, instead of directly using the output order of the bus enumeration tool, thereby maintaining consistency between the accelerator card device list and the logical registration order in the driver.
[0051] Finally, by completing the missing address fields and performing mapping lookup based on a unified address format, the differences in device address formats between the bus enumeration interface and logical device nodes can be eliminated. This ensures that the physical devices obtained from bus enumeration accurately correspond to the logical device index registered by the driver, reducing device matching errors caused by different address formats or device enumeration orders. It also provides a reliable device association basis for subsequent collection and aggregation of operational status data according to the logical device index.
[0052] Reference Figure 2 , Figure 2 This is a schematic diagram of a PCIe BDF format alignment and bidirectional mapping process provided in an embodiment of this application. Figure 2 The left side shows the logical device nodes on the virtual file system side and their corresponding complete BDF addresses. For example, the complete BDF address corresponding to logical device node dlc0 is 0000:4f:00.0, and the complete BDF address corresponding to logical device node dlc1 is 0000:52:00.0. Based on the logical device nodes and their complete BDF addresses, a busToDlc mapping table can be constructed with the complete BDF address as the lookup key and the logical device index as the mapping value.
[0053] Figure 2The right side shows the device records and their address processing procedures on the bus enumeration side. Taking the enumeration address 4f:00.0 as an example, after extracting the enumeration address from the corresponding device record and adding a field prefix, the complete BDF address 0000:4f:00.0 is obtained. Querying the busToDlc mapping table with this complete BDF address reveals that its corresponding logical device index is 0. After processing other enumeration addresses in the same way, the successfully matched accelerator cards can be sorted according to their logical device indices to form an ordered list of accelerator card devices. Figure 2 The example also illustrates the enumeration address, current link rate of 16GT / s, and current link width x16 corresponding to logical device index 0. This allows the accelerator card on the bus enumeration side to be stably associated with the logical device index registered by the driver, and the corresponding link status to be written to the corresponding device record.
[0054] Further optionally, in the above embodiments, after sorting the successfully matched accelerator cards according to the order of the logical device index and generating the accelerator card device list, the link status output information of the corresponding accelerator card can be obtained based on the second device identification information; the current link rate and current link width can be extracted from the link status output information; the current link rate and current link width can be associated with the logical device index of the corresponding accelerator card and written into the accelerator card device list.
[0055] In the above embodiments, after sorting the successfully matched accelerator cards according to the logical device index and generating an accelerator card device list, the second device identification information of each accelerator card can be used as a device positioning parameter to call the bus status query tool to obtain the link status output information of the corresponding accelerator card. The link status output information is used to indicate the current bus link status between the accelerator card and the target computing device, and may include the current link speed and the current link width.
[0056] For example, for an accelerator card with the second device identifier 4f:00.0, the corresponding Peripheral Component Interconnect Express (PCIe) link status can be queried based on this address. Then, 16GT / s and x16 can be extracted from the rate and width fields of the link status output information, respectively, and determined as the current link rate and width of the accelerator card. Based on the mapping relationship between the second device identifier and the logical device index, 4f:00.0 can be identified as corresponding to logical device index 0. Therefore, 16GT / s and x16 can be written into the device record corresponding to logical device index 0 in the accelerator card's device list. For other accelerator cards, the corresponding link status can be obtained in the same way and associated with the corresponding logical device index.
[0057] The current link rate represents the data transfer rate of the current bus link between the accelerator card and the target computing device, while the current link width represents the number of bus channels currently participating in data transfer. By associating these two values with the logical device index and writing them into the accelerator card device list, the physical device identifier, logical device index, and current bus link status of the accelerator card can be stored simultaneously in a unified device record. This facilitates querying the link operation status of each accelerator card according to the logical device index and provides a data basis for identifying link deceleration or abnormal link width.
[0058] As an optional embodiment, step 102, associating the accelerator cards in the accelerator card device list with device identity information, includes: calling a device information reading program to obtain the device identity information corresponding to each accelerator card; parsing the serial number information corresponding to the logical device index from the output of the device information reading program; and binding the serial number information with the corresponding accelerator card in the accelerator card device list according to the logical device index; wherein, the device identity information includes at least the serial number information used to distinguish different physical accelerator cards.
[0059] In the above embodiments, a device information reading program can be invoked to obtain the device identity information of each accelerator card in the target computing device. The output of the device information reading program can record the serial number information of each accelerator card according to the logical device index. For example, the output result may include PN[0]:<serial number A> and PN[1]:<serial number B>, where the value in square brackets represents the logical device index, and the subsequent fields represent the serial number of the corresponding accelerator card.
[0060] It is worth noting that the device information reading program is used to read device identity information from the device information storage area, driver interface, or firmware interface of the accelerator card. This program can access the corresponding accelerator card according to the logical device index and output the read device serial number according to a preset field format, so that the resource monitoring and management program can parse the logical device index and its corresponding serial number information from the output result. In this embodiment, the device information reading program can be the read_info tool, but its specific program name, access interface, and output format are not limited. Any program or tool that can obtain the device identity information of the accelerator card and determine the correspondence between the device identity information and the logical device index can be used as the device information reading program.
[0061] Furthermore, the output of the device information reading program can be scanned line by line to identify the serial number field and extract the logical device index and serial number information respectively. For example, after extracting the logical device index 0 and serial number A from PN[0]:<serial number A>, the device record corresponding to logical device index 0 can be found in the accelerator card device list, and the serial number A can be written into the device record. In the same way, each serial number information can be bound to the accelerator card associated with the corresponding logical device index.
[0062] The device identification information includes at least the serial number, which distinguishes different physical accelerator cards. Compared to the logical device index, which may be affected by the device enumeration order or driver registration process, the serial number identifies the specific physical accelerator card. By recording the serial number, logical device index, and physical device identifier in the same accelerator card device list, the corresponding physical accelerator card can be determined during status monitoring or abnormal device identification, facilitating maintenance personnel in locating devices that need to be inspected or replaced.
[0063] Reference Figure 3 , Figure 3 This is a schematic diagram of the overall system architecture provided in an embodiment of this application. Figure 3 The architecture shown includes a data acquisition layer, functional modules above the data acquisition layer, and a command-line parsing and routing layer. The data acquisition layer illustrates data sources such as the sysfs virtual file system, PCIe enumeration, sensor interfaces, and / proc. The middle section shows the device discovery and mapping module, the status monitoring and acquisition module, the fault diagnosis and scheduling module, and the topology visualization module. The top section shows the command-line parsing and routing layer, with examples illustrating getopt_long parameter parsing and output redirection control.
[0064] In this embodiment, the data acquisition layer can provide device identification information, resource status data, sensor data, device enumeration information, and process information to the corresponding functional modules. The device discovery and mapping module is used to establish the mapping relationship between physical device identifiers and logical device indices. The status monitoring and acquisition module is used to acquire and process the accelerator card's operating status data. The fault diagnosis and scheduling module is used to perform corresponding fault checks or management operations. The topology visualization module is used to generate interconnection topology status information based on interconnection detection results. The command line parsing and routing layer can call corresponding functions based on received parameters and control the output method of monitoring and management results. It should be noted that... Figure 3 This mainly illustrates the composition of each layer and functional module. The specific data processing relationships are subject to the method steps described in the embodiments of this application.
[0065] Step 103: Monitor the operating status data of each accelerator card in the accelerator card device list.
[0066] In this embodiment, the operational status data reflects the resource usage, hardware operating status, and interconnection communication status of the accelerator cards during model training, model inference, or multi-card collaborative computing. The resource monitoring and management program can access the corresponding virtual file system nodes, device files, status acquisition programs, or management interfaces based on the logical device indexes in the accelerator card device list to obtain the operational status data of each accelerator card.
[0067] The operational status data in this embodiment may include storage resource data, power consumption data, power domain operating parameters, clock data, firmware version data, interconnect detection data, and device enumeration status data. Storage resource data may include the amount of free memory, used memory, and total memory in high-bandwidth memory. Power consumption data may include the power consumption of a single accelerator card and the total power consumption of a device group consisting of multiple accelerator cards. Power domain operating parameters may include the operating current of the processor core power domain. Clock data may include the processor core clock frequency and the high-bandwidth memory clock frequency. Interconnect detection data may include the error status and link connectivity status of the accelerator card interconnect ports. Device enumeration status data may include information reflecting whether the device is responding normally, such as a device revision version field.
[0068] Different operational status data can come from different data sources. For example, storage resource data can be obtained from driver nodes in the virtual file system, power consumption data and power domain operating parameters can be obtained from sensor acquisition programs, clock data and firmware version data can be obtained from firmware information programs, interconnect detection data can be obtained from link detection programs, and device enumeration status data can be obtained from PCIe device enumeration information. Step 103 obtains this operational status data to provide input for step 104 to parse and perform state transitions according to the data source.
[0069] Specifically, the operating status data of each accelerator card can be obtained sequentially according to the logical device index in the accelerator card device list. For example, if the accelerator card device list includes logical device index 0 and logical device index 1, the corresponding virtual file system nodes can be accessed to obtain the raw memory status data of High Bandwidth Memory (HBM). Then, the single-card power consumption, total power consumption of the device group, and operating current of the target power domain corresponding to each logical device index can be obtained. The firmware information program is then called to obtain the processor core clock frequency, HBM clock frequency, and firmware version information of each accelerator card.
[0070] For the interconnection status between accelerator cards, a link detection program can be invoked to obtain the error status and link connectivity status of each accelerator card port. For example, the link detection program can output the error status value and link status value corresponding to port 1 of logical device index 0, which can then be used to determine the status of the interconnection link to that port in conjunction with a preset port connection relationship. For abnormal device status, PCIe device enumeration information can be obtained, and the device address and revision field corresponding to each accelerator card can be read to identify accelerator cards that are not responding correctly.
[0071] The runtime status data output from different data sources can temporarily retain their original format, and record the corresponding data source and the device address, logical device index, or group identifier that indicates the accelerator card to which the data belongs. For example, the output of the sensor acquisition program can carry the processor unit number, the output of the firmware information program can be grouped according to the processor unit, and the output of the link detection program can carry the logical device index and port number. After obtaining the above raw data in step 103, the resource monitoring and management program provides it to the parsing, calculation, conversion, or identification process corresponding to the data source in step 104.
[0072] By obtaining operational status data from different data sources according to the accelerator card device list, the monitoring scope can be kept consistent with the accelerator cards with established device mapping relationships. This provides a data foundation for subsequent aggregation of storage resources, power consumption, temperature, clock, firmware, interconnection, and abnormal status according to logical device indexes, reducing status data mismatch caused by inconsistent output order or identification methods of different management tools.
[0073] Step 104: Based on the data sources of the operation status data, perform resource status analysis, hot status estimation, interconnection status conversion, and abnormal status identification respectively. Use the mapping relationship as the correlation benchmark between the operation status data corresponding to different data sources. Collect the resource status analysis results, hot status estimation results, and abnormal status identification results into the device status records associated with the corresponding logical device index, and associate the interconnection status conversion results with the logical device indexes at both ends of the corresponding interconnection link to obtain the comprehensive status information corresponding to each accelerator card.
[0074] In this embodiment, the data source refers to the system interface, driver node, or status acquisition program from which the operational status data originates. Since different data sources may have different output formats and device identification methods, the resource monitoring and management program employs appropriate processing methods based on the data source to convert unstructured or scattered operational status data into status information that can be retrieved according to logical device indexes.
[0075] Resource status parsing is used to extract the resource usage status of the accelerator card from the runtime status data. For example, it can format and parse the high-bandwidth memory status data in the virtual file system to obtain the amount of free memory, used memory, and total memory; it can perform field matching on the output of the sensor acquisition program to obtain the power consumption of a single card and the total power consumption of the device group; and it can parse the grouped output of the firmware information program to obtain the processor core clock frequency, high-bandwidth memory clock frequency, and firmware version.
[0076] Thermal state estimation is used to determine the temperature state of the processor unit based on the power domain operating parameters of the accelerator card. In one implementation, the operating current of the target core power domain can be extracted from the output of the sensor acquisition program, and this operating current can be input into a pre-calibrated temperature calibration model to obtain the estimated temperature of the corresponding processor unit. For cases where multiple target power domain data exist for the same processor unit, anti-duplicate sampling processing can be performed to avoid duplicate data affecting the temperature estimation results.
[0077] Interconnect state transitions are used to convert port-level detection data into link-level states between accelerator cards. Accelerator cards can form a preset communication topology through high-speed inter-chip interconnect interfaces, with each interconnect link corresponding to a specific port on two accelerator cards. The resource monitoring and management program can determine the interconnect peer corresponding to an abnormal port based on pre-saved port connection relationships and update the state value of the corresponding interconnect link in the interconnect state matrix. Since an interconnect link involves two accelerator cards simultaneously, the interconnect state transition result can be associated with the logical device indices at both ends of the interconnect link, rather than being limited to a single logical device index.
[0078] Anomaly identification is used to determine whether an accelerator card is abnormal based on device enumeration status data. For example, when the revision version field in the device enumeration information presents a preset invalid value, this status can be identified as a revision version anomaly. The resource monitoring and management program extracts the physical device identifier of the device record where the anomaly identifier is located, and uses the mapping relationship established in step 102 to determine the corresponding logical device index, thereby identifying the specific accelerator card where the anomaly occurred.
[0079] Device status records are structured data records corresponding to a single logical device index. They can include the accelerator card's physical device identifier, device identity information, storage resource status, power consumption status, thermal status, clock status, firmware version status, and abnormal status. The mapping relationship serves as a reference between different data sources, meaning that status data belonging to the same physical accelerator card from different data sources are uniformly aggregated under that accelerator card's logical device index. This avoids determining the device to which status data belongs solely based on the output order of different tools.
[0080] The comprehensive status information includes device status records corresponding to a single accelerator card and interconnection link status records corresponding to two accelerator cards. Resource status, hot status, and abnormal status can be aggregated according to the logical device index of the corresponding accelerator card. Interconnection status can be associated according to the logical device indices at both ends of the interconnection link. In this way, the correspondence between different status data and the actual accelerator cards or interconnection links can be maintained.
[0081] As an optional embodiment, in step 104, resource status parsing is performed on storage resource data from the virtual file system to obtain the storage resource usage status corresponding to each accelerator card; power consumption status parsing is performed on power consumption data from the sensor acquisition program to obtain the single-card power consumption status and the total power consumption status corresponding to each accelerator card group; thermal status estimation is performed on power consumption data and power domain operating parameters from the sensor acquisition program to obtain the thermal status corresponding to each accelerator card; group parsing is performed on firmware output data from the firmware information program to obtain the clock status and firmware version status corresponding to each accelerator card; interconnection status conversion is performed on interconnection detection data from the link detection program to obtain the interconnection status between each accelerator card; abnormal status identification is performed on abnormal identifiers from device enumeration information to obtain the abnormal status corresponding to abnormal accelerator cards; the storage resource usage status, single-card power consumption status, total power consumption status, thermal status, clock status, firmware version status, interconnection status, and abnormal status are associated according to the corresponding logical device index or interconnection link to generate the comprehensive status information.
[0082] Specifically, for storage resource data originating from the virtual file system, the amount of free memory, used memory, and total memory can be parsed from the raw memory status data of the corresponding accelerator card to form the storage resource usage status. For example, after reading the corresponding raw memory status data from the virtual file system node corresponding to logical device index 0, the parsed amount of free memory, used memory, and total memory can be written into the device status record corresponding to logical device index 0.
[0083] For the output data from the sensor acquisition program, the power consumption of each accelerator card, the total power consumption of the accelerator card group, and the operating parameters of the target power domain can be extracted separately. The power consumption of a single card can be associated with the logical device index of the corresponding accelerator card to form the power consumption status of a single card; the total power consumption of a group can be associated with the multiple logical device indices contained in that group to form the total power consumption status. For example, the power consumption of a single card in processor unit 0 can be written into the device status record corresponding to logical device index 0, and the total power consumption of processor units 0 to 3 can be associated with the corresponding accelerator card group.
[0084] Thermal state estimation can be based on the target power domain operating parameters output by the sensor acquisition program. In one specific implementation, the core power domain operating current corresponding to each processor unit can be extracted from the output data of the sensor acquisition program, and the corresponding temperature state can be determined according to a pre-calibrated temperature relationship. The power consumption data output by the sensor acquisition program can be used to form the power consumption state of a single card and the total power consumption state of a group, while the target power domain operating current can be used to form the thermal state. Thus, different fields from the same sensor data source can be used for power consumption monitoring and thermal state estimation respectively.
[0085] For firmware output data originating from the firmware information program, the logical device index corresponding to each group of data can be determined according to the accelerator card group identifier or grouping order in the output data. The processor core clock frequency, memory clock frequency, and firmware version information can then be extracted from the corresponding group to form clock status and firmware version status. For example, if group 0 in the firmware output data corresponds to logical device index 0, the core clock frequency, memory clock frequency, and firmware version information in that group can be written into the device status record corresponding to logical device index 0.
[0086] For interconnect detection data originating from the link detection program, the port number, port error status, and link connectivity status of the accelerator card can be extracted, and the interconnect link corresponding to the abnormal port can be determined according to the preset port connection relationship. For example, when port 1 of logical device index 0 is connected to port 1 of logical device index 2, if an error or disconnection is detected in port 1 of logical device index 0, the interconnect link between logical device index 0 and logical device index 2 can be determined as an abnormal state, and this interconnection status can be associated with the logical device indices at both ends of the link.
[0087] For anomaly identifiers originating from device enumeration information, the enumeration address of the device record containing the anomaly identifier can be extracted, the format of the enumeration address can be normalized, and the corresponding logical device index can be determined through the aforementioned mapping relationship. For example, if a device record corresponding to a certain enumeration address contains a revision version anomaly identifier, the accelerator card corresponding to that enumeration address can be identified as an anomaly accelerator card, and the corresponding anomaly status can be written into the device status record corresponding to its logical device index.
[0088] After completing the above processing, the storage resource usage status, single-card power consumption status, thermal status, clock status, firmware version status, and abnormal status can be aggregated into the device status record associated with the corresponding logical device index. The total power consumption status of the group is associated with the corresponding accelerator card group, and the interconnect status is associated with the logical device indices at both ends of the corresponding interconnect link, thereby generating comprehensive status information. For example, the device status record corresponding to logical device index 0 may include the memory usage, single-card power consumption, temperature, clock frequency, firmware version, and abnormal status of the accelerator card; the interconnect link status between logical device index 0 and logical device index 2 can be recorded separately and associated with both logical device indices.
[0089] Finally, by processing the operational status data according to the data source and using logical device indexes or interconnection links as the correlation benchmark, the scattered status data provided by different system interfaces can be aggregated to the corresponding accelerator card or interconnection link, avoiding status mismatch due to different data formats, device numbers and output order. At the same time, it is convenient to uniformly query the single card resource status, group power consumption status and multi-card interconnection status.
[0090] As an optional embodiment, step 104 involves performing thermal state estimation on the power consumption data and power domain operating parameters derived from the sensor acquisition program. This includes: scanning the output results of the sensor acquisition program line by line; extracting the single-card power consumption, the total power consumption of the accelerator card group, and the target power domain operating current corresponding to the target processor unit in each accelerator card from the output results; performing anti-duplicate sampling processing on the target power domain operating current of the same accelerator card to ensure that the thermal state estimation of the same accelerator card uses the first matched target power domain operating current; and inputting the target power domain operating current into a preset temperature calibration model to obtain the temperature state of the corresponding accelerator card. The preset temperature calibration model represents the correspondence between the target power domain operating current and the temperature state, and is calibrated based on the target power domain operating current and the corresponding measured temperature collected under different operating conditions.
[0091] In this embodiment, an accelerator card may include one or more target processor units. A target processor unit may be a Tensor Processing Unit (TPU). When an accelerator card includes one target processor unit, there is a one-to-one correspondence between the accelerator card and the target processor unit; the target power domain operating current and temperature state of the target processor unit can be used as the target power domain operating current and temperature state of the corresponding accelerator card. When an accelerator card includes multiple target processor units, the target power domain operating current of each target processor unit can be extracted and the corresponding temperature state calculated, and then the temperature state of each target processor unit can be associated with its respective accelerator card.
[0092] For example, the output of the sensor acquisition program may include the single-card power consumption of target processor unit 0, the total power consumption of the groups to which target processor units 0 to 3 belong, and the operating current of the 0V8 core power domain of target processor unit 0. Based on the correspondence between target processor unit 0 and logic device index 0, the single-card power consumption and the target power domain operating current can be associated with logic device index 0, and the total power consumption of the groups can be associated with the accelerator card groups corresponding to logic device indices 0 to 3. The 0V8 core power domain is one example of a target power domain; the specific type of the target power domain can be determined based on the power supply structure and temperature calibration results of the accelerator card.
[0093] The output of the sensor acquisition program may contain multiple target power domain data corresponding to the same target processor unit. To avoid duplicate data being used in thermal state estimation, a sampling status flag can be set for each target processor unit. When scanning the output results line by line, upon first matching the target power domain operating current of a target processor unit, the operating current of that target power domain is recorded, and the sampling status flag corresponding to that target processor unit is updated. When the same target processor unit's target power domain operating current is matched again, it is no longer used as input for this thermal state estimation. When there is a one-to-one correspondence between the accelerator card and the target processor unit, the sampling status flag can also be set to correspond to the accelerator card.
[0094] For example, after initially extracting the operating current of the 0V8 core power domain of target processor unit 0 corresponding to logic device index 0 as 210.50 amps, this operating current can be recorded and the sampling status flag corresponding to target processor unit 0 can be updated to the sampled status. When the operating current of the 0V8 core power domain of target processor unit 0 is scanned again subsequently, this duplicate data can be ignored, so that the thermal state estimation of target processor unit 0 in this instance only uses the operating current matched the first time.
[0095] The preset temperature calibration model can be pre-calibrated based on the target power domain operating current and corresponding measured temperature collected under different operating conditions. Different operating conditions can include idle state and operating states under different computing loads. After calibration, the first matched target power domain operating current is input into the preset temperature calibration model to obtain the temperature state of the corresponding target processor unit. For example, inputting the aforementioned operating current of 210.50 amps into the preset temperature calibration model yields a temperature state of approximately 46.6 degrees Celsius. This example value is only used to illustrate the thermal state calculation process and does not constitute a limitation on the target power domain operating current, temperature state, or temperature calibration model.
[0096] By extracting power consumption data and target power domain operating current separately from the same output of the sensor acquisition program, the power consumption status of a single card, the total power consumption status of a group, and the temperature status of the target processor unit can be generated. By performing anti-duplication sampling according to the target processor unit, duplicate power domain data from the same target processor unit can be avoided from participating in the same thermal state estimation. Using a temperature calibration model calibrated based on measured data under different operating conditions, the temperature status of the target processor unit can be determined based on existing power domain operating parameters, and this temperature status can be accurately correlated to the corresponding accelerator card.
[0097] In this embodiment, a preset temperature calibration model is used to represent the correspondence between the target power domain operating current and the core temperature of the target processor unit. This model can be a calibration model obtained through statistical regression based on the target power domain operating current collected under different operating conditions and its corresponding measured temperature, before deploying the resource monitoring and management method. The input of the preset temperature calibration model is the target power domain operating current, and the output is the estimated temperature of the target processor unit.
[0098] During the model calibration phase, the target processor unit can be run in idle state and under different computational loads, and the target power domain operating current and measured temperature can be collected synchronously in each operating state to form multiple sets of calibration samples. Each set of calibration samples includes at least the target power domain operating current and measured temperature of the same target processor unit at the corresponding sampling time or within the corresponding stable operating range. The target power domain can be the power domain that supplies power to the processor core and whose operating current is correlated with the processor's heating state. For example, in this embodiment, the target power domain can be the 0V8 core power domain. The measured temperature can be obtained through temperature calibration equipment or a temperature sensing device in the accelerator card that can provide a reference temperature.
[0099] To ensure that the operating current and measured temperature form an effective calibration sample, they can be paired according to the same target processor unit, operating state, and sampling period. Considering the potential time delay in the transmission of processor power consumption changes to temperature changes, the operating current and measured temperature can also be collected after the target processor unit enters its corresponding operating state and has stabilized for a preset period. Alternatively, the operating current data and measured temperature data can be time-aligned based on the accelerator card's thermal response characteristics. The specific stabilization period or time alignment amount can be determined through calibration experiments.
[0100] In one specific implementation, the preset temperature calibration model can be a linear calibration model. The parameters of the preset temperature calibration model include: the target power domain operating current of the target processor unit, the calculated temperature determined based on the target power domain operating current, the ratio of change between the operating current and temperature, and the temperature calibration offset. The ratio of change and the temperature calibration offset can be determined using linear regression or least squares fitting based on multiple sets of calibration samples, ensuring that the overall deviation between the calculated temperature output by the model and the measured temperature in the corresponding calibration samples meets preset requirements. For example, in a set of experimental calibration results, the parameters of the linear calibration model could be: the ratio of change between the operating current and temperature is 0.19915, and the temperature calibration offset is 4.72066.
[0101] Accordingly, by inputting the operating current of the 0V8 core power domain of the target processor unit into this linear calibration model, the corresponding core temperature can be obtained. For example, when the operating current of the 0V8 core power domain of the target processor unit is 210.50 amps, the calculated temperature output by the model is approximately 46.6 degrees Celsius. The above model parameters and example values are specific implementation results obtained based on the corresponding accelerator card model and its test data, and do not constitute a limitation on other accelerator cards, target power domains, or model parameters.
[0102] In practical applications, model parameters can be pre-calibrated and stored according to the accelerator card model, target processor unit type, or hardware version. When different accelerator cards have different power supply structures, heat dissipation structures, or processor characteristics, corresponding calibration samples can be collected separately, and the corresponding model parameters can be redefined. When performing thermal state estimation, the corresponding model parameters can be read according to the accelerator card model, target processor unit identifier, or device identity information, and then the collected target power domain operating current can be input into the corresponding preset temperature calibration model.
[0103] After model fitting is complete, the preset temperature calibration model can be validated using validation samples not involved in model parameter determination. Specifically, the target power domain operating current from the validation samples can be input into the model, and the model's output calculated temperature can be compared with the corresponding measured temperature. When the deviation meets the preset temperature monitoring accuracy requirements, the corresponding model parameters are determined as usable parameters. When the deviation does not meet the preset requirements, calibration samples under different operating conditions can be added, and the model parameters can be redefined. The preset temperature monitoring accuracy requirements can be determined based on the accelerator card's temperature monitoring needs, over-temperature protection threshold, and allowable measurement error.
[0104] During the operational phase, the target power domain operating current of the target processor unit can be extracted from the output of the sensor acquisition program. Anti-duplicate sampling is then performed using the sampling status identifier corresponding to the target processor unit, ensuring that the same target processor unit uses the first matched target power domain operating current during this monitoring process. Subsequently, this operating current is substituted into the calibrated temperature calibration model to obtain the temperature state of the corresponding target processor unit, and this temperature state is associated with the logical device index of the accelerator card to which the target processor unit belongs.
[0105] It should be noted that the preset temperature calibration model in this embodiment is a numerical calibration model based on experimental samples to determine parameters. Its model structure, input data, output data, parameter meanings, parameter determination methods, and runtime calling process are all clearly defined. This model does not rely on undisclosed neural network structures or training processes. In practical applications, other calibration relationships can also be adopted based on the calibration results of the operating current and measured temperature, but the input, output, and parameter determination methods of the corresponding model should be clearly defined based on the calibration samples. Using the aforementioned preset temperature calibration model, the processor temperature can be calculated using the operating current collected by the existing power management devices on the accelerator card, providing temperature data for monitoring the thermal status of the accelerator card.
[0106] As an optional embodiment, step 104 involves performing resource status parsing on the storage resource data from the virtual file system, including: writing a target memory pool type corresponding to high-bandwidth memory to the memory pool type node corresponding to the accelerator card in the virtual file system to trigger the driver to switch to the corresponding memory statistics type; reading the memory status node corresponding to the target memory pool type in the virtual file system to obtain raw memory status data; performing format parsing on the raw memory status data to obtain the free memory amount, used memory amount, and total memory amount of the corresponding accelerator card; and generating the storage resource usage status based on the free memory amount, the used memory amount, and the total memory amount.
[0107] It is understandable that both memory pool type nodes and memory status nodes can be access nodes registered by the accelerator card driver in the virtual file system for each logical device. These nodes are used to transmit memory query parameters and memory statistics results between the user-mode resource monitoring and management functions and the kernel-mode accelerator card driver. Accelerator cards corresponding to different logical device indices can have their own memory pool type nodes and memory status nodes, so that the written query parameters and read statistics results correspond to the same accelerator card.
[0108] The memory pool type node is a control node used to receive memory pool type identifiers. Since an accelerator card can manage high-bandwidth memory, on-chip memory, or other types of storage resources, the driver can provide status data for different memory pools through the same set of statistics interfaces. Before reading the memory status, writing the target memory pool type to the memory pool type node is equivalent to specifying the memory pool to be queried. For example, type number 8, representing high-bandwidth memory, can be written to the mem_pool_type node corresponding to logical device index 0 to instruct the driver to prepare high-bandwidth memory statistics for logical device index 0. Type number 8 is a specific implementation of the driver interface convention and does not constitute a limitation.
[0109] The memory status node is a data node used to provide statistical results for a target memory pool. After the memory pool type node receives the target memory pool type, the driver can determine the memory pool to be statistically analyzed based on the target memory pool type and have the memory status node return the capacity and usage of that memory pool. For example, the mem_pool_free node can return the amount of free memory, used memory, and total memory for high-bandwidth memory. The specific name and output fields of the memory status node can be defined by the driver interface and are not limited to the example above.
[0110] Memory statistics type refers to the memory pool category that the driver currently uses to generate or provide memory usage statistics. It defines the statistical object corresponding to the memory state node, rather than representing a new memory resource. For example, when the memory statistics type is switched to high-bandwidth memory, subsequent data read from the memory state node corresponds to high-bandwidth memory. When the memory statistics type is switched to another memory pool type, the same memory state node can return statistics for that specific memory pool. Therefore, triggering a driver to switch to the corresponding memory statistics type can be understood as the driver setting the statistical object for subsequent memory state queries to that target memory pool based on the type of the target memory pool being written.
[0111] In the above steps, the memory pool type node and memory status node corresponding to the accelerator card in the virtual file system can be determined based on the logical device index of the accelerator card. Since the same accelerator card can include different types of storage resources, the data provided by the memory status node can be determined by the type identifier written to the memory pool type node. Therefore, before reading the usage status of High Bandwidth Memory (HBM), the target memory pool type corresponding to HBM can be written to the corresponding memory pool type node to trigger the driver to switch the memory statistics type to be queried to HBM.
[0112] For example, for the accelerator card corresponding to logical device index 0, the target memory pool type number can be written to the mem_pool_type node corresponding to that accelerator card. In one specific implementation, the target memory pool type number can be 8, used to represent an HBM memory pool. This number is only a specific example agreed upon with the current driver interface; in actual applications, it can be determined according to the correspondence between the memory pool type and the type number defined by the driver.
[0113] After writing to the target memory pool type, the `mem_pool_free` memory state node under the same logical device index is read to obtain the raw memory state data corresponding to HBM. For example, the raw memory state data could be: `freemem: 512.00MB / 512.00MB, used mem: 0.00MB / ` 512.00MB. The original memory status data can be formatted and parsed according to field names, field order, and data delimiters to extract free memory, used memory, and total memory. In the example above, the free memory is determined to be 512.00MB, the used memory to be 0.00MB, and the total memory to be 512.00MB. The parsing process can also retain the corresponding units of measurement for each value to avoid data confusion between different capacity units.
[0114] After parsing and obtaining the amount of free memory, used memory, and total memory, the data can be written into the device status record associated with the corresponding logical device index, based on the logical device node to which the memory status node belongs, thus forming the storage resource usage status of the accelerator card. For example, the HBM usage data can be associated with logical device index 0, so that it can be used in conjunction with the accelerator card's power consumption status, thermal status, clock status, and firmware version status to generate comprehensive status information.
[0115] Finally, by first writing the target memory pool type and then reading the corresponding memory status node, the driver can return statistical data for the specified storage resources, avoiding the mixing of data from different memory pools. Furthermore, associating the parsed free memory, used memory, and total memory with the corresponding logical device index can accurately reflect the HBM usage of each accelerator card and provide data for determining the available storage resources and resource occupancy status of the accelerator card.
[0116] As an optional embodiment, step 104 involves performing group parsing on the firmware output data from the firmware information program, including: scanning the output results of the firmware information program line by line; when an accelerator card group identifier is detected, determining the logical device index corresponding to the current group based on the accelerator card group identifier; when the accelerator card group identifier does not contain a logical device index, updating the current group count value according to a preset grouping order, and determining the current group count value as the logical device index corresponding to the current group; extracting the core clock frequency, memory clock frequency, and firmware version information respectively within the current group; and associating the core clock frequency, memory clock frequency, and firmware version information with the corresponding logical device index within the current group.
[0117] In the above steps, the output of the firmware information program can be scanned line by line, and the firmware information can be divided into groups corresponding to different accelerator cards according to the accelerator card group identifier. Whenever a new accelerator card group identifier is scanned, the firmware output data after that group identifier and before the next accelerator card group identifier is determined as the current group, so as to extract the firmware status information of the corresponding accelerator card from the current group.
[0118] When the accelerator card group identifier contains a device number, the logical device index corresponding to the current group can be determined based on the device number. For example, if the output of the firmware information program includes TPU[0], the core clock field, the memory clock field, and the firmware version field in sequence, the group can be identified as the firmware information group corresponding to logical device index 0 based on the number 0 in TPU[0].
[0119] When the accelerator card group identifier is used only to represent a new group and does not contain a device number, the group count value can be set according to the pre-determined group output order of the firmware information program. For example, the group count value can be initialized to a preset initial value, and each time a new accelerator card group identifier is scanned, the group count value is updated sequentially, and the updated group count value is determined as the logical device index corresponding to the current group. This processing method is based on the firmware information program outputting the firmware information of each accelerator card according to the logical device index order. For example, the first and second accelerator card groups that appear sequentially can correspond to logical device index 0 and logical device index 1, respectively.
[0120] After determining the logical device index corresponding to the current group, the fields within the current group can be scanned further, and the core clock frequency, memory clock frequency, and firmware version information can be extracted respectively. For example, the current group corresponding to logical device index 0 may include: text TPU Core clock: 1.20GHz HBM clock: 1.60GHz FirmwareVersion: FW_A01. By identifying the core clock field, memory clock field, and firmware version field, the core clock frequency of 1.20GHz, the memory clock frequency of 1.60GHz, and the firmware version information FW_A01 can be extracted and written to the device status record corresponding to logical device index 0. After scanning the next accelerator card group identifier, the logical device index corresponding to the current group is updated, and the firmware output data in the next group is parsed in the same way. The above field names and example values are only used to illustrate the parsing process and do not constitute a limitation.
[0121] Finally, by determining the logical device index to which the firmware output data belongs based on the accelerator card group identifier or the preset group order, the clock frequency and firmware version information in each group can be accurately associated with the corresponding accelerator card even when the firmware output results do not use a unified device identifier format. This reduces data mismatch caused by incorrect identification of group boundaries or output order and provides a data foundation for the subsequent formation of comprehensive status information according to the logical device index.
[0122] As an optional embodiment, step 104 involves performing an interconnection state transition on the interconnection detection data from the link detection program, including: obtaining a preset connection relationship between accelerator cards, the preset connection relationship representing the connection correspondence between ports of different accelerator cards; constructing an interconnection state matrix based on the preset connection relationship, the row index and column index of the interconnection state matrix corresponding to the logical device indexes of the accelerator cards, and the state values in the interconnection state matrix representing the interconnection link status between corresponding two logical device indices; extracting abnormal port information from the output results of the link detection program, the abnormal port information including at least one of port error information and link connectivity status information; determining the corresponding interconnection link in the preset connection relationship based on the abnormal port information; updating the state values in the interconnection state matrix corresponding to the logical device indices at both ends of the interconnection link; and obtaining the interconnection state between the accelerator cards based on the updated interconnection state matrix.
[0123] Specifically, the connection mappings between the ports of each accelerator card in the target computing device can be pre-saved and used as preset connection relationships. These preset connection relationships record the logical device indices and port numbers at both ends of an interconnect link, used to determine the interconnect link and its peer based on abnormal port information at either end. For example, in an eight-card interconnect system, the preset connection relationships can record that port 1 of logical device index 0 is connected to port 1 of logical device index 2, and that port 0 of logical device index 0 is connected to port 0 of logical device index 4. The number of accelerator cards, the number of ports, and the specific connection methods can be determined based on the actual interconnect topology.
[0124] Based on preset connection relationships, an interconnection state matrix corresponding to the number of accelerator cards can be constructed. Specifically, the number of accelerator cards participating in the interconnection can be determined first based on the accelerator card device list, and the row and column indices of the interconnection state matrix can be set according to the logical device index of each accelerator card, so that each position in the matrix corresponds to a group of accelerator cards. Subsequently, each port connection record in the preset connection relationship is traversed, the logical device indices at both ends of each interconnection link are extracted, and the corresponding position is located in the matrix, and the state value of that position is initialized to a normal state value. For non-directional interconnection links, the matrix positions after swapping row and column indices can also be set synchronously, so that the two directions record the same initial link state.
[0125] For example, if the accelerator card device list includes logical device indices 0 to 7, an interconnection state matrix corresponding to eight accelerator cards can be constructed. If the preset connection relationship records show that port 1 of logical device index 0 is connected to port 1 of logical device index 2, then the intersection of the row corresponding to logical device index 0 and the column corresponding to logical device index 2 can be set to a normal state value, and the intersection of the row corresponding to logical device index 2 and the column corresponding to logical device index 0 can be synchronously set to a normal state value. By traversing the remaining port connection records in the same way, the preset interconnection links in the matrix can be initialized.
[0126] For logical device index combinations not appearing in the preset connection relationships, the corresponding matrix positions can be set to a connectionless state to distinguish them from established but currently abnormal interconnect links. For matrix positions with the same row and column indices, since they correspond to the same accelerator card, they can be set to a connectionless state or the device's own state, and are not treated as valid interconnect links between accelerator cards. Thus, the interconnection status matrix can distinguish between three situations—preset link normal, preset link abnormal, and no preset connection—through different status values. The specific status values can be set according to implementation requirements.
[0127] After matrix initialization, link detection results can be used as the basis for updates. When an error or disconnection is detected at a port, the logical device indices at both ends of the interconnection link containing that port are determined according to the preset connection relationship, and the corresponding bidirectional status value in the matrix is updated from normal to abnormal. Through the above construction method, port-level connection relationships can be converted into link status data with logical device indices as coordinates, facilitating subsequent queries and topology output.
[0128] Furthermore, a link detection program can be invoked to perform detection on the interconnect ports of each accelerator card, and its output results can be scanned line by line. The output results of the link detection program can be grouped according to logical device index or processor unit number, and each group includes the error status and link connectivity status of the corresponding port. Port error information indicates whether a transmission error has occurred on the port, and link connectivity status information indicates whether a valid connection has been established on the corresponding port. In one specific implementation, when the port error status value is non-zero, the port can be identified as an error port. When the link connectivity status value is not a preset normal value, the port can be identified as a broken link port. The port status values and their judgment conditions can be determined according to the definition of the link detection interface.
[0129] After extracting the abnormal port information, the corresponding interconnection link and its peer can be queried in the preset connection relationship based on the logical device index and port number to which the abnormal port belongs. For example, if the link detection result indicates that port 1 of logical device index 0 has a port error or broken link, and the preset connection relationship indicates that this port is connected to port 1 of logical device index 2, then it can be determined that the interconnection link between logical device index 0 and logical device index 2 is in an abnormal state.
[0130] Once an abnormal interconnect link is identified, the status values corresponding to the logical device indices at both ends of that interconnect link in the interconnect status matrix can be updated. Since an interconnect link connects two accelerator cards simultaneously, the status values pointing from logical device index 0 to logical device index 2 and from logical device index 2 to logical device index 0 can be updated separately, ensuring that both directions in the interconnect status matrix indicate that the interconnect link is abnormal. For interconnect links where no port errors are detected and the link connectivity is normal, their normal status values can be retained.
[0131] After processing each abnormal port, the updated interconnection status matrix represents the current status of each interconnection link in the preset interconnection topology. Based on the status values in the interconnection status matrix, it can be determined whether the interconnection link between any two accelerator cards with a preset connection relationship is normal or abnormal, and the corresponding interconnection status is associated with the logical device indices at both ends of the link. Furthermore, the updated interconnection status matrix can also be used to generate interconnection topology status information.
[0132] By converting port-level detection results into link-level states between accelerator cards through preset connection relationships, the affected interconnect links and their two endpoints can be directly identified after an abnormal port is detected, avoiding the need to manually search for topology connections based on port numbers. Using an interconnect status matrix to uniformly record the status of each link also facilitates subsequent querying, updating, and presentation of multi-card interconnect topologies, thereby improving the efficiency of locating abnormal interconnect links.
[0133] As an optional embodiment, step 104, obtaining the interconnection status between the accelerator cards based on the updated interconnection status matrix, includes: obtaining a preset character layout template, the preset character layout template including multiple interconnection edge placeholders; determining the accelerator card port connection relationship corresponding to each interconnection edge placeholder; determining the interconnection link status corresponding to each interconnection edge placeholder based on the interconnection status matrix; configuring distinguishable display attributes for the corresponding interconnection edge placeholders according to the interconnection link status, and converting the interconnection edge placeholders into interconnection edge characters with the display attributes; and outputting interconnection topology status information according to the preset character layout template.
[0134] Specifically, a character layout template corresponding to the accelerator card interconnection topology can be pre-set. The character layout template represents the arrangement of each accelerator card and its interconnection links in the terminal canvas. Multiple interconnection edge placeholders can be set, each representing the corresponding character position of an interconnection link in the character layout template. When the same interconnection link spans multiple character positions, it can be represented by multiple interconnection edge placeholders. For example, the character layout template can be multi-line character data arranged according to a preset interconnection topology of multiple accelerator cards, used to specify the display position of each accelerator card node, interconnection edge, and interconnection edge placeholder in the terminal canvas. The character layout template itself describes the spatial arrangement of the topology; the interconnection edge placeholders it contains do not yet represent the actual operating state of the links and need to be replaced according to the updated interconnection state matrix.
[0135] Taking four accelerator cards as an example, in a four-accelerator card application scenario, the character layout template can be configured with corresponding node and interconnection edge positions according to the preset interconnection topology of the four accelerator cards. Each accelerator card node corresponds to logical device indices 0 to 3, and multiple interconnection edge placeholders are used to represent interconnection links between logical device indices 0 and 1, 0 and 2, 1 and 3, and 2 and 3, respectively. Each interconnection edge placeholder is pre-associated with the logical device indices and port numbers at both ends of the corresponding link. During output, the status of the corresponding link is queried according to the interconnection status matrix, and the placeholder is replaced with an interconnection edge character with corresponding display attributes. For example, green characters can represent normal links, and red characters can represent abnormal links; if color display is not supported, different character styles or text markers can be used to distinguish link statuses. For positions without actual connections, spaces can be used for replacement. When the same interconnection link occupies multiple character positions in the character layout, multiple placeholders can be set to jointly represent the link, and these placeholders can be associated with the same set of logical device indices. Therefore, the character layout template can convert the link states in the interconnection state matrix into a characterized display result corresponding to the actual multi-card topology. The node arrangement, number of placeholders, and display method described above are merely examples and do not constitute a limitation.
[0136] For example, in an eight-card interconnect system, letters A through U can be used as interconnect edge placeholders, and a correspondence table between each interconnect edge placeholder and the connection relationship of the accelerator card ports can be pre-established. For placeholder A, the interconnect link between port 1 of logical device index 0 and port 1 of logical device index 2 can be determined according to the correspondence table. Then, based on the status values corresponding to logical device index 0 and logical device index 2 in the interconnect status matrix, it can be determined whether the interconnect link is in a normal or abnormal state.
[0137] After determining the interconnection link status, display attributes can be configured for the corresponding interconnection edge placeholders to distinguish the link status, and the corresponding interconnection edge characters can be selected. Display attributes can include at least one of color, brightness, or character style. For example, when the link is normal, a green display attribute can be used to output horizontal, vertical, or diagonal lines corresponding to that position as the interconnection edge character. When the link is abnormal, a red display attribute can be used to output the corresponding interconnection edge character. For positions where no interconnection link exists, a space can be used to replace the corresponding placeholder. The specific display attributes and interconnection edge characters can be determined based on the display methods supported by the terminal and the extension direction of the interconnection edge in the template.
[0138] After processing the interconnection edge placeholders in the above manner, the corresponding placeholders in the character layout template can be replaced with interconnection edge characters that have corresponding display attributes to form an interconnection topology character template with completed state rendering. This interconnection topology character template is then output to obtain the interconnection topology status information. Therefore, the link status in the interconnection status matrix can be converted into a character topology corresponding to the actual multi-card connection relationship, allowing abnormal links to be displayed differently at preset topology locations, reducing the need for maintenance personnel to manually search for both ends of the link and compare topology relationships based on port detection text.
[0139] Reference Figure 4 , Figure 4 This is a visual schematic diagram of an eight-card TPU interconnect topology provided in an embodiment of this application. Figure 4 The eight TPU nodes corresponding to logical device indices 0 to 7 are displayed in terminal character format. The nodes in the top row are 3, 0, 4, and 7, respectively, and the nodes in the bottom row are 2, 1, 5, and 6, respectively. Interconnections between nodes are represented by horizontal, vertical, and diagonal lines, and the status of the interconnections is indicated by different colors. According to... Figure 4 As shown in the diagram, green interconnecting edges indicate that the link is connected, while red interconnecting edges indicate that the link is disconnected.
[0140] In this embodiment, the interconnection status matrix can be queried based on the logical device index corresponding to each interconnection edge, and display attributes can be configured for the corresponding interconnection edge placeholders in the character layout template based on the retrieved link status. After replacing all placeholders, the output is formed. Figure 4 The interconnection topology status information is shown. Figure 4 The specific port numbers corresponding to each interconnect edge are not shown. The correspondence between interconnect edges and accelerator card ports can be determined by pre-set connection relationships. This character-based display method can distinguish between normal and abnormal links at corresponding locations in a multi-card topology, reducing the need for manual searching of interconnect links based on port detection text.
[0141] As an optional embodiment, in step 104, after updating the state values corresponding to the logical device indices at both ends of the interconnection link in the interconnection state matrix, it is further possible to determine whether the interconnection links between valid accelerator cards in the accelerator card device list meet the preset physical connectivity conditions based on the updated interconnection state matrix; when the interconnection links between valid accelerator cards meet the preset physical connectivity conditions, aggregated communication verification parameters are generated based on the logical device indices of the valid accelerator cards; an aggregated communication verification program is called based on the aggregated communication verification parameters to perform aggregated communication operations on the valid accelerator cards; the aggregated communication verification result output by the aggregated communication verification program is obtained; and based on the interconnection state matrix and the aggregated communication verification result, the communication state between the valid accelerator cards is determined to be a physical interconnection link abnormal state, an aggregated communication function abnormal state, or a normal communication state.
[0142] It is worth noting that after updating the interconnection status matrix, valid accelerator cards can be identified first from the accelerator card device list. A valid accelerator card refers to one that has not been marked or filtered by the abnormal status identification results and can participate in multi-card communication verification. Subsequently, the interconnection links required for collective communication between valid accelerator cards are checked according to the interconnection status matrix to determine whether the preset physical connectivity conditions are met. These preset physical connectivity conditions may include the absence of port errors or broken links between the valid accelerator cards involved in the collective communication, or the availability of all interconnection paths required by the preset communication topology.
[0143] When the preset physical connectivity conditions are met, the logical device indexes of each valid accelerator card can be extracted, and the collective communication verification parameters can be generated according to the parameter format required by the collective communication verification program. For example, if the logical device indices corresponding to the valid accelerator cards are 0, 1, 2, and 3, these logical device indices can be passed as participating device parameters to the collective communication verification program to specify the range of accelerator cards involved in this verification.
[0144] The aggregated communication verification program can perform at least one aggregated communication operation, such as full reduction, broadcast, or full collection, between specified valid accelerator cards and output the execution results of each operation. The aggregated communication verification results can include participating devices, communication operation types, and success or failure status, reflecting whether the software communication function between accelerator cards is normal.
[0145] For example, if the interconnection status matrix indicates that all interconnections required by logical device indices 0 to 3 are normal, and the aggregated communication verification result shows that all aggregated communication operations are successful, the corresponding communication status can be determined as a normal communication status. If port errors or broken links exist in the interconnection status matrix, causing the preset communication topology to fail to meet physical connectivity conditions, it can be determined as a physical interconnection link abnormality, and aggregated communication verification can be discontinued. If the interconnection status matrix indicates that the relevant interconnections are normal, but the aggregated communication verification program outputs a communication failure result, it can be determined as an aggregated communication function abnormality.
[0146] By first checking the physical interconnect links and then verifying the aggregated communication function, the connectivity status at the port or link level can be combined with the execution status of the aggregated communication at the software level. This distinguishes whether the communication anomaly occurs in the physical interconnect link or the aggregated communication function, reducing the location deviation caused by relying solely on a single detection result for fault diagnosis, and providing a basis for checking the communication status before starting multi-card computing tasks.
[0147] As an optional embodiment, step 104 involves performing abnormal state identification on the abnormal identifiers derived from the device enumeration information, including: scanning the device enumeration information for a revision version abnormal identifier used to indicate device abnormality, wherein the revision version abnormal identifier is a preset invalid value in the device revision version field, and the preset invalid value is used to characterize an abnormal device configuration space read of the corresponding accelerator card; if the revision version abnormal identifier exists, extracting the device enumeration address containing the revision version abnormal identifier; performing format normalization processing on the device enumeration address, and determining the logical device index corresponding to the device enumeration address based on the mapping relationship; identifying the accelerator card corresponding to the logical device index as an abnormal accelerator card, and generating an abnormal state corresponding to the abnormal accelerator card.
[0148] Specifically, the device enumeration information can be scanned line by line, and the revision version field in the corresponding device record of each accelerator card can be read. The revision version field is used to indicate the hardware revision version of the PCIe device. When the target computing device cannot read the device configuration space of the accelerator card normally, this field may return a preset invalid value, forming a revision version anomaly identifier in the device enumeration information. For example, the preset invalid value can be all 1s, and the corresponding device enumeration record may contain the rev ff identifier. When a device record containing this identifier is scanned, the device enumeration address in that record can be extracted. For example, if the device enumeration information contains address 57:00.0 and the rev ff identifier, 57:00.0 is determined as the enumeration address of the abnormal device to be identified. The above-mentioned preset invalid value and anomaly identifier are only one specific implementation, and can be set according to the definition of the invalid revision version field in the device enumeration interface.
[0149] Since the device enumeration address may be a short bus device function address without a domain prefix, the corresponding domain prefix can be added according to the aforementioned format normalization method. For example, 57:00.0 can be converted to 0000:57:00.0, and then the mapping relationship between the physical device identifier and the logical device index can be queried using the format normalized address. If the mapping relationship records that 0000:57:00.0 corresponds to logical device index 2, then the accelerator card corresponding to logical device index 2 can be identified as an abnormal accelerator card, and the abnormal status can be read from the configuration space in its device status record.
[0150] Furthermore, the logical device index, physical device identifier, and device identity information of an abnormal accelerator card can be associated with the same abnormal device record for subsequent device marking, valid device filtering, or faulty device location. For device records where no revised version of the abnormal identifier is detected, their original device status can be retained, and other operational status data can be used to determine whether the accelerator card is functioning correctly.
[0151] By utilizing the revision anomaly identifier in the device enumeration information, accelerator cards with abnormal device configuration space reads can be automatically identified from the bus enumeration results. Combined with address format normalization and device mapping relationships, abnormal physical devices can be accurately converted into logical device indexes, reducing manual verification of enumeration addresses and logical numbers, and providing a basis for anomaly device filtering and physical device location.
[0152] Step 105: Bind the comprehensive status information to the logical device index of the corresponding accelerator card, and update the accelerator card device list according to the abnormal status identification result in the comprehensive status information to generate accelerator card resource monitoring and management results.
[0153] In the above steps, binding the comprehensive status information with the logical device index means using the logical device index as the retrieval identifier for the device status record, and writing the physical device identifier, device identity information, and operating status information of the corresponding accelerator card into the same device status record. For interconnection status information, the logical device indices at both ends of the interconnection link can be used together as the association identifier for the interconnection link status record.
[0154] Correlation updates refer to updating the status or validity of the corresponding device record in the accelerator card device list based on the anomaly identification results and their corresponding logical device indexes. For example, an anomaly flag can be written to the corresponding device record, or the logical device index of the anomaly accelerator card can be added to the exclusion set to filter the anomaly accelerator card when forming the valid accelerator card list. Simultaneously, the logical device index, physical device identifier, and device identity information of the anomaly accelerator card can be written to the anomaly device record so that maintenance personnel can identify the specific physical accelerator card that is experiencing an anomaly.
[0155] The accelerator card resource monitoring and management results are outputs formed by organizing the accelerator card device list and comprehensive status information. These results may include a list of valid accelerator cards, a list of abnormal accelerator cards, device status records for each accelerator card, and interconnection link status records between accelerator cards. Specifically, the accelerator card resource monitoring and management results may include one or more of the following: logical device index, physical device identifier, device serial number, PCIe link status, high-bandwidth memory usage status, power consumption status, temperature status, clock status, firmware version status, abnormal status, and interconnection topology status.
[0156] In practical applications, the results of accelerator card resource monitoring and management can be output through a command-line interface, structured data files, or a resource management platform. These results can be used for device status queries, locating abnormal devices, analyzing interconnection link faults, and performing device checks before multi-card computing tasks are initiated. Consequently, data provided by different system interfaces and status acquisition programs can be organized uniformly according to their corresponding logical device indexes or interconnections, reducing the need for maintenance personnel to manually verify different device numbers and status outputs.
[0157] As an optional embodiment, in step 105, the comprehensive status information is written into the status record associated with the corresponding logical device index according to the mapping relationship; the set of logical device indexes corresponding to the abnormal acceleration card is determined according to the abnormal status identification result; based on the set of logical device indexes, the abnormal acceleration card is filtered from the list of valid acceleration cards, and the logical device index and device identity information corresponding to the abnormal acceleration card are written into the abnormal device record; based on the list of valid acceleration cards, the abnormal device record, and the comprehensive status information corresponding to each acceleration card, the acceleration card resource monitoring and management result is generated.
[0158] In the above embodiments, the mapping relationship between physical device identifiers and logical device indexes can be used as a basis to write comprehensive status information formed from different data sources into the corresponding device status records. Each device status record uses the logical device index as a retrieval identifier and can record the physical device identifier, device identity information, storage resource usage status, power consumption status, thermal status, clock status, firmware version status, and abnormal status of the corresponding accelerator card. Interconnect status can be written into interconnect link status records jointly identified by the logical device indices of both ends of the interconnect link.
[0159] Based on the abnormal status identification results, the logical device indexes corresponding to each abnormal accelerator card can be extracted to form a logical device index set. This set distinguishes between valid accelerator cards that can continue participating in status monitoring or computing tasks and those exhibiting device abnormalities. When generating the valid accelerator card list, it is checked whether the logical device index corresponding to each device record belongs to the logical device index set. Accelerator cards belonging to this set are not written to the valid accelerator card list; the remaining accelerator cards are retained in the valid accelerator card list.
[0160] For example, if the accelerator card device list includes logical device indices 0 to 3, and the abnormal status identification result indicates that the accelerator card corresponding to logical device index 2 has a device configuration space read error, then logical device index 2 can be added to the logical device index set. When generating the valid accelerator card list, the accelerator cards corresponding to logical device indices 0, 1, and 3 are retained, while the abnormal accelerator card corresponding to logical device index 2 is filtered out. Simultaneously, logical device index 2, its corresponding physical device identifier, and device serial number are written to the abnormal device record. Abnormal status and other valid status information already obtained for this abnormal accelerator card can also be retained in the corresponding device status record.
[0161] Based on the list of valid accelerator cards, records of abnormal devices, and comprehensive status information, accelerator card resource monitoring and management results can be generated. The list of valid accelerator cards indicates currently available accelerator cards and their operating status; the records of abnormal devices indicate the logical device index, physical device identifier, device identity information, and abnormality type of the abnormal accelerator card; and the comprehensive status information provides the resource, power consumption, temperature, clock, firmware, and interconnection link status of each accelerator card. Accelerator card resource monitoring and management results can be output through a command-line interface, a structured data file, or a resource management platform.
[0162] By filtering abnormal accelerator cards from the list of valid accelerator cards and retaining their logical device index and device identity information separately, it is possible to prevent abnormal devices from being mixed with valid devices or from being selected as candidate devices for computing tasks. At the same time, by leveraging the correspondence between physical device identifiers, logical device indexes, and device identity information, the specific physical accelerator card can be directly located from the abnormal status, facilitating subsequent inspection or device replacement.
[0163] Further optionally, in the above embodiments, after writing the logical device index and device identity information corresponding to the abnormal accelerator card into the abnormal device record, a soft reset request for the corresponding abnormal accelerator card can be sent to the device management process according to the logical device index in the abnormal device record; the device management process performs a soft reset on the abnormal accelerator card according to the soft reset request; after the soft reset is completed, the device enumeration information of the target computing device is re-acquired, and the revised version abnormal identifier corresponding to the abnormal accelerator card is checked according to the re-acquired device enumeration information; if the revised version abnormal identifier disappears, the mapping relationship between the physical device identifier and the logical device index of the abnormal accelerator card is re-established, and the abnormal accelerator card is restored to the list of valid accelerator cards; if the revised version abnormal identifier does not disappear, the abnormal accelerator card is retained in the abnormal device record.
[0164] Specifically, after an abnormal device record is generated, a soft reset request can be sent to the device management process based on the logical device index recorded therein. The device management process can be a daemon process that resides on the target computing device and is used to manage the power status and reset operation of the accelerator card. It can identify the target abnormal accelerator card based on the logical device index in the soft reset request and perform the reset operation through the corresponding device management interface.
[0165] For example, if the abnormal device log records that the accelerator card corresponding to logical device index 2 has an abnormal device configuration space read, a soft reset request carrying logical device index 2 can be sent to the device management process. Upon receiving this request, the device management process writes soft reset control information to the accelerator card corresponding to logical device index 2 to clear its abnormal operating state and reinitializes the corresponding PCIe link. The soft reset process does not require restarting the target computing device; other valid accelerator cards can continue to operate.
[0166] After a soft reset, the device enumeration information of the target computing device is retrieved again. Based on the physical device identifier in the abnormal device record or the device address obtained from the re-enumeration, the device enumeration record of the corresponding accelerator card is searched. If the revised version abnormal identifier in the device enumeration record has disappeared, it indicates that the device configuration space of the corresponding accelerator card can be read normally. At this point, the first and second device identifier information of the accelerator card can be retrieved again. The corresponding device identifiers are then formatted and matched, the mapping relationship between the physical device identifier and the logical device index is re-established, and the accelerator card is restored to the list of valid accelerator cards.
[0167] For example, if the abnormal accelerator card corresponding to logical device index 2 contains the "rev ff" identifier in its enumeration record before a soft reset, but this identifier is no longer present in the re-acquired enumeration record after the soft reset, then the physical device identifier, logical device index, and device identity information of the accelerator card can be re-verified, and it can be removed from the abnormal device record and restored to the list of valid accelerator cards. After restoration, the operating status data of the accelerator card can be re-collected to update its device status record.
[0168] If the re-acquired device enumeration information still contains the corresponding revision version anomaly identifier, it indicates that the soft reset did not restore the accelerator card to normal operation. In this case, the accelerator card will remain in the abnormal device record and will continue to be filtered out of the list of valid accelerator cards to avoid using it as a candidate device for subsequent computing tasks.
[0169] In this way, by linking abnormal device identification, soft reset, re-detection after reset, and device list restoration, it is possible to attempt to recover abnormal accelerator cards without interrupting the operation of the target computing device. The decision on whether to restore its valid device identity is based on the actual enumeration status after reset, avoiding the determination of a device as usable simply because a soft reset request is issued, while reducing the downtime of devices caused by recoverable anomalies.
[0170] Optionally, for a target accelerator card in the accelerator card device list, obtain the device file path corresponding to the target accelerator card; use the device file path as the query object to obtain the process identifier information occupying the target accelerator card; read the process name file based on the process identifier information to obtain process name information; read the command line file based on the process identifier information to obtain raw command line data; replace the parameter separator in the raw command line data with spaces to obtain readable command line information; write the process name information and the readable command line information into the process status record associated with the corresponding logical device index, so as to incorporate the process occupancy status into the comprehensive status information.
[0171] In the above embodiments, the corresponding device file path can be determined based on the logical device index of the target accelerator card. The device file path is the user-mode access entry provided by the operating system of the target computing device for the accelerator card, and different logical device indices correspond to different device files. For example, the device file path corresponding to logical device index 0 can be / dev / dlc0, and the device file path corresponding to logical device index 1 can be / dev / dlc1.
[0172] Using the device file path of the target accelerator card as the query object, a file usage query tool can be invoked to obtain the process identification information of the currently open or using device file. The process identification information can be a process identifier (PID). For example, querying / dev / dlc0 yields PID 3256 and PID 4180, indicating that these two processes are currently using the target accelerator card corresponding to logical device index 0.
[0173] For each PID, the corresponding process name file and command line file in the process file system can be read separately. For example, reading `proc / 3256 / comm` will provide the process name information, and reading ` / proc / 3256 / cmdline` will provide the raw command line data for that process. The program name and various runtime parameters in the raw command line data are usually separated by null bytes. Therefore, the null byte parameter separators can be replaced with spaces to obtain readable command line information that is easier to display and recognize. For example, a complete command line content including the program name, task script name, model name, and batch processing parameters can be obtained.
[0174] After reading and conversion, the PID, process name information, and readable command-line information are written to the process status record corresponding to logical device index 0, and this process status record is merged into the comprehensive status information of the target accelerator card. If multiple processes are occupying the same target accelerator card, corresponding process information entries can be generated separately. If no occupying process is found, the process occupancy status of the target accelerator card can be recorded as idle. If the corresponding process exits during the query, causing the process name file or command-line file to become unreadable, the corresponding field can be recorded as unavailable without affecting the collection of other runtime status data.
[0175] By using the device file path as the query entry point, and combining it with the PID to read the process name and complete command line, the accelerator card usage can be accurately associated with the corresponding logical device index. This enables operations and maintenance personnel to determine which processes and computing tasks are currently using the specific accelerator card, providing a basis for device resource query, task anomaly troubleshooting, and accelerator card scheduling.
[0176] Reference Figure 5 , Figure 5 This is a schematic diagram of a DLC-SMI integrated status information panel output provided in an embodiment of this application. Figure 5 The panel displays eight TPU device records corresponding to logical device indices 0 to 7. The status fields in the panel include device identifier, PCIe link rate and link width, HBM idle amount, used amount and total amount, TPU core clock frequency, HBM clock frequency, power consumption, temperature and firmware version. Figure 5The total power consumption of the groups TPU0 to TPU3 and TPU4 to TPU7 is also shown.
[0177] exist Figure 5 In the example shown, each TPU device record displays a current PCIe link rate of 16GT / s, a current link width of x16, and a normal status. The total HBM for each TPU is displayed as 63360.00MB, and the used amount is displayed as 0.00MB. The figure also shows the core clock frequency, HBM clock frequency, single-card power consumption, and temperature for each TPU. The specific firmware version string is not clearly shown in the attached figure, and its content is not further limited. Figure 5 The lower section contains a process information area, whose fields include the TPU logical device index, process identifier, process name, and command-line information. In this example, devices corresponding to logical device indices 0 to 7 all show no active processes.
[0178] In this embodiment, status information from different data sources can be written into the corresponding device status record according to the logical device index, and a comprehensive status information panel can be generated according to a preset field order. For accelerator cards with occupied processes, the corresponding process identifier, process name, and readable command-line information can also be displayed in the process information area. For accelerator cards without occupied processes, a status of no active processes can be output. Figure 5 This is just an example of output; the actual output fields, arrangement, numerical precision, and output medium can be set according to monitoring requirements.
[0179] Optionally, the system can receive user-inputted output control parameters and determine the output target, output format, and monitoring method of the accelerator card resource monitoring and management results based on these parameters. The output control parameters can be passed in via command-line parameters or a management interface and include at least one of output target parameters, output format parameters, and cyclic monitoring parameters.
[0180] The output target parameter specifies the output location for monitoring and management results. When the output target parameter indicates file output, the corresponding target file can be opened, and the monitoring and management results originally intended for the terminal can be written to that target file. For example, the output target parameter can specify a status log file, and the list of valid accelerator cards, abnormal device records, and comprehensive status information can be written to that file. In practice, the output target of the standard output stream and standard error stream can be switched to the target file, and the original output target can be restored after the output is completed, thus without changing the original output process of each status information.
[0181] The output format parameter specifies the organization format of the monitoring and management results. When the output format parameter indicates a structured format output, data such as logical device index, physical device identifier, device identity information, storage resource status, power consumption status, temperature status, clock status, firmware version status, and abnormal status can be organized according to a preset field order. For example, the status records of each accelerator card can be output line by line using a comma-separated value format, and the field headers, device identifier prefixes, and data units can be retained based on the output format parameter. The above structured format and its field settings are only examples; in practical applications, other data formats that are easy to store or parse can also be used.
[0182] The cyclic monitoring parameter specifies whether to continuously update the accelerator card resource monitoring and management results, as well as the time interval between two adjacent monitoring sessions. For example, when the cyclic monitoring parameter specifies a refresh interval of 5 seconds, it can wait 5 seconds after completing one device identifier acquisition, status monitoring, comprehensive status generation, and result output before reacquiring the operating status data of each accelerator card and updating the monitoring and management results. The preset time interval can be set according to the rate of status change, real-time monitoring requirements, and the processing load of the target computing device.
[0183] During cyclic monitoring, user-triggered interrupt signals can be received. Upon receiving an interrupt signal, the cyclic stop flag can be updated, and the loop can exit after the current monitoring and output process is completed, thus avoiding direct interruption of processing while status data is being acquired or the target file is being written. For example, an interrupt signal generated by the control terminal can be used as a trigger condition to stop cyclic monitoring. The specific signal type and processing method can be determined based on the operating system of the target computing device.
[0184] By outputting control parameters, the same monitoring and management results can be adapted to different needs such as terminal viewing, file retention, and structured data processing without changing the status acquisition and processing logic. By re-acquiring and updating status information according to preset time intervals, the changes in the accelerator card's operating status can be continuously reflected. By responding to interrupt signals and orderly stopping the cyclic monitoring, the situation of incomplete output or abnormal file writing can be reduced.
[0185] In this embodiment, by normalizing and matching the device identification information under different system interfaces, a correspondence is established between the physical device identifier, logical device index, and device identity information of the accelerator card. This mapping relationship serves as the association benchmark for multi-source operational status data, aggregating resource status, hot status, and abnormal status to the device status record of the corresponding accelerator card, and associating interconnection status with the logical device indexes at both ends of the corresponding interconnection link. Based on this, the accelerator card device list is updated according to the abnormal status identification results, generating resource monitoring and management results that include valid devices, abnormal devices, and their comprehensive status. This reduces status mismatches caused by inconsistencies in device identifiers, device numbers, and data output order under different system interfaces, facilitating unified monitoring and location of resource usage, hot status, interconnection status, and abnormal status of each accelerator card.
[0186] The above describes a method for monitoring and managing accelerator card resources in the embodiments of this application. The following describes the accelerator card resource monitoring and management device that performs the above method.
[0187] See Figure 6 ,like Figure 6 The diagram shows a structural schematic of an accelerator card resource monitoring and management device. The accelerator card resource monitoring and management device in this embodiment can achieve the functions described above. Figure 1 The steps of the accelerator card resource monitoring and management method executed in the corresponding embodiments are described above. The functions implemented by the accelerator card resource monitoring and management device can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions, and the modules can be software and / or hardware. The accelerator card resource monitoring and management device may include various modules, and the functional implementation of each module can be referred to... Figure 1 The operations performed in the corresponding embodiments will not be described in detail here.
[0188] The acquisition module 601 is configured to acquire first device identification information representing the physical connection relationship of the accelerator card from the target computing device, and acquire second device identification information representing the logical registration relationship of the accelerator card. The mapping module 602 is configured to establish a mapping relationship between the physical device identifier and the logical device index of the accelerator card based on the format-normalized second device identifier information and the first device identifier information, to obtain an accelerator card device list, and associate the accelerator cards in the accelerator card device list with device identity information; The monitoring module 603 is configured to monitor the operating status data of each accelerator card in the accelerator card device list; based on the data source of the operating status data, it performs resource status analysis, hot status estimation, interconnection status conversion and abnormal status identification respectively; using the mapping relationship as the correlation benchmark between the operating status data corresponding to different data sources, it collects the resource status analysis results, hot status estimation results and abnormal status identification results into the device status record associated with the corresponding logical device index, and associates the interconnection status conversion results with the logical device indexes at both ends of the corresponding interconnection link to obtain the comprehensive status information corresponding to each accelerator card; The management module 604 is configured to bind the comprehensive status information with the logical device index of the corresponding accelerator card, and update the accelerator card device list according to the abnormal status identification result in the comprehensive status information, thereby generating accelerator card resource monitoring and management results.
[0189] In some implementations, the acquisition module 601, which acquires first device identification information representing the physical connection relationship of accelerator cards and second device identification information representing the logical registration relationship of accelerator cards from the target computing device, is configured to: traverse the logical device nodes corresponding to each accelerator card in the virtual file system of the target computing device; read the bus address information of the corresponding accelerator card from the logical device node and use the bus address information as the first device identification information; call the bus enumeration tool to acquire the enumeration output information of each accelerator card in the target computing device; extract the enumeration address information of the corresponding accelerator card from the enumeration output information and use the enumeration address information as the second device identification information; wherein the first device identification information and the second device identification information respectively correspond to the device identification presented by the same accelerator card under different system interfaces.
[0190] In some implementations, the mapping module 602 establishes a mapping relationship between the physical device identifier and the logical device index of the accelerator card based on the format-normalized second device identifier information and the first device identifier information, thereby obtaining an accelerator card device list, which is configured as follows: The first device identification information is indexed and extracted to obtain the logical device index corresponding to each first device identification information, and a first mapping table is established to represent the correspondence between the first device identification information and the logical device index. The second device identification information is normalized to make the second device identification information have the same address format as the first device identification information. The normalized second device identification information is matched with the first device identification information in the first mapping table to determine the logical device index corresponding to the second device identification information. The accelerator cards that successfully match are sorted according to the order of the logical device index to generate the accelerator card device list.
[0191] In some implementations, the mapping module 602 matches the format-normalized second device identifier information with the first device identifier information in the first mapping table to determine the logical device index corresponding to the second device identifier information, and is configured as follows: Extract the bus device function address without the field prefix from the enumerated output information; determine whether the bus device function address and the first device identification information have the same address field; if the bus device function address does not contain the field identifier of the first device identification information, add a preset field prefix to the bus device function address to obtain the format-normalized second device identification information; use the format-normalized second device identification information as the query key to perform a search operation in the first mapping table; if the first device identification information that is consistent with the format-normalized second device identification information is found in the first mapping table, then determine the logical device index corresponding to the first device identification information as the logical device index corresponding to the second device identification information.
[0192] In some implementations, after the mapping module 602 sorts the successfully matched accelerator cards according to the order of the logical device index and generates the accelerator card device list, it is further configured to: obtain the link status output information of the corresponding accelerator card based on the second device identification information; extract the current link rate and current link width from the link status output information; associate the current link rate and current link width with the logical device index of the corresponding accelerator card and write them into the accelerator card device list.
[0193] In some implementations, the mapping module 602 is configured to associate accelerator cards in the accelerator card device list with device identity information as follows: The device information reading program is invoked to obtain the device identity information corresponding to each accelerator card; the serial number information corresponding to the logical device index is parsed from the output of the device information reading program; according to the logical device index, the serial number information is bound to the corresponding accelerator card in the accelerator card device list; wherein, the device identity information includes at least the serial number information used to distinguish different physical accelerator cards.
[0194] In some implementations, the monitoring module 603 performs resource status analysis, hot status estimation, interconnection status transformation, and abnormal status identification based on the data source of the operational status data. Using the mapping relationship as the correlation benchmark between operational status data from different data sources, it aggregates the resource status analysis results, hot status estimation results, and abnormal status identification results into the device status records associated with the corresponding logical device indexes, and associates the interconnection status transformation results with the logical device indexes at both ends of the corresponding interconnection link to obtain the comprehensive status information corresponding to each accelerator card, which is configured as follows: Resource status parsing is performed on storage resource data from the virtual file system to obtain the storage resource usage status of each accelerator card; power consumption status parsing is performed on power consumption data from the sensor acquisition program to obtain the single-card power consumption status and the total power consumption status of the accelerator card group; thermal status estimation is performed on power consumption data and power domain operating parameters from the sensor acquisition program to obtain the thermal status of each accelerator card; group parsing is performed on firmware output data from the firmware information program to obtain the clock status and firmware version status of each accelerator card; interconnection status conversion is performed on interconnection detection data from the link detection program to obtain the interconnection status between each accelerator card; abnormal status identification is performed on abnormal identifiers from device enumeration information to obtain the abnormal status of the abnormal accelerator card; the storage resource usage status, single-card power consumption status, total power consumption status, thermal status, clock status, firmware version status, interconnection status, and abnormal status are associated according to the corresponding logical device index or interconnection link to generate the comprehensive status information.
[0195] In some implementations, the monitoring module 603 is configured to perform thermal state estimation on power consumption data and power domain operating parameters derived from the sensor acquisition program, as follows: The output of the sensor acquisition program is scanned line by line, and the single-card power consumption, the total power consumption of the accelerator card group, and the target power domain operating current of the target processor unit in each accelerator card are extracted from the output. Anti-duplicate sampling processing is performed on the target power domain operating current of the same accelerator card so that the thermal state estimation of the same accelerator card adopts the first matched target power domain operating current. The target power domain operating current is input into a preset temperature calibration model to obtain the temperature state of the corresponding accelerator card. The preset temperature calibration model is used to represent the correspondence between the target power domain operating current and the temperature state, and is calibrated according to the target power domain operating current collected under different working states and the corresponding measured temperature.
[0196] In some implementations, the monitoring module 603 is configured to perform resource status parsing on storage resource data originating from the virtual file system, and is configured to: Write the target memory pool type corresponding to high-bandwidth memory to the memory pool type node corresponding to the accelerator card in the virtual file system to trigger the driver to switch to the corresponding memory statistics type; read the memory status node corresponding to the target memory pool type in the virtual file system to obtain the raw memory status data; perform format parsing on the raw memory status data to obtain the free memory amount, used memory amount, and total memory amount of the corresponding accelerator card; generate the storage resource usage status based on the free memory amount, the used memory amount, and the total memory amount.
[0197] In some implementations, the monitoring module 603 is configured to perform packet parsing on firmware output data originating from the firmware information program, and is configured to: When an accelerator card group identifier is detected, the logical device index corresponding to the current group is determined based on the accelerator card group identifier. When the accelerator card group identifier does not contain a logical device index, the current group count value is updated according to a preset group order, and the current group count value is determined as the logical device index corresponding to the current group. The core clock frequency, memory clock frequency, and firmware version information are extracted from the current group. The core clock frequency, memory clock frequency, and firmware version information are associated with the logical device index corresponding to the current group.
[0198] In some implementations, the monitoring module 603 is configured to perform interconnect state transitions on interconnect detection data from the link detection program, and is configured to: A preset connection relationship between accelerator cards is obtained, which represents the connection correspondence between ports of different accelerator cards. An interconnection state matrix is constructed based on the preset connection relationship, where the row index and column index of the interconnection state matrix correspond to the logical device indexes of the accelerator cards, and the state values in the interconnection state matrix represent the interconnection link status between two corresponding logical device indices. Abnormal port information is extracted from the output of the link detection program, which includes at least one of port error information and link connectivity information. The corresponding interconnection link is determined in the preset connection relationship based on the abnormal port information. The state values in the interconnection state matrix corresponding to the logical device indices at both ends of the interconnection link are updated. The interconnection status between the accelerator cards is obtained based on the updated interconnection state matrix.
[0199] In some implementations, the monitoring module 603, based on the updated interconnection state matrix, obtains the interconnection state between the accelerator cards and is configured as follows: Obtain a preset character layout template, which includes multiple interconnect edge placeholders; determine the accelerator card port connection relationship corresponding to each interconnect edge placeholder; determine the interconnect link status corresponding to each interconnect edge placeholder based on the interconnect status matrix; configure distinguishable display attributes for the corresponding interconnect edge placeholders according to the interconnect link status, and convert the interconnect edge placeholders into interconnect edge characters with the display attributes; output interconnect topology status information according to the preset character layout template.
[0200] In some implementations, after updating the status values in the interconnection status matrix corresponding to the logical device indices at both ends of the interconnection link, the monitoring module 603 is further configured to: Based on the updated interconnection status matrix, determine whether the interconnection links between valid accelerator cards in the accelerator card device list meet the preset physical connectivity conditions; when the interconnection links between valid accelerator cards meet the preset physical connectivity conditions, generate aggregated communication verification parameters based on the logical device index of the valid accelerator cards; call the aggregated communication verification program based on the aggregated communication verification parameters to perform aggregated communication operations on the valid accelerator cards; obtain the aggregated communication verification result output by the aggregated communication verification program; and determine the communication status between the valid accelerator cards as a physical interconnection link abnormal state, an aggregated communication function abnormal state, or a normal communication state based on the interconnection status matrix and the aggregated communication verification result.
[0201] In some implementations, the monitoring module 603, which performs abnormal state identification on abnormal identifiers derived from device enumeration information, is further configured to: The system scans the device enumeration information for a revision error identifier used to indicate device abnormalities. The revision error identifier is a preset invalid value in the device revision field, which is used to characterize an abnormal read of the device configuration space of the corresponding accelerator card. If the revision error identifier exists, the system extracts the device enumeration address containing the revision error identifier. The system performs format normalization on the device enumeration address and determines the logical device index corresponding to the device enumeration address based on the mapping relationship. The system identifies the accelerator card corresponding to the logical device index as an abnormal accelerator card and generates an abnormal status corresponding to the abnormal accelerator card.
[0202] In some implementations, the management module 604 binds the comprehensive status information to the logical device index of the corresponding accelerator card, and updates the accelerator card device list based on the abnormal status identification results in the comprehensive status information, generating accelerator card resource monitoring and management results, which are configured as follows: Based on the mapping relationship, the comprehensive status information is written into the status record associated with the corresponding logical device index; the set of logical device indexes corresponding to the abnormal accelerator card is determined based on the abnormal status identification result; based on the set of logical device indexes, the abnormal accelerator card is filtered from the list of valid accelerator cards, and the logical device index and device identity information corresponding to the abnormal accelerator card are written into the abnormal device record; based on the list of valid accelerator cards, the abnormal device record, and the comprehensive status information corresponding to each accelerator card, the accelerator card resource monitoring and management result is generated.
[0203] In some implementations, after writing the logical device index and device identity information corresponding to the abnormal acceleration card into the abnormal device record, the management module 604 is further configured to: Based on the logical device index in the abnormal device record, a soft reset request for the corresponding abnormal accelerator card is sent to the device management process; the device management process performs a soft reset on the abnormal accelerator card according to the soft reset request; after the soft reset is completed, the device enumeration information of the target computing device is reacquired, and the revised version abnormal identifier corresponding to the abnormal accelerator card is checked based on the reacquired device enumeration information; if the revised version abnormal identifier disappears, the mapping relationship between the physical device identifier and the logical device index of the abnormal accelerator card is re-established, and the abnormal accelerator card is restored to the list of valid accelerator cards; if the revised version abnormal identifier does not disappear, the abnormal accelerator card is retained in the abnormal device record.
[0204] In some implementations, the management module 604 is further configured to: obtain the device file path corresponding to the target accelerator card in the accelerator card device list; obtain the process identifier information occupying the target accelerator card by querying the device file path; read the process name file based on the process identifier information to obtain process name information; read the command line file based on the process identifier information to obtain raw command line data; replace the parameter separator in the raw command line data with spaces to obtain readable command line information; and write the process name information and the readable command line information into the process status record associated with the corresponding logical device index to incorporate the process occupancy status into the comprehensive status information.
[0205] In this embodiment, the aforementioned device acquires first device identification information representing the physical connection relationship of accelerator cards and second device identification information representing the logical registration relationship of accelerator cards from the target computing device. The second device identification information is format-normalized and matched with the first device identification information to establish a mapping relationship between the physical device identification and logical device index of the accelerator cards, forming an accelerator card device list. Simultaneously, each accelerator card is associated with its corresponding device identity information. Further, the aforementioned device monitors the operating status data of each accelerator card in the accelerator card device list. Based on the source of the operating status data, it performs resource status parsing, hot status estimation, interconnection status conversion, and abnormal status identification. Using the mapping relationship as the association benchmark between different data sources, the resource status parsing results, hot status estimation results, and abnormal status identification results are aggregated into the device status record associated with the corresponding logical device index. The interconnection status conversion results are associated with the logical device indices at both ends of the corresponding interconnection link to form comprehensive status information. Based on this, the aforementioned device binds the comprehensive status information to the logical device index of the corresponding accelerator card and updates the accelerator card device list according to the abnormal status identification results, generating accelerator card resource monitoring and management results. This allows for the standardization of device identifiers and multi-source operational status data across different system interfaces, reducing discrepancies between device information and status information, and facilitating the monitoring and location of the accelerator card's resource status, thermal status, interconnection status, and abnormal status.
[0206] The above describes the accelerator card resource monitoring and management device in the embodiments of this application from the perspective of modular functional entities. The following describes the accelerator card resource monitoring and management device in the embodiments of this application from the perspective of hardware processing.
[0207] It should be noted that, Figure 6 The physical device corresponding to the acquisition module 601 shown can be a transceiver, radio frequency circuit, communication module, and input / output (I / O) interface, etc., and the physical device corresponding to the mapping module 602, monitoring module 603, and management module 604 can be a processor.
[0208] Figure 6 The devices shown can all have the following characteristics: Figure 7 The structure shown, when Figure 6 The accelerator card resource monitoring and management device shown has the following features: Figure 7 When the structure shown is used, Figure 7 The processor and transceiver in the device can perform the same or similar functions as the modules provided in the aforementioned device embodiments corresponding to this device. Figure 7 The memory storage processor in the memory needs to call computer programs when executing the above-mentioned accelerator card resource monitoring and management methods.
[0209] This application also relates to a chip system, which includes at least one processor and an interface circuit. The processor includes multiple vector storage units. The processor is used to execute instruction and / or data interaction through the interface circuit, so that the chip system executes the accelerator card resource monitoring and management method of any of the above embodiments.
[0210] In one possible implementation, the chip system may also directly include memory storing computer programs or computer instructions. For example, the memory may be volatile or non-volatile, or may include both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory may be random access memory (RAM), which serves as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DRRAM).
[0211] This application also relates to a processor, which includes multiple storage units for calling computer programs or computer instructions stored in the memory, so that the processor executes the accelerator card resource monitoring and management method described in any of the above embodiments.
[0212] For example, in the embodiments of this application, the processor is an integrated circuit chip with signal processing capabilities. For instance, the processor may be an FPGA, a general-purpose processor, a DSP, an ASIC, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, a SoC, a CPU, a network processor (NP), a microcontroller unit (MCU), a PLD, or other integrated chips, which can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application.
[0213] In one possible implementation, this application also provides a computer-readable storage medium storing program code that, when executed on a computer, causes the computer to perform the above-described method embodiments.
[0214] This application also provides a server; please refer to [link / reference]. Figure 8 , Figure 8 This is a schematic diagram of a server structure provided in an embodiment of this application. The server 1100 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 1122 (e.g., one or more processors) and memory 1132, and one or more storage media 1130 (e.g., one or more mass storage devices) for storing application programs 1142 or data 1144. The memory 1132 and storage media 1130 can be temporary or persistent storage. The program stored in the storage media 1130 may include one or more modules (not shown in the figure), each module may include a series of instruction operations on the server. Furthermore, the CPU 1122 may be configured to communicate with the storage media 1130 and execute the series of instruction operations in the storage media 1130 on the server 1100. Server 1100 may also include one or more power supplies 1126, one or more wired or wireless network interfaces 1150, one or more input / output interfaces 1158, and / or one or more operating systems 1141, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc.
[0215] The steps performed by the server in the above embodiments can be based on this Figure 8 The structure of server 1100 is shown. For example, in the above embodiment, it consists of... Figure 6 The steps performed by the accelerator card resource monitoring and management device shown can be based on this. Figure 8The server structure is shown. For example, the central processing unit 1122 performs the following operations by calling instructions from memory 1132: The input / output interface 1158 is used to obtain first device identification information representing the physical connection relationship of the accelerator card from the target computing device, and second device identification information representing the logical registration relationship of the accelerator card. The central processing unit 1122 establishes a mapping relationship between the physical device identifier and logical device index of the accelerator card based on the format-normalized second device identifier information and the first device identifier information, thereby obtaining an accelerator card device list and associating the accelerator cards in the accelerator card device list with device identity information; monitors the operating status data of each accelerator card in the accelerator card device list; performs resource status parsing, hot status estimation, interconnection status conversion, and abnormal status identification based on the data source of the operating status data, using the mapping relationship as the association benchmark between the operating status data corresponding to different data sources, and collects the resource status parsing results, hot status estimation results, and abnormal status identification results into the device status record associated with the corresponding logical device index, and associates the interconnection status conversion results with the logical device indexes at both ends of the corresponding interconnection link to obtain the comprehensive status information corresponding to each accelerator card; binds the comprehensive status information with the logical device index of the corresponding accelerator card, and updates the accelerator card device list according to the abnormal status identification results in the comprehensive status information, generating accelerator card resource monitoring and management results.
[0216] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0217] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and modules described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0218] In the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, apparatuses, or modules, and may be electrical, mechanical, or other forms.
[0219] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0220] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium.
[0221] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0222] The computer program product includes one or more computer instructions. When the computer program is loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., a solid-state disk (SSD)).
[0223] The technical solutions provided in the embodiments of this application have been described in detail above. Specific examples have been used in the embodiments of this application to illustrate the principles and implementation methods of the embodiments of this application. The description of the above embodiments is only for the purpose of helping to understand the methods and core ideas of the embodiments of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the embodiments of this application. Therefore, the content of this specification should not be construed as a limitation on the embodiments of this application.
Claims
1. A method for monitoring and managing accelerator card resources, characterized in that, The method includes: Obtain first device identification information from the target computing device to represent the physical connection relationship of the accelerator card, and obtain second device identification information to represent the logical registration relationship of the accelerator card; Based on the format-normalized second device identification information and the first device identification information, a mapping relationship between the physical device identification and logical device index of the accelerator card is established to obtain the accelerator card device list, and the accelerator cards in the accelerator card device list are associated with device identity information. Monitor the operating status data of each accelerator card in the accelerator card device list; based on the data source of the operating status data, perform resource status analysis, hot status estimation, interconnection status conversion and abnormal status identification respectively, use the mapping relationship as the correlation benchmark between the operating status data corresponding to different data sources, collect the resource status analysis results, hot status estimation results and abnormal status identification results into the device status record associated with the corresponding logical device index, and associate the interconnection status conversion results with the logical device indexes at both ends of the corresponding interconnection link to obtain the comprehensive status information corresponding to each accelerator card; The comprehensive status information is bound to the logical device index of the corresponding accelerator card, and the accelerator card device list is updated according to the abnormal status identification results in the comprehensive status information to generate accelerator card resource monitoring and management results.
2. The method for monitoring and managing accelerator card resources according to claim 1, characterized in that, The step of obtaining first device identification information from the target computing device to represent the physical connection relationship of the accelerator cards, and obtaining second device identification information to represent the logical registration relationship of the accelerator cards, includes: Traverse the logical device nodes corresponding to each accelerator card in the virtual file system of the target computing device; Read the bus address information of the corresponding accelerator card from the logical device node, and use the bus address information as the first device identification information; Call the bus enumeration tool to obtain the enumeration output information of each accelerator card in the target computing device; Extract the enumeration address information of the corresponding accelerator card from the enumeration output information, and use the enumeration address information as the second device identification information; The first device identification information and the second device identification information respectively correspond to the device identification presented by the same accelerator card under different system interfaces.
3. The method for monitoring and managing accelerator card resources according to claim 1, characterized in that, The mapping relationship between the physical device identifier and logical device index of the accelerator card is established based on the format-normalized second device identifier information and the first device identifier information to obtain the accelerator card device list, including: The first device identification information is indexed and extracted to obtain the logical device index corresponding to each first device identification information, and a first mapping table is established. The first mapping table is used to represent the correspondence between the first device identification information and the logical device index. The second device identification information is format-normalized to ensure that the second device identification information has the same address format as the first device identification information. The format-normalized second device identifier information is matched with the first device identifier information in the first mapping table to determine the logical device index corresponding to the second device identifier information; The successfully matched accelerator cards are sorted according to the logical device index to generate the accelerator card device list.
4. The method for monitoring and managing accelerator card resources according to claim 3, characterized in that, The step of matching the format-normalized second device identifier information with the first device identifier information in the first mapping table to determine the logical device index corresponding to the second device identifier information includes: Extract the bus device function address without the field prefix from the enumerated output information; Determine whether the bus device function address and the first device identification information have the same address field; If the bus device function address does not contain the field identifier of the first device identifier information, then a preset field prefix is added to the bus device function address to obtain the second device identifier information after format normalization; Using the normalized second device identifier information as the query key, a search operation is performed in the first mapping table; If a first device identifier that matches the second device identifier after format normalization is found in the first mapping table, then the logical device index corresponding to the first device identifier is determined as the logical device index corresponding to the second device identifier.
5. The accelerator card resource monitoring and management method according to claim 3, characterized in that, After sorting the successfully matched accelerator cards according to the logical device index to generate the accelerator card device list, the method further includes: Based on the second device identification information, obtain the link status output information of the corresponding acceleration card; Extract the current link rate and current link width from the link status output information; The current link rate and the current link width are associated with the logical device index of the corresponding accelerator card and written into the accelerator card device list.
6. The method for monitoring and managing accelerator card resources according to claim 1, characterized in that, Associating accelerator cards in the accelerator card device list with device identity information includes: The device information reading program is invoked to obtain the device identity information corresponding to each accelerator card; The serial number information corresponding to the logical device index is parsed from the output of the device information reading program. Based on the logical device index, the serial number information is bound to the corresponding accelerator card in the accelerator card device list; The device identification information includes at least the serial number information used to distinguish different physical accelerator cards.
7. The method for monitoring and managing accelerator card resources according to claim 1, characterized in that, The data sources based on the operational status data respectively perform resource status analysis, hot status estimation, interconnection status transformation, and abnormal status identification. Using the mapping relationship as the correlation benchmark between operational status data from different data sources, the resource status analysis results, hot status estimation results, and abnormal status identification results are aggregated into the device status records associated with the corresponding logical device indexes. Furthermore, the interconnection status transformation results are associated with the logical device indexes at both ends of the corresponding interconnection link to obtain the comprehensive status information corresponding to each accelerator card, including: Perform resource status parsing on storage resource data from the virtual file system to obtain the storage resource usage status of each accelerator card; Power consumption status analysis is performed on the power consumption data from the sensor acquisition program to obtain the single power consumption status of each accelerator card and the total power consumption status of the accelerator card group. Thermal state calculations are performed on the power consumption data and power domain operating parameters obtained from the sensor acquisition program to obtain the thermal state of each accelerator card. The firmware output data from the firmware information program is parsed in groups to obtain the clock state and firmware version state of each accelerator card. Perform interconnect state transitions on interconnect detection data from the link detection program to obtain the interconnect state between each accelerator card; Perform abnormal status identification on the abnormal identifiers derived from the device enumeration information to obtain the abnormal status corresponding to the abnormal acceleration card; The storage resource usage status, single card power consumption status, total power consumption status, thermal status, clock status, firmware version status, interconnection status, and abnormal status are associated according to the corresponding logical device index or interconnection link to generate the comprehensive status information.
8. The method for monitoring and managing accelerator card resources according to claim 7, characterized in that, The thermal state calculation of power consumption data and power domain operating parameters obtained from the sensor acquisition program includes: The output results of the sensor acquisition program are scanned line by line, and the single card power consumption, the total power consumption of the accelerator card group, and the target power domain operating current of the target processor unit in each accelerator card are extracted from the output results. Perform anti-repeated sampling processing on the target power domain operating current of the same accelerator card so that the thermal state estimation of the same accelerator card adopts the target power domain operating current matched for the first time. The target power domain operating current is input into a preset temperature calibration model to obtain the temperature state of the corresponding accelerator card. The preset temperature calibration model is used to represent the correspondence between the target power domain operating current and the temperature state, and is calibrated based on the target power domain operating current and the corresponding measured temperature collected under different operating states.
9. The method for monitoring and managing accelerator card resources according to claim 7, characterized in that, The process of performing resource status parsing on storage resource data originating from the virtual file system includes: Write the target memory pool type corresponding to high-bandwidth memory to the memory pool type node corresponding to the accelerator card in the virtual file system to trigger the driver to switch to the corresponding memory statistics type; Read the memory status node in the virtual file system corresponding to the target memory pool type to obtain the raw memory status data; The raw memory status data is formatted and parsed to obtain the free memory, used memory, and total memory of the corresponding accelerator card. The storage resource usage status is generated based on the amount of free memory, the amount of used memory, and the total amount of memory.
10. The method for monitoring and managing accelerator card resources according to claim 7, characterized in that, The step of performing packet parsing on firmware output data from the firmware information program includes: The output of the firmware information program is scanned line by line. When an accelerator card group identifier is detected, the logical device index corresponding to the current group is determined based on the accelerator card group identifier; when the accelerator card group identifier does not contain a logical device index, the current group count value is updated according to a preset group order, and the current group count value is determined as the logical device index corresponding to the current group. Extract the core clock frequency, memory clock frequency, and firmware version information within the current group; Associate the core clock frequency, the memory clock frequency, and the firmware version information with the corresponding logical device index within the current group.
11. The method for monitoring and managing accelerator card resources according to claim 7, characterized in that, The step of performing interconnection state transitions on interconnection detection data originating from the link detection program includes: Obtain the preset connection relationship between accelerator cards, wherein the preset connection relationship is used to represent the connection correspondence between different accelerator card ports; An interconnection state matrix is constructed based on the preset connection relationship. The row index and column index of the interconnection state matrix correspond to the logical device index of the accelerator card, respectively. The state value in the interconnection state matrix is used to represent the interconnection link status between two corresponding logical device indices. Extract abnormal port information from the output of the link detection program. The abnormal port information includes at least one of port error information and link connectivity status information. Based on the abnormal port information, the corresponding interconnection link is determined in the preset connection relationship; Update the state value in the interconnection state matrix corresponding to the logical device indices at both ends of the interconnection link; The interconnection status between the accelerator cards is obtained based on the updated interconnection status matrix.
12. The method for monitoring and managing accelerator card resources according to claim 11, characterized in that, The process of obtaining the interconnection state between the accelerator cards based on the updated interconnection state matrix includes: Obtain a preset character layout template, which includes multiple interconnected edge placeholders; Determine the connection relationship of the accelerator card port corresponding to each interconnection edge placeholder; The interconnection link status corresponding to each interconnection edge placeholder is determined based on the interconnection state matrix. Configure distinguishable display attributes for the corresponding interconnection edge placeholders according to the interconnection link status, and convert the interconnection edge placeholders into interconnection edge characters with the display attributes; Output interconnection topology status information according to the preset character layout template.
13. The method for monitoring and managing accelerator card resources according to claim 11, characterized in that, After updating the state value corresponding to the logical device indexes at both ends of the interconnection link in the interconnection state matrix, the method further includes: Based on the updated interconnection status matrix, determine whether the interconnection links between valid accelerator cards in the accelerator card device list meet the preset physical connectivity conditions; When the interconnection link between the valid accelerator cards meets the preset physical connectivity condition, a set of communication verification parameters is generated based on the logical device index of the valid accelerator cards. The collective communication verification program is invoked based on the aforementioned collective communication verification parameters to perform collective communication operations on the valid accelerator card; Obtain the set communication verification result output by the set communication verification program; Based on the interconnection state matrix and the aggregated communication verification results, the communication state between the valid accelerator cards is determined to be either a physical interconnection link abnormal state, an aggregated communication function abnormal state, or a normal communication state.
14. The method for monitoring and managing accelerator card resources according to claim 7, characterized in that, The step of performing abnormal state identification on abnormal identifiers derived from device enumeration information includes: Scan the device enumeration information to see if there is a revision version anomaly identifier used to indicate device anomalies. The revision version anomaly identifier is a preset invalid value in the device revision version field. The preset invalid value is used to characterize the device configuration space reading anomaly of the corresponding accelerator card. If the revision version anomaly identifier exists, extract the device enumeration address containing the revision version anomaly identifier; The device enumeration addresses are format-normalized, and the logical device index corresponding to the device enumeration addresses is determined based on the mapping relationship. The accelerator card corresponding to the logical device index is identified as an abnormal accelerator card, and an abnormal state corresponding to the abnormal accelerator card is generated.
15. The method for monitoring and managing accelerator card resources according to claim 1, characterized in that, The process of binding the comprehensive status information with the logical device index of the corresponding accelerator card, and updating the accelerator card device list based on the abnormal status identification results in the comprehensive status information to generate accelerator card resource monitoring and management results includes: According to the mapping relationship, the comprehensive status information is written into the status record associated with the corresponding logical device index; The set of logical device indexes corresponding to the abnormal acceleration card is determined based on the abnormal state identification results. Based on the logical device index set, abnormal acceleration cards are filtered from the list of valid acceleration cards, and the logical device index and device identity information corresponding to the abnormal acceleration card are written into the abnormal device record. Based on the list of valid accelerator cards, the records of abnormal devices, and the comprehensive status information corresponding to each accelerator card, the accelerator card resource monitoring and management results are generated.
16. The method for monitoring and managing accelerator card resources according to claim 15, characterized in that, After writing the logical device index and device identity information corresponding to the abnormal acceleration card into the abnormal device record, the method further includes: Based on the logical device index in the abnormal device record, a soft reset request for the corresponding abnormal acceleration card is sent to the device management process. The device management process performs a soft reset on the abnormal acceleration card based on the soft reset request; After the soft reset is completed, the device enumeration information of the target computing device is reacquired, and the revised version abnormality identifier corresponding to the abnormal accelerator card is checked based on the reacquired device enumeration information. If the revised version error identifier disappears, the mapping relationship between the physical device identifier and logical device index of the error accelerator card is re-established, and the error accelerator card is restored to the list of valid accelerator cards. If the revised version's anomaly identifier does not disappear, the anomaly acceleration card will be retained in the anomaly device record.
17. The method for monitoring and managing accelerator card resources according to claim 1, characterized in that, The method further includes: For the target accelerator card in the accelerator card device list, obtain the device file path corresponding to the target accelerator card; The process identifier information occupying the target accelerator card is obtained by querying the device file path; Based on the process identifier information, the process name file is read to obtain the process name information; Read the command line file based on the process identifier information to obtain the raw command line data; replace the parameter separators in the raw command line data with spaces to obtain readable command line information. The process name information and the readable command line information are written into the process status record associated with the corresponding logical device index, so as to incorporate the process occupancy status into the comprehensive status information.
18. A device for monitoring and managing accelerator card resources, characterized in that, The device includes: The acquisition module is configured to acquire first device identification information representing the physical connection relationship of the accelerator card from the target computing device, and to acquire second device identification information representing the logical registration relationship of the accelerator card. The mapping module is configured to establish a mapping relationship between the physical device identifier and the logical device index of the accelerator card based on the format-normalized second device identifier information and the first device identifier information, to obtain an accelerator card device list, and associate the accelerator cards in the accelerator card device list with device identity information. The monitoring module is configured to monitor the operating status data of each accelerator card in the accelerator card device list; based on the data source of the operating status data, it performs resource status analysis, hot status estimation, interconnection status conversion and abnormal status identification respectively; using the mapping relationship as the correlation benchmark between the operating status data corresponding to different data sources, it collects the resource status analysis results, hot status estimation results and abnormal status identification results into the device status record associated with the corresponding logical device index, and associates the interconnection status conversion results with the logical device indexes at both ends of the corresponding interconnection link to obtain the comprehensive status information corresponding to each accelerator card; The management module is configured to bind the comprehensive status information with the logical device index of the corresponding accelerator card, and update the accelerator card device list based on the abnormal status identification results in the comprehensive status information, thereby generating accelerator card resource monitoring and management results.
19. A chip, characterized in that, The chip includes a processor coupled to a transceiver for executing the accelerator card resource monitoring and management method as described in any one of claims 1-17.
20. A computer-readable storage medium, characterized in that, The method includes instructions that, when executed on a computer, cause the computer to perform the accelerator card resource monitoring and management method as described in any one of claims 1-17.