Server information collection, updating method and device, system, medium, product and equipment

By loading the target image into the server and automatically collecting hardware information, the problem of low efficiency and low accuracy of manual data collection is solved, the construction efficiency and scalability of large-scale clusters are improved, and the operation and maintenance complexity is reduced.

CN120915798BActive Publication Date: 2026-04-07CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-25
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In large-scale clusters, existing technologies rely on manual collection of server parameter information, which leads to low efficiency and low accuracy, affecting construction efficiency and scalability, and increasing operation and maintenance complexity.

Method used

By loading the target image into the target server, the system automatically collects hardware information, including the network card information of the NPU, and combines it with LLDP information to determine network topology parameters, and automatically updates the management database.

Benefits of technology

It enables efficient and accurate collection of server hardware information, improves the construction efficiency and scalability of large-scale clusters, and reduces the complexity of operation and maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120915798B_ABST
    Figure CN120915798B_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, system, medium, product, and device for server information collection and updating. The method includes: acquiring a target image; loading the acquired target image into a target server, wherein the target server includes a neural network processor (NPU); and collecting hardware information of the target server through the loaded target image, wherein the hardware information includes the network interface card (NIC) information of the NPU. This improves the efficiency and accuracy of server hardware information collection, thereby increasing the construction efficiency and scalability of large-scale clusters and reducing the operational complexity of the cluster.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, in particular to a server information acquisition and updating method and device, system, medium, product and equipment. BACKGROUND

[0002] In cloud computing or computing power network, a large number of servers are usually deployed to form a large-scale cluster to provide computing resources, so as to facilitate users to use the computing resources. However, such a large-scale cluster often needs to update the servers contained therein (such as replacing, reducing or increasing servers), and after each update, it is necessary to collect the relevant parameter information of the updated servers in the cluster to ensure the normal use of the cluster.

[0003] In the related art, the relevant parameter information of the updated servers is usually collected by artificial means, which is low in efficiency and accuracy, thereby seriously reducing the construction efficiency and scalability of the large-scale cluster and increasing the operation and maintenance complexity of the cluster. SUMMARY

[0004] To solve the above technical problems, the embodiments of the present application provide a server information acquisition and updating method and device, system, medium, product and equipment, which can improve the collection efficiency and accuracy of server hardware information, thereby improving the construction efficiency and scalability of the large-scale cluster and reducing the operation and maintenance complexity of the cluster.

[0005] In a first aspect, the embodiments of the present application provide a server information acquisition method, comprising:

[0006] obtaining a target image;

[0007] loading the obtained target image in a target server, wherein the target server comprises a neural network processor (NPU);

[0008] collecting hardware information of the target server through the loaded target image, wherein the hardware information comprises network card information of the NPU.

[0009] Optionally, the obtaining a target image comprises:

[0010] obtaining a temporary network address after the target server enters a specified startup mode;

[0011] obtaining the target image based on the temporary network address.

[0012] Optionally, the collecting hardware information of the target server through the loaded target image comprises:

[0013] acquire hardware information of the target server by running a preset agent service in the loaded target image.

[0014] Optionally, the preset agent service is configured to invoke a network management tool to acquire the hardware information of the target server during running of the preset agent service, the network management tool corresponding to the NPU.

[0015] Optionally, the target server is a bare-metal server.

[0016] Optionally, the hardware information further includes at least one of:

[0017] central processing unit (CPU) information;

[0018] memory information;

[0019] disk information.

[0020] Optionally, the network card information includes link layer discovery protocol (LLDP) information of the NPU, the NPU being a plurality of, the plurality of NPU communicating via a parameter plane network, and the method further includes:

[0021] determining network topology parameters related to the parameter plane network based on the LLDP information;

[0022] sending the network topology parameters to a management and control node corresponding to the target server.

[0023] Optionally, the network topology parameters include at least one of:

[0024] media access control (MAC) address of a network card of the NPU;

[0025] parameter information of a switch associated with the network card of the NPU;

[0026] interface parameter information of the network card of the NPU.

[0027] Optionally, the network topology parameters are used to instruct the management and control node to perform a preset operation on the target server, the preset operation including at least one of: node information management, port information management.

[0028] In a second aspect, an embodiment of the present application provides a server information updating method, applicable to a management and control node, and the method includes:

[0029] receiving an inspection request for a target server, wherein the target server includes an NPU;

[0030] determining a server self-check task in response to the inspection request;

[0031] based on the server self-checking task, control the target server to enter a specified startup mode, wherein the target server is configured to: after entering the specified startup mode, acquire a target image; load the acquired target image; and collect hardware information of the target server through the loaded target image, wherein the hardware information includes network card information of the NPU.

[0032] receive network topology parameters determined according to the hardware information from the target server, and perform a preset operation on the target server according to the network topology parameters, the preset operation including at least one of the following: node information management, port information management.

[0033] In a third aspect, an embodiment of the present application provides a server information collection system, comprising:

[0034] a management node configured to receive an inspection request for a target server; determine a server self-checking task in response to the inspection request; based on the server self-checking task, control the target server to enter a specified startup mode; and

[0035] the target server including an NPU, the target server being configured to, after entering the specified startup mode, acquire a target image; load the acquired target image; and collect hardware information of the target server through the loaded target image, wherein the hardware information includes network card information of the NPU.

[0036] In a fourth aspect, an embodiment of the present application provides a server information collection device, comprising:

[0037] an image acquisition module configured to acquire a target image;

[0038] an image loading module configured to load the acquired target image in a target server, wherein the target server includes a neural network processor (NPU);

[0039] a collection module configured to collect hardware information of the target server through the loaded target image, wherein the hardware information includes network card information of the NPU.

[0040] In a fifth aspect, an embodiment of the present application provides a server information updating device suitable for a management node, comprising:

[0041] a request receiving module configured to receive an inspection request for a target server, wherein the target server includes an NPU;

[0042] a task determining module configured to determine a server self-checking task in response to the inspection request;

[0043] The mode control module is used to control the target server to enter a specified startup mode based on the server self-test task. The target server is configured to: acquire a target image after entering the specified startup mode; load the acquired target image; and collect the hardware information of the target server through the loaded target image, wherein the hardware information includes the network card information of the NPU.

[0044] The management module is used to receive network topology parameters determined based on the hardware information from the target server, and to perform preset operations on the target server based on the network topology parameters. The preset operations include at least one of the following: node information management and port information management.

[0045] Sixthly, embodiments of this application provide a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in any of the preceding claims.

[0046] In a seventh aspect, embodiments of this application provide a computer program product, including computer instructions that, when executed by a processor, implement the steps of the method described in any of the preceding claims.

[0047] Eighthly, embodiments of this application provide a computer device including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the steps of the method described in any of the preceding claims.

[0048] In summary, the embodiments of this application have at least the following beneficial effects:

[0049] By using the embodiments of this application, the obtained target image is loaded into the target server containing the NPU, so that the hardware information of the target server can be automatically collected by loading the target image into the target server. This can improve the efficiency of collecting server hardware information. In particular, hardware information such as NPU network card information, which has a significant impact on the performance of servers and large-scale clusters, is also collected. This can improve the accuracy of collecting key hardware information of the server, thereby improving the construction efficiency and scalability of large-scale clusters and reducing the operation and maintenance complexity of the cluster. Attached Figure Description

[0050] Figure 1 This is a flowchart illustrating the server information collection method provided in an embodiment of this application;

[0051] Figure 2 This is a schematic diagram illustrating server information updates provided in an embodiment of this application;

[0052] Figure 3This is a flowchart illustrating the server information update method provided in an embodiment of this application;

[0053] Figure 4 This is a schematic diagram of the server information collection system provided in an embodiment of this application;

[0054] Figure 5 This is a schematic diagram of the server information collection device provided in the embodiments of this application;

[0055] Figure 6 This is a schematic diagram of the server information updating device provided in the embodiments of this application;

[0056] Figure 7 This is a schematic diagram of the computer device provided in the embodiments of this application. Detailed Implementation

[0057] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments / examples are only a part of the embodiments / examples of this application, and not all of the embodiments / examples. Based on the embodiments / examples in this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0058] In the description of this application, the terms "first," "second," "third," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first," "second," "third," etc., may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "multiple" means two or more. In the description of this application, the term "comprising" and its variations are open-ended, meaning "including but not limited to." The term "based on" means "at least partially based on." The term "according to" means "at least partially according to." The term "one embodiment / example" means "at least one embodiment / example"; the term "another embodiment / example" means "at least one additional embodiment / example"; the term "some embodiments / examples" means "at least some embodiments / examples."

[0059] In the description of this application, it should be noted that, unless otherwise expressly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.

[0060] In the description of this application, it should be noted that, unless otherwise defined, all technical and scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in this application is for the purpose of describing specific embodiments only and is not intended to limit the application. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.

[0061] The following explains some terms and concepts involved in the embodiments of this application.

[0062] A Neural Processing Unit (NPU) is a specialized hardware processor designed specifically to accelerate computations in Artificial Intelligence (AI) (such as neural networks in deep learning).

[0063] Link Layer Discovery Protocol Information (LLDP) refers to the set of data that network devices actively send and receive through the Link Layer Discovery Protocol to describe their own identity and capabilities.

[0064] Ironic refers to the bare metal management service in the OpenStack architecture. Its control plane components are generally deployed in the management nodes used to manage the corresponding servers. For example, ironic-api is the external interface gateway of Ironic, which receives all bare metal operation requests (such as inspection, deployment, and deletion) and forwards them to the backend components; ironic-conductor is the task scheduling core of Ironic, responsible for executing specific bare metal operations (such as controlling server power on / off and coordinating the inspection process); ironic-inspector is a dedicated service for bare metal hardware information detection, which can be used to trigger PXE boot, receive hardware data reported by ironic-python-agent, and process and standardize information.

[0065] In some example scenarios, with the increasing prevalence of cloud computing, cloud-based applications have become a trend. In the cloud computing field, Infrastructure as a Service (IaaS) technology is very mature and widely used in the industry. However, in some cases, users may require more control, more hardware access, higher performance, stronger security isolation, and / or the ability to choose their operating environment. Bare-metal servers can provide tenants (users) with a near-native computing experience, compensating for the significant performance degradation of virtualized instances in related technologies. Therefore, in cloud computing and computing power networks, bare-metal servers are provided to users as an important basic computing resource.

[0066] In related technologies, the configuration of the parameter plane network in bare metal server deployment heavily relies on manual operation. Typically, during the construction phase, hardware integration (HI) personnel manually collect the physical information of each NPU network interface card (NIC) and document it. Subsequently, software integration (SI) personnel, based on this document, manually create bare metal node-ports and enter this information into the Ironic bare metal management database. When the bare metal server is activated, the Ironic service passes this pre-entered hardware information as parameters to the Neutron network service to complete the parameter plane network configuration. This parameter plane network (or management network, control network) can provide a stable, isolated, and low-latency communication channel for the server and / or devices within the server (such as NPUs), and can also be used to manage and control the server at various stages of its lifecycle.

[0067] However, the aforementioned technologies are highly manual and inefficient. Each time new equipment is deployed, HI and SI personnel must repeatedly collect and input information, making the process cumbersome and severely hindering the construction efficiency of large-scale clusters. Furthermore, since a server may contain multiple NPU cards, an error in inputting information for any one NPU card can lead to network configuration failures for the entire machine. In other words, the accuracy of manual input is difficult to guarantee, resulting in a high error rate and poor reliability. Moreover, if an error in information causes a server to fail to activate or experience network anomalies, troubleshooting becomes extremely difficult, requiring repeated verification across teams (HI / SI / Operations), which is time-consuming and labor-intensive, severely impacting the construction cycle and operational experience, increasing operational complexity. Furthermore, because manual methods are difficult to adapt to the needs of rapid scaling and automated operations, the scalability of the cluster / server is affected.

[0068] Firstly, see [the following]Figure 1 The diagram shows a schematic flowchart of a server information collection method provided in an embodiment of this application. This server information collection method can be applied to a target server. The method includes S101-S103, as detailed below.

[0069] S101, Obtain the target image.

[0070] In some examples, the target image can be obtained by downloading the relevant files for the target image.

[0071] In some examples, the target image mentioned above can refer to a streamlined Linux system, a miniature operating system that can be used for hardware self-testing.

[0072] S102, Load the acquired target image into the target server, wherein the target server includes a neural network processor (NPU).

[0073] In some examples, the target server may integrate one or more NPUs.

[0074] In some examples, the target server can be configured to automatically acquire and load the target image after booting through the PrebooteXecution Environment (PXE). This PXE boot can be implemented by the management node through invoking relevant services. For instance, after receiving an inspection request from the administrator for the target server via a relevant Application Programming Interface (API), the management node can respond to the inspection request by controlling the target server to enter the PXE boot process using out-of-band management tools.

[0075] S103, collect the hardware information of the target server through the loaded target image, wherein the hardware information includes the network card information of the NPU.

[0076] In some examples, the acquired target image can be directly loaded into the target server, allowing the target server to perform a self-test (hardware self-test) using the loaded target image, thereby automatically collecting hardware information about the target server. Here, the target image may integrate self-test programs to control the hardware self-test, which can be configured to run automatically within the loaded target image, thus achieving automatic collection of hardware information about the target server.

[0077] In some examples, the target image loaded above can run in the memory of the target server. In this embodiment, the target image runs in memory, thus not relying on the target server's local disk system. This allows it to run on a brand-new server without an operating system deployed, while also avoiding interference with servers with deployed systems. This feature also enables the target image to be applied in parallel to batch collection on hundreds or thousands of servers, supporting the rapid construction of large-scale clusters. At the same time, the collection process is fully automated, reducing the maintenance costs (such as troubleshooting and configuration synchronization) caused by manual intervention and lowering the complexity of cluster maintenance.

[0078] In one optional implementation, obtaining the target image includes:

[0079] After the target server enters the specified startup mode, a temporary network address is obtained;

[0080] The target image is obtained based on the temporary network address.

[0081] In some examples, the specified boot mode can include the PXE boot described above.

[0082] In some examples, the aforementioned temporary network address may include a temporary Internet Protocol (IP) address, so that the target server can obtain a temporary IP address from a Dynamic Host Configuration Protocol (DHCP) server (see, for example, [link to relevant documentation]). Figure 2 The target server sends an address acquisition request to the DHCP-Server to request the temporary network address. Here, the DHCP-Server can be used to provide DHCP services for ironic's dnsmasq (a lightweight, open-source software that integrates DHCP and TFTP functionality).

[0083] See in some examples Figure 2 Once the target server obtains the temporary network address, it can be considered a node in the network, and thus obtain the target image from the Trivial File Transfer Protocol (TFTP) server through the temporary network address.

[0084] In this embodiment, the program code for the specified startup mode acts as a boot file, automatically guiding the target server to obtain and load the target image when the file is invoked. This achieves automated booting without manual intervention, allowing the target server to automatically load the target image from the network without manual operation, skipping the traditional manual startup steps. This fully automated process avoids the efficiency loss of manual operation on each machine in a large-scale cluster, directly improving data acquisition efficiency. Furthermore, the acquired target image corresponds to pre-defined startup parameters (such as image loading path, temporary network configuration, etc.) to ensure that all target servers enter the acquisition environment with the same process and parameters. This standardization avoids configuration deviations that may occur during manual operation (such as incorrect startup methods or missing parameters), laying the foundation for the accuracy of subsequent hardware information acquisition.

[0085] In one optional implementation, the step of collecting the hardware information of the target server through the loaded target image includes:

[0086] The hardware information of the target server is collected by running a preset proxy service in the loaded target image.

[0087] In some examples, the default proxy service may include an Ironic Python Agent (IPA) service, which, when run, can automatically collect hardware information of the target server. The loaded target image can automatically run this default proxy service on it.

[0088] In one optional implementation, the preset proxy service is used to invoke a network management tool during its operation to collect hardware information of the target server, the network management tool corresponding to the NPU.

[0089] In some examples, the network management tool can interact with the network controller on the NPU, actively sending LLDP packets (i.e., LLDP requests) into the network, and receiving LLDP responses carrying the LLDP information from peer devices (such as the switch associated with / communicated with the aforementioned NPU), thereby collecting the NPU's network interface card (NIC) information. For example, the NPU may include an Ascend NPU card, such as the Ascend 910B. Thus, the network management tool invoked by the default proxy service may include the hccn_tool tool for the Ascend NPU card, which can interact with the network controller on the NPU card and actively send LLDP packets into the network. It is easy to understand that, as described above, the network management tool here can actually be determined by the NPU (e.g., by the type and model of the NPU).

[0090] In one alternative implementation, the target server is a bare metal server.

[0091] It should be noted that a bare metal server is a physical dedicated server that runs directly on the hardware, generally without relying on any virtualization layer (such as a hypervisor), and does not share underlying computing resources with other tenants. In this embodiment, compared to traditional physical servers, bare metal servers represent a modern evolution of traditional physical servers in the cloud computing era, significantly improving performance, security, and programmable management capabilities.

[0092] In one alternative implementation, the hardware information further includes at least one of the following:

[0093] Central Processing Unit (CPU) information; in some examples, this CPU information may include key performance parameters such as CPU model, architecture, number of cores and / or frequency.

[0094] Memory information; in some examples, this memory information may include information such as the capacity, type, speed, and / or slot distribution of physical memory.

[0095] Disk information. In some examples, this disk information may include information such as the type, capacity, interface, and / or health status of the local storage device.

[0096] In one optional implementation, the network interface card (NIC) information includes the Link Layer Discovery Protocol (LLDP) information of the NPU, the number of NPUs is multiple, and the NPUs communicate with each other via a parameter plane network. The method further includes:

[0097] Based on the LLDP information, determine the network topology parameters related to the parametric plane network;

[0098] The network topology parameters are sent to the management node corresponding to the target server.

[0099] In some examples, the LLDP information of the NPU can be collected by actively sending LLDP packets (i.e., LLDP requests) into the network and receiving LLDP responses carrying LLDP information from peer devices (such as the switches associated with / communicated with the NPU mentioned above).

[0100] In some examples, the LLDP information can be parsed to extract specific, corresponding network topology parameters.

[0101] In one optional implementation, the network topology parameters include at least one of the following:

[0102] The NPU's network interface card (NIC) has a Media Access Control (MAC) address. In some examples, this MAC address is called a Media Access Control address, and can also be referred to as a LAN address, MAC address, Ethernet address, hardware address, or physical address. This MAC address is actually the adapter address or adapter identifier EUI-48, used to indicate the location of the network device. A MAC address is used to uniquely identify a NIC within a network. If a device has one or more NICs, each NIC needs and will have a unique MAC address.

[0103] The parameter information of the switch associated with the network card of the NPU; in some examples, the switch may include the uplink switch associated with the network card of the NPU, and the parameter information of the switch may include at least one of the following: information of the port where the switch is plugged in, the name of the switch, and the management IP address of the switch.

[0104] The interface parameter information of the NPU's network interface card. In some examples, this interface parameter information may include the port VLAN ID (PVID) configuration information of the NPU's network interface card interface, where VLAN refers to Virtual Local Area Network. This PVID configuration is used to set the PVID for the physical network port connected to the NPU (Neural Processor) network interface card to ensure that the port can correctly send and receive data frames of the specified VLAN when connected to the switch.

[0105] In one optional implementation, the network topology parameters are used to instruct the control node to perform a preset operation on the target server, the preset operation including at least one of the following: node information management and port information management.

[0106] In some examples, after obtaining the network topology parameters, these parameters can be reported to the management node; for example, see [link to example]. Figure 2This allows the IPA service in the target server (via API call) to call the ironic-inspector interface to report network topology parameters. The ironic-inspector in the management node can then use the ironicclient to call the ironic-api to update or create nodes and / or ports and / or portgroups in the management database, thus completing node and / or port information management. It should be noted that the creation described in this embodiment is a self-discovery process; it does not require prior node creation, but only requires configuring the bare metal server to boot from PXE.

[0107] Thus, this embodiment can achieve automated registration / updating of network-related information for this parameter in the management database.

[0108] Secondly, see Figure 3 The diagram shows a flowchart of a server information update method provided in an embodiment of this application. This server information update method is applicable to management nodes and includes steps S301-S304, as detailed below.

[0109] S301, Receive an inspection request for a target server, wherein the target server includes an NPU;

[0110] S302, in response to the inspection request, determine the server self-test task;

[0111] S303, based on the server self-test task, control the target server to enter a specified startup mode, wherein the target server is configured to: after entering the specified startup mode, acquire a target image; load the acquired target image; and collect the hardware information of the target server through the loaded target image, wherein the hardware information includes the network card information of the NPU;

[0112] S304, receive network topology parameters determined based on the hardware information from the target server, and perform a preset operation on the target server based on the network topology parameters. The preset operation includes at least one of the following: node information management and port information management.

[0113] In some examples, users (administrators) can input an inspection request for a target server through the ironic client, which will then be sent to the ironic-api, and the control node can receive the inspection request via the ironic-api.

[0114] In some examples, after receiving the inspection request, the ironic-api in the control node can determine the server self-check task corresponding to the inspection request (for example, based on the target server of the inspection request, determine the target server that needs to be inspected, and thus obtain the server self-check task for that target server). This ironic-api can be accessed via Remote Procedure Call (RPC). Figure 2 The RPC call in the process forwards the server self-test task to the ironic-conductor in the management node, so that the ironic-conductor can further call the ironic-inspector in the management node through API call to execute the server self-test task. The ironic-inspector can then call the ironic-api through the ironicclient to control the target server to enter a specified startup mode (such as PXE restart, i.e., the restart action) via the out-of-band control command of ipmitool.

[0115] In some examples, the ironic-inspector of the control node can also call the ironic-api via ironicclient to control the shutdown of the target server.

[0116] In one optional implementation, obtaining the target image includes:

[0117] After the target server enters the specified startup mode, a temporary network address is obtained;

[0118] The target image is obtained based on the temporary network address.

[0119] In one optional implementation, the step of collecting the hardware information of the target server through the loaded target image includes:

[0120] The hardware information of the target server is collected by running a preset proxy service in the loaded target image.

[0121] In one optional implementation, the preset proxy service is used to invoke a network management tool during its operation to collect hardware information of the target server, the network management tool corresponding to the NPU.

[0122] In one alternative implementation, the target server is a bare metal server.

[0123] In one alternative implementation, the hardware information further includes at least one of the following:

[0124] Central Processing Unit (CPU) information;

[0125] Memory information;

[0126] Disk information.

[0127] In one optional implementation, the network interface card (NIC) information includes the Link Layer Discovery Protocol (LLDP) information of the NPU, wherein there are multiple NPUs, and the NPUs communicate with each other via a parameter plane network, wherein the network topology parameters are related to the parameter plane network and determined by the LLDP information.

[0128] In one optional implementation, the network topology parameters include at least one of the following:

[0129] The media access control MAC address of the network interface card of the NPU;

[0130] Parameter information of the switch associated with the network interface card of the NPU;

[0131] The interface parameter information of the network card of the NPU.

[0132] Thirdly, see Figure 4 The diagram shows a schematic representation of a server information collection system provided in an embodiment of this application. The server information collection system includes:

[0133] Control node 401 is configured to receive an inspection request for target server 402; in response to the inspection request, determine a server self-check task; and based on the server self-check task, control target server 402 to enter a specified startup mode; and

[0134] The target server 402 includes an NPU and is configured to, after entering a specified boot mode, acquire a target image; load the acquired target image; and collect hardware information of the target server 402 through the loaded target image, wherein the hardware information includes the network interface card information of the NPU.

[0135] In one optional implementation, obtaining the target image includes:

[0136] After the target server 402 enters the specified startup mode, a temporary network address is obtained;

[0137] The target image is obtained based on the temporary network address.

[0138] In one optional implementation, the step of collecting the hardware information of the target server 402 through the loaded target image includes:

[0139] The hardware information of the target server 402 is collected by running a preset proxy service in the loaded target image.

[0140] In one optional implementation, the preset proxy service is used to invoke a network management tool during its operation to collect hardware information of the target server 402, the network management tool corresponding to the NPU.

[0141] In one alternative implementation, the target server 402 is a bare metal server.

[0142] In one alternative implementation, the hardware information further includes at least one of the following:

[0143] Central Processing Unit (CPU) information;

[0144] Memory information;

[0145] Disk information.

[0146] In one optional implementation, the network interface card (NIC) information includes the Link Layer Discovery Protocol (LLDP) information of the NPU, the number of NPUs is multiple, and the NPUs communicate with each other via a parameter plane network. The target server 402 is further configured as follows:

[0147] Based on the LLDP information, determine the network topology parameters related to the parametric plane network;

[0148] The network topology parameters are sent to the control node 401.

[0149] In one optional implementation, the network topology parameters include at least one of the following:

[0150] The media access control MAC address of the network interface card of the NPU;

[0151] Parameter information of the switch associated with the network interface card of the NPU;

[0152] The interface parameter information of the network card of the NPU.

[0153] In an optional implementation, the control node 401 is further configured to perform a preset operation on the target server 402 based on the network topology parameters, the preset operation including at least one of the following: node information management and port information management.

[0154] Fourthly, correspondingly, the embodiments of this application also provide a server information collection device, which can implement all the processes of the server information collection method provided in any of the embodiments of the first aspect above.

[0155] See Figure 5The diagram shows a schematic of the server information collection device 500 provided in an embodiment of this application. The server information collection device 500 includes:

[0156] Image acquisition module 501 is used to acquire the target image;

[0157] Image loading module 502 is used to load the acquired target image into the target server, wherein the target server includes a neural network processor (NPU);

[0158] The acquisition module 503 is used to acquire the hardware information of the target server through the loaded target image, wherein the hardware information includes the network card information of the NPU.

[0159] In one optional implementation, obtaining the target image includes:

[0160] After the target server enters the specified startup mode, a temporary network address is obtained;

[0161] The target image is obtained based on the temporary network address.

[0162] In one optional implementation, the step of collecting the hardware information of the target server through the loaded target image includes:

[0163] The hardware information of the target server is collected by running a preset proxy service in the loaded target image.

[0164] In one optional implementation, the preset proxy service is used to invoke a network management tool during its operation to collect hardware information of the target server, the network management tool corresponding to the NPU.

[0165] In one alternative implementation, the target server is a bare metal server.

[0166] In one alternative implementation, the hardware information further includes at least one of the following:

[0167] Central Processing Unit (CPU) information;

[0168] Memory information;

[0169] Disk information.

[0170] In one optional implementation, the network interface card (NIC) information includes the Link Layer Discovery Protocol (LLDP) information of the NPU, the number of NPUs is multiple, and the NPUs communicate with each other via a parameter plane network. The apparatus further includes:

[0171] The parameter determination module is used to determine the network topology parameters related to the parameter plane network based on the LLDP information.

[0172] The parameter sending module is used to send the network topology parameters to the management node corresponding to the target server.

[0173] In one optional implementation, the network topology parameters include at least one of the following:

[0174] The media access control MAC address of the network interface card of the NPU;

[0175] Parameter information of the switch associated with the network interface card of the NPU;

[0176] The interface parameter information of the network card of the NPU.

[0177] In one optional implementation, the network topology parameters are used to instruct the control node to perform a preset operation on the target server, the preset operation including at least one of the following: node information management and port information management.

[0178] Fifthly, correspondingly, embodiments of this application also provide a server information updating apparatus capable of implementing all processes of the server information updating method provided in any of the embodiments of the second aspect above.

[0179] See Figure 6 The diagram shows a schematic of the structure of a server information updating device 600 provided in an embodiment of this application. This server information updating device 600 is suitable for management nodes and includes:

[0180] The request receiving module 601 is used to receive an inspection request for a target server, wherein the target server includes an NPU;

[0181] Task determination module 602 is used to determine the server self-test task in response to the inspection request;

[0182] The mode control module 603 is used to control the target server to enter a specified startup mode based on the server self-test task. The target server is configured to: acquire a target image after entering the specified startup mode; load the acquired target image; and collect the hardware information of the target server through the loaded target image, wherein the hardware information includes the network card information of the NPU.

[0183] The management module 604 is used to receive network topology parameters determined based on the hardware information from the target server, and to perform preset operations on the target server based on the network topology parameters. The preset operations include at least one of the following: node information management and port information management.

[0184] In one optional implementation, obtaining the target image includes:

[0185] After the target server enters the specified startup mode, a temporary network address is obtained;

[0186] The target image is obtained based on the temporary network address.

[0187] In one optional implementation, the step of collecting the hardware information of the target server through the loaded target image includes:

[0188] The hardware information of the target server is collected by running a preset proxy service in the loaded target image.

[0189] In one optional implementation, the preset proxy service is used to invoke a network management tool during its operation to collect hardware information of the target server, the network management tool corresponding to the NPU.

[0190] In one alternative implementation, the target server is a bare metal server.

[0191] In one alternative implementation, the hardware information further includes at least one of the following:

[0192] Central Processing Unit (CPU) information;

[0193] Memory information;

[0194] Disk information.

[0195] In one optional implementation, the network interface card (NIC) information includes the Link Layer Discovery Protocol (LLDP) information of the NPU, wherein there are multiple NPUs, and the NPUs communicate with each other via a parameter plane network, wherein the network topology parameters are related to the parameter plane network and determined by the LLDP information.

[0196] In one optional implementation, the network topology parameters include at least one of the following:

[0197] The media access control MAC address of the network interface card of the NPU;

[0198] Parameter information of the switch associated with the network interface card of the NPU;

[0199] The interface parameter information of the network card of the NPU.

[0200] Sixthly, embodiments of this application provide a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in any of the preceding claims.

[0201] In a seventh aspect, embodiments of this application provide a computer program product, including computer instructions that, when executed by a processor, implement the steps of the method described in any of the preceding claims.

[0202] Eighthly, embodiments of this application provide a computer device including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the steps of the method described in any of the preceding claims.

[0203] See Figure 7 The computer device in this embodiment includes a processor 701, a memory 702, and a computer program stored in the memory 702 and executable on the processor 701, such as a server information acquisition or server information update program. When the processor 701 executes the computer program, it implements the steps in the various server information acquisition method embodiments or server information update method embodiments described above.

[0204] For example, the computer program may be divided into one or more modules / units, which are stored in the memory 702 and executed by the processor 701 to complete this application. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the computer device.

[0205] The computer device may be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device may include, but is not limited to, a processor 701 and a memory 702. Those skilled in the art will understand that the schematic diagram is merely an example of a computer device and does not constitute a limitation on the computer device. It may include more or fewer components than illustrated, or combine certain components, or different components. For example, the computer device may also include input / output devices, network access devices, buses, etc.

[0206] The processor 701 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor, or processor 701 can be any conventional processor. The processor 701 is the control center of the computer device, connecting various parts of the entire computer device through various interfaces and lines.

[0207] The memory 702 can be used to store the computer programs and / or modules. The processor 701 implements various functions of the computer device by running or executing the computer programs and / or modules stored in the memory 702 and calling the data stored in the memory 702. The memory 702 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory 702 may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0208] Wherein, if the modules / units integrated into the computer device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a non-transitory computer-readable storage medium. When the computer program is executed by the processor 701, it can implement the steps of the various method embodiments described above. Wherein, the computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0209] In summary, the embodiments of this application have at least the following beneficial effects:

[0210] By using the embodiments of this application, the obtained target image is loaded into the target server containing the NPU, so that the hardware information of the target server can be automatically collected by loading the target image into the target server. This can improve the efficiency of collecting server hardware information. In particular, hardware information such as NPU network card information, which has a significant impact on the performance of servers and large-scale clusters, is also collected. This can improve the accuracy of collecting key hardware information of the server, thereby improving the construction efficiency and scalability of large-scale clusters and reducing the operation and maintenance complexity of the cluster.

[0211] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary hardware platforms, or it can be implemented entirely by hardware. Based on this understanding, all or part of the technical solutions of this application that contribute to the background technology can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM (Read-Only Memory) / RAM (Random Access Memory), magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0212] The above description is the preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications are also considered to be within the scope of protection of this application.

Claims

1. A method for collecting server information, characterized in that, include: Obtain the target image; The process of obtaining the target image includes: obtaining a temporary network address after the target server enters a specified startup mode; obtaining the target image based on the temporary network address, wherein the target image has predefined startup parameters. The acquired target image is loaded into the target server, wherein the target server includes a neural network processor (NPU); The target server's hardware information is collected using the loaded target image, including the network interface card information of the NPU.

2. The method according to claim 1, characterized in that, The step of collecting hardware information of the target server through the loaded target image includes: The hardware information of the target server is collected by running a preset proxy service in the loaded target image.

3. The method according to claim 2, characterized in that, The preset proxy service is used to invoke a network management tool during its operation to collect hardware information of the target server, and the network management tool corresponds to the NPU.

4. The method according to claim 1, characterized in that, The target server is a bare metal server.

5. The method according to claim 1, characterized in that, The hardware information also includes at least one of the following: Central Processing Unit (CPU) information; Memory information; Disk information.

6. The method according to any one of claims 1-5, characterized in that, The network interface card (NIC) information includes the Link Layer Discovery Protocol (LLDP) information of the NPU. There are multiple NPUs, and they communicate with each other via a parameter plane network. The method further includes: Based on the LLDP information, determine the network topology parameters related to the parametric plane network; The network topology parameters are sent to the management node corresponding to the target server.

7. The method according to claim 6, characterized in that, The network topology parameters include at least one of the following: The media access control MAC address of the network interface card of the NPU; Parameter information of the switch associated with the network interface card of the NPU; The interface parameter information of the network card of the NPU.

8. The method according to claim 6, characterized in that, The network topology parameters are used to instruct the control node to perform preset operations on the target server, and the preset operations include at least one of the following: node information management and port information management.

9. A method for updating server information, characterized in that, Applicable to control nodes, the method includes: Receive an inspection request for a target server, wherein the target server includes an NPU; In response to the inspection request, a server self-check task is determined; Based on the server self-test task, the target server is controlled to enter a specified startup mode. The target server is configured to: acquire a target image after entering the specified startup mode; load the acquired target image; and collect hardware information of the target server through the loaded target image, including the network interface card information of the NPU. Acquiring the target image includes: acquiring a temporary network address after the target server enters the specified startup mode; and acquiring the target image based on the temporary network address, wherein the target image has predefined startup parameters. The system receives network topology parameters determined based on the hardware information from the target server and performs a preset operation on the target server based on the network topology parameters. The preset operation includes at least one of the following: node information management and port information management.

10. A server information acquisition system, characterized in that, include: The control node is configured to receive inspection requests for the target server; In response to the inspection request, a server self-check task is determined; Based on the server self-test task, control the target server to enter the specified startup mode; as well as The target server includes an NPU and is configured to acquire a target image after entering a specified boot mode. Load the acquired target image; The hardware information of the target server is collected through the loaded target image, wherein the hardware information includes the network card information of the NPU; The process of obtaining the target image includes: obtaining a temporary network address after the target server enters a specified startup mode; obtaining the target image based on the temporary network address, wherein the target image corresponds to predefined startup parameters.

11. A server information acquisition device, characterized in that, include: The image acquisition module is used to acquire the target image; The process of obtaining the target image includes: obtaining a temporary network address after the target server enters a specified startup mode; obtaining the target image based on the temporary network address, wherein the target image has predefined startup parameters. An image loading module is used to load the acquired target image into a target server, wherein the target server includes a neural network processor (NPU). The acquisition module is used to acquire the hardware information of the target server through the loaded target image, wherein the hardware information includes the network card information of the NPU.

12. A server information updating device, characterized in that, Applicable to control nodes, the device includes: A request receiving module is used to receive an inspection request for a target server, wherein the target server includes an NPU; The task determination module is used to determine the server self-check task in response to the check request; A mode control module is used to control the target server to enter a specified startup mode based on the server self-test task. The target server is configured to: acquire a target image after entering the specified startup mode; load the acquired target image; and collect hardware information of the target server through the loaded target image, including the network interface card information of the NPU. Acquiring the target image includes: acquiring a temporary network address after the target server enters the specified startup mode; and acquiring the target image based on the temporary network address. The target image has predefined startup parameters. The management module is used to receive network topology parameters determined based on the hardware information from the target server, and to perform preset operations on the target server based on the network topology parameters. The preset operations include at least one of the following: node information management and port information management.

13. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method described in any one of claims 1-9.

14. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the method described in any one of claims 1-9.

15. A computer device, characterized in that, The method includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements the method of any one of claims 1-9.