Method and device for checking and accepting server and related equipment

By adding large-model training related performance detection in the server acceptance process, the problem that the existing technology cannot determine whether the server meets the large-model training needs is solved, and a comprehensive evaluation and acceptance of the server's computing power and data transmission performance is achieved.

CN119988170APending Publication Date: 2025-05-13SHANGHAI INFINIGENCE AI INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510074918.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

It is difficult for the prior art to determine whether the server can meet the computing power and data transmission requirements of large model training, although the traditional acceptance method has passed the basic configuration acceptance of the server.

Method used

Provide a server acceptance method and device. By obtaining the configuration information of the target server, verifying whether it complies with the preset configuration information, and when complying, the performance related to large model training of the target server is detected, including disk read and write rates, network transmission rates, interconnection rates of the large model accelerator and hardware reliability, and finally determine whether the server passes acceptance based on the detection results.

Benefits of technology

Ensure the effectiveness of the server in the big model training scenario, and through detailed performance detection and acceptance processes, it can accurately determine whether the server meets the needs of big model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988170A_ABST
    Figure CN119988170A_ABST
Patent Text Reader

Abstract

The invention provides a method and device for checking and accepting a server and related equipment. The method for checking and accepting the server comprises the following steps: acquiring configuration information of a target server to be checked and accepted; verifying whether the configuration information accords with preset configuration information; under the condition that the configuration information conforms to the preset configuration information, target performance of the target server is detected, a performance detection result is obtained, and the target performance is one or more performance, related to large model training, of the target server; and evaluating whether the performance detection result meets a preset condition or not so as to determine whether the target server passes acceptance check or not.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a method, apparatus and related equipment for acceptance of a server. Background Art

[0002] With the development of artificial intelligence technology, large models built by deep neural networks have been widely used in various fields, such as natural language processing, computer vision, speech recognition, and recommendation systems. In the large model training scenario, a single server is required to provide the highest possible computing power and ensure that the huge amount of data generated during the training process can be quickly transferred between servers.

[0003] Since large-model training scenarios require high computing power from the server, and related technologies usually only accept the basic configuration of the server, even if the server passes the acceptance in the traditional way, it is still impossible to determine whether the server can meet the large-model training requirements. Summary of the invention

[0004] The purpose of the present disclosure is to provide a server acceptance method, apparatus and related equipment, which are used to further accept the performance of the server related to large model training based on the basic configuration of the acceptance server.

[0005] In a first aspect, an embodiment of the present disclosure provides a method for accepting a server, the method comprising:

[0006] Obtain the configuration information of the target server to be accepted;

[0007] Verifying whether the configuration information complies with preset configuration information;

[0008] In the case where the configuration information conforms to the preset configuration information, detecting the target performance of the target server to obtain a performance detection result, wherein the target performance is one or more performances of the target server related to large model training;

[0009] Evaluate whether the performance test result meets the preset conditions to determine whether the target server passes the acceptance.

[0010] In a second aspect, an embodiment of the present disclosure provides a device for an acceptance server, the device comprising:

[0011] A configuration information acquisition unit, used to acquire configuration information of a target server to be accepted;

[0012] A configuration information verification unit, used to verify whether the configuration information conforms to preset configuration information;

[0013] a target performance detection unit, configured to detect the target performance of the target server when the configuration information conforms to the preset configuration information, and obtain a performance detection result, wherein the target performance is one or more performances of the target server related to large model training;

[0014] The test result evaluation unit is used to evaluate whether the performance test result meets the preset conditions to determine whether the target server has passed the acceptance.

[0015] In a third aspect, the present disclosure provides an electronic device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the method described in the embodiments of the present disclosure.

[0016] In a fourth aspect, the present disclosure provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the method described in the embodiments of the present disclosure is implemented.

[0017] In a fifth aspect, the present disclosure provides a computer program product, comprising computer instructions, which, when executed by a processor, implement the method described in the embodiments of the present disclosure.

[0018] In the present disclosure, after verifying the basic configuration of the server, the performance of the server related to large model training can be further verified, so that it can be determined whether the server meets the large model training requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a flow chart of a method for accepting a server provided by an embodiment of the present disclosure;

[0020] Figure 2 is a schematic diagram of the structure of a device for acceptance server provided in an embodiment of the present disclosure;

[0021] Figure 3 It is a structural diagram of an electronic device capable of implementing the embodiments of the present disclosure. DETAILED DESCRIPTION

[0022] The following will be combined with the drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.

[0023] The present disclosure provides a method for accepting a server. Figure 1As shown, the method includes:

[0024] Step 101: Obtain configuration information of the target server to be accepted.

[0025] For example, the target server to be accepted may be a server produced / assembled by the user, a commercially available server purchased by the user through an external channel, or a server rented by the user. The present disclosure does not limit the specific method of obtaining the target server. For example, the configuration information of the target server to be accepted may be obtained by making the target server execute or making other devices (such as the device for accepting the server according to the embodiment of the present disclosure) execute a predetermined computer program on the target server.

[0026] The configuration information of the target server is used to represent the basic configuration information of the target server that needs to be accepted, such as related information of hardware and software. In some embodiments, the configuration information of the target server to be obtained includes one or more of the following: central processing unit (CPU) configuration information, memory configuration information, physical disk configuration information, logical disk configuration information, network adapter configuration information, large model accelerator configuration information, software configuration information related to large model training, and network configuration information. CPU configuration information may include, but is not limited to, the number, manufacturer, model, number of cores, etc. of CPUs, memory configuration information may include, but is not limited to, the type, number, capacity, interface type, etc. of memory modules, physical disk configuration information may include, but is not limited to, the type, number, capacity, etc. of physical disks, logical disk configuration information may include, but is not limited to, the number, size, etc. of logical disks, network adapter configuration information may include, but is not limited to, the type and number of network adapters, large model accelerator configuration information may include, but is not limited to, the type (e.g., graphics processing unit (GPU), neural network processing unit (NPU), field programmable gate array (FPGA), etc.), number, manufacturer, model, etc. of large model accelerators, software configuration information related to large model training may include, but is not limited to, the name and version of software related to large model training, network configuration information may include, but is not limited to, IP address, subnet mask, default gateway, DNS, Ethernet card hardware address, etc. Which configuration information to obtain may be appropriately set according to acceptance requirements or actual hardware / software installation conditions.

[0027] Step 102: Verify whether the configuration information of the target server conforms to the preset configuration information.

[0028] The preset configuration information is the basic configuration information that the target server is expected to have, for example, it can be the basic configuration information obtained when the target server is produced / assembled, purchased or rented, and specifically, for example, it can include one or more of the above-mentioned configuration information. For example, the preset configuration information can be set according to the hardware / software installed on the target server during production / assembly, or the preset configuration information can be set according to the hardware / software configuration list provided by the server provider during purchase / rental. Verifying whether the configuration information of the target server conforms to the preset configuration information can be understood as verifying whether the basic configuration of the target server reaches the basic configuration that the target server is expected to have. Verifying whether the configuration information of the target server conforms to the preset configuration information can be performed, for example, by comparing whether the configuration information of the target server is the same as the preset configuration information. For example, the target server can be made to execute or other devices (such as the device for acceptance of the server according to the embodiment of the present disclosure) can be made to execute a predetermined computer program to compare the configuration information of the target server with the preset configuration information to verify whether the configuration information of the target server conforms to the preset configuration information. If the configuration information of the target server obtained is exactly the same as the preset configuration information (the amount and content of the information are the same), it is determined that the configuration information of the target server conforms to the preset configuration information, and the basic configuration of the target server reaches the basic configuration that the target server is expected to have. On the contrary, if the configuration information of the target server obtained is at least partially different from the preset configuration information, it is determined that the configuration information of the target server does not conform to the preset configuration information, and the basic configuration of the target server does not reach the basic configuration that the target server is expected to have, that is, the target server acceptance fails, and the next step is no longer performed. It should be understood that it is also possible to verify whether the configuration information of the target server conforms to the preset configuration information by comparing whether the basic configuration indicated by the configuration information of the target server is higher than or equal to the preset configuration indicated by the preset configuration information. For example, when the memory capacity indicated by the configuration information of the target server is greater than the preset memory capacity indicated by the preset configuration information, it can also be considered that the configuration information of the target server conforms to the preset configuration information.

[0029] For example, assume that the preset configuration information is CPU configuration information a1, memory configuration information a2, and physical disk configuration information a3. If the configuration information of the target server obtained is also CPU configuration information a1, memory configuration information a2, and physical disk configuration information a3, then the configuration information of the target server is considered to be in compliance with the preset configuration information. If the configuration information of the target server obtained is CPU configuration information a1 and memory configuration information a2, then the configuration information of the target server is considered to be inconsistent with the preset configuration information because physical disk configuration information a3 is missing compared to the preset configuration information. If the configuration information of the target server obtained is CPU configuration information a1, memory configuration information a4, and physical disk configuration information a3, then the configuration information of the target server is considered to be inconsistent with the preset configuration information because the memory configuration information is different compared to the preset configuration information. Alternatively, if the memory capacity indicated by memory configuration information a4 is greater than the memory capacity indicated by memory configuration information a2, then the configuration information of the target server can also be considered to be consistent with the preset configuration information.

[0030] Step 103: When the configuration information of the target server meets the preset configuration information, the target performance of the target server is tested to obtain a performance test result, wherein the target performance is one or more performances of the target server related to the large model training.

[0031] The target performance may include, for example, one or more of the following: disk read / write rate, network transmission rate, interconnection rate of large model accelerator, hardware reliability of large model accelerator. The target performance to be tested may be appropriately set according to acceptance requirements or actual hardware / software installation conditions.

[0032] For the disk read / write rate, for example, the target server may be made to execute or other devices (such as the device for acceptance of the server according to the embodiment of the present disclosure) may be made to execute a predetermined computer program (such as the disk I / O test tool fio) to detect the disk read / write rate of the target server. For example, the target server may call a predetermined function to create a test file and cause the disk to perform I / O read / write operations on the test file, thereby detecting the disk read / write rate of the target server.

[0033] The network transmission rate may further include, for example: a first network transmission rate of the target server in the in-band management network, and a second network transmission rate of the target server in the high-performance network used for large model training. An in-band management network refers to a network in which the server uses the same link to transmit management control data (such as control instructions) and business application data (such as data related to large model training). The in-band management network can be used as a data interaction network outside the large model training process, for example, it can be used for data acquisition before large model training, training status feedback during training, and training result feedback after training. A high-performance network refers to a network that achieves the highest throughput (infinitely close to the network bandwidth) under low latency and low jitter. A high-performance network can be used as a server interconnection network in the large model training process, for example, it can transmit large batches of input data and multi-dimensional feature data between servers. By setting the network transmission rate of the target server in the in-band management network and the network transmission rate in the high-performance network used for large model training as the target performance, and detecting these two network transmission rates, the overall network transmission performance of the target server can be comprehensively determined from two aspects: outside the large model training process and during the large model training process, so that the detected network transmission performance related to the large model training is more comprehensive and reliable. For example, the target server can be made to execute or other devices (such as devices for acceptance servers according to embodiments of the present disclosure) can be made to execute a predetermined computer program on the target server to detect the above-mentioned first network transmission rate and second network transmission rate. For example, the ib_write_bw program can be used to test the network card interconnection and network transmission rate of a high-speed network card that supports Remote Direct Memory Access (RDMA) on the server.

[0034] The interconnection rate of the large model accelerator refers to the communication rate when multiple large model accelerators (such as GPU, NPU, FPGA, etc.) in the target server are interconnected. For example, the target server can be made to execute or other devices (such as the device for acceptance server according to the embodiment of the present disclosure) can be made to execute a predetermined computer program on the target server to detect the communication rate when multiple large model accelerators in the target server are interconnected. For example, the nccl-test program can be used to test the communication rate between accelerators on a target server with an NVIDIA general accelerator installed. Alternatively, when the target server uses an AMD general accelerator, the rccl-test program can be used to test the communication rate between accelerators; when the target server uses a Huawei general accelerator, the hccl-test program can be used to test the communication rate between accelerators.

[0035] The hardware reliability of a large model accelerator refers to the hardware reliability of a large model accelerator when performing large model training. The indicators for evaluating the hardware reliability of a large model accelerator may be, for example, state parameters of a large model accelerator in a target server when performing a predetermined computing task (e.g., a large model computing task for testing) at a predetermined utilization rate (e.g., above 90%). The state parameters may be, for example, the temperature or power consumption of a large model accelerator. For a predetermined computing task, for example, different large model computing tasks for testing may be set for different large model accelerators from different manufacturers, so that the computing tasks can fully utilize the computing resources of the corresponding large model accelerators to ensure the accuracy of the indicators of the hardware reliability detected. For example, the target server may be caused to execute or other devices (e.g., a device for accepting a server according to an embodiment of the present disclosure) may be caused to execute a predetermined computer program on the target server to detect the indicators of the hardware reliability of the large model accelerator. For example, the target server may be caused to run a gpu-burn program to perform a continuous high-intensity usage test on an NVIDIA general-purpose accelerator.

[0036] In addition, when multiple servers to be accepted form an interconnected cluster, each of the multiple servers can be understood as an interconnected node in the interconnected cluster. By sending a test instruction to the multiple interconnected nodes included in the interconnected cluster through the server as a control node, and recovering the test results fed back by the multiple interconnected nodes in response to the test instruction, the interconnection performance between the multiple servers can be detected. The interconnection performance can be, for example, the overall bandwidth of data communication between the multiple servers. The interconnection performance can also be detected as the target performance of the target server.

[0037] Step 104: Evaluate whether the performance test result of the target performance of the target server meets the preset conditions to determine whether the target server passes the acceptance.

[0038] According to the specific items included in the target performance, the preset conditions may include one or more of the following: the disk read / write rate of the target server is greater than or equal to the preset disk read / write rate threshold; the first network transmission rate of the target server is greater than or equal to the preset first network transmission rate threshold; the second network transmission rate of the target server is greater than or equal to the preset second network transmission rate threshold; the interconnection rate of the large model accelerator of the target server is greater than or equal to the preset accelerator interconnection rate threshold; the hardware reliability index of the large model accelerator of the target server is less than or equal to the preset hardware reliability threshold. In the case where the preset conditions include one or more of the above, the target server can be considered to have passed the acceptance only when the performance test results of the target performance all meet these one or more items.

[0039] As an example but not limitation, the disk read / write rate threshold may be, for example, 300 Mb / s, the first network transmission rate threshold may be, for example, 10 Gb / s, the second network transmission rate threshold may be, for example, 200 Gb / s, the accelerator interconnect rate threshold may be, for example, 90 Gb / s, and when the indicator of the hardware reliability of the large model accelerator is temperature, the hardware reliability threshold may be, for example, 85°C.

[0040] For ease of understanding, a specific example of using the aforementioned method for accepting a server to perform server acceptance is given below.

[0041] Assume that the target server to be accepted is equipped with an NVIDIA general-purpose accelerator and a high-speed network card that supports RDMA. First, make sure that the target server to be accepted has installed the operating system and device drivers required for acceptance and can be accessed locally or over the network by the device used to accept the server. The device used to accept the server can first obtain the configuration information of the CPU / memory module / disk / network adapter / large model accelerator of the target server, and then compare the obtained configuration information with the preset configuration information to verify whether these configuration information and the preset configuration information are the same. If different items are found, the acceptance is deemed to have failed and the acceptance process is terminated. When these configuration information and the preset configuration information are exactly the same, perform the following detection. For example, use the fio tool to detect the read and write rates of each physical disk on the target server to obtain the disk read and write rates; use the ib_write_bw program to test the network card interconnection and transmission rate of the high-speed network card that supports RDMA on the target server to obtain the first network transmission rate and the second network transmission rate; use the nccl-test tool to test the interconnection rate between accelerators on the target server to obtain the interconnection rate of the large model accelerator; use the gpu-burn tool to perform continuous high-intensity usage tests on the target server to obtain hardware reliability indicators (such as accelerator temperature).

[0042] Next, the obtained performance test results, i.e., the above-mentioned disk read / write rate, the first network transmission rate and the second network transmission rate, the interconnection rate of the large model accelerator and the index of the hardware reliability of the large model accelerator, are compared with the preset disk read / write rate threshold, the first network transmission rate threshold and the second network transmission rate threshold, the accelerator interconnection rate threshold and the hardware reliability threshold. If at least one of the obtained disk read / write rate, the first network transmission rate and the second network transmission rate, and the interconnection rate of the large model accelerator is less than the corresponding threshold, or the index of the hardware reliability is greater than the hardware reliability threshold, it is deemed that the acceptance has failed and the acceptance process is terminated; otherwise, it is deemed that the acceptance has succeeded.

[0043] Note that the aforementioned steps 101 and 102 may also be interchanged with the order of steps 103 and 104, that is, the target performance of the target server may be detected and evaluated first, and then the configuration information of the target server may be acquired and verified.

[0044] The present disclosure also provides a device for an acceptance server. Figure 2 As shown, the device 200 for accepting the server includes:

[0045] The configuration information acquisition unit 201 is used to acquire the configuration information of the target server to be accepted;

[0046] The configuration information verification unit 202 is used to verify whether the configuration information conforms to the preset configuration information;

[0047] A target performance detection unit 203 is used to detect the target performance of the target server when the configuration information conforms to the preset configuration information, and obtain a performance detection result, wherein the target performance is one or more performances of the target server related to large model training;

[0048] The test result evaluation unit 204 is used to evaluate whether the performance test result meets a preset condition to determine whether the target server has passed the acceptance.

[0049] The above-mentioned various units of the device 200 for accepting a server provided in the embodiment of the present disclosure can implement the corresponding operations of the aforementioned method for accepting a server, and will not be described again here to avoid repetition.

[0050] Note that the predetermined computer program mentioned in the present disclosure may be, for example, a script or program designed and generated by oneself, or a program compiled based on public source code, or a compiled executable program.

[0051] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.

[0052] Figure 3 A schematic block diagram of an example electronic device 300 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0053] like Figure 3 As shown, the device 300 includes a computing unit 301, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 302 or a computer program loaded from a storage unit 308 to a random access memory (RAM) 303. In RAM 303, various programs and data required for the operation of the device 300 can also be stored. The computing unit 301, ROM 302 and RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0054] A number of components in the device 300 are connected to the I / O interface 305, including: an input unit 306, such as a keyboard, a mouse, etc.; an output unit 307, such as various types of displays, speakers, etc.; a storage unit 308, such as a disk, an optical disk, etc.; and a communication unit 309, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 309 allows the device 300 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0055] The computing unit 301 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSP), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 301 performs the various methods and processes described above, such as the server acceptance method. For example, in some embodiments, the server acceptance method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 308. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 300 via the ROM 302 and / or the communication unit 309. When the computer program is loaded into the RAM 303 and executed by the computing unit 301, one or more steps of the server acceptance method described above may be performed. Alternatively, in other embodiments, the computing unit 301 may be configured to execute the server acceptance method in any other appropriate manner (eg, by means of firmware).

[0056] Various embodiments of the systems and techniques described above herein may be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs, which may be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, which may receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0057] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0058] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0059] As used herein, the term "machine-readable medium" refers to any computer program product, apparatus, and / or device (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.

[0060] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0061] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0062] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0063] The present disclosure also provides a computer program product, including computer instructions, which, when executed by a processor, implement the above Figure 1 The various processes of the method embodiment shown can achieve the same technical effect, and will not be described again here to avoid repetition.

[0064] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0065] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A method for accepting a server, characterized in that: The method comprises: Obtain the configuration information of the target server to be accepted; Verifying whether the configuration information complies with preset configuration information; In the case where the configuration information conforms to the preset configuration information, detecting the target performance of the target server to obtain a performance detection result, wherein the target performance is one or more performances of the target server related to large model training; Evaluate whether the performance test result meets the preset conditions to determine whether the target server passes the acceptance.

2. The method according to claim 1, characterized in that: The target performance includes one or more of the following: disk read and write rate, network transmission rate, interconnection rate of large model accelerator, and hardware reliability of large model accelerator.

3. The method according to claim 2, characterized in that The network transmission rates include: a first network transmission rate of the target server in the in-band management network, and The second network transmission rate of the target server in a high-performance network for large model training.

4. The method according to claim 3, characterized in that The interconnection rate of the large model accelerator is the communication rate when multiple large model accelerators in the target server are interconnected, and the indicator of the hardware reliability of the large model accelerator is the state parameter of the large model accelerator in the target server executing a predetermined computing task at a predetermined usage rate.

5. The method according to claim 4, characterized in that According to the target performance, the preset conditions include one or more of the following: The disk read / write rate of the target server is greater than or equal to a preset disk read / write rate threshold; The first network transmission rate of the target server is greater than or equal to a preset first network transmission rate threshold; The second network transmission rate of the target server is greater than or equal to a preset second network transmission rate threshold; The interconnection rate of the large model accelerator of the target server is greater than or equal to a preset accelerator interconnection rate threshold; The hardware reliability index of the large model accelerator of the target server is less than or equal to a preset hardware reliability threshold.

6. The method according to any one of claims 1 to 5, characterized in that The configuration information includes one or more of the following: central processing unit CPU configuration information, memory configuration information, physical disk configuration information, logical disk configuration information, network adapter configuration information, large model accelerator configuration information, software configuration information related to large model training, and network configuration information.

7. A device for acceptance server, characterized in that: The device comprises: A configuration information acquisition unit, used to acquire configuration information of a target server to be accepted; A configuration information verification unit, used to verify whether the configuration information conforms to preset configuration information; a target performance detection unit, configured to detect the target performance of the target server when the configuration information conforms to the preset configuration information, and obtain a performance detection result, wherein the target performance is one or more performances of the target server related to large model training; The test result evaluation unit is used to evaluate whether the performance test result meets the preset conditions to determine whether the target server has passed the acceptance.

8. An electronic device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the method according to any one of claims 1 to 6 when executed by the processor.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

10. A computer program product, characterized in that The method comprises computer instructions, which implement the method according to any one of claims 1 to 6 when executed by a processor.